Masahiro Kato, Taka Kato
arXiv 14 Aug 2026 · Artificial Intelligence
arXiv:2608.14528 · PDF · Extracted main text
This study investigates the methodological and theoretical properties of session handover in applications that use large language models. A task may continue in a new session when the context reaches the model's input limit, when the application restarts, or when another agent is asked to finish the task. The application must then decide which information from the earlier session to pass on. We formulate handover as the transfer of a task-relative in-context learning (ICL) state and distinguish exact recovery of earlier material from preservation of the target distribution. Under an exogeneity condition, predictive equivalence characterizes the coarsest deterministic sufficient handover and gives a fixed-length bit requirement. The analysis isolates the effects of the memory constraint, the writer, and the continuation procedure, and quantifies the cost of writing before the realized downstream query is known. We propose a three-part record that stores decisions and constraints exactly, uses task-justified statistics for repeated evidence, and retains original observations whose effect is not preserved by those statistics. Gaussian linear regression gives an exact finite-dimensional handover and finite-bit perturbation bounds, while nonparametric regression gives upper and lower bounds that relate memory to squared prediction error. These results provide a theory and method for deciding what a handover must retain and how its memory requirement depends on the continuation task.
appendix boundary found by appendix_command · 51% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Alliot Nagle, Adway Girish, Marco Bondaschi, Michael Gastpar, Ashok… (2024) Fundamental limits of prompt compression: A rate-distortion framework for black-box language models | 0.737 | 3 | 3 | 67% |
| 2 | Ricardo Baptista, Andrew Stuart, and Son Tran (2026) Large language models: A mathematical formulation, 2026 | 0.737 | 3 | 2 | 100% |
| 3 | Kazusato Oko, Yujin Song, Taiji Suzuki, and Denny Wu (2024) Pretrained transformer efficiently learns low-dimensional target functions in-context | 0.585 | 3 | 3 | 33% |
| 4 | Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei (2023) Transformers as statisticians: provable in-context learning with in-context algorithm selection | 0.511 | 2 | 2 | 50% |
| 5 | David Blackwell (1953) Equivalent comparisons of experiments | 0.511 | 2 | 2 | 50% |
| 6 | Musa Cim, Burak Topcu, Chita Das, and Mahmut Taylan Kandemir (2026) Parallel context compaction for long-horizon llm agent serving, 2026 | 0.511 | 2 | 2 | 50% |
| 7 | Ashwin Gerard Colaco and Nada Lahjouji (2026) What to keep, what to forget: A rate–distortion view of memory compaction in llms and agents, 2026 | 0.511 | 2 | 2 | 50% |
| 8 | Juno Kim and Taiji Suzuki (2024) Transformers learn nonlinear features in context: Nonconvex mean-field dynamics on the attention landscape | 0.511 | 2 | 2 | 50% |
| 9 | Andrew Semenov and Svyatoslav Dorofeev (2026) Beyond compaction: Structured context eviction for long-horizon agents, 2026 | 0.511 | 2 | 2 | 50% |
| 10 | Johannes Von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacrame… (2023) Transformers learn in-context by gradient descent | 0.511 | 2 | 2 | 50% |
Showing the top 10 of 21 scored citations.