Context Language Models Explained: Context as a File
A University of Washington and Meta Superintelligence Labs preprint lets an LLM edit its own context like a file. What it found, what it skipped, what to try.
IT
InterviewStack TeamResearch
|
Context Language Models Let the Model Rewrite Its Own Context
The paper Context Language Models (arXiv 2609.37725), by Rulin Shao, Pang Wei Koh, Luke Zettlemoyer, Mike Lewis and colleagues, lets an LLM edit its own context window like a file. It is a different work from ContextLM (arXiv 2510.20280) and from the 2015 Document-Context language models, which share the name.
It was submitted on 29 September 2026 by researchers at the University of Washington and Meta Superintelligence Labs, with one author each from MIT and Trillium Labs. On its benchmarks it matched or beat the baselines the authors tested, usually with less theoretical compute. We separate what the authors report from our own reading, and every number comes from the paper's text and tables.
What is a Context Language Model?
A Context Language Model (CLM) is an ordinary LLM that is allowed to edit its own context window like a file, instead of a harness deciding what to keep, summarize or drop. The paper proposes it as a way for long-running agents to manage their own working memory, and tests it zero-shot with existing models.
The Problem: Long-Running Agents Outgrow the Window
Watch what happens when an agent's history grows past the context limit and the oldest turns get cut.
Why long-running agents run out of context
Illustration
Context window0.4K of 32K tokens
32K limit
Turn 1User:Which award did Dr. Lee win in 2019?0.4K
Turn 2Agent:search("Dr. Lee 2019 award")0.2K
Turn 3Tool:Search results: 20 pages21K
Turn 4Tool:Fetched document: lee-bio.html9K
Turn 5Tool:More search results4K
31.6K of room left.
Lost: “Which award did Dr. Lee win in 2019?” The context now fits (13K), but the task is gone.
Turn 1: the user asks a question. It is tiny (0.4K tokens), but it is the whole point of the task.
Turn 2: the agent calls a search tool. Still tiny.
Turn 3: the search results come back at about 21K tokens. One tool result fills two thirds of the window.
Turn 4: the agent fetches a document (about 9K tokens). The window is now 96% full.
Turn 5: more results arrive. The context is 34.6K tokens, over the 32K limit. Something has to go.
The naive fix drops the oldest turns until it fits. Turn 1 goes first, and that was the question. The agent keeps working without knowing what it was asked.
Step 1 of 6. Turn 1: the user asks a question. It is tiny (0.4K tokens), but it is the whole point of the task.
Step 1 of 6
In short: an agent's context fills up fast with tool output. Dropping the oldest turns makes it fit again, but the oldest turn is usually the task itself.
What to notice: plain truncation can drop the user's original question. The authors built a diagnostic, ContextBench, with a 32K context limit and input volume up to 24 times that limit. None of the existing strategies they tested was perfect, even on simple synthetic tasks.
Three Ways to Handle a Full Window
Step through the same overflowing conversation under truncation, a fixed-rule summary, and model-managed context.
Three ways to cope when the context overflows
Illustration
Harness ruleat the limit, drop the oldest turns until it fitsat the limit, summarize all but the latest turncat context.mdsave note: award = Hale Prize (2019)summarize turns 3-4delete turn 5
live context
[1] user: Which award did Dr. Lee win in 2019?0.4K
[2] agent: search("Dr. Lee 2019 award")0.2K
[3] tool: Search results: 20 pages21K
[4] tool: Fetched document: lee-bio.html9K
[5] tool: More search results4K
-Removed: [1] user: Which award did Dr. Lee win in 2019?0.4K
-Removed: [2] agent: search("Dr. Lee 2019 award")0.2K
-Removed: [3] tool: Search results: 20 pages21K
[4] tool: Fetched document: lee-bio.html9K
[5] tool: More search results4K
No longer in context: [1] user: Which award did Dr. Lee win in 2019?out
No longer in context: [2] agent: search("Dr. Lee 2019 award")out
No longer in context: [3] tool: Search results: 20 pagesout
[4] tool: Fetched document: lee-bio.html9K
[5] tool: More search results4K
[1] user: Which award did Dr. Lee win in 2019?0.4K
[2] agent: search("Dr. Lee 2019 award")0.2K
[3] tool: Search results: 20 pages21K
[4] tool: Fetched document: lee-bio.html9K
[5] tool: More search results4K
-Removed: [1] user: Which award did Dr. Lee win in 2019?0.4K
-Removed: [2] agent: search("Dr. Lee 2019 award")0.2K
+Added: [S] summary: Turns 3-4: 20 results, bio confirms the award0.6K
[5] tool: More search results4K
[1] user: Which award did Dr. Lee win in 2019?0.4K
[N] note: award = Hale Prize (2019), from lee-bio.html0.1K
[2] agent: search("Dr. Lee 2019 award")0.2K
[S] summary: Turns 3-4: 20 results, bio confirms the award0.6K
-Removed: [5] tool: More search results4K
Context window34.6K of 32K tokens (over the limit) (over the limit)
32K limit
Pro: Cheap and simpleCon: Loses information, often the task itself
Pro: PredictableCon: Rules are fixed upfront, so the summary can drop what matters later
Pro: Adaptive: keeps what this task needsCon: Harder to debug and to cache
Same conversation: 34.6K tokens, over the 32K limit. This harness truncates.
The rule drops the oldest turns until the rest fits. Turns 1 to 3 go.
It fits at 13K tokens. Cheap, but the question and the search results are gone.
Same conversation, same 34.6K tokens. This harness summarizes instead.
A rule written in advance replaces turns 1 to 4 with one summary.
It fits at 4.8K tokens and the task survives, but the summary dropped the award name the bio contained.
Same conversation, but the context is a file the model can edit. The model decides what to keep.
First it pins the fact it found in the bio as a short note.
Then it replaces the two big tool results with a one-line summary. The note keeps the key detail safe.
Last, it deletes turn 5 (off-topic results). The question, the key fact and a summary remain: 1.3K tokens.
Step 1 of 3. Same conversation: 34.6K tokens, over the 32K limit. This harness truncates.
Step 1 of 3
In short: truncation is cheap but loses information, often the task itself. Fixed-rule summaries are predictable but cannot know which detail matters later. A model-managed context (CLM) edits itself to keep what the task needs, at the cost of harder debugging and caching.
What to notice: a fixed rule ("summarize at the threshold") can drop the one document ID the agent needs ten turns later, and action-based tools can offload data but could not evict it from the live context on demand. Model-managed context lets the model decide what stays.
How Does a Context Language Model Work?
A Context Language Model (CLM) keeps its live context mirrored in an ordinary file and gets Bash access to edit it. Edits sync back into what the model sees on its next turn, and if it does not edit, new tokens are simply appended as usual.
How a Context Language Model manages its context
Illustration
1Context as a file
2Model reads it
3Model writes an edit
4System syncs it back
5Model continues
Then back to step 1, every turn
/ctx/agent.md
[1]user: Which award did Dr. Lee win in 2019?
[3]tool: 20 search results21K
[4]tool: lee-bio.html9K
[5]tool: more search results4K
The model reads /ctx/agent.md:
Context window34.6K of 32K tokens (over the limit) (over the limit)
32K limit
edit.py (written by the model)
ctx = open(CONTEXT_FILE).read()
ctx = ctx.replace(turns_3_and_4,
"[S] bio confirms the award")
open(CONTEXT_FILE, "w").write(ctx)
After the sync, the live context is the edited file:
Context window5.2K of 32K tokens
32K limit
/ctx/agent.md
[1]user: Which award did Dr. Lee win in 2019?
[S]summary: bio confirms the award0.6K
[5]tool: more search results4K
+ [6]agent: answer("Hale Prize, 2019")
The live context is mirrored as an ordinary file. Its path is given in the system prompt.
The model reads the file, which is simply its own context, and sees it is over the limit.
It writes an ordinary Bash or Python edit. There is no special "compact" tool.
The system syncs the edited file back. It is now the live context the model sees on its next turn.
The model continues the task. If it does not edit the file, new tokens are appended as usual. Then the loop repeats.
Step 1 of 5. The live context is mirrored as an ordinary file. Its path is given in the system prompt.
Step 1 of 5
In short: the context is exposed as a file, the model reads it, writes a Bash or Python edit, the system syncs the edited file back as the live context, and the model continues. When it does not edit, new tokens are appended as in any LLM.
What to notice: there is no special "compact" tool, just ordinary shell or Python edits. The authors report behaviors nobody programmed in: a scoreboard updated through 163 in-place edits that kept context at 6-8K tokens, and a helper function the model wrote itself, then called 37 times. An agent swarm is several synced context files.
What Did the Context Language Models Paper Find?
Applied zero-shot, CLMs matched or beat the strongest baselines the authors compared against on most long-horizon agent benchmarks, usually with less compute. On BrowseComp-Plus with Qwen3.6-27B and a 32K context limit, CLM scored 59.4%, 11.4% (relative) above Codex-style summarization, with 21.5% fewer prefix-reuse FLOPs.
What the paper reports, benchmark by benchmark
From the paper
Qwen3.6-27B, 32K context limit, no training
Accuracy
+11.4%
relative gain; CLM scored 59.4%
Compute
21.5% fewer
prefix-reuse FLOPs
Accuracy, summarization = 100
Codex-style summarization
baseline
CLM
+11.4%
0120
FLOPs, summarization = 100
Codex-style summarization
baseline
CLM
21.5% fewer
0120
Caveat:The paper states this gain relative to summarization, so both charts are indexed to it. CLM also used 28.9% fewer FLOPs than MEM1. FLOPs are theoretical compute, not latency or cost.
Qwen3.6-27B, 32K context limit, 10-task subset
Score
+5%
44.6 vs 42.3 (0 to 100 scale)
Compute
59% fewer
179 vs 437 PFLOPs per trial
Score (0 to 100)
Codex-style summarization
42.3
CLM
44.6
0100
PFLOPs per trial
Codex-style summarization
437
CLM
179
0500
Caveat:Best of three seeds on a 10-task subset, with one open-weight model at 32K. At a 128K budget the edge narrows.
BrowseComp-Plus, 830 held-out questions, before and after training
Accuracy
+0.4 pts
42.5% vs 42.1%: a near tie
Compute
38.8% fewer
1.34 vs 2.19 PFLOPs per question
Accuracy %, faded = before training
Summary harness
34.7 → 42.1
CLM
28.8 → 42.5
0%60%
PFLOPs per question, faded = before
Summary harness
4.01 → 2.19
CLM
1.52 → 1.34
04.5
Caveat:Only CLM was trained with the efficiency reward; the summary harness got task reward alone, so the FLOPs gap is not purely about the architecture. Before training, CLM was behind (28.8% vs 34.7%).
BrowseComp-Plus, no training: CLM is 11.4% more accurate (relative) than Codex-style summarization and uses 21.5% fewer FLOPs.
12-hour EdgeBench, 10 tasks: CLM scores 44.6 vs 42.3 for summarization while using 59% fewer FLOPs (179 vs 437 PFLOPs per trial).
After RL on a 9B model, CLM and a summary harness also trained with RL end in a near tie (42.5% vs 42.1%), with CLM at 38.8% fewer FLOPs.
Benchmark 1 of 3. BrowseComp-Plus, no training: CLM is 11.4% more accurate (relative) than Codex-style summarization and uses 21.5% fewer FLOPs.
Benchmark 1 of 3
In short: with no training, CLM beat Codex-style summarization on BrowseComp-Plus (+11.4% relative, 21.5% fewer FLOPs) and on a 10-task 12-hour EdgeBench (44.6 vs 42.3, 179 vs 437 PFLOPs). After RL on Qwen3.5-9B it matched a summary harness also trained with RL (42.5% vs 42.1%) with 38.8% fewer FLOPs, though only CLM got the efficiency reward. FLOPs are theoretical compute.
EdgeBench-10, 12 hours (Qwen3.6-27B, 32K): 5% higher than Codex-style summarization (44.6 vs 42.3), at 179 vs 437 PFLOPs per trial.
TerminalBench 2.1: matches accuracy at 70% of the FLOPs. TBLite: 73.7% vs 67.0%, at 91% of the FLOPs.
Software World (24 hours, six agents, GPT-5.6-Sol, 272K): 65% greater downstream speedup than a summary-based swarm, at the same spend.
BrowseComp-Plus after RL (Qwen3.5-9B): 28.8% to 42.5%, against 34.7% to 42.1% for a summary harness.
Steering (held-out KV Store): accuracy rose from 38.3% to 74.2%, up to 35.9 points, with a Qwen3.6-27B agent and Claude Fable 5.1 proposing the skill. ContextBench is the authors' own synthetic diagnostic.
Suffix Cache Reuse: 35% less server-side compute than standard SGLang at matched performance on BrowseComp-Plus; about two thirds of the reuse comes from stripping reasoning tokens.
What to notice: every gain is against a specific baseline, and "59% fewer" FLOPs is theoretical compute, not a bill or latency. The 47.6% relative RL gain is over the model's own weak start (28.8%). Against the RL-trained summary harness, 42.5% vs 42.1% is matching (the paper's word), and only CLM got the efficiency reward, so its 38.8% FLOPs gap (1.34 vs 2.19 PFLOPs) is not purely architecture.
What Does the Context Language Models Paper Not Show?
It is a preprint, not yet peer reviewed, and the two headline benchmarks (BrowseComp-Plus and EdgeBench) use one open-weight model (Qwen3.6-27B) at a 32K context limit. Compute is measured in theoretical FLOPs, not latency. At a 128K budget on EdgeBench-10, plain CLM (47.3) did not beat summarization (47.8), though it used less compute.
Does the edge hold with a bigger context budget?
From the paper
EdgeBench-10, 12-hour tasks, score out of 100
CLM leads summarization
Score
Summarization437 PFLOPs
42.3
CLM179 PFLOPs
44.6
CLM + subagents
44.2
060
PFLOPs are theoretical compute per trial. Compute for CLM + subagents at 32K is not charted here.
At 32K, CLM beats summarization 44.6 to 42.3 while using 59% less compute.
At 128K, plain CLM (47.3) falls just behind summarization (47.8), though with less compute. Only CLM with subagents leads, at about the same compute.
At 32K, CLM beats summarization 44.6 to 42.3 while using 59% less compute.
In short: on 12-hour EdgeBench-10, CLM beats summarization at a 32K budget (44.6 vs 42.3). At 128K, plain CLM scores 47.3 against 47.8 for summarization (using 142 vs 222 PFLOPs), and only CLM with subagents leads (50.2, at 219 vs 222 PFLOPs). The edge narrows when the budget is larger.
What to notice: the edge narrows as the budget grows. At 128K, CLM with subagents reached 50.2 against 47.3 for plain CLM, at about the same compute as summarization (219 vs 222 PFLOPs); at 32K, subagents added little (44.2 vs 44.6). More gaps:
Capability dependence. In the RL setup the 9B model started at 28.8% against 34.7% for the summary harness, yet in a separate untrained 32K comparison it scored 39.9% vs 37.7%, so the gap depends on the setup. It also edited 1.4 times per TerminalBench 2.1 task against 2.6 for Qwen3.6-27B.
Safety is flagged, not solved. Editable context can carry prompt injections across turns; defenses are future work.
Missing. We found no latency or wall-clock results, and nothing on how hard these runs are to debug and audit.
Should You Rebuild Your Agent's Context Handling Around CLMs Now?
Our read: not yet, but it is worth prototyping. The evidence is zero-shot benchmark results plus one reinforcement learning run, with no production reliability data. A cheap first step is to let your agent summarize or drop parts of its own context under a clear instruction, and to log every edit so you can compare it against your current fixed rule.
One builder detail is easy to miss: editing the middle of a context invalidates the usual prefix cache.
Why editing the middle of the context costs more to serve
Illustration
Each block is a chunk of tokens in the context, oldest on the left.
edit ▾
Reused from cache
Computed now
Recomputed
Reused stale state (SCR)
Removed by the edit
14 blocks cached
Reused 14
SCR's extra reuse: 7.8% of prompt tokens (from the paper)
Reasoning stripped: 5.3
Edits: 2.5
A prompt cache keeps the computed states of tokens the model has already processed, so they are not computed again.
Normal append: new tokens go on the end. The whole cached prefix is reused, and only the 3 new blocks are computed.
Now the model edits the middle of its context: it summarizes 4 blocks into 1.
With a standard prefix cache, everything from the edit point onward is recomputed, even blocks that did not change.
Suffix Cache Reuse (SCR), the paper's serving fix, reuses the cached states of blocks that survived the edit, even though they were computed under the old prefix.
In the paper, about two thirds of SCR's extra reuse (5.3 of 7.8 points of prompt tokens) comes from stripping reasoning tokens, not from context edits.
Step 1 of 6. A prompt cache keeps the computed states of tokens the model has already processed, so they are not computed again.
Step 1 of 6
In short: appending reuses the whole cached prefix, but an edit in the middle invalidates the cache from the edit point onward. The paper's Suffix Cache Reuse reuses the stale cached states of surviving tokens; about two thirds of its extra reuse (5.3 of 7.8 points) comes from stripping reasoning tokens. SCR needs serving-stack changes (the authors patch SGLang). Block counts are illustrative.
What to notice: the authors add Suffix Cache Reuse (SCR), which reuses cached states for tokens that survive an edit. It needs serving-stack changes (they patch SGLang), and about two thirds of its extra reuse (5.3 of 7.8 points) comes from stripping reasoning tokens, not from edits. Our read: count cache cost, not just tokens, and benchmark on your own tasks and context limits. Fixed rules may stay the safer floor for small models.
How Could Context Language Models Come Up in an AI Engineer Interview?
Expect it as a design question: your agent's history outgrows the window, so what do you do? A strong answer compares truncation, fixed-rule summarization and model-managed context on information loss, predictability, cost, debuggability and cache behavior, then commits to a choice for the stated workload and says how it would be evaluated.
Our summary of the trade-offs. The debuggability and caching points are our read, not measurements from the paper.
Likely question: "Your agent's conversation no longer fits the context window. How do you handle it?"
Model answer line: "I would keep recent turns verbatim, summarize older ones with a rule I can inspect, pin the facts that must never drift, and only let the model edit its own context where I can log and replay every edit."
If a model can decide what to say, it can probably decide what to remember. The early numbers support that on long-horizon tasks, with real caveats about scale, cost accounting and safety. Read it as a prompt to test model-managed context on your own agents, not as a verdict.
Topics
context language modelspaper explainedLLM agentscontext engineeringAI researchAI engineer
Ready to practice?
Put what you've learned into practice with AI mock interviews and structured preparation guides.