Industry Insights7 min read

Context Language Models Explained: Context as a File

A University of Washington and Meta Superintelligence Labs preprint lets an LLM edit its own context like a file. What it found, what it skipped, what to try.

IT
InterviewStack TeamResearch
|

Context Language Models Let the Model Rewrite Its Own Context

The paper Context Language Models (arXiv 2609.37725), by Rulin Shao, Pang Wei Koh, Luke Zettlemoyer, Mike Lewis and colleagues, lets an LLM edit its own context window like a file. It is a different work from ContextLM (arXiv 2510.20280) and from the 2015 Document-Context language models, which share the name.

It was submitted on 29 September 2026 by researchers at the University of Washington and Meta Superintelligence Labs, with one author each from MIT and Trillium Labs. On its benchmarks it matched or beat the baselines the authors tested, usually with less theoretical compute. We separate what the authors report from our own reading, and every number comes from the paper's text and tables.

What is a Context Language Model?

A Context Language Model (CLM) is an ordinary LLM that is allowed to edit its own context window like a file, instead of a harness deciding what to keep, summarize or drop. The paper proposes it as a way for long-running agents to manage their own working memory, and tests it zero-shot with existing models.

The Problem: Long-Running Agents Outgrow the Window

Watch what happens when an agent's history grows past the context limit and the oldest turns get cut.

Why long-running agents run out of context

Illustration
Context window0.4K of 32K tokens
  1. Turn 1User: Which award did Dr. Lee win in 2019?0.4K

31.6K of room left.

Step 1 of 6. Turn 1: the user asks a question. It is tiny (0.4K tokens), but it is the whole point of the task.

1 of 6
In short: an agent's context fills up fast with tool output. Dropping the oldest turns makes it fit again, but the oldest turn is usually the task itself.

What to notice: plain truncation can drop the user's original question. The authors built a diagnostic, ContextBench, with a 32K context limit and input volume up to 24 times that limit. None of the existing strategies they tested was perfect, even on simple synthetic tasks.

Three Ways to Handle a Full Window

Step through the same overflowing conversation under truncation, a fixed-rule summary, and model-managed context.

Three ways to cope when the context overflows

Illustration
Harness ruleat the limit, drop the oldest turns until it fits
live context
  • [1] user: Which award did Dr. Lee win in 2019?0.4K
  • [2] agent: search("Dr. Lee 2019 award")0.2K
  • [3] tool: Search results: 20 pages21K
  • [4] tool: Fetched document: lee-bio.html9K
  • [5] tool: More search results4K
Context window34.6K of 32K tokens (over the limit)
Pro: Cheap and simpleCon: Loses information, often the task itself

Step 1 of 3. Same conversation: 34.6K tokens, over the 32K limit. This harness truncates.

1 of 3
In short: truncation is cheap but loses information, often the task itself. Fixed-rule summaries are predictable but cannot know which detail matters later. A model-managed context (CLM) edits itself to keep what the task needs, at the cost of harder debugging and caching.

What to notice: a fixed rule ("summarize at the threshold") can drop the one document ID the agent needs ten turns later, and action-based tools can offload data but could not evict it from the live context on demand. Model-managed context lets the model decide what stays.

How Does a Context Language Model Work?

A Context Language Model (CLM) keeps its live context mirrored in an ordinary file and gets Bash access to edit it. Edits sync back into what the model sees on its next turn, and if it does not edit, new tokens are simply appended as usual.

How a Context Language Model manages its context

Illustration
/ctx/agent.md
[1]user: Which award did Dr. Lee win in 2019?
[3]tool: 20 search results21K
[4]tool: lee-bio.html9K
[5]tool: more search results4K

Step 1 of 5. The live context is mirrored as an ordinary file. Its path is given in the system prompt.

1 of 5
In short: the context is exposed as a file, the model reads it, writes a Bash or Python edit, the system syncs the edited file back as the live context, and the model continues. When it does not edit, new tokens are appended as in any LLM.

What to notice: there is no special "compact" tool, just ordinary shell or Python edits. The authors report behaviors nobody programmed in: a scoreboard updated through 163 in-place edits that kept context at 6-8K tokens, and a helper function the model wrote itself, then called 37 times. An agent swarm is several synced context files.

What Did the Context Language Models Paper Find?

Applied zero-shot, CLMs matched or beat the strongest baselines the authors compared against on most long-horizon agent benchmarks, usually with less compute. On BrowseComp-Plus with Qwen3.6-27B and a 32K context limit, CLM scored 59.4%, 11.4% (relative) above Codex-style summarization, with 21.5% fewer prefix-reuse FLOPs.

What the paper reports, benchmark by benchmark

From the paper

Qwen3.6-27B, 32K context limit, no training

Accuracy

+11.4%

relative gain; CLM scored 59.4%

Compute

21.5% fewer

prefix-reuse FLOPs

Accuracy, summarization = 100

Codex-style summarization
baseline
CLM
+11.4%

FLOPs, summarization = 100

Codex-style summarization
baseline
CLM
21.5% fewer

Caveat:The paper states this gain relative to summarization, so both charts are indexed to it. CLM also used 28.9% fewer FLOPs than MEM1. FLOPs are theoretical compute, not latency or cost.

Benchmark 1 of 3. BrowseComp-Plus, no training: CLM is 11.4% more accurate (relative) than Codex-style summarization and uses 21.5% fewer FLOPs.

1 of 3
In short: with no training, CLM beat Codex-style summarization on BrowseComp-Plus (+11.4% relative, 21.5% fewer FLOPs) and on a 10-task 12-hour EdgeBench (44.6 vs 42.3, 179 vs 437 PFLOPs). After RL on Qwen3.5-9B it matched a summary harness also trained with RL (42.5% vs 42.1%) with 38.8% fewer FLOPs, though only CLM got the efficiency reward. FLOPs are theoretical compute.
  • EdgeBench-10, 12 hours (Qwen3.6-27B, 32K): 5% higher than Codex-style summarization (44.6 vs 42.3), at 179 vs 437 PFLOPs per trial.
  • TerminalBench 2.1: matches accuracy at 70% of the FLOPs. TBLite: 73.7% vs 67.0%, at 91% of the FLOPs.
  • Software World (24 hours, six agents, GPT-5.6-Sol, 272K): 65% greater downstream speedup than a summary-based swarm, at the same spend.
  • BrowseComp-Plus after RL (Qwen3.5-9B): 28.8% to 42.5%, against 34.7% to 42.1% for a summary harness.
  • Steering (held-out KV Store): accuracy rose from 38.3% to 74.2%, up to 35.9 points, with a Qwen3.6-27B agent and Claude Fable 5.1 proposing the skill. ContextBench is the authors' own synthetic diagnostic.
  • Suffix Cache Reuse: 35% less server-side compute than standard SGLang at matched performance on BrowseComp-Plus; about two thirds of the reuse comes from stripping reasoning tokens.

What to notice: every gain is against a specific baseline, and "59% fewer" FLOPs is theoretical compute, not a bill or latency. The 47.6% relative RL gain is over the model's own weak start (28.8%). Against the RL-trained summary harness, 42.5% vs 42.1% is matching (the paper's word), and only CLM got the efficiency reward, so its 38.8% FLOPs gap (1.34 vs 2.19 PFLOPs) is not purely architecture.

What Does the Context Language Models Paper Not Show?

It is a preprint, not yet peer reviewed, and the two headline benchmarks (BrowseComp-Plus and EdgeBench) use one open-weight model (Qwen3.6-27B) at a 32K context limit. Compute is measured in theoretical FLOPs, not latency. At a 128K budget on EdgeBench-10, plain CLM (47.3) did not beat summarization (47.8), though it used less compute.

Does the edge hold with a bigger context budget?

From the paper

EdgeBench-10, 12-hour tasks, score out of 100

CLM leads summarization

Score

Summarization437 PFLOPs
42.3
CLM179 PFLOPs
44.6
CLM + subagents
44.2

PFLOPs are theoretical compute per trial. Compute for CLM + subagents at 32K is not charted here.

At 32K, CLM beats summarization 44.6 to 42.3 while using 59% less compute.

In short: on 12-hour EdgeBench-10, CLM beats summarization at a 32K budget (44.6 vs 42.3). At 128K, plain CLM scores 47.3 against 47.8 for summarization (using 142 vs 222 PFLOPs), and only CLM with subagents leads (50.2, at 219 vs 222 PFLOPs). The edge narrows when the budget is larger.

What to notice: the edge narrows as the budget grows. At 128K, CLM with subagents reached 50.2 against 47.3 for plain CLM, at about the same compute as summarization (219 vs 222 PFLOPs); at 32K, subagents added little (44.2 vs 44.6). More gaps:

  • Capability dependence. In the RL setup the 9B model started at 28.8% against 34.7% for the summary harness, yet in a separate untrained 32K comparison it scored 39.9% vs 37.7%, so the gap depends on the setup. It also edited 1.4 times per TerminalBench 2.1 task against 2.6 for Qwen3.6-27B.
  • Safety is flagged, not solved. Editable context can carry prompt injections across turns; defenses are future work.
  • Missing. We found no latency or wall-clock results, and nothing on how hard these runs are to debug and audit.

Should You Rebuild Your Agent's Context Handling Around CLMs Now?

Our read: not yet, but it is worth prototyping. The evidence is zero-shot benchmark results plus one reinforcement learning run, with no production reliability data. A cheap first step is to let your agent summarize or drop parts of its own context under a clear instruction, and to log every edit so you can compare it against your current fixed rule.

One builder detail is easy to miss: editing the middle of a context invalidates the usual prefix cache.

Why editing the middle of the context costs more to serve

Illustration

Each block is a chunk of tokens in the context, oldest on the left.

  • Reused from cache
  • Computed now
  • Recomputed
  • Reused stale state (SCR)
  • Removed by the edit

14 blocks cached

Step 1 of 6. A prompt cache keeps the computed states of tokens the model has already processed, so they are not computed again.

1 of 6
In short: appending reuses the whole cached prefix, but an edit in the middle invalidates the cache from the edit point onward. The paper's Suffix Cache Reuse reuses the stale cached states of surviving tokens; about two thirds of its extra reuse (5.3 of 7.8 points) comes from stripping reasoning tokens. SCR needs serving-stack changes (the authors patch SGLang). Block counts are illustrative.

What to notice: the authors add Suffix Cache Reuse (SCR), which reuses cached states for tokens that survive an edit. It needs serving-stack changes (they patch SGLang), and about two thirds of its extra reuse (5.3 of 7.8 points) comes from stripping reasoning tokens, not from edits. Our read: count cache cost, not just tokens, and benchmark on your own tasks and context limits. Fixed rules may stay the safer floor for small models.

How Could Context Language Models Come Up in an AI Engineer Interview?

Expect it as a design question: your agent's history outgrows the window, so what do you do? A strong answer compares truncation, fixed-rule summarization and model-managed context on information loss, predictability, cost, debuggability and cache behavior, then commits to a choice for the stated workload and says how it would be evaluated.

Three options for context that outgrows the window: truncate, summarize with fixed rules, or let the model manage it, each with a plus and a minus

Our summary of the trade-offs. The debuggability and caching points are our read, not measurements from the paper.

Likely question: "Your agent's conversation no longer fits the context window. How do you handle it?"

Model answer line: "I would keep recent turns verbatim, summarize older ones with a rule I can inspect, pin the facts that must never drift, and only let the model edit its own context where I can log and replay every edit."

For the full set of trade-offs, see our context engineering interview questions and answers. To drill the topic, try the LLM question bank, or practice with AI mock interviews at InterviewStack.io.

Where This Leaves Us

If a model can decide what to say, it can probably decide what to remember. The early numbers support that on long-horizon tasks, with real caveats about scale, cost accounting and safety. Read it as a prompt to test model-managed context on your own agents, not as a verdict.

Topics

context language modelspaper explainedLLM agentscontext engineeringAI researchAI engineer

Ready to practice?

Put what you've learned into practice with AI mock interviews and structured preparation guides.