Blog / Engineering

Keeping Long Agent Turns Inside the Context Window

Our sub-agents could grow until the model refused them, stop early to save room, or miscount tokens in Russian. Three fixes that let long AI work finish.

Mythex Team · 2026-09-29 · 5 min read

A model can only read so much at once — its context window. Mythex's agent does long work: reading many files, running commands, handing bounded jobs to sub-agents. We found three ways that long work broke: sub-agents grew until the model refused every request, a sub-agent stopped early because it thought it was out of room, and our token count was badly wrong for Russian, Chinese and Japanese text. Here's what we changed.

Some terms first

  • Tokens are the chunks a model reads. Limits and costs are counted in them, not in characters or words.
  • Context window is the most a model can read in one request. We also set our own lower budget for each request, so costs and speed stay reasonable.
  • Tool results are what the agent gets back from its tools: the contents of a file it read, the output of a command it ran. In long work, they are most of what gets sent.
  • A sub-agent is a helper the main agent starts for a self-contained task, such as "check these chapters against the notes". It gets one instruction, then works through many tool calls on its own.

Problem 1: sub-agents grew until the model refused them

Every step, the agent sends the model the conversation so far. When that got too big, our code trimmed it — but only at the start of a user message, so that a tool call and its result were never split apart. It kept the first message, added a note saying earlier conversation was trimmed, and kept the newest whole turns that fit.

A sub-agent has exactly one user message: its instruction. There was nowhere to cut. So it sent its entire history on every step, growing each time, until the request was bigger than the model accepts at all — about a million tokens. In a test with 8 such sub-agents, all 8 were refused.

The fix: clear old tool results first

Requests now fit the budget in three steps, trying each only if the one before wasn't enough:

  1. Clear the oldest tool results. Their content is replaced with a short note: the result was cleared to keep the conversation within the model's context, and the agent can run the tool again if it needs it. A file read from an hour ago is usually the biggest and least useful thing in the request, and the file is still there to read again. The newest results stay whole.
  2. Keep whole user turns. As before: the first message, a note, and the newest complete turns that fit.
  3. Keep whole tool rounds. Inside one long turn — the only kind a sub-agent has — keep the first message, a note, and the newest complete rounds of "call a tool, get its result".

This only shapes what is sent. The agent's saved history keeps everything. In the same test, 0 of 8 were refused and all 8 finished.

Clearing in batches, so caching keeps working

Model providers can reuse work on the start of a request if it's identical to the last one — a prefix cache. If we cleared one more result on every call, the start of the request would change every call and the cache would never help.

So the oldest results are cleared four at a time, not one. The cleared part then stays the same over several calls. A result not much bigger than the note is left alone: clearing it would save almost nothing and still change the start of the request. How much stays whole is decided by size, not count: the newest results that fit in a quarter of the budget stay, and always at least one. An earlier version kept a fixed six whole, which was right for short command output but, with long file reads, left most of the budget to old results.

Refusals now say so

A request the model rejects as invalid is no longer retried; sending the same request again gets the same answer. If the conversation is too long for the model, the person sees: "This conversation got too long for the model to read in one go. Start a new chat for the next step — everything built so far stays." That's the person's to act on, not an outage, so it doesn't page us.

Problem 2: the agent stopped to "save room"

A sub-agent asked to read 40 long chapters stopped after 2, saying it wanted to save room. It had used about 100,000 tokens of a 400,000-token budget.

It had no way to know that old results would now be cleared for it. So we told it. The agent's instructions now say that when a conversation grows past what the model can take, the oldest tool results are cleared and the newest stay whole — so don't stop a long task to save room; keep going, and read a file again if you need what was cleared.

A fix the model can't see isn't a fix from the model's point of view. The mechanism only helps if the agent trusts it.

Problem 3: counting tokens as if everything were English

We estimated tokens without running a tokenizer, as characters divided by four. That's a reasonable rule of thumb for English. It isn't for other scripts: Russian runs about 2.2 characters a token, and Chinese or Japanese about one.

So a sub-agent reading Russian chapters sent a request of about 640,000 tokens while our estimate said it was under the 400,000 budget. The estimate could just as easily have put a request past the model's hard limit.

The estimate now counts by script: plain ASCII at four characters a token, Chinese, Japanese and Korean at one, and everything else — Cyrillic, Greek, Arabic, accented letters, emoji — at two. It's meant to err high. On one Russian chapter it now estimates 43,423 tokens against about 44,600 measured; before, it said 25,023. Since a message never changes once written, its estimate is worked out once and remembered, rather than on every step.

What we took from it

  • Trim where the size actually is. In agent work, that's tool results, not user messages.
  • Tell the agent about its own safety nets. Otherwise it plans around limits that no longer apply.
  • "Characters ÷ 4" is an English assumption. Anything that counts tokens for multilingual users needs to count by script, or use a real tokenizer.
  • Mind the cache. A request that changes at the front on every call costs more and runs slower.

If you're building AI features into your own app, our guides on what an LLM is and how to use LLM APIs cover context windows and tokens from the start. More on how sub-agents work in Mythex is in the sub-agents docs.

Keep reading

  • A Failed Republish Never Deletes a Working Backend — A failed republish could delete an app's live API, or leave its broken new version running. Three fixes to how Mythex undoes a publish that goes wrong.
  • Alerts That Fire Once, Not Once per Server — A once-a-day alert reached us five times in one afternoon. Why in-memory dedupe breaks with two servers and a deploy, and how shared state fixed it.
  • Moving to Bigger Workspaces Without Making Anyone Wait — A quarter of our workspaces were stuck at 1 GB after we moved to 2 GB. Our first fix made one person wait seven minutes. What we changed, and changed again.
  • Billing Databases by What They Actually Use — We used to estimate each app's database cost, and the estimate was wrong both ways. Now every database is billed from measured usage, hour by hour.
  • Catching Apps That Crash Right After Publishing — Three publishes passed our health check while their API crashed seconds into every boot. The cause: counting crashes in a list capped at five entries.
  • Keeping Image-Heavy Agent Turns Within Memory — An agent turn that kept looking at catalogue scans ran out of memory at 952 MB. Three separate causes, each holding images too long, and how we fixed them.

Start building free · Templates · Docs