Anyone who builds with reasoning models has hit the same wall. You give a model a genuinely hard task, turn reasoning effort all the way up, and watch it think. It maps dependencies, tests edge cases, backtracks, and refines its plan. Then, right as it starts writing the actual code or document, it hits the output token cap and stops dead. Sometimes it burns the entire output budget inside its thinking block and returns nothing at all.
For the last two years, the AI industry obsessed over the wrong number. Every model release was a race on input context: 128K, 200K, 1 million, 2 million tokens. We celebrated being able to drop an entire codebase, ten thousand pages of PDFs, or an hour of video into a prompt.
Meanwhile, the output limit barely moved. Even as input windows hit a million tokens, output stayed boxed into 32K or 64K. You could hand a model the Library of Congress, and it could only answer you on an index card.
When Google announced Gemini 4 Argon with a 1 million token output limit, up from 64K, the reaction online was telling. Half the comments assumed “1M output” was a typo for input context. The other half asked a reasonable question: who on earth wants to read a 1,400-page response from a chatbot?
Nobody does. And that is why the move is so disruptive.
Input is ROM, output is RAM
The mistake is thinking of output tokens as the text a human reads at the end. In a modern reasoning model, the output window is where the model thinks.
Input context is read-only memory. It holds the reference material you feed in before the model starts working. Output tokens are working memory and CPU cycles combined. Every token the model spends reasoning through a problem, testing a hypothesis, or discarding a bad approach comes out of the exact same output budget as the final answer.
Under a 64K ceiling, thinking and doing are locked in a zero-sum fight. The harder the problem, the more tokens the model needs to reason, and the fewer tokens it has left to produce the result. If an agent spends 52,000 tokens tracing a subtle bug across twenty files, it has 12,000 tokens left to write the fix.
Models adapt to that ceiling in ways every developer recognizes. They ration their thinking, leave // TODO: implement the rest comments in the middle of a file, or hand back a tidy summary of what you should go type yourself. We blamed the models for being lazy. Most of the time, they were just out of room.
Death by “context glue”
Because 64K was too small for real engineering work, we built an entire discipline of workarounds. Take any serious agent framework today and look at what the code actually does. A huge fraction of it is what people on Reddit rightly call “context glue”: splitting a large task into forty small prompts, passing summaries from step to step, and running continuation loops whenever a generation cuts off.
Every one of those handoffs is lossy compression.
In turn one, the model understands the subtle constraints of your system. By turn four, the summary has flattened those constraints into bullet points. By turn ten, the agent is writing code against a game-of-telephone version of its own earlier decisions. When long-horizon agents fail in practice, it is rarely because the model wasn’t smart enough to solve step seven. It fails because its state drifted across the seams between turns.
Giving a model 1 million output tokens removes the seams. When an agent has the headroom to think and generate across hundreds of thousands of tokens in a single trajectory, you can throw away most of the brittle state-machine scaffolding we spent the last two years writing.
What a million-token trajectory actually does
The most interesting detail in the Argon announcement wasn’t a benchmark score. It was a story about libgav1, an open-source AV1 video decoder.
An earlier effort had ported the decoder from C++ to Rust, but 32,000 lines of hand-tuned SIMD code were left behind because naïve Rust ports were too slow. Argon agents took those 32,000 lines, ran repeated rounds of profile-guided experiments, studied the compiler output, and rewrote the SIMD routines into safe Rust structured so the compiler would auto-vectorize them. The result was a memory-safe decoder running 2.7x faster than the prior Rust port, with identical video output.
Think about what happens inside a run like that. The model isn’t writing a 32,000-line essay in one shot. It is reading assembly and compiler diagnostics, trying an approach, seeing why the vectorizer rejected a loop, adjusting data layouts, and iterating until the numbers hold up. That kind of sustained, state-heavy loop is impossible when the model has to pack its bags and summarize itself every few thousand tokens.
The new bottleneck is verification
Of course, a 1M output window creates its own problems. At Argon’s introductory API pricing ($2 per million input tokens, $10 per million output tokens), a full million-token output trajectory costs $10. That is remarkably cheap if the trajectory replaces three days of unpleasant refactoring, and an expensive way to generate garbage if your prompt is vague.
More importantly, if a model can think for half a million tokens and hand you a 40,000-line diff, no human is going to read that diff line by line over coffee.
For the last two years, the bottleneck in applied AI was orchestration: how do we feed the model enough context and chop the work into small enough pieces that it doesn’t fall over? With 1M input and now 1M output, that era is closing. The bottleneck moves entirely to verification: compilers, type systems, property-based tests, and sandboxes that can prove a massive trajectory actually did what you asked.
Input context let AI read our systems. Output headroom lets it build them. Now we have to make sure our test suites are ready for the output.
Opinions are my own.
