The four runs side by side
How full the busiest workbench got, and how careful the result was. The dashed line is where work starts to suffer, around 70–80% of the window.
The workbench
An AI coding assistant can only hold so much at once: the request, the code it has read, its own reasoning, the tests. That space is its context window. Opus 5.5 has 1M tokens, but work stays sharp only up to roughly 70–80% of it.
The squeeze
When a job sprawls, the assistant can tell it is running out of room. It starts rationing: skimming files, assuming how code works, putting tests off. Each shortcut looks small. Together they are why big tasks come back needing an expert's review.
The team
An orchestrator plans at the top and gives each specialist one scoped brief. Every specialist starts with a fresh, mostly empty window, and only a short report comes back. The big job gets the same care a small one does.
Token counts and quality scores are illustrative. They show the shape of the effect, not benchmark results.