Labs · Benchmarks
Agentic coding, scored in schmeckles
My own hands-on ratings of AI coding models, measured in schmeckles. The schmeckle is a made-up unit — each score is my gut feel from daily use — so read them side by side, not against a ceiling.
| Model | Provider | Schmeckles | Status |
|---|---|---|---|
| Fable 5 | Anthropic | 105 | Suspended — pulled from service |
| Opus 4.8 | Anthropic | 98 | Settled |
| Opus 4.7 | Anthropic | 52 | Settled |
| Gippity 5.4 | OpenAI | 58 | Settled |
| Gemini 3.1 | 72 | Settled | |
| Opus 4.6 | Anthropic | 85 | Settled |
| Gemini 3 Flash | 58 | Settled | |
| Gippity 5.2 | OpenAI | 63 | Settled |
| Opus 4.5 | Anthropic | 80 | Settled |
| Gemini 3 Pro | 68 | Settled | |
| Gippity 5.1 | OpenAI | 60 | Settled |
| Haiku 4.5 | Anthropic | 69 | Settled |
| Sonnet 4.5 | Anthropic | 75 | Settled |
| Gippity 5 Codex | OpenAI | 38 | Settled |
| Gippity 5 | OpenAI | 40 | Settled |
| Opus 4.1 | Anthropic | 73 | Settled |
| Gemini 2.5 Deep Think | 48 | Settled | |
| o3-pro | OpenAI | 44 | Settled |
Notes
- Fable 5
Two weeks on, the suspension still reads to me as a policy and access call rather than a judgment on Fable 5 as a coding tool. Pulling a model three days after launch is the part that doesn’t sit right: anything this new ships with jailbreaks still waiting to be found, and the first few weeks in the open are exactly when a vendor finds and patches them. Fable 5 never got that window.
Fable 5 was the best model I’d used — it tops this board on quality alone — though it ran more cautious than 4.8, declining work I expected it to take on. Three days after launch the US government suspended it on national-security grounds: a policy and access-control call, not a verdict on the model as a coding tool. I’m recording the score as I found it, before access was pulled.
- Opus 4.8
Bumping my tentative score to 99. Since release, 4.8 has been the most dependable model I’ve used, and by a clear margin my default. It rarely loses the thread, holds focus across long tasks without wandering or looping back through bad patches, and manages its own context well. I’m still looking for work it can’t finish — features like /goal lean on that reliability and hold up.
Opus 4.8 pairs well with Claude Code’s new Dynamic Workflow: the harness lets it write a small orchestration script that fans subagents out in parallel, pipes results between stages, has some generate while others judge, then synthesizes. I’d been approximating this by hand with commands and skills, so spawning one from inside a skill is the part I value most. The one cost: “workflow” is now a loaded keyword, and I keep escaping a word I use all day.
- Opus 4.7
4.7 is the one release I couldn’t make work. It lost the thread constantly and made basic mistakes, and after enough wasted sessions I rolled back to 4.6 — taking its much smaller context window as the better trade. The only model here I actively stopped using.
- Gippity 5.4
I gave Gippity 5.4 a fair shot — back in Cursor I’d flip between it and other models to see where each one fit. It was decent at prose until it wasn’t: the OpenAI models tend to over-produce, padding answers with volume rather than substance. By the time 4.5 landed, I’d stopped reaching for them.
- Gemini 3.1
I still like Gemini 3.1 — I just don’t reach for it much in coding, only because Claude and Claude Code are where the power is for me right now. On value it’s hard to beat, and the quality-to-cost holds up. It doesn’t hit the ceiling I find in the top Claude models, but it stays a genuinely good option.
- Opus 4.6
By 4.6, the reliability that started with 4.5 had settled in. These were the first models I’d trust without watching closely — before them, output needed constant correcting and rewriting. They still missed on genuinely hard tasks, but the limits were clear, and stepping in where they fell short felt fine.
- Opus 4.5
4.5 was the first model where the quality crossed a line for me. Despite the high cost, it was the point I stopped worrying about constant re-prompting and rewriting, and trusted it enough to leave Cursor’s diff-review flow for editing directly in the file.
◊ schmeckles · scores are subjective & reflect my own hands-on use.