Roga Digital

Labs · Benchmarks

Agentic coding, scored in schmeckles

My own hands-on ratings of AI coding models, measured in schmeckles. The schmeckle is a made-up unit — each score is my gut feel from daily use — so read them side by side, not against a ceiling.

Provider
Status
Sort
View
AI coding models ranked by schmeckle score. Higher is better.
ModelProviderSchmecklesStatus
Fable 5Anthropic105Suspended — pulled from service
Opus 4.8Anthropic98Settled
Opus 4.7Anthropic52Settled
Gippity 5.4OpenAI58Settled
Gemini 3.1Google72Settled
Opus 4.6Anthropic85Settled
Gemini 3 FlashGoogle58Settled
Gippity 5.2OpenAI63Settled
Opus 4.5Anthropic80Settled
Gemini 3 ProGoogle68Settled
Gippity 5.1OpenAI60Settled
Haiku 4.5Anthropic69Settled
Sonnet 4.5Anthropic75Settled
Gippity 5 CodexOpenAI38Settled
Gippity 5OpenAI40Settled
Opus 4.1Anthropic73Settled
Gemini 2.5 Deep ThinkGoogle48Settled
o3-proOpenAI44Settled
= schmeckles settled score tentative — still forming an opinion suspended — pulled from service Anthropic OpenAI Google

Notes

Fable 5

Two weeks on, the suspension still reads to me as a policy and access call rather than a judgment on Fable 5 as a coding tool. Pulling a model three days after launch is the part that doesn’t sit right: anything this new ships with jailbreaks still waiting to be found, and the first few weeks in the open are exactly when a vendor finds and patches them. Fable 5 never got that window.

Fable 5 was the best model I’d used — it tops this board on quality alone — though it ran more cautious than 4.8, declining work I expected it to take on. Three days after launch the US government suspended it on national-security grounds: a policy and access-control call, not a verdict on the model as a coding tool. I’m recording the score as I found it, before access was pulled.

Opus 4.8

Bumping my tentative score to 99. Since release, 4.8 has been the most dependable model I’ve used, and by a clear margin my default. It rarely loses the thread, holds focus across long tasks without wandering or looping back through bad patches, and manages its own context well. I’m still looking for work it can’t finish — features like /goal lean on that reliability and hold up.

Opus 4.8 pairs well with Claude Code’s new Dynamic Workflow: the harness lets it write a small orchestration script that fans subagents out in parallel, pipes results between stages, has some generate while others judge, then synthesizes. I’d been approximating this by hand with commands and skills, so spawning one from inside a skill is the part I value most. The one cost: “workflow” is now a loaded keyword, and I keep escaping a word I use all day.

Opus 4.7

4.7 is the one release I couldn’t make work. It lost the thread constantly and made basic mistakes, and after enough wasted sessions I rolled back to 4.6 — taking its much smaller context window as the better trade. The only model here I actively stopped using.

Gippity 5.4

I gave Gippity 5.4 a fair shot — back in Cursor I’d flip between it and other models to see where each one fit. It was decent at prose until it wasn’t: the OpenAI models tend to over-produce, padding answers with volume rather than substance. By the time 4.5 landed, I’d stopped reaching for them.

Gemini 3.1

I still like Gemini 3.1 — I just don’t reach for it much in coding, only because Claude and Claude Code are where the power is for me right now. On value it’s hard to beat, and the quality-to-cost holds up. It doesn’t hit the ceiling I find in the top Claude models, but it stays a genuinely good option.

Opus 4.6

By 4.6, the reliability that started with 4.5 had settled in. These were the first models I’d trust without watching closely — before them, output needed constant correcting and rewriting. They still missed on genuinely hard tasks, but the limits were clear, and stepping in where they fell short felt fine.

Opus 4.5

4.5 was the first model where the quality crossed a line for me. Despite the high cost, it was the point I stopped worrying about constant re-prompting and rewriting, and trusted it enough to leave Cursor’s diff-review flow for editing directly in the file.

◊ schmeckles · scores are subjective & reflect my own hands-on use.