Frontier agents solved the Rails tasks.
Then the graders checked whether they knew Rails existed.
across 21 atomic Rails tasks and 504 runs, frontier agents solve most of the work, Opus 5 at 92%, but mostly by hand-rolling code: Rails API recall runs from 8% (DeepSeek) to 35% (Fable), and cost decouples from quality, with GPT-5.6 Luna clearing 73% for 91 cents total while Opus costs 132x the price for nineteen extra points.
Before you read on. Your call?
VERIFIED
8%
it holds, within its stated scope. The methodology is the strongest this desk has seen from an indie benchmark: one frozen harness, one bash tool, default settings, hidden behavior tests, three runs per model per task for $491, and the corpus plus harness going open source. The honest limiters are in the report itself: tests check behavior, so hand-rolled fixes pass like idiomatic ones, and six of 21 tasks are solved by every run of every model. Also on the record: Fable 5 would lead at ~95% but went zero for three on the one task worded like a pen-test report.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
Receipts
- Supports rubyonrails.org:
Rails API recall runs from 8% (DeepSeek) to 35% (Fable). Those same two models are 1.4× apart on score and four times apart on recall: knowing Rails separates the field far more sharply than solving the tasks does.
- Supports rubyonrails.org:
Luna ran all 63 of its runs for 91 cents and beat two models that cost more. Sol costs 36× Luna and buys eleven points. Opus costs 132× and buys nineteen.
- Context rubyonrails.org:
Fable would lead the field at ~95%, except it went zero for three on the one task worded like a pen-test report.
- Context rubyonrails.org:
3 runs per model per task, 504 runs total, $491
- Refutes rubyonrails.org:
the tests check behavior, not implementation, so a hand-rolled fix passes just like an idiomatic one
- Context artificialanalysis.ai:
61 Artificial Analysis Intelligence Index
- Context microsoft.ai:
Despite this, our model is toe-to-toe with Claude Opus 4.6 on SWE-Bench Pro.
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.