Subscribe

Frontier agents solved the Rails tasks.

Then the graders checked whether they knew Rails existed.

Issue 814 August 20267 receipts4 min

across 21 atomic Rails tasks and 504 runs, frontier agents solve most of the work, Opus 5 at 92%, but mostly by hand-rolling code: Rails API recall runs from 8% (DeepSeek) to 35% (Fable), and cost decouples from quality, with GPT-5.6 Luna clearing 73% for 91 cents total while Opus costs 132x the price for nineteen extra points.

Before you read on. Your call?

it holds, within its stated scope. The methodology is the strongest this desk has seen from an indie benchmark: one frozen harness, one bash tool, default settings, hidden behavior tests, three runs per model per task for $491, and the corpus plus harness going open source. The honest limiters are in the report itself: tests check behavior, so hand-rolled fixes pass like idiomatic ones, and six of 21 tasks are solved by every run of every model. Also on the record: Fable 5 would lead at ~95% but went zero for three on the one task worded like a pen-test report.

8%RAILS API RECALL FLOOR
91CENTS FOR LUNA'S ENTIRE 63-RUN BENCHMARK
132xOPUS COST PREMIUM FOR NINETEEN MORE POINTS

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

Say this in tomorrow's meeting“Finally a benchmark with receipts: frozen harness, published runs, stated limits. The models solve Rails tasks while mostly not using Rails, and after the first dollar, price stops predicting score.”

Receipts

  1. Supports rubyonrails.org: Rails API recall runs from 8% (DeepSeek) to 35% (Fable). Those same two models are 1.4× apart on score and four times apart on recall: knowing Rails separates the field far more sharply than solving the tasks does.
  2. Supports rubyonrails.org: Luna ran all 63 of its runs for 91 cents and beat two models that cost more. Sol costs 36× Luna and buys eleven points. Opus costs 132× and buys nineteen.
  3. Context rubyonrails.org: Fable would lead the field at ~95%, except it went zero for three on the one task worded like a pen-test report.
  4. Context rubyonrails.org: 3 runs per model per task, 504 runs total, $491
  5. Refutes rubyonrails.org: the tests check behavior, not implementation, so a hand-rolled fix passes just like an idiomatic one
  6. Context artificialanalysis.ai: 61 Artificial Analysis Intelligence Index
  7. Context microsoft.ai: Despite this, our model is toe-to-toe with Claude Opus 4.6 on SWE-Bench Pro.

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.