Shared run
started Jul 26, 11:18 UTC · took 468m 37s · pass@1 · tokens 14.7M in / 4.1M out · cost $284.66
benchmark my first benchmark based on neon's source code
{
"api": "v1",
"caps": {
"agentWallClockMin": 60
},
"driver": "cloud-agent",
"harness": "cloud-agent",
"rulesIncluded": true,
"promptTemplate": 2,
"cursorAgentVersion": null
}completed: 224
Score vs cost
Y = judge pass rate — passed over judged + agent failures (timeouts, empty diffs, agent errors, budget: FAILs, D48) · X = mean charged cost per passed attempt over passed attempts recording a cost, reversed — right = cheaper · hover a point for exact numbers, tokens included · a line joins one model's thinking levels, weakest to strongest; the family is named once beside its line · the dashed crosshair marks Opus 4.8 high, the strongest family default on the plot
Cost is the charged amount from Cursor's usage API — what the run actually bills on the shared service key (D9). The raw pre-adjustment amount is stored alongside it and shown in the run cost block; the two diverge in both directions by model, so every headline number here is the one that bills.
Model comparison
| Task | claude-opus-4-8-low | claude-opus-4-8-medium | claude-opus-4-8 | claude-opus-4-8-xhigh | claude-opus-5-low | claude-opus-5-medium | claude-opus-5 | claude-opus-5-xhigh | claude-sonnet-5-low | claude-sonnet-5-medium | claude-sonnet-5 | claude-sonnet-5-xhigh | claude-sonnet-4 | gemini-3-flash |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| PR #12668 · ycb_neon_pr12668 | PASS | PASS | PASS | PASS | FAIL | PASS | PASS | PASS | FAIL | PASS | PASS | PASS | FAIL | PASS |
| PR #12684 · ycb_neon_pr12684 | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS |
| PR #12688 · ycb_neon_pr12688 | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS |
| PR #12694 · ycb_neon_pr12694 | PASS | PASS | PASS | FAIL | FAIL | PASS | FAIL | PASS | FAIL | PASS | FAIL | FAIL | FAIL | FAIL |
| PR #12705 · ycb_neon_pr12705 | FAIL | PASS | PASS | PASS | PASS | FAIL | FAIL | FAIL | FAIL | PASS | PASS | PASS | FAIL | FAIL |
| PR #12727 · ycb_neon_pr12727 | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | FAIL | FAIL | PASS | FAIL | FAIL | FAIL |
| PR #12745 · ycb_neon_pr12745 | FAIL | PASS | PASS | PASS | FAIL | FAIL | PASS | PASS | FAIL | FAIL | PASS | PASS | FAIL | FAIL |
| PR #12754 · ycb_neon_pr12754 | PASS | PASS | PASS | PASS | PASS | FAIL | PASS | FAIL | FAIL | PASS | FAIL | PASS | FAIL | FAIL |
| PR #12772 · ycb_neon_pr12772 | FAIL | FAIL | FAIL | PASS | PASS | PASS | FAIL | FAIL | FAIL | PASS | FAIL | PASS | FAIL | FAIL |
| PR #12777 · ycb_neon_pr12777 | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS |
| PR #12782 · ycb_neon_pr12782 | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS |
| PR #12785 · ycb_neon_pr12785 | FAIL | PASS | PASS | PASS | FAIL | FAIL | FAIL | FAIL | FAIL | PASS | PASS | PASS | FAIL | PASS |
| PR #12793 · ycb_neon_pr12793 | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | FAIL |
| PR #12819 · ycb_neon_pr12819 | FAIL | PASS | PASS | PASS | FAIL | FAIL | PASS | PASS | PASS | PASS | PASS | PASS | FAIL | FAIL |
| PR #12826 · ycb_neon_pr12826 | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS |
| PR #12899 · ycb_neon_pr12899 | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS | PASS |
Powered by YourCodingBench — benchmark coding models on your own repo's merged PRs. sign in
judge verdicts (pass rate)
claude-opus-4-8
claude-opus-5
claude-sonnet-5
Tokens & cost — agent usage
18.9M total — 14.7M in / 4.1M out
224 of 224 terminal attempts reporting
$284.66 charged — mean $1.27/attempt
224 of 224 terminal attempts reporting
raw (pre-adjustment) $252.28 — recorded, not billed
Agent-phase usage only, per-model split on the cards below. Judge-call usage is not measured today. Cost is the charged amount from Cursor's usage API — what the run bills on the shared service key (D9); the raw pre-adjustment amount is recorded alongside it and can be higher or lower. Missing usage reads n/a — never a silent 0.
claude-opus-4-8 — thinking levels
claude-opus-4-8-lowthinking low
16 attempts · avg duration 2m 50s
tokens 1.2k in / 185k out · mean 11.6k/attempt
cost $14.52 charged · mean $0.91/attempt
claude-opus-4-8-mediumthinking medium
16 attempts · avg duration 3m 14s
tokens 1.3k in / 197k out · mean 12.4k/attempt
cost $16.17 charged · mean $1.01/attempt
claude-opus-4-8
16 attempts · avg duration 3m 31s
tokens 1.5k in / 240k out · mean 15.1k/attempt
cost $19.27 charged · mean $1.20/attempt
claude-opus-4-8-xhighthinking xhigh
16 attempts · avg duration 7m 29s
tokens 1.9k in / 417k out · mean 26.2k/attempt
cost $28.98 charged · mean $1.81/attempt
claude-opus-5 — thinking levels
claude-opus-5-lowthinking low
16 attempts · avg duration 2m 30s
tokens 1.1k in / 133k out · mean 8.4k/attempt
cost $12.95 charged · mean $0.81/attempt
claude-opus-5-mediumthinking medium
16 attempts · avg duration 5m 21s
tokens 1.6k in / 276k out · mean 17.4k/attempt
cost $21.92 charged · mean $1.37/attempt
claude-opus-5
16 attempts · avg duration 7m 18s
tokens 1.9k in / 399k out · mean 25.1k/attempt
cost $30.03 charged · mean $1.88/attempt
claude-opus-5-xhighthinking xhigh
16 attempts · avg duration 13m 37s
tokens 2.6k in / 659k out · mean 41.4k/attempt
cost $51.92 charged · mean $3.25/attempt
claude-sonnet-5 — thinking levels
claude-sonnet-5-lowthinking low
16 attempts · avg duration 2m 58s
tokens 1.3k in / 155k out · mean 9.8k/attempt
cost $8.96 charged · mean $0.56/attempt
claude-sonnet-5-mediumthinking medium
16 attempts · avg duration 3m 50s
tokens 1.5k in / 199k out · mean 12.5k/attempt
cost $11.66 charged · mean $0.73/attempt
claude-sonnet-5
16 attempts · avg duration 3m 17s
tokens 1.4k in / 181k out · mean 11.4k/attempt
cost $11.35 charged · mean $0.71/attempt
claude-sonnet-5-xhighthinking xhigh
16 attempts · avg duration 5m 3s
tokens 2.1k in / 333k out · mean 20.9k/attempt
cost $19.28 charged · mean $1.20/attempt
claude-sonnet-4
16 attempts · avg duration 4m 26s
tokens 616 in / 225k out · mean 14.1k/attempt
cost $21.02 charged · mean $1.31/attempt
gemini-3-flash
16 attempts · avg duration 4m 13s
tokens 14.7M in / 537k out · mean 952k/attempt
cost $16.62 charged · mean $1.04/attempt
pass rate = passed / (judged + agent failures) — a timeout, empty diff, agent error, or blown budget counts as a FAIL without a verdict (D48); infra failures and owner-cancelled attempts are excluded from the rate and annotated. Verdicts by gpt-5.6-luna · reasoning=none (pinned, D22) · verdict rubric v3 (D44). PASS = a correct solution to the same problem as the reference, equivalent or a valid alternative; FAIL = anything less. The only scorer since D26 (judge-only grading).