Shared run
started Jul 23, 19:56 UTC · took 81m 4s · pass@1 · tokens 14.5M in / 1.8M out · cost —
benchmark davidgomes/yourcodingbench
{
"api": "v1",
"caps": {
"agentWallClockMin": 30
},
"driver": "cloud-agent",
"harness": "cloud-agent",
"rulesIncluded": true,
"promptTemplate": 2,
"cursorAgentVersion": null
}timeout: 2 · completed: 67 · empty_diff: 1
Score vs tokens
Y = judge pass rate (passed/judged, D27) · X = mean tokens per passed attempt (input + output) over passed attempts reporting usage, reversed — right = fewer tokens · hover a point for exact numbers
Model comparison
judge verdicts (pass rate)
| Task | gpt-5.4-mini | gpt-5.4-nano | claude-sonnet-4 | kimi-k2.6 | grok-4.5 |
|---|---|---|---|---|---|
| PR #11 · ycb_yourcodingbench_pr11 | FAIL | FAIL | FAIL | —timeout | PASS |
| PR #13 · ycb_yourcodingbench_pr13 | FAIL | FAIL | FAIL | FAIL | FAIL |
| PR #17 · ycb_yourcodingbench_pr17 | FAIL | FAIL | FAIL | FAIL | FAIL |
| PR #2 · ycb_yourcodingbench_pr2 | FAIL | FAIL | FAIL | FAIL | PASS |
| PR #24 · ycb_yourcodingbench_pr24 | FAIL | FAIL | FAIL | FAIL | FAIL |
| PR #29 · ycb_yourcodingbench_pr29 | FAIL | FAIL | FAIL | FAIL | FAIL |
| PR #31 · ycb_yourcodingbench_pr31 | FAIL | FAIL | FAIL | FAIL | PASS |
| PR #35 · ycb_yourcodingbench_pr35 | FAIL | FAIL | FAIL | FAIL | FAIL |
| PR #37 · ycb_yourcodingbench_pr37 | FAIL | FAIL | FAIL | FAIL | FAIL |
| PR #40 · ycb_yourcodingbench_pr40 | FAIL | FAIL | FAIL | —timeout | FAIL |
| PR #5 · ycb_yourcodingbench_pr5 | FAIL | FAIL | FAIL | FAIL | PASS |
| PR #6 · ycb_yourcodingbench_pr6 | PASS | PASS | PASS | PASS | PASS |
| PR #7 · ycb_yourcodingbench_pr7 | FAIL | FAIL | FAIL | FAIL | PASS |
| PR #8 · ycb_yourcodingbench_pr8 | FAIL | FAIL | FAIL | —empty_diff | FAIL |
Powered by YourCodingBench — benchmark coding models on your own repo's merged PRs. sign in
Tokens — agent usage
16.3M total — 14.5M in / 1.8M out
68 of 70 terminal attempts reporting — attempts that failed before usage capture (and reference attempts) carry no usage
Agent-phase tokens only, per-model split on the cards below. Judge-call usage is not measured today. Missing usage reads n/a — never a silent 0.
gpt-5.4-mini
14 attempts · avg duration 2m 45s
tokens 4.6M in / 337k out · mean 353k/attempt
gpt-5.4-nano
14 attempts · avg duration 3m 45s
tokens 6.1M in / 390k out · mean 464k/attempt
claude-sonnet-4
14 attempts · avg duration 3m 11s
tokens 896 in / 199k out · mean 14.3k/attempt
kimi-k2.6
14 attempts · avg duration 5m 46s
tokens 1.3M in / 494k out · mean 153k/attempt · 12 of 14 attempts reporting
3 attempts failed — timeout: 2 · empty_diff: 1 (no diff produced; excluded from the pass rate, never counted as a judge FAIL)
grok-4.5
14 attempts · avg duration 2m 54s
tokens 2.5M in / 342k out · mean 203k/attempt
judge = judge_verdict pass rate (passed/judged) — gpt-5.6-luna · reasoning=none (pinned, D22) · verdict rubric v2 (D27: PASS = a correct solution to the same problem as the reference, equivalent or a valid alternative; FAIL = anything less). The only scorer since D26 (judge-only grading).