A coding bench from your own commits
Run the models you actually use against work that already shipped. Public leaderboards average someone else's repo. This one is yours.
Invite-only. Ask a member for an invite
Signing in asks for public repository access only. Private repos are opt-in later.
Don't know a member? Leave an email. We write when a spot opens.
the plot scrolls sideways to the cheap end
Y = judge pass rate, passed over judged + agent failures (timeouts, empty diffs, agent errors, budget: FAILs, D48) · X = average charged cost per task over the same attempts, reversed — right = cheaper · hover a point for exact numbers, tokens included · a line joins one model's thinking levels, weakest to strongest; the family is named once beside its line · the dashed crosshair marks Grok 4.6, the strongest family default on the plot · a starred family name means not every thinking level it offers is on the chart · verdict rubric v5 (D54)
How it works
- 01
Connect GitHub
Connect a repo. Commits become tasks. Private repos stay opt-in until you say otherwise.
- 02
Run models
Pick the models. One attempt per task per model. The model is the only variable.
- 03
Compare
Pass rates, tokens, and a score-vs-cost plot. Every attempt keeps its transcript.
How we grade
Each benchmark uses one of these. We do not blend the two.
Golden-diff judge
An LLM compares the model's patch to the reference gold. Benches from your PRs and commits use this. The scorer name is judge_verdict.
Harbor tests
Held-out tests, the way OG DeepSWE grades. Fail-to-pass (F2P) tests must pass. Pass-to-pass (P2P) tests must not fail. The scorer name is deepswe_verdict.
What you get

