A coding bench from your own commits

Run the models you actually use against work that already shipped. Public leaderboards average someone else's repo. This one is yours.

Invite-only. Ask a member for an invite

Signing in asks for public repository access only. Private repos are opt-in later.

Don't know a member? Leave an email. We write when a spot opens.

Cost vs pass rateDavid DistSystemsBench · completed Aug 13, 2026
0%20%40%60%80%100%$25$20$15$10$5$0judge pass rateAverage cost per task90.7%$2.98Grok 4.6 high*Grok 4.6Gemini 3.1 ProClaude Sonnet 5 default (high)*GPT-5.6 Terra default (medium)*GPT-5.6 Sol default (medium)*Claude Fable 5 default (high)*

the plot scrolls sideways to the cheap end

Y = judge pass rate, passed over judged + agent failures (timeouts, empty diffs, agent errors, budget: FAILs, D48) · X = average charged cost per task over the same attempts, reversed — right = cheaper · hover a point for exact numbers, tokens included · a line joins one model's thinking levels, weakest to strongest; the family is named once beside its line · the dashed crosshair marks Grok 4.6, the strongest family default on the plot · a starred family name means not every thinking level it offers is on the chart · verdict rubric v5 (D54)

Live numbers from a public run, not a screenshot. David DistSystemsBench Extreme · 16 models. Each dot is one model: what a task costs it, against how often it passed. Cheaper to the right, stronger up, so the top-right corner wins.

How it works

  1. 01

    Connect GitHub

    Connect a repo. Commits become tasks. Private repos stay opt-in until you say otherwise.

  2. 02

    Run models

    Pick the models. One attempt per task per model. The model is the only variable.

  3. 03

    Compare

    Pass rates, tokens, and a score-vs-cost plot. Every attempt keeps its transcript.

How we grade

Each benchmark uses one of these. We do not blend the two.

  • Golden-diff judge

    An LLM compares the model's patch to the reference gold. Benches from your PRs and commits use this. The scorer name is judge_verdict.

  • Harbor tests

    Held-out tests, the way OG DeepSWE grades. Fail-to-pass (F2P) tests must pass. Pass-to-pass (P2P) tests must not fail. The scorer name is deepswe_verdict.

What you get

Run, model comparison4 models · 16 tasks · captured Jul 24, 2026
Model comparison panel from a completed run: judge pass-rate bars for four models, led by claude-opus-4-8 at 15/16 passed, then claude-opus-5 at 11/15, gemini-3-flash at 7/16, and claude-sonnet-4 at 5/16.
Pass rates side by side. One bar per model, passed over judged on every bar, and a transcript behind every attempt.
Benchmark, task listneondatabase/neon · 16 tasks
Task list of a benchmark built from merged pull requests in the neondatabase/neon repository: rows with PR number, title, merge date, and diff size — each row is one benchmark task.
Shipped changes become the tasks. Here, 16 merged PRs from the open-source neondatabase/neon repository. Each keeps its prompt and the real reference diff.