Scientific software engineering benchmark

SWE-bench
Science

Measuring whether coding agents can repair repository-level scientific software while preserving its scientific contracts.

Tasks
0
Repositories
0
Domains
0
01

Leaderboard

Pass@1 versus mean token consumption per task. Select a point or model to inspect the configuration.

Claude-Opus-5 (max)47.90% Pass@17.048M input
Interactive scatter plot of model Pass@1 against mean token use.

All configurations are evaluated on the same 119 tasks. Token counts are per-task means; same-color lines connect different reasoning depths of the same model.

02

Model results

9 configurations
Rank
01Claude Code96.64%75.11%68.60%97.37%47.90%38.46%65.31%27.78%
02Codex99.16%78.82%72.30%97.66%46.22%36.54%59.18%38.89%
03Claude Code100.00%73.16%65.77%96.58%42.02%26.92%57.14%44.44%
04Kimi Code98.32%66.34%57.55%94.94%35.29%25.00%44.90%38.89%
05Codex94.12%63.61%53.81%97.53%31.93%17.31%46.94%33.33%
06Codex93.28%61.89%51.09%94.92%24.37%11.54%36.73%27.78%
07Claude Code98.32%61.41%52.34%95.74%23.53%19.23%26.53%27.78%
08Claude Code100.00%58.77%47.67%95.02%19.33%11.54%28.57%16.67%
09Codex96.64%51.79%38.33%95.16%14.29%5.77%24.49%11.11%

Pass@1 requires every applicable private test to pass. Issue, Expert, and Engineering report Pass@1 for the three task paradigms.