Scientific software engineering benchmark
SWE-bench
Science
Measuring whether coding agents can repair repository-level scientific software while preserving its scientific contracts.
- Tasks
00 - Repositories
00 - Domains
00
01
Leaderboard
Pass@1 versus mean token consumption per task. Select a point or model to inspect the configuration.
All configurations are evaluated on the same 119 tasks. Token counts are per-task means; same-color lines connect different reasoning depths of the same model.
02
Model results
| Rank | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| 01 | Claude Code | 96.64% | 75.11% | 68.60% | 97.37% | 47.90% | 38.46% | 65.31% | 27.78% | |
| 02 | Codex | 99.16% | 78.82% | 72.30% | 97.66% | 46.22% | 36.54% | 59.18% | 38.89% | |
| 03 | Claude Code | 100.00% | 73.16% | 65.77% | 96.58% | 42.02% | 26.92% | 57.14% | 44.44% | |
| 04 | Kimi Code | 98.32% | 66.34% | 57.55% | 94.94% | 35.29% | 25.00% | 44.90% | 38.89% | |
| 05 | Codex | 94.12% | 63.61% | 53.81% | 97.53% | 31.93% | 17.31% | 46.94% | 33.33% | |
| 06 | Codex | 93.28% | 61.89% | 51.09% | 94.92% | 24.37% | 11.54% | 36.73% | 27.78% | |
| 07 | Claude Code | 98.32% | 61.41% | 52.34% | 95.74% | 23.53% | 19.23% | 26.53% | 27.78% | |
| 08 | Claude Code | 100.00% | 58.77% | 47.67% | 95.02% | 19.33% | 11.54% | 28.57% | 16.67% | |
| 09 | Codex | 96.64% | 51.79% | 38.33% | 95.16% | 14.29% | 5.77% | 24.49% | 11.11% |
Pass@1 requires every applicable private test to pass. Issue, Expert, and Engineering report Pass@1 for the three task paradigms.