Model
One code-generation call, followed by execution. No iterative tool or feedback loop.
Leaderboard · September 2026
Model and Model + Harness results on 1,000 econometric replication tasks.
Scroll the table to see all metrics →
| # | Model | Harness | Full replication | Partial replication | Execution success | Coefficient direction | Significance level |
|---|---|---|---|---|---|---|---|
| 1 | GPT-5.6 Sol | DeepAgents | 51.8% | 69.9% | 91.0% | 87.9% | 78.7% |
| 2 | Claude Opus 4.8 | DeepAgents | 39.2% | 62.1% | 83.4% | 79.3% | 69.6% |
| 3 | Gemini 3.1 Pro | DeepAgents | 39.1% | 63.2% | 85.0% | 81.1% | 71.8% |
| 4 | Kimi K3 | DeepAgents | 39.0% | 63.5% | 86.7% | 80.4% | 70.9% |
| 5 | GPT-5.6 Sol | — | 33.8% | 43.0% | 55.7% | 53.7% | 48.5% |
| 6 | Qwen3.7-Max | DeepAgents | 24.8% | 37.9% | 52.3% | 48.8% | 43.3% |
| 7 | Gemini 3.1 Pro | — | 24.6% | 36.6% | 54.3% | 50.6% | 43.7% |
| 8 | Kimi K3 | — | 23.0% | 36.6% | 52.6% | 49.4% | 43.1% |
| 9 | Qwen3.7-Max | — | 19.8% | 31.5% | 48.6% | 45.4% | 39.6% |
| 10 | Claude Opus 4.8 | — | 18.4% | 30.5% | 46.7% | 43.7% | 37.0% |
| 11 | DeepSeek V4 Pro | DeepAgents | 18.3% | 27.0% | 35.2% | 33.8% | 30.8% |
| 12 | DeepSeek V4 Pro | — | 12.9% | 19.5% | 28.1% | 26.0% | 23.6% |
Published results · 17 Sep 2026
Research comparison. Historical run settings differ; these scores are not an equal-cost comparison.
02 / Research
Our goal is to understand how models and harnesses can improve econometric replication, and turn evaluation evidence into better systems and training signals.
01
Study how different models and harness architectures work together, and which combinations improve replication performance.
02
Explore a feedback-driven cycle of automatic harness design, evaluation and refinement, using execution trajectories and replication outcomes to guide each update.
03
Investigate how validated agent trajectories, tool feedback and replication outcomes can support model post-training, including trajectory distillation and reinforcement learning.
One code-generation call, followed by execution. No iterative tool or feedback loop.
The current results use DeepAgents, with up to six model calls and four trial tool executions. The final program is replayed and scored.
Published result archives: Hugging Face PR #4.
Full replication comes from the published evaluation summaries (local-paper-v1). The other four metrics come from archived per-task scores (hf-leaderboard-v1). These scoring profiles have not been verified against the official scorer. Unknown and invalid outcomes remain in the 1,000-task denominator without earning successes.
The six models share Selected_1000 at dataset revision 59f9512. Execution environments, time limits and recovery protocols differ across historical batches. See the accepted archives for per-run settings and evidence status.