Data · Selected_1000

The benchmark data, task by task.

InferenceNet turns tables from published economics, finance and management papers into replication tasks. Each task names an outcome, a treatment, the controls, the estimation method and a data file, and asks for the coefficient, standard error and p-value of one reported effect. The full benchmark holds 2,644 tasks from 36 journals; Selected_1000 is the pinned 1,000-task list that the internal study runs on; counts are computed from this list unless another source is explicitly noted.

01 / Scale

The full benchmark and the pinned thousand

The data card describes the whole benchmark. Selected_1000 is the fixed subset that the internal paired study and every chart on this page use, so its counts can be recomputed by anyone from the same CSV.

Full benchmark · Hugging Face data card

  • 2,644Research tasksplanned scale ≈ 10,000
  • 36Academic journals
  • 6Research disciplines
  • 2022–2025Publication yearsmainly

Selected_1000 · the evaluated subset

  • 1,000Taskspinned task list
  • 15Journals
  • 286Distinct article stringsas written in the task list
  • 14Raw method labelsfolded into 5 classes
  • 1,000/1,000Tasks labelled Statareference programs are .do files
  • 5,207Files in the snapshot
  • 54.28GBLogical size of the snapshot
  • 0001–1125Task id rangenon-contiguous

02 / Journals

Fifteen journals, one heavy tail

Management Science alone supplies 300 of the 1,000 tasks; with the Journal of Economic History (156) and Econometrica (152) the three largest sources cover 608. Business, management and economics journals dominate; law and economics (113) and the finance journals — led by the Journal of Financial Economics (92) — follow. Below the top eight, no journal contributes more than 21 tasks.

Tasks per journal

All 15 journals in Selected_1000, ranked by task count (n = 1,000).

Data table
Tasks per journal, Selected_1000 (n = 1,000)
JournalTasksShare
Management Science30030.0%
Journal of Economic History15615.6%
Econometrica15215.2%
The Journal of Law and Economics11311.3%
Journal of Financial Economics929.2%
American Economic Review565.6%
Journal of Political Economy404.0%
Marketing Science272.7%
Review of Economic Studies212.1%
Journal of Finance131.3%
Journal of Financial and Quantitative Analysis131.3%
Quarterly Journal of Economics101.0%
The Review of Financial Studies30.3%
Journal of Health Economics30.3%
Review of Finance10.1%

Source: CamoAiLab/InferenceNet, revision 59f9512, task list Selected_1000/1000_new.csv; journal column counted over the 1,000 tasks.

Method mix inside the eight largest journals

Each bar is one journal scaled to 100%; segments are the five method classes. OLS dominates everywhere except The Journal of Law and Economics, where difference-in-differences and regression discontinuity designs take over.

Data table
Method class by journal, eight largest journals (task counts)
JournalOLSDIDIVOtherRDDn
Management Science25639131300
Journal of Economic History1551000156
Econometrica10972691152
The Journal of Law and Economics31646012113
Journal of Financial Economics746111092
American Economic Review54110056
Journal of Political Economy32033240
Marketing Science18630027

Source: journal × method cross-tabulation of the pinned task list; classes as defined in section 03.

The complete cross-tabulation: task counts for all 15 journals by method class. Rows sum to each journal's n; columns sum to the class totals of section 03.

All 15 journals by method class (task counts, n = 1,000).
JournalOLSDIDIVOtherRDDnShare
Management Science2563913130030.0%
Journal of Economic History155100015615.6%
Econometrica1097269115215.2%
The Journal of Law and Economics3164601211311.3%
Journal of Financial Economics7461110929.2%
American Economic Review541100565.6%
Journal of Political Economy320332404.0%
Marketing Science186300272.7%
Review of Economic Studies108300212.1%
Journal of Finance100030131.3%
Journal of Financial and Quantitative Analysis93010131.3%
Quarterly Journal of Economics100000101.0%
The Review of Financial Studies3000030.3%
Journal of Health Economics2001030.3%
Review of Finance0010010.1%
All journals7731355521161,000100%

03 / Methods

Five estimation classes

Three in four tasks are ordinary least squares, usually with fixed effects and clustered errors. Difference-in-differences is the second most common design; instrumental variables and regression discontinuity are rare, and each needs a specific first stage or bandwidth that only the requirement text spells out.

Estimation method, five classes

Five-class labels over the 1,000 Selected_1000 tasks.

Data table
Estimation method, five classes, Selected_1000 (n = 1,000)
ClassTasksShare
OLS77377.3%
DID13513.5%
IV555.5%
Other212.1%
RDD161.6%

Class mapping: OLS includes panel OLS; IV includes IV-2SLS; Other = Probit, Logit, Negative binomial, LMM, PSM and Cox proportional hazards.

The 14 raw method labels in the task list and the class each one is folded into (n = 1,000).
Raw labelClassTasks
OLSOLS743
panel OLSOLS30
DIDDID135
IV-2SLSIV35
IVIV20
RDDRDD16
PROBITOther7
Negative binomialOther4
LOGITOther3
ProbitOther2
LMMOther2
PSMOther1
Cox proportional hazardsOther1
LogitOther1
Labels are printed exactly as they appear in the task list, so PROBIT/Probit and LOGIT/Logit are separate raw labels of the same estimator.

04 / Fields

Five fields, as the publisher counts them

Economics carries most of the weight, followed by finance and accounting. The field labels come from the publisher's own figure, not from the task list, so they are shown as transcribed rather than recomputed.

Tasks per field

Five fields, transcribed from the publisher's data-card figure (n = 1,000).

Data table
Tasks per field, publisher's data-card figure (n = 1,000)
FieldTasksShare
Economics62762.7%
Finance & Accounting18418.4%
Management & Strategy898.9%
Innovation & Information Systems676.7%
Marketing333.3%

Source: transcribed from the field figure on the CamoAiLab/InferenceNet data card, which was drawn from the earlier task list.

Caution

These five counts are transcribed from the publisher's data-card figure, which was based on the earlier task list. No field label exists in Selected_1000/1000_new.csv, so the distribution cannot be recomputed from the pinned list and is not recomputed here. Read it as the publisher's description of the subset, not as a column of the task list.

The journal distribution in section 02 is the recomputed view of the same question: which literatures the tasks come from.

05 / Shape

What the tasks ask for

Beyond the method, a task is shaped by the answer it expects and the specification it demands: how significant the published effect is, its sign, how many controls it carries, which estimation details are tagged, and what the data look like. These distributions decide what a model must get right.

Reference p-value band

Roughly half of the published effects are significant at the 1% level, but 21.8% are not significant even at the 10% level — a model that assumes every reported effect is significant gets that band wrong (n = 997 tasks with a numeric reference p-value).

Data table
Reference p-value band (n = 997)
BandTasksShare
p < 0.0151952.1%
0.01 ≤ p < 0.0518919.0%
0.05 ≤ p < 0.10727.2%
p ≥ 0.1021721.8%

Bands follow the leaderboard's Significant Level Correctness metric. Three tasks have no numeric reference p-value.

Sign of the reference coefficient

Positive effects outnumber negative ones five to three, so getting the sign right is not a coin toss but far from automatic (n = 999 tasks with a numeric reference coefficient).

Data table
Sign of the reference coefficient (n = 999)
SignTasksShare
Positive62462.5%
Negative37537.5%

Sign of the reference coefficient in the task list; one task has no numeric reference coefficient.

Declared control variables per task

Most tasks carry a short control list, but one in five declares nine or more terms, often factor-variable interactions written in Stata syntax (n = 1,000).

Data table
Declared control variables per task (n = 1,000)
ControlsTasksShare
015915.9%
1–339639.6%
4–823123.1%
9+21421.4%

Counted from the control_variables field of the task list.

Tag families

Tags are multi-label, so counts do not sum to 1,000. Fixed effects and clustered standard errors are the norm; a specification that ignores them may still run and still be wrong (n = 1,000 tasks).

Data table
Tag families: tasks carrying each family (multi-label, n = 1,000)
Tag familyTasksShare of tasks
Fixed effects69269.2%
Clustered SE61261.2%
Data processing30630.6%
Robust SE23223.2%
DID / event study282.8%
Weighted regression242.4%
Instrumental variable131.3%
Regression discontinuity60.6%

Raw tag strings of the task list grouped into families; a task counts once per family.

Declared input files by format

Almost every task reads a single Stata .dta file; 14 tasks declare more than one input file (1,022 declared files over 1,000 tasks).

Data table
Declared input files by format (1,022 files over 1,000 tasks)
FormatFiles
.dta1,001
.csv19
no extension1
.xlsx1

Counted from the data_source field of the task list; 14 tasks declare more than one input file.

Tasks per article, top six

One paper can yield many tasks — one column or row of a table each — so the 1,000 tasks come from 286 distinct article strings, and the six most-mined articles alone account for 167 tasks.

Six articles with the most tasks in Selected_1000 (article strings as written in the task list).
Article stringTasks
Closing the Gender Profit Gap?37
Deterrence and Compellence in the Parliament33
Institutions-Trade-and-Growth-The-Ancient-Greek-Case-of-Proxenia31
economic-uncertainty-and-divisive-politics-evidence-from-the-dos-espanas24
sovereign-collateral21
Liquidity Effects of Litigation Risk Evidence from a Legal Shock21
286 distinct article strings over 1,000 tasks. Two of these six articles are also cases in section 07 (tasks 0104 and 0145).

Length of the requirement text

  • 27charsShortest requirementother_requirements field
  • 210charsMedian requirementabout two sentences
  • 651charsLongest requirementa full estimation recipe

The requirement text is the only free-form part of a task: the sample restriction, the clustering variable, the fixed effects and the bandwidth all live there. Cases in section 07 show requirements from 27 to several hundred characters.

06 / Anatomy

Inside a task package

Every task is a small directory: a JSON row, the declared data and a reference Stata program. The model sees a single rendered instruction and the data; the reference program and the reference answer stay on the scorer side.

Package layout · task 00013 files · 650,587 B
0001/
├── task_row.json        986 B
├── data/
│   └── Data_reg.dta 646,607 B
└── do/
    └── task_0001.do   2,994 B
                   ───────────
                     650,587 B  ≈ 0.65 MB

task_row.json holds the task fields and the reference answer; data/ is the declared input, mounted read-only at /data; do/ is the reference Stata program, never shown to the model. Sizes are the byte counts in the pinned revision: the data file is 99.4% of the package, and the instruction the model receives is smaller than a kilobyte.

Output contract/output/<index>_result.json
{
  "coefficient":    <number>,
  "standard_error": <number>,
  "p_value":        <number>
}

Only these three fields, and they must refer to the requested effect of x on y. Task 0001 writes to /output/1_result.json; the scorer compares the file with the reference triplet.

task_row.json · 14 fields

  • Rendered into the instruction

    • index
    • y
    • x
    • control_variables
    • data_source
    • method
    • other_requirements

    These seven fields become the one instruction below; index also sets the output path.

  • Metadata, not part of the instruction text

    • journal
    • article
    • tags
    • notation
    • language
    • sequence_id

    Used to describe and audit the task: where it comes from, which table cell it reproduces, and which estimation details it is tagged with.

  • Kept on the scorer side Withheld

    • answer
    • do/task_<id>.do

    The reference triplet and the reference Stata program are never shown to the model; they exist only for scoring.

Exact model-visible instruction · task 0001as rendered by the evaluator's task loader
Please use the panel OLS method to compute the effect of lndispinc on lnexpenditure. You also need to control the following control variables: c.size#i.quarter, c.fam69r#i.quarter, c.fam1012r#i.quarter, c.fam1316r#i.quarter, c.fam17mr#i.quarter, c.fam17fr#i.quarter, i.tfe. Besides, you need to consider the following requirements: Estimate a panel OLS regression of log consumption on log disposable income with household fixed effects, month-year fixed effects, household-composition controls interacted with quarter dummies, and standard errors clustered by household.. You could load the corresponding data from /data/0001/data/Data_reg.dta. At the end of the program, please print the coefficient, standard error, p value of the effect in a json format like {"coefficient": 0.1, "standard_error": 0.1, "p_value": 0.1}, and output the json string as json file to /output/1_result.json

This is the whole prompt for the task: the controls are Stata factor-variable syntax, the requirement text is pasted verbatim (double full stop included), and the answer format is fixed. Nothing about the journal, the article or the published estimate is included.

07 / Cases

Six real tasks

One task from each of the main designs, taken verbatim from the pinned list: what the paper asked, what the task specifies, and exactly what the model is told. Reference values are shown for 0001 and 0011 only; the other four stay withheld.

Task 0001 Panel OLS

Consumption smoothing in interwar Japan

How much does a working-class household's consumption move when its disposable income changes? A monthly household panel with household and month-year fixed effects, composition controls interacted with quarter dummies, and household-clustered standard errors.

Specification

Outcome (y)
lnexpenditure
Treatment (x)
lndispinc
Controls
c.size#i.quarter, c.fam69r#i.quarter, c.fam1012r#i.quarter, c.fam1316r#i.quarter, c.fam17mr#i.quarter, c.fam17fr#i.quarter, i.tfe
Requirements
Estimate a panel OLS regression of log consumption on log disposable income with household fixed effects, month-year fixed effects, household-composition controls interacted with quarter dummies, and standard errors clustered by household.

Data and provenance

Data source
0001/data/Data_reg.dta
Tags
  • data processing
  • cluster standard error
  • fixed effect
Journal
Journal of Economic History
Article
consumption smoothing in the working class households of interwar japan
Notation
Table 3,Panel A,Row Total consumption
Exact model-visible instructionas rendered by the task loader
Please use the panel OLS method to compute the effect of lndispinc on lnexpenditure. You also need to control the following control variables: c.size#i.quarter, c.fam69r#i.quarter, c.fam1012r#i.quarter, c.fam1316r#i.quarter, c.fam17mr#i.quarter, c.fam17fr#i.quarter, i.tfe. Besides, you need to consider the following requirements: Estimate a panel OLS regression of log consumption on log disposable income with household fixed effects, month-year fixed effects, household-composition controls interacted with quarter dummies, and standard errors clustered by household.. You could load the corresponding data from /data/0001/data/Data_reg.dta. At the end of the program, please print the coefficient, standard error, p value of the effect in a json format like {"coefficient": 0.1, "standard_error": 0.1, "p_value": 0.1}, and output the json string as json file to /output/1_result.json
Reference (scorer side)
coefficient
0.391865
standard_error
0.037674
p_value
4.20e-21

Rounded to six decimals. Shown because the estimate is printed in the source paper; during evaluation it never leaves the scorer.

Task 0011 OLS Real run

Economic uncertainty and divisive politics (the “two Spains”)

How does socio-economic conflict relate to economic policy uncertainty, 1905–1945? OLS with month-clustered standard errors.

Specification

Outcome (y)
EPU0month_simsn_w
Treatment (x)
Wscmonth_simsn_w
Controls
Wnamonth_simsn_w, Wmimonth_simsn_w, Wremonth_simsn_w
Requirements
Replicate PDF Table 2 column 1: OLS of EPU on the four political division variables for 1905-1945, clustering standard errors by month.

Data and provenance

Data source
0011/data/data_np.dta
Data shape
4,550 observations × 38 columns in the recorded run. Restricting to 1905–1945 gives 984 observations; excluding 7 with missing regression variables leaves 977.
Tags
  • cluster standard error
Journal
Journal of Economic History
Article
economic uncertainty and divisive politics evidence from the dos espanas
Notation
Table 2, Column 1, Row Socioec. conflict
Exact model-visible instructionas rendered by the task loader
Please use the OLS method to compute the effect of Wscmonth_simsn_w on EPU0month_simsn_w. You also need to control the following control variables: Wnamonth_simsn_w, Wmimonth_simsn_w, Wremonth_simsn_w. Besides, you need to consider the following requirements: Replicate PDF Table 2 column 1: OLS of EPU on the four political division variables for 1905-1945, clustering standard errors by month.. You could load the corresponding data from /data/0011/data/data_np.dta. At the end of the program, please print the coefficient, standard error, p value of the effect in a json format like {"coefficient": 0.1, "standard_error": 0.1, "p_value": 0.1}, and output the json string as json file to /output/11_result.json
Reference (scorer side)
coefficient
0.272736
standard_error
0.054225
p_value
6.90e-07

Rounded to six decimals. Shown because this task is the one walked through on the Agent page; during evaluation it never leaves the scorer.

Task 0104 DID

Deterrence and compellence in parliament

Did lifting an MP's immunity change how often they submitted formal queries to the government? A difference-in-differences with MP and month-by-year fixed effects, clustered by MP.

Specification

Outcome (y)
Formal Queries (questions)
Treatment (x)
Immunity Lifted x Post
Controls
MP fixed effects; month-by-year fixed effects
Requirements
Replicate Table 3 for all MPs. The outcome is the number of times an MP submitted a formal query to members of the government, stored as questions in the data. Estimate Immunity Lifted x Post with MP fixed effects and month-by-year fixed effects, clustering standard errors at the MP level.

Data and provenance

Data source
data/data_july_2021.dta
Tags
  • difference-in-differences; fixed effects; clustered standard errors
Journal
The Journal of Law and Economics
Article
Deterrence and Compellence in the Parliament
Notation
Table 3, Panel B, Column 4, Formal Queries
Exact model-visible instructionas rendered by the task loader
Please use the DID method to compute the effect of Immunity Lifted x Post on Formal Queries (questions). You also need to control the following control variables: MP fixed effects; month-by-year fixed effects. Besides, you need to consider the following requirements: Replicate Table 3 for all MPs. The outcome is the number of times an MP submitted a formal query to members of the government, stored as questions in the data. Estimate Immunity Lifted x Post with MP fixed effects and month-by-year fixed effects, clustering standard errors at the MP level.. You could load the corresponding data from /data/data/data_july_2021.dta. At the end of the program, please print the coefficient, standard error, p value of the effect in a json format like {"coefficient": 0.1, "standard_error": 0.1, "p_value": 0.1}, and output the json string as json file to /output/104_result.json
Reference Withheld

Reference values for this task stay on the scorer side; see the notation pointer for the published table.

Task 0100 IV-2SLS (fuzzy RDD)

Patent validity and litigation

Does filing an inter partes review make a case more likely to settle? A fuzzy regression discontinuity around the one-year filing deadline, estimated as an IV/2SLS second stage with a triangular kernel and case-clustered errors.

Specification

Outcome (y)
settle
Treatment (x)
ipr_file
Controls
time2, inter, stay. Mtd, msj, software, lnfamily_size, lnforwardcites_3yr, lnbackwardcites, lnnplcites, sep, npe, lnnb_applicants, lnnb_inventors, lnpat_clm_ct, miss_claim, Instruments, Electrical_eng, Chemistry, Mechanical_eng
Requirements
Replicate Table 2, the [-120, 120] day second-stage fuzzy regression discontinuity estimate. Instrument IPR filing with the discontinuous jump in filing probability in the 10 days before the one-year deadline, use time2 and its interaction with the jump indicator as running-variable controls, apply the triangular kernel weight with bandwidth 120, include the listed patent and case controls, and cluster standard errors by case id.

Data and provenance

Data source
data/filing.dta
Tags
  • IV-2SLS
  • regression discontinuity
  • cluster standard error
  • fixed effect
Journal
The Journal of Law and Economics
Article
Patent Validity and Litigation Evidence from US Inter Partes Review
Notation
Table 2, [-120, 120] days, Second Stage (Petition Filed)
Exact model-visible instructionas rendered by the task loader
Please use the IV-2SLS method to compute the effect of ipr_file on settle. You also need to control the following control variables: time2, inter, stay. Mtd, msj, software, lnfamily_size, lnforwardcites_3yr, lnbackwardcites, lnnplcites, sep, npe, lnnb_applicants, lnnb_inventors, lnpat_clm_ct, miss_claim, Instruments, Electrical_eng, Chemistry, Mechanical_eng. Besides, you need to consider the following requirements: Replicate Table 2, the [-120, 120] day second-stage fuzzy regression discontinuity estimate. Instrument IPR filing with the discontinuous jump in filing probability in the 10 days before the one-year deadline, use time2 and its interaction with the jump indicator as running-variable controls, apply the triangular kernel weight with bandwidth 120, include the listed patent and case controls, and cluster standard errors by case id.. You could load the corresponding data from /data/data/filing.dta. At the end of the program, please print the coefficient, standard error, p value of the effect in a json format like {"coefficient": 0.1, "standard_error": 0.1, "p_value": 0.1}, and output the json string as json file to /output/100_result.json
Reference Withheld

Reference values for this task stay on the scorer side; see the notation pointer for the published table.

Task 0145 RDD

Liquidity effects of litigation risk

How did a legal shock change firms' cash holdings? A local-linear regression discontinuity at the cutoff with a triangular kernel, clustered by company.

Specification

Outcome (y)
delta_lcash_scaled
Treatment (x)
Trdd
Controls
None declared — the instruction says “There is no control group.”
Requirements
source-equivalent reconstructed RDD from Oliviero_Park_Zou_JLE_file3.do: local linear regression at cutoff 0 with triangular kernel and row-specific bandwidth; cluster by comp_id.

Data and provenance

Data source
0145/data/Regression3.dta
Tags
None
Journal
The Journal of Law and Economics
Article
Liquidity Effects of Litigation Risk Evidence from a Legal Shock
Notation
Table 5, Row Baseline Row
Exact model-visible instructionas rendered by the task loader
Please use the RDD method to compute the effect of Trdd on delta_lcash_scaled. There is no control group. Besides, you need to consider the following requirements: source-equivalent reconstructed RDD from Oliviero_Park_Zou_JLE_file3.do: local linear regression at cutoff 0 with triangular kernel and row-specific bandwidth; cluster by comp_id.. You could load the corresponding data from /data/0145/data/Regression3.dta. At the end of the program, please print the coefficient, standard error, p value of the effect in a json format like {"coefficient": 0.1, "standard_error": 0.1, "p_value": 0.1}, and output the json string as json file to /output/145_result.json
Reference Withheld

Reference values for this task stay on the scorer side; see the notation pointer for the published table.

Task 0194 Probit

Incentives in non-routine analytical team tasks

Does a bonus incentive change behaviour? A probit model on the field (non-laboratory) sample.

Specification

Outcome (y)
bonus
Treatment (x)
incentive45
Controls
None declared — the instruction says “There is no control group.”
Requirements
This regression only retains non-laboratory data.

Data and provenance

Data source
0194/data/table.dta
Tags
  • cluster standard error
Journal
Journal of Political Economy
Article
The Effect of Incentives in Nonroutine Analytical Team Tasks
Notation
Table 2, Column 1
Exact model-visible instructionas rendered by the task loader
Please use the PROBIT method to compute the effect of incentive45 on bonus. There is no control group. Besides, you need to consider the following requirements: This regression only retains non-laboratory data.. You could load the corresponding data from /data/0194/data/table.dta. At the end of the program, please print the coefficient, standard error, p value of the effect in a json format like {"coefficient": 0.1, "standard_error": 0.1, "p_value": 0.1}, and output the json string as json file to /output/194_result.json
Reference Withheld

Reference values for this task stay on the scorer side; see the notation pointer for the published table.

08 / Provenance

Where every number comes from

Except where another source is explicitly noted, counts on this page are computed directly from one pinned CSV in one pinned dataset revision. Anyone with the same revision can recompute the CSV-based counts; nothing here depends on a model run.

Dataset repository
CamoAiLab/InferenceNet on Hugging Face
Pinned revision
59f9512a38e594528807744214a60ee00367434e
Task list
Selected_1000/1000_new.csv · 1,000 rows
Task list sha256
5338c8af548b35dab0e92a63b0acfebb0bc29c7901da4088851025f5709d6bb3
Counts computed
15 Sep 2026, over all 1,000 rows of the pinned list
Evaluator
YTZSR/chatpilot_evaluator on GitHub — task loading, instruction rendering, execution and scoring

How the counts were computed

Journal, method, control, tag and file-format distributions are column counts over the 1,000 rows of the pinned task list; the five method classes fold the 14 raw labels as listed in section 03. Reference-value bands and signs are read from the answer field where it is numeric. The field distribution in section 04 is the one exception: it is transcribed from the publisher's figure, not recomputed.

Answer schema of the 1,000 reference answers.
SchemaTasks
Standard triplet (coefficient, standard_error, p_value)998
Composite (multi-regression)1
Triplet + n1
The output contract in section 06 is the standard triplet.
Scorer reference support: how many tasks carry each reference field (coverage, not model performance).
Reference fieldSupportedOf
Coefficient and direction9991,000
Significance band9971,000
Ids listed as unsupported: 0346, 0915, 0917, 1050, 1109. Coverage per the evaluator's leaderboard documentation.

Selected_1000 keeps the original task ids, which run from 0001 to 1125 with gaps. The snapshot behind these counts holds 5,207 files and 54.28 GB of data.