Data · Selected_1000
The benchmark data, task by task.
InferenceNet turns tables from published economics, finance and management papers into replication tasks. Each task names an outcome, a treatment, the controls, the estimation method and a data file, and asks for the coefficient, standard error and p-value of one reported effect. The full benchmark holds 2,644 tasks from 36 journals; Selected_1000 is the pinned 1,000-task list that the internal study runs on; counts are computed from this list unless another source is explicitly noted.
01 / Scale
The full benchmark and the pinned thousand
The data card describes the whole benchmark. Selected_1000 is the fixed subset that the internal paired study and every chart on this page use, so its counts can be recomputed by anyone from the same CSV.
Full benchmark · Hugging Face data card
- 2,644Research tasksplanned scale ≈ 10,000
- 36Academic journals
- 6Research disciplines
- 2022–2025Publication yearsmainly
Selected_1000 · the evaluated subset
- 1,000Taskspinned task list
- 15Journals
- 286Distinct article stringsas written in the task list
- 14Raw method labelsfolded into 5 classes
- 1,000/1,000Tasks labelled Statareference programs are .do files
- 5,207Files in the snapshot
- 54.28GBLogical size of the snapshot
- 0001–1125Task id rangenon-contiguous
02 / Journals
Fifteen journals, one heavy tail
Management Science alone supplies 300 of the 1,000 tasks; with the Journal of Economic History (156) and Econometrica (152) the three largest sources cover 608. Business, management and economics journals dominate; law and economics (113) and the finance journals — led by the Journal of Financial Economics (92) — follow. Below the top eight, no journal contributes more than 21 tasks.
Tasks per journal
All 15 journals in Selected_1000, ranked by task count (n = 1,000).
Data table
| Journal | Tasks | Share |
|---|---|---|
| Management Science | 300 | 30.0% |
| Journal of Economic History | 156 | 15.6% |
| Econometrica | 152 | 15.2% |
| The Journal of Law and Economics | 113 | 11.3% |
| Journal of Financial Economics | 92 | 9.2% |
| American Economic Review | 56 | 5.6% |
| Journal of Political Economy | 40 | 4.0% |
| Marketing Science | 27 | 2.7% |
| Review of Economic Studies | 21 | 2.1% |
| Journal of Finance | 13 | 1.3% |
| Journal of Financial and Quantitative Analysis | 13 | 1.3% |
| Quarterly Journal of Economics | 10 | 1.0% |
| The Review of Financial Studies | 3 | 0.3% |
| Journal of Health Economics | 3 | 0.3% |
| Review of Finance | 1 | 0.1% |
Source: CamoAiLab/InferenceNet, revision 59f9512, task list Selected_1000/1000_new.csv; journal column counted over the 1,000 tasks.
Method mix inside the eight largest journals
Each bar is one journal scaled to 100%; segments are the five method classes. OLS dominates everywhere except The Journal of Law and Economics, where difference-in-differences and regression discontinuity designs take over.
Data table
| Journal | OLS | DID | IV | Other | RDD | n |
|---|---|---|---|---|---|---|
| Management Science | 256 | 39 | 1 | 3 | 1 | 300 |
| Journal of Economic History | 155 | 1 | 0 | 0 | 0 | 156 |
| Econometrica | 109 | 7 | 26 | 9 | 1 | 152 |
| The Journal of Law and Economics | 31 | 64 | 6 | 0 | 12 | 113 |
| Journal of Financial Economics | 74 | 6 | 11 | 1 | 0 | 92 |
| American Economic Review | 54 | 1 | 1 | 0 | 0 | 56 |
| Journal of Political Economy | 32 | 0 | 3 | 3 | 2 | 40 |
| Marketing Science | 18 | 6 | 3 | 0 | 0 | 27 |
Source: journal × method cross-tabulation of the pinned task list; classes as defined in section 03.
The complete cross-tabulation: task counts for all 15 journals by method class. Rows sum to each journal's n; columns sum to the class totals of section 03.
| Journal | OLS | DID | IV | Other | RDD | n | Share |
|---|---|---|---|---|---|---|---|
| Management Science | 256 | 39 | 1 | 3 | 1 | 300 | 30.0% |
| Journal of Economic History | 155 | 1 | 0 | 0 | 0 | 156 | 15.6% |
| Econometrica | 109 | 7 | 26 | 9 | 1 | 152 | 15.2% |
| The Journal of Law and Economics | 31 | 64 | 6 | 0 | 12 | 113 | 11.3% |
| Journal of Financial Economics | 74 | 6 | 11 | 1 | 0 | 92 | 9.2% |
| American Economic Review | 54 | 1 | 1 | 0 | 0 | 56 | 5.6% |
| Journal of Political Economy | 32 | 0 | 3 | 3 | 2 | 40 | 4.0% |
| Marketing Science | 18 | 6 | 3 | 0 | 0 | 27 | 2.7% |
| Review of Economic Studies | 10 | 8 | 3 | 0 | 0 | 21 | 2.1% |
| Journal of Finance | 10 | 0 | 0 | 3 | 0 | 13 | 1.3% |
| Journal of Financial and Quantitative Analysis | 9 | 3 | 0 | 1 | 0 | 13 | 1.3% |
| Quarterly Journal of Economics | 10 | 0 | 0 | 0 | 0 | 10 | 1.0% |
| The Review of Financial Studies | 3 | 0 | 0 | 0 | 0 | 3 | 0.3% |
| Journal of Health Economics | 2 | 0 | 0 | 1 | 0 | 3 | 0.3% |
| Review of Finance | 0 | 0 | 1 | 0 | 0 | 1 | 0.1% |
| All journals | 773 | 135 | 55 | 21 | 16 | 1,000 | 100% |
03 / Methods
Five estimation classes
Three in four tasks are ordinary least squares, usually with fixed effects and clustered errors. Difference-in-differences is the second most common design; instrumental variables and regression discontinuity are rare, and each needs a specific first stage or bandwidth that only the requirement text spells out.
Estimation method, five classes
Five-class labels over the 1,000 Selected_1000 tasks.
Data table
| Class | Tasks | Share |
|---|---|---|
| OLS | 773 | 77.3% |
| DID | 135 | 13.5% |
| IV | 55 | 5.5% |
| Other | 21 | 2.1% |
| RDD | 16 | 1.6% |
Class mapping: OLS includes panel OLS; IV includes IV-2SLS; Other = Probit, Logit, Negative binomial, LMM, PSM and Cox proportional hazards.
| Raw label | Class | Tasks |
|---|---|---|
| OLS | OLS | 743 |
| panel OLS | OLS | 30 |
| DID | DID | 135 |
| IV-2SLS | IV | 35 |
| IV | IV | 20 |
| RDD | RDD | 16 |
| PROBIT | Other | 7 |
| Negative binomial | Other | 4 |
| LOGIT | Other | 3 |
| Probit | Other | 2 |
| LMM | Other | 2 |
| PSM | Other | 1 |
| Cox proportional hazards | Other | 1 |
| Logit | Other | 1 |
| Labels are printed exactly as they appear in the task list, so PROBIT/Probit and LOGIT/Logit are separate raw labels of the same estimator. | ||
04 / Fields
Five fields, as the publisher counts them
Economics carries most of the weight, followed by finance and accounting. The field labels come from the publisher's own figure, not from the task list, so they are shown as transcribed rather than recomputed.
Tasks per field
Five fields, transcribed from the publisher's data-card figure (n = 1,000).
Data table
| Field | Tasks | Share |
|---|---|---|
| Economics | 627 | 62.7% |
| Finance & Accounting | 184 | 18.4% |
| Management & Strategy | 89 | 8.9% |
| Innovation & Information Systems | 67 | 6.7% |
| Marketing | 33 | 3.3% |
Source: transcribed from the field figure on the CamoAiLab/InferenceNet data card, which was drawn from the earlier task list.
Caution
These five counts are transcribed from the publisher's data-card figure, which was based on the earlier task list. No field label exists in Selected_1000/1000_new.csv, so the distribution cannot be recomputed from the pinned list and is not recomputed here. Read it as the publisher's description of the subset, not as a column of the task list.
The journal distribution in section 02 is the recomputed view of the same question: which literatures the tasks come from.
05 / Shape
What the tasks ask for
Beyond the method, a task is shaped by the answer it expects and the specification it demands: how significant the published effect is, its sign, how many controls it carries, which estimation details are tagged, and what the data look like. These distributions decide what a model must get right.
Reference p-value band
Roughly half of the published effects are significant at the 1% level, but 21.8% are not significant even at the 10% level — a model that assumes every reported effect is significant gets that band wrong (n = 997 tasks with a numeric reference p-value).
Data table
| Band | Tasks | Share |
|---|---|---|
| p < 0.01 | 519 | 52.1% |
| 0.01 ≤ p < 0.05 | 189 | 19.0% |
| 0.05 ≤ p < 0.10 | 72 | 7.2% |
| p ≥ 0.10 | 217 | 21.8% |
Bands follow the leaderboard's Significant Level Correctness metric. Three tasks have no numeric reference p-value.
Sign of the reference coefficient
Positive effects outnumber negative ones five to three, so getting the sign right is not a coin toss but far from automatic (n = 999 tasks with a numeric reference coefficient).
Data table
| Sign | Tasks | Share |
|---|---|---|
| Positive | 624 | 62.5% |
| Negative | 375 | 37.5% |
Sign of the reference coefficient in the task list; one task has no numeric reference coefficient.
Declared control variables per task
Most tasks carry a short control list, but one in five declares nine or more terms, often factor-variable interactions written in Stata syntax (n = 1,000).
Data table
| Controls | Tasks | Share |
|---|---|---|
| 0 | 159 | 15.9% |
| 1–3 | 396 | 39.6% |
| 4–8 | 231 | 23.1% |
| 9+ | 214 | 21.4% |
Counted from the control_variables field of the task list.
Tag families
Tags are multi-label, so counts do not sum to 1,000. Fixed effects and clustered standard errors are the norm; a specification that ignores them may still run and still be wrong (n = 1,000 tasks).
Data table
| Tag family | Tasks | Share of tasks |
|---|---|---|
| Fixed effects | 692 | 69.2% |
| Clustered SE | 612 | 61.2% |
| Data processing | 306 | 30.6% |
| Robust SE | 232 | 23.2% |
| DID / event study | 28 | 2.8% |
| Weighted regression | 24 | 2.4% |
| Instrumental variable | 13 | 1.3% |
| Regression discontinuity | 6 | 0.6% |
Raw tag strings of the task list grouped into families; a task counts once per family.
Declared input files by format
Almost every task reads a single Stata .dta file; 14 tasks declare more than one input file (1,022 declared files over 1,000 tasks).
Data table
| Format | Files |
|---|---|
| .dta | 1,001 |
| .csv | 19 |
| no extension | 1 |
| .xlsx | 1 |
Counted from the data_source field of the task list; 14 tasks declare more than one input file.
Tasks per article, top six
One paper can yield many tasks — one column or row of a table each — so the 1,000 tasks come from 286 distinct article strings, and the six most-mined articles alone account for 167 tasks.
| Article string | Tasks |
|---|---|
| Closing the Gender Profit Gap? | 37 |
| Deterrence and Compellence in the Parliament | 33 |
| Institutions-Trade-and-Growth-The-Ancient-Greek-Case-of-Proxenia | 31 |
| economic-uncertainty-and-divisive-politics-evidence-from-the-dos-espanas | 24 |
| sovereign-collateral | 21 |
| Liquidity Effects of Litigation Risk Evidence from a Legal Shock | 21 |
| 286 distinct article strings over 1,000 tasks. Two of these six articles are also cases in section 07 (tasks 0104 and 0145). | |
Length of the requirement text
- 27charsShortest requirementother_requirements field
- 210charsMedian requirementabout two sentences
- 651charsLongest requirementa full estimation recipe
The requirement text is the only free-form part of a task: the sample restriction, the clustering variable, the fixed effects and the bandwidth all live there. Cases in section 07 show requirements from 27 to several hundred characters.
06 / Anatomy
Inside a task package
Every task is a small directory: a JSON row, the declared data and a reference Stata program. The model sees a single rendered instruction and the data; the reference program and the reference answer stay on the scorer side.
0001/
├── task_row.json 986 B
├── data/
│ └── Data_reg.dta 646,607 B
└── do/
└── task_0001.do 2,994 B
───────────
650,587 B ≈ 0.65 MB
task_row.json holds the task fields and the reference answer; data/ is the declared input, mounted read-only at /data; do/ is the reference Stata program, never shown to the model. Sizes are the byte counts in the pinned revision: the data file is 99.4% of the package, and the instruction the model receives is smaller than a kilobyte.
{
"coefficient": <number>,
"standard_error": <number>,
"p_value": <number>
}
Only these three fields, and they must refer to the requested effect of x on y. Task 0001 writes to /output/1_result.json; the scorer compares the file with the reference triplet.
task_row.json · 14 fields
-
Rendered into the instruction
- index
- y
- x
- control_variables
- data_source
- method
- other_requirements
These seven fields become the one instruction below;
indexalso sets the output path. -
Metadata, not part of the instruction text
- journal
- article
- tags
- notation
- language
- sequence_id
Used to describe and audit the task: where it comes from, which table cell it reproduces, and which estimation details it is tagged with.
-
Kept on the scorer side Withheld
- answer
- do/task_<id>.do
The reference triplet and the reference Stata program are never shown to the model; they exist only for scoring.
Please use the panel OLS method to compute the effect of lndispinc on lnexpenditure. You also need to control the following control variables: c.size#i.quarter, c.fam69r#i.quarter, c.fam1012r#i.quarter, c.fam1316r#i.quarter, c.fam17mr#i.quarter, c.fam17fr#i.quarter, i.tfe. Besides, you need to consider the following requirements: Estimate a panel OLS regression of log consumption on log disposable income with household fixed effects, month-year fixed effects, household-composition controls interacted with quarter dummies, and standard errors clustered by household.. You could load the corresponding data from /data/0001/data/Data_reg.dta. At the end of the program, please print the coefficient, standard error, p value of the effect in a json format like {"coefficient": 0.1, "standard_error": 0.1, "p_value": 0.1}, and output the json string as json file to /output/1_result.json
This is the whole prompt for the task: the controls are Stata factor-variable syntax, the requirement text is pasted verbatim (double full stop included), and the answer format is fixed. Nothing about the journal, the article or the published estimate is included.
07 / Cases
Six real tasks
One task from each of the main designs, taken verbatim from the pinned list: what the paper asked, what the task specifies, and exactly what the model is told. Reference values are shown for 0001 and 0011 only; the other four stay withheld.
Task 0001 Panel OLS
Consumption smoothing in interwar Japan
How much does a working-class household's consumption move when its disposable income changes? A monthly household panel with household and month-year fixed effects, composition controls interacted with quarter dummies, and household-clustered standard errors.
Specification
- Outcome (y)
lnexpenditure- Treatment (x)
lndispinc- Controls
c.size#i.quarter, c.fam69r#i.quarter, c.fam1012r#i.quarter, c.fam1316r#i.quarter, c.fam17mr#i.quarter, c.fam17fr#i.quarter, i.tfe- Requirements
- Estimate a panel OLS regression of log consumption on log disposable income with household fixed effects, month-year fixed effects, household-composition controls interacted with quarter dummies, and standard errors clustered by household.
Data and provenance
- Data source
0001/data/Data_reg.dta- Tags
- data processing
- cluster standard error
- fixed effect
- Journal
- Journal of Economic History
- Article
- consumption smoothing in the working class households of interwar japan
- Notation
- Table 3,Panel A,Row Total consumption
Exact model-visible instruction
Please use the panel OLS method to compute the effect of lndispinc on lnexpenditure. You also need to control the following control variables: c.size#i.quarter, c.fam69r#i.quarter, c.fam1012r#i.quarter, c.fam1316r#i.quarter, c.fam17mr#i.quarter, c.fam17fr#i.quarter, i.tfe. Besides, you need to consider the following requirements: Estimate a panel OLS regression of log consumption on log disposable income with household fixed effects, month-year fixed effects, household-composition controls interacted with quarter dummies, and standard errors clustered by household.. You could load the corresponding data from /data/0001/data/Data_reg.dta. At the end of the program, please print the coefficient, standard error, p value of the effect in a json format like {"coefficient": 0.1, "standard_error": 0.1, "p_value": 0.1}, and output the json string as json file to /output/1_result.json
- coefficient
- 0.391865
- standard_error
- 0.037674
- p_value
- 4.20e-21
Rounded to six decimals. Shown because the estimate is printed in the source paper; during evaluation it never leaves the scorer.
Task 0011 OLS Real run
Economic uncertainty and divisive politics (the “two Spains”)
How does socio-economic conflict relate to economic policy uncertainty, 1905–1945? OLS with month-clustered standard errors.
Specification
- Outcome (y)
EPU0month_simsn_w- Treatment (x)
Wscmonth_simsn_w- Controls
Wnamonth_simsn_w, Wmimonth_simsn_w, Wremonth_simsn_w- Requirements
- Replicate PDF Table 2 column 1: OLS of EPU on the four political division variables for 1905-1945, clustering standard errors by month.
Data and provenance
- Data source
0011/data/data_np.dta- Data shape
- 4,550 observations × 38 columns in the recorded run. Restricting to 1905–1945 gives 984 observations; excluding 7 with missing regression variables leaves 977.
- Tags
- cluster standard error
- Journal
- Journal of Economic History
- Article
- economic uncertainty and divisive politics evidence from the dos espanas
- Notation
- Table 2, Column 1, Row Socioec. conflict
Exact model-visible instruction
Please use the OLS method to compute the effect of Wscmonth_simsn_w on EPU0month_simsn_w. You also need to control the following control variables: Wnamonth_simsn_w, Wmimonth_simsn_w, Wremonth_simsn_w. Besides, you need to consider the following requirements: Replicate PDF Table 2 column 1: OLS of EPU on the four political division variables for 1905-1945, clustering standard errors by month.. You could load the corresponding data from /data/0011/data/data_np.dta. At the end of the program, please print the coefficient, standard error, p value of the effect in a json format like {"coefficient": 0.1, "standard_error": 0.1, "p_value": 0.1}, and output the json string as json file to /output/11_result.json
- coefficient
- 0.272736
- standard_error
- 0.054225
- p_value
- 6.90e-07
Rounded to six decimals. Shown because this task is the one walked through on the Agent page; during evaluation it never leaves the scorer.
Task 0104 DID
Deterrence and compellence in parliament
Did lifting an MP's immunity change how often they submitted formal queries to the government? A difference-in-differences with MP and month-by-year fixed effects, clustered by MP.
Specification
- Outcome (y)
Formal Queries (questions)- Treatment (x)
Immunity Lifted x Post- Controls
MP fixed effects; month-by-year fixed effects- Requirements
- Replicate Table 3 for all MPs. The outcome is the number of times an MP submitted a formal query to members of the government, stored as questions in the data. Estimate Immunity Lifted x Post with MP fixed effects and month-by-year fixed effects, clustering standard errors at the MP level.
Data and provenance
- Data source
data/data_july_2021.dta- Tags
- difference-in-differences; fixed effects; clustered standard errors
- Journal
- The Journal of Law and Economics
- Article
- Deterrence and Compellence in the Parliament
- Notation
- Table 3, Panel B, Column 4, Formal Queries
Exact model-visible instruction
Please use the DID method to compute the effect of Immunity Lifted x Post on Formal Queries (questions). You also need to control the following control variables: MP fixed effects; month-by-year fixed effects. Besides, you need to consider the following requirements: Replicate Table 3 for all MPs. The outcome is the number of times an MP submitted a formal query to members of the government, stored as questions in the data. Estimate Immunity Lifted x Post with MP fixed effects and month-by-year fixed effects, clustering standard errors at the MP level.. You could load the corresponding data from /data/data/data_july_2021.dta. At the end of the program, please print the coefficient, standard error, p value of the effect in a json format like {"coefficient": 0.1, "standard_error": 0.1, "p_value": 0.1}, and output the json string as json file to /output/104_result.json
Reference values for this task stay on the scorer side; see the notation pointer for the published table.
Task 0100 IV-2SLS (fuzzy RDD)
Patent validity and litigation
Does filing an inter partes review make a case more likely to settle? A fuzzy regression discontinuity around the one-year filing deadline, estimated as an IV/2SLS second stage with a triangular kernel and case-clustered errors.
Specification
- Outcome (y)
settle- Treatment (x)
ipr_file- Controls
time2, inter, stay. Mtd, msj, software, lnfamily_size, lnforwardcites_3yr, lnbackwardcites, lnnplcites, sep, npe, lnnb_applicants, lnnb_inventors, lnpat_clm_ct, miss_claim, Instruments, Electrical_eng, Chemistry, Mechanical_eng- Requirements
- Replicate Table 2, the [-120, 120] day second-stage fuzzy regression discontinuity estimate. Instrument IPR filing with the discontinuous jump in filing probability in the 10 days before the one-year deadline, use time2 and its interaction with the jump indicator as running-variable controls, apply the triangular kernel weight with bandwidth 120, include the listed patent and case controls, and cluster standard errors by case id.
Data and provenance
- Data source
data/filing.dta- Tags
- IV-2SLS
- regression discontinuity
- cluster standard error
- fixed effect
- Journal
- The Journal of Law and Economics
- Article
- Patent Validity and Litigation Evidence from US Inter Partes Review
- Notation
- Table 2, [-120, 120] days, Second Stage (Petition Filed)
Exact model-visible instruction
Please use the IV-2SLS method to compute the effect of ipr_file on settle. You also need to control the following control variables: time2, inter, stay. Mtd, msj, software, lnfamily_size, lnforwardcites_3yr, lnbackwardcites, lnnplcites, sep, npe, lnnb_applicants, lnnb_inventors, lnpat_clm_ct, miss_claim, Instruments, Electrical_eng, Chemistry, Mechanical_eng. Besides, you need to consider the following requirements: Replicate Table 2, the [-120, 120] day second-stage fuzzy regression discontinuity estimate. Instrument IPR filing with the discontinuous jump in filing probability in the 10 days before the one-year deadline, use time2 and its interaction with the jump indicator as running-variable controls, apply the triangular kernel weight with bandwidth 120, include the listed patent and case controls, and cluster standard errors by case id.. You could load the corresponding data from /data/data/filing.dta. At the end of the program, please print the coefficient, standard error, p value of the effect in a json format like {"coefficient": 0.1, "standard_error": 0.1, "p_value": 0.1}, and output the json string as json file to /output/100_result.json
Reference values for this task stay on the scorer side; see the notation pointer for the published table.
Task 0145 RDD
Liquidity effects of litigation risk
How did a legal shock change firms' cash holdings? A local-linear regression discontinuity at the cutoff with a triangular kernel, clustered by company.
Specification
- Outcome (y)
delta_lcash_scaled- Treatment (x)
Trdd- Controls
- None declared — the instruction says “There is no control group.”
- Requirements
- source-equivalent reconstructed RDD from Oliviero_Park_Zou_JLE_file3.do: local linear regression at cutoff 0 with triangular kernel and row-specific bandwidth; cluster by comp_id.
Data and provenance
- Data source
0145/data/Regression3.dta- Tags
- None
- Journal
- The Journal of Law and Economics
- Article
- Liquidity Effects of Litigation Risk Evidence from a Legal Shock
- Notation
- Table 5, Row Baseline Row
Exact model-visible instruction
Please use the RDD method to compute the effect of Trdd on delta_lcash_scaled. There is no control group. Besides, you need to consider the following requirements: source-equivalent reconstructed RDD from Oliviero_Park_Zou_JLE_file3.do: local linear regression at cutoff 0 with triangular kernel and row-specific bandwidth; cluster by comp_id.. You could load the corresponding data from /data/0145/data/Regression3.dta. At the end of the program, please print the coefficient, standard error, p value of the effect in a json format like {"coefficient": 0.1, "standard_error": 0.1, "p_value": 0.1}, and output the json string as json file to /output/145_result.json
Reference values for this task stay on the scorer side; see the notation pointer for the published table.
Task 0194 Probit
Incentives in non-routine analytical team tasks
Does a bonus incentive change behaviour? A probit model on the field (non-laboratory) sample.
Specification
- Outcome (y)
bonus- Treatment (x)
incentive45- Controls
- None declared — the instruction says “There is no control group.”
- Requirements
- This regression only retains non-laboratory data.
Data and provenance
- Data source
0194/data/table.dta- Tags
- cluster standard error
- Journal
- Journal of Political Economy
- Article
- The Effect of Incentives in Nonroutine Analytical Team Tasks
- Notation
- Table 2, Column 1
Exact model-visible instruction
Please use the PROBIT method to compute the effect of incentive45 on bonus. There is no control group. Besides, you need to consider the following requirements: This regression only retains non-laboratory data.. You could load the corresponding data from /data/0194/data/table.dta. At the end of the program, please print the coefficient, standard error, p value of the effect in a json format like {"coefficient": 0.1, "standard_error": 0.1, "p_value": 0.1}, and output the json string as json file to /output/194_result.json
Reference values for this task stay on the scorer side; see the notation pointer for the published table.
08 / Provenance
Where every number comes from
Except where another source is explicitly noted, counts on this page are computed directly from one pinned CSV in one pinned dataset revision. Anyone with the same revision can recompute the CSV-based counts; nothing here depends on a model run.
- Dataset repository
- CamoAiLab/InferenceNet on Hugging Face
- Pinned revision
59f9512a38e594528807744214a60ee00367434e- Task list
Selected_1000/1000_new.csv· 1,000 rows- Task list sha256
5338c8af548b35dab0e92a63b0acfebb0bc29c7901da4088851025f5709d6bb3- Counts computed
- 15 Sep 2026, over all 1,000 rows of the pinned list
- Evaluator
- YTZSR/chatpilot_evaluator on GitHub — task loading, instruction rendering, execution and scoring
How the counts were computed
Journal, method, control, tag and file-format distributions are column counts over the 1,000 rows of the pinned task list; the five method classes fold the 14 raw labels as listed in section 03. Reference-value bands and signs are read from the answer field where it is numeric. The field distribution in section 04 is the one exception: it is transcribed from the publisher's figure, not recomputed.
| Schema | Tasks |
|---|---|
| Standard triplet (coefficient, standard_error, p_value) | 998 |
| Composite (multi-regression) | 1 |
| Triplet + n | 1 |
| The output contract in section 06 is the standard triplet. | |
| Reference field | Supported | Of |
|---|---|---|
| Coefficient and direction | 999 | 1,000 |
| Significance band | 997 | 1,000 |
Ids listed as unsupported: 0346, 0915, 0917, 1050, 1109. Coverage per the evaluator's leaderboard documentation. | ||
Selected_1000 keeps the original task ids, which run from 0001 to 1125 with gaps. The snapshot behind these counts holds 5,207 files and 54.28 GB of data.