From the research bench

What is reproducibility? Can another analyst reconstruct the result?

Reproducibility asks whether another analyst can rebuild a result from its inputs, code, environment, and procedure. It is not replication.

1,075 words 5 min read 2 references  readers

Management summary

Reproducibility asks whether another analyst can reconstruct a result from the specified question, input versions, procedures, code, computational environment, and expected outputs. It is different from replication with new data and different again from external validity. Banzi et al. provide a current consensus view of core reproducibility items across planning, methods, data collection, management and analysis, and dissemination. The Open Science Collaboration shows why a replication needs a declared success criterion rather than one universal percentage. This article builds a synthetic reconstruction worksheet with rerun, mismatch, unavailable, and confirmed states. The worksheet and commercial translation are author synthesis. A successful rerun does not establish substantive truth, causal identification, measurement validity, or transfer to a new target.

Keywords: Reproducibility · Reproducible Research · Replication · Computational Environment · Audit Trail · Research Design

On this page

A result can be clear enough to read and still impossible to rebuild. The data file is missing. A manual exclusion was made in a spreadsheet. The analysis used a library version nobody recorded. The reported chart has no rule for which rows entered it.

Reproducibility asks whether another analyst can reconstruct a result from the materials and procedure specified for it. The object is the path from question to output, not the confidence or importance of the conclusion.

The evidence-review stopping-rule article asks when a claim has earned decision weight. The measurement-invariance article asks whether a result can travel to a named target. This page asks the earlier question: can another analyst rerun the result in the first place?

What does reproducibility mean?

The practical contract names:

ComponentRequired questionExample failure
Question or estimandWhat was the analysis trying to calculate or compare?The chart has no target quantity
InputsWhich data files, versions, filters, and access conditions were used?A current table replaces the historical snapshot
ProcedureWhich transformations, exclusions, and decisions occurred in what order?A manual step is remembered but not recorded
CodeWhich analysis script and parameters produced the output?A notebook cell or formula is missing
EnvironmentWhich runtime, libraries, configurations, and random state matter?A package update changes the result
OutputWhat exact table, estimate, chart, or file should appear?“Same result” has no tolerance or comparison rule

Table 1What does reproducibility mean?

Source: Table from this essay. Sources and interpretation are given in the article.

View exhibit page

Banzi et al. (2026) present a consensus list of core reproducibility items across planning, methods, data collection, management and analysis, and dissemination. The source is a research-consensus view, not a validated business scorecard. Its value here is the breadth of the reconstruction contract: reproducibility begins before the final calculation.

Is reproducibility the same as replication?

No. Keep the tests separate:

TestData and procedureQuestion answeredWhat it does not answer
ReproducibilitySame declared materials and procedureCan the result be reconstructed?Whether the result is true or general
ReplicationNew observation or data, with an intendedly similar designDoes a result recur under a new observation?Whether the original workflow was fully reconstructable
External validitySource result and a named targetDoes the conclusion transfer?Whether the source analysis can be rerun

Table 2Is reproducibility the same as replication?

Source: Table from this essay. Sources and interpretation are given in the article.

View exhibit page

The Open Science Collaboration re-ran 100 psychology studies and reported more than one replication success measure. The lesson for a commercial review is not a percentage to import. It is that the question “did it replicate?” needs a declared criterion, population, outcome, and comparison.

What does a reconstruction worksheet look like?

The six rows below are synthetic. They illustrate readiness states, not a reproducibility rate for a company, discipline, or dataset.

IDQuestion or outputInputs and data packageCode and environmentExpected outputRerun state and disposition
R-01Cohort rate at day 14Snapshot V3 and eligibility rule heldScript S-12; runtime recorded50.0%Confirmed within stated tolerance
R-02Revenue bridge by lineInvoice file held; product map version unclearScript S-18; package lock missingFive-effect bridgeMismatch risk; hold release
R-03Survey mean by localeData package and codebook heldAnalysis script held; exclusion manualTable of means and countsRerun pending exclusion log
R-04Forecast error by horizonBaseline and actual files heldNotebook and library versions heldError table and plotConfirmed; no claim of causality
R-05Customer segment modelAccess restricted; derived file onlyScript available; random seed absentModel metricsUnavailable input; cannot reconstruct
R-06Treatment contrastPreregistered dataset and treatment rule heldScript and environment heldEffect estimate and intervalReconstructed; replication still open

Figure 1The synthetic reconstruction worksheet

The rows are illustrative. Confirmed reconstruction, replication, and causal validity remain separate statuses.

Source: Author's synthetic worksheet grounded in Banzi et al. (2026) and the Open Science Collaboration (2015); all rows and states are illustrative.

View exhibit page

R-02 is not “almost reproducible” because its final chart looks plausible. A missing product-map version can change the line population. R-05 is not solved by publishing model metrics when the input cannot be reconstructed. R-06 is reproducible as a workflow, while the replication question remains open because it needs new observation.

Which decisions belong in the workflow record?

Keep these decisions visible rather than burying them in prose:

  • the unit, target quantity, and analysis window;
  • the source file, snapshot, version, and access condition;
  • every filter, exclusion, join, transformation, and manual edit;
  • the analysis script, parameter file, runtime, library versions, and random seed where relevant;
  • the expected output, comparison tolerance, and known nondeterminism;
  • the person or process that approved a change and the date it took effect;
  • the result of the rerun, including mismatch and unavailable states.

Preregistration can record selected questions or analysis decisions before observation. It does not make an analysis immune to change. A changed decision can be legitimate when it is dated, explained, and distinguished from the original plan.

What can a successful rerun prove?

It can show that the specified path produces the specified output under the specified environment. It cannot, by itself, prove that the input is accurate, the construct is valid, the model is correctly specified, the comparison is causal, or the result transfers to a new target.

That boundary is why a reconstruction worksheet should contain both a status and a disposition. A confirmed rerun can release the method for review. A mismatch can open a debugging question. An unavailable input can block a stronger claim. None of those statuses is a substantive outcome.

How should a team test reproducibility?

  1. Freeze the question, unit, outcome, and observation window.
  2. Assemble the exact input versions and data package.
  3. Replay every transformation, exclusion, and manual step.
  4. Run the declared script in the declared environment.
  5. Compare the output to the stated result with a declared tolerance or exact-match rule.
  6. Record mismatch, unavailable input, environment drift, or successful reconstruction.
  7. Treat a new-data test as replication and a target change as external-validity work.

Reproducibility is evidence about the reconstruction path. It is not a universal score, a guarantee of truth, or a substitute for a new study.

References

  1. Banzi, R., Varga, M., Gelsleichter, Y. A., Vinatier, C., Moher, D., & Naudet, F. (2026). An international consensus on core reproducibility items in research. PLoS Biology, 24(4), e3003726. https://doi.org/10.1371/journal.pbio.3003726
  2. Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. https://doi.org/10.1126/science.aac4716

Pass it on

Share this essay

If it was useful to you, it is probably useful to someone on your team.

Download as PDF

A complete document: title page, contents, sources, and the citation on the last page.

Sinan Isoglu

About the author

Sinan Isoglu, MBA (Quantic)

Commercial growth leader, lecturer and doctoral researcher

Sinan Isoglu is a commercial growth leader, lecturer and doctoral researcher. His work spans go-to-market, pricing and revenue operations; his doctoral research at EM Normandie examines sales and marketing integration after cross-border M&A. He lectures on marketing and growth at IU International University of Applied Sciences.

Credentials

  • Doctoral researcher, EM Normandie Business School
  • MBA, Quantic School of Business and Technology
  • Lecturer, IU International University of Applied Sciences

Writes on

  • Go-to-market
  • Pricing
  • Revenue operations
  • AI in commerce
  • Cross-border growth

The track

The test behind this question.

This piece sits in the research track: the stricter standard applied to the patterns practice produces.

Comments

Join the thinking.

Comment on the piece, or select a passage above to quote it directly.

Leave a comment

Comments are read and approved personally before they appear. Your name and comment are stored for publication. See the Privacy note.