This is a Preprint and has not been peer reviewed. This is version 2 of this Preprint.
More reliable inference from open raw data: a Monte Carlo comparison of raw-data and aggregate-data meta-analysis under publication bias
Downloads
Authors
Abstract
While the benefits of open data are often discussed, they are rarely quantified. Here, we provide the first simulation-based quantification of what raw data from all conducted primary studies, including unpublished ones, can add to meta-analysis under publication bias, and we introduce a tool that helps researchers determine when this approach is most beneficial. Classical meta-analysis (CMA) relies on published results and is thus vulnerable to publication bias and p-hacking. We developed a Monte Carlo simulation framework that compares CMA with raw data meta-analysis (RDMA) across varying values of true effect sizes, heterogeneity, and bias levels. In our simulation we start from a population of conducted primary studies. Some of them get published, subject to p-hacking and publication selection. CMA then samples from these, while RDMA samples raw data from the full population of primary studies, irrespective of whether they are published or not, drawing the same number of studies as CMA. Thus, the two differ only in which studies they sample, not how many. Using ecology as a case study, we demonstrate that RDMA outperforms CMA in most scenarios. When true effects are small and bias is severe, RDMA reduces relative mean absolute error in effect size estimate by 54 to 76%. Under moderate bias, reductions reach 32 to 61%. For medium true effects, reductions are 48 to 71% and 2 to 35%, respectively. RDMA maintained reliable confidence interval coverage (93 to 95%) across all scenarios; CMA fell as low as 21%. Crucially, RDMA's errors reflect natural sampling variation by construction, while CMA's reflect systematic bias that persists regardless of the sample size; in practice, selective availability of raw data would attenuate this contrast. The benefits of RDMA reflect its access to an unbiased, non-p-hacked set of primary studies, and not the recomputation of effect sizes from the raw data of the published literature alone. We provide a decision tool for meta-analysts across disciplines to calculate RDMA's benefits. Our results provide quantitative evidence that access to raw data from all conducted studies, including those that remain unpublished, improves meta-analytic accuracy in the presence of publication bias, strengthening the case not merely for open data but for complete open data, that is, the archiving of raw data irrespective of statistical significance.
DOI
https://doi.org/10.32942/X2DM3P
Subjects
Ecology and Evolutionary Biology, Life Sciences
Keywords
meta-analysis, publication bias, raw data, simulation, effect size estimation, open science
Dates
Published: 2026-04-22 15:35
Last Updated: 2026-07-29 19:05
Older Versions
License
CC BY Attribution 4.0 International
Additional Metadata
Conflict of interest statement:
None.
Data and Code Availability Statement:
Data availability: All simulation code, parameter estimation scripts, input data, and complete results are available in the OSF repository: https://doi.org/10.17605/OSF.IO/VEBYM.
Language:
English
Metrics
Views: 330
Downloads: 94
There are no comments or no comments have been made public for this article.