Skip to main content
More reliable inference from open raw data: a Monte Carlo comparison of raw-data and aggregate-data meta-analysis under publication bias

More reliable inference from open raw data: a Monte Carlo comparison of raw-data and aggregate-data meta-analysis under publication bias

This is a Preprint and has not been peer reviewed. This is version 2 of this Preprint.

Add a Comment

You must log in to post a comment.


Comments

There are no comments or no comments have been made public for this article.

Downloads

Download Preprint

Authors

Danijela Žanko, Sunčana Geček, Azra Tafro, Livia Puljak, Antica Culina

Abstract

While the benefits of open data are often discussed, they are rarely quantified. Here, we provide the first simulation-based quantification of what raw data from all conducted primary studies, including unpublished ones, can add to meta-analysis under publication bias, and we introduce a tool that helps researchers determine when this approach is most beneficial. Classical meta-analysis (CMA) relies on published results and is thus vulnerable to publication bias and p-hacking. We developed a Monte Carlo simulation framework that compares CMA with raw data meta-analysis (RDMA) across varying values of true effect sizes, heterogeneity, and bias levels. In our simulation we start from a population of conducted primary studies. Some of them get published, subject to p-hacking and publication selection. CMA then samples from these, while RDMA samples raw data from the full population of primary studies, irrespective of whether they are published or not, drawing the same number of studies as CMA. Thus, the two differ only in which studies they sample, not how many. Using ecology as a case study, we demonstrate that RDMA outperforms CMA in most scenarios. When true effects are small and bias is severe, RDMA reduces relative mean absolute error in effect size estimate by 54 to 76%. Under moderate bias, reductions reach 32 to 61%. For medium true effects, reductions are 48 to 71% and 2 to 35%, respectively. RDMA maintained reliable confidence interval coverage (93 to 95%) across all scenarios; CMA fell as low as 21%. Crucially, RDMA's errors reflect natural sampling variation by construction, while CMA's reflect systematic bias that persists regardless of the sample size; in practice, selective availability of raw data would attenuate this contrast. The benefits of RDMA reflect its access to an unbiased, non-p-hacked set of primary studies, and not the recomputation of effect sizes from the raw data of the published literature alone. We provide a decision tool for meta-analysts across disciplines to calculate RDMA's benefits. Our results provide quantitative evidence that access to raw data from all conducted studies, including those that remain unpublished, improves meta-analytic accuracy in the presence of publication bias, strengthening the case not merely for open data but for complete open data, that is, the archiving of raw data irrespective of statistical significance.

DOI

https://doi.org/10.32942/X2DM3P

Subjects

Ecology and Evolutionary Biology, Life Sciences

Keywords

meta-analysis, publication bias, raw data, simulation, effect size estimation, open science

Dates

Published: 2026-04-22 15:35

Last Updated: 2026-07-29 19:05

Older Versions

License

CC BY Attribution 4.0 International

Additional Metadata

Conflict of interest statement:
None.

Data and Code Availability Statement:
Data availability: All simulation code, parameter estimation scripts, input data, and complete results are available in the OSF repository: https://doi.org/10.17605/OSF.IO/VEBYM.

Language:
English

Metrics

Views: 330

Downloads: 94