Skip to main content
When and how to integrate probability and nonprobability samples to monitor biodiversity

When and how to integrate probability and nonprobability samples to monitor biodiversity

This is a Preprint and has not been peer reviewed. This is version 2 of this Preprint.

Add a Comment

You must log in to post a comment.


Comments

There are no comments or no comments have been made public for this article.

Downloads

Download Preprint

Authors

Rob James Boyd, Oliver L. Pescott , Simon Rolph, Diana E. Bowler, Nick J. B. Isaac, Francesca Mancini, David Roy, Gesa von Hirschheydt, Michael Pocock

Abstract

1. Species occurrence and abundance data are generated by probability or nonprobability sampling. Under probability sampling, sites are selected according to a known design (e.g. stratified random sampling). Under nonprobability sampling—for example, the collection of data by citizen scientists at sites of their own choosing—they are not. Whereas probability samples tend to be relatively small but representative of the wider landscape, nonprobability samples are often much larger but less representative. 2. We evaluate when and how probability and nonprobability samples can be integrated to estimate temporal change in species occupancy. Using computer simulation, we compare four estimators that are applied to a small probability sample, a larger nonprobability sample, or both. 3. Scenarios vary in terms of baseline occupancy and how it changes over time, the size of the nonprobability sample, and the dependence of site inclusion in the nonprobability sample on occupancy. To isolate the effects of sampling design, we assume occupancy is measured without error at sampled sites. 4. Estimators based solely on nonprobability samples were generally biased and did not achieve nominal confidence interval coverage. Model-based integration improved precision but suffered from bias unless the dependence of site inclusion in the nonprobability sample on occupancy was stable over time—a condition that is difficult to verify in practice. 5. A design-based integrated estimator, which has not previously been considered in ecology, consistently performed well. It was approximately unbiased and achieved nominal confidence interval coverage in all scenarios, and it had low mean squared error in most scenarios. 6. The desirable statistical properties of the design-based integrated estimator arise from the sampling design rather than the ecological process and are therefore expected to apply regardless of study variable (e.g. occupancy, abundance, species richness). Extensions that account for measurement error are discussed but not tested. 7. Design-based integration provides a more reliable basis for combining probability and nonprobability samples than the model-based estimators considered here, and we recommend it as the default approach. More elaborate model-based estimators may perform better in some settings, but their performance depends on assumptions that are difficult to verify in practice.

DOI

https://doi.org/10.32942/X2KM4M

Subjects

Life Sciences

Keywords

citizen science, integrated distribution model, joint likelihood, species abundance, data defect correlation

Dates

Published: 2026-09-22 02:22

Last Updated: 2026-09-22 02:23

Older Versions

License

CC BY Attribution 4.0 International

Additional Metadata

Conflict of interest statement:
None.

Data and Code Availability Statement:
The code needed to fully reproduce our analysis is available at https://doi.org/10.5281/zenodo.22816909.

Language:
English

Metrics

Views: 35

Downloads: 1