This is a Preprint and has not been peer reviewed. This is version 1 of this Preprint.
Assessing veracity of web-based citizen science projects: An audit study of eBird
Downloads
Authors
Abstract
1. Uncertainty in data accuracy and lack of diversity among participants are two common concerns in citizen science projects. But what if the methods used to simultaneously maximize data quality and user participation lead to unbalanced user experiences and biased data? While many studies focus on advancing techniques for analyzing large and noisy datasets, methods are needed to test the veracity of the systems that produce those data.
2. Here we present a web-based audit study, a standard and increasingly common method in computer and social sciences but novel to ecology, to test the veracity of methods used in web-based citizen science data collection and curation. We tested for several types of bias in eBird’s data quality control process, with the intent to provide recommendations to improve representation and scientific veracity in the world’s largest citizen science project. Following a basic design for auditing web-based platforms, we systematically submitted 19,491 fictitious bird observations from 196 simulated eBirders who varied by region, sex, and race, and we assessed how those variables were associated with eBird’s real acceptance or rejection of the fictitious observations.
3. We found near-total confirmation bias, in that eBird accepts 100% of false observations that fit eBird’s a priori expectations for species’ spatio-temporal ranges, while accepting 0 – 29% of true observations that do not fit eBird’s expectations. Hence, the eBird data quality control process does not validate data, but rather it produces illusory quality through forced conformity. We show how this problem contributes to a larger, self-perpetuating cycle of bias inherent to the design and function of eBird. We cannot report results from geographic or demographic analyses (see SI).
4. Regarding species’ spatio-temporal ranges, and changes thereto, we conclude that the eBird database almost entirely reflects eBird’s predetermined expectations, and changes that eBird makes to those expectations over time. This irreparably detaches the eBird dataset from ecological reality because no statistical technique can reconstitute patterns that are perpetually and permanently removed and replaced during data collection and curation. eBird is an undeniably magnificent, user-friendly birdwatching application, but as citizen science, its design obviates the citizen and feigns science. The ecological and evolutionary sciences could benefit from similar independent investigations of other web-based research projects.
DOI
https://doi.org/10.32942/X2W101
Subjects
Research Methods in Life Sciences
Keywords
algorithm auditing, big data, confirmation bias, discrimination, ornithology
Dates
Published: 2026-08-05 02:28
Last Updated: 2026-08-05 02:28
License
CC-BY Attribution-No Derivatives 4.0 International
Additional Metadata
Conflict of interest statement:
None
Data and Code Availability Statement:
Not applicable
Language:
English
Metrics
Views: 3297
Downloads: 15
Comment #325 Henry Michael Streby @ 2026-08-07 13:39
The shooting of the messenger and misinterpretation of the intent and conclusions is not unexpected from those deeply tied to eBird. But thank you again to those who continue to reach out with constructive comments and encouragement, including those currently working for eBird. Alex, the "hard to believe" part will be told in its entirety, with all the receipts.
Comment #324 Alexander Lin-Moore @ 2026-08-07 13:14
The author(s) of this preprint aim to explore the trustworthiness and implementation of data quality in eBird, a worldwide, public data collection platform for bird observation. Such an investigation is to the benefit of the ecological sciences, and addresses common concerns of community science projects and their applicability to academic research. Unfortunately such a goal is not thoroughly investigated in this report. What has instead occurred appears to be the submission of a single packet of test data to uncover the operation of eBird as a platform. In most cases, when the platform’s data quality controls were triggered, such faked data was caught and excluded from output; in other words, a demonstration that the system works as intended and as advertised. A more straightforward approach to this conclusion may simply have been to ask the administrators of eBird, or even current local reviewers, about the platform's operation. The decision to rely on probing a community science platform with fake data in order to understand how it works is also reflected in the introduction, which infers the process of eBird review and record-keeping (occasionally incorrectly), when a simple request for information would have definitively and accurately outlined the process. As it is, this investigation does not probe an unknown mechanism or underlying bias, but rather attempts to reconstruct the widely available review process of a single, deliberately-designed community science platform.
The involvement of sociological parameters into this study is admirable in concept, and the closest match to the author’s stated goal of performing a web-based audit study. Regrettably this attempt was faulty in execution, to the extent that it had to be excluded entirely from the report - though such exclusion does not preclude its description as central to the project’s aims. Flaws in this approach are clear from this report's own description of experimental design. Furthermore the claim that such work would not require IRB input is extraordinary given that the central conceit of the study lies in the intentional deception of volunteer eBird reviewers with regard to not only submitted data but the demographics of the users providing said data. Input from sociologists familiar with such a data collection approach would have benefited this attempt, and indeed the successful application of such a demographic approach would make this study more valuable sociologically, not ecologically. As it is, mention of this attempted demographic study exists in the preprint primarily as an explanation for the author’s stated long-term persecution by Cornell University. Without passing judgement on the validity of these claims, it should go without saying that the explicit attempts of this preprint to publicly paint Cornell as an unduly litigious aggressor, and the author as a victim of conspiracy on the part of the Cornell Lab of Ornithology, are a disservice to what work has been done here, and degrade the image of community science platforms and the credibility of the ornithology field at large.
The conclusions drawn from the successfully-collected data are overreaching in the extreme. Put bluntly, the only data presented in this study to inform the 11 pages of discussion are four numbers: the total number of fake entries of expected/unexpected species submitted to eBird, and the proportion of these fake entries rejected by the platform. All conclusions related to eBird's self-perpetuating bias, reporting of spatiotemporal ranges, bot farming, fraudulent use of the platform (this study notwithstanding), or relationship to "true" ecological conditions are pure conjecture, extrapolated from these four numbers, seemingly based in a fundamental and long-standing disagreement the platform's approach to data collection and storage. No investigations are carried out in support of any claims made, which is doubly disappointing given the clear opportunities to compare eBird data to other data sources in the context of the issues raised. For example, that the eBird data filtering system could obfuscate or delay the detection of range expansion is an interesting idea, and ample modern examples of range expansion in birds would seemingly provide an ideal opportunity to compare eBird to non-eBird data. Such a comparison is ignored in favour of an entirely fabricated and hypothetical diagram outlining perceived data quality catastrophe theorised by the authors (Fig. 2). What arises in the discussion is not an interpretation of the data collected, but a manifesto of eBird’s sins and hubris regarding their own data collection and record-keeping, and the hyperbolic disaster that apparently inevitably awaits such practice.
I do not know Dr. Streby, but a quick search through his work and familiarity with some of his co-authors seems to demonstrate an admirable publication record and a solid basis of rigorous science. This preprint however is anything but: a conspiratorial screed against a public data collection platform masquerading as an audit study, belying a long-term grudge against the Cornell Lab of Ornithology, one which is all-but stated on the preprint’s front page to be based in pushback to this very work. This insult-laden opening statement explains that the sole named author has suffered "enduring professional and personal harm" by Cornell University for attempting this work, but it is difficult to imagine a single piece of material more discrediting to one’s reputation as a serious scientist than this manuscript. I struggle to describe this text using any term other than ‘hysterical’, and am stunned that cooler heads did not prevail after it was typed out. As it exists, this preprint is in no way suitable or appropriate for submission for peer review.
Comment #322 Henry Michael Streby @ 2026-08-06 20:48
Thank you to all the ornithologists who've reached out directly to express support and encouragement. Special thanks to the eBird reviewers who have offered suggestions to clarify and emphasize certain points. This is what pre-prints are for.
Comment #321 Henry Michael Streby @ 2026-08-06 20:47
Thank you to all the ornithologists who've reached out directly to express support and encouragement. Special thanks to the eBird reviewers who have offered suggestions to clarify and emphasize certain points. This is what pre-prints are for.
Comment #320 John Smith @ 2026-08-06 17:42
The lack of effort and knowledge that went into this, coupled with the many assumptions made, are truly impressive. I can see why most of the authors would choose to remain anonymous. Well done!