Skip to main content
Spatial structure before algorithmic complexity: evidence-based guidance for marine species distribution modelling where data are limited

Spatial structure before algorithmic complexity: evidence-based guidance for marine species distribution modelling where data are limited

This is a Preprint and has not been peer reviewed. This is version 1 of this Preprint.

Add a Comment

You must log in to post a comment.


Comments

There are no comments or no comments have been made public for this article.

Downloads

Download Preprint

Supplementary Files

Authors

Ebenezer Afrifa-Yamoah , Alexandre C Siquerira, Yaw Kwaafo Awuah-Mensah, Francky Fouedjio, Ute Mueller

Abstract

Aim Species distribution models increasingly underpin marine conservation decisions, but most evidence on model performance comes from data rich settings unlike the sparse, clustered and irregularly sampled records typical of marine systems. We synthesise published evidence to establish which modelling choices govern predictive reliability when data are limited and translate the result into operational guidance.

Location Global; marine and coastal systems.

Methods We screened 330 records and retained 95 studies meeting three criteria: peer reviewed marine application, explicit treatment of spatial autocorrelation, and an available software implementation. Twelve studies reporting direct method comparisons, validation optimism assessments or sample size experiments were extracted quantitatively; the remaining 83 informed a structured narrative synthesis. Performance metrics are not interconvertible across studies, so we retained within study contrasts and directional consensus.

Results Four findings recur across the corpus. Adding explicit spatial structure to a model improves prediction more consistently than changing algorithm family. The spatial coverage of sampling constrains predictive reliability more than the number of records, so clustered data generalise poorly regardless of sample size. Validation must match the inference target: random cross validation is defensible for prediction within sampled conditions but inflates apparent accuracy for prediction beyond them. Minimum data requirements are method specific and modest, with roughly 50 to 100 observations sufficient for tree-based classification given well dispersed sampling.

Main conclusions Predictive reliability under data limitation depends more on how space and sampling are handled than on algorithmic sophistication. We provide a decision framework linking data type, inference goal and data volume to method and validation choice, and identify metric harmonisation and standardised reporting as prerequisites for quantitative synthesis in this field.

DOI

https://doi.org/10.32942/X2JM49

Subjects

Ecology and Evolutionary Biology, Life Sciences

Keywords

conservation biogeography, data limitation, machine learning, marine conservation, model transferability, sample size, spatial autocorrelation, spatial cross validation, species distribution models

Dates

Published: 2026-09-23 13:40

Last Updated: 2026-09-23 13:40

License

CC BY Attribution 4.0 International

Additional Metadata

Language:
English

Metrics

Views: 3

Downloads: 0