This is a Preprint and has not been peer reviewed. This is version 2 of this Preprint.
Using large language models to identify reporting challenges for Target 6 implementation: a short validation against human assessment
Downloads
Authors
Abstract
National reports and related biodiversity policy documents contain important information on implementation barriers, but extracting comparable evidence across many countries is timeconsuming. We tested whether large language models (LLMs) could provide a reliable firstpass synthesis of reported challenges relevant to Target 6 of the Kunming–Montreal Global Biodiversity Framework. Human assessor scores for 50 countries were compared with structured outputs from three LLMs: Claude, ChatGPT and Gemini. The aim was not to test exact score reproduction, but to determine whether LLMs captured the same relative patterns in reported challenge categories. Claude showed the strongest alignment with human assessment, especially at the level of challenge-category ranking. Claude-derived scores were therefore used to summarise average challenge patterns across 126 reports. The results suggest that LLM scoring is useful for identifying broad thematic patterns in reporting challenges, but should not be used for precise country-level ranking.
DOI
https://doi.org/10.32942/X2N96W
Subjects
Artificial Intelligence and Robotics, Biodiversity, Research Methods in Life Sciences
Keywords
biodiversity reporting, Target 6, invasive alien species, large language models, validation, policy synthesis, Global Biodiversity Framework
Dates
Published: 2026-08-13 23:09
License
CC BY Attribution 4.0 International
Additional Metadata
Conflict of interest statement:
None
Data and Code Availability Statement:
Data and analytical code associated with this preprint are publicly available in the GBF_IAS_Challenges GitHub repository: https://github.com/AgentschapPlantentuinMeise/GBF_IAS_Challenges. The repository contains scripts for retrieving and processing CBD Seventh National Report content, applying the Target 6 challenge-scoring framework with LLMs, standardising model outputs, comparing LLM and human assessor scores, calculating agreement statistics, and generating figures and summary tables. The underlying national reports are public documents available through the Convention on Biological Diversity reporting system.
Language:
English
Metrics
Views: 1
Downloads: 0
There are no comments or no comments have been made public for this article.