
Scientists across many fields are becoming dependent on powerful data sources and analytical tools that they often cannot inspect, verify or fully understand, according to a new study, raising concerns about the future reproducibility and trustworthiness of scientific research.
For the study, published in BioScience, an international team of researchers examined how emerging technologies—including artificial intelligence, satellite imagery, digital sensors and online data platforms—are reshaping the way scientists study the natural world, while often hiding the processes behind their results.
The team calls these systems "black boxes" because the companies that build them frequently limit outside access to the algorithms, training data and internal workings involved for proprietary or commercial reasons.
The researchers identified several types of black boxes now common in ecology and conservation research. Large language models and other AI tools are increasingly used to analyze massive datasets, interpret satellite imagery and model ecosystems, yet researchers often have little or no access to the data used to train these systems or insight into how they generate specific outputs.
Similar problems extend to remote sensing products built on proprietary processing, wildlife tracking devices that release only processed locations rather than raw data, and social media platforms whose hidden algorithms and shifting policies can introduce unknown biases into biodiversity research. Even social surveys are affected, as private companies increasingly manage participant recruitment and data quality with limited transparency.
“Modern scientific tools are also becoming so technically complex that users, and in some cases even their developers, may struggle to fully scrutinize and understand how they operate,” said Karen Anderson, professor at the University of Exeter and co-author of the study.
The growing dependence on black box technologies is further strengthened by the publish-or-perish culture, as well as the need to more effectively cope with growing datasets and urgent environmental crises.
To address the problem, the authors recommend prioritizing open-source software and hardware, benchmarking proprietary tools against transparent datasets, comparing results across multiple methods, and thoroughly documenting training data, pipelines and tool limitations.
“It is also necessary to intensify efforts toward open science, including regulations that would improve researchers' access to digital platforms and their underlying data,” said study author Michael Bertram of the Swedish University of Agricultural Sciences and Stockholm University.
Data by University of Exeter