Using AI for biodiversity + climate change
Background
I first started thinking about using AI for biodiversity + climate change over a year ago.
The first goal was to see if AI could properly transcribe natural history specimen labels, a super thorny problem that is endless, expensive, and important for biodiversity research.
This exploration eventually turned into a homegrown CRUD web app for specimen transcription. Without using any of today’s agentic stack, it worked brilliantly as a homegrown harness consisting of a three-shot prompt, a hardcoded tool call, and a homemade “eval”.

Does this say Mießler, Zießler, or Kießler? A quick human Google search (or AI tool call) easily shows that it’s Oscar Mießler, a German botanist.
In the past year, we’ve come a long way from this homegrown web app. Today, with nothing but a specimen sheet and the directive to “please transcribe”, models have basically saturated the field. This makes sense as specimen digitization is typically a job done by undergradute-level researchers or volunteers.
My concern is that as models become better than our best PhDs in subjects like math, we don’t have enough investment in the natural sciences. If we want to tackle the hardest open problems facing humanity around climate change and biodiversity loss, models need to get much, much better at this kind of work.
And we don’t get that for free. Humans need to make sure that happens.
Next steps for biodiversity + AI
While I can’t comment on whether or not we should be training smarter AI models (although I do think we should consider a coordinated slowdown), it’s obvious that we are. Given that, I believe that we have an opportunity to use this technology to attempt to solve our most difficult and urgent human problems.
In other words, if we want models to help us with climate + biodiversity loss, we should be actively working towards that future now.
Here are a few ideas based on what I’ve learned.
1) Build a really good biodiversity + climate benchmark and publish a leaderboard for frontier models
Benchmarks are tests to see how good models are at something. However, biology evals are heavily skewed towards bio risk today.
While alignment is really important, I think there’s space in this field for people to work on multiple classes of human existential risk. Labs are working hard to make AI great at software engineering, lawyer-ing, banking and other money-making fields alongside safety – why not natural sciences? Why not also build evals that measure how good models are at identifying a tree’s species and approximate age based on a photo and a location? Modeling weather patterns? Measuring air quality?
If we don’t have good evals, we can’t measure whether new models are actually getting better at these skills. Publicly ranking progress against this benchmark also has a flywheel effect as leaderboards provide strong incentives to climb the ranks.
2) Leverage high-quality climate and biodiversity data in training
What if we trained on the millions of natural history specimens that we’ve already digitized? Climate data? Going even further, could we use this as a way to unlock the world’s undigitized collections?
That is, there’s a treasure trove of previously unseen, high-quality biodiversity data locked in the basements in academic institutions everywhere. Could these institutions negotiate access to these collections in exchange for 1) money and 2) making the data accessible to the broader scientific community?
Could this be a way to simultaneously fund chronically underfunded natural history museums, increase the amount of data available for research, and make models great at doing this work at the same time?
3) Increase scientific adoption
AI is already smart enough to greatly accelerate climate + biodiversity work. But the field doesn’t seem to be broadly leveraging these tools.
This is where my skill set is useful, so my plan is to continue empowering and resourcing researchers that are already using AI or are AI-curious.
Currently I’m working on:
- Maintaining the Herbarium AI Leaderboard so people know which models work best for transcription.
- Making sure that natural science researchers have access to the tokens they need to do their work.
- Creating a set of super easy, skill-backed directions for how to digitize herbarium specimens with normal consumer AI apps.
- Sidebar: Do you know about best practices for digitizing marine, fossil, insect, plant, bryophyte, or other specimen types covered by Darwin Core? I could use some help. Let me know!!!
Conclusion
The vast majority of us care about life on Earth, but it’s not always clear how to contribute. Software engineers aren’t familiar with the science, scientists don’t know about AI, and AI safety researchers are too worried about loss of control at the moment to think about much else.
But I think we have to find a way to start doing something, somehow, anyway. In personal terms, I’m just a software engineer. I don’t know how to make frontier evals and I have no direct experience in biodiversity + climate science. So I need your help!
If you’d like to get in touch, email me at [email protected].



Reward hacking rate by intervention type, n=100 (lower is better)