R2H
insights

What the Newest AI Tools for Science Agree On

By Ravza Nazli Muyesseroglu, MD, MSc Scientific Lead & Co-Founder, R2H·July 31, 2026

Three new tools for science, three different problems, one shared idea underneath. A look at what Proto, Biomni and Claude Science quietly agree on.

What the Newest AI Tools for Science Agree On

Something worth noticing is happening in AI for science. Over the last few weeks, a handful of new tools have arrived: Proto from Arc Institute, Biomni out of Stanford (now a company called Phylo), and Anthropic's Claude Science, which scientists have just started putting to the test. They do quite different things. But underneath, they're built on the same quiet idea.

What Each One Actually Does

Start with Proto. There are hundreds of powerful AI tools in biology now, for proteins, for DNA, for RNA, but they mostly can't talk to each other. Each was built by a different group, with its own setup, so before you can even start, you lose days just getting two of them to run side by side. Proto is a shared language that lets these tools work together. And it pays off: on a design task where older methods screened thousands of candidates to find one that worked, Proto got there testing only tens.

Where Proto connects the tools to each other, Biomni puts them in the hands of the researcher. It's an AI lab partner you talk to in plain language. You hand it a task, like analyse this dataset and suggest a hypothesis, and it decides which tools it needs, runs them, and comes back with a result. Underneath sits a language model connected to over 150 specialised tools and dozens of scientific databases. The model itself isn't the point. What matters is that it knows which tool to reach for and when.

Claude Science is Anthropic's version of the same thing, aimed at scientists. It's one workspace that connects the tools and databases a researcher already relies on, all behind a single assistant you talk to in plain language. Scientists have just started testing it in their own work. Their verdicts are revealing, and we'll get to them.

It's Not the Model. It's the Wiring.

Notice what none of these is built around. Not one of them is a single standout model meant to outperform everything else. Proto connects existing AI tools through a shared language. Biomni runs a set of existing tools on the researcher's behalf. Claude Science brings existing tools and databases together in one place. In each case, the interesting part isn't the model. It's the coordination.

That's the thread running through all three. The heart of each one isn't a model built to be smarter than the rest; it's the layer that gets different tools and models to work together. For a while, the assumption was that progress meant a bigger, more capable model. These take a different route, and it's turning out to be a hard and useful problem in its own right: not building the cleverest model, but making the ones we have work as a system.

Where the Human Comes Back In

But coordination only takes you so far, because none of these tools is finished, and their weak spots are as revealing as their strengths. Biomni is strong at things like database queries and sequence analysis, yet its own creators note it still struggles with the nuanced clinical judgment and deep reasoning that come with experience. The early verdicts on Claude Science echo that: fast and unusually transparent, but shaky on reasoning, unable to reach paywalled papers, and held back by safety guardrails that sometimes block ordinary research. Even Proto's team says plainly that it works best paired with lab testing and a researcher's own interpretation.

Line those up and the same gap shows through all three: what they still hand back to a human is judgment. One scientist testing Claude Science put it better than any spec sheet could. The hard part, he said, is no longer predicting the next word. It's doubting it. And doubt is close to the centre of how science actually works, knowing when an answer that looks right deserves a second look. That's the one thing these systems can't yet do on their own.

The Direction Is Becoming Clear

Last week's argument here was that the useful question isn't how smart an AI is, but what it's connected to and whether a researcher can check it. This week, the same idea turned up on its own, in three separate places at once. An argument is always more convincing when it arrives from somewhere other than the people making it.

It's also, in practice, the thinking behind Clarisyn. There's no model trained from scratch, because the advantage was never going to be there. The design takes strong existing tools and connects them so their outputs can be cross-checked rather than taken on faith, with the scientist kept in the loop throughout: able to see how an answer was reached, and to bring in real expertise when a question calls for it. That last part matters most. The tools in this piece are all moving the same way, and the useful response isn't to claim you saw it coming. It's to keep integrating what works, and to stay honest about where a scientist is still needed.

Accelerate Your Translational Research

Experience how Clarisyn transforms complex scientific problems into structured, evidence-based reports.

Get Started with Clarisyn