Skip to main content
All posts
Agents5 min read

Let an AI agent search the public DNA databases first

Roughly 950 AI agents searched public DNA databases for 21 hours and flagged an enzyme system in viruses that infect bacteria. The searching is what an AI agent that runs bioinformatics pipelines from public data does, and Elnora does it for a lab's questions.

Carmen Kivisild

Carmen Kivisild

Founder and CEO

Roughly 950 AI agents spent more than 21 hours reading a public database of DNA sequences and came back with an enzyme system that no one had described before: it sits in viruses that infect bacteria, next to a long array of repeated DNA. An AI agent that runs bioinformatics pipelines from public data is software that does the sequence-reading work a scientist would otherwise do at a screen: it takes a question, searches the record in bulk and flags what a person should examine. Anthropic, the company that built the agents, reported the find on 23 September 2026 as the first result from its new molecular biology laboratory in the San Francisco Bay Area.

What an AI agent is

The idea comes from AI research of the past few years. The language models behind modern AI, programs trained on very large bodies of text to predict the next word, learned to follow instructions written in ordinary language, and researchers then gave them tools to call: run a search, open a database, execute code. A model with tools stops answering questions and starts finishing jobs, which is the problem agents were invented to solve. A chat reply points you at the answer, and an agent goes and gets it.

Ask a phone's assistant for a pharmacy that opens on a Saturday, and it runs the search, opens the pages, reads the opening hours and comes back with the ones that fit, each step taken on its own while you wait. A bank's payment software reads every card transaction as it happens and flags the unusual few for a fraud investigator, who decides what the exceptions mean. The Anthropic search is that division of labour applied to DNA, with the agents reading the databases and the scientists deciding what was worth taking to the bench.

What the agents found

The team gave the agents a starting question: look for proteins that might work together with reverse transcriptases, the enzymes that copy RNA back into DNA. Working through more than 200,000 of those enzyme sequences, the agents gathered about 3,500 candidate arrangements of genes and narrowed those to 20 for closer study, consuming about 210 million tokens, the pieces of text language models read and write as they work. They reported what they were finding to one another and chose what to examine next on their own, and the candidate that came out of the narrowing was a reverse transcriptase gene with a long run of short repeats beside it, first in a giant virus and then in other viruses that infect bacteria. Anthropic calls the system ART, short for array-associated reverse transcriptases.

An enzyme gene beside a repeat array is the layout of a bacterial CRISPR system: repeats alternating with segments copied from viruses that attacked the cell, and, when the same virus returns, RNA transcribed from those segments guiding a cutting enzyme to the invader's DNA. For the viral repeats no cutting partner is known, and the evidence that they act as an immune memory is thin, so the find is a lead whose function remains unknown. The company says its scientists confirmed the system in the laboratory, with every physical experiment run by people at biosafety levels 1 and 2, and its chief executive said the model made the find mostly, though not entirely, on its own. The write-up is a preprint awaiting peer review, specialists in genome mining have pointed out that their field has searched sequence databases for unusual arrays for years, and a separate team had reported a somewhat similar system earlier. Where the credit sits is contested: people set the research question, curated the sequences and ran the experiments, and the software found the pattern and prioritised it.

Why the reading matters in biology

The field's tools have a history of arriving this way: CRISPR itself came out of oddities in the microbial record, repeat arrays noticed in bacterial genomes long before anyone knew what they were for, understood as an immune memory decades later and turned into genome-editing tools after that. Finding such things has depended on a person reading the record with enough patience to notice one strange entry, and the record has grown past what a person can read: the database the agents searched holds sequences for billions of proteins. The agents covered more than 200,000 enzyme sequences in 21 hours, a volume one person reading entries one at a time cannot reach, and what the flagged candidates needed was a question worth asking and a bench to test them on. The company's own life sciences lead said the rest of the work, characterising the system and understanding its function, needs experiments in the physical world.

How Elnora runs the search

A scientist planning a characterisation run around a reverse transcriptase from a phage asks Elnora, the AI agent for lab work, what the public record knows about the genes that sit beside it. Elnora resolves the accession to its authoritative entry in the public sequence archives, among them GenBank, RefSeq and SRA, writes the comparison code itself and runs it in a sandbox with the scientific computing stack already installed. It searches the literature alongside, verifies every reference it cites, and every answer ends with a source list that names the database record or paper behind each statement, so the entry behind any line can be opened directly. Where the question needs the deposited data itself, accessions from GEO and SRA run straight into single-cell quality control and variant calling, and the output lands beside the method that produced it.

Where a pattern in that report looks worth testing, Elnora's hypothesis skill proposes an explanation and rates each candidate by the evidence behind it, and the scientist chooses which one deserves the next run. Elnora writes the protocol for it, and on connected instruments it holds until a person approves the move, the same division of labour the Anthropic laboratory used: software read and drafted, people decided and ran the experiments. The readout returns to the knowledge base beside the protocol it came from, so the next search starts from the results and the failed attempts together.

Elnora's part of the job is the reading, the ranking and the drafting. The question worth asking, the bench and the judgement of what a result means stay with the scientist.

The reading of the public DNA record in bulk now runs in software. What a flagged pattern turns out to do, and whether it deserves a week at the bench, stay with the scientist.

Questions this post answers

Did AI agents discover a new CRISPR system in September 2026?
No. Roughly 950 agents spent more than 21 hours searching public DNA data and flagged a previously undescribed enzyme system in viruses that infect bacteria, with a repeat array that resembles a bacterial CRISPR array. Its function is unknown and no cutting partner enzyme is known, so a lead is the honest description.
What is an AI agent, in plain terms?
Software that takes a whole job and runs it on its own: it decides what to examine next, calls tools such as a search or a database, and comes back with the result. A chat reply points you at the answer, and an agent goes and gets it, which is what the Anthropic search relied on.
Can an AI agent run bioinformatics pipelines from public data?
Yes. Elnora runs single-cell quality control and variant calling directly on public accessions from GEO and SRA, in the same conversation that drafts the protocol, and every answer ends with a source list naming the literature identifier behind each claim. The scientist still judges whether the analysis fits the system in front of them.
Did the AI agents run the laboratory experiments themselves?
No. Human scientists designed and performed every physical experiment, at biosafety levels 1 and 2 with no work on pathogens that infect humans, while the agents did the computational search. The company's chief executive said the model made the find mostly, though not entirely, on its own, and the preprint is awaiting peer review.
Carmen Kivisild

Carmen Kivisild

Founder and CEO, Elnora

PhD in molecular biology, ran a wet lab before founding Elnora. Writes about running a company with agents and about what scientists actually need from them.

Keep reading