Skip to main content
All posts
Data4 min read

Test the AI hypothesis at your own bench and keep the data

An AI model proposes ideas from a published record skewed towards success. An AI agent for drug discovery lab work earns its place on the validation experiment: designed well, run at the bench and kept, negatives included.

Carmen Kivisild

Carmen Kivisild

Founder and CEO

An AI model can propose a drug target or a compound in minutes, and the value of that proposal is decided at the bench, in an experiment the lab runs and records itself. For an AI agent for drug discovery lab work, the useful job is getting that validation experiment right on the first run and keeping its result, the misses included, because a lab's own tested results are the data the public record leaves out.

Why the public record is skewed

The skew has a name: publication bias. Researchers in psychology and medicine began measuring it in the 1970s, when they noticed that a question could look settled in the journals while the studies that found no effect sat unpublished in someone's filing cabinet, which is why one of them called it the file drawer problem. In one plain sentence, the published record over-represents results that worked, because positive results are the ones that get written up, submitted and accepted.

A phone's photo album shows the same effect. Most people delete the blurred shots and the grey afternoons, so anyone scrolling through the album a year later sees a sunny, well-framed holiday, and the album is an accurate record of what was kept and a skewed record of the week. Restaurant ratings work the same way: reviews come mostly from customers who had a very good or a very bad meal, so the average star count describes the people who chose to write, and the diners who found the meal fine and went home are missing from the number.

An AI model is software that learns patterns from a large collection of past examples and uses them to score a new case, and it inherits whatever skew those examples carry. A model trained on published biology has seen far more compounds reported as active than compounds reported as inactive, and it scores new candidates on that basis.

Why it matters in biology

Drug development shows what happens when plausible ideas meet the test. About 13% of drug candidates that enter Phase I trials reach launch, and historically roughly 62% of compounds entering Phase II failed and about 45% of those entering Phase III failed, with lack of efficacy and safety problems each behind roughly 30% of the losses. Those candidates had passed years of preclinical work, so a lack of efficacy in patients means an idea that looked right in the literature and in the models failed when it was tested in people.

AI has sped up the front of that process, shortening early discovery by roughly 30 to 70% on current estimates, while the probability that a clinical candidate reaches approval has stayed within the industry's usual range of 8 to 12%. Faster ideas have left the odds where they were, and the step that moves the odds is the experiment that tests the idea in the lab's own cells, assays and hands.

That experiment has to be the right one: a dose range wide enough to find the potency, controls that rule out the obvious alternative explanation, and enough replicates to trust the curve, even when a smaller design would be quicker to run. Its result, positive or negative, is data that exists nowhere else, and a negative result from a well-designed run is the kind of example the public record is shortest of.

How Elnora fits

A scientist has a list of twenty compounds that a model trained on public bioactivity data scored as likely inhibitors of their target enzyme. They ask Elnora to design a dose-response experiment for the eight highest-scoring compounds, to add two the model scored as inactive as a check on the model itself, and to plan it for the lab's connected liquid handler and plate reader.

Elnora first searches the lab's knowledge base with that question in full, and the passages that come back, each tied to the file it came from, include an earlier run of the same assay and the concentration range that worked in it. It returns a plate map recording what each well holds, with the ten compounds in dilution series, vehicle and reference inhibitor controls, and replicates, and its critic agent ranks the weak points in the design and proposes a fix for each while the plate can still change. Elnora then writes the run file the liquid handler and plate reader accept, the run is rehearsed in simulation, and a scientist approves it before the plate is set up.

When the plate reader finishes, its readout file binds to the plate map, which ties each reading to the compound and concentration in that well, so a surprising value can be checked against its sample. The readings go straight into a dose-response fit for each compound, and the scientist gets a fitted curve, a potency estimate and the measured wells it was fitted from. Say three of the eight top-scoring compounds give a clean curve and the rest stay flat: the flat curves and the three hits go into the knowledge base together, stored with the protocol for this plate, so the next design, and the next question anyone in the lab asks, starts from what this lab measured.

In that example Elnora searches, designs, plans and analyses, and the scientist chooses which compounds deserve a plate, runs the bench work and judges what the curves mean.

The point

A model trained on published data proposes ideas from a record skewed towards success, and the experiment that tests them is where their value is settled. The lab that designs that experiment carefully and keeps its results, failures included, owns the data that makes its next prediction better.

Questions this post answers

What is publication bias in scientific research?
Publication bias is the tendency of the published record to over-represent results that worked, because positive findings are the ones that get written up, submitted and accepted, while studies that found no effect stay unpublished. An AI model trained on published data inherits that skew, having seen far more compounds reported as active than compounds reported as inactive.
Why do drug candidates still fail when AI speeds up discovery?
AI shortens early discovery by roughly 30 to 70% on current estimates, and the probability that a clinical candidate reaches approval has stayed around 8 to 12%. Historically about 62% of compounds entering Phase II failed, with lack of efficacy and safety problems each behind roughly 30% of the losses, and both show up only when the idea is tested.
How can an AI agent help validate a hypothesis in a drug discovery lab?
Elnora designs the validation experiment and analyses what comes back. It returns a plate map with dilution series, controls and replicates, writes the run file for connected instruments once a scientist approves the run, fits a dose-response curve with its potency value from the plate readout, and files the hits and the flat curves in the lab knowledge base beside the protocol.
Should a lab keep negative results from validation experiments?
Yes. A negative result from a well-designed run is the kind of example public databases are shortest of, because results that found nothing are rarely published. Filed beside the protocol that produced it, a flat dose-response curve tells the next scientist which compounds and concentrations were already tested and what the lab measured.
Carmen Kivisild

Carmen Kivisild

Founder and CEO, Elnora

PhD in molecular biology, ran a wet lab before founding Elnora. Writes about running a company with agents and about what scientists actually need from them.

Keep reading