One in 20 animal therapies reached approval. So ask three questions.
Most readers are told to scoff at the phrase or to swallow the claim whole. Three plain questions about numbers, comparisons and size do better than either.
Somewhere this week a headline will tell you that a molecule, a diet or a bacterium protects against something you would rather not get. The finding will be real. The experiment will have been done. And the words "in mice" will be in the ninth paragraph, if they appear at all.
My position is simple. "In mice" is a starting point, not a verdict. It does not mean the result is worthless, and it does not mean the result will help you. It means a long chain of evidence has one link in it. Your job as a reader is not to dismiss the link or to treat it as the whole chain, but to ask three questions that tell you how well it was forged. How many animals, and what counted as the unit? Compared with what? And how big was the effect, in something that matters?
What the chain looks like
Start with a scale check, because without one people tend to swing between awe and contempt. In 2024 Benjamin Ineichen and colleagues published an umbrella review in PLOS Biology, a review of reviews, covering English-language systematic reviews of how animal findings carry over to humans. It took in 122 reviews, 54 human diseases and 367 therapeutic interventions [1].
The authors estimated that 50% of those therapies progressed from animal studies to any human study. About 40% reached a randomised controlled trial. About 5% won regulatory approval [1]. The median journey took five years to a first human study, seven to a randomised trial and 10 to approval [1].
Now the fine print, which matters. Those figures describe the interventions the review found, pooled across many diseases and many research histories. They are not the odds for any one mouse study, and you should not quote "five percent" at a friend who has just told you about a promising result in rodents. They are a picture of the whole landscape, the way a national average height tells you nothing about the person standing in front of you.
The review also reported 86% concordance between positive results in animal and clinical studies in its meta-analysis [1]. That is a different thing again. It concerns agreement among the positive findings the authors assessed, not the share of animal-tested treatments that work in people [1]. Mixing up those two numbers is the most common way to get this subject wrong, in either direction.
The chain has real successes in it. The review notes that animal studies helped lay the groundwork for multiple-sclerosis drugs including mitoxantrone and glatiramer acetate. It also notes the stroke drug NXY-059, which showed promise in animals and then failed in human trials [1]. Both stories are true. Neither tells you which one your headline is.
Question one: how many, and of what?
The first question sounds like arithmetic, but it is really about independence. A study that says it used 40 mice sounds like it has 40 data points. It may have far fewer.
The ARRIVE guidelines, the reporting checklist for animal research, say readers need the exact number of experimental units in each group and the total number of animals used [5]. The experimental unit is the thing that was truly assigned to a treatment. It may be a single animal. It may be a litter, or a cage, depending on the design [5].
Here is an invented example, to show the logic and nothing more. Suppose a researcher puts a compound in the drinking water of two cages of 10 mice and leaves two other cages on plain water. There are 40 animals, but the treatment was applied to cages, and mice sharing a cage share a good deal besides the water. The honest count of independent units is four, not 40. A paper that reports only "n = 40" has hidden the number you most needed. ARRIVE asks authors to distinguish animal counts from units such as cages for exactly this reason [5].
The same guidance asks authors to explain how they chose their sample size [5]. A small group is not a disqualification. It is a reason to hold the result loosely. A big, bold claim resting on a handful of units should make you sit up straighter, not lean in.
Question two: compared with what?
Everyone knows a study needs a control. Fewer people ask what the control group was given, and that is where a good deal of the story hides.
A control is not automatically a group that gets nothing. ARRIVE's own explanation says the right comparison depends on the research question. It might be a placebo, a sham procedure, a different method of delivering the treatment, or the same animals measured over time, serving as their own controls [4]. There are negative controls, which help test whether a difference was really caused by the intervention, and positive controls, which help show that an expected effect would be detectable in the setup at all [4].
So when you read that treated mice did better than untreated mice, ask what "untreated" meant. If the treated animals were injected and the others were left alone, you may be reading the effect of being injected. A sham procedure exists to take that explanation off the table. When a report does not say what the comparison group received, you cannot judge the comparison, and a result you cannot judge is not yet a result you can use.
Question three: how big, and does it matter?
The third question is the one headlines are worst at. A reported difference alone does not tell you whether the outcome matters for health. You have to read it against what was measured, the comparison group and the statistical methods used [2][3][4].
Another invented illustration. Imagine treated mice show a blood marker that is a few percent lower than in the comparison group, and the paper reports the gap as statistically significant. That tells you a difference was probably not an accident of sampling. It does not tell you the mice lived longer, felt better or avoided disease. A marker is a measurement near the thing you care about, not the thing itself. Ask whether the outcome is one a person would notice, and how large the change is in the units the researchers used, not just whether it cleared a threshold.
I will not give you a magic number here. No single size works for every experiment, and anyone who offers you one is selling something. What you can do is refuse to accept the word "significant" as a synonym for "large" or "important." ARRIVE asks authors to report outcome measures, statistical methods and results precisely so that a reader can make that judgment [2][3]. Use it.
Why the dull parts of a methods section matter
Behind the three questions sit two habits of good design that rarely make the news. Randomisation means the animals are assigned to groups by a method that does not depend on the researcher's hunches. Blinding means the people assessing the outcomes do not know which group is which. Both appear in the ARRIVE checklist, alongside sample size, inclusion and exclusion criteria, outcome measures and statistical methods [2][3].
The reason is human and unglamorous. People who hope for a result tend to find it. A technician who knows which mice got the drug may, without meaning to, handle them differently or score them more kindly. Randomisation and blinding are the cheap defences. When a paper does not say whether it used them, you are left to guess, and the guess should be cautious.
There is a hard fact underneath this. The authors of ARRIVE 2.0, published in 2020, wrote that adherence to the earlier guidelines had been inconsistent and that the expected improvements in research quality had not been achieved [2]. The checklist exists because reporting was patchy. Better reporting does not make a result true, but it lets you see what you are being asked to believe. The review's authors drew a related lesson from the low approval rate, which they said points to weaknesses in the design of animal studies and early clinical trials [1].
The best objection
The strongest objection to all this is that I am teaching suspicion, and suspicion is cheap. If every mouse result is met with a raised eyebrow, the reader learns nothing, and the research that did lead to a drug gets lumped in with the research that did not.
That is fair, and the data on translation back it up. The review's authors concluded that animal-to-human translation may be more successful than some earlier estimates suggested [1]. Half of the therapies it examined reached a human study. Animal work did help lay the ground for real multiple-sclerosis drugs [1]. Dismissing everything "in mice" would throw out the chain along with its weakest links.
But notice what the objection actually asks of you. It does not ask you to believe more. It asks you to be specific. The three questions are not a way of saying no. They are a way of saying how much. A well-designed study with a clear comparison and a large, meaningful effect deserves more of your attention than a small one with none of those, and the questions are how you tell them apart. Being sceptical of a bad mouse study and being interested in a good one are the same skill.
What to do with this
Next time someone tells you a study says something is good for you, ask them the three questions, gently, the way you would ask what the weather is like where they are. How many animals, and what was the unit? Compared with what? How big was the change, in something that matters?
Then keep the scale in mind. Across the 367 interventions in that review, about one in 20 made it to approval, and the typical road took a decade [1]. That is not a forecast for the mouse in your headline. It is a reminder that a result in an animal is an opening move, and most openings are not the whole game.
What to watch for next is simple. Look for the methods section, or for a news report that quotes numbers from it. If the story gives you the group sizes, the comparison and the size of the effect, someone did the work for you. If it gives you only a verdict, the sentence you actually need is the one in the ninth paragraph, and you now know how to read it.
Every edition in brief, three times a day, on our Telegram channel.
Spotted an error? Tell the editors
- Analysis of animal-to-human translation shows that only 5% of animal-tested therapeutic interventions obtain regulatory approval for human applications | PLOS Biology journals.plos.org
- The ARRIVE guidelines 2.0: Updated guidelines for reporting animal research | PLOS Biology journals.plos.org
- The ARRIVE guidelines 2.0 | ARRIVE Guidelines arriveguidelines.org
- 1a. Study design - Explanation | ARRIVE Guidelines | ARRIVE Guidelines arriveguidelines.org
- 2a. Sample size - Explanation | ARRIVE Guidelines | ARRIVE Guidelines arriveguidelines.org




