“AI or Not AI” Is No Longer the Question in Laboratory Animal Research

By Stefano Gaburro

 

A familiar question still circulates at conferences and in grant review panels. Will artificial intelligence replace animals in research? The framing is comfortable because it sorts the world into two camps. It is also wrong. AI is not a substitute for an in vivo model, and an in vivo model is not a defence against computation. The two operate on different parts of the same problem. The real question is no longer whether to use AI. It is whether our data are good enough for AI to mean anything.





Let’s begin with the assumption that AI is a replacement technology quietly retiring the animal in biomedical research. That belief misreads what these tools actually do. In laboratory animal science, the dominant uses of machine learning today are not a replacement at all. They are refinement and reduction, in the precise sense of the 3Rs, as I will further explain.

What AI is actually doing in the vivarium

Consider the daily health check, the oldest task in the animal facility. Technicians inspect cages briefly, usually during the rodents' sleep phase, and subtle signs of distress are easy to miss behind enrichment material. A multicentric study analysing locomotion from three institutions with the same sensor technology showed that machine-learning could identify animals in distress three to six days before verifiable clinical signs or death, with an accuracy of 66 to 80 percent at day minus three and 80 to 91 percent at day minus six [1]. Thus, continuous monitoring of locomotion outperformed human observation in critical health monitoring. This is refinement, through earlier detection of subclinical cases, as well as reduction, through more objective, quantifiable and replicable humane – and scientific – endpoints, with fewer animals being lost to undetected decline. Computer vision can similarly allow for automated pain recognition, through convolutional neural networks capable of reading facial expressions and body language to produce more objective welfare assessments than subjective grimace scoring by an operator [2]. These are genuine contributions, yet still conditional, as there are issues concerning data scarcity, the absence of robust ground truth, and unresolved ethical questions in automated pain-recognition [2].

Behavioural analysis tells the same story. Manual scoring is slow, subjective, and poorly reproducible, with inter-rater reliability that few laboratories report and fewer defend [3]. Open-source pose estimation and supervised classifiers now detect and categorize rodent behaviours with a consistency that human observers cannot match across sites and sessions [3]. Protocols built on DeepLabCut and SimBA convert raw video into pose data and then into behavioural classifiers for such tasks as operant self-administration, work that previously depended on hand-coded lever press counts [4].

A constraint holding back automated assessment was group housing. Single-animal video tracking matured years ago, yet mice should be housed socially, unless there is a strong justification for single-housing, and hence the algorithm will lose track of which animal is which, making the whole group within each cage collapse into one averaged replicate. A recent system addressed this directly. Smart Lids mounted above standard cages, paired with a multi-organism tracker and a purpose-designed ear tag, achieved over 97 percent multi-animal tracking accuracy and returned 21 health metrics covering activity, feeding, drinking, aggression, and sleep, at a runtime cost under 100 dollars per month [5]. None of this removes the animal. It removes the noise around the animal, preserves the individual inside the social group and hence also our ability to treat each individual animal as an independent replicate (aka “experimental unit”), preventing pseudoreplication and sample size overestimation, thus preserving statistical power and furthering the R for Reduction.

From welfare check to regulatory endpoint

The more consequential shift is the move of machine-learning-powered automated systems from husbandry into the regulated core of drug safety. Safety pharmacology studies conducted under ICH S7A protect first-in-human volunteers from acute harm to the cardiovascular, respiratory, and central nervous systems. They work for acute, life-threatening effects but are weaker at detecting subtle, progressive ones, partly because central nervous system assessments require removing the animal from its home cage, depend on human observation, use discontinuous endpoints, and are run during the light phase when rodents are least active [6]. Continuous home cage monitoring mitigates each of these weaknesses through automation, quantitative outputs, and integration into longer repeat-dose studies, which is why there is now a documented call to update ICH S7A to leverage digital measures [6].

Here AI is not replacing the animal model, but rather making it more translational, objective, and human-relevant. The authors are candid about a key caveat, which is that many of these sensors rely on machine learning algorithms trained on human-selected examples, so human bias can re-enter an otherwise unbiased pipeline unless it is controlled through observer redundancy, over-reads, and analytical validation against the detected event [6]. The technology therefore does not dismiss the need for rigorous data collection, analysis and validation, but rather critically depends on it.



Where AI sits on the replacement axis

AI also enters the replacement conversation, though indirectly, by making non-animal models (NAMs) credible enough to qualify. “New approach methodologies” (another designation for NAMs), which encompass in silico, in vitro, and organotypic methods, increasingly depend on computation to be usable. Imaging-based in vitro assays generate high-content data that only machine learning can analyse at scale. Piergiovanni et al [7] hence have placed validation, inter-laboratory transferability, and data reliability at the centre of regulatory acceptance.

NAMs face other barriers that have little to do with the algorithms. A cross-sector analysis grouped the obstacles into five recurring themes – perception, regulatory acceptance, scientific and technical limitations, education, and financial constraints – and noted that adoption remains slow and awareness outside the specialist community remains low [8]. An earlier review reached the same diagnosis from the chemical safety side, where cultural inertia and regulatory expectation, not raw capability, have prevented the uptake of better methods [9].

The regulatory signal is widely misread

It is tempting to read recent policy as the moment the binary animal vs. non-animal methods resolved in favour of replacement. That reading is not accurate. The FDA Modernization Act 2.0 amended the Federal Food, Drug, and Cosmetic Act to permit non-animal methods as alternatives, but it did not abolish animal testing, and it did not declare any specific method qualified for any specific context. The literature that celebrates the law is often loose on this point, describing reduced requirements when the actual change is permissive language that basically opens a door [10,11]. This is an important distinction, given that a door that is open still requires a model that can walk through it, and that model must be validated for a defined context of use.

And this is where the enthusiasm for organoids and organ-on-chip technology meets reality. While these systems mimic human physiology with increasing fidelity and offer real advantages in disease modelling and screening, regulatory issues persist around their reliability, scientific robustness, and applicability [11]. However, there is promise in such fields as patient-derived induced pluripotent stem cell platforms, which have enabled clinical-trial-in-a-dish precision medicine strategies for testing cardiovascular drugs [10], with the proposed approach being the combination of microphysiological systems, clinical omics and artificial intelligence rather than betting on any single method. [10]

The constraint nobody wants to name

If “AI or not AI” is not the question, what is it, then? The rather unglamorous answer is metadata, provenance, and reporting quality. A systematic review of reporting guidelines in medical AI found 26 standards published between 2009 and 2023, varying in consensus quality, breadth, and target research phase, with persistent gaps that undermine validity and reproducibility [12]. The rate of development of AI and AI-reliant tools has outrun the capability of community guidelines, consensus documents (e.g. standardized, agreed-upon protocols) and regulations to provide a framework for scientific robustness and regulatory safeguards under which to base our trust in them. An algorithm trained on poorly annotated, non-reproducible in vivo data will produce confident outputs that translate no better than the unreliable endpoints they were built on.

All of this reframes the old translation debate. Preclinical work often fails to predict clinical outcomes for systemic reasons, weak endpoints, underpowered studies, misaligned incentives, and poor metadata, and not necessarily because animals are inherently poorly predictive models. AI will not only fail to fix any of these problems by itself, but will moreover amplify whatever it is given. Feed it continuous, well-annotated home cage data and it will sharpen phenotyping and identify and flag health and welfare decline days early. Feed it retrospective, inconsistently labelled study archives and it will manufacture noise with a veneer of spurious precision. FAIR (findable, accessible, interoperable, reusable) and reliable data is therefore not a luxury layered on top of the science producing it, but a precondition for AI to add value at all, and it has to be planned, captured at acquisition, beforehand, and not reconstructed afterwards.

A more mature position

The implications of the tension between the promise and caveats of AI use in biomedical research do not call for a compromise, as animal research is not obsolete and NAMs are not a default replacement. AI should neither be considered a threat to the first nor a saviour, delivering the second. The honest framing is that the type of model used must not govern preclinical research strategy, but rather the context of its use, with regard to the biological question at hand, the extent to which the data can answer it, and the standard evidence that decision requires. Only then does it make sense to ask which combination of in vivo, in vitro, and in silico methods fits.

It is time to retire this binary framing of the issue. The laboratory of the next decade will not choose between AI and animals, but run continuous digital monitoring alongside in vitro and in silico models, with machine learning being trained on all of it, and governed by the discipline of standardized, reusable data. Commercial platforms from The Jackson Laboratory and others already package computer vision home cage monitoring as validated infrastructure rather than experiment [13]. Thus, the organizations that will succeed will not be the ones that adopted AI first, but the ones whose data were reliable enough in the first place, making their adoption of AI improve everything.

References
  1. Eswaraka J, Gommet C, Diomaiuta D, Rigamonti M, Rosati G, Gaburro S, Zwick M, Begoud L, Warot X, Doenlen R. Enhanced health evaluation in mice using continuous home-cage monitoring and machine learning: A multicentric study. Lab Animals 2026;55,275–281. https://doi.org/10.1038/s41684-026-01745-2
  2. Chiavaccini L, Gupta A, Chiavaccini G. From facial expressions to algorithms: A narrative review of animal pain recognition technologies. Frontiers in Veterinary Science. 2024;11,1436795. https://doi.org/10.3389/fvets.2024.1436795
  3. Isik S, Unal G. Open-source software for automated rodent behavioral analysis. Frontiers in Neuroscience. 2023;17,1149027. https://doi.org/10.3389/fnins.2023.1149027
  4. Pereira-Sanabria LF, Voutour LS, Kaufman VJ, Reeves CA,
    Bal AS, Maureira F, Arguello AA. Analysis of operant self-administration behaviors with supervised machine learning: Protocol for video acquisition and pose estimation analysis using DeepLabCut and Simple Behavioral
    Analysis. eNeuro. 2025;Feb6;12(2):ENEURO.0031-24.2024. https://doi.org/10.1523/ENEURO.0031-24.2024
  5. Delalic S, Kaca M, Alimsijah P, Weber N, Selmanovic E, Galindez M, Marquez G, Balmaceda F, Delalic E, Bekkaye I, Bakija L, Kurtagic-Pasalic M, Agic E, Anderson D, Wagers A, Florea M. Smart Lids for deep multi-animal phenotyping in standard home cages. Front Behav Neurosci. 2026;19:1696654. https://doi.org/10.3389/fnbeh.2025.1696654
  6. Berridge BR, Liu CN, Cherian AK, Lainee P, DaSilva JK, Anger LT, Foley CM, Bratcher-Petersen N. Non-invasive, home cage digital monitoring for improved safety pharmacology assessments in drug development: ICH S7A update considerations. J Pharmacol Toxicol Methods. 2026;140:
    108430. https://doi.org/10.1016/j.vascn.2026.108430
  7. Piergiovanni M, Mennecozzi M, Barale-Thomas E, Danovi D, Dunst S, Egan D, Fassi A, Hartley M, Kainz P, Koch K, Le Devedec SE, Mangas I, Miranda E, Nyffeler J, Pesenti E, Ricci F, Schmied C, Schreiner A, Stokar-Regenscheit N, Whelan M, et al. Bridging imaging-based in vitro methods from biomedical research to regulatory toxicology. Arch Toxicol. 2025;99(4):1271–1285. https://doi.org/10.1007/s00204-024-03922-z
  8. Oyetade OB, Allen DG, Carder J, Farley-Dawson EA, Reinke EN, Marko S, Kleinstreuer NC, Hogberg HT. Ways to broaden the awareness, consideration and adoption of new approach methodologies (NAMs). ALTEX. 2025;42(4):714–726. https://doi.org/10.14573/altex.2505281
  9. Sewell F, Alexander-White C, Brescia S, Currie RA, Roberts R, Roper C, Vickers C, Westmoreland C, Kimber I. New approach methodologies (NAMs): Identifying and overcoming hurdles to accelerated adoption. Toxicol Res (Camb). 2024;13(2):tfae044. https://doi.org/10.1093/toxres/tfae044
  10. Wu X, Swanson K, Yildirim Z, Liu W, Liao R, Wu JC. Clinical trials in-a-dish for cardiovascular medicine. Eur Heart J. 2024;45(40):4275–4290. https://doi.org/10.1093/eurheartj/ehae519
  11. Zhou L, Huang J, Li C, Gu Q, Li G, Li ZA, Xu J, Zhou J, Tuan RS. Organoids and organs-on-chips: Recent advances, applications in drug development, and regulatory challenges. Med. 2025; 6(4):100667. https://doi.org/10.1016/j.medj.2025.100667
  12. Kolbinger FR, Veldhuizen GP, Zhu J, Truhn D, Kather JN. Reporting guidelines in medical artificial intelligence: A systematic review and meta-analysis. Commun Med (Lond). 2024;4(1):71. https://doi.org/10.1038/s43856-024-00492-0
  13. The Jackson Laboratory. The Jackson Laboratory and Allentown unveil a revolutionary AI-driven home-cage monitoring solution set to transform preclinical research [Press release]. 2024 Nov 3. https://www.jax.org/news-and-insights/2024/november/envision-announcement

Download
“AI or Not AI” Is No Longer the Question in Laboratory Animal Research.pdf