Medical research depends on something deceptively simple: getting the right biological sample, from the right person, at the right time, under the right conditions.
Blood, plasma, tissue, saliva, urine and cerebrospinal fluid underpin drug discovery, biomarker development, diagnostics, genomic research and clinical trials. Yet the process of obtaining them remains fragmented across patient registries, hospitals, laboratories, biobanks, courier networks and data systems that were never designed to talk to each other.
The consequence is familiar to anyone who has tried to assemble a well-characterised cohort: long timelines, high costs, and a persistent gap between the samples that exist and the samples a study can actually use.
Artificial intelligence is often introduced into this problem at a single point, usually patient identification. The more interesting opportunity is architectural. Applied across the full specimen lifecycle, AI can connect identification, consent, collection, logistics, processing, quality control and analysis into a coordinated supply chain rather than a relay of disconnected handoffs.
This article sets out where that opportunity is real today, where it remains speculative, and what organisations should do first.
In September 2022, we facilitated the collection of plasma from 100 people in the Democratic Republic of Congo who had recovered from mpox. Today that pool, filled into roughly 1,900 ampoules and held at −20°C, is the proposed First WHO International Standard for mpox antibodies, the primary calibrant against which laboratories worldwide will express antibody potency in International Units.
The WHO Expert Committee on Biological Standardization report on that work documents what it produced. Eighteen laboratories across eleven countries and four WHO regions tested the candidate standard against a blinded panel. For the two DRC convalescent pools, inter-laboratory variability in neutralising antibody titres fell from 61-fold and 37-fold to 8-fold once results were expressed relative to the standard.
What is biospecimen sourcing, and why is it so difficult?
Biospecimen sourcing is the process of identifying, consenting, collecting, processing, transporting, storing and distributing human biological samples for research.
The difficulty is that scientific suitability is far narrower than physical availability. A researcher who needs 500 blood samples from patients with a given disease usually needs something considerably more specific:
- A particular disease stage
- A defined age or demographic profile
- A specific treatment history and biomarker status
- Collection before treatment initiation, sometimes at a defined time of day
- Specified collection tubes and anticoagulant
- A bounded processing window
- Controlled temperature throughout transport
- A limited number of freeze-thaw cycles
- Complete clinical metadata
- Consent that permits the intended secondary use
A sample can therefore be physically available and scientifically unusable. This is why specimen quality cannot be separated from provenance, and why the ISBER Best Practices, Fifth Edition place such weight on standardised, documented handling from collection through distribution.
The operative question has shifted. It is no longer "can we find a sample?" but "can we find a sample whose provenance and metadata make it fit for the scientific question we are asking?"
How can AI improve biospecimen sourcing?
Conventional sourcing queries an inventory against inclusion criteria. An AI-enabled approach can reason across clinical notes, laboratory results, imaging, medication records, genomic data and prior research participation, then rank candidates by how well they satisfy both the biological and the operational requirements of a study.
The difference is substantive. Instead of:
Find 500 patients with Parkinson's disease.
the system can resolve:
Find patients within three years of diagnosis, not yet started on a specific therapy, with a recent neurological assessment, within a feasible collection radius, and likely to yield plasma meeting our proteomic assay requirements.
The same logic applies in reverse. Assay requirements can be translated into a machine-readable specimen specification (matrix, anticoagulant, minimum volume, maximum processing delay, storage temperature, freeze-thaw ceiling, required metadata fields), which then becomes the search key against existing collections.
This is the shift from inventory matching to fitness-for-purpose matching, and it is where most of the near-term value sits.
Prioritising the samples that matter
Not every available specimen carries equal scientific weight. A composite score combining clinical relevance, sample quality, metadata completeness, rarity, geographic representation and downstream assay suitability allows a rare, deeply annotated sample to outrank a large but thinly characterised collection.
That distinction matters most precisely when specimens are scarce or expensive to collect.
Can AI improve collection and transport?
Research no longer has to originate inside a major academic medical centre. The FDA's guidance on decentralised clinical trial activities recognises trial activities conducted away from traditional sites, including in participants' homes and through local providers.
That regulatory framing enables a distributed collection network: participant, local phlebotomist, regional laboratory, courier, central biorepository. Scheduling, collector availability, courier selection, routing, processing deadlines and laboratory capacity become a single optimization problem rather than five separate ones, with the system continuously re-estimating whether a specimen will still arrive inside its acceptable processing window.
Predictive cold-chain management
Combining temperature telemetry, location, transit time, weather, sample type and processing requirements allows a system to estimate the probability that a specimen will remain fit for its intended assay, and to flag risk while the sample is still in motion rather than after it arrives.
How can AI improve biospecimen quality control?
A structured digital specimen record captures the variables that determine downstream reliability:
|
Variable |
Example |
|
Donor identity |
Verified |
|
Collection time |
08:42 |
|
Collection method |
Standardised per SOP |
|
Time to processing |
24 minutes |
|
Temperature excursion |
None recorded |
|
Volume |
3.2 mL |
|
Haemolysis index |
Low |
|
Freeze-thaw cycles |
1 |
|
Metadata completeness |
97% |
|
Assay suitability |
High |
Across thousands of specimens and several years, models can surface interactions that manual review will not: a particular combination of processing delay, storage temperature and freeze-thaw history that degrades proteomic signal, for instance, even where each individual variable sits within SOP tolerance.
Traditional quality control asks whether a sample followed the SOP. Predictive quality control asks how likely that sample is to produce reliable results for a specific assay. These are not the same question, and the second is more useful.
None of this works without disciplined pre-analytical annotation. Organisations building toward it should adopt SPREC (the Standard PREanalytical Code) for handling variables, BRISQ for reporting, and MIABIS for interoperable sample and collection descriptions. Models cannot correct variables nobody recorded.
What is a digital biobank?
A digital biobank is not a database of freezer locations. It links the physical specimen to the full body of information that gives it meaning: collection and processing records, storage history, molecular data, whole-slide images, longitudinal clinical outcomes and environmental context.
Recent reviews of AI in biobanking describe exactly this convergence, with machine learning supporting automated quality control, multimodal integration and predictive analytics, while well-annotated biospecimens simultaneously supply the training data those models require.
The practical effect is a different kind of query result. Rather than:
1,423 samples available.
a researcher receives:
327 samples meet the biological criteria; 214 have complete longitudinal follow-up; 168 clear the assay quality threshold; 74 are paired pre- and post-treatment specimens.
How do digital twins relate to biospecimens?
A digital twin is a computational representation of an individual, integrating clinical records, laboratory results, genomics, imaging, wearable data, biomarkers and longitudinal outcomes to model disease progression and treatment response.
The connection to biospecimens runs in both directions.
The sample informs the twin. Protein expression, immune activity, metabolic state, genetic variation and pathogen signatures derived from a specimen update the model's representation of that patient.
The twin informs the next sample. Where a model identifies its own largest source of uncertainty, it can indicate which single measurement would most reduce it. Instead of collecting to a fixed schedule, collection becomes adaptive:
patient → specimen → data → model → prediction → optimal next specimen
For longitudinal studies with scarce or invasive sampling, this is a meaningful efficiency gain.
However,
Digital twins are advancing quickly, but their limits deserve stating plainly, particularly by anyone building in this space.
The validity of a predictive patient model depends on exchangeability between the population it was trained on and the population it is applied to. A model trained predominantly on North American and European trial and registry data will not transport cleanly to a cohort with different comorbidity burden, standard of care, competing mortality risk, genetic background or measurement practice. Applied without adjustment, it does not merely lose accuracy; it introduces structured bias.
Regulators have been explicit that these applications require a defined context of use and risk-proportionate validation. The FDA's draft guidance on AI in regulatory decision-making for drugs and biologics (January 2025) and its earlier draft guidance on externally controlled trials (February 2023) are the relevant reference points.
The honest position today is that digital twins are most defensible as tools for study design, sampling strategy and hypothesis prioritisation, and least mature as substitutes for control-arm patients or for physical measurement.
What does this mean for research for places like Africa?
The case is strongest where disease burden is high, populations are geographically distributed, and existing research infrastructure is thin. WHO estimates that roughly 40% of the global neglected tropical disease burden falls in Africa, and emphasises the role of laboratory networks in identifying disease foci, tracking prevalence and detecting drug resistance.
But the standard AI-for-research playbook does not transfer directly, and it is worth being clear about why.
The EHR assumption does not hold
Most published AI recruitment tooling assumes a decade of digitised free-text clinical notes. Across much of sub-Saharan Africa that substrate does not exist in the same form. What exists instead is different, and in some respects better suited to biospecimen research:
- Health and Demographic Surveillance System (HDSS) sites, which follow geographically bounded populations through repeated household visits, several for more than 25 years. These offer complete denominators, low loss to follow-up and longitudinal depth that few Western EHR systems match.
- DHIS2, used for routine health information nationally across more than 80 countries, valuable for burden mapping and site feasibility.
- Laboratory information systems, frequently the most complete individual-level electronic record in a facility.
- Paper registers and community health worker records, still the largest single source of clinical information.
The highest-value AI applications in this context are therefore different from those in high-income settings. Structured extraction from paper and semi-structured documents; probabilistic record linkage in the absence of universal patient identifiers; predictive cold-chain management, where the constraint binds hardest; and federated approaches that allow models to be trained across institutions without moving raw data. A multi-country chest imaging study across eight African countries has demonstrated the feasibility of the latter, alongside a candid account of the infrastructure, connectivity and governance obstacles that remain (Fabila et al., 2025).
From sample collection to adaptive surveillance
Consider a region reporting an unexplained rise in febrile illness with neurological features. Community health workers collect standardised specimens; the system records symptoms, age, location, collection date, sample type, environmental conditions and laboratory results.
Pattern detection across that population can identify emerging clusters and, more usefully, indicate where additional sampling would yield the most information. Extended further, the same architecture supports a model of a community or disease ecosystem integrating human cases, pathogen sequences, vector data, rainfall and temperature, population movement, healthcare access and laboratory capacity.
The questions it helps answer are operational: where should we sample next, which communities carry the greatest uncertainty, which specimens should be prioritised for sequencing, where should mobile laboratories be deployed.
This is not a replacement for epidemiologists or laboratory scientists. It is a way of directing scarce resources toward the samples that will generate the most information.
The H3Africa consortium's experience building biorepositories in Nigeria, Uganda and South Africa is instructive: ethics approvals and material transfer agreements routinely took months, particularly at institutions unfamiliar with biospecimen, data and benefit-sharing frameworks (Genome Medicine, 2023).
Add the Nagoya Protocol, national data protection legislation such as POPIA, Kenya's Data Protection Act and Nigeria's NDPA, and a well-founded regional sensitivity to extractive research practice. Any system that moves samples or data across borders will be evaluated on governance before it is evaluated on model performance.
What should organisations do now?
The wrong starting question is "how do we add AI to our biobank?" The better one is "where are we losing biological information, time or sample quality across the specimen lifecycle?"
A practical sequence:
- Map the lifecycle end to end. Identification, consent, collection, processing, transport, storage, retrieval, analysis, metadata capture, data integration. Locate where information is lost rather than where tasks are slow.
- Fix the metadata layer first. Adopt SPREC, BRISQ and MIABIS. This is unglamorous and it is the constraint on everything downstream.
- Instrument the cold chain. Temperature and transit telemetry is inexpensive and generates the training data for predictive quality control.
- Start with extraction and linkage, not prediction. Getting existing records into structured, linkable form delivers more value in year one than any model.
- Define context of use before deploying predictive models. Specify the population, the decision the model informs, and the validation evidence required, in line with current regulatory guidance.
For organisations building this infrastructure, ISBER Best Practices, Fifth Edition remains the foundational reference, and ISBER's Essentials of Biobanking course covers planning, establishment, maintenance and access.
Where this leads
The future of biospecimen sourcing is unlikely to be a single centralised biobank. It is more likely to be a distributed network: a participant recruited in their community, sampled at home or a local clinic, with processing at a regional laboratory, sequencing at a specialised centre, and results connected to a longitudinal model of that individual and that population.
AI's role in that architecture is orchestration and prioritisation. Not more samples, but a clearer answer to which samples matter, why, where they are, whether they are reliable, and what should be collected next.
For pharmaceutical developers, that means faster biomarker discovery and better-powered studies. For academic researchers, it means fragmented biological resources become discoverable and usable. For infectious disease programmes, it means scattered clinical samples can function as a continuously updating surveillance network.
Frequently asked questions
What is AI-powered biospecimen sourcing? The use of machine learning, natural language processing and data integration to identify appropriate donors and specimens, match them against research requirements, optimize collection and logistics, and assess sample quality and fitness for a specific downstream assay.
How can AI improve biobank management? Current and emerging applications include sample discovery, metadata extraction from unstructured records, automated quality control, inventory and storage optimisation, prediction of specimen degradation, workflow automation, and multimodal analysis linking specimens to imaging and molecular data.
Can AI predict whether a biospecimen is usable? Potentially, yes. Models can combine collection, processing, storage and sample characteristics to estimate suitability for a given assay. Such predictions must be validated for their intended use and treated as a complement to, not a replacement for, established laboratory quality procedures.
How do digital twins relate to biospecimens? Biospecimens supply the biological observations that update a computational model of a patient, and the model can in turn indicate which additional measurement would most reduce remaining uncertainty. The relationship is bidirectional, and the models require validation against the specific population they are applied to.
Why does this matter for infectious disease research? Biological samples carry information about pathogens, immunity, disease progression and transmission. Linking them to geography, clinical data and environmental conditions turns individual specimens into inputs for responsive surveillance and adaptive sampling strategies.
What standards should a biobank adopt before investing in AI? ISBER Best Practices for overall operations, SPREC for pre-analytical variable coding, BRISQ for reporting, and MIABIS for interoperable sample and collection descriptions. Model performance is bounded by metadata quality.
At Infiuss Health, we work at the intersection of decentralised research, AI and biospecimen infrastructure. Our PROBE platform connects clinical research operations, participant management and research data, and our Digital Patient Twin work focuses on integrating clinical, molecular and imaging data into predictive models.
If your organisation is exploring AI-enabled biospecimen sourcing, decentralised sample collection, digital biobanking or patient digital twins, get in touch to discuss how this applies to your research programme.
