Bioinformatics
Bioinformatics is the branch of science that builds computational methods, databases, and software for storing, searching, and interpreting biological data, above all the sequences of DNA, RNA, and proteins and the three-dimensional structures those sequences fold into. Paulien Hogeweg and Ben Hesper began using the word in 1970 and defined it as the study of informatic processes in biotic systems. By Hogeweg's own account the term became mainstream in the late 1980s, narrowing in practice to mean the development and use of computational methods for the management and analysis of sequence data.[1]
The field exists because biological measurement outruns human reading speed by many orders of magnitude. GenBank release 272, dated June 2026, held 7,618,210,921,117 bases across 264,214,354 sequence records, plus a further 49 trillion bases in roughly 5 billion whole-genome shotgun records.[2] UniProt release 2026_02 contained 149,810,139 protein entries, of which only 575,503 sit in the manually reviewed Swiss-Prot section.[3] The Protein Data Bank held 257,181 experimentally determined and integrative structures against 1,062,058 computed structure models in July 2026.[4] Those ratios, roughly 259 automatically annotated protein entries for every curated one and roughly four predicted structures for every solved one, describe the working conditions of the discipline. Manual curation cannot keep pace with data acquisition, so most annotation is produced computationally.
Machine learning has been part of bioinformatics since the 1990s, when hidden Markov models became standard for gene finding and for modelling protein families. What changed after 2020 is that learned models became the default estimator for several biological quantities rather than one option among many. AlphaFold did this for protein structure, and the pattern has since repeated for regulatory genomics, variant interpretation, and single-cell analysis.
Origins
The first collected protein sequence database originated in the early 1960s with Margaret Dayhoff's work on protein evolution, and became the collection later maintained as the Protein Information Resource.[5] From it Dayhoff derived the percent accepted mutation (PAM) substitution matrices, which score how likely one amino acid is to replace another over evolutionary time.[6] Substitution matrices remain the scoring layer underneath most sequence comparison.
Two algorithms defined the next stage. Saul Needleman and Christian Wunsch published a general dynamic programming method for comparing two protein sequences end to end in 1970.[7] Temple Smith and Michael Waterman modified it in 1981 to find the best-scoring local subsequence match rather than forcing a global alignment, which matters because homology is usually confined to domains rather than whole proteins.[8] Both are exact but cost time proportional to the product of the sequence lengths, which made them impractical against growing databases.
The heuristic answer was BLAST, published in 1990 by Stephen Altschul, Warren Gish, Webb Miller, Eugene Myers, and David Lipman. BLAST approximates optimal local alignment by finding short high-scoring word matches and extending them, and its authors gave a statistical theory for the resulting maximal segment pair scores so that a hit could be assigned significance. The original paper reported it to be "an order of magnitude faster than existing sequence comparison tools of comparable sensitivity."[9] BLAST turned database search into an everyday operation and remains a standard reference method.
Sequence alignment and search
Alignment remains the load-bearing operation. Pairwise dynamic programming handles two sequences exactly. Multiple sequence alignment extends this to families and feeds phylogenetic inference and profile construction. Profile hidden Markov models generalize a family alignment into a position-specific probabilistic model, which detects remote homologs that pairwise scoring misses.
Short-read sequencing forced a different tradeoff. Aligning hundreds of millions of 100-base reads to a 3-billion-base reference is not a search problem so much as an indexing problem. Heng Li and Richard Durbin's BWA, published in 2009, used the Burrows-Wheeler transform to build a compressed full-text index of the reference and made whole-genome read mapping routine on commodity hardware.[10]
Structure search followed the same trajectory once predicted structures became abundant. Foldseek, published online in 2023, encodes tertiary amino acid interactions as a sequence over a structural alphabet, converting three-dimensional comparison back into a string problem. It reported computation times four to five orders of magnitude lower than Dali, TM-align, and CE while retaining 86, 88, and 133 percent of their respective sensitivities.[11] That speedup is what makes searching hundreds of millions of predicted structures feasible.
Databases
Most primary bioinformatics data sits in a small number of public repositories run by publicly funded institutes and offered free of charge.
| Resource | Content | Reported scale |
|---|---|---|
| GenBank (NCBI) | Annotated nucleotide sequences | 264,214,354 records / 7.6 trillion bases; WGS division 4,994,330,522 records / 49.1 trillion bases (release 272, June 2026)[2] |
| UniProtKB | Protein sequences and functional annotation | 149,810,139 entries, 575,503 reviewed in Swiss-Prot (release 2026_02, 10 June 2026)[3] |
| Protein Data Bank | Experimentally determined biomolecular structures | 256,789 experimental plus 392 integrative structures; 1,062,058 computed structure models held alongside (July 2026)[4] |
| AlphaFold Protein Structure Database | Predicted protein structures | Over 200 million predictions[12] |
| NAR Online Molecular Biology Database Collection | Catalogue of public biological databases | 2,173 entries after the 2026 review[13] |
The 2026 Nucleic Acids Research database issue carried 182 papers, 84 of them describing new resources, and its editors removed 319 outdated or discontinued entries from the collection in a single review cycle. They also relaxed the journal's stance on access restrictions, noting that funding problems now reach even major established databases.[13] Attrition and funding fragility are ordinary features of the ecosystem.
Genomics and sequencing analysis
The Human Genome Project's first reference sequence cost somewhere between $500 million and $1 billion to generate, by the National Human Genome Research Institute's own accounting. A high-quality draft still cost about $14 million in 2006, and had fallen below $1,500 by late 2015.[14] That collapse in per-genome cost is what created the analysis bottleneck that modern bioinformatics addresses.
A standard resequencing pipeline maps reads to a reference, then calls variants. Variant calling was traditionally handled by statistical models with hand-built error terms. DeepVariant, published in Nature Biotechnology in 2018, reframed it as image classification: read pileups are rendered as images and a convolutional neural network classifies each candidate site. The paper reported that the model generalized across genome builds, sequencing platforms, and even species, so non-human projects could reuse human truth sets.[15]
Reference genomes themselves improved. The Telomere-to-Telomere consortium published T2T-CHM13 in 2022, a 3.055 billion base pair assembly with no gaps covering every chromosome except Y. It added close to 200 million base pairs that earlier references had left unresolved, including all centromeric satellite arrays and the short arms of the five acrocentric chromosomes.[16]
Population-scale sequencing followed. The UK Biobank consortium reported whole-genome sequences for 490,640 participants in 2025, and noted that most disease associations were still observed primarily in participants of European ancestry, with some strong signals in African and Asian ancestry groups.[17] Ancestry imbalance in reference cohorts is a recurring source of bias in downstream variant interpretation.
Phylogenetics
Phylogenetics reconstructs evolutionary relationships from sequence data using parsimony, maximum likelihood, or Bayesian inference over substitution models. IQ-TREE 2, released in 2020, is a widely used maximum-likelihood package built for genome-scale alignments and a large model catalogue.[18]
Phylogenetic inference also underpins genomic epidemiology. Nextstrain, published in 2018, combines a pathogen sequence database, an evolutionary analysis pipeline, and an interactive visualization layer to track pathogen spread in close to real time, joining sequences with geographic and host metadata.[19]
Structural bioinformatics
Structure prediction has been evaluated since 1994 by CASP, a biennial blind experiment in which groups model targets whose experimental structures are not yet public. CASP16 ran from May to September 2024, drawing roughly 100 groups and more than 80,000 models across about 300 targets.[20]
The protein folding problem, predicting a three-dimensional structure from amino acid sequence alone, had been open for more than 50 years by the AlphaFold authors' own description. AlphaFold2, from Google DeepMind, won CASP14 in 2020 and was published in Nature in 2021; its authors reported that the network could "regularly predict protein structures with atomic accuracy even in cases in which no similar structure is known."[21] Google DeepMind reports more than 3 million users of AlphaFold across over 190 countries and more than 40,000 citations of the 2021 paper.[22] An independent three-track network, RoseTTAFold, was published in Science the same year by Minkyung Baek and colleagues, operating on sequence, distance-map, and coordinate representations at once; its authors described it as producing accuracies approaching those of DeepMind in CASP14 rather than matching them.[23]
The 2024 Nobel Prize in Chemistry went half to David Baker for computational protein design and one quarter each to Demis Hassabis and John Jumper for protein structure prediction.[24]
AlphaFold 3, published in 2024 by Google DeepMind and Isomorphic Labs, moved to a diffusion-based architecture that predicts complexes containing proteins, nucleic acids, small molecules, ions, and modified residues in one framework.[25] Google reported at least a 50 percent improvement over existing methods for protein interactions with other molecule types, and released model code and weights for academic use in November 2024 after initially offering only a hosted server.[26]
Openly licensed alternatives followed quickly. Boltz-2, released as a 2025 preprint under a permissive open license, extended structure prediction toward drug discovery by predicting binding affinity alongside structure. Its authors described it as the first AI model to approach free-energy perturbation performance on small molecule binding affinity while being at least 1000 times more computationally efficient than that method.[27] CASP17 is running through 2026, with 130 targets released by mid-July against a 31 July target deadline, and results due at a December 2026 meeting in Rome. Its categories include immune complexes, organic ligand complexes, nucleic acids, conformational ensembles, difficult protein structures, and accuracy estimation.[28]
Machine learning and foundation models
Protein sequence data suited self-supervised learning well, because evolutionary constraint leaves statistical regularities that a sequence-modelling objective can recover without labels. A protein language model trained this way learns representations that transfer to structure and function tasks. ESM-2, scaled to 15 billion parameters, drove ESMFold, which predicts structure directly from a single sequence without a multiple sequence alignment. Its authors used it to build the ESM Metagenomic Atlas, predicting over 617 million metagenomic protein structures with more than 225 million predicted confidently.[29] ESM3, published in Science in 2025, extended the approach to a generative model over sequence, structure, and function jointly; its authors generated a fluorescent protein with 58 percent sequence identity to known fluorescent proteins.[30]
AlphaMissense adapted AlphaFold to variant interpretation, combining structural context with evolutionary conservation to classify 89 percent of possible human missense variants as likely benign or likely pathogenic without training directly on clinical labels.[31]
DNA has proved harder, because regulatory signal is sparse and long-range. Evo 2, from the Arc Institute, Stanford, and NVIDIA, was trained on about 9 trillion base pairs spanning all domains of life, at model sizes up to 40 billion parameters, using the StripedHyena 2 architecture to reach a 1 million token context at single-nucleotide resolution. Model parameters, training code, and the OpenGenome2 dataset were all released openly.[32][33] AlphaGenome, announced by Google DeepMind in June 2025 and published in Nature in January 2026, takes up to 1 million base pairs of input and predicts thousands of regulatory tracks at base resolution, including splice junctions, RNA expression, chromatin accessibility, and contact maps. The June 2025 announcement reported that it outperformed the best external models on 22 of 24 sequence prediction evaluations and matched or exceeded them on 24 of 26 variant effect prediction evaluations; the published paper puts the second figure at 25 of 26.[34][35] It is available for non-commercial research through an API.[35]
Single-cell transcriptomics attracted the same treatment. Geneformer was pretrained on roughly 30 million single-cell transcriptomes and fine-tuned for network biology tasks including therapeutic target identification in cardiomyopathy.[36] scGPT followed in Nature Methods in 2024, aimed at single-cell multi-omics.[37] NVIDIA's BioNeMo framework packages several of these, including ESM-2, Geneformer, and Evo 2, as GPU-optimized training and inference recipes.[38]
Software ecosystem
Two open source projects carry most day-to-day analysis. Biopython, first described in 2009, provides Python modules for sequence handling, file parsing, and interfaces to external tools; version 1.87 was released on 30 March 2026.[39][40] Bioconductor, described in 2004, distributes R packages for genomic data analysis on a twice-yearly release cycle. Release 3.23, dated 29 April 2026, shipped 2,418 software packages, 928 annotation packages, 437 experiment data packages, and 28 workflows against R 4.6.[41][42]
Above the libraries sit workflow managers. Nextflow, published in 2017, expresses a pipeline as dataflow between isolated processes and infers parallelism from each process's declared inputs and outputs.[43] The same script runs unchanged under SLURM, LSF, PBS, HTCondor, Kubernetes, AWS, Google Cloud, or Azure, with dependencies pinned in Docker or Singularity containers and intermediate results checkpointed so a failed run resumes from the last successful step.[44] Reproducibility is the motivating concern: a genomics pipeline typically chains a dozen tools whose versions and parameters all affect the result.
Limitations and criticism
Independent evaluation has been less flattering to the newest models than their announcements. A 2025 Genome Biology study evaluated scGPT and Geneformer in zero-shot settings across five datasets and found that they could be outperformed by simpler approaches including scVI, Harmony, and plain highly variable gene selection.[45] A 2025 Nature Methods paper compared five foundation models and two other deep learning methods on predicting transcriptome changes after genetic perturbation and reported that none outperformed simple linear baselines.[46] Both papers argue the same point: benchmark design in computational biology often fails to include the boring comparator that would reveal how much the learned model actually adds.
Agentic systems are further behind than demonstrations suggest. BixBench, a 2025 benchmark of more than 50 real bioinformatics analysis scenarios with nearly 300 open-answer questions, found frontier models reaching about 17 percent accuracy in the open-answer setting and no better than random on multiple choice.[47]
Data hygiene problems are older and more mundane. Excel silently converts some gene symbols into dates and floating-point numbers, and a 2016 survey found the resulting errors in roughly one fifth of papers carrying supplementary Excel gene lists.[48] A 2021 follow-up covering 2014 to 2020 found them in 30.9 percent of 11,117 such articles, a higher rate than before the original warning.[49]
Generative protein design also created a biosecurity gap. A 2025 Science paper showed that open source protein design software could produce variants of proteins of concern that evaded the screening tools used by nucleic acid synthesis providers; the authors developed and deployed patches before publication.[50]
Structural prediction has its own boundaries. AlphaFold-class models return a single static structure with per-residue confidence, not the conformational ensembles, allosteric transitions, or mutation-induced stability changes that much of molecular biology turns on. CASP17 runs a separate conformational ensembles category alongside single-structure targets, including targets whose conformation shifts after a single point mutation.[28]
Recent developments
Between 2025 and mid-2026 the field absorbed several results at once. Evo 2 appeared in Nature in 2026 after its 2025 preprint,[32] AlphaGenome was published in Nature in January 2026,[34] and the UK Biobank released whole-genome sequences for 490,640 participants.[17] In a collaboration between EMBL-EBI, Google DeepMind, NVIDIA, and Seoul National University, the AlphaFold database added roughly 2.2 million high-confidence homodimeric and 79,000 heterodimeric complex structures, with about 31 million predictions from the effort available for bulk download.[12]
The other visible shift is toward AI agents that run analyses rather than models that score sequences. Biomni, described in a 2025 preprint, wires a large language model to 105 biomedical software tools, 150 specialized biological tools, and 59 databases using retrieval-augmented planning and code execution.[51] Anthropic launched Claude for Life Sciences on 20 October 2025 with connectors to Benchling, BioRender, PubMed, Synapse, and 10x Genomics, and reported a Protocol QA score of 0.83 for Claude Sonnet 4.5 against a stated human baseline of 0.79.[52] The measured gap between these demonstrations and independent benchmarks such as BixBench remains wide.
See also
References
- ^Hogeweg P. "The roots of bioinformatics in theoretical biology." PLoS Computational Biology 7(3):e1002021, March 2011. journals.plos.org/...article
- ^NCBI. "GenBank and WGS Statistics" (release 272, June 2026). ncbi.nlm.nih.gov/...statistics
- ^UniProt Consortium. "UniProt release notes, release 2026_02, 10-Jun-2026." ftp.uniprot.org/...relnotes.txt
- ^RCSB Protein Data Bank. Homepage structure counts. rcsb.org
- ^Barker WC, George DG, Mewes HW, Pfeiffer F, Tsugita A. "The PIR-International databases." Nucleic Acids Research 21(13):3089-3092, 1993. pubmed.ncbi.nlm.nih.gov/8332528
- ^Mount DW. "Using PAM Matrices in Sequence Alignments." Cold Spring Harbor Protocols, 2008. PMID 21356854. pubmed.ncbi.nlm.nih.gov/21356854
- ^Needleman SB, Wunsch CD. "A general method applicable to the search for similarities in the amino acid sequence of two proteins." Journal of Molecular Biology 48(3):443-453, 1970. pubmed.ncbi.nlm.nih.gov/5420325
- ^Smith TF, Waterman MS. "Identification of common molecular subsequences." Journal of Molecular Biology 147(1):195-197, 1981. pubmed.ncbi.nlm.nih.gov/7265238
- ^Altschul SF, Gish W, Miller W, Myers EW, Lipman DJ. "Basic local alignment search tool." Journal of Molecular Biology 215(3):403-410, 1990. pubmed.ncbi.nlm.nih.gov/2231712
- ^Li H, Durbin R. "Fast and accurate short read alignment with Burrows-Wheeler transform." Bioinformatics 25(14):1754-1760, 2009. pubmed.ncbi.nlm.nih.gov/19451168
- ^van Kempen M, Kim SS, Tumescheit C, Mirdita M, Lee J, Gilchrist CLM, Soding J, Steinegger M. "Fast and accurate protein structure search with Foldseek." Nature Biotechnology 42(2):243-246, February 2024 (published online 8 May 2023). pubmed.ncbi.nlm.nih.gov/37156916
- ^AlphaFold Protein Structure Database, EMBL-EBI and Google DeepMind. alphafold.ebi.ac.uk
- ^"The 2026 Nucleic Acids Research database issue and the online molecular biology database collection" (editorial). Nucleic Acids Research 54(D1):D1. academic.oup.com/...8035161
- ^National Human Genome Research Institute. "The Cost of Sequencing a Human Genome." genome.gov/...Sequencing-Human-Genome-cost
- ^Poplin R, Chang PC, Alexander D, et al. "A universal SNP and small-indel variant caller using deep neural networks." Nature Biotechnology 36(10):983-987, 2018. pubmed.ncbi.nlm.nih.gov/30247488
- ^Nurk S, Koren S, Rhie A, Rautiainen M, et al. "The complete sequence of a human genome." Science 376(6588):44-53, 2022. pubmed.ncbi.nlm.nih.gov/35357919
- ^UK Biobank Whole-Genome Sequencing Consortium. "Whole-genome sequencing of 490,640 UK Biobank participants." Nature 645(8081):692-701, September 2025. pubmed.ncbi.nlm.nih.gov/40770095
- ^Minh BQ, Schmidt HA, Chernomor O, Schrempf D, Woodhams MD, von Haeseler A, Lanfear R. "IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era." Molecular Biology and Evolution 37(5):1530-1534, 2020. pubmed.ncbi.nlm.nih.gov/32011700
- ^Hadfield J, Megill C, Bell SM, Huddleston J, Potter B, Callender C, Sagulenko P, Bedford T, Neher RA. "Nextstrain: real-time tracking of pathogen evolution." Bioinformatics 34(23):4121-4123, 2018. pubmed.ncbi.nlm.nih.gov/29790939
- ^Protein Structure Prediction Center. "CASP16." predictioncenter.org/casp16
- ^Jumper J, Evans R, Pritzel A, et al. "Highly accurate protein structure prediction with AlphaFold." Nature 596(7873):583-589, 2021. pubmed.ncbi.nlm.nih.gov/34265844
- ^Google DeepMind. "AlphaFold." deepmind.google/...alphafold
- ^Baek M, DiMaio F, Anishchenko I, et al. "Accurate prediction of protein structures and interactions using a three-track neural network." Science 373(6557):871-876, 2021. pubmed.ncbi.nlm.nih.gov/34282049
- ^Royal Swedish Academy of Sciences. "The Nobel Prize in Chemistry 2024." kva.se/...the-nobel-prize-in-chemistry-2024
- ^Abramson J, Adler J, Dunger J, et al. "Accurate structure prediction of biomolecular interactions with AlphaFold 3." Nature 630(8016):493-500, 2024. pubmed.ncbi.nlm.nih.gov/38718835
- ^Google. "AlphaFold 3 predicts the structure and interactions of all of life's molecules." 8 May 2024, updated November 2024. blog.google/...ind-isomorphic-alphafold-3-ai-model
- ^Passaro S, Corso G, Wohlwend J, et al. "Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction." bioRxiv preprint, 18 June 2025. pubmed.ncbi.nlm.nih.gov/40667369
- ^Protein Structure Prediction Center. "CASP17." predictioncenter.org/casp17
- ^Lin Z, Akin H, Rao R, et al. "Evolutionary-scale prediction of atomic-level protein structure with a language model." Science 379(6637):1123-1130, 2023. pubmed.ncbi.nlm.nih.gov/36927031
- ^Hayes T, Rao R, Akin H, et al. "Simulating 500 million years of evolution with a language model." Science 387(6736):850-858, 2025. pubmed.ncbi.nlm.nih.gov/39818825
- ^Cheng J, Novati G, Pan J, et al. "Accurate proteome-wide missense variant effect prediction with AlphaMissense." Science 381(6664):eadg7492, 2023. pubmed.ncbi.nlm.nih.gov/37733863
- ^Brixi G, Durrant MG, Ku J, et al. "Genome modelling and design across all domains of life with Evo 2." Nature 652(8112):1349-1361, 2026. pubmed.ncbi.nlm.nih.gov/41781614
- ^Arc Institute. "Evo 2: Genome modeling and design across all domains of life." 19 February 2025. arcinstitute.org/...evo2
- ^Avsec Z, et al. "Advancing regulatory variant effect prediction with AlphaGenome." Nature 649(8099):1206-1218, January 2026. pubmed.ncbi.nlm.nih.gov/41606153
- ^Google DeepMind. "AlphaGenome: AI for better understanding the genome." 25 June 2025. deepmind.google/...better-understanding-the-genome
- ^Theodoris CV, Xiao L, Chopra A, et al. "Transfer learning enables predictions in network biology." Nature 618(7965):616-624, 2023. pubmed.ncbi.nlm.nih.gov/37258680
- ^Cui H, Wang C, Maan H, Pang K, Luo F, Duan N, Wang B. "scGPT: toward building a foundation model for single-cell multi-omics using generative AI." Nature Methods 21(8):1470-1480, 2024. pubmed.ncbi.nlm.nih.gov/38409223
- ^NVIDIA. "BioNeMo Framework documentation." docs.nvidia.com/...latest
- ^Cock PJA, Antao T, Chang JT, et al. "Biopython: freely available Python tools for computational molecular biology and bioinformatics." Bioinformatics 25(11):1422-1423, 2009. pubmed.ncbi.nlm.nih.gov/19304878
- ^Biopython project homepage (version 1.87, 30 March 2026). biopython.org
- ^Gentleman RC, Carey VJ, Bates DM, et al. "Bioconductor: open software development for computational biology and bioinformatics." Genome Biology 5(10):R80, 2004. pubmed.ncbi.nlm.nih.gov/15461798
- ^Bioconductor. "Bioconductor 3.23 Released." 29 April 2026. bioconductor.org/...bioc_3_23_release
- ^Di Tommaso P, Chatzou M, Floden EW, Prieto Barja P, Palumbo E, Notredame C. "Nextflow enables reproducible computational workflows." Nature Biotechnology 35(4):316-319, 2017. pubmed.ncbi.nlm.nih.gov/28398311
- ^Nextflow project homepage. "Nextflow: A DSL for parallel and scalable computational pipelines." nextflow.io
- ^Kedzierska KZ, Crawford L, Amini AP, Lu AX. "Zero-shot evaluation reveals limitations of single-cell foundation models." Genome Biology 26(1):101, 2025. pubmed.ncbi.nlm.nih.gov/40251685
- ^Ahlmann-Eltze C, Huber W, Anders S. "Deep-learning-based gene perturbation effect prediction does not yet outperform simple linear baselines." Nature Methods 22(8):1657-1661, August 2025. pubmed.ncbi.nlm.nih.gov/40759747
- ^Mitchener L, Laurent JM, Andonian A, et al. "BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology." arXiv:2503.00096, 28 February 2025. arxiv.org/...2503.00096
- ^Ziemann M, Eren Y, El-Osta A. "Gene name errors are widespread in the scientific literature." Genome Biology 17(1):177, 2016. pubmed.ncbi.nlm.nih.gov/27552985
- ^Abeysooriya M, Soria M, Kasu MS, Ziemann M. "Gene name errors: Lessons not learned." PLoS Computational Biology 17(7):e1008984, 2021. pubmed.ncbi.nlm.nih.gov/34329294
- ^Wittmann BJ, Alexanian T, Bartling C, et al. "Strengthening nucleic acid biosecurity screening against generative protein design tools." Science 390(6768):82-87, 2025. pubmed.ncbi.nlm.nih.gov/41037625
- ^Huang K, Zhang S, Wang H, et al. "Biomni: A General-Purpose Biomedical AI Agent." bioRxiv preprint, 2 June 2025. pubmed.ncbi.nlm.nih.gov/40501924
- ^Anthropic. "Claude for Life Sciences." 20 October 2025. anthropic.com/...claude-for-life-sciences
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
v1 · 3,672 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independent adversarial fact-check at creation (wanted175 campaign, 2026-07-24): every claim verified against primary sources by a dedicated verification agent; corrections applied before publication.
Cite this page: AI Wiki. "Bioinformatics." aiwiki.ai, updated 24 Jul 2026, fact-checked 24 Jul 2026. CC BY 4.0. https://aiwiki.ai/wiki/bioinformatics