If your model of biological AI development assumes that scale remains the dominant performance driver, that the trajectory toward larger parameter counts and larger training corpora continues to produce the strongest models in protein language modelling and genomic foundation modelling, the 2025 benchmark data shows the assumption has broken. In 2024, the largest protein language model was the 98-billion-parameter ESM3. In 2025, MSAPairformer (111 million parameters, three orders of magnitude smaller) surpassed previous state-of-the-art methods on ProteinGym at a fraction of the training and parameter budget. In genomics, GPN-Star (200 million parameters) outperformed Evo 2 (40 billion parameters) on multiple variant effect prediction tasks. The scaling assumption that produced ESM3 and Evo 2 has been overtaken by a different research direction, and the new direction is producing better results.
The protein language model size trajectory is one of the clearest data signals in the biological AI literature. The progression from 2020 to 2024:
- 2020: ProGen 1.2B parameters; ProtBert 0.42B
- 2022: ProGen2 6.4B; ProtT5 1.2B; ESM2 15B
- 2023: OpenCRISPR-1 (proseLM) 6.4B
- 2024: ESM3 98B (the largest of the trajectory)
The 2025 inflection: SHIVER 0.04B (40 million parameters); PoET-2 0.18B; Proteina 0.40B; ProGen3 46B; MSA-Pairformer 0.11B (111 million parameters); E1 0.60B; MSA-Transformer 0.10B. The 2025 cohort includes some models at the prior parameter scale (ProGen3 at 46B) but the strongest performance results come from the smaller models, particularly MSAPairformer and the Profluent E1 series, both of which substantially outperformed the larger predecessors on ProteinGym, the comprehensive benchmark for protein fitness prediction and design.
The ProteinGym benchmark results show the performance inversion clearly:
- MSA-Transformer (100M): 0.45 Spearman correlation
- ESM-2 (150M): 0.39
- ESM-2 (3B): 0.41 (larger but only marginally better)
- ESM-2 (650M): 0.41
- ESMC (300M): 0.41
- E1 (Single Sequence) (600M): 0.42
- MSAPairformer (111M): 0.45, equal to MSA-Transformer despite being smaller
- E1 (Retrieval Augmented) (600M): 0.48, the highest
The 2025 winners are MSAPairformer (small, no retrieval) and E1 with retrieval augmentation. Both achieve their performance through training method and architecture innovation rather than parameter count. The scaling-only approach (ESM-2 3B vs ESM-2 150M producing similar ProteinGym scores) has clear diminishing returns in this domain.
The trajectory: protein language model development has shifted from "make it bigger" to "train it better with better data and methods." The implication for the biological AI research community is substantial. Compute resources can be redirected to method innovation rather than parameter scaling. Training data curation matters more than training data volume. Specialised models can compete with or exceed generalist models in specific tasks.
The genomic foundation model results echo the pattern. The benchmark comparison:
- Enformer (250M): 0.38 AUPRC
- Borzoi (190M): 0.40
- Evo 2 (40B): 0.53, substantial parameter count but modest improvement over Borzoi
- GPN-Star (200M): 0.75, smallest of the four, highest performance
GPN-Star, focused on functional and regulatory genomics, outperformed Evo 2 (which has 200x the parameters) on multiple variant effect prediction tasks. The trajectory observation: scale alone is not sufficient in genomic foundation modelling. Training method and data curation are the determining factors.
Three structural drivers explain the shift away from scale dominance in biological AI.
The first driver: data is the bottleneck, not compute. Biological data is expensive to generate, fragmented across institutions, and limited in volume compared to natural language data. Cofolding models now represent all structure types in the Protein Data Bank, meaning the model can train on essentially all of the experimentally available structural data. Further performance improvement requires new data sources (distilled datasets of AI-predicted structures, combined training across structural and binding data) rather than larger models trained on the same data.
The second driver: domain specialisation pays. The ESM-C series demonstrated that smaller models geared toward a single task (representation learning) could be successful without the full feature set of the ESM3 family. Task-specific training and architecture choices produce better task-specific performance than general-purpose foundation models. The "one big model for everything" approach that has dominated general-language modelling has been less successful in biology.
The third driver: retrieval augmentation is highly effective in biological contexts. The Profluent E1 series achieves its strongest performance with retrieval-augmented generation, where the model accesses external biological data sources during inference rather than encoding all knowledge in parameters. The approach is well-suited to biology where the underlying data (protein structures, sequence databases, experimental measurements) is highly structured and growing rapidly. RAG architectures benefit from the structured data more than general-purpose architectures.
The trajectory observation: the future of biological AI is plausibly characterised by smaller, more specialised, retrieval-augmented models trained on curated multi-source datasets, not by larger general-purpose foundation models. The architectural and methodological direction has changed in a structurally meaningful way in 2025.
Three implications follow for organisations and researchers operating in biological AI in 2026.
The first implication: compute investment strategy needs to shift. Funding plans that anticipated continued scaling (100B+ parameter models becoming standard in biological AI) should be reconsidered. The competitive frontier has moved toward training method innovation, data curation, and retrieval architectures. Compute resources may be better deployed in these directions than in larger model training.
The second implication: smaller biotech firms have a more competitive position than the scaling trajectory suggested. Firms that cannot afford 100B-parameter model training can still compete on protein language modelling and genomic foundation modelling through method innovation and data curation. The barrier to entry on capability has lowered substantially as the scaling assumption broke.
The third implication: data infrastructure investment is the high-leverage strategic move. The PDB's 50-year history of structural data accumulation, the recent Tahoe-100M single-cell sequencing dataset, the BaseData 9.8 billion gene metagenomic dataset, and the various distilled datasets from AlphaFold and successor models all represent training data that produces capability advantages. Organisations that invest in data infrastructure (curation, quality, integration) will outperform organisations that invest only in model training.
The trajectory: through 2026-2028, biological AI is likely to continue this direction. Smaller, specialised, well-trained models will likely continue to set the performance frontier. Larger general-purpose models will continue to be developed but their performance advantage will be narrower than scale alone would suggest. The strategic implications for research funding, biotech investment, and biological AI deployment all follow from this trajectory. Plans that engage with the actual data observation (that scale is no longer the dominant performance driver in biological AI) will be aligned with the visible direction of the field.
Discussion