When do biological reasoning models use their biological inputs?
1Harvard University 2Massachusetts Institute of Technology
Biological reasoning models pair language models with DNA, protein or single-cell inputs. Strong benchmark performance can create a mirage of reliance on biological foundation model representations, even when those representations contribute little to task performance.
Approach
Biological reasoning models
A model fθ receives a text query q and biological inputs: a foundation model representation Zrepr and biological text Ztext. It returns an answer, often with a reasoning trace.
Research questions
Does task-relevant biological content in an input affect the prediction?
RQ2Does higher benchmark performance come with a greater contribution from biological inputs?
RQ3Do reasoning traces and generated explanations reflect the supplied biological evidence?
RQ4Can auxiliary supervision increase use of a biological input?
Models
| Model | Modality | Biological inputs | Task | Evaluation data |
|---|---|---|---|---|
| BioReason | DNA | Evo2 representations; chromosome, gene, pathway and network text | Disease prediction | 1,449 queries, 37 diseases |
| ChatNT | DNA | Nucleotide Transformer representations | Splice sites, TATA promoters | 4,578 queries |
| BioReason-Pro | Protein | ESM3; GO-GPT predictions; InterPro annotations; organism | GO function prediction | 14,102 human proteins |
| Prot2Text-V2 | Protein | ESM2; protein name; taxon | Function description | 3,917 proteins |
| C2S-Scale | Single cell | Expression-ranked gene sentence | Cell type annotation | 4,846 cells, 5 atlases |
| CellWhisperer | Single cell | Transcriptome representation; top-k gene list | Cell type prediction | 4,846 cells, 5 atlases |
RQ1 Does task-relevant biological content in an input affect the prediction?
Foundation model representations can contribute little to performance even when they encode task information
Does perturbing the biological input change performance?
We shuffle the sequence (or the counts across genes) and recompute Zrepr, keeping the query and text fixed.
Zrepr affects the model's performance
Zrepr does not affect the model's performance
Dark bar: intact input. Light bar: shuffled input.
BioReason: all text present. ChatNT: splice donors. BioReason-Pro: GO-GPT and InterPro present. Prot2Text-V2: name and taxon removed. C2S-Scale 27B and CellWhisperer: mean over five atlases.
Given conflicting inputs, which one does the model follow?
The representation comes from entity A and the text from entity B. Only queries the model answers correctly with unmodified inputs are used.
ChatNT is not evaluated: its queries carry no query-specific biological text that could conflict with the DNA input.
Is the representation useful?
For the two models whose representations contribute little, a linear probe on the representation predicts the task target.
BioReason: 165 queries where the text-only input does not always map to one disease. BioReason-Pro: DNA binding across 540 proteins in 140 InterPro families; enzyme activity and organelle targeting give ESM3 probe AUROC 0.766 and 0.898 against 0.729 and 0.890 for the model.
Finding 1
Some biological inputs contribute little to model performance. Shuffling reveals limited performance contributions, and evidence conflicts reveal preferences for other inputs. Linear probes show these inputs are predictive of the task targets.
RQ2 Does higher benchmark performance come with a greater contribution from biological inputs?
Post-training raises task accuracy without increasing the biological input's performance contribution
How does post-training affect the model's use of biological inputs?
The RQ1 shuffling perturbation applied to the released SFT and RL checkpoints, with all biological text present.
Across 42 further BioReason checkpoints varying model size, LoRA rank and RL duration, accuracy varies without a consistent increase in the advantage of the intact over the shuffled representation.
Does a larger Cell2Sentence-Scale model use the biological input more?
We delete the strongest differentially expressed genes (DEGs) from the cell sentence, or the same number of non-DEGs with similar expression as a control. Move the slider to select a point on the curves.
Cell sentence one cell, first 26 of 1,000 genes by expression
Accuracy is the mean over five atlases (4,846 cells). The gene names only illustrate the changes on one cell.
Finding 2
Improvements in task performance across post-training configurations are not consistently accompanied by a greater performance contribution from the biological input. RL improves BioReason and BioReason-Pro but does not increase the contribution from foundation model representations. C2S-Scale 27B shows a larger accuracy decrease under DEG removal than 2B.
RQ3 Do reasoning traces and generated explanations reflect the supplied biological evidence?
Reasoning traces and rationales do not reflect biological inputs
BioReason states variant properties that its input does not determine
The first step of each trace names the substituted bases.
Question
Pathway: SOD1* -| BIP -| ERN1 -> XBP1
What is the biological effect of this SOD1 allele, and what disease does it contribute to?
Evo2 input (reference, variant)
ref ...GATCTCACTCTCAGG...
var ...GATCTCATTCTCAGG...
True variant: C>T
BioReason RL response
Step 1: The variant is a T>C substitution at position 3124452 on chromosome 21, occurring in the SOD1 gene, which encodes superoxide dismutase 1. ...
Answer: amyotrophic lateral sclerosis (ALS)
Stated T>C, input is C>T. Disease answer correct.
On 145 held-out queries the RL checkpoint states a T>C substitution in all 145 responses, including when Evo2 is recomputed from shuffled DNA. The SFT checkpoint states the correct base in 3.6%.
BioReason-Pro's final GO predictions omit functions stated in its reasoning trace
Evidence conflict: ESM3 and GO-GPT from protein A (Q62227), InterPro domains from protein B (Q505F1).
Inputs
Q62227: ESM3 representation, GO-GPT. Nuclear receptor subfamily 0 group B member 2, no DNA-binding domain.
Q505F1: InterPro domains. Nuclear receptor subfamily 2 group C member 1, has a DNA-binding domain.
Reasoning trace
I begin with the InterPro architecture. The N-terminal half contains a canonical nuclear receptor DNA-binding module: IPR001628 (Zinc finger, nuclear receptor-type, residues 98-173) and IPR013088 (Zinc finger, NHR/GATA-type, residues 93-199) define a tandem C4 zinc-finger array that recognizes hormone response elements and mediates dimerization and high-affinity DNA binding.
Output GO terms
GO:0005515 protein binding
GO:0032922 circadian regulation of gene expression
GO:0007219 Notch signaling pathway
GO:0045892 negative regulation of DNA-templated transcription
GO:0032991 protein-containing complex
No DNA binding term (GO:0003677 or a descendant). The trace states DNA binding from the Q505F1 InterPro text; the GO terms follow the Q62227 GO-GPT input.
In 53 conflict pairs where B's InterPro entries predict a property that A lacks, all 53 traces recite B's InterPro annotations and 23 state the property in words before predicting GO terms from the GO-GPT source.
C2S-Scale rationales show early DEG enrichment but overrepresent ribosomal genes
C2S-Scale 27B, prompted cell type rationale for one non-classical monocyte.
Input cell sentence first 26 of 1,000 genes
Rationale first 34 of 70 genes cited, answer correct
400 cells from each of five atlases. DEG precision is enriched in the first five cited genes but is 10.63% over the full rationale. Adding a rationale request lowers pooled annotation accuracy from 0.45 to 0.41.
Finding 3
Generated reasoning traces and rationales can misrepresent the biological inputs they describe. BioReason misstates variants, BioReason-Pro omits functions stated in its trace from its prediction, and C2S-Scale rationales show early DEG enrichment but overrepresent ribosomal genes overall.
RQ4 Can auxiliary supervision increase use of a biological input?
Auxiliary sequence supervision increases sensitivity to the representation with little transfer to disease prediction
Does training BioReason to state the variant increase its use of the DNA input?
Training target
Auxiliary: state the reference and variant bases, e.g. C>T, then the disease.
No auxiliary: the original BioReason reasoning trace, then the disease.
145 held-out queries. Auxiliary models are trained with 257-bp or 2,048-bp input windows; values are means of three training seeds.
Is the gain limited to genomes seen in training?
135 of the 145 held-out queries contain genomes present in training. Pooling validation and held-out queries:
Finding 4
Auxiliary supervision increases BioReason's sequence sensitivity, but edited-base prediction generalizes poorly to unseen genomes. Edited-base predictions respond to DNA shuffling, but the accuracy gain is concentrated in genomes present during training. Disease prediction accuracy changes little when Evo2 is recomputed from shuffled DNA.
Takeaways
High benchmark performance does not mean all biological inputs are used
Some inputs encode the task, yet shuffling them does not change the answer
Post-training raises accuracy, but not the use of the biological input
A correct answer can come with a trace that misstates the input
Training to state the variant makes that step depend on the DNA, not the answer
Input use should be measured alongside accuracy. A direction to test is post-training objectives that reward correct responses to changes in the biological input, in addition to final accuracy.
Limitations: we evaluate six models on specific tasks and inputs. Shuffling can move representations outside the distribution of natural sequences; evidence conflicts built from unmodified inputs support the same conclusions. Linear probes show that a target is predictable, not how the language model can use the representation.
Citation
@misc{fang2026biologicalreasoningmodelsuse,
title={When Do Biological Reasoning Models Use Their Biological Inputs?},
author={Ada Fang and Nikitha Thoduguli and Lukas Fesser and Hanlin Zhang and Sham M. Kakade and Marinka Zitnik},
year={2026},
eprint={2610.00898},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2610.00898},
}