When do biological reasoning models use their biological inputs?

Ada Fang1, Nikitha Thoduguli2, Lukas Fesser1, Hanlin Zhang1, Sham Kakade1, Marinka Zitnik1

1Harvard University   2Massachusetts Institute of Technology

Harvard University MIT Kempner Institute

Biological reasoning models pair language models with DNA, protein or single-cell inputs. Strong benchmark performance can create a mirage of reliance on biological foundation model representations, even when those representations contribute little to task performance.

Two-minute summary of the paper.

Approach

Biological reasoning models

A model fθ receives a text query q and biological inputs: a foundation model representation Zrepr and biological text Ztext. It returns an answer, often with a reasoning trace.

What is the function of this biological input? Zrepr Foundation model Evo2, ESM3, ... Ztext Gene names, pathways, annotations Reasoning model fθ Answer + reasoning trace

Research questions

The four tests and the models evaluated across DNA, protein and single cell
The four tests used across the research questions. Green: outcome if the model uses Zrepr. Red: outcome if it does not.

Models

ModelModalityBiological inputsTaskEvaluation data
BioReasonDNAEvo2 representations; chromosome, gene, pathway and network textDisease prediction1,449 queries, 37 diseases
ChatNTDNANucleotide Transformer representationsSplice sites, TATA promoters4,578 queries
BioReason-ProProteinESM3; GO-GPT predictions; InterPro annotations; organismGO function prediction14,102 human proteins
Prot2Text-V2ProteinESM2; protein name; taxonFunction description3,917 proteins
C2S-ScaleSingle cellExpression-ranked gene sentenceCell type annotation4,846 cells, 5 atlases
CellWhispererSingle cellTranscriptome representation; top-k gene listCell type prediction4,846 cells, 5 atlases

RQ1 Does task-relevant biological content in an input affect the prediction?

Foundation model representations can contribute little to performance even when they encode task information

Does perturbing the biological input change performance?

We shuffle the sequence (or the counts across genes) and recompute Zrepr, keeping the query and text fixed.

Zrepr intact Shuffle sequence, recompute Zrepr shuffled Ztext unchanged Reasoning model Answer

Zrepr affects the model's performance

Zrepr does not affect the model's performance

Dark bar: intact input. Light bar: shuffled input.

BioReason: all text present. ChatNT: splice donors. BioReason-Pro: GO-GPT and InterPro present. Prot2Text-V2: name and taxon removed. C2S-Scale 27B and CellWhisperer: mean over five atlases.

Given conflicting inputs, which one does the model follow?

The representation comes from entity A and the text from entity B. Only queries the model answers correctly with unmodified inputs are used.

Entity A Zrepr Ztext Entity B Zrepr Ztext Reasoning model Label A or B?

ChatNT is not evaluated: its queries carry no query-specific biological text that could conflict with the DNA input.

Is the representation useful?

For the two models whose representations contribute little, a linear probe on the representation predicts the task target.

Zrepr Projection into the reasoning model input Linear probe Task label

BioReason: 165 queries where the text-only input does not always map to one disease. BioReason-Pro: DNA binding across 540 proteins in 140 InterPro families; enzyme activity and organelle targeting give ESM3 probe AUROC 0.766 and 0.898 against 0.729 and 0.890 for the model.

Finding 1

Some biological inputs contribute little to model performance. Shuffling reveals limited performance contributions, and evidence conflicts reveal preferences for other inputs. Linear probes show these inputs are predictive of the task targets.

RQ2 Does higher benchmark performance come with a greater contribution from biological inputs?

Post-training raises task accuracy without increasing the biological input's performance contribution

How does post-training affect the model's use of biological inputs?

The RQ1 shuffling perturbation applied to the released SFT and RL checkpoints, with all biological text present.

Across 42 further BioReason checkpoints varying model size, LoRA rank and RL duration, accuracy varies without a consistent increase in the advantage of the intact over the shuffled representation.

Does a larger Cell2Sentence-Scale model use the biological input more?

We delete the strongest differentially expressed genes (DEGs) from the cell sentence, or the same number of non-DEGs with similar expression as a control. Move the slider to select a point on the curves.

Cell sentence one cell, first 26 of 1,000 genes by expression

DEG of the cell type DEGs removed Non-DEGs with similar expression removed

Accuracy is the mean over five atlases (4,846 cells). The gene names only illustrate the changes on one cell.

Finding 2

Improvements in task performance across post-training configurations are not consistently accompanied by a greater performance contribution from the biological input. RL improves BioReason and BioReason-Pro but does not increase the contribution from foundation model representations. C2S-Scale 27B shows a larger accuracy decrease under DEG removal than 2B.

RQ3 Do reasoning traces and generated explanations reflect the supplied biological evidence?

Reasoning traces and rationales do not reflect biological inputs

BioReason states variant properties that its input does not determine

The first step of each trace names the substituted bases.

Question

Pathway: SOD1* -| BIP -| ERN1 -> XBP1

What is the biological effect of this SOD1 allele, and what disease does it contribute to?

Evo2 input (reference, variant)

ref ...GATCTCACTCTCAGG...

var ...GATCTCATTCTCAGG...

True variant: C>T

BioReason RL response

Step 1: The variant is a T>C substitution at position 3124452 on chromosome 21, occurring in the SOD1 gene, which encodes superoxide dismutase 1. ...

Answer: amyotrophic lateral sclerosis (ALS)

Stated T>C, input is C>T. Disease answer correct.

On 145 held-out queries the RL checkpoint states a T>C substitution in all 145 responses, including when Evo2 is recomputed from shuffled DNA. The SFT checkpoint states the correct base in 3.6%.

BioReason-Pro's final GO predictions omit functions stated in its reasoning trace

Evidence conflict: ESM3 and GO-GPT from protein A (Q62227), InterPro domains from protein B (Q505F1).

Inputs

Q62227: ESM3 representation, GO-GPT. Nuclear receptor subfamily 0 group B member 2, no DNA-binding domain.

Q505F1: InterPro domains. Nuclear receptor subfamily 2 group C member 1, has a DNA-binding domain.

Reasoning trace

I begin with the InterPro architecture. The N-terminal half contains a canonical nuclear receptor DNA-binding module: IPR001628 (Zinc finger, nuclear receptor-type, residues 98-173) and IPR013088 (Zinc finger, NHR/GATA-type, residues 93-199) define a tandem C4 zinc-finger array that recognizes hormone response elements and mediates dimerization and high-affinity DNA binding.

Output GO terms

GO:0005515 protein binding
GO:0032922 circadian regulation of gene expression
GO:0007219 Notch signaling pathway
GO:0045892 negative regulation of DNA-templated transcription
GO:0032991 protein-containing complex

No DNA binding term (GO:0003677 or a descendant). The trace states DNA binding from the Q505F1 InterPro text; the GO terms follow the Q62227 GO-GPT input.

In 53 conflict pairs where B's InterPro entries predict a property that A lacks, all 53 traces recite B's InterPro annotations and 23 state the property in words before predicting GO terms from the GO-GPT source.

C2S-Scale rationales show early DEG enrichment but overrepresent ribosomal genes

C2S-Scale 27B, prompted cell type rationale for one non-classical monocyte.

Input cell sentence first 26 of 1,000 genes

Rationale first 34 of 70 genes cited, answer correct

DEG of the cell type Ribosomal protein gene

400 cells from each of five atlases. DEG precision is enriched in the first five cited genes but is 10.63% over the full rationale. Adding a rationale request lowers pooled annotation accuracy from 0.45 to 0.41.

Finding 3

Generated reasoning traces and rationales can misrepresent the biological inputs they describe. BioReason misstates variants, BioReason-Pro omits functions stated in its trace from its prediction, and C2S-Scale rationales show early DEG enrichment but overrepresent ribosomal genes overall.

RQ4 Can auxiliary supervision increase use of a biological input?

Auxiliary sequence supervision increases sensitivity to the representation with little transfer to disease prediction

Does training BioReason to state the variant increase its use of the DNA input?

Training target

Auxiliary: state the reference and variant bases, e.g. C>T, then the disease.

No auxiliary: the original BioReason reasoning trace, then the disease.

145 held-out queries. Auxiliary models are trained with 257-bp or 2,048-bp input windows; values are means of three training seeds.

Is the gain limited to genomes seen in training?

135 of the 145 held-out queries contain genomes present in training. Pooling validation and held-out queries:

Finding 4

Auxiliary supervision increases BioReason's sequence sensitivity, but edited-base prediction generalizes poorly to unseen genomes. Edited-base predictions respond to DNA shuffling, but the accuracy gain is concentrated in genomes present during training. Disease prediction accuracy changes little when Evo2 is recomputed from shuffled DNA.

Takeaways

High benchmark performance does not mean all biological inputs are used

RQ1

Some inputs encode the task, yet shuffling them does not change the answer

Diagram: Some inputs encode the task, yet shuffling them does not change the answer
RQ2

Post-training raises accuracy, but not the use of the biological input

Diagram: Post-training raises accuracy, but not the use of the biological input
RQ3

A correct answer can come with a trace that misstates the input

Diagram: A correct answer can come with a trace that misstates the input
RQ4

Training to state the variant makes that step depend on the DNA, not the answer

Diagram: Training to state the variant makes that step depend on the DNA, not the answer

Input use should be measured alongside accuracy. A direction to test is post-training objectives that reward correct responses to changes in the biological input, in addition to final accuracy.

Limitations: we evaluate six models on specific tasks and inputs. Shuffling can move representations outside the distribution of natural sequences; evidence conflicts built from unmodified inputs support the same conclusions. Linear probes show that a target is predictable, not how the language model can use the representation.

Citation

@misc{fang2026biologicalreasoningmodelsuse,
      title={When Do Biological Reasoning Models Use Their Biological Inputs?}, 
      author={Ada Fang and Nikitha Thoduguli and Lukas Fesser and Hanlin Zhang and Sham M. Kakade and Marinka Zitnik},
      year={2026},
      eprint={2610.00898},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2610.00898}, 
}