BioReason: Incentivizing Multimodal Biological Reasoning within a DNA-LLM Model

Bo Wang (Sensetime) · Chris Maddison (University of Toronto) · Adibvafa Fallahpour (NVIDIA, Arc Institute, Vector Institute, UHN, UofT) · Andrew Magnuson (University of Toronto) · Purav Gupta (Vector Institute) · Shihao Ma (University of Toronto) · Jack Naimer (École Polytechnique Fédérale de Lausanne) · Arnav Shah (Vector Institute) · Haonan Duan (Department of Computer Science, University of Toronto) · Omar Ibrahim (University Health Network) · Hani Goodarzi (Arc Institute)
biologically coherentdeep biological reasoningdna foundation modelsgenomic datainterpretable aikegg-based disease pathwayslarge language modellogical deductionsmechanistic aimulti-step reasoningreinforcement learningsupervised fine-tuningtransparent explanationsunseen biological entitiesvariant effect prediction

Unlocking deep and interpretable biological reasoning from complex genomic data remains a major AI challenge limiting scientific progress. While current DNA foundation models excel at representing sequences, they struggle with multi-step reasoning and lack transparent, biologically meaningful explanations. BioReason addresses this by tightly integrating a DNA foundation model with a large language model (LLM), enabling the LLM to directly interpret and reason over genomic information. Through supervised fine-tuning and reinforcement learning, BioReason learns to produce logical, biologically coherent deductions. It achieves major performance gains, boosting KEGG-based disease pathway prediction accuracy from 86% to 98% and improving variant effect prediction by an average of 15% over strong baselines. BioReason can reason over unseen biological entities and explain its decisions step by step, offering a transformative framework for interpretable, mechanistic AI in biology. All data, code, and checkpoints are available at [https://github.com/bowang-lab/BioReason](https://github.com/bowang-lab/BioReason).