Tx-LLM: Supporting therapeutic development with large language models
drug-discoverybioinformaticsllmgoogle-deepmind
Abstraction: PaLM-2 fine-tuned on 66 drug discovery tasks across full pipeline
Key points:
- Tx-LLM is fine-tuned from PaLM-2 on 66 datasets from the Therapeutic Data Commons (TDC), spanning target identification to clinical trial approval
- Achieves competitive performance on 43 of 66 tasks and exceeds state-of-the-art specialist models on 22; drug development typically costs $1-2B and 10-15 years
- Particularly strong on tasks combining small molecules with textual data (e.g., predicting clinical trial approval given drug + disease name)
- Positive transfer observed between protein and small-molecule tasks despite their differences — multi-task training helps across domains
- Predictions binned into integers 0-1000 to stabilize numeric output, addressing known LLM struggles with math
- Current limitations: not instruction-tuned for natural language explanations; integration with Gemini family planned
Connections: Google · Google Deepmind · Palm 2 · Large Language Models · Fine Tuning
Source: https://research.google/blog/tx-llm-supporting-therapeutic-development-with-large-language-models/