Learning to Replicate Expert Judgment in Financial Tasks
fine-tuningfinancial-aillm-evaluationdifferentiated-intelligencethinking-machines
Abstraction: Fine-tuned LLM beats frontier models at financial info triage
Key points:
- Thinking Machines Lab (2026-06-30) trained a proprietary model to do investor "information triage" — filtering/interpreting/segmenting financial documents to surface investment-relevant signal.
- On six real information-filtering tasks, frontier models (Gemini, Claude, GPT variants) averaged only ~50% accuracy with a plain prompt; strong expert-written prompts + reframing (e.g. 3-way relevant-and-interesting / relevant-but-uninteresting / irrelevant) lifted them to mid-70s, still under the ~80% bar investors trust.
- Automatic prompt-optimization gave no further gains; newer models barely improve per dollar (GPT 5.4 costs 43% more than 5.2 for marginal accuracy).
- With high-quality human annotations, their fine-tuned model outperformed all tested frontier models on accuracy and recall at a fraction of the cost.
- Frames a vision of "differentiated intelligence" — smaller models tuned to a specific organization's judgment/taste rather than general frontier scale.
Connections: Thinking Machines Lab · GPT-5 · Claude · Gemini · Fine Tuning · Large Language Models · LLM Evaluation · Prompt Engineering
Source: https://thinkingmachines.ai/news/learning-to-replicate-expert-judgment-in-financial-tasks/