instruction-tuned models
- AtmosSci-Bench: Evaluating the Recent Advance of Large Language Model for Atmospheric Science
- Beyond Oracle: Verifier-Supervision for Instruction Hierarchy in Reasoning and Instruction-Tuned LLMs
- ComPO: Preference Alignment via Comparison Oracles
- Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model