response quality
- CARE: Decoding-Time Safety Alignment via Rollback and Introspection Intervention
- DynamicRAG: Leveraging Outputs of Large Language Model as Feedback for Dynamic Reranking in Retrieval-Augmented Generation
- Scalable Best-of-N Selection for Large Language Models via Self-Certainty
- Transcending Cost-Quality Tradeoff in Agent Serving via Session-Awareness
- Weaver: Shrinking the Generation-Verification Gap by Scaling Compute for Verification