response generation
- AdvPrefix: An Objective for Nuanced LLM Jailbreaks
- Deliberation on Priors: Trustworthy Reasoning of Large Language Models on Knowledge Graphs
- Explainable Reinforcement Learning from Human Feedback to Improve Alignment
- Leveraging robust optimization for llm alignment under distribution shifts
- Lookahead Routing for Large Language Models