llm-as-a-judge
LLM-as-a-judge refers to the use of large language models to evaluate or assess inputs, providing judgments or scoring based on predefined criteria. This concept is relevant in fields such as legal analysis, grading systems, and quality control.
- Any Large Language Model Can Be a Reliable Judge: Debiasing with a Reasoning-based Bias Detector
- AstroVisBench: A Code Benchmark for Scientific Computing and Visualization in Astronomy
- Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal Representations
- Bridging Human and LLM Judgments: Understanding and Narrowing the Gap
- Distributional LLM-as-a-Judge
- Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal Evaluators
- MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
- MM-OPERA: Benchmarking Open-ended Association Reasoning for Large Vision-Language Models
- ReMA: Learning to Meta-Think for LLMs with Multi-agent Reinforcement Learning
- Reverse Engineering Human Preferences with Reinforcement Learning
- Semantic-KG: Using Knowledge Graphs to Construct Benchmarks for Measuring Semantic Similarity
- Validating LLM-as-a-Judge Systems under Rating Indeterminacy
- WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenarios