experimental evaluations
A systematic approach to assess the performance and effectiveness of AI models or methods through controlled experiments, typically involving comparisons against baselines.
- Alleviating Hallucinations in Large Language Models through Multi-Model Contrastive Decoding and Dynamic Hallucination Detection
- ComPO: Preference Alignment via Comparison Oracles
- Enhancing 3D Reconstruction for Dynamic Scenes
- FairDD: Fair Dataset Distillation
- Learning to Watermark: A Selective Watermarking Framework for Large Language Models via Multi-Objective Optimization
- LoRASuite: Efficient LoRA Adaptation Across Large Language Model Upgrades
- Multi-Objective Hyperparameter Selection via Hypothesis Testing on Reliability Graphs
- Robust Federated Finetuning of LLMs via Alternating Optimization of LoRA
- Safely Learning Controlled Stochastic Dynamics
- Smooth and Flexible Camera Movement Synthesis via Temporal Masked Generative Modeling