robustness assessment
- Fix False Transparency by Noise Guided Splatting
- MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
- Transferable Black-Box One-Shot Forging of Watermarks via Image Preference Models
- Worse than Zero-shot? A Fact-Checking Dataset for Evaluating the Robustness of RAG Against Misleading Retrievals