detectability
- $\mathcal{X}^2$-DFD: A framework for e$\mathcal{X}$plainable and e$\mathcal{X}$tendable Deepfake Detection
- Learning to Watermark: A Selective Watermarking Framework for Large Language Models via Multi-Objective Optimization
- Multi-agent KTO: Enhancing Strategic Interactions of Large Language Model in Language Game
- SNEAKDOOR: Stealthy Backdoor Attacks against Distribution Matching-based Dataset Condensation
- Shallow Diffuse: Robust and Invisible Watermarking through Low-Dim Subspaces in Diffusion Models
- Stealthy Yet Effective: Distribution-Preserving Backdoor Attacks on Graph Classification
- Theoretically Grounded Framework for LLM Watermarking: A Distribution-Adaptive Approach