OVERT: A Benchmark for Over-Refusal Evaluation on Text-to-Image Models
benign promptscustomized evaluation datafaithfulnessgeneration frameworkharmful promptsover-refusalovertprompt rewritingsafety alignmentsafety-related categoriessafety–utility trade-offsynthetic evaluation datat2i modelstext-to-imageuser-defined policies
Text-to-Image (T2I) models have achieved remarkable success in generating visual content from text inputs. Although multiple safety alignment strategies have been proposed to prevent harmful outputs, they often lead to overly cautious behavior