SPQR: A Standardized Benchmark for Modern Safety Alignment Methods in Text-to-Image Diffusion Models

2511.19558v1 cs.CR, cs.AI, cs.CV, cs.LG 2025-11-26

Авторы:

Mohammed Talha Alam, Nada Saadi, Fahad Shamshad, Nils Lukas, Karthik Nandakumar, Fahkri Karray, Samuele Poppi

Abstract

Text-to-image diffusion models can emit copyrighted, unsafe, or private content. Safety alignment aims to suppress specific concepts, yet evaluations seldom test whether safety persists under benign downstream fine-tuning routinely applied after deployment (e.g., LoRA personalization, style/domain adapters). We study the stability of current safety methods under benign fine-tuning and observe frequent breakdowns. As true safety alignment must withstand even benign post-deployment adaptations, we introduce the SPQR benchmark (Safety-Prompt adherence-Quality-Robustness). SPQR is a single-scored metric that provides a standardized and reproducible framework to evaluate how well safety-aligned diffusion models preserve safety, utility, and robustness under benign fine-tuning, by reporting a single leaderboard score to facilitate comparisons. We conduct multilingual, domain-specific, and out-of-distribution analyses, along with category-wise breakdowns, to identify when safety alignment fails after benign fine-tuning, ultimately showcasing SPQR as a concise yet comprehensive benchmark for T2I safety alignment techniques for T2I models.

Ссылки и действия

Читать на arXiv Скачать PDF

Дополнительные ресурсы:

SPQR: A Standardized Benchmark for Modern Safety Alignment Methods in Text-to-Image Diffusion Models

Авторы:

Abstract

Ссылки и действия

Связанные статьи

GEO-Detective: Unveiling Location Privacy Risks in Images with LLM Agents

BackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine-...

EditTrack: Detecting and Attributing AI-assisted Image Editing

Taught Well Learned Ill: Towards Distillation-conditional Backdoor Attack

Of-SemWat: High-payload text embedding for semantic watermarking of AI-generated...

Навигация