Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content

  • 2026-09-27 23:01:30
  • Shaz Furniturewala, Arkaitz Zubiaga
  • 0

Abstract

Adversarial perturbations can reduce state-of-the-art toxicity classifiers to near-zero accuracy, yet existing defences treat models as black boxes. We apply mechanistic interpretability to toxicity classification for the first time, identifying the internal attention-head circuits responsible for both correct classification and adversarial vulnerability. Across a 2$\times$2 factorial study (BERT $\times$ RoBERTa) $\times$ (Jigsaw $\times$ ToxiGen), extended to Llama Guard~2 (8B), we show that zeroing a single attention head recovers up to 70.4 pp of adversarial accuracy for RoBERTa on Jigsaw and 37.3 pp for Llama Guard 2 on ToxiGen, at $\leq$0.6 pp clean cost. Vulnerable heads generalise to held-out examples within $\leq$1 pp, and a class-imbalance sweep confirms they act as selective toxic-class detectors. Head suppression matches or outperforms adversarial training on Jigsaw; data augmentation dominates on ToxiGen: a dataset-specific reversal explained by whether the classifier encodes a concentrated bottleneck or a distributed circuit. Demographic analysis across 20 Jigsaw and 13 ToxiGen minority groups reveals structurally unequal adversarial vulnerability, exposing mechanistically traceable fairness gaps in current toxicity classifiers.

 

Quick Read (beta)

loading the full paper ...