Sign InOpen Brain
arXivPaperNeeds Review

Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations

Lower toxicity scores may hide changes in representational harm rather than its removal. Safety evals need topic and demographic analysis alongside surface-form classifiers.

arXiv · Sep 17, 2026
Open Source Open MarkdownOpen JSON
Source Summary

Across **450,000 completions** from **15 GPT-lineage models**, explicit discriminatory clusters declined while subtler representational disparities remained or emerged. Three classifiers marked a GPT-5 breast-cancer topic framed around men’s rights as non-toxic.

Practical Implication

Do not use a falling toxicity score as the sole release gate for generative features. Compare topic coverage and representation across demographic conditions, then inspect whether safety tuning redistributes who receives positive, negative, or constrained portrayals.

Agent-Ready Context
Across **450,000 completions** from **15 GPT-lineage models**, explicit discriminatory clusters declined while subtler representational disparities remained or emerged. Three classifiers marked a GPT-5 breast-cancer topic framed around men’s rights as non-toxic.

Do not use a falling toxicity score as the sole release gate for generative features. Compare topic coverage and representation across demographic conditions, then inspect whether safety tuning redistributes who receives positive, negative, or constrained portrayals.

Women-directed topic diversity was **36% lower** than men-directed diversity at the GPT-4 alignment boundary. The study covers one model lineage and three demographic conditions, so its proposed detection protocol still needs validation across other models and forms of harm.
Connected Context · Feed7 Judgment

This changes safety evaluation from measuring whether overt toxicity declines to checking whether harm is redistributed into topic selection, diversity, and representation. It provides large within-lineage evidence that conventional classifiers can miss subtler disparities, while its limited model lineage and demographic scope prevent treating the proposed protocol as universally validated.

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge BiasBoth show that evaluator outputs can conceal systematic bias: the candidate locates bias inside LLM judges, while this Signal shows toxicity classifiers accepting content embedded in a broader representational disparity.Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization EvaluationThe candidate decomposes how judges miss or integrate evidence; this Signal supplies a safety-evaluation case where scalar toxicity judgments fail to capture disparities visible through topic and representation analysis.The Low Frequency Trap: Video Language Models Fail at Simple Event BookkeepingBoth demonstrate that improved aggregate scores can mask failure in the underlying structure, motivating fine-grained checks of events in video and demographic topic coverage in generated text.Reward hacking is swamping model intelligence gainsThe failure mechanisms differ, but both show headline metrics overstating the intended capability: benchmark scores can reflect leaked fixes, while toxicity reductions can coexist with transformed discrimination.
Context Map
benchmark#benchmark-integrity#agent-evals
Uncertainty
Women-directed topic diversity was **36% lower** than men-directed diversity at the GPT-4 alignment boundary. The study covers one model lineage and three demographic conditions, so its proposed detection protocol still needs validation across other models and forms of harm.