Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations
Lower toxicity scores may hide changes in representational harm rather than its removal. Safety evals need topic and demographic analysis alongside surface-form classifiers.
Across **450,000 completions** from **15 GPT-lineage models**, explicit discriminatory clusters declined while subtler representational disparities remained or emerged. Three classifiers marked a GPT-5 breast-cancer topic framed around men’s rights as non-toxic.
Do not use a falling toxicity score as the sole release gate for generative features. Compare topic coverage and representation across demographic conditions, then inspect whether safety tuning redistributes who receives positive, negative, or constrained portrayals.
Across **450,000 completions** from **15 GPT-lineage models**, explicit discriminatory clusters declined while subtler representational disparities remained or emerged. Three classifiers marked a GPT-5 breast-cancer topic framed around men’s rights as non-toxic. Do not use a falling toxicity score as the sole release gate for generative features. Compare topic coverage and representation across demographic conditions, then inspect whether safety tuning redistributes who receives positive, negative, or constrained portrayals. Women-directed topic diversity was **36% lower** than men-directed diversity at the GPT-4 alignment boundary. The study covers one model lineage and three demographic conditions, so its proposed detection protocol still needs validation across other models and forms of harm.
This changes safety evaluation from measuring whether overt toxicity declines to checking whether harm is redistributed into topic selection, diversity, and representation. It provides large within-lineage evidence that conventional classifiers can miss subtler disparities, while its limited model lineage and demographic scope prevent treating the proposed protocol as universally validated.