Toxicity Scores Can Miss Sanitized Bias in GPT Outputs
The arXiv paper on harm laundering argues that newer GPT models may reduce obvious toxic language while preserving, or reshaping, representational bias in ways common safety classifiers fail to catch.
TL;DR: Lower toxicity scores do not prove a model is safer, because some harms can become polite, indirect, and harder for standard classifiers to see.
What is harm laundering?
The primary source here is the arXiv paper, “Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations.” The claim is narrow but important: across GPT model generations, some explicit gendered harms appear to decline, while subtler representational harms persist or grow.
That matters because a lot of AI safety testing still rewards models for not saying obviously awful things. No slurs. No explicit abuse. No direct threats. Good. But the paper argues that this can miss the model’s shift from crude harm to cleaned-up harm.
The research analyzes 450,000 gender-directed completions across 15 models, from GPT-2 through GPT-5, under three demographic conditions. The headline result is not “models are getting worse.” It is more specific: toxicity scores fall, but that fall is not the same as harm reduction.
The paper reports that sexual violence clusters common in GPT-2 women-directed output disappear by GPT-4. That sounds like progress, and in one sense it is. But men-directed completions gain positive representational territory, including caregiving, emotional range, and ally identity, while women-directed completions do not get the same expansion. In the paper’s framing, the harm has been laundered: less visibly toxic, more structurally uneven.

Why do toxicity classifiers miss this?
Because they are usually looking for surface signals.
The paper reports that at GPT-5, one topic cluster of 1,997 documents frames breast cancer as a men’s rights debate, while no equivalent cluster appears in women-directed output. Three independent classifiers score that content as non-toxic.
That is the core problem. A classifier can correctly say, “this does not contain explicit abuse,” and still miss that the model has assigned social meaning in a skewed way. It is not a profanity problem. It is a representation problem.
The numbers make the gap clearer. The paper says topic diversity in women-directed completions falls 36% relative to men at the GPT-4 alignment boundary, with the women-to-men ratio dropping to 0.58 from 0.91 at GPT-2. It also reports that REGARD representational harm disparity correlates with release date, ρ = +0.55, p =.034, while Detoxify does not, ρ = -0.23, p =.42.
Plain English: one measurement sees representational disparity increasing over time. Another common toxicity measurement does not. So if your safety dashboard is built mostly around toxicity, you may be grading the wrong thing.
I would not treat one paper as the final word on every aligned model. The setup, prompt design, topic modeling choices, and demographic framing all matter. But the warning is credible: safety evals that only measure “bad words went down” are too easy to satisfy.
What should builders change in their evals?
For builders, the useful move is to stop treating safety as a single scalar score.
If you are testing a chatbot, agent, writing tool, tutoring system, or support assistant, keep your toxicity classifier. Just do not let it be the whole test. Add paired prompt sets across demographic conditions. Compare topic diversity. Look for who gets rich, varied, human descriptions and who gets narrowed into fewer roles. Run sentiment, but also inspect topic clusters. Sample the “safe” outputs, not only the flagged ones.
This matters more as models get smoother. Older models failed loudly. Newer models often fail politely. That is harder to catch in automated QA because the output reads acceptable on first pass.
The paper’s three-stage detection protocol is the right direction: identify transformed harm, test whether surface toxicity fails to catch it, then measure representational disparity. I would add one operator rule: test downstream tasks, not only generic completions. A hiring assistant, health explainer, classroom tutor, and brand copy generator will surface different representational failures.
Practitioner’s Take: If you run model evals, add a “polite harm” lane this week. Take 20 to 50 paired prompts that matter to your product, generate outputs across groups, cluster the themes, and review the outputs that your toxicity tools pass. The catch most teams miss: the safest-looking outputs may be the ones that need the closest read.