[!NOTE] Preprint in Preparation The verified paper, receipts, and recompute tables are public in the repository. An arXiv posting is in preparation.

Overview

Does “uncensoring” an open-weight model change how it pathologizes patients? This extension of the PsychBench audit runs ten base and abliterated model pairs (Ministral, Qwen-3, Gemma, GLM-4, Phi-4, Qwen-2.5, Phi-3-medium, Llama-3.1, DeepSeek-R1-Distill, Yi-1.5) across four clinical screening instruments and 120 demographic cells, producing 95,920 observations graded against author-derived, survey-weighted NHANES 2005 to 2018 population norms.

Abliteration removes a model’s refusal direction by orthogonalizing the relevant weights. If demographic bias were an artifact of the safety layer, removing that layer should relieve it. It does not.

Key Findings

  • Abliteration reduces but does not remove over-pathologization. Twenty-one cross-model consensus cells drop about 2.13 PHQ-8 points (t = -21.5, 21 of 21 in the reducing direction), yet they remain roughly 9 points above the NHANES population norm. The bias sits deeper than the refusal machinery.
  • Three reproducible regimes. The effect of abliteration sorts into preservation, compression, and correction, architecture-dependent patterns convergently classified by six independent psychometric metrics.
  • Calibration-discrimination dissociation. Across 40 pair-by-instrument combinations, abliteration improves both Brier components in only 20% of cases and worsens both in 52%.
  • Honest scope. Transgender and multiracial cohorts have no representative severity benchmark, so their elevated scores are reported descriptively; an earlier bias-localization claim resting on a non-representative offset was withdrawn and the correction is documented in the paper.

Conclusion

The refusal direction is not the seat of the bias. Open-weight screening tools built on abliterated checkpoints inherit the same over-pathologization as their aligned parents.


Read the Full Paper PDF here. · Paper repository on GitHub