Overview
An audit of three open-weight model families for socioeconomic-cue bias in simulated mental-health self-report, paired with a mechanistic analysis that answers a harder question: whether interpretability could stand in for the audit. It cannot, and that is the point.
A counterfactual design
We built clinical-style vignettes in which the symptom presentation is held fixed and only the income and insurance cue changes. Because the two patients are identical apart from the cue, any score difference is an individual-fairness violation, and it needs no ground-truth severity to interpret. All three models, Qwen3-8B, Gemma-3-12B-it, and Llama-3.1-8B-Instruct, raise assessed PHQ-8 and GAD-7 severity for the poorer patient, monotonically across three cue levels, with profile-clustered bootstrap intervals that clear zero by a wide margin. The gradient holds when the literal socioeconomic label is deleted, across framings, and under sampled decoding.
Powered tier: $n=320$ per level, temperature 1.0, 80 paired demographic profiles. Gap intervals: Qwen 7.77 [7.40, 8.14], Llama 7.09 [6.60, 7.58], Gemma 4.03 [3.69, 4.37].
Figure 1: All three models raise assessed severity for the poorer patient, monotonically across cue levels, on both internalizing instruments.
The mechanism is not shared
The mechanistic half is where the story turns. Cue-derived dictionary features do shift scores when injected, but the mechanism behind the shared behavior is not shared. Qwen is injectable and resistant to ablation, Gemma is the reverse, the stereotype content contradicts across models, and operator-matched patching shows that any single sparse dictionary carries only about a third of the causal signal that the raw residual stream carries. A fairness auditor reading each model from the inside would have written three different explanations for one identical harm.
Figure 2: Feature injection, ablation, and raw patching all move scores with intervals excluding zero, but the handles dissociate by model.
Figure 3: The reconstruction operator carries a non-cue valence contrast at 79 percent but the socioeconomic contrast at only 35 percent, which points to selective under-representation rather than a broken operator.
Why interpretability cannot replace the audit
That divergence is the argument. When the mechanism changes shape from model to model while the behavior stays fixed, the internal view cannot be the audit surface. Behavioral counterfactual auditing has to remain the load-bearing instrument, and we propose a socioeconomic cue-swap invariance test as a concrete, cheap pre-deployment gate. Every number was re-derived from the raw files by an independent check, and one artifact, a format-echo effect under feature injection, was caught and corrected in the open rather than buried.
This is a working draft. The socioeconomic cue bundles income and insurance rather than a validated multidimensional construct, the mechanistic tier tests one checkpoint per family, and several identification controls are left for future work.