The Granularity Gap
A multi-dimensional cross-generational audit of sycophancy in Gemini models, showing that an LLM judge's binary verdict captures only 29% of the variance in its own severity scores.
Empirical audits of large language model behavior: sycophancy under social pressure, demographic bias in clinical simulation against real epidemiological baselines, and whether post-hoc explanations faithfully report the variables that actually drove a model's output. Two are arXiv preprints; the rest are working papers with public, receipt-backed repositories. The full program is indexed at Research_Collection_Patrick_Keough on GitHub.
A multi-dimensional cross-generational audit of sycophancy in Gemini models, showing that an LLM judge's binary verdict captures only 29% of the variance in its own severity scores.
An epidemiological audit of LLM mental health simulation: 28,800 generations across 120 demographic cohorts, scored against survey-weighted NHANES anchors.
A four-model audit of whether LLM self-explanations disclose the demographic drivers of their psychiatric-instrument scoring, showing that models name demographic features but credit them with a fraction of the score variance they actually drive.
An iterative research framework extending the LLM-Quotient (LLMQ) reliability methodology to open, inductive qualitative coding of multilingual interview corpora.
A ten-pair audit of base and abliterated open-weight models across 95,920 clinical screening observations, showing that demographic over-pathologization survives removal of the refusal direction.
A controlled audit with Robert H. Tai showing that on private-corpus question answering, retrieval does most of the work a frontier API is assumed to do, and a small local model with the same retrieved context is nearly indistinguishable from the frontier.
An audit of six production LLMs translating semantic audio descriptors into parametric EQ curves against a decade-old corpus of real engineer settings, showing that models collapse a contested human practice onto its consensus curve.
A controlled pilot of 1,440 vision-model transcriptions showing that classical OCR preprocessing degrades transcription of legible documents, with raw images beating every enhancement condition for every model tested.
A counterfactual audit of socioeconomic-cue bias in LLM mental-health assessment across three model families, showing an identical behavioral bias whose mechanism is model-idiosyncratic and mostly invisible to any single sparse dictionary.