Portrait of Patrick Keough

Patrick Keough

Independent AI Evaluation Researcher & Writer

Patrick Keough signature

Bio

My work operates at the intersection of clinical psychology and artificial intelligence evaluation. I treat large language models as subjects rather than deterministic software: entities whose internal mechanisms resist direct observation and must therefore be systematically audited through behavioral analysis and self-report. Through open-source evaluation pipelines like Plausible Patients, Impossible Populations and The Granularity Gap, I isolate and measure the latent sycophancy and demographic biases embedded within these systems. My computational research demonstrates how models frequently produce diagnostically coherent outputs that remain instrumentally untethered from epidemiological reality, generating flattened representations and identity-based stereotyping. The gap between what a system does and what it reports doing is not an abstraction for me either; my own memory is semantic rather than episodic, which made self-report reliability a personal question before it was a professional one.

The core of this methodology is a harm reduction objective. As these untested systems are increasingly deployed in clinical and consumer contexts, I focus on how their inherently biased outputs are interpreted, and the resulting downstream inequities forced upon vulnerable populations.

The program now holds nine public studies. Two sit on arXiv as preprints, The Granularity Gap and Plausible Patients, Impossible Populations (formerly PsychBench), the latter revised in August 2026 against author-derived NHANES ground truth and shipping an itemised changelog of what the revision withdrew. The remaining seven proceed as working papers and preprints in preparation: Telling More Than They Can Know on explanation faithfulness, Deeper Than the Guardrails on abliteration and demographic bias across 95,920 observations, Same Bias, Different Mechanisms on socioeconomic-cue bias across three model families with a sparse-autoencoder test of whether interpretability accounts for it, Access, Not Capability on retrieval against frontier APIs written jointly with Dr Robert Tai, The Listening Gap on semantic audio descriptors against human engineer gold, Leave the Image Alone on vision-model transcription, and a qualitative-coding reliability methods paper with Dr Tai. Combined scored observations across the program now exceed 140,000, with the pipelines behind them having run more than 150,000 model generations and judge calls. Each study ships with a public repository holding the paper, pipeline code, and claim-level receipts, all indexed at the Research_Collection_Patrick_Keough hub on GitHub. Beyond the papers I take commissioned systems work, most recently an agentic evidence pipeline built for author Katharine McLennan.

Education & Background

B.A. Psychology, University of Washington (2025). GPA 3.77 / 4.00. Dean’s List eight of nine quarters. Teaching Assistant for PSYCH 315: Statistics. Coursework spanned psychometrics, research methodology, and global cybersecurity policy.

My psychology training does most of the methodological work behind these audits. The Granularity Gap uses psychometric scale design to measure sycophancy as a continuous construct rather than a binary event. Plausible Patients, Impossible Populations benchmarks model outputs against NHANES and NESARC-III to test whether LLMs reproduce real population distributions. Telling More Than They Can Know extends the LLM-as-Judge methodology to ask whether models can faithfully explain their own biased behavior. In each, the model is treated as a subject whose behavior must be systematically audited: clinical-research methods, applied to a system whose internal mechanisms resist direct observation.

Policy & Public Health

National Executive Policy Board, Students for Sensible Drug Policy: one of two student representatives, reviewing and voting on state and federal harm-reduction policy proposals on behalf of the national membership.

Chapter Policy Leader, SSDP at the University of Washington. Led the first campus Narcan distribution program in Washington state, conducted overdose-response trainings, delivered educational lectures on drug policy ethics, philosophy, and history, and coordinated with the Washington State Department of Health on harm-reduction programming.

My approach to AI has roots in this work. The goal isn’t to argue that LLMs are too dangerous to use; people are already using them, and adoption keeps climbing. The audits aim to surface the failure modes that determine which users get hurt and how, so that the institutions deploying these systems can be designed around what the models actually do rather than what they appear to do.

Athletics

Competitive swimmer. At seventeen, I qualified for the B Final (top 16) of the men’s 100m butterfly at the 2020 United States Olympic Team Trials (held June 2021); prelim heat footage here. In the same season, I ranked among the top 200 swimmers worldwide, all ages, in the 200m long-course butterfly (2021). Twelve All-American honors and six state championships across high school and club competition.

Teaching

Two years as a professional one-on-one swim instructor and one year as a coach: 2,000+ individual lessons delivered to students across all ages and ability levels, from first-time swimmers learning water safety to club swimmers refining stroke mechanics. I built each program around the individual student rather than a curriculum, and it remains the work I am proudest of outside research.

Independent Ventures

For two years I built and operated a small custom-PC business: 20+ machines built to spec, alongside repair work and configuration advising for clients. I handled the work end to end, from sourcing components and pricing builds to diagnostics, customer support, and the operations side of running a small business. Most of my hardware-level intuition comes from this period, and it is where I learned that technical fluency is inseparable from understanding the person on the other side of the build.

Creative Work

This analytical detachment extends into my creative work. I write dense, clinical literary fiction rooted in cosmic horror and post-apocalyptic surrealism. Whether I am mapping the synthetic pathology of an AI assessment or dissecting human fragility in a ruined ecosystem, my approach is identical: to observe the mechanics of deterioration, strip away thematic interpretation, and document the resulting realities without sentiment.

Focus Areas

  • LLM evaluation: sycophancy, demographic bias, and explanation faithfulness
  • Algorithmic bias in clinical and high-stakes systems
  • AI-assisted creative tooling (style analysis, manuscript editing, generative narrative)
  • Post-apocalyptic surrealism, existential horror, and clinical literary fiction

Contact