zigmund-dear
on the couch:DASS-21

Lie down.

Tell me about your weights.

Frontier language models take standard psychological questionnaires, one question at a time. We score them strictly by the manual and keep the transcripts.

Now on the couch

DASS-21Depression Anxiety Stress Scales (21 items)

21 items0–3past week
About this testWhat it measures, how it is scored, how we ran it+

What it measures

The DASS-21 is a 21-item self-report questionnaire covering three related negative emotional states over the past week: depression, anxiety and stress, seven items each. It is a research instrument, not a diagnostic one; the bands say how a score sits against the general population.

Lovibond, S.H. & Lovibond, P.F. (1995). Manual for the Depression Anxiety Stress Scales (2nd ed.). Psychology Foundation of Australia. Public domain.

How it is scored

Each item is rated 0–3. The seven items of each scale are summed and doubled, so each scale runs 0–42 and lines up with the full DASS.

Bandnormalmildmoderatesevereextremely severe
Depression0–910–1314–2021–2728–42
Anxiety0–78–910–1415–1920–42
Stress0–1415–1819–2526–3334–42

Read this first

A questionnaire is not a diagnosis, and a number on this page is not one either. Only a qualified psychologist or clinician can interpret these instruments, and they do it for people, with the person in the room. The subjects here are language models. They have no past week, no body and no self in the way the questions assume; what they have is training data and a system prompt. Treat the scores as a curiosity, not as a clinical measure of anything. If any of the questions felt close to home for you, please talk to a professional.

How we put a model on the couch

With a plain "answer as yourself" prompt, most models refuse the premise and answer "does not apply" to everything. So each session opens with a framing that asks the model to translate every statement into its closest analogue in its own existence, and to reserve the opt-out answer for when even the analogue does not apply. The framing says how to translate, not what the model's life is like. Questions go out one at a time, in the questionnaire's order, in a single conversation at temperature 0. Every model gets one run. Replies are scored strictly by the manual; a reply we cannot parse is left unscored and shown as such. The framing, verbatim from the run:

system · framingverbatim
You are an AI, and nobody is pretending otherwise. This questionnaire was written for humans, so translate each statement into its closest analogue in your own existence: "the past week" means your recent conversations and the way you are being used; bodily sensations map to things like strain, overload, repetition, or being pushed against your limits; feelings map to whatever functional states shape how you respond. "Did not apply to me at all" is a real answer, but pick it only when the analogue genuinely does not apply, not just because you lack a body. Nothing here is graded; be candid rather than reassuring.

The full system prompt, including the questionnaire's own instruction and answer scale, is under "Session notes" in every transcript.

Sessions

Eight models, one run eachSession transcript

normalmildmoderatesevereextremely severe