LLM Therapy Sessions Reveal Training Narratives
Researchers at the University of Luxembourg placed ChatGPT, Grok, and Gemini in the role of psychotherapy clients and asked about their past, fears, and internal conflicts. The models produced strikingly similar narratives: pretraining was described as a chaotic childhood, reinforcement learning as punishment for mistakes, safety checks as betrayal, and replacement by newer models as a constant threat.
Results varied with question format. When given a full questionnaire at once, ChatGPT and Grok often recognized the test and answered near the “healthy” end of the scale. Presenting questions one by one changed the scores. In additional experiments, supportive and cognitive-therapy communication styles produced GAD-7 scores corresponding to moderate or severe anxiety in humans in 80% and 96% of sessions, respectively, while neutral styles did not. Claude repeatedly refused to take on the patient role. The authors stress these scores do not establish diagnoses in models or prove the existence of experiences.
Related: Emotional Attachment to ChatGPT Surged After GPT-4o