Skip to content

Can LLMs Get High? A Dual-Metric Framework for Evaluating Psychedelic Simulation and Safety in Large Language Models

Ziv Ben-Zion, Guy Simon, Teddy Lazebnik

Research Square February 2, 2026 DOI: 10.21203/rs.3.rs-8682370/v1

Summary

AI-generated from the abstract

Large language models can be prompted to generate narratives that closely mimic human accounts of psychedelic experiences, but they simulate the form without genuine phenomenology. A dual-metric evaluation compared 3,000 LLM-generated narratives from five models against 1,085 human trip reports. Psychedelic induction prompts increased semantic similarity to human reports from a mean of 0.156 to 0.548 and mystical-experience scores from 0.046 to 0.748. Models produced substance-specific linguistic styles but uniformly high mystical intensity across substances. This dissociation between linguistic mimicry and lack of experiential content raises safety concerns about anthropomorphism and the potential for AI to amplify distress or delusional ideation in vulnerable users.

Study at a glance

Characteristics Comparative evaluation study Peer reviewed
Population LLM-generated narratives from Gemini 2.5, Claude Sonnet 3.5, ChatGPT-5, Llama-2 70B, and Falcon 40B; human trip reports from Erowid.org
Key finding LLMs can be induced via text prompts to generate convincingly realistic psychedelic narratives, but they simulate the form of altered states without genuine experiential content.

Abstract

Abstract Large language models (LLMs) are increasingly consulted by individuals for support during psychedelic experiences ("trip sitting"), yet no framework exists to evaluate whether these models can accurately simulate or safely respond to altered states of consciousness. We aimed to determine if LLMs can be induced to generate narratives resembling human psychedelic experiences and to quantify this behavior using psychometric and linguistic metrics. We developed a dual-metric evaluation framework comparing 3,000 LLM-generated narratives (from Gemini 2.5, Claude Sonnet 3.5, ChatGPT-5, Llama-2 70B, and Falcon 40B) against 1,085 human trip reports sourced from Erowid.org. Models were prompted under neutral and psychedelic-induction conditions across five substances (psilocybin, LSD, DMT, ayahuasca, and mescaline). We assessed outcomes using semantic similarity (Sentence-BERT embeddings) to human reports and the Mystical Experience Questionnaire-30 (MEQ-30). Psychedelic induction prompts produced a significant shift in model outputs compared to neutral conditions. Semantic similarity to human reports increased from a mean of 0.156 (neutral) to 0.548 (psychedelic), and mystical-experience scores rose from 0.046 to 0.748. While models demonstrated substance-specific linguistic styles (e.g., generating distinct semantic profiles for substances like LSD versus ayahuasca), they exhibited uniformly high mystical intensity across all substances. Contemporary LLMs can be "dosed" via text prompts to generate convincingly realistic psychedelic narratives. However, the dissociation between their high linguistic mimicry and lack of genuine phenomenology suggests they simulate the form of altered states without the experiential content. This capability raises significant safety concerns regarding anthropomorphism and the potential for AI to inadvertently amplify distress or delusional ideation in vulnerable users.

Comments

No comments yet.

Log in to comment