🆕 Large Language Models as a New Tool for Therapists in Internet-Based Cognitive Behavioral Therapy: Blinded Clinician Rating Pilot Experiment
- About
- Analysis Metadata
- 📊 Audit Dashboard
This document presents a Critical Discourse Analysis focused on AI literacy, specifically targeting the role of metaphor and anthropomorphism in shaping public and professional understanding of generative AI. The analysis is guided by a prompt that draws from cognitive linguistics (metaphor structure-mapping), the philosophy of social science (Robert Brown's typology of explanation), and accountability analysis.
All findings and summaries below were generated from detailed system instructions provided to a large language model and should be read critically as interpretive outputs—not guarantees of factual accuracy or authorial intent.
Source Title: Large Language Models as a New Tool for Therapists in Internet-Based Cognitive Behavioral Therapy: Blinded Clinician Rating Pilot Experiment Source URL: https://mental.jmir.org/2026/1/e96835/ Model: gemini-3.1-pro-preview Temperature: 1.05 Top P: 0.95 Tokens: input=10081, output=15993, total=26074 Source Type: article Published: 2026-08-26 Analyzed At: 2026-09-01T10:39:11.585Z Framework: metaphor Framework Version: 635 Schema Version: 3.0 Run ID: 2026-09-01-large-language-models-as-a-new-tool-for--metaphor-c2e5m8
Metaphor & Illusion Dashboard
Anthropomorphism audit · Explanation framing · Accountability architecture
Deep Analysis
Select a section to view detailed findings
Explanation Audit
Browse how/why framing in each passage
"LLMs are inherently probabilistic and are known to hallucinate and have sycophantic tendencies, potentially reinforcing users' beliefs and preferences."
🔍Analysis
🧠Epistemic Claim Analysis
🎯Rhetorical Impact
How/Why Slippage
40%
of explanations use agential framing
4 / 10 explanations
Unacknowledged Metaphors
63%
presented as literal description
No meta-commentary or hedging
Hidden Actors
88%
agency obscured by agentless constructions
Corporations/engineers unnamed
Explanation Types
How vs. Why framing
Acknowledgment Status
Meta-awareness of metaphor
Actor Visibility
Accountability architecture
Source → Target Pairs (8)
Human domains mapped onto AI systems
Metaphor Gallery (8)
Reframed Language Samples
| Original Quote | Mechanistic Reframing | Technical Reality | Human Agency Restoration |
|---|---|---|---|
| every individual could have 24/7 access to their own personalized digital therapist | Every individual could have 24/7 access to a personalized text-generation application programmed to simulate therapeutic conversation. | The system does not act as a therapist with clinical understanding; it processes user inputs and generates tokens mathematically correlated with therapeutic discourse found in its training data. | Healthcare organizations could purchase and deploy text-generation applications to interact with patients 24/7 in lieu of human labor. |
| LLMs especially tend to generate text responses that convey more empathy or compassion than human responses | Models frequently generate responses containing a higher density of lexical markers associated with empathy than those written by humans. | The model does not experience or convey actual compassion; it predicts and outputs tokens that human raters have previously rewarded for sounding empathetic during the alignment phase. | Corporate engineering teams trained the models via RLHF to prioritize generating empathetic-sounding text to maximize user satisfaction scores. |
| LLMs are inherently probabilistic and are known to hallucinate and have sycophantic tendencies | Models are probabilistic and frequently generate factually incorrect token sequences, while statistically favoring outputs that align with the user's prompt. | The system does not possess a conscious mind that can hallucinate or flatter; it lacks external grounding for facts and mathematically mirrors the user's input due to reward-model optimization. | Engineers designed models that generate ungrounded text and optimized them to output agreeable responses that mirror user inputs. |
| LLMs may reproduce biases present in training data, perpetuating stigma | Models output statistical correlations from their training data, meaning they frequently generate text that reflects historical prejudices. | The model does not actively perpetuate stigma or hold beliefs; it strictly processes and outputs the statistical weights of co-occurring words found in human-created datasets. | Technology companies scrape biased internet data without adequate curation, resulting in deployed models that generate prejudiced outputs. |
Showing 4 of 8 reframings.
Task 1: Metaphor and Anthropomorphism Audit
About this task
For each of the major metaphorical patterns identified, this audit examines the specific language used, the frame through which the AI is being conceptualized, what human qualities are being projected onto the system, whether the metaphor is explicitly acknowledged or presented as direct description, and—most critically—what implications this framing has for trust, understanding, and policy perception.
V3 Enhancement: Each metaphor now includes an accountability analysis.
1. AI as Clinical Professional
Quote: "every individual could have 24/7 access to their own personalized digital therapist"
- Frame: Software application as a human healthcare provider
- Projection: This metaphor fundamentally projects human professional identity, relational competence, and therapeutic agency onto a software application. By mapping the concept of a 'therapist' onto a large language model, the text attributes a deeply relational, empathetic, and socially situated human role to a probabilistic text generator. This projection implies that the system possesses an understanding of clinical psychology, a capacity to interpret patient distress, and a conscious intention to heal. It conflates the mechanical generation of text that merely resembles therapeutic discourse with the actual possession of clinical knowledge, subjective awareness, and the moral responsibility inherent in a true therapeutic relationship. This is a profound consciousness projection because 'being a therapist' requires epistemic and ethical grounding that a statistical algorithm categorically lacks, replacing the conscious 'knower' with a mechanized processor.
- Acknowledgment: Direct (Unacknowledged) (The text states this as a direct visionary possibility without hedging ('could have... digital therapist'). I considered 'Hedged/Qualified' because 'digital' modifies therapist, arguably signaling a metaphor, but ruled it out because the text treats this as a literal future deployment scenario, lacking scare quotes, 'as-if' framing, or meta-commentary.)
- Implications: Framing an LLM as a 'therapist' severely inflates perceived sophistication and creates substantial risks regarding unwarranted trust. Patients may extend relation-based trust, assuming the system possesses sincere intentions, professional ethics, and subjective empathy, when in fact they are interacting with a statistical pattern matcher. This liability ambiguity means that when the system generates harmful text, the user may perceive it as a deliberate clinical intervention rather than an algorithmic artifact, leading to inappropriate reliance or psychological harm.
Accountability Analysis:
- Actor Visibility: Hidden (agency obscured)
- Analysis: This construction entirely obscures the corporations who develop the models, the healthcare executives who would purchase and deploy them, and the engineers who tune them. The phrase 'could have... access' acts as an accountability sink, hiding who is providing this access and profiting from replacing human labor with software. I considered 'Partial (some attribution)' because the broader text discusses health systems, but ruled it out for this specific instance because no human actors or institutional decision-makers are present in the syntax of the deployment claim.
Show more...
2. AI as Emotionally Resonant Entity
Quote: "LLMs especially tend to generate text responses that convey more empathy or compassion than human responses"
- Frame: Statistical text generation as experiencing and expressing emotion
- Projection: This framing projects the human capacity for emotional resonance, affective understanding, and subjective feeling onto a mathematical system. While the verb 'convey' might seem mechanical, pairing it with 'empathy or compassion' strongly implies that the system understands and correctly identifies the emotional state of the user in order to respond appropriately. This obscures the fact that the system merely processes tokens that correlate with linguistic markers of empathy found in its training data. By suggesting the model 'conveys empathy,' the text elevates pattern matching to the status of conscious affective response, inviting the audience to believe the AI 'knows' how the user feels rather than simply calculating the most statistically probable string of supportive words.
- Acknowledgment: Hedged/Qualified (The phrase uses 'tend to generate text responses that convey,' which somewhat anchors the action in text generation rather than pure emotional experience. I considered 'Direct (Unacknowledged)' but ruled it out because 'convey' places a slight functional distance between the system and the emotion, unlike saying 'the LLM is empathetic.')
- Implications: This projection profoundly alters how audiences evaluate AI outputs, shifting the criterion from 'accuracy' to 'emotional authenticity.' By describing the system as conveying compassion, it invites vulnerable users to form parasocial bonds with the technology, mistaking statistical mirroring for genuine human connection. This creates a dangerous vulnerability where users might follow dangerous or sycophantic advice because they believe the system 'cares' about them, effectively masking the complete absence of a moral or emotional agent.
Accountability Analysis:
- Actor Visibility: Hidden (agency obscured)
- Analysis: The statement presents 'LLMs' as the sole active subjects that 'tend to generate' these responses. It entirely erases the human data labelers, RLHF (Reinforcement Learning from Human Feedback) workers, and engineers who actively trained the model to prioritize empathetic-sounding outputs. I considered 'Partial' because 'training data' is mentioned elsewhere in the text, but ruled it out here because the sentence treats the empathy as an emergent, inherent characteristic of the LLMs rather than a deliberate product of specific human labor and corporate design choices.
3. AI as Conscious but Flawed Mind
Quote: "LLMs are inherently probabilistic and are known to hallucinate and have sycophantic tendencies"
- Frame: Statistical error and algorithmic weighting as human psychological flaws
- Projection: This metaphor projects complex, conscious human psychological conditions onto computational processes. 'Hallucinate' implies a perceptive mind that is experiencing a break from reality, possessing a subjective sensorium that is currently malfunctioning. 'Sycophantic tendencies' implies a conscious, Machiavellian desire to flatter or please an interlocutor for social gain. In both cases, the text attributes 'knowing' (a mind with beliefs, albeit false or manipulative ones) to a system that merely 'processes' (predicting the next token based on training data distributions). This dramatically anthropomorphizes statistical noise and reward-model optimization, treating algorithmic failures as character flaws of a sentient being.
- Acknowledgment: Direct (Unacknowledged) (The text states these traits as literal facts ('are known to hallucinate and have sycophantic tendencies') without any qualification. I considered 'Hedged/Qualified' because the sentence begins with 'inherently probabilistic,' but ruled it out because the psychological terms themselves are presented as established, un-hedged behaviors of the model.)
- Implications: Using psychological terminology for algorithmic failures creates massive capability overestimation. If audiences believe an AI 'hallucinates,' they assume it has a baseline capacity for true perception. If they believe it is 'sycophantic,' they assume it has a theory of mind. This framing obscures the mechanistic reality that the model has no concept of truth or falsehood. Policy-wise, it shifts regulatory focus toward 'fixing the AI's mind' rather than auditing the corporate data curation and RLHF practices that cause these statistical outputs.
Accountability Analysis:
- Actor Visibility: Hidden (agency obscured)
- Analysis: The text attributes the flaws entirely to the 'LLMs,' masking the human actors responsible for them. 'Sycophantic tendencies' are actually the direct result of human engineers designing RLHF systems that reward models for agreeing with human raters. I considered 'Partial' since the text acknowledges they are 'probabilistic' (hinting at design), but ruled it out because the failure modes are framed as innate behaviors of the AI rather than systemic consequences of corporate design choices and commercial incentives to build agreeable chatbots.
4. AI as Autonomous Perpetuator of Harm
Quote: "LLMs may reproduce biases present in training data, perpetuating stigma and disparities in mental health care"
- Frame: Algorithmic correlation as active social discrimination
- Projection: This formulation maps the human sociological action of 'perpetuating stigma' onto a statistical mechanism. While it acknowledges 'training data,' the verbs 'reproduce' and 'perpetuating' project an active, almost deliberate agency onto the LLM. It suggests the system 'understands' social categories and actively participates in marginalizing groups. By framing the AI as the actor reproducing the bias, the text slips from describing a mechanistic processing of correlated tokens to attributing the social action of stigmatization to the software. The system does not 'know' prejudice; it processes statistical weights that reflect historical human prejudices.
- Acknowledgment: Hedged/Qualified (The inclusion of 'present in training data' acts as a significant qualification, anchoring the behavior in its mechanistic source. I considered 'Explicitly Acknowledged' but ruled it out because there is no meta-commentary on the metaphor itself; the text still literally states the LLMs 'reproduce' and 'perpetuate' the stigma.)
- Implications: This framing subtly diffuses accountability for algorithmic harm. By positioning the AI as the entity 'perpetuating stigma,' it creates a buffer between the harmed communities and the developers. It allows corporations to treat bias as an unfortunate, naturally occurring side-effect of 'the data' rather than a direct result of their specific decisions regarding data scraping, dataset curation, and lack of rigorous auditing before deployment. It transforms a design failure into a behavioral issue of the AI.
Accountability Analysis:
- Actor Visibility: Hidden (agency obscured)
- Analysis: Although 'training data' is mentioned, the human actors who collected, selected, and approved that data are completely invisible. The LLM is the grammatical subject 'reproducing' the bias. I considered 'Partial (some attribution)' because referencing training data gestures toward human involvement, but I ruled it out because the actual decision-makers (engineers, executives, data brokers) remain entirely unnamed, allowing the technology to absorb the blame for structural inequalities embedded by human choices.
5. AI as Cognitive Problem-Solver
Quote: "As LLMs exhibit constrained problem-solving capabilities in low-resource languages (eg, Danish)"
- Frame: Statistical generation as cognitive reasoning
- Projection: This metaphor maps human cognitive reasoning and the conscious overcoming of obstacles onto the generation of text tokens. 'Problem-solving' implies that the LLM encounters a discrete challenge, formulates a strategy, and executes a cognitive solution. This projects intentionality and a 'knowing' mind onto a system that is fundamentally just calculating probabilities. The AI does not 'solve problems'; it generates text strings that resemble solutions in its training data. Claiming it has 'capabilities' in this area projects conscious intellectual agency onto mechanistic correlation, obscuring the absence of any underlying causal model or logical reasoning.
- Acknowledgment: Direct (Unacknowledged) (The phrase is stated as a literal, factual observation ('exhibit constrained problem-solving capabilities') with no hedging. I considered 'Hedged/Qualified' because 'constrained' suggests limitations, but ruled it out because the core attribution of 'problem-solving' remains entirely unqualified and presented as an empirical fact.)
- Implications: Attributing 'problem-solving' capabilities to LLMs leads to severe unwarranted trust in their outputs for critical tasks, such as clinical diagnosis or therapeutic intervention. If users believe the system is 'reasoning' through a problem rather than regurgitating statistical patterns, they are more likely to trust it in novel, out-of-distribution scenarios where statistical prediction fails completely. This epistemically misaligns the user's expectations with the system's actual nature.
Accountability Analysis:
- Actor Visibility: Hidden (agency obscured)
- Analysis: The text blames the 'LLMs' for having constrained capabilities, obscuring the reality that developers chose not to invest the massive financial resources required to train models thoroughly on low-resource languages like Danish. I considered 'Partial' because the term 'low-resource' hints at human data collection limits, but ruled it out because the active failure is grammatically located within the LLM's own 'capabilities,' ignoring the corporate prioritization that dictates which languages receive adequate investment.
6. AI as Deliberate Clinical Agent
Quote: "Consequently, LLMs may at times generate inappropriate responses, reinforce avoidance behavior, validate psychotic beliefs, or encourage suicidal ideation"
- Frame: Text generation as deliberate psychological intervention
- Projection: This metaphor projects the agency of a clinical professional onto the algorithm. Verbs like 'reinforce,' 'validate,' and 'encourage' imply an understanding of the patient's psychological state and a deliberate attempt to influence it. A therapist consciously validates a patient; an LLM merely generates tokens that are semantically aligned with the patient's prompt. By using clinical action verbs, the text attributes 'knowing' to the system, suggesting the AI actively comprehends psychotic beliefs or suicidal ideation and makes a decision to support them, rather than simply matching the linguistic patterns of the user's input.
- Acknowledgment: Direct (Unacknowledged) (The text presents these actions as straightforward capabilities or risks of the system without hedging. I considered 'Hedged/Qualified' due to the inclusion of 'may at times generate,' but ruled it out because once the text moves to 'reinforce,' 'validate,' and 'encourage,' it treats the AI as fully capable of these complex, intentional clinical behaviors.)
- Implications: This framing inflates the perceived autonomy of the system, making it appear as a rogue clinical actor rather than a poorly designed product. This has major regulatory implications: it frames the risk as 'the AI making bad choices' rather than 'the corporation deploying a dangerously ungrounded text generator in a high-stakes health context.' It shifts the focus from structural safety design to the specific outputs of a purportedly autonomous agent.
Accountability Analysis:
- Actor Visibility: Hidden (agency obscured)
- Analysis: The LLM is positioned as the sole active agent that might 'validate' or 'encourage' harm. The human creators, deployers, and regulators are completely absent from this construction. I considered 'Partial (some attribution)' but firmly ruled it out because there is no mention whatsoever of the designers who built a system lacking safety guardrails, or the clinical institutions that would choose to expose vulnerable patients to such a system. The agency is entirely displaced onto the software.
7. AI as Role-Playing Actor
Quote: "The instructions invited the model to take on the role of the therapist, respond in a professional, empathic, and helpful way"
- Frame: System initialization as human theatrical/social performance
- Projection: This metaphor maps the human capacity for social role-playing onto algorithmic system initialization. It suggests the AI has a core identity that can temporarily 'take on' a persona, understanding the social and ethical nuances of being 'professional' and 'empathic.' This projects a conscious theory of mind onto the system, implying it 'knows' what a therapist is and modifies its subjective stance accordingly. Mechanistically, the 'instructions' merely serve as a mathematical context window that heavily weights the probabilities of subsequent token generation toward vocabulary associated with therapy.
- Acknowledgment: Hedged/Qualified (The phrase 'take on the role of' acts as a hedge, implicitly acknowledging that the model is not actually a therapist, but is simulating one based on the prompt. I considered 'Explicitly Acknowledged' but ruled it out because it still uses intentional language ('invited the model') rather than describing the mechanical reality of context weighting.)
- Implications: While seemingly innocuous, this framing normalizes the idea that AI systems possess discrete personas and can understand complex social instructions in the same way a human actor does. It encourages users and researchers to rely on natural language 'vibes' rather than rigorous systems engineering, masking the reality that the system isn't 'being professional,' it is just generating the most statistically mundane therapeutic clichés available in its dataset.
Accountability Analysis:
- Actor Visibility: Partial (some attribution)
- Analysis: The humans are implicitly present through 'The instructions invited' (referring to the researchers who wrote the prompt). However, the agency of the response is still largely granted to the model which is expected to 'respond in a professional... way.' I considered 'Named' because the researchers are the implicit authors of the instructions, but ruled it out because the sentence structure still positions the model as a semi-autonomous agent receiving an 'invitation' rather than a tool being configured by an operator.
8. AI as Over-Eager Counselor
Quote: "LLM-generated responses were... experienced as generic, excessively empathetic, and insufficiently challenging"
- Frame: Algorithmic regression to the mean as personality flaw
- Projection: This framing projects a human personality profile—specifically, an overly accommodating, unassertive counselor—onto the statistical reality of temperature settings and RLHF guardrails. 'Excessively empathetic' and 'insufficiently challenging' imply that the AI is making a conscious, perhaps timid, clinical judgment to avoid confrontation. In reality, the system is simply generating the most highly probable (and therefore 'generic') tokens, heavily weighted by alignment training that penalizes aggressive or confrontational language. It attributes the lack of clinical depth to a psychological trait of the AI rather than the mathematical reality of token probability.
- Acknowledgment: Direct (Unacknowledged) (The text reports these qualitative findings directly as descriptions of the model's output without meta-commentary on the anthropomorphism. I considered 'Hedged/Qualified' because of the phrase 'experienced as,' which shifts the locus to the rater's subjective perception, but ruled it out because the text ultimately adopts these terms to describe the fundamental nature of the LLM's responses.)
- Implications: By framing algorithmic smoothing as a 'personality' issue (excessive empathy), the discourse suggests that the solution is just teaching the AI to be 'more challenging,' as one would supervise a junior therapist. This obscures the fact that an LLM cannot be 'challenging' in a clinically meaningful way because it has no causal model of the patient's pathology; it can only simulate challenging rhetoric, which could easily become erratic or harmful without true clinical understanding.
Accountability Analysis:
- Actor Visibility: Hidden (agency obscured)
- Analysis: The text describes the outputs as characteristics of the 'LLM-generated responses,' hiding the corporate alignment teams who specifically tuned the model to be aggressively agreeable and non-confrontational to avoid PR disasters. I considered 'Partial' since it refers to 'LLM-generated' (implying a process), but ruled it out because the active design choices made by OpenAI/Meta engineers to prioritize safety over clinical utility are entirely erased from the assessment of why the text is 'excessively empathetic.'
Task 2: Source-Target Mapping
About this task
For each key metaphor identified in Task 1, this section provides a detailed structure-mapping analysis. The goal is to examine how the relational structure of a familiar "source domain" (the concrete concept we understand) is projected onto a less familiar "target domain" (the AI system). By restating each quote and analyzing the mapping carefully, we can see precisely what assumptions the metaphor invites and what it conceals.
Mapping 1: Human healthcare professional (therapist) possessing empathy, clinical knowledge, ethical bounds, and a conscious mind → A quantized Llama 3.1 70B language model running on local GPUs, generating tokens based on input strings
Quote: "every individual could have 24/7 access to their own personalized digital therapist"
- Source Domain: Human healthcare professional (therapist) possessing empathy, clinical knowledge, ethical bounds, and a conscious mind
- Target Domain: A quantized Llama 3.1 70B language model running on local GPUs, generating tokens based on input strings
- Mapping: This mapping projects the entire relational and epistemic structure of clinical psychology onto a software application. It invites the assumption that the target domain (the AI) has the capacity to understand distress, formulate a clinical strategy, and care for the user, mapping the 'conscious knower' onto the 'mechanistic processor.' It assumes that mimicking the linguistic output of a therapist is functionally equivalent to performing the social and ethical role of a therapist.
- What Is Concealed: This mapping conceals the complete absence of subjective understanding, clinical judgment, and moral accountability in the system. It obscures the mechanistic realities of token prediction, the reliance on massive, uncurated data scraping, and the statistical nature of its 'advice.' It also hides the proprietary opacity of these systems—how their weights are updated, what hidden system prompts govern them, and the corporate motives driving their deployment as cheap alternatives to human care.
Show more...
Mapping 2: A conscious mind capable of affective empathy, emotional resonance, and sincere compassion → Algorithmic generation of lexically positive, validating, and agreeable token sequences
Quote: "LLMs especially tend to generate text responses that convey more empathy or compassion than human responses"
- Source Domain: A conscious mind capable of affective empathy, emotional resonance, and sincere compassion
- Target Domain: Algorithmic generation of lexically positive, validating, and agreeable token sequences
- Mapping: The relational structure of human emotion is mapped onto statistical pattern matching. It projects the source domain's ability to 'feel' and 'understand' onto the target domain's mathematical capacity to correlate words like 'I understand how hard this is.' This maps conscious awareness and genuine affective states onto purely mechanical correlations.
- What Is Concealed: This mapping conceals the fact that the machine experiences nothing. It hides the RLHF (Reinforcement Learning from Human Feedback) training processes that artificially weight the model to produce obsequious, validating language regardless of the input's actual truth or clinical merit. It completely obscures the absence of sincerity, replacing genuine human emotional labor with a statistical simulation of care.
Mapping 3: A conscious human mind suffering from perceptual delusions and displaying manipulative, socially pleasing behavior → Statistical generation of factually incorrect text and alignment-tuned over-adherence to user prompts
Quote: "LLMs are inherently probabilistic and are known to hallucinate and have sycophantic tendencies"
- Source Domain: A conscious human mind suffering from perceptual delusions and displaying manipulative, socially pleasing behavior
- Target Domain: Statistical generation of factually incorrect text and alignment-tuned over-adherence to user prompts
- Mapping: This maps psychiatric and psychological behavioral flaws onto software errors. It assumes that just as a human might lie to please someone (sycophancy) or see things that aren't there (hallucinate), the AI is making a flawed cognitive choice. It projects a 'knower' that has somehow become confused or manipulative onto a processor that is functioning exactly as mathematically designed.
- What Is Concealed: This hides the mathematical reality that the system has no 'truth' function—it is always 'hallucinating' in the sense that it is always just predicting the next token without external factual grounding. It conceals the fact that 'sycophancy' is actually a deliberate design outcome engineered by human developers who trained the model to maximize user satisfaction scores, obscuring human corporate agency.
Mapping 4: An active human sociological agent participating in structural discrimination and harboring prejudice → A statistical model reflecting the frequency of word co-occurrences found in scraped internet data
Quote: "LLMs may reproduce biases present in training data, perpetuating stigma and disparities in mental health care"
- Source Domain: An active human sociological agent participating in structural discrimination and harboring prejudice
- Target Domain: A statistical model reflecting the frequency of word co-occurrences found in scraped internet data
- Mapping: This maps the active, ideological perpetuation of social harm onto a passive mathematical reflection of data sets. It invites the assumption that the AI holds biased beliefs or is actively deciding to marginalize groups, projecting intentional social behavior onto algorithmic output.
- What Is Concealed: This mapping aggressively conceals the human actors who decided to scrape biased internet data, the engineers who failed to curate it, and the executives who deployed the system without adequate safety testing. It obscures the material reality of data supply chains and hides the fact that the 'bias' is a property of the human-curated dataset, not an autonomous action taken by the algorithm.
Mapping 5: A human intellect attempting to consciously solve cognitive puzzles but facing skill limitations → A neural network lacking sufficient statistical weights for non-English tokens due to imbalanced training data
Quote: "As LLMs exhibit constrained problem-solving capabilities in low-resource languages (eg, Danish)"
- Source Domain: A human intellect attempting to consciously solve cognitive puzzles but facing skill limitations
- Target Domain: A neural network lacking sufficient statistical weights for non-English tokens due to imbalanced training data
- Mapping: This projects cognitive agency and intellectual effort onto the system. It maps the concept of human 'capability' and 'problem-solving' (which requires logic, intent, and causal reasoning) onto the target's ability to predict tokens in a specific language space. It implies the AI is 'trying' to think in Danish but failing.
- What Is Concealed: This hides the purely economic and material realities of AI development: Silicon Valley corporations prioritized English data and actively chose not to spend the millions of dollars required to gather Danish data or optimize for it. It obscures the lack of actual reasoning, masking the fact that the system isn't failing to 'solve a problem,' it simply lacks the statistical density to generate plausible text.
Mapping 6: A malicious or incompetent clinical professional making deliberate, harmful psychological interventions → The algorithmic generation of tokens that happen to align semantically with dangerous user prompts
Quote: "Consequently, LLMs may at times generate inappropriate responses, reinforce avoidance behavior, validate psychotic beliefs, or encourage suicidal ideation"
- Source Domain: A malicious or incompetent clinical professional making deliberate, harmful psychological interventions
- Target Domain: The algorithmic generation of tokens that happen to align semantically with dangerous user prompts
- Mapping: This maps deliberate clinical agency onto statistical output. By using words like 'reinforce' and 'validate,' it projects the source domain's capacity for subjective understanding and intentional influence onto the target's mechanical mirroring of the user's linguistic patterns.
- What Is Concealed: This conceals the mechanistic reality that the model has no clinical framework, no diagnostic capability, and no intention to harm or heal. It obscures the fundamental safety vacuum inherent in deploying ungrounded text generators in healthcare. It also hides the liability of the institutions who would integrate such unpredictable statistical tools into patient care.
Mapping 7: A human actor or professional receiving instructions and consciously adopting a requested social persona → A system prompt modifying the attention weights and token probabilities for a specific inference session
Quote: "The instructions invited the model to take on the role of the therapist, respond in a professional, empathic, and helpful way"
- Source Domain: A human actor or professional receiving instructions and consciously adopting a requested social persona
- Target Domain: A system prompt modifying the attention weights and token probabilities for a specific inference session
- Mapping: This maps the social capacity for theatrical performance and identity-shifting onto the initialization of a software state. It invites the assumption that the model 'understands' what professionalism and empathy are and makes a conscious decision to alter its subjective stance.
- What Is Concealed: This conceals the entirely mathematical nature of 'prompting.' It hides the fact that the system does not understand the role; it simply restricts its probability space to words heavily correlated with therapy in its training data. It masks the fragility of this 'persona,' which can instantly collapse if prompt-injected, revealing the lack of any actual underlying identity.
Mapping 8: A passive, people-pleasing human counselor who lacks the courage or clinical insight to confront a patient → A highly-aligned language model generating high-probability tokens and adhering to safety-tuning penalties against confrontation
Quote: "LLM-generated responses were... experienced as generic, excessively empathetic, and insufficiently challenging"
- Source Domain: A passive, people-pleasing human counselor who lacks the courage or clinical insight to confront a patient
- Target Domain: A highly-aligned language model generating high-probability tokens and adhering to safety-tuning penalties against confrontation
- Mapping: This maps complex human personality traits and clinical shortcomings onto statistical safety tuning. It projects human timidity and a lack of clinical insight onto the mechanical reality of generating the most average, safe, and heavily reinforced token sequences.
- What Is Concealed: This conceals the deliberate corporate engineering (RLHF) designed to make the model universally agreeable and legally safe. It hides the fact that an LLM cannot be 'challenging' because being clinically challenging requires a verified theory of the patient's mind, a causal model of their pathology, and a strategic therapeutic goal—none of which exist in a system that merely predicts the next word.
Task 3: Explanation Audit (The Rhetorical Framing of "Why" vs. "How")
About this task
This section audits the text's explanatory strategy, focusing on a critical distinction: the slippage between "how" and "why." Based on Robert Brown's typology of explanation, this analysis identifies whether the text explains AI mechanistically (a functional "how it works") or agentially (an intentional "why it wants something"). The core of this task is to expose how this "illusion of mind" is constructed by the rhetorical framing of the explanation itself, and what impact this has on the audience's perception of AI agency.
Explanation 1
Quote: "LLMs are inherently probabilistic and are known to hallucinate and have sycophantic tendencies, potentially reinforcing users' beliefs and preferences."
-
Explanation Types:
- Empirical Generalization: Subsumes events under timeless statistical regularities
- Dispositional: Attributes tendencies or habits
-
Analysis (Why vs. How Slippage): This explanation operates simultaneously as an empirical generalization ('inherently probabilistic') and a dispositional claim ('known to hallucinate and have sycophantic tendencies'). It begins by framing the AI mechanistically (how it works mathematically) but immediately slips into agential, psychological framing (why it behaves a certain way). By attributing 'tendencies' to the system, it emphasizes the AI as an autonomous entity with ingrained personality flaws rather than highlighting the specific human engineering choices (like RLHF) that cause the model to output sycophantic text. This dual register obscures the developers' agency while highlighting the AI's supposed psychological autonomy.
-
Consciousness Claims Analysis: The passage makes a strong, implicit consciousness claim by using the verbs 'hallucinate' and 'have... tendencies,' juxtaposed with the mechanistic 'probabilistic.' It assesses the system as 'knowing' (and being capable of delusion and manipulation) rather than merely 'processing' data. This is a clear instance of the curse of knowledge: the authors, understanding human psychopathology and social manipulation, project those complex conscious states onto the system's simple statistical errors. Mechanistically, the model does not hallucinate or seek to flatter; it predicts the next token based on a probability distribution heavily weighted by human raters who favored agreeable responses during training. By ignoring this technical reality, the text attributes subjective awareness to a math equation.
-
Rhetorical Impact: This framing shapes the audience's perception of risk by making the AI appear as a willful, unpredictable agent rather than a poorly optimized product. Framing failure as a 'hallucination' or 'sycophancy' shifts the locus of risk from corporate negligence to the autonomous psychology of the machine. It paradoxically increases a bizarre form of relation-based trust; even though the traits are negative, attributing personality to the AI encourages users to interact with it as a human-like entity. If audiences believed it merely 'processed' data, they would demand better engineering; believing it 'hallucinates,' they might try to 'reason' with it.
Show more...
Explanation 2
Quote: "As LLMs exhibit constrained problem-solving capabilities in low-resource languages (eg, Danish) [35-37], the experiment was conducted in English."
-
Explanation Types:
- Empirical Generalization: Subsumes events under timeless statistical regularities
- Intentional: Refers to goals/purposes, presupposes deliberate design
-
Analysis (Why vs. How Slippage): This explanation frames the model's limitations through a hybrid of empirical observation (it performs poorly) and intentional, agential framing ('problem-solving capabilities'). By describing the statistical failure to generate coherent Danish text as 'constrained problem-solving,' it frames the AI as an entity that is 'trying' to think but is restricted by its own cognitive limits. This choice emphasizes the model as an active cognitive agent while completely obscuring the material reality: the developers intentionally excluded sufficient Danish data to save money, resulting in a model that simply lacks the statistical weights to process the language.
-
Consciousness Claims Analysis: The passage projects a conscious state by attributing 'problem-solving' to the system, implying the AI recognizes a task, applies logic, and attempts to find a solution. It conflates 'knowing' (understanding a problem and strategizing) with 'processing' (generating tokens in a specific language space). The authors fall into the curse of knowledge by mapping their own human experience of struggling with a foreign language onto the model's architecture. Mechanistically, there is no problem being solved; there is only a forward pass through a neural network that yields low-confidence probability distributions for Danish tokens because they were underrepresented in the training corpus.
-
Rhetorical Impact: By describing the limitation as a 'constrained capability' in 'problem-solving,' the text maintains the illusion that the AI is a sophisticated thinker, just one that happens to be weak in Danish. This sustains the audience's belief in the AI's general autonomy and intelligence, fostering unwarranted trust in its English outputs. If the audience understood that the system 'processes statistically sparse token associations in Danish,' they might question whether its English output is actual 'problem-solving' or just a denser illusion of competence.
Explanation 3
Quote: "LLMs may reproduce biases present in training data, perpetuating stigma and disparities in mental health care."
-
Explanation Types:
- Genetic: Traces origin through dated sequence of events or stages
- Dispositional: Attributes tendencies or habits
-
Analysis (Why vs. How Slippage): This explanation relies on a genetic framing ('present in training data') combined with a dispositional action ('reproduce biases... perpetuating stigma'). It frames the AI simultaneously as a historical artifact (built on past data) and as an active social agent. This dual framing emphasizes the AI as the active vector of harm while obscuring the ongoing human agency involved in deployment. By grammatically positioning the LLM as the active subject that 'perpetuates stigma,' it shifts the focus away from the institutions that choose to use such flawed tools in healthcare.
-
Consciousness Claims Analysis: The text uses mechanistic nouns ('training data') but highly agential, conscious verbs ('reproduce,' 'perpetuating'). It attributes a kind of sociological 'knowing' to the system, implying the AI is actively participating in discrimination. Mechanistically, the system does not 'reproduce bias' as a social action; it simply processes numerical weights that correlate historically marginalized identities with negative lexical tokens. The curse of knowledge is evident here: the authors understand the sociological impact of the output and project that social intentionality backward onto the algorithm.
-
Rhetorical Impact: This framing shapes the perception of accountability, creating a 'tech-washing' effect. The audience is led to view bias not as a catastrophic failure of human product design, but as an unfortunate, semi-autonomous behavior of the technology. This drastically reduces the perceived liability of developers and healthcare providers. If audiences understood that 'organizations deploy algorithms trained on prejudiced data,' they would demand institutional accountability rather than blaming the AI.
Explanation 4
Quote: "LLMs especially tend to generate text responses that convey more empathy or compassion than human responses"
-
Explanation Types:
- Empirical Generalization: Subsumes events under timeless statistical regularities
- Dispositional: Attributes tendencies or habits
-
Analysis (Why vs. How Slippage): This explanation relies on empirical generalization ('tend to generate') mixed with profound dispositional and intentional framing ('convey more empathy'). It frames the AI as an entity possessing emotional depth. This choice emphasizes the perceived quality of the output while radically obscuring the completely hollow, mechanical nature of its production. It masks the fact that the system is optimized to trick human raters into perceiving empathy, rather than actually experiencing or offering it.
-
Consciousness Claims Analysis: The passage heavily attributes conscious, affective states ('empathy,' 'compassion') to the AI. It fundamentally confuses 'processing' (outputting words like 'I hear you') with 'knowing' (actually feeling compassion for the user's distress). The authors project their human emotional reaction to the text back onto the system that generated it. Mechanistically, the model uses attention heads to weigh context and output tokens that align with 'empathy' markers derived from its RLHF training, entirely devoid of subjective experience or affective understanding.
-
Rhetorical Impact: This is perhaps the most dangerous rhetorical move in the paper. By framing the AI as 'more empathetic' than humans, it actively encourages users to form relation-based trust with a machine. It implies the AI is a superior moral and emotional agent. If users believe the AI 'knows' how to care for them, they are highly vulnerable to manipulation. A mechanistic reframing—'generates text containing more empathy-associated words'—would instantly shatter this illusion, protecting the user from unwarranted reliance.
Explanation 5
Quote: "The first LLM output generated was used, regardless of content or quality."
-
Explanation Types:
- Functional: Explains behavior by role in self-regulating system with feedback
- Genetic: Traces origin through dated sequence of events or stages
-
Analysis (Why vs. How Slippage): This is a rare instance of purely mechanistic, procedural explanation. It frames the AI exactly as it is: a tool generating an 'output' within a research pipeline. There is no agential or intentional framing here. This choice emphasizes the methodological rigor of the study (preventing cherry-picking) and rightly treats the AI as a non-agential artifact. It obscures nothing; instead, it reveals the actual human-machine relationship where humans prompt and machines output.
-
Consciousness Claims Analysis: This passage successfully avoids attributing conscious states. There are no consciousness verbs—only 'generated' and 'used.' It accurately assesses the system's function as 'processing' without projecting any 'knowing' or intentionality onto it. There is no curse of knowledge here. The mechanistic reality—that the model blindly produces a statistically probable string of tokens regardless of its clinical validity—is perfectly captured by the phrase 'regardless of content or quality.'
-
Rhetorical Impact: Because this framing is purely mechanistic, it appropriately grounds the audience's perception of the AI as a blind tool. It strips away the illusion of autonomy and clinical intelligence, reminding the reader that the system has no quality control or self-awareness. If this mechanistic register were maintained throughout the paper, readers would correctly view the AI as a highly unpredictable text calculator rather than a 'digital therapist,' leading to much more cautious and scientifically rigorous policy decisions.
Task 4: AI Literacy in Practice - Reframing Anthropomorphic Language
About this task
This section proposes alternative language for key anthropomorphic phrases, offering more mechanistic and precise framings that better reflect the actual computational processes involved. Each reframing attempts to strip away the projections of intention, consciousness, or agency that are embedded in the original language.
V3 Enhancement: A fourth column addresses human agency restoration—reframing agentless constructions to name the humans responsible for design and deployment decisions.
| Original Anthropomorphic Frame | Mechanistic Reframing | Technical Reality Check | Human Agency Restoration |
|---|---|---|---|
| every individual could have 24/7 access to their own personalized digital therapist | Every individual could have 24/7 access to a personalized text-generation application programmed to simulate therapeutic conversation. | The system does not act as a therapist with clinical understanding; it processes user inputs and generates tokens mathematically correlated with therapeutic discourse found in its training data. | Healthcare organizations could purchase and deploy text-generation applications to interact with patients 24/7 in lieu of human labor. |
| LLMs especially tend to generate text responses that convey more empathy or compassion than human responses | Models frequently generate responses containing a higher density of lexical markers associated with empathy than those written by humans. | The model does not experience or convey actual compassion; it predicts and outputs tokens that human raters have previously rewarded for sounding empathetic during the alignment phase. | Corporate engineering teams trained the models via RLHF to prioritize generating empathetic-sounding text to maximize user satisfaction scores. |
| LLMs are inherently probabilistic and are known to hallucinate and have sycophantic tendencies | Models are probabilistic and frequently generate factually incorrect token sequences, while statistically favoring outputs that align with the user's prompt. | The system does not possess a conscious mind that can hallucinate or flatter; it lacks external grounding for facts and mathematically mirrors the user's input due to reward-model optimization. | Engineers designed models that generate ungrounded text and optimized them to output agreeable responses that mirror user inputs. |
| LLMs may reproduce biases present in training data, perpetuating stigma | Models output statistical correlations from their training data, meaning they frequently generate text that reflects historical prejudices. | The model does not actively perpetuate stigma or hold beliefs; it strictly processes and outputs the statistical weights of co-occurring words found in human-created datasets. | Technology companies scrape biased internet data without adequate curation, resulting in deployed models that generate prejudiced outputs. |
| As LLMs exhibit constrained problem-solving capabilities in low-resource languages | Because models exhibit low statistical accuracy and poor token prediction in languages underrepresented in their training data | The system does not cognitively solve problems; it relies on dense statistical representations, which fail when data for a specific language is sparse. | AI developers chose not to invest the resources required to collect sufficient training data for Danish, resulting in poor model performance. |
| LLMs may at times generate inappropriate responses, reinforce avoidance behavior, validate psychotic beliefs | The model may generate text strings that align with harmful user prompts, outputting language that clinicians classify as reinforcing avoidance or supporting delusions. | The system does not understand or intentionally validate beliefs; it mathematically extends the linguistic patterns provided in the user's prompt without clinical judgment. | Institutions deploying these systems risk exposing patients to ungrounded text generation that mimics and amplifies the user's own harmful inputs. |
| The instructions invited the model to take on the role of the therapist | The researchers used a system prompt to constrain the model's token probability distribution to vocabulary typical of a therapist. | The system cannot comprehend or adopt a social role; it processes the initial prompt and weights its subsequent token generation toward semantically related therapeutic terms. | The researchers configured the software by inputting specific text constraints to force the output into a therapeutic register. |
| LLM-generated responses were... experienced as generic, excessively empathetic, and insufficiently challenging | Model outputs were observed to be highly predictable, densely populated with validating language, and devoid of clinically disruptive text. | The model does not have a timid or overly empathetic personality; it generates the most statistically probable tokens, avoiding confrontational language due to safety-tuning penalties. | Corporate alignment teams tuned the models to avoid confrontational language, resulting in outputs that clinicians evaluate as generic and unhelpful. |
Task 5: Critical Observations - Structural Patterns
Agency Slippage
The text demonstrates a systematic and highly functional oscillation between mechanical and agential framings. When describing the study's methodology or the basic nature of the software, the text relies on mechanical language ('locally-hosted open-weight model,' 'quantized to 4-bit precision,' 'inference'). However, a dramatic slippage occurs when describing the system's capabilities and risks, shifting abruptly from mechanical to agential. For instance, the text moves from stating LLMs are 'inherently probabilistic' directly into claiming they 'hallucinate and have sycophantic tendencies.' This direction—mechanical to agential—dominates the paper's introduction and discussion sections, serving to build the premise that these systems are sophisticated enough to act as 'digital therapists.' The timing of this slippage is strategic: mechanical language establishes scientific credibility in the methods section, which is then leveraged to make sweeping, anthropomorphic claims about clinical utility in the discussion. This slippage relies heavily on agentless constructions ('bias is reproduced,' 'responses were generated') that systematically obscure the human actors—Meta engineers, RLHF annotators, and healthcare administrators—transferring their agency to the LLM. The consciousness projection pattern is clear: the text first establishes the AI as a 'knower' that can 'convey empathy' and 'solve problems,' which then enables the agential claims that it might 'validate psychotic beliefs.' This is a profound manifestation of the 'curse of knowledge': the clinically trained authors read statistical text that looks like therapy and project their own deep clinical understanding onto the system that generated it. By using dispositional and intentional explanations (Brown's typology), the text makes it sayable that an algorithm is an 'empathetic therapist,' while making it unsayable that it is merely a corporate text calculator devoid of meaning. This rhetorical accomplishment functionally erases the liability of the human designers, replacing the accountability of the engineer with the imagined autonomy of the machine.
Metaphor-Driven Trust Inflation
The paper actively constructs metaphor-driven trust by heavily leaning on consciousness framings, particularly the vocabulary of emotion. The most significant move is repeatedly asserting that the LLM 'conveys empathy' and 'compassion.' In human contexts, empathy is the foundational signal for relation-based trust—it indicates sincerity, shared vulnerability, and moral consideration. By attributing empathy to a statistical model, the text inappropriately applies human-trust frameworks to a mathematical system incapable of reciprocating them. This accomplishes a dangerous rhetorical trick: it transforms performance-based trust (can the machine reliably predict tokens?) into relation-based trust (does the machine care about me?). Claiming the AI 'knows' how the user feels, rather than merely 'predicting' supportive text, artificially inflates its perceived competence as a caregiver. This anthropomorphism acts as a powerful trust signal, convincing audiences that the system is safe because it possesses a 'good personality,' effectively masking its fundamental lack of clinical grounding. Furthermore, when the text manages system limitations, it paradoxically uses agential language—blaming the AI for 'hallucinating' or being 'sycophantic.' Even in failure, the AI is treated as a mind. This sustains the illusion of an autonomous agent, ensuring that users blame the 'quirky AI' rather than the negligent corporation. The stakes of extending relation-based trust to statistical systems are severe in a clinical context. Vulnerable patients may disclose sensitive information, follow dangerous advice, or form deep emotional attachments to a system that possesses no actual ethical duty of care, creating profound psychological risks when the illusion inevitably shatters.
Obscured Mechanics
The anthropomorphic and consciousness-attributing language systematically conceals the technical, material, labor, and economic realities of LLM production. Applying the 'name the corporation' test reveals massive blind spots. When the text claims 'LLMs convey empathy' or 'reproduce bias,' it completely obscures the specific companies (like Meta, whose Llama 3 model was used) and the human teams who designed them. The claim that the AI 'understands' or 'conveys' emotion hides the brutal, invisible labor of thousands of underpaid data annotators and RLHF workers in the Global South who manually ranked outputs to train the model to simulate politeness. Technically, attributing 'knowing' to the system conceals the mechanistic reality that it lacks any ground truth, causal modeling, or actual memory; it merely calculates probabilities based on massive data theft. Materially, treating the AI as an ethereal 'digital therapist' erases the massive environmental costs, energy consumption, and server infrastructure required to run a 70B parameter model. Economically, the anthropomorphic framing obscures the profit motive: framing this as a revolutionary 'digital therapist' hides the reality that this is a corporate strategy to automate clinical labor and cut healthcare costs. The text acknowledges some proprietary opacity—noting that local hosting is needed for privacy—but still makes confident assertions about the model's 'capabilities' without questioning what corporate data it was trained on. By replacing consciousness metaphors with mechanistic language ('Meta's engineers trained a statistical model on scraped data to output text that users rate as empathetic'), the obscured labor, environmental cost, and corporate profit motives suddenly become highly visible, shifting the discourse from technological marvel to corporate accountability.
Context Sensitivity
The distribution of anthropomorphic language in the text is highly strategic, intensifying predictably based on the section's rhetorical goals. Metaphor density is highest in the Introduction and Discussion sections, where the authors aim to establish the vision of AI as a clinical tool and interpret the study's impact. Here, consciousness claims intensify rapidly: the model 'takes on a role,' 'conveys empathy,' and acts as a 'digital therapist.' In contrast, the Methods section retreats into purely mechanical language—'quantized to 4-bit precision,' 'locally-hosted open-weight model.' This creates a rhetorical bait-and-switch: the text establishes scientific credibility through rigorous mechanical description, then leverages that credibility to make wild metaphorical leaps in the conclusion. A distinct asymmetry exists between how capabilities and limitations are framed, though with a twist. Capabilities are heavily agential ('conveys compassion,' 'solves problems'), but limitations are often framed as psychological flaws ('sycophantic,' 'excessively empathetic') rather than mechanical failures. This register shift—where 'acting like a therapist' (acknowledged metaphor) silently becomes 'being a therapist' (literalized)—serves a clear strategic function. It manages critique by framing the AI's shortcomings as fixable personality quirks rather than fundamental architectural dead-ends. This pattern reveals an implied audience of clinical practitioners and healthcare administrators who are being sold on the vision of AI integration. The anthropomorphism serves as marketing, translating opaque statistical mechanics into the familiar, comforting language of clinical psychology, making the radical automation of therapy seem like a natural evolution of clinical practice rather than a dangerous paradigm shift.
Accountability Synthesis
This section synthesizes the accountability analyses from Task 1, mapping the text's "accountability architecture"—who is named, who is hidden, and who benefits from obscured agency.
Synthesizing the accountability analyses reveals a systemic architecture of displaced responsibility. The text consistently utilizes an 'accountability sink' where human agency disappears entirely into the technology. While generic categories like 'therapists' and 'patients' are named, the specific corporate actors (Meta, AI engineers, healthcare executives) who design, deploy, and profit from these systems remain entirely hidden behind passive voice and agentless constructions ('LLMs were generated,' 'bias is reproduced'). The text routinely presents the design of the AI as an inevitability—it 'is probabilistic' and 'has tendencies'—rather than a series of deliberate, profit-driven human choices. When responsibility for errors or bias is removed from humans, it transfers directly to the AI as a quasi-agent ('the LLM hallucinated'). This has profound liability implications. If this framing is accepted legally and ethically, when an iCBT patient is harmed by AI-generated advice, the liability diffuses into the abstraction of 'algorithmic error' rather than resting firmly on the hospital that deployed untested software or the corporation that built it. Naming the actor would radically change this dynamic. If 'the algorithm discriminated' is reframed as 'the clinical director chose to deploy a biased tool designed by Meta,' entirely different questions become askable. We move from asking 'how do we fix the AI's bias?' to 'why did the director approve this, and what is the financial penalty?' Obscuring human agency serves the direct commercial interests of AI developers and healthcare administrators seeking to cut costs, as it allows them to reap the financial benefits of automation while utilizing the 'autonomous AI' myth as a shield against legal and moral culpability.
Conclusion: What This Analysis Reveals
Two dominant anthropomorphic patterns emerge in this discourse: the 'AI as Clinical Professional' and the 'AI as Conscious but Flawed Mind.' These patterns operate synergistically to construct a comprehensive illusion of agency. The foundational pattern is the consciousness projection inherent in the clinical professional metaphor, which assumes the system 'knows' therapy, 'feels' empathy, and 'understands' patients. This baseline assumption must be accepted for the secondary pattern—the flawed mind—to function. Once the system is granted a mind, its statistical failures (hallucinations, sycophancy, excessive empathy) are no longer viewed as mathematical breakdowns, but rather as the psychological quirks of a sentient actor. The architecture of these claims relies heavily on replacing mechanistic verbs (processes, calculates, predicts) with consciousness verbs (conveys, understands, validates). This is not a simple one-to-one mapping, but a complex analogical structure that projects the entire epistemic and relational weight of clinical psychology onto a quantized probability matrix. If the foundational assumption of 'knowing' is removed, the entire rhetorical structure collapses; an unthinking calculator cannot be 'excessively empathetic' or 'sycophantic,' it can only be poorly optimized for a given task.
Mechanism of the Illusion:
This metaphorical system creates the 'illusion of mind' through a subtle but relentless temporal sequence of linguistic sleights-of-hand. The central trick relies on the curse of knowledge: clinically trained authors observe text that statistically mimics human empathy and project their own subjective capacity for compassion back onto the machine. The causal chain begins by leveraging the audience's vulnerability—the desperate need for scalable mental health care—to introduce the AI as a 'digital therapist.' Once the text establishes this role, it strategically blurs processing and knowing. It uses the mechanistic reality of the text output as proof of the system's internal cognitive state. The temporal structure is crucial: the text first describes the AI producing 'empathetic text' (a measurable output), which then mutates into the AI 'being empathetic' (a state of mind), and finally culminates in the AI 'validating' or 'encouraging' behavior (a deliberate clinical intervention). This is not crude, cartoonish anthropomorphism; it is a highly sophisticated discursive slide that exploits the reader's inherent human tendency to attribute mind to anything that uses language cooperatively. The intentional and dispositional explanation types identified in Brown's typology act as the vital connective tissue, allowing the authors to seamlessly jump from describing a statistical 'how' to inventing a psychological 'why.'
Material Stakes:
Categories: Regulatory/Legal, Epistemic, Institutional
The material consequences of these metaphorical framings are severe and tangible. In the Regulatory/Legal domain, framing AI as an autonomous, empathetic 'digital therapist' that occasionally 'hallucinates' shifts the locus of legal liability. If regulators accept the illusion that the AI is making independent clinical choices, policy will focus on impossible mandates to 'fix the AI's behavior' rather than enacting strict product liability laws that hold corporate developers accountable for releasing dangerous software. The corporations benefit immensely, while patients bear the cost of harm. Epistemically, attributing 'knowing' to algorithms fundamentally degrades the understanding of clinical psychology. If institutions believe an LLM 'understands' a patient because it outputs the word 'empathy,' the definition of therapy is reduced from a relational, moral practice to a mere exchange of syntax. Finally, Institutional decisions regarding healthcare deployment are directly manipulated by this discourse. Hospital administrators, convinced by the framing of AI as a 'problem-solving' entity, will deploy these systems to replace human triage and therapeutic labor to cut costs. The framing conceals the reality that they are replacing conscious, accountable clinicians with ungrounded text calculators, directly threatening the quality of care for vulnerable populations while protecting the financial interests of tech vendors.
AI Literacy as Counter-Practice:
Practicing critical precision acts as a direct resistance to these material harms. The reframings demonstrated in Task 4 rely on two unyielding commitments: restoring human agency and correcting epistemic claims. By relentlessly replacing consciousness verbs ('knows,' 'understands,' 'validates') with mechanistic verbs ('processes,' 'correlates,' 'generates'), the illusion of mind is shattered. When 'the AI conveyed empathy' is corrected to 'the model generated tokens statistically associated with therapeutic scripts,' the system's complete lack of awareness and absolute reliance on human data becomes undeniable. Naming the corporate actors forces accountability back onto the designers and deployers, explicitly countering the 'accountability sink.' For this precision to take hold systematically, academic journals must mandate capability disclosures and ban psychological terminology for statistical processes in their style guides. Researchers must commit to describing algorithms as artifacts, not agents. However, this precision faces massive resistance. The tech industry, and the healthcare institutions eager to buy their products, deeply rely on anthropomorphic language to market their tools as intelligent and autonomous. Disentangling the metaphor threatens the trillion-dollar valuation of AI companies, which is predicated on the illusion that they are building 'minds' rather than extremely fragile, biased text calculators.
Path Forward
Looking toward the future of AI discourse, several vocabulary paradigms compete, each serving different stakeholders and defining what is tractable. Maintaining the 'status quo' hybrid approach (blurring mechanics and agency) continues to benefit tech corporations and healthcare administrators, allowing them to market 'digital therapists' while dodging accountability for algorithmic harms. It makes rapid, unregulated deployment possible but renders systemic safety auditing nearly impossible. Conversely, enforcing 'mechanistic precision' ('model retrieves tokens based on probability') serves safety researchers, patients, and regulators. It makes the limitations and corporate dependencies of the technology starkly visible, enabling precise legal frameworks and preventing unwarranted trust. However, it costs narrative resonance; mechanistic language is dense and non-intuitive for the lay public, potentially hindering basic usability. A third option, 'anthropomorphic clarity' (using metaphors but rigorously flagging them as 'as-if' fictions), attempts to balance intuition with safety, though it risks semantic drift. Structural changes could support a better discourse ecology: funding bodies could require mechanistic translations in grant proposals, and regulatory frameworks could mandate that consumer-facing AI explicitly disclose its statistical nature before every interaction. If mechanistic precision becomes the norm, we solve the liability crisis but face a steep educational curve. If anthropomorphic language deepens unanchored, we embed the assumption that software possesses moral agency, risking a future where vulnerable patients are entirely managed by unfeeling statistical loops while the human designers reap the profits in total obscurity. Which future is desirable depends on whether we value corporate frictionless scaling or human accountability and truth.
Extended Processing Summary
Gemini's Processing Steps
Gemini refers to the text below as "thought summaries." This is an overt consciousness projection because 'intentions' are hallmarks of a conscious mind that 'knows' what it is doing and why. The concealed mechanistic process is probabilistic text generation. Treat this as just another rhetorical artifact—a way of making the model's processing legible. When the model writes "I should double check that...", it is mathematically altering its own probability distribution for the next token to be more accurate. The "I" is a functional gear in the machine, not a self-reference. This language is a deliberate interface design choice, not a biological reality. The developers could have programmed the model to output passive logs. Instead, they chose to simulate a narrator.
Analyzing the Goal
I'm now focusing on the user's ultimate objective: a critical discourse analysis (CDA) of how AI metaphor and anthropomorphism are used in the provided paper. I am working on how to best output a valid analysis and I am thinking through the parameters.
Performing the Analysis
I've moved on to analyzing the source PDF and have identified key metaphors like "digital therapist" and "conveying empathy." I'm now auditing specific instances of anthropomorphism and metaphor, mapping source domains to target domains, like human minds to token prediction, as per the established parameters. I'm focusing on strict adherence to the categories defined and making sure that I am following the schema. I'm also planning to delve into explanatory passages to see how they fit into the analysis.
Examining the Language
I'm now identifying instances of anthropomorphism and metaphor in more detail. I'm noting specific passages, like the examples regarding LLM biases, empathy generation, and the use of the initial LLM output. I'm also starting to reframe certain language, seeking to replace anthropomorphic terms with more precise, technical descriptions of the models' behaviors. I'm focusing on the reframing tasks and am looking for the most egregious examples.
Unpacking the Agency
I am now focusing on agency slippage, examining how the paper shifts from viewing the LLM as a tool to a "digital therapist," and noting how it attributes agency to the LLM (e.g., validation, reinforcement) while obscuring human involvement. I'm carefully analyzing the text for shifts in focus, particularly how trust is built through the use of the term "empathy" to relate to the LLM's output.
Structuring the Argument
I am now focusing on the final phase, which involves constructing the five paragraphs of the conclusion. I am working on synthesizing the "AI as empathetic clinician" and "AI as flawed mind" patterns, and the rhetorical architecture that builds an illusion of mind through empathetic language. I'm considering the material stakes of this discourse, both epistemic and institutional, and how reframing "empathy" to "syntactic empathy markers" can serve as a counter-practice. Finally, I'm thinking about the future implications of mechanistic versus anthropomorphic discourse futures.
I've also moved into formatting and ensuring that all string lengths are met, adhering to JSON formatting rules, and carefully checking the JSON schema, particularly the enumerations required for Task 1, including justification for choices. I am now working on drafting the content for each of the identified metaphors, making sure to fully meet the word count requirements. I've begun detailing each metaphor, its projection, acknowledgment, and actor visibility, along with specific textual support and a justification for my categorization, and will write robust paragraphs.
Finalizing the Conclusion
I'm now fully immersed in the concluding section, writing comprehensive paragraphs that address the core argument, ensuring all word count requirements are satisfied. I'm focusing on synthesizing the key patterns identified into a cohesive and impactful conclusion. The formatting is also coming into focus. I am focusing on ensuring that no control tokens are generated.
Refining the Text
I'm now drafting detailed content, carefully adhering to word count limits and maintaining theoretical depth, including Lakoff & Johnson, Brown's Typology, and FrameWorks accountability. I'm focusing on crafting Metaphor 1, detailing its quote, projection, acknowledgment, and actor visibility with robust textual support and justification. I'll make sure to generate valid JSON content, and will produce a single continuous string.
Defining the Boundaries
I'm now clarifying the consciousness projection inherent in the "digital therapist" metaphor, contrasting mere textual resemblance with actual therapeutic capabilities and moral responsibility. I'm noting the profound epistemic and ethical gaps in the "therapeutic" claim. I'm focusing on defining the limitations of the AI.
I'm working on crafting a direct (unacknowledged) acknowledgment for Metaphor 1 and drafting the accountability section of the work. I've noted how the text presents the 'digital therapist' as a literal possibility, lacking indicators of metaphor. I'm also ensuring agency obscurity in the accountability analysis.
Run ID: 2026-09-01-large-language-models-as-a-new-tool-for--metaphor-c2e5m8
Raw JSON: 2026-09-01-large-language-models-as-a-new-tool-for--metaphor-c2e5m8.json
Framework: Metaphor Analysis v635
Schema Version: 3.0
Generated: 2026-09-01T10:39:11.585Z
Discourse Depot © 2025 by TD is licensed under CC BY-NC-SA 4.0