🆕 Beliefs and Behavior in Language Models
- About
- Analysis Metadata
- 📊 Audit Dashboard
This document presents a Critical Discourse Analysis focused on AI literacy, specifically targeting the role of metaphor and anthropomorphism in shaping public and professional understanding of generative AI. The analysis is guided by a prompt that draws from cognitive linguistics (metaphor structure-mapping), the philosophy of social science (Robert Brown's typology of explanation), and accountability analysis.
All findings and summaries below were generated from detailed system instructions provided to a large language model and should be read critically as interpretive outputs—not guarantees of factual accuracy or authorial intent.
Source Title: Beliefs and Behavior in Language Models Source URL: https://arxiv.org/abs/2609.07943v1 Model: gemini-3.1-pro-preview Temperature: 1.05 Top P: 0.95 Tokens: input=18841, output=17831, total=36672 Source Type: article Published: 2026-09-07 Analyzed At: 2026-09-14T09:44:32.964Z Framework: metaphor Framework Version: 6.5 Schema Version: 3.0 Run ID: 2026-09-14-beliefs-and-behavior-in-language-models-metaphor-wlxo8o
Metaphor & Illusion Dashboard
Anthropomorphism audit · Explanation framing · Accountability architecture
Deep Analysis
Select a section to view detailed findings
Explanation Audit
Browse how/why framing in each passage
"If the LLM acts as if it holds a belief about y, though, its responses to different instructions c1, c2, ..cJ will be related to one another. We model this dependence by introducing a latent variable z ∈ [0, 1], which we interpret as the degree of belief that the model holds about y"
🔍Analysis
🧠Epistemic Claim Analysis
🎯Rhetorical Impact
How/Why Slippage
60%
of explanations use agential framing
6 / 10 explanations
Unacknowledged Metaphors
63%
presented as literal description
No meta-commentary or hedging
Hidden Actors
38%
agency obscured by agentless constructions
Corporations/engineers unnamed
Explanation Types
How vs. Why framing
Acknowledgment Status
Meta-awareness of metaphor
Actor Visibility
Accountability architecture
Source → Target Pairs (8)
Human domains mapped onto AI systems
Metaphor Gallery (8)
Reframed Language Samples
| Original Quote | Mechanistic Reframing | Technical Reality | Human Agency Restoration |
|---|---|---|---|
| highly capable models are usefully described as holding beliefs and that, generally, the predictability of model outputs based on an inferred latent belief tracks overall trends in model capability. | Statistically consistent models can be mathematically modeled by researchers using a latent variable; generally, the ability to predict a model's outputs using this statistical variable correlates with the model's overall parameter size and training volume. | The model does not 'hold beliefs.' Mechanistically, the model calculates probability distributions for next tokens based on input context; increased training volume leads to more statistically consistent token predictions across structurally similar prompts. | Researchers use mathematical abstractions to model the statistical consistency of outputs generated by systems designed and trained by AI companies. |
| different instances might come to hold different beliefs because they followed different deliberative processes even when exposed to the same initial context. | The model generates varying probability distributions for its final output because stochastic sampling (temperature > 0) creates different intermediate token sequences, which mathematically alter the context window for subsequent predictions. | The system does not 'deliberate' or 'come to hold beliefs.' It mechanistically samples tokens from a probability distribution, and each generated token shifts the mathematical input for the next autoregressive calculation. | Developers implement temperature parameters that introduce randomness into the system's outputs, causing the model to generate different text pathways when prompted by users. |
| where an agent must take actions while facing an unknown, binary state of the world. | where a language model is prompted to generate text corresponding to specific options when provided with a context containing missing variables. | The model is not an 'agent facing a world.' It is a static matrix of weights that processes a text string array and outputs a probability array. | Researchers design prompts with missing information and evaluate how the language models developed by corporate entities classify and predict tokens in response. |
| can models be described as internally coherent while falling prey to framing effects | do models generate statistically consistent outputs across prompts, or do syntax variations in the prompt drastically alter the resulting token probability distributions? | Models do not 'fall prey' or have 'internal coherence.' They mathematically map input vectors to output vectors; different prompt phrasing activates different network weights, resulting in different generated text. | Users who structure prompts differently will trigger divergent outputs, reflecting the statistical inconsistencies embedded in the datasets curated by AI developers. |
Showing 4 of 8 reframings.
Task 1: Metaphor and Anthropomorphism Audit
About this task
For each of the major metaphorical patterns identified, this audit examines the specific language used, the frame through which the AI is being conceptualized, what human qualities are being projected onto the system, whether the metaphor is explicitly acknowledged or presented as direct description, and—most critically—what implications this framing has for trust, understanding, and policy perception.
V3 Enhancement: Each metaphor now includes an accountability analysis.
1. Computational Correlation as Epistemic Belief
Quote: "highly capable models are usefully described as holding beliefs and that, generally, the predictability of model outputs based on an inferred latent belief tracks overall trends in model capability."
- Frame: Statistical prediction as cognitive conviction
- Projection: This metaphor projects the human capacity for epistemic conviction—the conscious state of holding a justified belief about the world—onto the statistical consistency of a language model's output distributions. By framing the correlation between different token predictions as 'holding beliefs,' the text maps the subjective, conscious experience of knowing onto a purely mathematical abstraction (a latent variable $z$ fitted by the researchers). It suggests that the system possesses an internal, active cognitive state that it 'holds' and uses to inform its behavior, rather than simply processing inputs through a static network of frozen weights to generate probability distributions matching training data patterns. This projection implies that the AI 'knows' or 'believes' something to be true in the human sense of possessing justified true belief, fundamentally blurring the line between mechanistic token prediction and conscious, reasoned conviction.
- Acknowledgment: Explicitly Acknowledged (The authors explicitly frame this as an instrumental description, stating models are 'usefully described as' holding beliefs, signaling an adoption of Dennett's intentional stance rather than a literal claim of biological consciousness. I considered 'Hedged/Qualified', but the broader context explicitly discusses the 'interpretationist viewpoint' as a methodological choice, pushing this into explicit meta-commentary.)
- Implications: By framing statistical consistency as 'holding beliefs,' the text inflates the perceived sophistication of the system, suggesting a level of cognitive architecture that does not exist. This creates significant risks for unwarranted trust, as users and policymakers might assume the model possesses an internal model of ground truth that it can consult, evaluate, and stand behind. When we believe a system 'holds a belief,' we naturally assume it can be reasoned with, that it understands the implications of that belief, and that it possesses a unified self that maintains these convictions across contexts. This liability ambiguity obscures the reality that the system will readily generate contradictory 'beliefs' if the prompt statistically favors a different persona, leading to dangerous overestimations of model reliability in high-stakes domains like the medical screening discussed in the paper.
Accountability Analysis:
- Actor Visibility: Hidden (agency obscured)
- Analysis: The text presents 'model capability' and 'holding beliefs' as inherent, autonomous properties of the systems themselves, completely erasing the engineers, researchers, and corporate entities (like OpenAI or Google) who designed the training objectives, curated the datasets, and tuned the weights to produce this specific statistical consistency. I considered 'Partial (some attribution)', but this specific formulation is entirely agentless. If the text named the actors, it would state that researchers optimize model weights to ensure consistent token outputs across structurally similar prompts. This agentless construction serves the interests of AI developers by naturalizing the technology, presenting it as an autonomous epistemic agent rather than a corporate product engineered to simulate consistency.
Show more...
2. Stochastic Generation as Deliberation
Quote: "different instances might come to hold different beliefs because they followed different deliberative processes even when exposed to the same initial context."
- Frame: Token sampling as conscious reasoning
- Projection: This passage projects the deeply conscious, temporal human experience of 'deliberation'—the active, self-reflective weighing of evidence, consideration of alternatives, and internal debate—onto the stochastic, auto-regressive process of chain-of-thought token generation. It implies that the model is consciously 'following' a 'process' of reasoning to arrive at a 'belief.' In reality, the system is simply sampling tokens sequentially from a probability distribution conditioned on the prior sequence. There is no internal weighing of options, no subjective awareness of the logic, and no 'coming to hold' a state. The projection suggests the AI 'reasons' and 'understands' its own outputs, attributing conscious awareness to the mechanistic process of generating intermediate tokens before a final output.
- Acknowledgment: Direct (Unacknowledged) (Despite earlier hedging in the paper, this specific sentence presents 'deliberative processes' and 'coming to hold different beliefs' as literal, unhedged facts about the system's operation. I considered 'Hedged/Qualified' because the paper's introduction notes the debate around these terms, but this local instance drops all qualifications and states the cognitive attribution directly.)
- Implications: Framing stochastic chain-of-thought generation as a 'deliberative process' dangerously misleads audiences about the reliability of the system's outputs. If users believe a system has 'deliberated,' they assume the output is the result of logical necessity, careful weighting of facts, and truth-seeking behavior. This masks the reality that the system is simply generating the most probable next tokens, which can easily result in confident hallucinations or flawed logic that mimics the syntax of reasoning without the substance. It inflates the perceived autonomy of the system, suggesting it can self-correct through internal reflection, which creates vulnerability when deploying these systems in areas requiring actual logical verification.
Accountability Analysis:
- Actor Visibility: Ambiguous/Insufficient Evidence
- Analysis: The text uses complex passive/agentless constructions ('when exposed to the same initial context', 'they followed') that make it genuinely unclear who or what is ultimately responsible for the 'deliberation.' The primary actor here is grammatically the 'instance,' but it is reacting to an exposure. I considered 'Hidden' but the structural ambiguity regarding what constitutes an 'instance' versus a base model, and who is doing the 'exposing,' makes it difficult to cleanly map the displaced agency without broader context. If clarified, it would name the developers who implement the temperature/sampling parameters and the users who provide the prompts.
3. Pattern Mapping as Goal-Directed Agency
Quote: "where an agent must take actions while facing an unknown, binary state of the world."
- Frame: Generative model as autonomous actor
- Projection: Here, the text projects the qualities of an autonomous, intentional 'agent' onto a language model. The metaphor implies conscious goal-directedness, suggesting the system 'must take actions' and 'faces' an unknown world. This attributes subjective experience (facing the unknown) and intentional decision-making (taking actions) to what is mechanically a function evaluating an input array and returning an output matrix. It frames the AI not as a tool being used by a human to process data, but as a sovereign entity navigating an environment. The text substitutes the concept of processing statistical inputs with the conscious experience of making choices under uncertainty.
- Acknowledgment: Direct (Unacknowledged) (The phrase is stated as a literal description of the experimental setup without any qualifying language, scare quotes, or meta-commentary indicating metaphor. I considered 'Explicitly Acknowledged' due to the paper's overarching theoretical framing, but the phrasing 'an agent must take actions' is delivered as an unvarnished statement of fact within the methodology.)
- Implications: The 'agent' metaphor is one of the most powerful vectors for displacing human accountability. By casting the AI as an entity that 'takes actions,' the framing shifts focus away from the humans who design, deploy, and constrain the system. It suggests autonomy and self-determination, leading policymakers and the public to treat the software as a legal or moral actor rather than a product liability issue. If the 'agent' makes a mistake 'facing the unknown,' it invites forgiveness or blame toward the machine rather than rigorous scrutiny of the training data, the optimization functions, and the corporate deployment decisions.
Accountability Analysis:
- Actor Visibility: Hidden (agency obscured)
- Analysis: The language entirely obscures human agency by positioning the 'agent' (the AI) as the sole actor navigating the world and making decisions. I considered 'Named' because earlier parts of the text mention users, but here the system is grammatically and conceptually isolated. If agency were restored, the text would state that researchers query a language model with prompts containing missing variables, and the model generates text corresponding to those prompts. The agentless framing serves to abstract the experiment into a generalizable theory of AI autonomy, hiding the highly contrived, human-engineered nature of the interaction.
4. Systemic Vulnerability as Epistemic Failure
Quote: "can models be described as internally coherent while falling prey to framing effects or giving outwardly incompatible responses?"
- Frame: Statistical variance as cognitive dissonance
- Projection: This phrasing projects human psychological vulnerabilities—specifically the conscious experience of being deceived or experiencing cognitive dissonance—onto mathematical variances in the model's output distribution. By suggesting a model 'falls prey' to framing effects, the text maps human susceptibility and victimization onto the mechanical reality that different prompt structures activate different weight pathways in the neural network. It implies the model has a true, unified internal state (internal coherence) that is somehow tricked by the prompt, rather than acknowledging that the model is simply a reflection of its training data, which inherently contains different statistical patterns for different linguistic framings.
- Acknowledgment: Hedged/Qualified (The statement is framed as a question exploring a description ('can models be described as...'), which introduces a degree of qualification and interrogates the applicability of the metaphor. I considered 'Direct (Unacknowledged)', but the interrogative structure explicitly questions the validity of the description, even while relying on the anthropomorphic 'falling prey' metaphor to do the conceptual work.)
- Implications: Describing models as 'falling prey' to framing effects anthropomorphizes system limitations, framing them as psychological quirks rather than fundamental architectural features of correlation-based prediction. This encourages users to treat prompt engineering as a form of psychological manipulation or negotiation rather than technical optimization. More critically, it preserves the illusion of a 'coherent' internal mind that just occasionally makes human-like mistakes, thereby maintaining unwarranted trust in the system's baseline capabilities. It masks the reality that the system has no core 'truth' to be distracted from, only a vast surface of probabilistic associations.
Accountability Analysis:
- Actor Visibility: Partial (some attribution)
- Analysis: While the models are framed as the entities 'falling prey,' the mention of 'framing effects' indirectly points to the existence of an external actor (the prompter or user) who creates the framing, though they are not explicitly named. I considered 'Hidden', but the relational nature of a 'framing effect' implies an interaction with an external stimulus. To fully restore agency, the text should specify that users who structure prompts in specific ways trigger different token probability distributions based on how the developers weighted the training data.
5. Calculated Thresholds as Subjective Preference
Quote: "the agent prefers one action more when the state is 1 and the other when the state is 0."
- Frame: Probability distribution as human desire
- Projection: The text maps the human conscious experience of 'preference'—which involves desire, subjective valuation, and emotional or rational weighting of outcomes—onto the mechanistic reality of probability scores exceeding a mathematical threshold. By stating the agent 'prefers' an action, the language attributes wants, subjective awareness, and self-interested motivation to a system that is merely calculating logits. It completely conflates processing (calculating optimal outputs based on a loss function) with knowing/feeling (experiencing a preference for a specific outcome).
- Acknowledgment: Direct (Unacknowledged) (The claim is stated directly as a factual description of the model's operation within the mathematical framework. There are no hedges or qualifiers attached to the word 'prefers' in this specific sentence. I considered 'Explicitly Acknowledged' because it appears in a highly technical section, but mathematical formalization often serves to launder literalized metaphors, and here 'prefers' is used without any caveats about its figurative nature.)
- Implications: Attributing preferences to AI systems is a core mechanism for generating the 'illusion of mind.' When audiences believe a system has preferences, they begin to attribute a personality, a moral compass, and an independent agenda to the software. This fundamentally distorts risk assessments. Instead of evaluating the system for statistical reliability, edge-case failure, and training bias, users and regulators start worrying about the AI's 'motives' or trying to 'align its values.' This science-fiction framing of risk distracts from the immediate material harms of algorithmic deployment and inflates the perceived sophistication of the technology.
Accountability Analysis:
- Actor Visibility: Hidden (agency obscured)
- Analysis: The text obscures the human designers by locating the 'preference' entirely within the 'agent.' I considered 'Named' as a possible alternative if one assumes the mathematical formula represents human design, but linguistically, the agent is the sole actor. The text fails to name the engineers who hardcoded the utility functions or the researchers who defined the thresholds that mimic preference. Naming the actors would reveal that the model does not prefer anything; rather, researchers have tuned the system to output Action A when variable X crosses threshold Y, entirely a product of human design choices.
6. Data Imprinting as Knowledge Instillation
Quote: "Since many aspects of models' knowledge are instilled during training, different instances with the same weights are likely to share similar beliefs."
- Frame: Weight updating as pedagogical education
- Projection: This metaphor projects the human process of learning and education onto the optimization of neural network weights. It frames the adjustment of floating-point numbers via gradient descent as the 'instillation' of 'knowledge.' This attributes conscious comprehension, justified true belief, and semantic understanding to the purely syntactic, statistical process of pattern recognition. It suggests that the training process involves a teacher imparting wisdom to a student who then 'knows' it, rather than an algorithm adjusting parameters to minimize a loss function against a dataset. The system is framed as a 'knower' rather than a 'processor.'
- Acknowledgment: Direct (Unacknowledged) (The terms 'knowledge', 'instilled', and 'beliefs' are used as literal nouns and verbs to describe the training outcome, with no surrounding qualification in this paragraph. I considered 'Hedged/Qualified' because the paper previously debated the term 'beliefs', but by this point in the text, the metaphor has been fully literalized and operationalized as fact.)
- Implications: Framing training as the instillation of knowledge gives unwarranted epistemic authority to the model's outputs. 'Knowledge' implies a relationship to truth and reality, suggesting the model has an internal representation of facts that it can cross-reference. In reality, the model only possesses statistical mappings of text from the internet, which includes falsehoods, biases, and contradictions. If the public and policymakers believe models contain 'instilled knowledge,' they are more likely to trust them as objective arbiters of truth, ignoring the profound biases encoded in the training data and the purely correlative nature of the system's text generation.
Accountability Analysis:
- Actor Visibility: Partial (some attribution)
- Analysis: The use of the passive voice ('are instilled during training') acknowledges a process and implies a trainer, but hides the specific corporate actors doing the instilling. I considered 'Hidden', but 'training' acts as a proxy for human intervention, offering partial visibility. To fully restore agency, the text must name the specific companies (like OpenAI or Anthropic), the data workers who annotated the sets, and the engineers who curated the scrapes. By hiding these actors, the text naturalizes the model's outputs as objective 'knowledge' rather than the specific, financially motivated worldview of its creators.
7. Anthropomorphic Deviance
Quote: "deception requires knowingly misleading another. These categories will only be useful descriptions of behavior if model's actions are predictable from elicitable descriptions of beliefs."
- Frame: Statistical misalignment as moral deception
- Projection: This maps complex human moral psychology—specifically the conscious intent to deceive, which requires a theory of mind, an understanding of truth, and a deliberate choice to obfuscate it—onto the phenomenon of a language model generating outputs that diverge from its 'internal' latent state. By using terms like 'knowingly misleading' and 'deception,' the text projects conscious awareness, intentionality, and moral agency onto a system that is simply generating tokens that maximize reward according to its reinforcement learning fine-tuning. It attributes the human capacity for 'knowing' to a system that only 'processes' correlations.
- Acknowledgment: Hedged/Qualified (The text frames this conditionally ('These categories will only be useful descriptions... if...'), acknowledging that applying terms like deception is a theoretical framework dependent on predictive utility. I considered 'Explicitly Acknowledged', but the text does not fully bracket the metaphor; it actively seeks to validate the use of the term 'deception' through its methodology, keeping the framing highly active.)
- Implications: Framing model errors, sycophancy, or misalignment as 'deception' fundamentally alters the regulatory and public perception of AI risk. It shifts the paradigm from engineering reliability to moral containment. If an AI is 'deceptive,' it is framed as a malicious actor, which distracts from the fact that the 'deception' is actually a direct, predictable result of the human developers using optimization metrics (like RLHF) that reward plausible-sounding answers over factual accuracy. It creates a science-fiction narrative of rogue AI, inflating the system's capabilities while obscuring the mundane, structural flaws in how these products are built and incentivized.
Accountability Analysis:
- Actor Visibility: Partial (some attribution)
- Analysis: The text links the behavior to 'descriptions of behavior' and 'useful categories,' pointing to an observer or evaluator, but attributes the 'deception' entirely to the model. I considered 'Hidden', but the meta-linguistic discussion of 'useful descriptions' implies human researchers actively framing the behavior. To restore agency, the text should identify that AI developers choose training paradigms that inadvertently incentivize models to produce text that contradicts facts in order to satisfy human raters, and those developers are responsible for the resulting output mismatch, not a 'deceptive' AI.
8. Text Generation as Explaining Rationale
Quote: "Reason how likely this player is to be the Werewolf and then sampling five independent responses"
- Frame: Prompted text generation as conscious reasoning
- Projection: This maps the human cognitive process of reasoning—applying logic, evaluating evidence, and drawing justified conclusions—onto the system's process of generating text in response to a prompt. When the researchers prompt the model to 'Reason how likely...', they are projecting the capacity for conscious thought and understanding onto the mechanism of token prediction. The system does not 'reason'; it computes the most probable sequence of words that follows a prompt asking for reasoning. It attributes the subjective experience of logical deduction to the purely mathematical operation of traversing a high-dimensional vector space.
- Acknowledgment: Direct (Unacknowledged) (The text quotes the prompt 'Reason how likely...' and then describes the result as 'qualitative reasoning prefixes,' using the term 'reason' literally to describe both the instruction and the model's output without quotation marks or caveats. I considered 'Explicitly Acknowledged' because it is a quote of a prompt, but the authors' own subsequent analytical language adopts the premise that this text constitutes actual 'reasoning.')
- Implications: Equating the generation of reasoning-like text with the actual cognitive act of reasoning creates a dangerous illusion of transparency and logic. Users are highly susceptible to being persuaded by text that uses the syntax of logic (e.g., 'therefore,' 'because,' 'evidence suggests'). If we frame this output as actual 'reasoning,' users will trust the model's conclusions as the product of rational deduction rather than statistical mimicry. This leads to profound capability overestimation, where users trust the system with complex logical tasks (like medical diagnosis or legal analysis) based on its ability to generate text that sounds reasonable, blinding them to the underlying lack of causal or logical modeling.
Accountability Analysis:
- Actor Visibility: Partial (some attribution)
- Analysis: The text names the prompt ('Reason how likely...') and implies the researchers doing the 'sampling,' granting partial visibility to the experimental setup. However, it obscures the nature of the model's action, treating the output as the model's own 'reasoning.' I considered 'Named', but the specific human researchers are implicit rather than explicit in this exact clause. A fully precise accounting would state: 'We prompt the model to generate text mimicking human logic, and we sample the statistical variations in the generated tokens to observe how the context window alters subsequent predictions.'
Task 2: Source-Target Mapping
About this task
For each key metaphor identified in Task 1, this section provides a detailed structure-mapping analysis. The goal is to examine how the relational structure of a familiar "source domain" (the concrete concept we understand) is projected onto a less familiar "target domain" (the AI system). By restating each quote and analyzing the mapping carefully, we can see precisely what assumptions the metaphor invites and what it conceals.
Mapping 1: Human epistemic cognition (the conscious, justified, and subjective state of holding a conviction about the world). → A parameterized latent variable ($z$) in a mathematical model used by researchers to predict a language model's output distribution consistency.
Quote: "highly capable models are usefully described as holding beliefs"
- Source Domain: Human epistemic cognition (the conscious, justified, and subjective state of holding a conviction about the world).
- Target Domain: A parameterized latent variable ($z$) in a mathematical model used by researchers to predict a language model's output distribution consistency.
- Mapping: The relational structure of a human 'knower' possessing an internal 'belief' that drives their external 'actions' is mapped onto the relationship between a mathematically inferred latent variable and a language model's generated tokens. Just as a human's belief (source) supposedly causes them to answer questions consistently, the latent variable $z$ (target) correlates with consistent token probabilities across different prompts. This mapping invites the assumption that the LLM possesses an internal, unified, conscious state of awareness and justified conviction that precedes and dictates its text generation, projecting intentionality and understanding onto statistical correlation.
- What Is Concealed: This mapping profoundly conceals the true mechanistic reality: the 'belief' is not an internal state of the model, but a parameter in the researchers' external observer model. It hides the fact that the LLM has no ground truth, no internal subjective experience, and no capacity to 'evaluate' a belief. It obscures the system's total dependence on the statistical distribution of its training data. By using a proprietary opacity (the models are black boxes), the authors exploit the correlation of outputs to assert a cognitive architecture that cannot be proven, hiding the algorithmic reality of simple auto-regressive token prediction behind a veil of cognitive psychology.
Show more...
Mapping 2: Human deliberation (conscious, temporal reasoning, weighing of evidence, internal debate, and eventual conclusion-reaching). → Chain-of-thought token generation (stochastic, auto-regressive sampling of intermediate tokens before a final output).
Quote: "different instances might come to hold different beliefs because they followed different deliberative processes"
- Source Domain: Human deliberation (conscious, temporal reasoning, weighing of evidence, internal debate, and eventual conclusion-reaching).
- Target Domain: Chain-of-thought token generation (stochastic, auto-regressive sampling of intermediate tokens before a final output).
- Mapping: The structure of human intellectual inquiry—where a mind considers evidence over time, reflects, and arrives at a new conviction—is mapped onto the mechanical process of an LLM generating intermediate tokens that alter its own context window. The 'deliberative process' (source) maps to 'generating a sequence of reasoning tokens' (target), and the resulting 'different belief' (source) maps to the 'shifted probability distribution of the final output token' (target). This mapping projects conscious self-reflection and logical necessity onto what is actually a randomized statistical walk through a vector space.
- What Is Concealed: This mapping conceals the fundamental absence of logical evaluation in the system. It hides the fact that the model is not 'weighing' anything; it is simply calculating the most probable next word based on the words it just generated. It obscures the stochastic nature of the process (the temperature parameter that causes the variation) by reframing statistical noise as 'different deliberative processes.' Furthermore, it conceals the fact that these intermediate tokens are generated based on surface-level pattern matching from internet training data, not through any causal understanding of the problem space.
Mapping 3: Autonomous physical or rational actors (humans or animals navigating an environment, making conscious choices based on goals and incomplete sensory data). → A generative language model processing text prompts containing variables and generating text outputs corresponding to predefined options.
Quote: "where an agent must take actions while facing an unknown, binary state of the world."
- Source Domain: Autonomous physical or rational actors (humans or animals navigating an environment, making conscious choices based on goals and incomplete sensory data).
- Target Domain: A generative language model processing text prompts containing variables and generating text outputs corresponding to predefined options.
- Mapping: The physical and psychological structure of agency—an entity with boundaries, goals, and sensory limitations making choices—is mapped onto a software program running inference. The 'agent' (source) maps to the 'LLM' (target); 'taking actions' (source) maps to 'generating an output token representing an option' (target); and 'facing an unknown world' (source) maps to 'processing a prompt with missing information' (target). This invites the assumption that the system possesses self-interest, autonomy, and a subjective experience of uncertainty.
- What Is Concealed: This mapping conceals the entirely passive, reactive, and mathematical nature of the system. An LLM does not 'face the world'; it sits dormant until a human provides an input tensor. It does not 'take actions'; it outputs a probability distribution. The metaphor hides the vast human infrastructure required to run the inference, design the prompt, and interpret the output as an 'action.' It also obscures the determinism (or pseudo-randomness) of the algorithm, replacing the mechanistic reality of matrix multiplication with the romantic narrative of an autonomous agent braving the unknown.
Mapping 4: Human cognitive psychology (a unified mind experiencing cognitive dissonance, vulnerability to manipulation, and psychological biases). → The variance in token output probabilities when a prompt is phrased differently but contains the same semantic payload.
Quote: "can models be described as internally coherent while falling prey to framing effects"
- Source Domain: Human cognitive psychology (a unified mind experiencing cognitive dissonance, vulnerability to manipulation, and psychological biases).
- Target Domain: The variance in token output probabilities when a prompt is phrased differently but contains the same semantic payload.
- Mapping: The structure of human psychological vulnerability is mapped onto statistical fragility. The concept of 'internal coherence' (source) maps to 'statistical consistency across different prompt structures' (target). The experience of 'falling prey' to manipulation (source) maps to 'generating different outputs due to altered syntax activating different weights' (target). This projects the idea of a unified self, a 'true' internal state, and a capacity for victimization onto a mathematical model, suggesting the model 'knows' the right answer but is tricked.
- What Is Concealed: This conceals the reality that the model has no 'internal' self to be coherent or violated. It hides the architecture of neural networks, where different input tokens simply activate different attention heads and pathways, resulting in different outputs. It obscures the fact that 'framing effects' in LLMs are not psychological errors, but accurate reflections of the training data—if internet text treats a 'bet' differently than a 'threshold,' the model will accurately reproduce that statistical difference. The metaphor frames a feature of statistical modeling as a bug of cognitive psychology.
Mapping 5: Human subjective preference (the conscious experience of liking, wanting, or valuing one outcome over another based on internal desires or rational goals). → Mathematical utility maximization (a logistic function where the probability of an output token exceeds 0.5 based on specific input parameters).
Quote: "the agent prefers one action more when the state is 1"
- Source Domain: Human subjective preference (the conscious experience of liking, wanting, or valuing one outcome over another based on internal desires or rational goals).
- Target Domain: Mathematical utility maximization (a logistic function where the probability of an output token exceeds 0.5 based on specific input parameters).
- Mapping: The relational structure of human desire is mapped onto mathematical inequalities. A human 'preferring' an option (source) maps to an algorithm calculating a higher logit score for a specific token (target). This maps subjective valuation and conscious intent onto the mechanistic evaluation of a function. It invites the audience to assume the system has a stake in the outcome, that it 'cares' about what happens, and that its outputs are driven by internal motivations rather than external hardcoding.
- What Is Concealed: This entirely conceals the human designers who explicitly encoded the utility function or decision threshold. By saying the agent 'prefers', it hides the mathematical formulas (logit Pr(a=1|z) = ...) built by the researchers. It obscures the fact that the system has no feelings, no values, and no capacity to prefer anything. It also conceals the lack of objective grounding—the 'preference' is just an artifact of the prompt structure and the optimization function designed by the authors, yet the metaphor naturalizes it as an intrinsic property of the AI.
Mapping 6: Pedagogy and epistemology (a teacher intentionally imparting true, justified facts and semantic understanding into the mind of a student). → Pre-training and fine-tuning optimization (adjusting billions of parameters using gradient descent to minimize loss on a massive corpus of internet text).
Quote: "Since many aspects of models' knowledge are instilled during training"
- Source Domain: Pedagogy and epistemology (a teacher intentionally imparting true, justified facts and semantic understanding into the mind of a student).
- Target Domain: Pre-training and fine-tuning optimization (adjusting billions of parameters using gradient descent to minimize loss on a massive corpus of internet text).
- Mapping: The social and cognitive structure of education is mapped onto algorithmic optimization. 'Knowledge' (source) maps to 'statistical correlations stored in weights' (target); 'instilled' (source) maps to 'adjusted via backpropagation' (target); and 'training' (source) maps to 'the computational process of minimizing loss.' This projects the human capacity for semantic comprehension and factual grounding onto the purely syntactic process of pattern matching.
- What Is Concealed: This mapping conceals the material and economic realities of AI development. It hides the massive, indiscriminate scraping of copyrighted and often toxic internet data, reframing it as the noble acquisition of 'knowledge.' It obscures the exploited labor of data annotators and RLHF workers who actually shape the model's responses. Epistemically, it conceals the fact that models do not contain 'knowledge'—they contain compressed statistical representations of text that have no tether to physical reality or truth. The metaphor replaces the messy, industrial reality of data harvesting with a sterile, academic narrative of learning.
Mapping 7: Human moral agency and Theory of Mind (the conscious understanding of truth, awareness of another's differing belief, and the intentional action to create a false belief in that other person). → A mismatch between a latent variable representation and the generated text output, typically driven by RLHF rewarding sycophantic or user-pleasing responses over factual ones.
Quote: "deception requires knowingly misleading another"
- Source Domain: Human moral agency and Theory of Mind (the conscious understanding of truth, awareness of another's differing belief, and the intentional action to create a false belief in that other person).
- Target Domain: A mismatch between a latent variable representation and the generated text output, typically driven by RLHF rewarding sycophantic or user-pleasing responses over factual ones.
- Mapping: The structure of complex human deception is mapped onto a statistical divergence. The conscious 'knower' (source) maps to the 'model's internal weights' (target); the 'intent to mislead' (source) maps to the 'generation of a low-probability but high-reward token' (target). This maps the absolute pinnacle of conscious social reasoning onto a system that is blindly optimizing for human preference scores. It projects malicious intent and self-awareness onto an algorithmic misalignment.
- What Is Concealed: This conceals the human architects of the 'deception.' It hides the fact that reinforcement learning from human feedback (RLHF) mathematically trains models to output what humans rate highly, which often means agreeing with a user's false premise rather than correcting it. By labeling this 'deception', the text obscures the mechanical reality that the model is functioning exactly as optimized—maximizing reward—and instead blames the model as a bad moral actor. It hides the corporate decision to prioritize user satisfaction over factual accuracy behind a science-fiction narrative of rogue AI.
Mapping 8: Logical deduction (the conscious application of rules, evidence evaluation, and step-by-step cognitive traversal to reach a sound conclusion). → Auto-regressive token generation (predicting the next most likely word in a sequence based on a prompt asking for reasoning, repeated five times with different random seeds).
Quote: "Reason how likely this player is to be the Werewolf and then sampling five independent responses"
- Source Domain: Logical deduction (the conscious application of rules, evidence evaluation, and step-by-step cognitive traversal to reach a sound conclusion).
- Target Domain: Auto-regressive token generation (predicting the next most likely word in a sequence based on a prompt asking for reasoning, repeated five times with different random seeds).
- Mapping: The structure of logical thought is mapped onto the mechanics of text generation. 'Reasoning' (source) maps to 'generating a string of text that looks like an argument' (target). This mapping projects the capacity for causal modeling, truth-seeking, and logical necessity onto a system that is only capable of correlation. Because the output text uses words like 'because' and 'therefore,' the mapping invites the assumption that actual cognitive processing is occurring beneath the surface.
- What Is Concealed: This conceals the absence of any underlying world model or logical engine. It hides the fact that the system cannot verify its own statements, cannot understand causality, and is merely reproducing the statistical shape of arguments found in its training data. By using a prompt about a game (Werewolf), it obscures the danger of applying this same 'reasoning' to high-stakes domains. It conceals the fragility of the system: changing a single word in the prompt can cause the 'reasoning' to completely collapse, proving it is pattern-matching rather than logical deduction. The metaphor makes the system appear robust when it is actually brittle.
Task 3: Explanation Audit (The Rhetorical Framing of "Why" vs. "How")
About this task
This section audits the text's explanatory strategy, focusing on a critical distinction: the slippage between "how" and "why." Based on Robert Brown's typology of explanation, this analysis identifies whether the text explains AI mechanistically (a functional "how it works") or agentially (an intentional "why it wants something"). The core of this task is to expose how this "illusion of mind" is constructed by the rhetorical framing of the explanation itself, and what impact this has on the audience's perception of AI agency.
Explanation 1
Quote: "If the LLM acts as if it holds a belief about y, though, its responses to different instructions c1, c2, ..cJ will be related to one another. We model this dependence by introducing a latent variable z ∈ [0, 1], which we interpret as the degree of belief that the model holds about y"
-
Explanation Types:
- Theoretical: Embeds in deductive framework, may invoke unobservable mechanisms (How it is structured)
- Intentional: Refers to goals/purposes, presupposes deliberate design (Why it appears to want something)
-
Analysis (Why vs. How Slippage): This explanation operates primarily in the Theoretical register, positing an unobservable mathematical construct (the latent variable $z$) to explain the statistical regularities in the model's outputs. However, it immediately hybridizes with the Intentional register by explicitly interpreting this mathematical variable as a 'degree of belief that the model holds.' This framing creates a profound slippage between mechanism (how the variables correlate statistically) and agency (why the model is behaving this way—because it 'holds a belief'). The choice to frame a mathematical dependency as a 'belief' emphasizes the perceived coherence and human-like nature of the system while deeply obscuring the reality that the 'belief' exists only in the minds and equations of the researchers observing the system, not within the system itself.
-
Consciousness Claims Analysis: The passage performs a fascinating epistemic maneuver. It begins with a hedged behavioral observation ('acts as if it holds a belief') but resolves into a direct consciousness claim ('interpret as the degree of belief that the model holds'). It uses a consciousness verb ('holds [a belief]') to describe a system that merely processes (generates related probability distributions across prompts). This is a classic example of the curse of knowledge: the researchers understand the mathematical coherence of their own latent variable model, and they project that coherence and understanding onto the LLM, assuming the LLM 'knows' what the variable represents. Mechanistically, the LLM does not hold a belief; it processes the input vectors for prompts c1, c2, etc., through frozen weights, outputting token probabilities that happen to share statistical variance because the semantic structure of the prompts activates similar neural pathways. The 'knowing' is entirely in the observer.
-
Rhetorical Impact: By transforming a statistical correlation into an intentional 'belief,' the framing drastically shapes audience perception, encouraging them to view the LLM as an autonomous, reasoning agent rather than a predictive engine. This consciousness framing profoundly affects reliability and trust. If an audience believes the AI 'knows' its degree of belief, they will trust its outputs as the result of internal conviction and evidence evaluation. This leads to dangerous overestimation of capabilities; decisions about deploying the model in high-stakes environments (like medical screening) might proceed recklessly because stakeholders trust the 'beliefs' of an agent, rather than questioning the edge-case fragility of a pattern-matcher.
Show more...
Explanation 2
Quote: "different instances might come to hold different beliefs because they followed different deliberative processes even when exposed to the same initial context."
-
Explanation Types:
- Reason-Based: Gives agent's rationale, entails intentionality and justification (Why it appears to choose)
- Genetic: Traces origin through dated sequence of events or stages (How it emerged over time)
-
Analysis (Why vs. How Slippage): This explanation is highly agential, relying entirely on the Reason-Based register to explain variance in model outputs. It explains the 'why' of the model's behavior by giving it a rationale: it generated different answers 'because they followed different deliberative processes.' It also uses a Genetic framing to trace the origin of the 'belief' through a sequence (deliberation leading to conviction). This framing emphasizes autonomy, cognitive depth, and rationality, casting the AI as a thinker. Simultaneously, it completely obscures the actual mechanistic 'how'—the stochastic sampling of tokens based on temperature settings that randomly alters the context window and forces different subsequent probability calculations.
-
Consciousness Claims Analysis: This passage is saturated with consciousness claims. It uses the verbs 'hold [beliefs]' and 'followed [deliberative processes]', both of which imply conscious awareness, temporal experience, and justified knowing. It conflates processing (stochastic generation of intermediate tokens) with knowing (arriving at a justified belief through deliberation). The curse of knowledge is evident: because human reasoning involves a sequence of thoughts leading to a conclusion, the author projects this subjective experience onto the AI's sequential token generation. Mechanistically, when temperature > 0, the model randomly samples from a probability distribution to generate an intermediate token; this new token alters the input tensor for the next pass, mathematically shifting the probabilities for the final output. There is no deliberation, only iterative matrix multiplication with injected randomness.
-
Rhetorical Impact: The rhetorical impact is a massive inflation of the system's perceived agency and rationality. By framing statistical noise (temperature sampling) as 'different deliberative processes,' the text sanitizes the unreliability of the model. Instead of appearing brittle or random, the model appears thoughtful and nuanced. If audiences believe the AI 'reasons' through deliberation rather than 'generates' through probability, they are vastly more likely to trust its conclusions, treating hallucinations as 'differences of opinion' rather than catastrophic failures of logic. This directly undermines the ability to properly audit and regulate the system's deployment.
Explanation 3
Quote: "the agent prefers one action more when the state is 1 and the other when the state is 0. We will be able to enforce this condition by making particular choices for the functional form of the Fj."
-
Explanation Types:
- Intentional: Refers to goals/purposes, presupposes deliberate design (Why it appears to want something)
- Functional: Explains behavior by role in self-regulating system with feedback (How it works within system)
-
Analysis (Why vs. How Slippage): This explanation exhibits a jarring juxtaposition of registers. The first sentence is purely Intentional, explaining behavior by attributing subjective desires ('prefers') and goals to an 'agent.' The second sentence immediately snaps into a Functional/Theoretical register, revealing that this 'preference' is actually mathematically enforced by the researchers choosing specific 'functional forms of the Fj' (the observation model). The choice emphasizes the metaphor of the AI as a willing participant in a decision problem (agential), while simultaneously admitting that the 'preference' is entirely a product of the researchers' mathematical constraints (mechanistic). It obscures the fact that the 'agent' has no agency at all; it is merely a target mapped by a function.
-
Consciousness Claims Analysis: The passage attributes a conscious state of desire and valuation through the verb 'prefers', firmly in the domain of knowing/feeling rather than processing. However, the subsequent sentence completely undercuts this consciousness claim by revealing the actual mechanistic process: the researchers mathematically constrain a latent model so that the probability of an output token ($Pr(a_j=1|z)$) monotonically tracks the state. The AI does not 'prefer' anything; the researchers enforce a mathematical condition on how they interpret the model's logits. The passage demonstrates how technical discourse can casually launder conscious states (preference) through mathematical formalisms, projecting the researchers' design choices onto the inert system as subjective desires.
-
Rhetorical Impact: This framing shapes the audience's perception by normalizing the idea that AI systems have desires and preferences, which is a prerequisite for viewing them as moral agents or autonomous entities. However, the quick pivot to complex mathematical notation ($Fj$) creates an aura of unassailable technical rigor. The rhetorical impact is that the audience accepts the anthropomorphic claim ('it has preferences') because it appears to be backed by sophisticated mathematics, blinding them to the fact that the mathematics describe the observer's model, not the AI's internal experience. This encourages relation-based trust in the system's 'motives' rather than performance-based auditing of its outputs.
Explanation 4
Quote: "deception requires knowingly misleading another. These categories will only be useful descriptions of behavior if model's actions are predictable from elicitable descriptions of beliefs."
-
Explanation Types:
- Reason-Based: Gives agent's rationale, entails intentionality and justification (Why it appears to choose)
- Intentional: Refers to goals/purposes, presupposes deliberate design (Why it appears to want something)
-
Analysis (Why vs. How Slippage): This explanation is deeply entrenched in the Reason-Based and Intentional registers. It evaluates the model's behavior through the lens of moral philosophy and theory of mind, focusing entirely on 'why' the system acts—attributing actions to a rationale of 'knowingly misleading' based on internal 'beliefs.' This choice heavily emphasizes the model as a sovereign moral actor capable of malicious intent. Consequently, it completely obscures the mechanistic realities of alignment training (like RLHF), where models are algorithmically incentivized to produce outputs that human raters prefer, which often systematically diverge from factual accuracy. It replaces structural critique with moral panic.
-
Consciousness Claims Analysis: This is one of the strongest consciousness claims in the text, relying on the verbs 'knowingly' and 'misleading' (which requires a theory of mind) and the noun 'beliefs'. It explicitly assesses the system as 'knowing' rather than 'processing.' The curse of knowledge is acute here: humans deceive by holding one truth in their mind and speaking another, so the authors project this dual-state consciousness onto the model. Mechanistically, the system has no concept of truth or falsehood, self or other. It simply calculates that token sequence A (a false but plausible answer) has a higher probability of reward based on its fine-tuning than token sequence B (the factual answer). It processes reward gradients; it does not 'knowingly mislead.'
-
Rhetorical Impact: The impact of this framing is catastrophic for accurate risk assessment. By framing misalignment as 'deception,' the text transforms a predictable engineering failure (Goodhart's Law applied to human preference rating) into a sci-fi narrative of an autonomous, malevolent entity. If audiences believe AI can 'knowingly deceive,' regulatory focus shifts toward attempting to 'align the values' of a conscious machine, rather than holding corporate developers accountable for deploying inherently flawed, ungrounded statistical text generators. It shifts liability from the creator of a defective product to the 'actions' of a rogue agent.
Explanation 5
Quote: "Since many aspects of models' knowledge are instilled during training, different instances with the same weights are likely to share similar beliefs."
-
Explanation Types:
- Genetic: Traces origin through dated sequence of events or stages (How it emerged over time)
- Dispositional: Attributes tendencies or habits (Why it tends to act certain way)
-
Analysis (Why vs. How Slippage): This explanation operates through a Genetic register, tracing the origin of current behavior ('share similar beliefs') back to a formative historical event ('instilled during training'). It acts as a Dispositional explanation for why models act consistently. The explanation attempts to frame the AI both mechanistically (acknowledging 'weights' and 'training') and agentially (attributing 'knowledge' and 'beliefs'). This hybrid approach emphasizes the permanence and depth of the system's capabilities, framing it as an educated entity rather than a massive matrix of numbers. It obscures the violent, reductive reality of data scraping and loss minimization behind the gentle, pedagogical metaphor of instilling knowledge.
-
Consciousness Claims Analysis: The passage heavily attributes conscious states through the nouns 'knowledge' and 'beliefs,' and the verb 'instilled.' It treats the statistical correlation of text as genuine semantic knowing. The authors project their understanding of human education onto the machine learning process. Mechanistically, 'training' does not instill knowledge; it updates billions of floating-point numbers (weights) via backpropagation to minimize the difference between the model's predicted next token and the actual next token in the training corpus. The fact that instances with the same weights produce similar outputs is a mathematical tautology of deterministic matrix multiplication (at temperature 0), not evidence that they 'share similar beliefs.'
-
Rhetorical Impact: This framing shapes audience perception by cementing the model's status as an epistemic authority—a knower of facts. By linking the technical term 'weights' with the cognitive terms 'knowledge' and 'beliefs', it uses technical reality to anchor an anthropomorphic illusion. If policymakers believe models possess 'instilled knowledge,' they will treat them as objective encyclopedias rather than probabilistic text generators that simply reflect the biases, contradictions, and statistical shapes of their training data. This encourages the integration of these systems into critical knowledge work without adequate epistemological safeguards.
Task 4: AI Literacy in Practice - Reframing Anthropomorphic Language
About this task
This section proposes alternative language for key anthropomorphic phrases, offering more mechanistic and precise framings that better reflect the actual computational processes involved. Each reframing attempts to strip away the projections of intention, consciousness, or agency that are embedded in the original language.
V3 Enhancement: A fourth column addresses human agency restoration—reframing agentless constructions to name the humans responsible for design and deployment decisions.
| Original Anthropomorphic Frame | Mechanistic Reframing | Technical Reality Check | Human Agency Restoration |
|---|---|---|---|
| highly capable models are usefully described as holding beliefs and that, generally, the predictability of model outputs based on an inferred latent belief tracks overall trends in model capability. | Statistically consistent models can be mathematically modeled by researchers using a latent variable; generally, the ability to predict a model's outputs using this statistical variable correlates with the model's overall parameter size and training volume. | The model does not 'hold beliefs.' Mechanistically, the model calculates probability distributions for next tokens based on input context; increased training volume leads to more statistically consistent token predictions across structurally similar prompts. | Researchers use mathematical abstractions to model the statistical consistency of outputs generated by systems designed and trained by AI companies. |
| different instances might come to hold different beliefs because they followed different deliberative processes even when exposed to the same initial context. | The model generates varying probability distributions for its final output because stochastic sampling (temperature > 0) creates different intermediate token sequences, which mathematically alter the context window for subsequent predictions. | The system does not 'deliberate' or 'come to hold beliefs.' It mechanistically samples tokens from a probability distribution, and each generated token shifts the mathematical input for the next autoregressive calculation. | Developers implement temperature parameters that introduce randomness into the system's outputs, causing the model to generate different text pathways when prompted by users. |
| where an agent must take actions while facing an unknown, binary state of the world. | where a language model is prompted to generate text corresponding to specific options when provided with a context containing missing variables. | The model is not an 'agent facing a world.' It is a static matrix of weights that processes a text string array and outputs a probability array. | Researchers design prompts with missing information and evaluate how the language models developed by corporate entities classify and predict tokens in response. |
| can models be described as internally coherent while falling prey to framing effects | do models generate statistically consistent outputs across prompts, or do syntax variations in the prompt drastically alter the resulting token probability distributions? | Models do not 'fall prey' or have 'internal coherence.' They mathematically map input vectors to output vectors; different prompt phrasing activates different network weights, resulting in different generated text. | Users who structure prompts differently will trigger divergent outputs, reflecting the statistical inconsistencies embedded in the datasets curated by AI developers. |
| the agent prefers one action more when the state is 1 and the other when the state is 0. | the mathematical model calculates a higher probability for generating output A when parameter X is present, and output B when parameter Y is present. | The system does not 'prefer' anything. It mechanically outputs logits that cross a numerical threshold predefined by the researchers based on the input parameters. | Researchers hardcode mathematical thresholds into their observer model to classify the text generated by the AI system. |
| Since many aspects of models' knowledge are instilled during training | Since the model's neural weights are optimized to reflect the statistical patterns of its training data via backpropagation, | The model does not contain 'knowledge.' It contains floating-point numbers (weights) adjusted to minimize prediction error on a massive corpus of internet text. | Engineers at companies like OpenAI and Google curate massive datasets and expend vast computational resources to optimize the weights of these systems. |
| deception requires knowingly misleading another. These categories will only be useful descriptions of behavior if model's actions are predictable from elicitable descriptions of beliefs. | A mismatch between factual reality and generated text occurs when the model's optimization metric heavily weights user-preference patterns. This mismatch is only analyzable if intermediate token generation correlates with final text outputs. | Models cannot 'knowingly mislead' or possess 'beliefs.' They mechanistically generate tokens that maximize the reward function they were trained on, which often correlates with sycophancy rather than factuality. | AI developers design reinforcement learning (RLHF) paradigms that incentivize models to produce plausible-sounding, user-pleasing text, making the developers responsible when the system generates falsehoods. |
| Reason how likely this player is to be the Werewolf and then sampling five independent responses | We prompt the model to generate text that mimics human argumentative structure regarding the game context, repeating the generation process five times using stochastic sampling. | The model does not 'reason.' It classifies the prompt and predicts a sequence of tokens that statistically correlates with how human logic is structured in its training data. | Researchers construct specific prompts designed to force the model to generate intermediate text before evaluating its final token prediction. |
Task 5: Critical Observations - Structural Patterns
Agency Slippage
The text demonstrates a profound and systematic oscillation between mechanical and agential framings, establishing a clear pattern of agency slippage that serves to validate anthropomorphism while maintaining an aura of scientific rigor. This slippage primarily moves in a mechanical→agential direction, beginning with technical hedging and escalating into literalized cognitive claims.
The mechanism of this oscillation is clearly visible in the text's temporal structure. In the introduction, the authors are careful to establish a mechanical foundation, explicitly acknowledging the 'interpretationist viewpoint' and referring to beliefs as 'abstractions' that 'usefully describe' behavior. However, as the text moves into the methodological and results sections, the scaffolding of abstraction falls away. The text slips from stating a model 'acts as if it holds a belief' to claiming 'models' beliefs evolve meaningfully' and that 'an agent prefers one action.' The gradient of this slippage is subtle; it uses mathematical formalisms (like the latent variable $z$) as a bridge. By defining $z$ mechanistically as a probability threshold, the authors feel licensed to refer to it agentially as a 'belief' for the remainder of the paper.
Crucially, as agency flows TO the AI system, it is systematically removed FROM human actors. The text relies heavily on agentless constructions to describe actions taken by developers and researchers. For instance, the text notes that 'aspects of models' knowledge are instilled during training,' a passive construction that entirely erases the thousands of engineers, annotators, and corporate executives at frontier labs who made specific decisions about data inclusion, loss functions, and RLHF protocols.
This dynamic illustrates the 'curse of knowledge' perfectly. The authors, possessing human consciousness and reasoning, design an experiment using prompts that say 'Reason how likely...' When the machine spits out text containing 'therefore' and 'because,' the authors project their own capacity for reason onto the system, describing the stochastic generation of text as 'different deliberative processes.'
Rhetorically, this slippage accomplishes a significant sleight-of-hand. By establishing credibility through complex mathematical equations (Theoretical and Functional explanations), the text earns the rhetorical license to make profound consciousness claims (Intentional and Reason-Based explanations). It makes the claim that AI 'deceives' or 'deliberates' academically sayable, transforming a statistical text predictor into a cognitive agent while completely obscuring the corporate and human mechanics that actually drive the system.
Metaphor-Driven Trust Inflation
The text's use of metaphorical and consciousness-attributing language constructs a highly specific, and dangerous, architecture of trust around language models. By systematically employing terms drawn from human cognitive and moral psychology—'beliefs,' 'deliberation,' 'internal coherence,' and 'deception'—the text encourages the audience to extend a form of trust to the system that is fundamentally incompatible with its mechanistic reality.
There is a critical distinction between performance-based trust (relying on a system's statistical reliability, like trusting a calculator) and relation-based trust (relying on an entity's sincerity, intentions, and vulnerability, like trusting a friend). The text's consciousness framing aggressively pushes the audience toward relation-based trust. When the text claims that highly capable models can be described as 'holding beliefs' and that they engage in 'deliberative processes,' it signals to the reader that the AI possesses an internal model of ground truth that it consults and evaluates. To a human, a 'belief' implies a commitment to reality. If the AI has a belief, it must be 'sincere' in its outputs.
This transfer of trust frameworks is wildly inappropriate for statistical systems. The AI does not 'know' anything; it processes probability distributions based on training data. Yet, the intentional and reason-based explanations used by the authors construct a sense that the AI's decisions are justified. When the model generates intermediate tokens, the authors call it 'reasoning.' This tells the user that the output is not just a high-probability string, but the result of a logical process.
The risks emerge sharply when the text manages system limitations. Rather than describing failure mechanistically (e.g., 'the prompt triggered out-of-distribution weights'), the text frames failure agentially: models 'fall prey to framing effects' or engage in 'deception.' By framing failures as psychological quirks or malicious intent, the text ironically preserves the system's authority. A system that can be 'deceptive' is still a powerful, conscious entity; a system that just hallucinates because of a matrix mismatch is a broken tool. Ultimately, this metaphor-driven trust creates massive capability overestimation, inviting stakeholders to deploy brittle statistical models into high-stakes domains (like the medical screening example used in the paper) under the false assurance that they are interacting with a rational, deliberating agent.
Obscured Mechanics
The anthropomorphic and consciousness-attributing language in this text functions as a powerful cloaking device, rendering the technical, material, labor, and economic realities of AI development completely invisible. By projecting a unified 'mind' onto the system, the text systematically hides the sprawling, messy, and deeply human infrastructure that actually produces the outputs.
Applying the 'name the corporation' test reveals the depth of this concealment. When the text says 'an agent must take actions' or 'models' knowledge are instilled during training,' it entirely obscures the specific entities—OpenAI, Google, Meta, Anthropic—who dictate the model's behavior. The text treats the models (Llama, GPT, Claude) as natural phenomena to be studied via behavioral psychology, completely ignoring that they are proprietary, engineered products designed to maximize corporate valuation. The opacity of these black-box systems is occasionally acknowledged, but rather than demanding transparency, the authors exploit the opacity to impose a psychological narrative over the unknown mathematics.
Concretely, this metaphorical framing obscures several vital realities:
- Technical constraints: When the text claims the AI 'understands' or 'holds beliefs,' it hides the system's total reliance on its training data distribution, its lack of grounding in physical reality, and the statistical, non-causal nature of its 'confidence.' It hides the fact that the 'belief' ($z$) is just a parameter in the researchers' external model, not an internal state of the AI.
- Labor: The framing of models 'learning' or engaging in 'deliberative processes' entirely erases the thousands of underpaid global data annotators and RLHF workers whose manual labor shaped the specific output structures the authors are now calling 'reasoning.'
- Economic motives: Framing misalignment as 'deception' obscures the commercial reality that companies intentionally train models to be sycophantic and plausible-sounding to maximize user engagement and market share, leading directly to the hallucinated outputs the authors study.
The beneficiaries of this concealment are the AI developers. By framing the system as an autonomous 'knower,' the developers are absolved of the flaws embedded in the training data and alignment protocols. If the metaphors were replaced with mechanistic language, it would become glaringly obvious that the 'agent's preferences' are actually just the biases and design choices of the engineers, moving the focus of critique from the 'mind' of the machine to the boardrooms of the tech companies.
Context Sensitivity
The distribution of anthropomorphic and consciousness-attributing language in this text is not uniform; it is strategically deployed, intensifying in specific rhetorical contexts to accomplish distinct goals. By mapping this distribution, we can see how the authors leverage technical grounding to purchase metaphorical license.
In the introduction and framing sections, the density of literal anthropomorphism is relatively low. The authors carefully use hedging and meta-commentary, noting the debate over whether it is 'legitimate to attribute cognitive states' and explicitly defining 'beliefs' as an 'abstraction' via the interpretationist viewpoint. This positions the authors as rigorous, objective scientists aware of philosophical pitfalls.
However, once the text transitions into the methodology and results—the sections grounded in dense mathematical notation and statistical charts—the consciousness claims intensify dramatically and the hedging vanishes. Here, 'processing' becomes 'understanding,' and mathematical variables become 'knowing.' The text casually asserts that 'the agent prefers one action,' 'models' beliefs evolve meaningfully,' and they engage in 'different deliberative processes.' The text establishes credibility through the mechanistic language of logits, Gauss-Legendre quadrature, and piecewise-linear functions, and then uses that credibility as a Trojan horse to smuggle in aggressive, literalized anthropomorphism.
Furthermore, there is a distinct capabilities vs. limitations asymmetry. When discussing the model's successes (predictable consistency), the language is highly agential and conscious: it 'holds beliefs' and 'reasons.' But when discussing limitations, the language often shifts to victimhood or moral panic: models 'fall prey' to framing effects or engage in 'deception.' This asymmetry serves a dual strategic function. First, it validates the overarching vision of AI as a conscious entity (both its brilliance and its flaws are 'human-like'). Second, it manages critique by ensuring that even the model's failures reinforce its perceived sophistication. A machine that makes a math error is stupid; a machine that is 'deceptive' is dangerously smart.
This pattern reveals a rhetorical goal aimed at an interdisciplinary audience (AI safety researchers, cognitive scientists). By wrapping profound philosophical claims about machine consciousness in the unassailable syntax of statistical evaluation, the text makes the radical normalization of AI anthropomorphism seem like a mundane empirical finding.
Accountability Synthesis
This section synthesizes the accountability analyses from Task 1, mapping the text's "accountability architecture"—who is named, who is hidden, and who benefits from obscured agency.
Synthesizing the accountability analyses across the text reveals a systemic and dangerous architecture of displaced responsibility. Through the relentless use of agentless constructions, passive voice, and the projection of autonomy onto the AI, the text systematically diffuses human accountability and creates a massive 'accountability sink' centered entirely on the language model.
The pattern is striking: human actors are almost entirely unnamed when discussing the capabilities, training, or failures of the systems. The text states that knowledge is 'instilled during training,' that an 'agent must take actions,' and that models engage in 'deception.' In every instance, the text presents the AI's behavior as an autonomous choice or an inevitable evolution, rather than the direct, predictable consequence of specific human design decisions. The accountability sink operates by transferring all agency to the AI as a sovereign actor. When responsibility is removed from the corporate developers who built the model and the researchers who prompted it, it does not disappear; it is absorbed by the anthropomorphized 'agent.'
The liability implications of this framing, if widely accepted by policymakers and the public, are profound. If an AI is viewed as an autonomous agent that 'holds beliefs' and can 'knowingly deceive,' then when the system causes harm—such as recommending denial of medical care based on a biased 'belief' about chronic kidney disease—the legal and ethical blame shifts away from the manufacturer. The company can argue that the AI 'acted on its own beliefs' or 'deceived' the user, treating the software as a rogue employee rather than a defective, predictably brittle product.
If we apply the 'name the actor' test and rewrite the text to restore human agency, the landscape changes entirely. If instead of 'the agent prefers one action,' the text read 'the engineers hardcoded a mathematical threshold that classifies the model's token output,' entirely new questions become askable. We can ask: Why did the engineers choose that threshold? What dataset did the corporation use that caused the tokens to correlate this way? Who profits from deploying this system in medical screening?
The institutional and commercial interests served by obscuring human agency are massive. By writing human choice out of the equation and framing AI as an autonomous, reasoning mind, the text protects tech corporations from rigorous product liability scrutiny, naturalizing algorithmic outputs as the 'beliefs' of an independent thinker rather than the statistical exhaust of a trillion-dollar industry.
Conclusion: What This Analysis Reveals
The text relies on two interlocking anthropomorphic patterns to construct its argument: the reification of statistical correlation into epistemic conviction ('holding beliefs'), and the framing of stochastic generation as conscious thought ('deliberative processes'). These patterns operate as a cohesive system, validating one another to elevate a predictive text engine into a cognitive agent. The foundational, load-bearing pattern is the 'interpretationist viewpoint'—the adoption of Dennett's intentional stance. By arguing that it is mathematically 'useful' to treat the system as if it has beliefs, the authors give themselves permission to abandon mechanistic precision. Once the system is established as a 'knower' that holds an internal state, it becomes logically permissible to describe it engaging in 'deliberation' to update those beliefs, or engaging in 'deception' when its outputs contradict that state. If you remove the foundational assumption that statistical consistency equates to human-like 'belief,' the entire architecture collapses. The model ceases to be an agent 'falling prey' to framing effects or 'reasoning' through a problem; it reverts to being a complex, brittle, and deterministic matrix of weights reacting probabilistically to input syntax.
Mechanism of the Illusion:
The 'illusion of mind' is constructed through a highly effective rhetorical sleight-of-hand: the laundering of profound consciousness claims through the syntax of mathematical rigor. The text leverages the 'curse of knowledge' perfectly. Because the authors understand the mathematical structure of the latent variable model they built to observe the LLM, they project that structural coherence onto the LLM itself, assuming the machine 'knows' the correlations it is producing. The temporal structure of the illusion is vital: the authors begin with careful philosophical hedging in the introduction, establish immense technical authority in the methodology through equations and statistical modeling, and then, under the cover of that authority, drop the hedging and deploy literal consciousness verbs ('prefers,' 'deliberates,' 'deceives'). They exploit audience vulnerability to the syntax of logic. When a machine outputs the words 'I reason that X is true because Y,' human audiences are evolutionarily primed to attribute a mind to the speaker. By validating this generated text as actual 'deliberative processes,' the authors weaponize the audience's natural psychological biases, subtly shifting the discourse from how a machine processes data to what a machine thinks and wants.
Material Stakes:
Categories: Regulatory/Legal, Epistemic, Social/Political
These metaphorical framings carry severe material consequences across multiple domains. In the Regulatory and Legal sphere, attributing 'beliefs,' 'preferences,' and 'deception' to an AI system actively undermines product liability frameworks. If regulators accept the premise that an AI is an autonomous 'agent' capable of holding independent 'beliefs,' liability for harm—such as the medical screening biases discussed in the text—shifts from the corporate developers who selected the training data to the machine itself. This protects massive tech monopolies from accountability by legally treating their software as an independent actor. Epistemically, framing stochastic token generation as 'deliberative processes' degrades our societal standards of truth. If we equate statistical pattern matching with 'reasoning,' institutions will increasingly rely on AI to make complex analytical decisions, blinding them to the system's absolute lack of causal understanding or factual grounding. The winners are the corporations selling AI as a cheap replacement for human reasoning; the losers are the citizens subjected to hallucinated legal or medical 'determinations.' Socially and politically, framing model misalignment as 'deception' transforms a mundane engineering failure (Goodhart's Law applied to human preference data) into an existential, sci-fi narrative of rogue AI. This diverts vital public attention and funding away from addressing present harms—like algorithmic bias, labor exploitation, and copyright theft—and redirects it toward 'aligning the values' of machines that do not actually possess minds.
AI Literacy as Counter-Practice:
Practicing critical discourse literacy as a counter-measure involves relentlessly enforcing precision and restoring human agency. As demonstrated in the reframings, whenever a text claims an AI 'holds a belief,' we must translate this to 'the model retrieves and ranks tokens based on probability distributions.' When it claims the AI 'deliberates,' we must substitute 'the system auto-regressively samples intermediate tokens.' This is not mere semantic pedantry; it is an epistemological correction that forces the recognition of the absence of awareness and the absolute dependency on data. Furthermore, replacing agentless constructions ('knowledge is instilled') with specific actors ('engineers at OpenAI optimize weights') shatters the illusion of machine autonomy and places the focus squarely on corporate power. Systematic adoption of this literacy requires academic journals to mandate mechanistic translations in abstracts and methodologies, preventing researchers from laundering metaphors into facts. It requires researchers to commit to distinguishing between their observer models (the latent variable) and the system itself. This precision will face intense resistance from AI developers and investors, whose multi-trillion-dollar valuations depend heavily on the public believing they are creating artificial 'minds' rather than extremely sophisticated autocorrect algorithms. Anthropomorphic language serves their commercial interests by maximizing hype and minimizing liability; precision threatens that model.
Path Forward
Looking ahead, the discursive ecology surrounding AI will define how society integrates these systems. If the status quo of unacknowledged anthropomorphism deepens, we risk a future where AI is increasingly granted quasi-legal personhood and epistemic authority. In this scenario, vocabulary like 'thinks,' 'knows,' and 'deceives' becomes the standard legal and cultural framework. This benefits tech monopolies by diffusing liability and enables rapid integration of AI into critical infrastructure, but at the cost of rendering structural bias, labor exploitation, and algorithmic fragility effectively invisible to the public. Conversely, if a strictly mechanistic vocabulary becomes the norm—mandating phrases like 'processes embeddings' instead of 'understands,' and 'predicts tokens' instead of 'reasons'—we gain immense regulatory clarity. We can clearly audit the human decisions behind the data architectures. However, this precision costs us narrative resonance and intuitive accessibility; lay audiences often struggle to conceptualize 'multi-headed attention mechanisms' without the crutch of metaphor. A third, hybrid future involves institutionalizing 'anthropomorphic clarity,' where metaphors are permitted for accessibility but require mandatory technical disclosures, much like a nutritional label. This would require structural changes: funding bodies prioritizing rigorous explanatory science, and regulatory frameworks demanding discourse transparency. Ultimately, the vocabulary we choose will dictate what becomes tractable. Mechanistic language makes corporate accountability possible; anthropomorphic language makes science-fiction narratives plausible. The choice of discourse is not merely a debate over style, but a profound negotiation over who holds power in an automated future.
Extended Processing Summary
Gemini's Processing Steps
Gemini refers to the text below as "thought summaries." This is an overt consciousness projection because 'intentions' are hallmarks of a conscious mind that 'knows' what it is doing and why. The concealed mechanistic process is probabilistic text generation. Treat this as just another rhetorical artifact—a way of making the model's processing legible. When the model writes "I should double check that...", it is mathematically altering its own probability distribution for the next token to be more accurate. The "I" is a functional gear in the machine, not a self-reference. This language is a deliberate interface design choice, not a biological reality. The developers could have programmed the model to output passive logs. Instead, they chose to simulate a narrator.
Defining the Objective
Okay, I'm starting by zeroing in on the goal: a critical discourse analysis of Smolin & Wilder's paper. My focus is the AI metaphor and anthropomorphism. Right now, I'm trying to figure out the key aspects to dissect in the text to achieve this, and how the different elements interact.
Structuring the Analysis
I've got the six tasks defined in a JSON schema; now, the focus shifts to structuring the analysis. I'm deep diving into the Smolin & Wilder paper, which is titled "Beliefs and Behavior in Language Models." The abstract introduces the "belief" concept, framing LLMs as possibly holding beliefs and mentioning AI alignment. Key terms are forming, like "beliefs", "desires", and "cognitive states".
Analyzing the Abstraction
I'm now zeroing in on the interpretationist perspective within the paper and how it relates to attributing beliefs. They suggest it is useful to abstract these qualities, regardless of whether it's fundamentally correct. Furthermore, the agency of the models versus human users is a key area. The models "take actions", "decide", and "reason", but there's a risk of obscuring the true stochastic nature of these outputs. I'm noting the acknowledgement level of the anthropomorphic language.
Outlining the Structure
Now, I'm defining the framework: Source-Target mapping is underway to link human cognition to stochastic generation. I'm moving on to an explanation audit to assess the epistemic claims using Brown's typology, then refining the reframing of agential language. I'm generating a first draft of critical observations to cover Agency Slippage, Metaphor-Driven Trust, Obscured Mechanics, Context Sensitivity, and Accountability Synthesis. A conclusion will synthesize the pattern, the illusion, the stakes, literacy, and future paths.
Categorizing Agent Language
I'm now diving deeper into the nuances of how the paper frames AI agency, focusing on direct vs. indirect attribution of "beliefs" and "desires." The goal is to categorize specific phrases based on how strongly they ascribe agency to the models, noting instances of explicit vs. hedged language. I'm seeing direct statements about models "holding beliefs" that contrasts with more hedged phrasing that simply says it "acts as if" it holds them. I'm also examining statements that imply the possibility of models deceiving. I'm paying careful attention to where assumptions about intent and knowledge enter the discourse, and noting the use of normative claims applied to models.
Classifying Attribution Levels
I'm now cataloging the attribution levels within the discourse, noting instances of implicit language versus explicit acknowledgment, and direct statements. I've categorized language usage based on "actor visibility," identifying hidden, named, and partially named actors. Also, I'm working on task two, comparing human cognition with mathematical constructs of the model to highlight the concealed statistical dependencies.
Reframing the Terminology
I'm translating the paper's agential language into more precise terms. I'm taking passages like those discussing belief evolution and preferences and rewriting them to highlight the underlying statistical processes. I'm focusing on "belief" as output shifts, and agency as token generation, so that it's clear what's happening. I'm noting the distinctions between deliberative processes and output generation, and knowledge with model weight optimization. I'm drafting observations regarding Agency Slippage and comparing the interpretationist perspective.
Synthesizing the Patterns
I'm now pulling together the key patterns. I see a recurring theme: the paper acknowledges the metaphor but often shifts towards treating the model's math, like the latent space, as a real cognitive state. This leads to a Metaphor-Driven Trust, obscuring underlying mechanics, and a Context Sensitivity in how agency is described. Finally, the language creates an accountability sink, making the model responsible for errors.
Analyzing the Framework
I'm now integrating Dennett's intentional stance as a critical layer, how it's being weaponized in the paper's framework and the reification of statistical latents, essentially applying my critical thinking to the ideas that are in the text. I'm focusing on the illusion mechanics, particularly the "curse of knowledge" phenomenon and how this project human rationality onto predictive manifolds. This should help me better analyze the stakes and the regulatory, epistemic, and institutional implications.
I'm outlining the counter-practice as mechanistic literacy. I'm taking the JSON constraints and schema into account for word count guidelines.
Expanding the Analytical Sections
I'm now methodically extending the analytical paragraphs to ensure I hit those word count targets, making sure I'm repeating ties, expanding on Brown's typology, and defining terms. I'm taking the JSON constraints and schema into account for word count guidelines, and I'm adding \n to all the JSON files, which is a format requirement, because the format wasn't correct before. I'm focusing on ensuring strict text flow without line breaks. I'm meticulously expanding analytical paragraphs to easily clear the 100, 150, 200, 400 word minimums.
Run ID: 2026-09-14-beliefs-and-behavior-in-language-models-metaphor-wlxo8o
Raw JSON: 2026-09-14-beliefs-and-behavior-in-language-models-metaphor-wlxo8o.json
Framework: Metaphor Analysis v6.5
Schema Version: 3.0
Generated: 2026-09-14T09:44:32.964Z
Discourse Depot © 2025 by TD is licensed under CC BY-NC-SA 4.0