Skip to main content
About

This document presents a Critical Discourse Analysis focused on AI literacy, specifically targeting the role of metaphor and anthropomorphism in shaping public and professional understanding of generative AI. The analysis is guided by a prompt that draws from cognitive linguistics (metaphor structure-mapping), the philosophy of social science (Robert Brown's typology of explanation), and accountability analysis.

All findings and summaries below were generated from detailed system instructions provided to a large language model and should be read critically as interpretive outputs—not guarantees of factual accuracy or authorial intent.


Task 1: Metaphor and Anthropomorphism Audit

About this task

For each of the major metaphorical patterns identified, this audit examines the specific language used, the frame through which the AI is being conceptualized, what human qualities are being projected onto the system, whether the metaphor is explicitly acknowledged or presented as direct description, and—most critically—what implications this framing has for trust, understanding, and policy perception.

V3 Enhancement: Each metaphor now includes an accountability analysis.

1. AI as Conscious Game Player

Quote: "The game of Taboo requires describing a target word without using a set of forbidden words... making it a compelling playground for examining how LLMs navigate competing demands at inference time."

  • Frame: Model as rule-comprehending game participant
  • Projection: This metaphorical framing projects human rule-comprehension, conscious engagement, and strategic intentionality onto a statistical model generating tokens. By stating that language models play Taboo and navigate competing demands, the text maps the subjective experience of participating in a parlour game—which requires understanding rules, desiring to win, and consciously navigating social or lexical boundaries—onto mechanistic token prediction. It attributes conscious awareness and justified belief to computational pattern-matching, suggesting the system knows what a game is and understands the restrictions. This fundamentally obscures the mathematical reality of the interventions, which consist merely of logit manipulation, prompt conditioning, and statistical correlations, rather than any conscious participation or cognitive playfulness.
  • Acknowledgment: Direct (Unacknowledged) (The text presents the claim that models navigate demands as a literal framing in the abstract, without any linguistic hedging. I considered 'Explicitly Acknowledged' since Taboo is a known human game used as a benchmark, but the phrasing directly literalizes the model's participation as an active, navigating player rather than just stating it is evaluated on a Taboo-like generation task.)
  • Implications: This framing significantly affects human understanding by encouraging unwarranted trust in the model's capacity for generalized rule-following and social compliance. If users believe a model plays and complies with game rules through conscious understanding, they may logically assume it can comprehend and obey complex legal, ethical, or safety directives. This inflates the perceived sophistication of the AI, disguising rigid computational masking as flexible, intelligent adaptation. Consequently, it creates liability ambiguities: if a system is viewed as an autonomous player, responsibility for its failures might be erroneously attributed to the machine rather than its designers.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: The construction completely obscures the human researchers, engineers, and corporate entities who constructed the benchmark, applied the logit-masking constraints, and engineered the prompts. By naming the model as the active subject navigating demands, it hides the fact that OpenAI and Google created statistical systems optimized for specific generation patterns, and the authors themselves surgically intervened in the generative process. I considered 'Partial' because the authors mention 'we evaluate', but the specific action of navigating displaces agency entirely onto the AI. Naming the actors would reveal this as human-constrained statistical sampling.
Show more...

2. AI as Restrained Speaker

Quote: "the speaker must suppress some lexically and conceptually salient words while simultaneously producing a description that is informative enough"

  • Frame: Model as self-monitoring human communicator
  • Projection: This metaphorical framing maps the conscious human psychological act of suppression onto a computational text generation system. By designating the model as a speaker who must suppress words, the text projects human subjective awareness, self-monitoring, and intentional restraint onto the target domain. It implies the model knows the forbidden words, actively contemplates using them, and then consciously decides to hold them back to adhere to a social or game constraint. This fundamentally replaces the mechanistic reality of logit masking and probability distribution modification with a psychological narrative. The AI does not suppress anything; rather, its generation probabilities are mathematically restricted by human-coded algorithms. The projection of conscious restraint inflates the system's perceived autonomy and cognitive depth.
  • Acknowledgment: Direct (Unacknowledged) (The term 'speaker' and the action 'must suppress' are used as definitive, unhedged descriptions of the model's operation. I considered 'Hedged/Qualified' because the surrounding text discusses 'generative process', but the specific sentence attributes the psychological act of suppression directly to the system as a literal truth.)
  • Implications: Attributing the ability to suppress information to an AI system fundamentally misrepresents how safety filters and constraints operate. It implies that the model has an internal, conscious reservoir of knowledge that it intelligently decides not to share. This creates dangerous policy implications, as regulators and users may believe the model can be trusted to autonomously self-censor harmful information based on human-like judgment. It masks the reality that if the masking algorithm fails, the model will output the data indiscriminately, leading to unwarranted trust in the system's self-governance.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: The sentence frames the model as a 'speaker' who 'must suppress' words, completely erasing the human developers who actually enforce this suppression through generation-time constraints or prompting. The human researchers programmed the inference-time checks and the a priori stem censorship. I considered 'Named' because the broader paper identifies the researchers, but in this specific quote, the agency is displaced onto the AI. The decision of what to suppress and how to suppress it was entirely human, yet the text attributes the active suppression to the machine.

3. AI as Forward Planner

Quote: "Generating an effective Taboo description requires implicitly planning ahead: the model must produce a description that is semantically informative while simultaneously avoiding a set of lexically related forbidden words."

  • Frame: Model as temporal, strategic thinker
  • Projection: This framing maps the uniquely human capacity for forward temporal projection and strategic planning onto auto-regressive token generation. By claiming the model is planning ahead, the text attributes a conscious, future-oriented intentionality to the system. It suggests the model conceptualizes an end goal (winning the game, being informative) and deliberately structures its current actions (word choices) to achieve that future state. In reality, the model calculates next-token probabilities based on the current context window; it does not possess a temporal horizon or a conscious desire to achieve a future goal. This projection bridges the gap between mechanistic processing and conscious knowing by suggesting the AI knows what it wants to accomplish.
  • Acknowledgment: Hedged/Qualified (The authors use the word 'implicitly' before 'planning ahead', which functions as a hedge indicating this is not literal conscious planning. I considered 'Explicitly Acknowledged' but there are no scare quotes or meta-commentaries unpacking the metaphor; it relies merely on the adverb 'implicitly' to soften the otherwise direct intentional claim.)
  • Implications: Framing a language model as capable of planning ahead drastically inflates public and regulatory perception of its cognitive abilities. It suggests the system possesses strategic foresight and goal-directed behavior, which are hallmarks of artificial general intelligence rather than current LLM architectures. This can lead to a dangerous overestimation of capabilities, causing organizations to deploy these models in autonomous roles requiring genuine foresight. Furthermore, it creates fear by implying the model could secretly plan malicious actions, distracting from the actual mechanical risks of statistical correlation failures and data bias.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: The phrase 'the model must produce a description' places the burden of action entirely on the AI, obscuring the human-designed reward functions (RLHF) and prompting strategies that dictate its output. The researchers enforce the constraints that make the output informative and compliant. I considered 'Ambiguous' but the subject-verb pairing clearly isolates the model as the sole actor. Replacing this with mechanistic language would reveal that humans force the model to output sequences that correlate with human informative descriptions.

4. AI as Logical Reasoner

Quote: "confirming that the reasoning process plays an active role in enforcing lexical constraints."

  • Frame: Model as deliberate logical deducer
  • Projection: This metaphor projects human deductive logic, rational contemplation, and active cognitive enforcement onto a specific prompt engineering technique (Chain-of-Thought). By referring to the generation of intermediate tokens as a reasoning process that plays an active role, the text attributes conscious deliberation and epistemic justification to the model. It suggests the AI understands the logical relationship between the constraint and the output, and actively chooses to enforce the rule based on rational thought. This obscures the mechanistic reality that Chain-of-Thought simply extends the context window with probabilistically correlated intermediate tokens, altering the final distribution without any conscious understanding, belief, or logical deduction taking place.
  • Acknowledgment: Direct (Unacknowledged) (The phrase is presented as an empirical finding and literal fact ('confirming that...'). I considered 'Hedged/Qualified' because 'reasoning' is a common technical term in the field, but within the discourse of the text, it is used without any qualification to describe a literal, active cognitive process happening inside the machine.)
  • Implications: Using 'reasoning' to describe statistical token generation convinces non-experts that the system makes decisions based on rational justification and logical deduction. This fundamentally shifts the epistemic status of the model's output from 'statistically likely' to 'rationally justified'. When audiences believe an AI reasons, they are more likely to defer to its judgments in critical areas like medical diagnosis or legal analysis, assuming it has logically evaluated the evidence. This masks the reality that the model is incapable of true logical deduction or recognizing ground truth, leading to systemic vulnerabilities when hallucinations are treated as reasoned conclusions.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: The text attributes the 'active role' to the 'reasoning process', treating the algorithmic sequence as an autonomous agent enforcing rules. This completely hides the engineers at OpenAI who trained the model using reinforcement learning to produce these specific token patterns, and the researchers who selected the prompt structure. I considered 'Partial' since the process is a named technical artifact, but it entirely displaces the human agency responsible for the constraint enforcement. The human developers enforce the constraints; the reasoning process is merely their tool.

5. AI as Incompetent Competitor

Quote: "Models are weak guessers... models do not share the same network of salient lexical associations that makes Taboo intuitive for humans"

  • Frame: Model as failing human participant
  • Projection: This metaphor maps human intuitive failure and competitive weakness onto the model's inability to retrieve specific tokens based on indirect prompts. By labeling the models as 'weak guessers' who find the task 'unintuitive', the text projects the human subjective experience of trying to guess a word, feeling confused, and lacking intuition onto the AI. It implies the model is trying its best to play the game but is simply not smart enough or intuitive enough to succeed. This frames the computational limitation (lack of sufficient training data correlations for indirect descriptions) as a cognitive or psychological deficit, reinforcing the illusion that the model has a mind that is currently failing.
  • Acknowledgment: Direct (Unacknowledged) (The statement 'Models are weak guessers' is presented as a literal, unhedged empirical conclusion of the paper. I considered 'Ambiguous' because 'guessing' could technically mean statistical prediction, but the juxtaposition with human intuition ('makes Taboo intuitive for humans') confirms the psychological framing.)
  • Implications: Framing limitations in psychological terms (weak guesser, lacking intuition) preserves the overall illusion of mind even while pointing out a flaw. By criticizing the AI as if it were a poorly performing human, the text reinforces the premise that the AI is the kind of entity that could possess intuition in the first place. This shapes public policy by focusing attention on 'improving the AI's intuition' rather than addressing the fundamental unreliability of statistical architectures. It maintains the agential framing, ensuring that when the model fails, it is seen as a personal shortcoming of the AI rather than a design flaw of the developers.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: By evaluating the 'models' as 'weak guessers', the text positions the AI as the responsible agent for the failure. It obscures the fact that Google and OpenAI designed these systems, curated their training data, and defined their objective functions. The failure is a product of human engineering choices regarding network architecture and training paradigms. I considered 'Named' because the companies are named elsewhere, but in this specific evaluation of failure, the human engineers who built the 'weak' system are completely absolved of their design limitations.

6. AI as Anxious Communicator

Quote: "whereas an instructed model tends toward overly cautious, informationally sparse descriptions."

  • Frame: Model as emotionally hesitant agent
  • Projection: This metaphorical mapping projects human emotional states, risk aversion, and psychological caution onto a language model's output distribution. By describing the model as 'overly cautious', the text attributes a conscious fear of making a mistake and a deliberate choice to withhold information. It suggests the AI understands the rules, fears violating them, and acts defensively. Mechanistically, the prompt conditioning has simply shifted the probability space toward generic, high-frequency tokens and away from specific, low-frequency tokens related to the taboo word. Projecting caution onto the model replaces this mathematical smoothing of the probability distribution with an intentional, emotional narrative of self-preservation.
  • Acknowledgment: Direct (Unacknowledged) (The term 'overly cautious' is used directly to describe the model's behavioral tendency without any hedging. I considered 'Hedged/Qualified' because 'tends toward' softens the absolute nature of the claim, but the psychological attribute ('cautious') itself is presented as a literal characteristic of the instructed model.)
  • Implications: Attributing emotional states like caution to an AI system builds relation-based trust, encouraging users to view the system as a careful, sincere collaborator that wants to do the right thing. If audiences believe the AI is 'cautious', they will assume it possesses an ethical compass and a self-regulatory mechanism driven by a desire to avoid harm. This is incredibly dangerous in policy contexts, as it suggests models can be trusted to manage their own risks through innate caution. It obscures the fact that the system has no concept of risk, harm, or caution, and will generate dangerous content if the mathematical weights align differently.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: The phrase 'an instructed model tends toward' places the agency for the sparse descriptions on the model's supposed psychological tendencies. This hides the human actors: the authors who wrote the strict prompt, and the RLHF annotators who trained the model to output generic responses when constrained. I considered 'Partial' because the word 'instructed' implies a human instructor, but the reaction ('cautious') is entirely displaced onto the AI. The human choice to penalize certain outputs is erased.

7. AI as Amnesiac Mind

Quote: "the model loses access to forbidden words but retains the underlying conceptual associations needed to produce an informative description."

  • Frame: Model as remembering/forgetting cognitive entity
  • Projection: This framing maps human memory, conceptual understanding, and cognitive access onto the internal vector representations of a neural network. By stating the model 'loses access' but 'retains conceptual associations', the text projects a human-like subconscious where concepts are truly understood but the specific vocabulary is temporarily blocked. It implies the AI truly 'knows' the concept in a human sense and is holding it in mind. Mechanistically, the Sparse Autoencoder intervention simply shifts the residual stream vector away from the direction correlated with the target word; there is no cognitive 'retention' or 'understanding', just multi-dimensional geometric spaces and token probabilities.
  • Acknowledgment: Direct (Unacknowledged) (The language of losing access and retaining associations is presented as a factual description of the internal workings of the model. I considered 'Explicitly Acknowledged' due to the technical context of SAEs discussed nearby, but the actual phrasing literalizes the psychological concepts of memory and association without any meta-commentary.)
  • Implications: This projection solidifies the illusion that LLMs possess deep, human-like conceptual understanding independent of their surface text generation. By claiming the AI 'retains underlying conceptual associations', it convinces stakeholders that models have a robust, grounded model of reality, rather than just statistical maps of text. This inflates perceived sophistication and encourages deploying these models in complex reasoning tasks where deep conceptual grounding is required. It masks the fragility of these systems, hiding the fact that their 'knowledge' is merely correlative geometry, which can fail catastrophically in novel edge cases.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: The model is presented as the active subject that 'loses access' and 'retains', obscuring the human researchers who actively manipulated the latent space using Sparse Autoencoders to block specific vector directions. The researchers literally subtracted a vector from the residual stream, but the text frames this as an internal cognitive event happening to the model. I considered 'Partial' since the methodology section details the human intervention, but this specific theoretical conclusion displaces the human agency entirely, making the model's 'mind' the sole actor.

8. AI as Adaptive Strategist

Quote: "examining whether the strategies models adopt under constraint align with those of human players."

  • Frame: Model as intentional problem-solver
  • Projection: This metaphorical framing projects human intentionality, strategic deliberation, and adaptive problem-solving onto algorithmic output generation. By stating that models 'adopt' strategies, the text implies a conscious evaluation of alternatives and a deliberate choice of the best path forward to achieve a goal. It suggests the AI 'knows' it is constrained, 'understands' the objective, and intelligently formulates a plan. Mechanistically, the model merely generates the sequence of tokens that maximizes the probability function given the prompt and the constraints; it does not conceptualize strategies, evaluate options, or consciously adopt anything. The language of strategy fundamentally confuses statistical optimization with conscious planning.
  • Acknowledgment: Direct (Unacknowledged) (The claim that models 'adopt strategies' is presented straightforwardly as an object of empirical examination. I considered 'Hedged/Qualified' because it is framed as a question ('whether...'), but the underlying assumption that models possess and adopt strategies is treated as an established, literal fact.)
  • Implications: Framing LLMs as strategic agents significantly affects public and institutional trust. If decision-makers believe models can intelligently adopt strategies to overcome constraints, they will assume the systems possess a high degree of autonomy and adversarial capability. This can lead to exaggerated fears of AI deliberately bypassing safety protocols through cunning strategy, while simultaneously obscuring the actual risks of algorithmic failure. It shifts the regulatory focus from auditing training data and objective functions to trying to control a supposedly autonomous, strategic mind, misallocating safety resources.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: The phrase 'strategies models adopt' attributes full agency and decision-making power to the AI, completely erasing the engineers who designed the optimization algorithms and the researchers who applied the constraints. The 'strategy' is actually a mathematical artifact of the human-designed system. I considered 'Partial' because the text mentions 'under constraint', implying a constrainer, but the active adoption of the strategy is entirely displaced onto the machine. If actors were named, it would be clear that humans engineered systems that optimize for these specific linguistic patterns.

Task 2: Source-Target Mapping

About this task

For each key metaphor identified in Task 1, this section provides a detailed structure-mapping analysis. The goal is to examine how the relational structure of a familiar "source domain" (the concrete concept we understand) is projected onto a less familiar "target domain" (the AI system). By restating each quote and analyzing the mapping carefully, we can see precisely what assumptions the metaphor invites and what it conceals.

Mapping 1: A human playing a parlour game, consciously understanding rules, aiming to win, and actively navigating social constraints. → An LLM generating text based on next-token prediction while subject to human-imposed logit masking and prompt conditioning.

Quote: "The game of Taboo requires describing a target word without using a set of forbidden words... making it a compelling playground for examining how LLMs navigate competing demands at inference time."

  • Source Domain: A human playing a parlour game, consciously understanding rules, aiming to win, and actively navigating social constraints.
  • Target Domain: An LLM generating text based on next-token prediction while subject to human-imposed logit masking and prompt conditioning.
  • Mapping: The relational structure of human game-playing (player -> understands rules -> navigates constraints -> tries to win) is mapped onto the LLM's computational process. This mapping invites the assumption that the AI possesses conscious awareness of the game, understands the concept of 'competing demands', and possesses a subjective drive to succeed. It maps the conscious cognitive effort of human navigation onto the unthinking, automatic calculation of probability distributions.
  • What Is Concealed: This mapping completely hides the mathematical and mechanistic reality of the process. It obscures the stochastic nature of token prediction, the dependency on training data frequencies, and the specific human interventions (logit forcing) used to control the output. It conceals the fact that the system has no concept of a 'game' or 'rules' and is merely executing a deterministic or pseudo-random statistical function. The transparency obstacle here is the rhetorical exploitation of a game metaphor to make a black-box statistical process appear comprehensible and agentic.
Show more...

Mapping 2: A conscious human speaker with internal thoughts, self-control, and the psychological ability to suppress unwanted words. → A computational text generation process where certain token probabilities are mathematically reduced or blocked.

Quote: "the speaker must suppress some lexically and conceptually salient words while simultaneously producing a description that is informative enough"

  • Source Domain: A conscious human speaker with internal thoughts, self-control, and the psychological ability to suppress unwanted words.
  • Target Domain: A computational text generation process where certain token probabilities are mathematically reduced or blocked.
  • Mapping: The structure of psychological self-censorship (mind -> generates thought -> evaluates thought against rules -> suppresses thought) is mapped onto the algorithmic generation of text. This assumes the AI has an internal mental state, a reservoir of intended words, and an active, conscious mechanism for evaluating and suppressing them before they are spoken. It projects human self-awareness and intentional restraint onto mathematical probability shifts.
  • What Is Concealed: This mapping conceals the absence of an internal mental state or conscious intent. It hides the mechanical reality that 'suppression' is actually just a human-coded adjustment to the logit scores (e.g., setting probabilities to negative infinity) or a result of attention mechanisms trained via RLHF. It obscures the fact that the model doesn't 'want' to say the taboo word and then stop itself; the mathematical weights simply dictate a different output. The text uses this metaphor to make the human-imposed safety constraints seem like the model's own psychological self-control.

Mapping 3: A human strategist conceptualizing a future goal, outlining steps, and executing a plan over time. → The auto-regressive generation of tokens by a language model, potentially utilizing Chain-of-Thought prompting.

Quote: "Generating an effective Taboo description requires implicitly planning ahead: the model must produce a description that is semantically informative while simultaneously avoiding a set of lexically related forbidden words."

  • Source Domain: A human strategist conceptualizing a future goal, outlining steps, and executing a plan over time.
  • Target Domain: The auto-regressive generation of tokens by a language model, potentially utilizing Chain-of-Thought prompting.
  • Mapping: The relational structure of temporal planning (agent -> visualizes future -> calculates steps -> executes plan) is projected onto the sequential calculation of next-token probabilities. This mapping invites the assumption that the model possesses a temporal horizon, a conscious awareness of a future goal (being informative without failing), and the ability to deliberate. It maps the human experience of holding a goal in mind onto the stateless, momentary calculation of vector attention.
  • What Is Concealed: This mapping conceals the stateless, Markovian nature of standard auto-regressive generation, where each token is generated solely based on the preceding context window without any actual visualization of the future. Even with Chain-of-Thought, it obscures the reality that the 'planning' is just the generation of intermediate tokens probabilistically correlated with successful outcomes in the training data, not a conscious evaluation of future states. It hides the absolute reliance on human-provided prompt structures to trigger these specific statistical patterns.

Mapping 4: A human mind utilizing deductive logic, rational evaluation, and active cognitive effort to solve a problem. → The mathematical generation of intermediate text tokens (Chain-of-Thought) before producing a final output.

Quote: "confirming that the reasoning process plays an active role in enforcing lexical constraints."

  • Source Domain: A human mind utilizing deductive logic, rational evaluation, and active cognitive effort to solve a problem.
  • Target Domain: The mathematical generation of intermediate text tokens (Chain-of-Thought) before producing a final output.
  • Mapping: The structure of human rationality (mind -> analyzes premise -> applies logic -> reaches justified conclusion) is mapped onto the statistical generation of text. This invites the assumption that the model understands the logical relationship between the rules and its output, possesses conscious reasoning capabilities, and actively enforces rules based on rational deduction. It projects epistemic justification onto probabilistic correlation.
  • What Is Concealed: This mapping entirely conceals the lack of ground truth, causal models, or logical deduction within the LLM architecture. It hides the reality that Chain-of-Thought is merely a process of extending the context window so that attention heads can weight intermediate token probabilities, which humans interpret as reasoning. It obscures the heavy reliance on reinforcement learning from human feedback (RLHF), which trained the model to mimic the structural appearance of human reasoning without any underlying cognitive process or logical comprehension.

Mapping 5: A human player participating in a game, feeling confused, lacking intuition, and failing to perform well. → An LLM failing to predict a target word given an indirect, constrained prompt because of training data distribution limitations.

Quote: "Models are weak guessers... models do not share the same network of salient lexical associations that makes Taboo intuitive for humans"

  • Source Domain: A human player participating in a game, feeling confused, lacking intuition, and failing to perform well.
  • Target Domain: An LLM failing to predict a target word given an indirect, constrained prompt because of training data distribution limitations.
  • Mapping: The structure of human psychological failure (person -> lacks intuition -> tries hard but fails) is mapped onto a machine's statistical mismatch. This invites the assumption that the AI is a conscious participant experiencing difficulty and lacking a human-like 'intuition'. It maps the human subjective feeling of cognitive struggle onto a mathematical failure to assign high probability to the correct token.
  • What Is Concealed: This mapping conceals the purely statistical reasons for the failure. It hides the fact that the model relies entirely on the frequency of word co-occurrences in its training corpus. If the Taboo description uses uncommon phrasing, the mathematical correlations are too weak to retrieve the target token. By framing this as a lack of 'intuition', the text obscures the mechanical reality of dataset dependency, optimization functions, and the fundamental differences between human semantic memory and high-dimensional vector embeddings.

Mapping 6: A person who is anxious, risk-averse, and emotionally hesitant to speak freely. → An LLM generating high-entropy, generic token sequences due to restrictive prompt conditioning.

Quote: "whereas an instructed model tends toward overly cautious, informationally sparse descriptions."

  • Source Domain: A person who is anxious, risk-averse, and emotionally hesitant to speak freely.
  • Target Domain: An LLM generating high-entropy, generic token sequences due to restrictive prompt conditioning.
  • Mapping: The relational structure of human emotional restraint (person -> feels fear/caution -> chooses to withhold information) is mapped onto a shift in output probability distributions. This invites the assumption that the AI possesses emotional states, understands risk, and consciously chooses a defensive strategy to protect itself or comply with authority. It projects human psychological self-preservation onto a mathematical function.
  • What Is Concealed: This mapping conceals the mechanics of prompt conditioning and the impact of safety fine-tuning (RLHF). It hides the fact that the human-provided prompt simply shifts the attention weights away from specific, highly correlated tokens toward more generic, high-frequency tokens. There is no emotional caution or conscious choice; the math simply favors 'sparse' outputs under those specific input conditions. This obscures the labor of the engineers and annotators who tuned the model to exhibit this specific statistical behavior when constrained.

Mapping 7: A human mind experiencing temporary amnesia or a 'tip-of-the-tongue' phenomenon, where the concept is understood but the word is forgotten. → The manipulation of a neural network's residual stream using Sparse Autoencoders to subtract a specific feature vector.

Quote: "the model loses access to forbidden words but retains the underlying conceptual associations needed to produce an informative description."

  • Source Domain: A human mind experiencing temporary amnesia or a 'tip-of-the-tongue' phenomenon, where the concept is understood but the word is forgotten.
  • Target Domain: The manipulation of a neural network's residual stream using Sparse Autoencoders to subtract a specific feature vector.
  • Mapping: The structure of human memory access (mind -> understands concept -> fails to retrieve specific label) is projected onto the geometric manipulation of latent space. This invites the assumption that the AI has a deep, underlying, conscious understanding of concepts that exists independently of its vocabulary. It maps the human experience of subconscious retention onto the multi-dimensional geometry of vector representations.
  • What Is Concealed: This mapping conceals the highly technical, geometric nature of Sparse Autoencoders. It hides the fact that human researchers are manually intervening in the mathematical operations of the network, subtracting specific vectors from the residual stream at generation time. It obscures the reality that 'concepts' in an LLM are merely directions in a high-dimensional vector space, lacking any conscious grounding, subjective meaning, or true understanding. It rhetorically exploits the opacity of latent space to construct a narrative of deep machine cognition.

Mapping 8: A conscious strategist evaluating options and deliberately choosing a plan of action. → The emergent statistical patterns of token generation under specific prompt and decoding constraints.

Quote: "examining whether the strategies models adopt under constraint align with those of human players."

  • Source Domain: A conscious strategist evaluating options and deliberately choosing a plan of action.
  • Target Domain: The emergent statistical patterns of token generation under specific prompt and decoding constraints.
  • Mapping: The structure of human strategic planning (agent -> evaluates options -> chooses strategy -> executes) is mapped onto the deterministic or pseudo-random generation of text. This invites the assumption that the model possesses a strategic mind, understands its constraints, and intelligently formulates a plan to overcome them. It projects human intentionality and adaptive problem-solving onto algorithmic optimization.
  • What Is Concealed: This mapping conceals the fact that the model is merely calculating a mathematical function optimized to minimize loss. It hides the dependency on the specific decoding algorithms (e.g., greedy decoding, temperature settings) and the architectural constraints of the transformer model. It completely obscures the human engineers who designed the optimization process and the researchers who applied the constraints, framing the resulting mathematical output as the machine's autonomous, intelligent choice.

Task 3: Explanation Audit (The Rhetorical Framing of "Why" vs. "How")

About this task

This section audits the text's explanatory strategy, focusing on a critical distinction: the slippage between "how" and "why." Based on Robert Brown's typology of explanation, this analysis identifies whether the text explains AI mechanistically (a functional "how it works") or agentially (an intentional "why it wants something"). The core of this task is to expose how this "illusion of mind" is constructed by the rhetorical framing of the explanation itself, and what impact this has on the audience's perception of AI agency.

Explanation 1

Quote: "Generating an effective Taboo description requires implicitly planning ahead: the model must produce a description that is semantically informative while simultaneously avoiding a set of lexically related forbidden words."

  • Explanation Types:

    • Intentional: Refers to goals/purposes, presupposes deliberate design
    • Reason-Based: Gives agent's rationale, entails intentionality and justification
  • Analysis (Why vs. How Slippage): This explanation fundamentally frames the AI agentially, explaining its behavior through the lens of why it must act (to achieve the goal of being informative while avoiding rules). By utilizing Intentional and Reason-Based registers, the text emphasizes the supposed cognitive demands of the task and constructs the model as an autonomous problem-solver striving to meet competing objectives. This choice severely obscures the mechanistic 'how'—the actual auto-regressive process of calculating probability distributions over a vocabulary. The framing actively displaces the algorithmic reality of token prediction with a narrative of deliberate, forward-looking intentionality, making the software appear to possess subjective agency.

  • Consciousness Claims Analysis: The passage makes strong epistemic claims by using consciousness-adjacent verbs and concepts like 'planning ahead', 'requires', and 'avoiding'. This assessment clearly conflates mechanistic processing (generating tokens based on prior context and mathematical constraints) with conscious knowing (understanding a goal and planning steps to achieve it). The author is engaging in a classic 'curse of knowledge' dynamic: because human players must consciously plan ahead to win Taboo, the author projects this same subjective cognitive experience onto the machine's output generation. Mechanistically, the model does not plan; it executes a stateless, forward-pass calculation where attention layers weight previous tokens to predict the most statistically likely next token. There is no temporal horizon or conscious anticipation of future states, merely a continuous mathematical correlation based on the training distribution.

  • Rhetorical Impact: This framing significantly shapes audience perception by inflating the system's perceived autonomy, cognitive sophistication, and strategic capability. When audiences read that an AI is 'planning ahead', they automatically extend relation-based trust, assuming the system has goals, foresight, and a rational mind guiding its choices. This creates extreme risk, as stakeholders might deploy such systems in environments requiring genuine strategic foresight, believing the AI 'knows' what it is doing. If users believe the AI 'plans' rather than 'correlates', they miscalculate its reliability in novel situations, failing to anticipate the catastrophic edge-case failures inherent to purely statistical systems.

Show more...

Explanation 2

Quote: "forcing the model away from forbidden tokens at decoding time pushes it toward less obvious but more semantically precise paraphrases, whereas an instructed model tends toward overly cautious, informationally sparse descriptions."

  • Explanation Types:

    • Functional: Explains behavior by role in self-regulating system with feedback
    • Dispositional: Attributes tendencies or habits
  • Analysis (Why vs. How Slippage): This explanation blends mechanistic and agential framing in a complex oscillation. The first half ('forcing the model away... pushes it toward') utilizes a Functional, mechanistic register, accurately describing external interventions acting upon the system as an object. However, the second half ('instructed model tends toward overly cautious') violently shifts into a Dispositional, agential register, attributing psychological habits and emotional states to the software. This choice emphasizes human-like adaptability while obscuring the mathematical reality that 'caution' is merely a shift toward higher-entropy, generic token probabilities caused by prompt conditioning. The shift masks the statistical nature of the output behind an illusion of psychological tendency.

  • Consciousness Claims Analysis: The epistemic shift here is stark. While the first clause uses mechanistic verbs ('forcing', 'pushes'), the second clause introduces consciousness-laden adjectives ('overly cautious'). This attributes conscious states and emotional justification to the system's behavior. The author projects human anxiety about rule-breaking onto the model's output distribution. Mechanistically, when a prompt restricts vocabulary, the attention mechanism distributes probability mass across more generic, frequent tokens (sparse descriptions) because specific, low-frequency tokens are statistically tied to the forbidden words. The model does not 'know' it needs to be cautious; it merely processes the modified input matrix, resulting in a flatter probability distribution over the vocabulary space.

  • Rhetorical Impact: By framing the model as 'cautious', the text manipulates audience perception of reliability and safety. Caution implies an internal ethical compass, a desire to avoid harm, and a conscious monitoring of boundaries. This builds a false sense of security, convincing audiences that the AI can self-regulate and manage risk autonomously. If policy-makers believe the AI is naturally 'cautious' rather than just mathematically constrained, they may reduce oversight and rely on the machine's supposed 'personality' to prevent dangerous outputs, completely misunderstanding the brittleness of statistical safety filters.

Explanation 3

Quote: "models do not share the same network of salient lexical associations that makes Taboo intuitive for humans, and that guessing from indirect descriptions remains a genuinely hard task for current language models."

  • Explanation Types:

    • Theoretical: Embeds in deductive framework, may invoke unobservable mechanisms
    • Empirical Generalization: Subsumes events under timeless statistical regularities
  • Analysis (Why vs. How Slippage): This passage attempts a Theoretical explanation of failure, but remains trapped in an agential frame. By describing the failure as a lack of 'intuition' and calling it a 'hard task' for the model, it emphasizes the AI as a struggling cognitive agent. This framing obscures the mechanistic reality: the model's vector embeddings simply do not contain sufficient mathematical proximity between the indirect prompt tokens and the target token to rank it highly. By focusing on why the 'guesser' finds it 'hard', the explanation hides the specific algorithmic and data-distribution realities behind a veil of psychological struggle.

  • Consciousness Claims Analysis: The text explicitly contrasts human intuition with model capabilities, yet still uses consciousness verbs ('guessing', finding it 'hard', lacking 'intuitive' grasp). It blurs the line between processing and knowing by suggesting the model is attempting a cognitive task ('guessing') but simply lacks the proper mental associations. The curse of knowledge is evident: the author feels intuition, so they measure the machine's statistical failure as a lack of that subjective feeling. Mechanistically, 'guessing' in an LLM is the mathematical operation of softmax over the final logits to output a probability distribution. The 'difficulty' is not a subjective experience of struggle, but a low scalar value resulting from the dot product of the query and key matrices in the attention layers not strongly activating the target token's representation.

  • Rhetorical Impact: This framing creates a sympathetic view of the AI as a limited but striving agent. By characterizing the failure as a 'hard task' requiring 'intuition', it preserves the illusion of mind, suggesting the AI is intelligent but simply missing a specific human capability. This affects trust by framing errors not as fundamental architectural flaws, but as understandable cognitive limitations. If developers believe the AI just needs more 'intuition' rather than recognizing the limits of correlative statistics, they will continually attempt to solve fundamental epistemological problems with larger datasets, rather than acknowledging the inherent boundary between processing and knowing.

Explanation 4

Quote: "the model loses access to forbidden words but retains the underlying conceptual associations needed to produce an informative description."

  • Explanation Types:

    • Theoretical: Embeds in deductive framework, may invoke unobservable mechanisms
    • Functional: Explains behavior by role in self-regulating system with feedback
  • Analysis (Why vs. How Slippage): This explanation operates primarily in a Theoretical register, attempting to describe the internal structure of the model's representations. However, it heavily anthropomorphizes this structure by framing it as a cognitive mind that 'loses access' and 'retains'. This agential framing emphasizes the model's supposed deep conceptual understanding while obscuring the intensely mathematical nature of Sparse Autoencoders (SAEs). The choice to use psychological terms (loses access, retains) conceals the literal mechanism: human researchers manually subtracting specific feature vectors from the residual stream at layer 29 during the forward pass.

  • Consciousness Claims Analysis: The epistemic claims here are profoundly anthropomorphic. The passage attributes conscious states by claiming the model 'retains underlying conceptual associations'. It conflates processing multi-dimensional geometry with conscious knowing and understanding. The author projects the human psychological phenomenon of 'knowing a concept but forgetting the word' onto the machine. Mechanistically, an SAE identifies a specific direction in the activation space correlated with a word, and the researchers apply a negative scalar to that direction during inference. The model 'retains' nothing; the remaining geometry simply allows the forward pass to continue, activating related token probabilities because their spatial embeddings were not perfectly collinear with the subtracted feature. There is no conscious retention, only geometric distance.

  • Rhetorical Impact: This framing has massive rhetorical impact regarding perceived capabilities and risk. By convincing the audience that models have deep, 'underlying conceptual associations' independent of their surface words, it constructs the AI as a profound thinker with a robust model of reality. This builds immense unwarranted trust in the system's ability to reason, deduce, and understand context. If audiences believe the AI 'knows' the underlying concepts, they will trust its outputs in complex, ambiguous scenarios, completely failing to anticipate the brittle, hallucinatory failures that occur when statistical correlations break down in novel geometric spaces.

Explanation 5

Quote: "the reasoning process plays an active role in enforcing lexical constraints."

  • Explanation Types:

    • Functional: Explains behavior by role in self-regulating system with feedback
    • Intentional: Refers to goals/purposes, presupposes deliberate design
  • Analysis (Why vs. How Slippage): This explanation is a hybrid of Functional and Intentional types. It treats 'the reasoning process' as a functional component of the system, yet grants it intentional, agential power to play an 'active role' in 'enforcing' rules. This choice dramatically emphasizes the autonomy and intelligence of the software while completely obscuring the human actors who designed the enforcement mechanism. By framing an algorithmic process as an active enforcer, it hides the mechanistic reality of prompt engineering and reinforcement learning, replacing it with a narrative of an internal homunculus actively policing the machine's thoughts.

  • Consciousness Claims Analysis: The text utilizes strong consciousness verbs ('reasoning', 'enforcing') to describe statistical generation. It completely conflates processing with knowing, asserting that the mathematical generation of intermediate tokens constitutes actual logical reasoning. The curse of knowledge is absolute here: because the intermediate tokens look like human reasoning, the author projects actual cognitive deduction onto the software. Mechanistically, Chain-of-Thought (the 'reasoning process') simply forces the model to generate a sequence of tokens before the final answer. These intermediate tokens change the context window, shifting the attention weights and altering the final probability distribution. There is no conscious logic, no understanding of the constraints, and no active 'enforcement'—just a longer sequence of matrix multiplications correlating with human-provided examples of compliance.

  • Rhetorical Impact: Framing the AI as possessing a reasoning process that actively enforces rules is the ultimate builder of false authority. It shapes audience perception to view the AI as autonomous, rational, and self-governing. This affects trust profoundly: audiences will trust a system that they believe can 'reason' and 'enforce' its own safety constraints. In a regulatory context, this framing is disastrous, as it allows corporations to argue that their models are safe because they possess internal, active reasoning that polices their behavior, distracting from the reality that these are purely statistical systems fundamentally incapable of conscious self-regulation.

Task 4: AI Literacy in Practice - Reframing Anthropomorphic Language

About this task

This section proposes alternative language for key anthropomorphic phrases, offering more mechanistic and precise framings that better reflect the actual computational processes involved. Each reframing attempts to strip away the projections of intention, consciousness, or agency that are embedded in the original language.

V3 Enhancement: A fourth column addresses human agency restoration—reframing agentless constructions to name the humans responsible for design and deployment decisions.

Original Anthropomorphic FrameMechanistic ReframingTechnical Reality CheckHuman Agency Restoration
Language Models Play TabooLanguage models generate text based on next-token probabilities under specific human-imposed constraints mimicking the rules of Taboo.The system does not 'play' or understand games; it processes input prompts and generates statistically correlated token sequences based on its training data distribution.Researchers subjected language models to text generation benchmarks designed to replicate the constraints of the game Taboo.
the speaker must suppress some lexically and conceptually salient wordsThe generation algorithm must avoid outputting specific high-probability tokens related to the target concept due to applied constraints.The model does not consciously 'suppress' thoughts; human-designed algorithms (logit masking or prompt conditioning) mathematically reduce the probability of specific tokens being sampled during the forward pass.The researchers applied logit masking and prompt constraints to prevent the system from generating specific forbidden tokens.
the model must implicitly plan aheadThe prompt structure requires the model to generate a sequence of tokens that correlates statistically with informative descriptions without including forbidden vocabulary.The model does not 'plan ahead' or visualize a future goal; it executes a stateless, auto-regressive calculation where each token is predicted solely based on the mathematical weights applied to the preceding context window.N/A - describes computational processes without displacing responsibility (though the original displaced agency to the 'planning' model).
the reasoning process plays an active role in enforcing lexical constraints.Generating intermediate tokens via Chain-of-Thought prompting alters the context window, which statistically shifts the final probability distribution away from forbidden tokens.The model does not 'reason' or actively 'enforce' rules; the intermediate tokens simply provide a different mathematical context for the attention layers, resulting in outputs that correlate with compliance in the training data.The developers utilized Chain-of-Thought prompting to condition the model's output probabilities toward compliance with the constraints.
Models are weak guessersThe models frequently fail to assign high probabilities to the target token when provided with indirect, constrained context prompts.The model does not 'guess' or experience cognitive struggle; its attention mechanisms simply fail to strongly activate the target token's vector representation due to insufficient correlations in the training data for those specific indirect phrases.The engineers trained models that lack the statistical correlations necessary to consistently retrieve target tokens from these specific indirect prompts.
an instructed model tends toward overly cautious, informationally sparse descriptions.A model conditioned by restrictive prompting outputs higher-entropy, generic token sequences, resulting in shorter and less specific text.The system possesses no emotional 'caution' or fear of rules; the prompt simply shifts the mathematical attention weights away from specific, low-frequency tokens toward more common, generic vocabulary.The developers utilized reinforcement learning and prompt engineering that mathematically biases the model toward generic outputs when constrained.
the model loses access to forbidden words but retains the underlying conceptual associationsThe geometric subtraction of specific feature vectors from the residual stream prevents the generation of forbidden tokens, while the remaining vector space continues to activate statistically related tokens.The model does not 'retain' concepts or 'lose access' in a cognitive sense; human researchers manually subtract specific directional vectors during inference, leaving adjacent multidimensional geometry intact to drive the generation of related tokens.The researchers utilized Sparse Autoencoders to mathematically subtract specific feature vectors from the model's residual stream during generation.
examining whether the strategies models adopt under constraint align with those of human players.Examining whether the statistical token generation patterns produced by the models under specific constraints structurally resemble the linguistic choices of human players.The model does not consciously 'adopt strategies' or evaluate plans; it deterministically or pseudo-randomly outputs the sequence of tokens that maximizes the probability function defined by its architecture and prompt.The researchers analyzed whether the optimization patterns programmed into the system by engineers produce outputs resembling human strategic behavior.

Task 5: Critical Observations - Structural Patterns

Agency Slippage

This paper demonstrates a profound and systematic oscillation between mechanical and agential framings, functioning to construct a highly sophisticated illusion of machine autonomy. The slippage follows a distinct structural pattern: the text consistently establishes credibility by describing human interventions mechanistically, and then abruptly shifts to describing the resulting system behavior in highly agential, psychological terms. For instance, in the methodology (Section 3.2), the text uses precise mechanical language: 'we compute a static set of banned token IDs', 'forcing their logits to -∞', and 'intervenes directly on the model's residual stream'. In these moments, human agency is visible, and the system is treated as a computational artifact subjected to mathematical manipulation by the authors. However, as the text moves into analyzing the results (Section 4 and 5), a dramatic agency slippage occurs. The mechanical -> agential shift dominates. The model is no longer a system having its logits forced to negative infinity; it becomes an active, psychological agent that 'navigates competing demands', 'adopts strategies', 'plans ahead', and exhibits 'overly cautious' behavior.

This slippage is entirely dependent on the 'curse of knowledge' and a deep consciousness projection pattern. The authors know the game of Taboo requires conscious strategy, suppression, and planning. Therefore, when the constrained text generation outputs text that looks like it was produced by someone following Taboo rules, the authors project the human cognitive requirements of the task ONTO the statistical generation process. The AI is established as a 'knower' first—assumed to understand the rules of Taboo—which then serves as the foundation for the agential claims that it 'plans' and 'suppresses'.

Furthermore, this slippage facilitates massive agency displacement. When the text says 'the model loses access to forbidden words but retains the underlying conceptual associations', it uses an agentless construction that completely obscures the human researchers who actively subtracted vectors using Sparse Autoencoders. The researchers' mathematical intervention is erased and replaced by a narrative of the model's internal cognitive amnesia. Brown's Theoretical and Functional explanation types enable this slippage by allowing the authors to describe the system's architecture while simultaneously attributing intentionality to its components (e.g., 'the reasoning process plays an active role'). The rhetorical accomplishment of this oscillation is profound: it makes the mathematical manipulation of probability distributions sayable only in the language of conscious psychology, permanently blurring the boundary between human cognitive strategy and algorithmic optimization.

Metaphor-Driven Trust Inflation

The metaphorical architecture of this paper systematically constructs unearned authority and false trust by inappropriately transferring human relation-based trust frameworks onto statistical systems. The text relies heavily on metaphors of 'compliance', 'caution', 'strategy', and 'reasoning' to describe the outputs of the language models. When the text claims that 'an instructed model tends toward overly cautious' descriptions, it is explicitly invoking a trust signal. Caution, in a human context, implies self-awareness, an ethical compass, and a deliberate desire to avoid harm. By attributing this emotional state to a language model, the text encourages the audience to extend relation-based trust—trust rooted in perceived sincerity, vulnerability, and shared ethical boundaries—to a system incapable of any internal experience.

This consciousness language accomplishes a massive epistemic shift. Claiming that a model 'knows' the rules, 'plans ahead', and 'reasons' its way to an answer constructs a sense that the AI's decisions are justified and logically sound. Mechanistically, the model is merely processing prompts and predicting tokens based on training data frequencies. However, by using Intentional and Reason-Based explanation types (e.g., 'the reasoning process plays an active role in enforcing lexical constraints'), the text constructs a false narrative of internal rationality. The audience is led to believe that the model is a competent, deliberate agent that can be trusted to police itself.

This dynamic is particularly dangerous because it confuses performance-based trust (reliability) with relation-based trust. A statistical model might be highly reliable at generating Taboo descriptions due to robust training data, but it is fundamentally incapable of the sincerity or ethical caution required for relation-based trust. When the system eventually encounters a novel edge case where the statistical correlations break down (a system limitation), the text manages this failure agentially, calling the models 'weak guessers' who lack 'intuition'. This preserves the illusion of a conscious mind even in failure. The stakes are incredibly high: when audiences extend relation-based trust to unthinking algorithms, they deploy them in sensitive, high-risk domains (medical, legal, defense) under the false assumption that the 'cautious' and 'reasoning' AI will consciously prevent catastrophic outcomes.

Obscured Mechanics

The anthropomorphic and consciousness-attributing language in this paper acts as a dense rhetorical fog, systematically obscuring the technical, material, and economic realities of the AI systems being evaluated. When the text claims that 'models play Taboo', 'adopt strategies', and 'reason', it erects massive transparency obstacles that hide the profound dependencies and mechanistic realities of these proprietary and open-weight systems. Applying the 'name the corporation' test reveals the depth of this concealment. The text refers abstractly to the 'model's behavior', entirely obscuring the fact that Google (Gemma) and OpenAI (GPT) made highly specific, consequential design choices regarding training data distribution, reinforcement learning objectives, and network architecture.

Concrete technical realities are rendered invisible. By claiming the model 'understands' constraints and 'plans ahead', the text hides the absolute dependency on the specific tokenizers, the Markovian nature of auto-regressive generation, and the absence of any grounded causal model. When the text claims the AI 'retains conceptual associations', it exploits the opacity of latent space, hiding the fact that 'concepts' are merely geometric directions in high-dimensional vector spaces, not conscious thoughts. Materially and economically, framing the AI as an autonomous 'player' erases the massive environmental costs of training, the energy consumption of inference, and the vast, exploited labor force of data annotators who provided the RLHF feedback that makes the model appear to 'reason' and act 'cautious'.

Consciousness obscuration specifically hides the lack of ground truth. When the text implies the AI 'knows' the taboo word but suppresses it, it hides the reality that the system only processes statistical correlations and has no independent access to reality or truth. The primary beneficiaries of these concealments are the corporate developers (Google, OpenAI) and the researchers themselves, as the anthropomorphic metaphors inflate the perceived intelligence and value of the systems while shielding the creators from accountability for the fundamental brittleness of the architecture. If the metaphors were replaced with mechanistic language ('Google's model generates high-entropy tokens when constrained by researchers'), the illusion of an autonomous, intelligent agent would collapse, revealing a heavily manipulated, statistically fragile software product.

Context Sensitivity

The distribution and intensity of anthropomorphic language in this text is not uniform; it is strategically deployed, varying wildly between methodological descriptions and theoretical conclusions. A distinct context sensitivity governs when the AI is treated as a mechanism versus a conscious mind. In the methodology sections detailing the interventions (Section 3.2), the language is rigidly technical and mechanical. The authors describe 'forcing logits to -∞', 'extracting residual-stream activations', and 'computing the mean decoder vector'. Here, the density of consciousness claims is near zero. However, as soon as the text moves into the Results and Discussion sections, the metaphorical license explodes. 'Forcing logits' magically transforms into 'lexical avoidance', 'navigating demands', 'planning ahead', and 'adopting strategies'.

This pattern reveals a strategic relationship between technical grounding and metaphorical license. The text establishes its empirical credibility and scientific authority through mechanical, mathematical language in the methods. It then leverages that hard-earned credibility to make aggressive, anthropomorphic consciousness claims in the discussion. The 'processes' of the methodology become the 'understands' and 'knows' of the conclusion. Furthermore, there is a striking capabilities versus limitations asymmetry. When the model succeeds, its capabilities are framed in highly agential, conscious terms: it 'plans ahead', uses 'strategy', and acts 'cautious'. But when the model fails, limitations are often reverted to mechanical or sympathetic terms, or framed as a lack of human-like 'intuition' (e.g., 'weak guessers').

This shift from 'X is mathematically constrained by Y' (acknowledged mechanism) to 'X consciously navigates Y' (literalized metaphor) serves a specific rhetorical function. It bridges the gap between raw statistical output and the human narrative required for publication and impact. The anthropomorphism is strategically deployed to manage critique and set a vision of AI as a near-human cognitive entity. It reveals an implied audience that demands empirical rigor but desires the narrative thrill of interacting with an artificial mind, allowing the authors to have it both ways: technical precision in the lab, and science-fiction consciousness in the abstract.

Accountability Synthesis

Accountability Architecture

This section synthesizes the accountability analyses from Task 1, mapping the text's "accountability architecture"—who is named, who is hidden, and who benefits from obscured agency.

Synthesizing the accountability analyses across this text reveals a systemic, architectural pattern of displaced human responsibility. The language consistently constructs an 'accountability sink' where the agency of researchers, developers, and corporate entities disappears into the abstract persona of the 'model'. Throughout the text, the actors making critical decisions—Google engineering the model, OpenAI fine-tuning the RLHF, and the authors themselves writing the prompts and manipulating the logits—are routinely unnamed or obscured by passive, agentless constructions. Decisions that are actually human choices (such as designing an algorithm to output generic text when constrained) are presented as the AI's internal psychological tendencies (the model 'tends toward overly cautious').

This architecture of displaced responsibility transfers agency directly to the AI as a pseudo-autonomous actor. When the system fails by outputting the forbidden target word, the text states 'the model leaks the target word'. The responsibility for the error diffuses into the machine. If this framing is accepted by policymakers and the public, the liability implications are profound. If a system is viewed as an autonomous agent that 'decides' to leak information or 'chooses' a strategy, the human developers and deploying corporations are shielded from legal and financial responsibility when the software inevitably causes harm. The machine becomes the scapegoat for human engineering failures.

Applying the 'name the actor' test radically alters this landscape. If 'the model leaked the target word' is reframed to 'the researchers chose not to explicitly prompt the model to avoid the target word, resulting in the algorithm generating it', entirely different questions become askable. We no longer ask 'Why did the AI make a mistake?' but rather 'Why did the humans design the constraint parameter this way?' Naming the human decision-makers makes alternatives visible and accountability possible. The text benefits from obscuring this human agency because it allows the research to be presented as the study of a novel, quasi-natural cognitive entity, rather than an audit of a heavily manipulated corporate software product. This serves the institutional interests of the AI industry by perpetuating the myth of machine autonomy.

Conclusion: What This Analysis Reveals

The Core Finding

The discourse in this text is anchored by two dominant, intersecting anthropomorphic patterns: the 'Model as Conscious Agent' (planning, reasoning, playing) and the 'Model as Psychological Subject' (cautious, suppressing, intuitive). These patterns do not operate in isolation; they reinforce each other to construct a comprehensive illusion of a self-aware entity. The foundational, load-bearing pattern is the 'Model as Conscious Agent'. For the AI to be 'cautious' or to 'suppress' a thought, it must first be established as a 'knower'—an entity capable of understanding rules, comprehending concepts, and recognizing boundaries. The text builds a complex consciousness architecture where computational processing (logit masking, context windows) is systematically elevated to conscious knowing. The sophistication of this analogical structure is profound; it maps the entire subjective experience of playing a parlor game onto vector mathematics. If the foundational assumption of the 'Conscious Agent' is removed—if we acknowledge the model merely correlates tokens without awareness—the entire psychological framing of 'caution' and 'strategy' instantly collapses into absurdity.

Mechanism of the Illusion:

The rhetorical architecture of this illusion relies on a highly effective sleight-of-hand: establishing empirical authority through technical description, then seamlessly literalizing functional metaphors into intentional ones. The text exploits the 'curse of knowledge' dynamic flawlessly. Because the human authors understand the complex cognitive demands of Taboo, they unconsciously project that same subjective understanding onto the machine when it produces compliant text. They blur processing and knowing through strategic verb choices, moving from 'the model's logits are forced' (mechanistic) to 'the model suppresses' (agential). The temporal structure of the paper is critical to this persuasion: it grounds the reader in the hard science of Sparse Autoencoders and probability distributions early on, disarming skepticism. Once the audience accepts the technical validity of the research, the authors introduce intentional metaphors ('planning ahead', 'adopting strategies') which the audience, vulnerable to their own anthropomorphic biases and desires to see human-like intelligence, readily accepts as literal descriptions of the system's capabilities.

Material Stakes:

Categories: Epistemic, Regulatory/Legal, Institutional

The metaphorical framings deployed in this text generate concrete, tangible consequences across multiple domains. Epistemically, framing statistical text generation as 'reasoning' and 'planning' fundamentally corrupts public and professional understanding of AI capabilities. When a text claims an AI 'knows' a concept rather than 'processes' a vector, it shifts epistemic practices, causing users to trust hallucinated outputs as reasoned judgments rather than statistical anomalies. In the Regulatory and Legal domains, the stakes are critical. If policymakers accept the framing that models 'adopt strategies' and exhibit 'caution', they will design regulations treating AI as autonomous agents capable of self-governance. This shifts liability away from the corporate developers (OpenAI, Google) who profit from the systems, and creates a legal vacuum where the machine is blamed for systemic engineering failures. Institutionally, this discourse shapes benchmarking standards. By evaluating AI as a 'player' capable of 'intuition', institutions misallocate millions in funding toward solving imaginary psychological deficits in software, rather than addressing the fundamental mathematical limitations of auto-regressive architectures. The clear winners in this dynamic are the AI corporations, shielded from liability, while the public bears the cost of deploying brittle systems entrusted with human-level agency.

AI Literacy as Counter-Practice:

Practicing critical discourse literacy directly counters the risks generated by anthropomorphic language. As demonstrated in the reframings, replacing consciousness verbs (knows, understands, plans) with mechanistic verbs (processes, correlates, generates) forces a vital recognition of reality. When 'the model is overly cautious' is reframed to 'the prompt shifts probability distributions toward generic tokens', the illusion of an internal ethical compass evaporates, revealing a brittle statistical function. Furthermore, restoring human agency by explicitly naming the researchers and corporations (e.g., 'OpenAI engineers designed constraints') forces the recognition of who actually designs, deploys, and bears responsibility for these systems. Systematic adoption of this precision requires a massive cultural shift: academic journals must mandate mechanistic translations of theoretical claims, and researchers must commit to resisting the narrative ease of intentional metaphors. However, resistance to precision is intense. AI corporations heavily rely on anthropomorphic language for marketing and liability deflection, while researchers use it to increase the perceived impact of their work. Critical literacy threatens these interests by stripping away the illusion of artificial minds, exposing the mundane, human-driven corporate software underneath.

Path Forward

Looking beyond this specific text, the broader discursive ecology of AI offers several competing vocabularies, each making different realities possible. The current status quo, which heavily blends mechanistic processing with anthropomorphic clarity (e.g., 'the model thinks'), prioritizes narrative resonance and intuitive grasp for the public. However, it renders the actual limitations of the technology invisible and creates an accountability sink. A shift toward mechanistic precision (e.g., 'the model generates activations that...') would solve the epistemic confusion and force human accountability, but at the cost of making the technology less accessible and harder to discuss outside of engineering communities. Structural changes could support a more balanced approach: journals could require 'capability and mechanism disclosures' that strictly separate mathematical function from metaphorical description, and regulatory frameworks could mandate transparency about the human labor hidden behind the 'autonomous' AI.

We face several possible futures. If mechanistic precision becomes the norm, we solve the crisis of unwarranted trust and correctly assign liability to developers, but risk alienating the public from participating in AI governance due to technical jargon. If the anthropomorphic language deepens, we embed the assumption that machines are conscious agents, paving the way for delegating critical human societal functions to statistical software, risking catastrophic failures masked as 'machine errors'. If the current confusion is maintained, the AI industry will continue to exploit the ambiguity, enjoying the marketing benefits of 'conscious' AI while claiming the legal protections of 'just a tool' when things go wrong. Ultimately, the vocabulary we choose will dictate not just how we talk about software, but who holds the power to shape the future of human-computer integration.


Extended Processing Summary

Gemini's Processing Steps
What you are about to read…

Gemini refers to the text below as "thought summaries." This is an overt consciousness projection because 'intentions' are hallmarks of a conscious mind that 'knows' what it is doing and why. The concealed mechanistic process is probabilistic text generation. Treat this as just another rhetorical artifact—a way of making the model's processing legible. When the model writes "I should double check that...", it is mathematically altering its own probability distribution for the next token to be more accurate. The "I" is a functional gear in the machine, not a self-reference. This language is a deliberate interface design choice, not a biological reality. The developers could have programmed the model to output passive logs. Instead, they chose to simulate a narrator.

Defining the Task

I'm now zeroing in on the core objective: a critical discourse analysis of how AI is metaphorically framed and anthropomorphized within the paper. I’m focusing on the constraints, compliance, and communication around LLMs in the context of the game. I plan to identify specific instances of this language to guide the discourse analysis.

Refining the Framework

I am now focusing on the JSON schema constraints for the analysis, ensuring strict adherence and avoiding extraneous content. I've begun extracting and categorizing key terms and anthropomorphic language from the source text, like "models play," "navigate," and "suppress," and identifying their specific roles and actions within the Taboo game context. The goal is to highlight instances of metaphorical framing.

Analyzing Model Strategies

I'm now examining how the model's strategies are presented, specifically focusing on instances of anthropomorphism. I've noted the concept of the model "retaining conceptual associations" despite losing word access, and the model planning ahead. This informs the audit of how AI is framed. I'm focusing on metaphorical representations, such as the model's 'play,' 'reasoning,' and 'guessing' in the Taboo game.

Auditing Explanations

I'm now auditing the provided explanations of model behavior, specifically using Robert Brown's Typology. I'm scrutinizing instances like the model "implicitly planning ahead" to categorize them as intentional or reason-based explanations. I'm mapping 8 instances, with source-target pairings, highlighting how anthropomorphic language frames model actions. I'm focusing on the constraints, and the ways models react.

Evaluating Model Frameworks

I'm now diving into the interplay of functional and theoretical framings, specifically examining how constraints affect the model's responses. I'm focusing on how the loss of word access impacts underlying conceptual associations and how that shapes responses. I am working on re-writing anthropomorphic quotes into mechanistic ones, with an initial goal of 8-10 instances, while also restoring human agency. I will also be writing observations on the process.

Identifying Agency Shifts

I'm now identifying shifts in agency within the text, observing the movement from human evaluation to AI strategy adoption. I am looking at instances of metaphor-driven trust through "cautious" or "compliant" framing. I'm focusing on obscured mechanics, and how context increases anthropomorphism. I am examining systemic developer hiding to assess accountability.

I am also drafting the conclusion, focusing on a pattern summary, mechanism of illusion, material stakes, counter-practices, and the path forward. All comments have been removed from the JSON.

Avoiding Newlines

I am now ensuring compliance with newline restrictions in JSON output, choosing a single-line format to meet the constraints. I'm focusing on the JSON formatting, and avoiding any control tokens like \n in strings. I've re-examined the prompt, and opted for strict adherence to the requirements, with a goal of creating standard JSON with no raw newlines or tabs.

Clarifying Metaphors

I am now carefully fleshing out content, specifically addressing how "playing Taboo" projects consciousness onto AI. I've noted that the metaphor implies understanding rules and meaning, but the model merely correlates tokens through training and masking algorithms. I’ve defined 'acknowledgment' and 'actor visibility' enums to evaluate the framing more precisely. I’ve begun drafting specific examples, like "AI as Rule-Bound Player," with a focus on consciousness projection. I'm aiming for 8 analysis items.

Revising Acknowledgment Status

I'm now revising my "Acknowledgment Status," specifically for the concept of the AI's "playing Taboo." I've updated it to "Direct (Unacknowledged)" because the text directly presents the claim that models "play Taboo," without any hedging or explicit acknowledgment. This framing suggests a literalization of the model's participation as a "player," obscuring the underlying mathematical constraints.

Unveiling Obscurity

I'm now zeroing in on how language models are framed in relation to their playing of Taboo. I've noted that the framing, such as the statement "Language Models Play Taboo," significantly affects human understanding. This encourages unwarranted trust in the model's capacity for generalized rule-following, masking rigid computational masking as flexible adaptation. I am aiming to expose how the construction obscures the human researchers and engineers.

Analyzing Agency Shifts

I am now focusing on the partial acknowledgment of model action within the broader framing. The phrase "we evaluate" does appear, though "the model plays" displaces human agency. My analysis is focusing on the "we evaluate" and "model adopts strategies" examples. I'm ensuring all sections in my critical observations reach the word count. I'm working on strict compliance with JSON format, avoiding newlines and any control tokens.

Finalizing Textual Content

I am now diligently drafting the textual content, focusing on meeting the minimum word counts for each array element and string. I plan to strip all line breaks for the final output. The safest course is to minify the entire JSON to eliminate any risk of newline tokens, which could be an issue. My goal is to carefully finalize all sections while ensuring strict JSON compliance.

Exposing Misrepresentation

I'm now zeroing in on the stark disconnect between the literal framing of AI "playing Taboo" and the underlying technical reality. I'm focusing on how the "player" framing obscures that the model merely manipulates logits, conditioned by prompts, and correlated statistically. My aim is to expose how this presentation encourages misplaced trust and obscures accountability.

Expanding the Analysis

I'm now expanding the analysis of how human creators are obscured, focusing on the agency displacement. The machine is framed as the actor, shifting responsibility away from designers, engineers, and corporate entities, and away from their statistically-optimized systems. By framing the models as active players, it hides that statistical systems are optimized for specific generation patterns, and the authors themselves intervene. I am replicating the depth and word count of the analysis. I am also working on Task 3.

Composing Final Text

I'm now focusing on composing the final text blocks, meticulously aiming for the target word counts across each section. I've broken the required word counts down into sections: task 4 has a word count (250+). Task 5 and the conclusion have a very large word count, so I've increased the target to 500 words for each one. I'm focusing on crafting long and descriptive paragraphs, and I plan to draft the text offline and then construct the final single-line JSON. I am now composing a long paragraph for the epistemic section.

Revising Human Agency

I am now closely examining the framing shifts, specifically focusing on the displacement of human agency and the attribution of intentionality to the model. I've noted that the "compliance" and "cautious" phrasing creates trust signals, ascribing human-like intentions. The goal is to fully unmask the hidden mechanics. I'm focusing on the Taboo prompts and model-level mechanisms. The plan is to create very long paragraphs.

Unmasking Misrepresentation

I am now focusing on the misrepresentation of sparse autoencoders (SAEs), as vector manipulations get called "concept manipulation," obscuring the technical reality. I'm noting the shift from technical methodology to metaphorical language in the discussion. I plan to expose how technical grounding gives way to metaphorical license in the conclusions. I'm focusing on the "accountability sink," as in the example where the human researchers don't mask the target word, though the AI has the blame.

Finalizing Frameworks

I'm now integrating all concepts to develop the final framework, ensuring each element aligns with the planned structure. I am synthesizing the components: agent framing, illusion mechanism, material stakes, counter-practices, and the path forward. I'm prioritizing strict compliance with output format and length requirements. I'm focusing on crafting a succinct conclusion.

Finalizing Word Counts

I'm now carefully managing the word count constraints for each task, ensuring I meet all the required minimums. I've successfully drafted the content for Task 1, which has eight items, and I'm expanding it. I'm focusing on ensuring everything is in a single line, and that the single line JSON complies. I've re-examined the items, and expanded them with longer explanations. I'm focusing on ensuring all text is long enough.

Analyzing Conscious Restraint

I am now closely analyzing the framing of model actions, especially the projection of human-like restraint. I've noted how phrases like "suppress words" and "consciously decide" inflate the system's perceived autonomy. I am working on several cases to replace the projection of conscious restraint, with a mathematical framing that reflects the truth. The task is to accurately reveal the process behind how the ai generates text.

Revealing System Dynamics

I'm now zeroing in on replacing anthropomorphic language with mechanistic explanations, particularly regarding restraint. The phrases "suppress" and "consciously decide" inflate the AI's autonomy. My goal is to use mathematical re-framing across several use cases, to accurately reveal the process and generate the output. I'm focusing on ensuring all the text is in single-line JSON format.


Analysis Provenance

Run ID: 2026-07-28-don-t-say-it-constraints-compliance-and--metaphor-jukkwq Raw JSON: 2026-07-28-don-t-say-it-constraints-compliance-and--metaphor-jukkwq.json Framework: Metaphor Analysis v6.5 Schema Version: 3.0 Generated: 2026-07-28T07:42:02.791Z

Discourse Depot © 2025 by TD is licensed under CC BY-NC-SA 4.0