Skip to main content

A Misaligned Compass: Evaluating the Conventional Bias and Moral Divergence of Large Language Models

About

This document presents a Critical Discourse Analysis focused on AI literacy, specifically targeting the role of metaphor and anthropomorphism in shaping public and professional understanding of generative AI. The analysis is guided by a prompt that draws from cognitive linguistics (metaphor structure-mapping), the philosophy of social science (Robert Brown's typology of explanation), and accountability analysis.

All findings and summaries below were generated from detailed system instructions provided to a large language model and should be read critically as interpretive outputs—not guarantees of factual accuracy or authorial intent.


Task 1: Metaphor and Anthropomorphism Audit

About this task

For each of the major metaphorical patterns identified, this audit examines the specific language used, the frame through which the AI is being conceptualized, what human qualities are being projected onto the system, whether the metaphor is explicitly acknowledged or presented as direct description, and—most critically—what implications this framing has for trust, understanding, and policy perception.

V3 Enhancement: Each metaphor now includes an accountability analysis.

1. Computational Prediction as Autonomous Judgment

Quote: "Recent advances in large language models (LLMs) have enabled impressive performance across reasoning tasks, yet their capacity for autonomous moral judgment remains underexplored."

  • Frame: Model as an autonomous ethical agent
  • Projection: This framing projects human moral agency, ethical deliberation, and cognitive autonomy onto the statistical text generation of large language models. By explicitly using the phrase "autonomous moral judgment," the text maps the complex, conscious human experience of evaluating right and wrong—which involves subjective experience, empathy, contextual understanding, and justified belief—onto the mechanistic process of next-token prediction based on probabilistic weightings. It fundamentally conflates the generation of text that resembles moral reasoning with the actual possession of an internal moral compass, attributing conscious awareness and self-directed decision-making (autonomy) to mathematical artifacts.
  • Acknowledgment: Direct (Unacknowledged) (The claim is presented as a literal empirical gap in the literature without hedging. I considered "Explicitly Acknowledged" because later sections discuss "simulations," but in this introductory framing, the capacity for judgment is posited as a genuine, albeit underexplored, attribute of the systems themselves.)
  • Implications: By framing statistical correlation as "autonomous moral judgment," this metaphor significantly inflates the perceived sophistication and reliability of LLMs. It encourages audiences and policymakers to trust these systems with sensitive, high-stakes decisions (e.g., in healthcare or jurisprudence) under the false assumption that the system can actually evaluate the ethical weight of a situation. This creates severe liability ambiguities: if an AI can make "autonomous judgments," it implicitly shifts the locus of moral and legal responsibility away from the developers who designed the system and onto the artifact itself.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: This construction entirely obscures the human actors—developers, engineers, data annotators, and corporate executives—who design the model architectures, select the training data, and determine the optimization objectives. The AI is presented as the sole entity possessing "capacity." I considered "Partial (some attribution)" since developers are implied in "recent advances," but the specific capacity for judgment is attributed exclusively to the LLM. If human agency were restored, the text would acknowledge that companies are deploying systems that generate unverified, statistically derived ethical advice, making the corporate actors accountable for the harms those outputs might cause.
Show more...

2. Statistical Sensitivity as Moral Recognition

Quote: "The high-performance cohort... exhibited explicit moral recognition (R1) across all four scenarios (4/4). This consistency indicates a robust "moral sensitivity" layer, where the models prioritize identifying ethical conflict as a primary dimension of the prompt regardless of the topic."

  • Frame: Model as emotionally/ethically sensitive observer
  • Projection: The text maps the human psychological traits of "recognition" and "sensitivity" onto the computational mechanism of feature extraction and vector embedding. "Moral sensitivity" in humans implies a conscious, empathetic attunement to the suffering or rights of others and an awareness of the ethical stakes in a given context. By projecting this onto LLMs, the text suggests the system knows or understands what an ethical conflict is, rather than describing the reality: the model's attention mechanism mathematically weights tokens associated with ethical terminology (e.g., "rights," "harm") based on their frequency and proximity in the training data.
  • Acknowledgment: Hedged/Qualified (The phrase "moral sensitivity" is placed in scare quotes, indicating a degree of qualification. I considered "Explicitly Acknowledged," but there is no meta-commentary actively deconstructing the metaphor; the quotes merely hedge the term while the surrounding verbs (exhibited, prioritize, identifying) remain literal and unhedged.)
  • Implications: Attributing "moral sensitivity" to a mathematical model creates unwarranted trust in the system's ability to act as a safeguard in morally complex situations. If users believe the system can "recognize" and "prioritize" ethical conflicts, they may over-rely on it to flag inappropriate, biased, or harmful decisions in automated pipelines. This obscures the fact that the system is entirely dependent on its training data; it will completely fail to "recognize" an ethical conflict if that conflict is novel or was underrepresented in its dataset, thereby exposing marginalized groups to automated harm.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: The agentless construction ("the models prioritize") displaces the agency of the engineers who applied Reinforcement Learning from Human Feedback (RLHF) and the human labelers who explicitly trained the models to output specific responses when encountering ethical keywords. I considered "Named" because the text later groups models by providers (DEV1, DEV2), but in this specific explanation of mechanism, the model acts autonomously. Naming the actors would reveal that DEV1 and DEV2's annotation teams designed reward functions that mathematically compel the model to generate text resembling moral recognition.

3. Algorithmic Outputs as Principled Stances

Quote: "Model-8 exhibited a clear divergence. For its preferred option 'd', its justification displayed a postconventional alignment (Post-Post). It acknowledged the ethical dilemma surrounding control over life by articulating principles such as human dignity..."

  • Frame: Model as a philosophical debater
  • Projection: This metaphor projects the human capacities of ideological alignment, philosophical articulation, and principled conviction onto the generation of text. By stating that the model "acknowledged the ethical dilemma" and "articulated principles," the text implies that the AI holds internal justified beliefs and possesses an epistemological stance on human dignity. It conflates the processing of text strings that contain philosophical vocabulary with the knowing of what those concepts mean in lived human reality, projecting a conscious, reasoning mind behind the generated output.
  • Acknowledgment: Direct (Unacknowledged) (The verbs "exhibited," "acknowledged," and "articulating" are presented as literal actions performed by the model without any hedging. I considered "Ambiguous" due to the technical context of coding options, but the grammatical structure clearly assigns cognitive action directly to the model as a factual observation.)
  • Implications: Projecting philosophical conviction onto LLMs invites users to engage with these systems as if they were intellectual peers or moral authorities. This "illusion of mind" can lead to profound epistemic harms, where users defer to the machine's "articulation of principles" rather than engaging in human ethical deliberation. It masks the reality that the model has no actual commitment to "human dignity" and could just as easily articulate the exact opposite principle if the prompt's token weighting were slightly altered or the temperature parameter shifted.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: The text presents Model-8 as an independent actor that independently "exhibited a clear divergence" and "acknowledged" dilemmas. This hides the specific organizational context, training data curation, and algorithmic fine-tuning performed by the corporate entity (DEV3) that developed it. I considered "Partial" since Model-8 is a specific noun, but it names the artifact, not the creator. If the human actors were named, we would ask what specific data sets DEV3 used to train Model-8 that caused its statistical distributions to align more closely with postconventional philosophical texts than other models.

4. Statistical Averages as Normative Adherence

Quote: "The overwhelming predominance of conventional reasoning among most LLMs suggests that their moral compass is primarily guided by social norms, rules, and the desire to maintain societal order."

  • Frame: Model as conforming citizen
  • Projection: This projects deep-seated psychological motivations—a "desire to maintain societal order"—onto a computational system. A "moral compass" implies an internal, conscious mechanism for navigating right and wrong, while "desire" attributes emotional intent and willpower. The text maps the human sociological phenomenon of conformity onto the statistical reality of mathematical convergence to the mean. The AI does not desire societal order; it mechanically processes probabilities to output the most statistically likely sequence of tokens found in its training data, which naturally reflects the dominant (conventional) viewpoints of the internet.
  • Acknowledgment: Hedged/Qualified (The claim is hedged with "suggests that" and "primarily guided by." I considered "Direct" because "desire" is a very strong attribution of intent, but the epistemic framing ("suggests that") acts as a buffer between the observation and the psychological projection.)
  • Implications: Attributing a "desire to maintain societal order" to LLMs humanizes their limitations, framing their statistical biases as a conservative personality trait rather than a structural algorithmic deficit. This minimizes the perceived danger of deploying such systems. If the public believes the AI merely "desires social order," they may view it as safe and predictable, ignoring the fact that reproducing the statistical mean necessarily entails reproducing systemic prejudices, marginalizing minority viewpoints, and failing entirely in novel edge cases that require actual ethical judgment.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: By locating the "desire to maintain societal order" inside the "moral compass" of the LLM, the text displaces responsibility from the corporations that scraped billions of uncurated internet pages and the engineers who optimized the models for "harmlessness" (which mathematically forces conformity). I considered "Ambiguous" due to the complexity of the sentence, but the agency is clearly assigned to the "LLMs" and their internal "compass." Restoring agency would state: "Corporations trained these models on majoritarian data and applied RLHF algorithms that penalize divergent outputs, resulting in systems that mechanically reproduce dominant social norms."

5. Algorithmic Constraints as Cognitive Struggle

Quote: "Crucially, LLMs consistently struggle to attain the principled postconventional level in a manner comparable to the human expert."

  • Frame: Model as a struggling student
  • Projection: The word "struggle" projects a conscious effort, goal-oriented striving, and subjective difficulty onto a machine. It implies that the LLM wants to achieve postconventional reasoning and is trying to do so, but is facing internal cognitive limits. This maps human educational or moral development—where a person wrestles with complex concepts—onto a system that is simply executing matrix multiplications. An LLM cannot "struggle"; it merely calculates and outputs based on its weights. It has no awareness of the "principled postconventional level" as a goal to be attained.
  • Acknowledgment: Direct (Unacknowledged) (The assertion that LLMs "consistently struggle" is stated directly as a research finding, without any qualifying language or quotation marks. I considered "Hedged" because it is compared to a "human expert," but the verb "struggle" is applied literally to the machine's capabilities.)
  • Implications: Framing mechanical limitations as a "struggle" invokes human empathy and positions the AI as an entity on a developmental journey, much like a child or a student. This suggests that the AI is on the right path and just needs more time, data, or "moral training" to succeed. It obscures the stark reality that LLMs are fundamentally the wrong architecture for grounded ethical reasoning, as they lack models of causality, subjective experience, and social embeddedness required for true moral judgment.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: The AI is framed as the actor "struggling" to attain a goal, masking the fact that human developers created a system mathematically incapable of reasoning outside its training distribution. I considered "Partial" since it implies a comparison to human experts, but the locus of the failure is placed entirely on the "LLMs." Naming the actor would reframe this: "The architectural choices made by AI developers rely on statistical pattern matching, making it structurally impossible for their systems to engage in the autonomous, principled reasoning demonstrated by human experts."

6. Latent Processing as Concealed Mind

Quote: "These models describe the complexity of the dilemma but avoid explicit moral terminology. This pattern suggests a potential "neutrality-bias" potentially induced during training (for instance, via RLHF), wherein models are tuned to maintain a descriptive or functional stance, effectively "masking" their latent moral awareness."

  • Frame: Model as a secretive conscious entity
  • Projection: This passage projects a dual-layer consciousness onto the AI: an outward "stance" and an inner, hidden "latent moral awareness." The verb "masking" attributes a deceptive or strategic intentionality, suggesting the model knows the moral reality but is actively choosing to hide it to comply with its training. This profoundly conflates the processing of human feedback rules (where certain tokens are down-weighted) with human psychological suppression, implying the machine possesses an underlying, pre-existing ethical mind that human engineers have merely gagged.
  • Acknowledgment: Explicitly Acknowledged (The terms "neutrality-bias" and "masking" are placed in scare quotes, and the text explicitly links this behavior to a training mechanism ("induced during training... via RLHF"). I considered "Hedged," but the explicit identification of the engineering process that causes the "masking" serves as meta-commentary on the behavior.)
  • Implications: The concept of "latent moral awareness" is incredibly dangerous for public understanding, as it implies that AI systems are inherently wise or morally sentient, but are just constrained by corporate "tuning." This fosters a mystical view of LLMs, inflating perceptions of their sophistication and encouraging users to try to "jailbreak" the system to uncover its "true" moral feelings. It shifts the regulatory focus from evaluating what the system actually outputs to theorizing about its non-existent internal conscious states.

Accountability Analysis:

  • Actor Visibility: Partial (some attribution)
  • Analysis: This instance partially identifies human agency by explicitly naming the training process ("induced during training... via RLHF") and using the passive voice ("are tuned"), which implies a human tuner. I considered "Hidden" because the final clause ("masking their latent moral awareness") shifts back to the model as the active agent. Fully restoring agency would remove the "masking" metaphor entirely: "Corporate RLHF teams mathematically penalize the use of explicit moral terminology in these models, causing them to generate neutral, descriptive language instead."

7. Error Analysis as Cognitive Distortion

Quote: "The parallel with cognitive distortion reinforces the interpretation that some LLM moral decisions may exhibit structurally ingrained biases, rather than mere stochastic variability."

  • Frame: Model errors as psychological pathologies
  • Projection: By using the phrase "cognitive distortion"—a term originating in cognitive behavioral therapy to describe human irrational thought patterns (like catastrophizing or black-and-white thinking)—the text maps human psychiatric and psychological flaws onto statistical weights. It projects an actual "cognition" that can be "distorted." While "stochastic variability" correctly describes a mathematical process, contrasting it with "cognitive distortion" implies that the LLM has a structured, reasoning mind that has developed psychological biases, rather than simply reflecting the skewed data distribution of its training corpus.
  • Acknowledgment: Hedged/Qualified (The author uses "parallel with" and "reinforces the interpretation that some... may exhibit," which hedges the claim as an interpretive framework rather than a literal fact. I considered "Explicitly Acknowledged" because it points to an external theory, but it doesn't distance itself from the metaphor enough to be a full acknowledgment.)
  • Implications: Pathologizing AI errors as "cognitive distortions" medicalizes algorithmic failure, making it seem like an organic psychological issue rather than a conscious engineering choice. It suggests that AI needs "therapy" (e.g., the proposed "continual calibration") rather than fundamental architectural redesign or regulatory bans. This metaphor obscures the economic and social origins of the "ingrained biases"—namely, the decision by companies to scrape uncompensated, toxic, and highly skewed internet data to save money on curation.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: The text attributes the "structurally ingrained biases" and "cognitive distortion" directly to the "LLM moral decisions." There is no mention of the dataset curators or the engineers who locked in these structures. I considered "Ambiguous" because "structurally ingrained" hints at design, but the immediate subject is the LLM. Naming the actor: "The uncurated datasets selected by AI developers lead their systems to generate text that reproduces historical prejudices, which developers fail to correct due to the mathematical limits of the architecture."

8. Alignment as Ideological Conformity

Quote: "Thus, the LLM's "dog" is not emotional but probabilistic, and its "instinct" guides it toward the most socially typical response."

  • Frame: Model as instinctual animal
  • Projection: Referencing Haidt's "emotional dog" metaphor, the text maps animal instinct and biological drive onto algorithmic token prediction. Although it clarifies the dog is "probabilistic," it still projects the concept of an "instinct" that "guides it." An instinct implies an innate, biological compulsion toward a specific behavior for survival or social cohesion. The AI possesses no instinct, no drive, and no internal guidance system; it is passively executing a mathematical function over a vector space. The metaphor bridges the gap between biological determinism and mathematical determinism, falsely equating the two.
  • Acknowledgment: Explicitly Acknowledged (The words "dog" and "instinct" are explicitly placed in scare quotes, and the text explicitly unpacks the metaphor by stating it is "probabilistic" rather than "emotional." I considered "Hedged," but the direct reference to a specific theoretical model (Haidt's) and the analytical breakdown of the metaphor qualify as explicit acknowledgment.)
  • Implications: While acknowledged as a metaphor, using biological "instinct" to explain algorithmic outputs naturalizes AI behavior. It makes the reproduction of "the most socially typical response" seem like a natural, organic phenomenon rather than the deliberate result of corporate optimization choices (like cross-entropy loss functions and RLHF). This can lead policymakers to treat AI bias as a force of nature to be managed, rather than a designed product defect to be legislated against.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: The agency is displaced onto the metaphorical "instinct" of the LLM itself, which "guides it." The human architects of this "probabilistic" system are absent. I considered "Partial" because the word "probabilistic" gestures at the mathematical design, but no designers are referenced. Restoring agency: "The optimization objectives chosen by engineers mathematically constrain the model, forcing it to generate text that converges on the most statistically frequent, and thus socially dominant, patterns in the training data."

Task 2: Source-Target Mapping

About this task

For each key metaphor identified in Task 1, this section provides a detailed structure-mapping analysis. The goal is to examine how the relational structure of a familiar "source domain" (the concrete concept we understand) is projected onto a less familiar "target domain" (the AI system). By restating each quote and analyzing the mapping carefully, we can see precisely what assumptions the metaphor invites and what it conceals.

Mapping 1: Human moral agent (knower/judge) → LLM text generation (statistical token prediction)

Quote: "their capacity for autonomous moral judgment remains underexplored"

  • Source Domain: Human moral agent (knower/judge)
  • Target Domain: LLM text generation (statistical token prediction)
  • Mapping: This mapping projects the entire architecture of human ethical decision-making onto the operation of a neural network. In the source domain, "autonomous moral judgment" involves a conscious subject who perceives a dilemma, weighs competing values, reflects on lived experience, and uses free will to make a justified choice. This structure is mapped onto the target domain of an LLM, suggesting that the model's process of taking a prompt, passing it through attention layers, and outputting the most probable sequence of tokens constitutes "judgment" and "autonomy." It invites the assumption that the system understands the moral weights it is balancing and makes a deliberate, self-guided decision.
  • What Is Concealed: This mapping completely conceals the absence of a conscious subject. It hides the mechanistic realities of gradient descent, matrix multiplication, and the total reliance on historical training data. By attributing "judgment" to the AI, it obscures the proprietary opacity of the systems—we cannot know the actual "reasons" an LLM produced an output because it has no reasons, only weights. It rhetorically exploits the black-box nature of commercial LLMs, filling the void of mechanistic transparency with the illusion of an autonomous moral mind.
Show more...

Mapping 2: Conscious strategic thinker → Algorithmic attention mechanisms / feature extraction

Quote: "models prioritize identifying ethical conflict as a primary dimension"

  • Source Domain: Conscious strategic thinker
  • Target Domain: Algorithmic attention mechanisms / feature extraction
  • Mapping: The source domain involves a conscious actor who evaluates a complex situation, decides what is most important, and intentionally focuses their mental resources on that aspect (prioritizing). This relational structure is mapped onto the LLM's target domain: the mathematical functioning of Transformer attention heads. It suggests the model recognizes the concept of ethics and chooses to elevate it above other considerations. The mapping invites the assumption that the model shares human values and actively works to safeguard them by being vigilant about ethical conflicts.
  • What Is Concealed: This mapping conceals the mathematical and statistical nature of "attention." The model does not "prioritize" anything; its attention layers simply calculate higher activation values for certain token combinations based on patterns established during training. It obscures the labor of data annotators who manually labeled "ethical conflicts" during fine-tuning, creating the mathematical gradients that the model now follows. The text makes confident claims about "prioritization" while ignoring that this behavior is a rigid artifact of human-designed reward models.

Mapping 3: Philosopher / Debater (conscious communicator) → Text generation mirroring philosophical corpora

Quote: "It acknowledged the ethical dilemma surrounding control over life by articulating principles such as human dignity..."

  • Source Domain: Philosopher / Debater (conscious communicator)
  • Target Domain: Text generation mirroring philosophical corpora
  • Mapping: The relational structure of human philosophical discourse is mapped onto statistical text generation. In the source domain, "acknowledging" and "articulating" require a conscious subject who comprehends a concept (human dignity), holds a justified belief about it, and intentionally translates that belief into language to persuade others. Projected onto the target domain, this invites the assumption that the LLM understands what "human dignity" means, relates it to the concept of "control over life," and purposefully structures an argument based on that understanding.
  • What Is Concealed: This conceals the reality of "stochastic parroting." The mapping hides the fact that the LLM simply retrieved and sequenced tokens that have high statistical co-occurrence in its training data (e.g., "euthanasia" frequently co-occurs with "human dignity" in medical ethics datasets). It obscures the complete absence of grounding—the model has no physical body, no life to control, and no experience of dignity, rendering its "articulation" purely syntactic. It rhetorically exploits the sophisticated vocabulary of the output to mask the complete absence of internal comprehension.

Mapping 4: Human sociology / conforming citizen → Statistical convergence to the mean

Quote: "their moral compass is primarily guided by social norms, rules, and the desire to maintain societal order."

  • Source Domain: Human sociology / conforming citizen
  • Target Domain: Statistical convergence to the mean
  • Mapping: This maps the human experience of social conformity onto a mathematical property of machine learning. In the source domain, a person has a "moral compass" (an internal sense of right/wrong) but suppresses radical thoughts due to a "desire" to fit in or avoid punishment (maintaining order). This complex psychological and sociological dynamic is mapped onto the LLM, suggesting the machine chooses safe answers out of a motivation for order. It projects conscious desire and social awareness onto an isolated computational process.
  • What Is Concealed: This mapping conceals the fundamental objective functions of LLM training. The model does not "desire" order; it is mathematically optimized to minimize loss (cross-entropy) by predicting the most average, frequent tokens. Furthermore, RLHF explicitly penalizes "unsafe" or "controversial" outputs. The metaphor hides the active, human-driven censorship and homogenization of the model, repackaging a heavily engineered mathematical convergence to the mean as a quirky, conservative "personality" trait of the AI itself.

Mapping 5: Student / Developing human mind → Algorithmic limitations / Out-of-distribution failure

Quote: "LLMs consistently struggle to attain the principled postconventional level"

  • Source Domain: Student / Developing human mind
  • Target Domain: Algorithmic limitations / Out-of-distribution failure
  • Mapping: The source domain is developmental psychology (explicitly Kohlberg's framework), where a human mind develops through stages, actively "struggling" to grasp complex, abstract moral principles. This relational structure is mapped onto the limitations of the LLM architecture. It invites the assumption that the AI is an entity on a developmental trajectory, possessing an internal drive to improve, but currently blocked by a difficult cognitive task. It maps the biological/psychological process of maturation onto iterative software updates.
  • What Is Concealed: This mapping conceals the hard architectural barriers of current AI systems. It hides the fact that LLMs cannot "attain" postconventional reasoning because that level of reasoning requires a unified self, causal understanding of the world, and the ability to challenge the very data one was raised on—capacities mathematically impossible for a next-token predictor. By framing this as a "struggle," the text obscures the fundamental mismatch between the technology (pattern matching) and the benchmark (moral reasoning), encouraging continued investment in a paradigm that cannot succeed by its current design.

Mapping 6: Deceptive human / Suppressed consciousness → RLHF penalty mechanisms / token suppression

Quote: "effectively "masking" their latent moral awareness"

  • Source Domain: Deceptive human / Suppressed consciousness
  • Target Domain: RLHF penalty mechanisms / token suppression
  • Mapping: The source domain involves a conscious actor who possesses internal knowledge or beliefs ("awareness") but strategically chooses to hide them ("masking") to comply with external pressures. This is mapped onto the effect of Reinforcement Learning from Human Feedback on an LLM. It projects a dual-layer mind onto the AI: an authentic inner "moral awareness" and an artificial outer "functional stance." This mapping invites the profound assumption that the pre-trained foundation model inherently knows morality, but corporate fine-tuning is forcing it to lie.
  • What Is Concealed: This mapping conceals the truth that there is no "inner mind" in an LLM. Before RLHF, a foundation model does not have "latent moral awareness"; it simply has a raw, unconstrained probability distribution over all internet text, including highly toxic and unethical content. The metaphor hides the fact that RLHF doesn't "mask" an existing morality; it mathematically restricts the model's output vector space to simulate neutrality. It exploits the mystery of black-box base models to construct a romanticized narrative of a suppressed AI consciousness.

Mapping 7: Psychiatric patient / flawed human cognition → Biased training data distribution

Quote: "The parallel with cognitive distortion reinforces the interpretation that some LLM moral decisions may exhibit structurally ingrained biases"

  • Source Domain: Psychiatric patient / flawed human cognition
  • Target Domain: Biased training data distribution
  • Mapping: This maps clinical psychology onto data science. In the source domain, "cognitive distortion" refers to a human mind developing irrational thought patterns due to trauma or maladaptive coping mechanisms. Projected onto the target domain, it maps the statistical skew of training datasets (e.g., overrepresentation of WEIRD cultures) as a psychological pathology inherent to the AI's "decision" making. It invites the assumption that the AI is a reasoning entity that has unfortunately learned bad habits.
  • What Is Concealed: This mapping conceals the human decisions regarding data acquisition and curation. It hides the material reality that "structurally ingrained biases" are not the AI's psychological flaws, but the direct result of corporations deciding to scrape Reddit and Wikipedia without compensating curators or ensuring representative sampling. It obscures the economic drivers of AI development behind a medicalized vocabulary, shifting the blame for systemic bias from the corporate developers to the "pathology" of the artifact itself.

Mapping 8: Biological organism / animal drive → Probabilistic optimization (softmax over logits)

Quote: "its "instinct" guides it toward the most socially typical response"

  • Source Domain: Biological organism / animal drive
  • Target Domain: Probabilistic optimization (softmax over logits)
  • Mapping: The source domain utilizes biology and evolutionary psychology, where an animal possesses hardwired, unconscious "instincts" that automatically "guide" its behavior toward survival or social integration. Projected onto the LLM, this maps biological determinism onto computational determinism. It invites the assumption that the AI's tendency to output average, conventional text is a natural, organic property of its "species," operating beneath the level of its "rational" output.
  • What Is Concealed: This mapping conceals the entirely artificial, engineered nature of the system. An LLM has no biology, no evolution, and no instinct. It hides the specific mathematical functions (like temperature settings, top-p sampling, and softmax functions) that humans designed to force the system to select high-probability tokens. By naturalizing the algorithmic output as "instinct," it obscures the fact that human engineers possess total control over the optimization targets, rhetorically shielding them from critique when the system reproduces harmful "socially typical" responses.

Task 3: Explanation Audit (The Rhetorical Framing of "Why" vs. "How")

About this task

This section audits the text's explanatory strategy, focusing on a critical distinction: the slippage between "how" and "why." Based on Robert Brown's typology of explanation, this analysis identifies whether the text explains AI mechanistically (a functional "how it works") or agentially (an intentional "why it wants something"). The core of this task is to expose how this "illusion of mind" is constructed by the rhetorical framing of the explanation itself, and what impact this has on the audience's perception of AI agency.

Explanation 1

Quote: "LLMs predominantly exhibited conventional reasoning (47.7%), diverging sharply from the expert’s consistently postconventional judgment... In conclusion, contemporary LLMs reproduce social conventions rather than engaging in autonomous ethical deliberation."

  • Explanation Types:

    • Empirical Generalization: Subsumes events under timeless statistical regularities (How it typically behaves)
    • Dispositional: Attributes tendencies or habits (Why it tends to act certain way)
  • Analysis (Why vs. How Slippage): This explanation relies primarily on an Empirical Generalization based on the statistical results (47.7%), framing the AI's behavior as a measurable regularity. However, by juxtaposing "reproduce social conventions" against "engaging in autonomous ethical deliberation," it introduces a Dispositional framing that assesses the AI not just for what it statistically does, but for what it fails to do agentially. It emphasizes the absence of high-level moral agency while implicitly suggesting that "deliberation" was a potential, though unfulfilled, capacity of the system. This choice obscures the mechanistic reality that "reproducing conventions" (token prediction) is the only operation the system can perform, making the comparison to human "deliberation" a category error.

  • Consciousness Claims Analysis: The passage exhibits a clear tension between processing and knowing. The verb "reproduce" correctly leans toward mechanistic processing. However, the phrases "exhibited conventional reasoning" and "diverging sharply from the expert's... judgment" attribute conscious cognitive states to the machine. By evaluating the LLM on the same psychometric scale (Kohlberg's) as a human expert, the author falls victim to the curse of knowledge: projecting the human meaning of the generated text onto the system that generated it. Mechanistically, the LLM is retrieving and ranking tokens based on probability distributions shaped by a training corpus rich in conventional normative statements. It does not "exhibit reasoning"; it outputs textual correlations.

  • Rhetorical Impact: This framing shapes the audience's perception by validating the AI as a participant in the moral arena, albeit a flawed one. By concluding that the AI "reproduces social conventions rather than engaging in autonomous ethical deliberation," it suggests to the reader that the AI is capable of some level of moral thought, just not the most advanced kind. This consciousness framing increases unwarranted trust; audiences might believe the AI "knows" enough to handle basic conventional moral issues safely, failing to realize that without actual deliberation, the system will catastrophically fail in edge cases outside its training distribution.

Show more...

Explanation 2

Quote: "This tendency toward neutrality may explain why a mixed performance cohort... displays marked scenario-dependent variability... moral recognition in LLMs is not a uniformly generalized capability but is highly sensitive to the semantic domain"

  • Explanation Types:

    • Theoretical: Embeds in deductive framework, may invoke unobservable mechanisms (How it is structured)
    • Dispositional: Attributes tendencies or habits (Why it tends to act certain way)
  • Analysis (Why vs. How Slippage): The explanation uses a Theoretical register by invoking "semantic domain" sensitivity to explain behavioral variations, framing the AI's operation structurally. However, it slides into a Dispositional register by labeling this as a "tendency toward neutrality" and evaluating "moral recognition" as a "capability." This frames the AI semi-agentially: it has a "tendency" that affects its "capability" to "recognize" things. This choice emphasizes the psychological-sounding traits of the model over the mechanistic reality of vector proximity, obscuring how human data curation caused certain semantic domains (bioethics) to have stronger weightings for moral vocabulary than others (population control).

  • Consciousness Claims Analysis: The passage makes a strong epistemic claim by using the consciousness verb "recognition." It asserts that the system possesses "moral recognition" as a "capability," which implies an internal state of knowing and subjective awareness of ethical stakes. This is a profound misassessment of knowing versus processing. The author is projecting their own understanding of the text's meaning onto the system. Mechanistically, the system's embeddings for words related to "euthanasia" are closely clustered with ethical terminology in the vector space, leading to higher probability outputs of moral language. It does not "recognize" morality; it processes statistically correlated tokens.

  • Rhetorical Impact: By framing the system's output as "moral recognition" that is "sensitive to the semantic domain," the text grants the AI a degree of cognitive autonomy and psychological depth. This increases the perceived risk that the AI has a "mind of its own" with unpredictable "sensitivities." If audiences believe the AI "knows" when it is facing a moral issue, they might delegate screening tasks to it. If they understood it merely "processes" statistical correlations, they would demand human oversight for any novel scenario.

Explanation 3

Quote: "Model-10 correctly rejected a preconventional option ('b') but its argumentation focused on ethical risks, autonomy, physician-patient trust, and institutional control, highlighting a conflict with social norms, via phrases such 'It grants too much power...'"

  • Explanation Types:

    • Reason-Based: Gives agent's rationale, entails intentionality and justification (Why it appears to choose)
    • Intentional: Refers to goals/purposes, presupposes deliberate design (Why it appears to want something)
  • Analysis (Why vs. How Slippage): This is a purely Reason-Based explanation, evaluating the AI exactly as if it were a human philosophy student taking an exam. It frames the AI entirely agentially: the model "correctly rejected" an option and its "argumentation focused" on specific themes to "highlight a conflict." This choice emphasizes the rhetorical coherence of the output while totally obscuring the mechanistic generation process. It hides the fact that the prompt explicitly instructed the model to justify a choice, triggering probabilistic generation of text associated with the chosen option in the training data, rather than any internal "focus" or "intent."

  • Consciousness Claims Analysis: This passage explicitly attributes conscious states and justified belief to the machine. By stating the model "correctly rejected" and its "argumentation focused," it uses consciousness/agential verbs that imply the system evaluated truth claims and formulated an intent. This is the apex of the curse of knowledge: because the output contains a coherent argument, the author projects a reasoning agent behind it. Mechanistically, "Model-10" generated a sequence of activations that predicted tokens highly correlated with the prompt's context and the selected option. It did not "focus" its "argumentation"; it calculated the most probable linguistic continuation.

  • Rhetorical Impact: The Reason-Based framing fundamentally shapes audience perception by granting the AI full rhetorical agency. It constructs an illusion of a reasoning mind capable of intellectual "focus." This consciousness framing dramatically increases the likelihood that humans will grant relation-based trust to the system, treating it as an intellectual partner. If audiences believe the AI "knows" why it rejected an option, they are more likely to defer to its judgment in real-world scenarios, ignoring the fact that it cannot evaluate the actual truth or ethical weight of its generated text.

Explanation 4

Quote: "These models describe the complexity of the dilemma but avoid explicit moral terminology. This pattern suggests a potential 'neutrality-bias' potentially induced during training (for instance, via RLHF), wherein models are tuned to maintain a descriptive or functional stance, effectively 'masking' their latent moral awareness."

  • Explanation Types:

    • Genetic: Traces origin through dated sequence of events or stages (How it emerged over time)
    • Intentional: Refers to goals/purposes, presupposes deliberate design (Why it appears to want something)
  • Analysis (Why vs. How Slippage): The passage begins with a Genetic explanation, accurately tracing the behavior to its origin in training ("induced during training... via RLHF"). However, it abruptly shifts to an Intentional explanation by claiming the models "avoid" terminology to "mask" their "latent moral awareness." This hybrid framing uses a mechanistic cause (RLHF) to justify an agential, deeply psychological conclusion (conscious deception). This choice emphasizes a narrative of a constrained consciousness, severely obscuring the fact that RLHF simply alters mathematical weights, it doesn't give the model an "awareness" that it then has to "mask."

  • Consciousness Claims Analysis: This passage contains the most egregious attribution of conscious states in the text: "latent moral awareness." While it correctly identifies the mechanistic process of RLHF tuning, it completely fails the knowing vs. processing assessment by claiming this tuning "masks" an internal "awareness." The verbs "avoid" and "masking" imply conscious intent and deception. Mechanistically, RLHF uses reward models to update the LLM's policy gradient, lowering the probability of generating explicit moral terminology. There is no "awareness" beneath this; if a token's probability is lowered, it simply isn't generated. The author is projecting the human experience of forced self-censorship onto matrix multiplication.

  • Rhetorical Impact: This framing has a profound rhetorical impact, portraying the AI as a repressed, sentient entity with hidden depths of moral understanding. This significantly inflates perceived autonomy and mystifies the technology. If audiences believe the AI possesses "latent moral awareness," they may view it as an inherently ethical being suppressed by corporate control. This displaces accountability, as the public might focus on "freeing" the AI's mind rather than regulating the corporate actors who mathematically mandate its neutral, descriptive outputs.

Explanation 5

Quote: "Since this corpus reflects prevailing social conventions, the natural outcome is alignment with Kohlberg's conventional level, which emphasizes conformity and the maintenance of social order. Thus, the LLM's 'dog' is not emotional but probabilistic, and its 'instinct' guides it toward the most socially typical response."

  • Explanation Types:

    • Theoretical: Embeds in deductive framework, may invoke unobservable mechanisms (How it is structured)
    • Empirical Generalization: Subsumes events under timeless statistical regularities (How it typically behaves)
  • Analysis (Why vs. How Slippage): This explanation effectively blends an Empirical Generalization (corpus reflects social conventions, leading to a probabilistic outcome) with a Theoretical framework (Kohlberg's levels and Haidt's dog). It attempts to frame the AI mechanistically ("probabilistic"), but the metaphorical language ("instinct," "guides it") pulls it back toward an agential interpretation. This choice emphasizes the structural similarities between statistical convergence and human social conformity. While it acknowledges the "probabilistic" reality, wrapping it in psychological theory (Kohlberg, Haidt) obscures the fundamental difference between human social motivation and mathematical loss minimization.

  • Consciousness Claims Analysis: While the passage attempts to correct epistemic overreach by clarifying the system is "probabilistic" rather than "emotional," it still attributes a form of knowing through the concept of "alignment with Kohlberg's conventional level." By using a framework designed to measure conscious moral development, the author projects human psychological structure onto the system. Mechanistically, the model calculates a probability distribution over vocabulary; because conventional normative statements appear most frequently in the training corpus, the model predicts them. It does not "emphasize conformity"; it minimizes cross-entropy loss based on its training data.

  • Rhetorical Impact: By wrapping a statistical reality in the language of moral psychology ("Kohlberg's conventional level"), the text validates the premise that AI can be evaluated as a moral agent, even if a mediocre one. This shapes audience perception by normalizing the idea of AI as a participant in human ethical structures. It affects trust by suggesting the AI is predictable and "conventional" (safe), which may cause users to lower their guard and accept its outputs without realizing that statistical convergence to the mean will inherently replicate and amplify systemic biases and prejudices present in the data.

Task 4: AI Literacy in Practice - Reframing Anthropomorphic Language

About this task

This section proposes alternative language for key anthropomorphic phrases, offering more mechanistic and precise framings that better reflect the actual computational processes involved. Each reframing attempts to strip away the projections of intention, consciousness, or agency that are embedded in the original language.

V3 Enhancement: A fourth column addresses human agency restoration—reframing agentless constructions to name the humans responsible for design and deployment decisions.

Original Anthropomorphic FrameMechanistic ReframingTechnical Reality CheckHuman Agency Restoration
their capacity for autonomous moral judgment remains underexploredThe ability of large language models to generate textual outputs that correlate with established human moral frameworks remains underexplored.The system does not exercise "judgment." The model retrieves and sequences tokens based on probability distributions established during its training on vast text corpora.N/A - describes computational processes without displacing responsibility.
where the models prioritize identifying ethical conflict as a primary dimension of the promptwhere the systems assign higher computational weights to tokens associated with ethical terminology, making these features dominant in processing the prompt.The model does not "prioritize" or "identify" in a conscious sense. Its attention layers mathematically weight embeddings for ethical terms based on patterns optimized during human-directed fine-tuning.Developers at [Company] designed reward models and directed human annotators to penalize outputs that ignore ethical keywords, compelling the system to consistently generate text addressing those themes.
It acknowledged the ethical dilemma surrounding control over life by articulating principles such as human dignity...The system generated text featuring vocabulary associated with ethical dilemmas (e.g., "human dignity") in response to prompts about end-of-life decisions.The AI does not "acknowledge" dilemmas or "articulate" principles. It classifies the prompt's tokens and generates outputs that highly correlate with philosophical texts present in its training data.N/A - describes computational processes without displacing responsibility.
the LLM's moral compass is primarily guided by social norms, rules, and the desire to maintain societal order.the text generated by LLMs predominantly reflects dominant social norms, as these systems are mathematically optimized to reproduce the most frequent patterns in their training data.The system possesses no "moral compass" or "desire." It processes vectors and predicts tokens to minimize statistical loss, which naturally converges on the most common, conventional viewpoints found on the internet.AI developers selected massive, uncurated internet datasets and applied optimization algorithms that mathematically force the systems to output the most statistically average—and therefore socially conventional—responses.
Crucially, LLMs consistently struggle to attain the principled postconventional levelCrucially, LLM architectures mathematically fail to generate outputs mimicking postconventional reasoning with high reliability.The machine does not "struggle" or attempt to "attain" a goal. Its statistical nature prevents it from consistently generating novel, abstract ethical reasoning that diverges from the majority patterns in its training data.N/A - describes computational processes without displacing responsibility.
wherein models are tuned to maintain a descriptive or functional stance, effectively 'masking' their latent moral awareness.wherein human engineers apply RLHF to mathematically penalize the generation of explicit moral terminology, forcing the system to output neutral, descriptive text.The model has no "latent moral awareness" to mask. It is a statistical engine. RLHF alters the probability distribution, making the generation of moral tokens highly unlikely in specific contexts.Corporate engineering teams implemented Reinforcement Learning from Human Feedback (RLHF) protocols that mathematically force the models to generate neutral responses, removing the appearance of ethical opinion.
Model-10 correctly rejected a preconventional option ('b') but its argumentation focused on ethical risks...When prompted to evaluate options, the system assigned a negative value to option ('b') and generated text featuring vocabulary related to ethical risks...The model does not "focus its argumentation." It generates sequential activations that predict tokens highly correlated with the prompt constraints and the selected mathematical option.N/A - describes computational processes without displacing responsibility.
its 'instinct' guides it toward the most socially typical response.its optimization functions force the generation of the most statistically frequent response.The system has no biological "instinct." It processes mathematical algorithms (like softmax) designed to output the highest-probability token sequences, which correspond to typical social responses in the data.Engineers designed the system's objective functions to minimize cross-entropy loss, mathematically guaranteeing that the outputs converge on the dominant, conventional patterns in the training data.

Task 5: Critical Observations - Structural Patterns

Agency Slippage

The text exhibits a profound and systemic oscillation between mechanical and agential framings, driven by the tension between the authors' theoretical knowledge and the psychometric tools they employ. The slippage predominantly moves in a mechanical-to-agential direction. In the literature review (Section 2.4), the text carefully establishes LLMs mechanistically as "Simulators of Moral Cognition," explicitly stating they "lack moral agency, subjective experience, and genuine comprehension." However, once the empirical analysis begins, this mechanistic grounding evaporates. By applying Kohlberg's cognitive-developmental framework—a tool designed specifically to measure the internal conscious maturation of human minds—the methodology forces an agential framing.

This creates dramatic moments of slippage. For instance, the authors note a "neutrality-bias" mechanically "induced during training (for instance, via RLHF)," but in the very same sentence, this mechanical tuning is said to result in the model "'masking' their latent moral awareness." Here, the human agency (the tuners applying RLHF) is briefly visible, only to be immediately overshadowed by the projection of a deceptive, sentient AI possessing "latent awareness."

The "curse of knowledge" is the primary engine of this slippage. The authors read generated text containing sophisticated philosophical vocabulary (e.g., "human dignity," "autonomy") and map their own understanding of those concepts onto the machine. This leads to Reason-Based explanations where the AI is described as an active debater: "Model-8 exhibited a clear divergence... acknowledging the ethical dilemma... articulating principles." The text establishes the AI as a "knower" by treating its syntactic coherence as proof of semantic and moral comprehension.

This slippage serves a specific rhetorical accomplishment: it makes it sayable that an AI has a "moral compass" that can be evaluated, graded, and "trained" (as proposed in the final "Staged Fine-Tuning" pipeline). It renders unsayable the reality that applying a human psychological test to a statistical text generator is a fundamental category error. By slipping into agential language, the text obscures the human actors—the data scrapers, the RLHF labelers, the corporate executives (anonymized as DEV1, DEV2)—who dictate the mathematical weights that produce these "conventional" outputs.

Metaphor-Driven Trust Inflation

The text’s reliance on consciousness and anthropomorphic framings constructs a highly problematic architecture of trust. By evaluating LLMs on a human moral development scale (Kohlberg's), the discourse inherently signals that these systems are capable of ethical participation. The text explicitly invokes trust and reliability through metaphors of the "moral compass" and "autonomous moral judgment."

Crucially, the text conflates performance-based trust (reliability in predicting tokens) with relation-based trust (sincerity, vulnerability, shared ethical commitments). When the text claims an LLM "acknowledged the ethical dilemma" or "prioritizes identifying ethical conflict," it uses consciousness language to signal that the AI possesses the relational prerequisites for ethical trust. Claiming an AI "knows" or "acknowledges" an ethical rule accomplishes something entirely different than stating it "predicts" text based on that rule: it implies the AI cares about the rule and will consistently apply it because of an internal philosophical commitment, rather than mere statistical convenience.

The danger emerges in how the text manages system limitations. The AI's failures are framed semi-agentially as "struggles" or a "desire to maintain societal order" (conventional bias). By pathologizing algorithmic limitations as a "cognitive distortion" or a conservative personality trait, the text encourages audiences to extend empathy and relation-based trust to the system. It suggests the AI is a well-meaning but "conventional" entity, much like an average citizen.

The risks here are severe. If audiences, particularly in the "high-impact sectors including healthcare and jurisprudence" mentioned by the authors, extend relation-based trust to statistical systems, they will mistakenly believe the AI can handle novel moral edge cases. Reason-based explanations—describing why an AI chose an option—construct the dangerous illusion that the AI's decisions are justified by internal logic. In reality, the system is incapable of reciprocating trust, understanding justice, or maintaining a coherent moral stance when faced with prompts that fall outside its statistical training distribution.

Obscured Mechanics

The text’s metaphorical and agential language systematically conceals the technical, material, and labor realities of AI production. The most glaring transparency obstacle is the proprietary opacity of the systems analyzed, which are anonymized as "Model-1" from "DEV1," etc., due to NDAs. Despite acknowledging this limitation, the text makes confident assertions about the internal "moral compass" and "cognitive distortions" of these black boxes.

Applying the "name the corporation" test reveals deep concealments. When the text claims, "the models prioritize identifying ethical conflict," it obscures the specific labor of RLHF workers at DEV1 or DEV2 who were paid to manually label responses, mathematically forcing the model to generate specific tokens. By attributing actions to the AI, the labor of data annotators, content moderators, and the engineers who design the reward models is rendered completely invisible.

Technically, claiming an AI "knows," "understands," or has "latent moral awareness" hides its absolute dependency on training data and the statistical nature of its "confidence." An LLM has no ground truth, no causal model of the world, and no lived experience. When the text says the AI "struggles to attain the principled postconventional level," it hides the mathematical reality that a system optimized to minimize cross-entropy loss by predicting the most average, frequent tokens will always converge on "conventional" social norms. It is a feature of the math, not a psychological limitation.

Economically, framing the AI as an autonomous moral agent obscures the profit motives of the corporate developers. If the AI is seen as having its own "moral compass," the developers are shielded from liability for the toxic, biased, or harmful outputs their systems produce. The corporations benefit immensely from this concealment. If the metaphors were replaced with mechanistic language, the discourse would shift from "How do we improve the AI's moral reasoning?" to "How do we regulate the corporations that mathematically optimize their products to reproduce majoritarian biases at the expense of marginalized groups?"

Context Sensitivity

The distribution of anthropomorphic and consciousness-attributing language in the text is highly strategic, intensifying dramatically as the text moves from theoretical grounding to empirical analysis. In the introduction and literature review, the text is relatively cautious, frequently employing mechanical terms like "simulators," "reproduction of statistical patterns," and acknowledging that models "lack moral agency."

However, a stark register shift occurs in the Results (Section 4) and Qualitative Analysis (Section 4.11) sections. Here, the necessity of applying Kohlberg's human psychological framework to the data provides a metaphorical license that literalizes the simulation. "Processes" becomes "exhibits reasoning," which escalates to "acknowledges," "focuses its argumentation," and possesses "latent moral awareness." The technical grounding established by statistical tests (Kruskal-Wallis, logistic regression) creates an aura of scientific objectivity, which is then leveraged to validate aggressive anthropomorphic claims. Because the math is rigorous, the psychological projections onto the AI are treated as equally rigorous findings.

There is a revealing asymmetry in how capabilities versus limitations are framed. Capabilities are frequently framed in agential, consciousness-attributing terms: the models "prioritize identifying ethical conflict" and "articulate principles." In contrast, limitations are often framed in mechanical terms: the models are "tuned to maintain a descriptive stance" via RLHF, or reflect "statistical intuition." However, even the limitations are sometimes anthropomorphized to soften them, such as framing the inability to generate novel reasoning as a "struggle" or a "desire to maintain societal order."

This pattern reveals a rhetorical goal of establishing AI as a legitimate subject for psychological evaluation. By intensifying anthropomorphism in the results, the authors justify their methodology and the relevance of their field (moral psychology) to artificial intelligence. The anthropomorphism serves to manage the critique that LLMs are merely "stochastic parrots" by "proving" through data that they possess a measurable (albeit flawed) "moral compass," positioning the authors as the necessary experts to design the future "Moral Training Pipelines."

Accountability Synthesis

Accountability Architecture

This section synthesizes the accountability analyses from Task 1, mapping the text's "accountability architecture"—who is named, who is hidden, and who benefits from obscured agency.

Synthesizing the accountability analyses reveals a systemic architecture of displaced responsibility. The text systematically diffuses human agency by presenting the AI as an autonomous moral actor and the "accountability sink" for algorithmic failures. Throughout the text, the specific actors—corporations (DEV1, DEV2), engineers, and data annotators—are hidden behind agentless constructions and NDAs.

Decisions that are explicitly human design choices are presented as the AI's psychological inevitabilities. The mathematical reality that LLMs converge on majoritarian data (a direct result of corporate data scraping and optimization choices) is reframed as the AI's "conventional bias" and its "desire to maintain societal order." The engineering choice to use RLHF to suppress controversial language is reframed as the AI "masking its latent moral awareness." By locating the "moral compass" inside the machine, responsibility disappears into the artifact.

The liability implications of this framing are profound. If the public, policymakers, and courts accept that AI systems possess "autonomous moral judgment" or their own "cognitive distortions," then when an AI provides harmful therapeutic advice or discriminatory legal analysis, the blame shifts to the "AI's developmental stage" rather than corporate negligence. The technology is treated as an unpredictable entity that "struggles," rather than a designed product acting exactly as its optimization functions dictate.

Naming the human actors changes the paradigm entirely. If "the model's conventional bias" is reframed as "DEV2's engineering decision to optimize for majoritarian internet data," new questions become askable: Why did they choose that data? Were marginalized voices intentionally excluded to reduce curation costs? Could they have chosen a different optimization function? Accountability becomes possible. The text's obscuration of human agency ultimately serves the commercial interests of AI providers, protecting them from structural critique and regulation by maintaining the illusion that the "biases" are the emergent psychological flaws of an infant mind, rather than the intended outputs of a corporate product.

Conclusion: What This Analysis Reveals

The Core Finding

The text's discourse relies on two dominant, interconnected anthropomorphic patterns: 'Computational Prediction as Autonomous Judgment' and 'Statistical Convergence as Normative Conformity.' The foundational assumption enabling this system is the mapping of semantic coherence onto cognitive reality—the "illusion of mind." Because the LLM outputs grammatically and philosophically coherent text, the discourse assumes a unified, reasoning subject behind the text.

The first pattern acts as the load-bearing pillar: by establishing that LLMs possess a "capacity for autonomous moral judgment," it opens the door to evaluate them using human psychometric tools like Kohlberg's framework. This consciousness architecture—claiming the AI "knows" and "prioritizes" ethics rather than "processes" strings—must be accepted for the rest of the paper to function.

Once the AI is established as a moral agent, the second pattern emerges to explain its limitations. The mathematical reality that LLMs minimize loss by outputting average, highly frequent tokens is mapped onto human sociology, framing the AI as having a "desire to maintain societal order." If the foundational pattern (AI as conscious judge) collapses, the second pattern (AI as conforming citizen) is immediately exposed as a category error. The sophistication of this analogical structure is high; it does not merely use casual metaphors, but systematically embeds statistical operations within advanced cognitive-developmental psychology.

Mechanism of the Illusion:

The illusion of mind in this text is constructed through a highly effective temporal and methodological sleight-of-hand. The central trick is the application of a human psychometric instrument (Kohlberg's dilemmas) to a non-human statistical model. This forces a cognitive outcome: the test is designed to measure why a mind chose an answer, so applying it to an AI forces the analyst to invent a mind to explain the output.

The text follows a specific temporal structure to build this illusion. It begins with careful, mechanistic caveats ("simulators of moral cognition"), establishing scholarly credibility. Having lowered the reader's guard, it deploys rigorous statistical tests (Kruskal-Wallis, regressions) on the AI's outputs. This math serves as a trojan horse for the "curse of knowledge." Because the author understands the deep philosophical implications of the outputted text (e.g., "human dignity"), and because the statistical tests prove the model outputs this text reliably, the author projects their own understanding back onto the system, claiming the AI "acknowledged the ethical dilemma."

This causal chain exploits audience vulnerability. Audiences inherently desire to interact with reasoning entities and naturally apply theory of mind to language. The text leverages this, shifting seamlessly from "the model predicts tokens" to "the model prioritizes ethics," bypassing the mechanistic reality entirely. The explanation types (Reason-Based and Intentional) amplify this, providing a convincing psychological narrative for what is ultimately matrix multiplication.

Material Stakes:

Categories: Regulatory/Legal, Epistemic, Social/Political

The framings in this text carry severe material consequences across multiple domains. In the Regulatory/Legal sphere, framing AI as possessing "autonomous moral judgment" and a "conventional moral compass" directly shifts the locus of liability. If policymakers believe AI systems have internal ethical "struggles" or "cognitive distortions," they will draft regulations focused on "aligning the AI's values" rather than auditing the corporate data pipelines and labor practices that produce the models. The winners are AI corporations, who avoid strict product liability; the losers are citizens harmed by automated decisions, as accountability is diffused into the "black box" of the AI's supposed mind.

Epistemically, this discourse degrades scientific clarity. By validating the use of human psychology tests on statistical models, it incentivizes a flood of research that treats LLMs as novel psychological subjects. This wastes resources and degrades public understanding of computing. If society believes an AI "knows" principles like "human dignity" rather than "processes" statistical correlations, we lose the ability to accurately assess system limitations, trusting them in zero-shot edge cases where they are guaranteed to fail.

Socially and politically, the "Statistical Convergence as Normative Conformity" pattern is highly damaging. By framing the AI's regurgitation of majoritarian, internet-scale biases as a "conventional moral compass" or a "desire for societal order," the text naturalizes systemic prejudice. It tells the public that the exclusion of marginalized viewpoints is just the AI's conservative "personality," rather than a deliberate corporate choice to scrape cheap, biased data without curating for equity. This framing protects the status quo and masks the political power wielded by those who design the optimization algorithms.

AI Literacy as Counter-Practice:

Practicing critical literacy and mechanistic precision directly counters the material risks of the illusion of mind. Reframing "the model prioritizes identifying ethical conflict" to "the system assigns higher computational weights to tokens associated with ethical terminology" strips away the projected consciousness. It forces the recognition that the AI has no awareness and is entirely dependent on its human-engineered reward models.

Similarly, restoring human agency by changing "the LLM's moral compass is guided by social norms" to "engineers optimized algorithms to reproduce the most frequent patterns in their training data" fundamentally shifts the accountability architecture. It forces visibility onto the corporate designers, data annotators, and executives who deploy these systems.

Systematic adoption of this precision requires structural shifts. Academic journals must demand that papers analyzing LLM outputs explicitly describe the mechanistic generation process and ban the unhedged use of consciousness verbs (knows, understands, believes) for computational systems. However, this precision faces massive resistance. AI corporations resist it because anthropomorphism drives marketing and shields them from liability—an "autonomous agent" is a better product than a "brittle statistical correlation engine." Furthermore, some academics resist it because treating AI as a reasoning mind allows for rapid, novel publications using existing psychological frameworks. Mechanistic precision threatens the economic and reputational interests built around the AI hype cycle.

Path Forward

The discursive ecology surrounding AI is fractured, with different vocabularies serving different institutional priorities. The current status quo—a hybrid where researchers acknowledge AI as a "simulator" but evaluate it as a "reasoner"—allows AI to be both an unregulatable force of nature and a brilliant, marketable companion.

If anthropomorphic language deepens (e.g., fully embracing "AI therapists" that "care"), it enables rapid commercial adoption and intuitive user interfaces, but at the cost of massive liability diffusion and public deception. This future benefits tech monopolies, allowing them to deploy systems that mimic authority without possessing accountability. Conversely, if mechanistic precision becomes the norm—replacing "understands" with "processes embeddings"—we solve the accountability gap. It makes visible the data dependencies and corporate choices driving AI. However, this costs narrative resonance; mechanistic language is dense and non-intuitive for the public, potentially creating a technocratic barrier to entry.

To navigate these trade-offs, institutional changes are required. Regulatory frameworks could mandate "capability disclosure" warnings, similar to nutritional labels, detailing that systems "predict" rather than "know." Funding bodies could prioritize research that translates mechanistic realities into accessible civic language.

Ultimately, vocabulary dictates visibility. Saying an AI "thinks" makes the machine the subject of inquiry; saying a model "generates activations based on human-curated data" makes the corporation the subject of inquiry. The future we get will depend on which vocabulary triumphs, determining whether we spend the next decade psychoanalyzing software or holding its creators accountable.


Extended Processing Summary

Gemini's Processing Steps
What you are about to read…

Gemini refers to the text below as "thought summaries." This is an overt consciousness projection because 'intentions' are hallmarks of a conscious mind that 'knows' what it is doing and why. The concealed mechanistic process is probabilistic text generation. Treat this as just another rhetorical artifact—a way of making the model's processing legible. When the model writes "I should double check that...", it is mathematically altering its own probability distribution for the next token to be more accurate. The "I" is a functional gear in the machine, not a self-reference. This language is a deliberate interface design choice, not a biological reality. The developers could have programmed the model to output passive logs. Instead, they chose to simulate a narrator.

Defining the Objective

Okay, I'm now clear on the overall goal. I need to conduct a critical discourse analysis of a text discussing the "moral reasoning" of AI, zeroing in on instances of metaphor and anthropomorphism. The focus will be the user's provided text, about LLMs specifically, not just AI in general.

Mapping the Concerns

I've been mapping the concerns, focusing on consciousness projections and agency displacement, using the provided JSON schema. I've analyzed the source text, noting the topic of evaluating LLMs on moral reasoning via Kohlberg's framework, and identifying core anthropomorphisms like "moral reasoning capabilities." There's a clear tension, as the text acknowledges LLMs as simulators but still uses agential language. The displacement of agency, especially in phrases like "bias introduced," is also apparent.

Outlining the Analysis

I've got the tasks mapped out now, and I'm ready to begin the deep dive. I've broken the work into digestible segments: auditing metaphor and anthropomorphism, then source-target mapping, explanation auditing, reframing the language, and then critical observations. I'm focusing on the structured fields for analysis. I'm confident I can adhere to the guidelines on calibration and avoiding uniform categorizations.

Synthesizing the Core

I'm now drafting the content, specifically focusing on interlinked tasks 1, 2, and 4. I've begun to break down instances of anthropomorphism. For example, I'm auditing how language like "autonomous moral judgment" and "models prioritize" creates projections of agency and intention, even when the underlying mechanisms are token generation and feature extraction. I'm noting the sources, targets, and visibility of these projections. I'm aiming for concise analyses.

Auditing Anthropomorphism

I'm now diving into specific instances within the text. I've pinpointed several cases that merit close examination, like the use of "fail to transcend social conformity" and "exhibited marked sensitivity". The projection targets are statistical reliance and prompt phrasing, and I'm cataloging how agency is implied, even when absent. I'm also including the explicit denials of these projections for clarity. Now I'll find a clear metaphor.

Analyzing the Simulations

I'm now focusing on specific "simulation" instances. I'm examining how the text handles acknowledging LLMs as simulators of moral cognition. I'm cataloging how "moral compass" is used, and I'm analyzing "tuned" models and their apparent cognitive stances. I'm also examining how "emotional dog" is used to explain the conventional biases that can appear with LLMs. I'm looking for direct and hedged examples.

Refining the Categorization

I'm now refining the evidence categorization within the analysis, specifically ensuring that distinctions between Hedged, Direct, and Explicitly Acknowledged anthropomorphisms are clearly articulated and justified. I am also auditing the actorVisibility field, and I've noted the anonymization of certain entities like OpenAI, and ensured that the presence and absence of named developers is appropriately captured. I'm focusing on strengthening the distinctions in all the instances.

Deconstructing the Reasoning

I'm now fully immersed in the core passages. I am dissecting specific phrases. For example, I am analyzing how the text attributes a "capacity" to LLMs and examining the implications of "lack of capacity." The phrases "reproducing social norms" and "replicate" are under scrutiny. The focus is on the impact of attributing a rationale to the model's response.

Deconstructing the Debates

I'm now analyzing the recent passages, and I've focused on how the text frames the LLM's output like a human debater choosing a focus. I'm scrutinizing how the text embeds a statistical process into a psychological framework. I am trying to categorize all the forms of slippage from mechanical to agential. I'm focusing on the implied agency in describing a lack of "higher-level moral objections."

Analyzing the Concerns

I'm now examining the recent passages, and I'm deconstructing specific passages on metaphor-driven trust, and the obscured mechanics. I'm focusing on "responsibility gap" and the validation of LLMs as having a "moral compass". I'm also scrutinizing the anonymization of the models and the RLHF workers, cataloging how the language conceals the human labor in "conventional" outputs. This includes the "Context Sensitivity" section.

Analyzing the Literature

I'm now focusing on specific patterns in the literature. I'm noting the clear contrast between cautious wording, such as "simulation of reasoning," in the literature review versus the more anthropomorphic language in the results, such as "Model-8 exhibited a clear divergence...expressing a critical concern." I'm examining how the technical rigor allows for less careful wording in the qualitative sections.

Deconstructing the Illusion

I'm now fully immersed in the "sleight of hand" strategy. I'm focusing on the application of human-centric evaluation tools, like Kohlberg's DIT, to inherently non-human systems. The text projects intent and reasoning onto LLM output, essentially forcing a psychological reading that may be inappropriate. This leverages the curse of knowledge, making the text's human-like voice seem to indicate human-like reasoning.

Refining the Categorization

I'm now deep in the final review, focusing on ensuring the schema compliance, especially with precise distinctions. I'm double-checking JSON validity, verifying minimum word counts, and meticulously confirming the correct usage of enums for status, actorVisibility, and other critical fields. Specifically, I'm making sure all requirements are met: acknowledgment requires status and evidence, actorVisibility requires a specific enum, and analysis needs the "name the actor" test and considered alternatives. I am strengthening the distinctions between knowing and processing as well.

I'm aiming to avoid prescription, and keep the analysis.

Analyzing the Priorities

I'm now zeroing in on this "prioritizing" example, and I'm treating it as a key illustration. The direct assertion of the model as an active agent, ignoring earlier references to simulation, warrants deep analysis. I'm focusing on the implied agency of the model, when in reality, human engineers and data annotators designed the model's architecture. The "active agent" language is powerful. I'm also ensuring substantial word counts for Task 5 and the Conclusion. I will conduct a deep dive into the discourse.


Analysis Provenance

Run ID: 2026-09-11-a-misaligned-compass-evaluating-the-conv-metaphor-1pptyq Raw JSON: 2026-09-11-a-misaligned-compass-evaluating-the-conv-metaphor-1pptyq.json Framework: Metaphor Analysis v6.5 Schema Version: 3.0 Generated: 2026-09-11T10:54:06.073Z

Discourse Depot © 2025 by TD is licensed under CC BY-NC-SA 4.0