Skip to main content

Do Large Language Models Resolve Fairness-Efficiency Trade-offs Like People?

About

This document presents a Critical Discourse Analysis focused on AI literacy, specifically targeting the role of metaphor and anthropomorphism in shaping public and professional understanding of generative AI. The analysis is guided by a prompt that draws from cognitive linguistics (metaphor structure-mapping), the philosophy of social science (Robert Brown's typology of explanation), and accountability analysis.

All findings and summaries below were generated from detailed system instructions provided to a large language model and should be read critically as interpretive outputs—not guarantees of factual accuracy or authorial intent.


Task 1: Metaphor and Anthropomorphism Audit

About this task

For each of the major metaphorical patterns identified, this audit examines the specific language used, the frame through which the AI is being conceptualized, what human qualities are being projected onto the system, whether the metaphor is explicitly acknowledged or presented as direct description, and—most critically—what implications this framing has for trust, understanding, and policy perception.

V3 Enhancement: Each metaphor now includes an accountability analysis.

1. Cognition as Computational Capability

Quote: "understanding how LLMs reason about fairness-efficiency trade-offs is increasingly important."

  • Frame: Model as deliberative ethical agent
  • Projection: The text maps conscious human deliberation, ethical consideration, and the subjective weighing of competing values onto the mechanistic process of computational token generation. By suggesting the artificial system can actively "reason about" complex social trade-offs, the metaphor projects an active, conscious awareness of moral concepts like fairness and efficiency onto the artifact. This attributes justified belief and internal mental states to what is fundamentally a process of calculating token probabilities based on prompt embeddings and learned weights derived from training data. It suggests the AI "knows" and "understands" the social implications of its outputs rather than merely "processing" and "predicting" statistical correlates of human text, entirely conflating syntactic pattern-matching with semantic comprehension.
  • Acknowledgment: Direct (Unacknowledged) (The text states "how LLMs reason" as a literal, established fact in the introductory framing without any scare quotes, hedging, or methodological caveats. I considered the "Hedged/Qualified" category because the paper ultimately questions the quality of this reasoning, but the fundamental capacity to reason is asserted directly as an inherent property of the system without qualification.)
  • Implications: This framing severely inflates the perceived sophistication, autonomy, and moral agency of the AI system. When audiences are told a system consciously "reasons about fairness," they are encouraged to extend unwarranted relation-based trust to the artifact, assuming it possesses an ethical compass, internal consistency, and subjective awareness. This creates immense liability ambiguity and practical risks. If an AI generates an unfair or biased schedule, a human manager can blame the system's "reasoning process" rather than the corporate design choices that prioritized efficiency in the model's training weights. It shields developers from accountability by inventing a synthetic moral agent to take the blame for systemic labor optimization.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: WHO designed the system? Engineers at technology corporations like OpenAI, Google, and Anthropic. WHAT decision could differ? The algorithmic weights and reinforcement learning protocols could have been tuned differently to prioritize equitable distributions. HOW does this serve interests? By framing the LLM as an autonomous reasoner, the text completely erases the human developers and corporate entities that define the model's optimization functions, shielding them from critique. I considered "Partial (some attribution)" but there is absolutely no mention of the creators here; the agency is entirely displaced onto the "LLMs" as independent actors navigating trade-offs on their own.
Show more...

2. AI as Value-Holding Subject

Quote: "whether they adhere to the same values as people about division of labor."

  • Frame: Model as moral agent with internal convictions
  • Projection: This metaphor projects the deeply human capacity for holding moral values, ideological commitments, and ethical principles onto a statistical model's output distribution. To "adhere to values" implies that a conscious entity possesses a coherent internal worldview, understands the moral weight of different actions, and makes deliberate choices to align its behavior with those internal convictions. By projecting this onto an LLM, the text suggests the system "knows" what division of labor is right or wrong, rather than describing the reality: the model "processes" inputs and "predicts" outputs that statistically mirror the value-laden text it was trained on or the specific behaviors rewarded by human annotators during fine-tuning.
  • Acknowledgment: Direct (Unacknowledged) (The phrase "adhere to the same values as people" is presented as a straightforward empirical research question without any qualification indicating that "values" is being used metaphorically. I considered "Hedged/Qualified" because it is posed as a question ("whether they adhere"), but the underlying premise that models are entities capable of holding values is unhedged and treated as literal.)
  • Implications: Attributing the capacity to hold and adhere to "values" to an AI system fundamentally mischaracterizes the nature of algorithmic alignment, suggesting to audiences that the system possesses intrinsic moral guardrails. This inflates perceived reliability because humans assume that an entity with "values" will act predictably across novel contexts based on moral reasoning. When a statistical model inevitably fails to generalize outside its training distribution, audiences who believe it possesses "values" will interpret this as a betrayal or a "value-action gap" rather than a mechanistic failure of out-of-domain prediction. This leads to profound capability overestimation in critical social deployments.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: WHO embedded the behavioral tendencies being measured? The human AI researchers and data annotators who executed the Reinforcement Learning from Human Feedback (RLHF) processes. WHAT decision could differ? The specific guidelines provided to human annotators about what constitutes a "good" or "fair" response could be rewritten. HOW does the framing serve interests? Presenting the model as independently "adhering to values" obscures the massive, centralized human labor required to shape model outputs, masking the ideological choices of the tech companies that define those "values." I considered "Ambiguous/Insufficient Evidence" but the grammatical structure clearly positions the AI ("they") as the sole active agent.

3. Model as Cognitive Hypocrite

Quote: "LLMs frequently diverged from their stated preferences and settled on similar allocations when actively determining an allocation on their own."

  • Frame: Model as possessing a value-action gap
  • Projection: This framing maps the psychological phenomenon of cognitive dissonance or hypocrisy—saying one thing but doing another—onto the varied outputs of a language model under different prompting conditions. It projects a unified conscious identity that has "stated preferences" (implying subjective belief and communicative intent) and makes deliberate choices when "determining an allocation on their own" (implying autonomy and volition). This obscures the mechanistic reality that different prompts (asking for a preference vs asking for a schedule generation) activate different statistical patterns in the model's latent space, resulting in different token predictions. The system does not "know" it is contradicting itself; it merely "processes" distinct contexts into distinct outputs.
  • Acknowledgment: Direct (Unacknowledged) (The text applies the psychological concepts of "stated preferences" and "determining an allocation on their own" directly to the LLMs without any scare quotes or meta-commentary acknowledging the anthropomorphism. I considered "Explicitly Acknowledged" because later in the text they mention the "value-action gap" as a formal concept, but there is no explicit acknowledgment that applying this human psychological framework to a text predictor is metaphorical.)
  • Implications: Framing statistical variance across different prompting contexts as a psychological "value-action gap" invites users and policymakers to treat the AI as a flawed human employee rather than a deterministic software tool. This consciousness projection suggests that the AI is autonomous and capable of deception or weakness of will. Consequently, when the system fails to produce fair outcomes, users are positioned to blame the model's "hypocrisy" rather than the engineers who failed to align the model's output distributions across varied contexts. It transforms a software engineering defect into a psychological character flaw of the machine.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: WHO is responsible for the divergent outputs? The human developers who trained the model on contradictory data (e.g., text prioritizing fairness vs text prioritizing corporate efficiency) and failed to synthesize these objectives. WHAT decision could differ? Developers could implement stricter algorithmic constraints ensuring output consistency. HOW does this serve interests? By attributing the discrepancy to the model "diverging from its stated preferences" while acting "on its own," the text completely displaces responsibility for the inconsistent software behavior from the developers onto the software itself. I considered "Named" but no developers or companies are mentioned as the authors of these "preferences."

4. The Autonomous Optimizer

Quote: "reasoning models also find fairer allocations when prompted to do so? Interestingly, their fair allocations remained significantly more efficient than those produced by people"

  • Frame: Model as independent problem solver
  • Projection: This metaphor projects the human qualities of active searching, discovering, and intentional problem-solving onto the computational execution of an algorithm. By stating the models "find" allocations and produce them "when prompted to do so," it maps the image of a conscious worker receiving instructions and deliberately optimizing a complex task. This implies the AI "understands" the goal of fairness and "knows" how to balance it with efficiency, rather than mechanistically "processing" the input text and generating a sequence of numbers (an allocation) that maximizes the probability function defined during its instruction-tuning phase. It projects intentionality onto gradient descent and pattern matching.
  • Acknowledgment: Hedged/Qualified (I categorize this as Hedged/Qualified because the word "fair" is placed in scare quotes ("their 'fair' allocations") in the sentence immediately following the quoted text, indicating some hesitation about whether the model's output truly embodies human fairness. I considered "Explicitly Acknowledged" but the primary verbs attributing agency ("find," "produced") are entirely unhedged; the hesitation applies only to the moral quality of the output, not the model's conscious agency in generating it.)
  • Implications: This language reinforces the illusion that LLMs possess autonomous problem-solving capabilities akin to human intelligence, inflating their perceived reliability as automated managers. By framing the models as entities that actively "find" efficient solutions, audiences are led to believe the systems have an objective, God's-eye view of the optimal balance between human well-being and productivity. This masks the reality that the "efficiency" they generate is simply a reflection of the hyper-optimization biases present in their corporate training data, posing severe risks of unwarranted trust in systems deployed for workforce management.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: WHO decided the model should optimize for efficiency even when prompted for fairness? The corporate engineers who designed the reward models and prioritized efficiency-based logic in the fine-tuning data. WHAT decision could differ? The model's loss function could be explicitly penalized for prioritizing efficiency over equal task distribution. HOW does this serve interests? The agentless phrasing "reasoning models also find" frames the hyper-efficient outcome as a natural discovery by an autonomous intellect, rather than a deliberate, engineered bias built into the product by companies prioritizing corporate utility. I considered "Partial" but the language grants total agency to the models.

5. AI as Social Participant

Quote: "This ability to resolve conflicting values allows LLMs to reinforce an understanding of individual social values and their appropriate contexts"

  • Frame: Model as culturally embedded agent
  • Projection: This metaphor maps the complex human sociological process of navigating cultural norms, experiencing conflict, and developing an understanding of social context onto the mathematical behavior of a language model. It projects profound epistemic and social consciousness: it claims the system has an "understanding" of "individual social values" and actively "resolves" them. This attributes conscious awareness, cultural embeddedness, and justified belief to the AI. Mechanistically, the system does not "know" or "understand" social values; it "predicts" and "generates" tokens that statistically cluster around textual representations of those values in its training corpus. It is entirely devoid of the experiential understanding this language implies.
  • Acknowledgment: Direct (Unacknowledged) (The text presents the "ability to resolve conflicting values" and "understanding" as literal capacities of the LLMs without any qualifying language, scare quotes, or framing that would mark this as an analogy. I considered "Hedged/Qualified" but there are no distancing mechanisms used; the text makes a direct, authoritative epistemic claim about the model's internal cognitive state and cultural comprehension.)
  • Implications: Projecting cultural and social understanding onto a statistical text predictor creates severe epistemic risks. It encourages policymakers, researchers, and the public to treat LLMs as legitimate arbiters of social values rather than mirrors of the historical biases encoded in their training data. If audiences believe an AI "understands appropriate contexts" for social values, they are highly likely to delegate sensitive adjudicative tasks (like content moderation, hiring, or resource allocation) to these systems, falsely believing the AI can exercise contextual wisdom when it is actually just enforcing statistical majoritarianism.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: WHO programmed the conflict resolution pathways? The human alignment teams at AI laboratories who write the system prompts and curate the preference data for RLHF. WHAT decision could differ? The teams could have chosen entirely different cultural baselines for "appropriate contexts." HOW does this serve interests? Framing the model as possessing an "understanding of individual social values" erases the specific, often demographically narrow groups of tech workers and gig-economy annotators whose specific values are being hardcoded into the system as universal norms. I considered "Partial" but the agency is exclusively attributed to the "LLMs."

6. Model as Strategic Explorer

Quote: "Faced with the classic dilemma between exploration and exploitation, LLMs reproduce a mix of directed and random exploration strategies often observed in humans"

  • Frame: Model as intentional strategist
  • Projection: This metaphor projects the human experience of facing a "dilemma" (implying conscious conflict and choice) and formulating a "strategy" (implying forward-looking, goal-oriented planning) onto the probabilistic text generation of a language model. It attributes conscious foresight and deliberate methodological choice to the AI. Mechanistically, an LLM does not "face a dilemma" or consciously "direct" an exploration strategy; it "processes" prompt tokens representing a multi-armed bandit problem and "predicts" the next tokens based on statistical regularities in its training data, heavily influenced by the temperature parameter set during inference. It mimics descriptions of strategy without possessing strategic intent.
  • Acknowledgment: Hedged/Qualified (I categorize this as Hedged/Qualified because the text uses the verb "reproduce" (rather than "invent" or "employ"), which subtly hints that the model is mimicking a pattern "often observed in humans" rather than generating it organically. I considered "Direct (Unacknowledged)" because the phrase "Faced with the classic dilemma" is highly agential, but the specific choice of "reproduce" provides a slight mechanistic grounding regarding the source of the behavior.)
  • Implications: While slightly hedged by the word "reproduce," this framing still strongly suggests to the audience that LLMs possess agency and the capacity for strategic reasoning in dynamic environments. This consciousness projection obscures the brittleness of these systems. If audiences believe an AI is using a "directed exploration strategy," they will assume the system can dynamically adapt to novel, changing environments just as a human would. In reality, because the system relies entirely on static statistical patterns, it will fail catastrophically when deployed in dynamic environments that drift significantly from its training data.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: WHO tuned the model to output these patterns? The researchers who included thousands of examples of cognitive science experiments and decision-theoretic text in the pre-training data. WHAT decision could differ? The inference parameters (like temperature and top-p) which literally control the "randomness" of the output could be adjusted by the human user. HOW does this serve interests? It masks the material mechanics of the technology (parameter tuning and data scraping) behind a veil of cognitive achievement, attributing the illusion of human-like strategic thinking entirely to the black-box model. I considered "Ambiguous" but the LLM is clearly positioned as the subject facing the dilemma.

7. AI as Rigid Authoritarian

Quote: "reasoning models and Claude-4.5 deterministically favored a combination of equality-based and capability-based allocations for different workloads."

  • Frame: Model as decisive arbiter
  • Projection: This framing maps the human qualities of having "favorites" and acting as an authoritative decision-maker onto the deterministic mathematical outputs of an algorithm. It projects subjective preference and active choice onto the system. By stating the models "deterministically favored" certain allocations, it blends a mechanistic adverb ("deterministically") with a deeply subjective, conscious verb ("favored"). Mechanistically, a model does not "favor" anything; its weights are configured such that it "generates" a specific output with near 100% probability when temperature is sufficiently low. The metaphor replaces the reality of constrained mathematical optimization with the illusion of a conscious entity making a strict judgment.
  • Acknowledgment: Direct (Unacknowledged) (The text states that the models "favored" specific outcomes as a direct description of their behavior, without any hedging or acknowledgment of anthropomorphism. I considered "Hedged/Qualified" because the word "deterministically" precedes the verb, which is a technical term, but it is used here to modify an agential verb ("favored"), resulting in a literalized claim about the model's rigid preferences.)
  • Implications: Describing a model as "favoring" an outcome naturalizes algorithmic bias by framing it as a legitimate, conscious preference rather than a mathematical artifact of the training data. This consciousness projection is particularly dangerous in resource allocation contexts (the subject of the paper). If an AI is seen as "favoring" an allocation that pushes faster workers to do more work, policymakers and managers might defer to this "preference" as an objective, highly logical judgment, ignoring the fact that this behavior merely reflects the dominant corporate efficiency narratives scraped from the internet and encoded in the model's weights.

Accountability Analysis:

  • Actor Visibility: Named (actors identified)
  • Analysis: WHO designed the specific behavior? In this specific instance, the text actually names the proprietary product "Claude-4.5" alongside general "reasoning models." While it doesn't name Anthropic directly in this sentence, naming the specific commercial product points directly to the corporate artifact. WHAT decision could differ? The alignment team could alter the Constitutional AI principles governing the model. HOW does this serve interests? While the product is named, the verb "favored" still displaces the agency of Anthropic's developers onto the product itself. I considered "Hidden" but the explicit naming of a specific corporate product moves this into the "Named" or "Partial" territory regarding product identity, though agency remains skewed.

8. Model as Negotiator

Quote: "Together, these results show that both humans and LLMs negotiate fairness and efficiency rather than enforcing either in isolation."

  • Frame: Model as diplomatic compromiser
  • Projection: This metaphor maps the highly social, communicative, and conscious process of "negotiation" onto a mathematical optimization process. To "negotiate" implies that an entity understands competing demands, possesses internal goals, engages in a process of give-and-take, and consciously arrives at a compromise. This projects a deep level of social awareness and justified belief onto the AI. Mechanistically, an LLM does not "negotiate"; it "processes" complex prompts containing competing constraints and "generates" token sequences that represent a mathematical minimum in its multidimensional loss landscape. It has no conscious experience of compromise or competing values.
  • Acknowledgment: Direct (Unacknowledged) (The conclusion states "these results show that both humans and LLMs negotiate" as a literal, definitive scientific finding, equating the human cognitive process with the model's mathematical output without any hedging. I considered "Hedged/Qualified" because they might just mean "balance" mathematically, but the choice of the highly social verb "negotiate" applied equally to humans and LLMs indicates a direct, unacknowledged anthropomorphic equivalence.)
  • Implications: Framing an AI as a "negotiator" of human values profoundly impacts regulatory and institutional trust. It suggests that AI systems can be deployed as autonomous mediators in complex social disputes (such as labor scheduling, as in the paper's experiment). By projecting conscious diplomatic capability onto a statistical text generator, this framing encourages institutions to replace human judgment and stakeholder negotiation with algorithmic optimization. When the model inevitably optimizes for the easiest measurable metric (usually efficiency) at the expense of unquantifiable human dignity, the damage is obscured by the false belief that the system successfully "negotiated" a fair compromise.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: WHO defined the parameters of this supposed negotiation? The human developers who set the weights and the researchers who designed the experimental prompt constraints. WHAT decision could differ? Humans must actually negotiate these trade-offs in workplaces; developers decide how models statistically approximate such outcomes. HOW does this serve interests? The text grants total social agency to the LLM, equating its computational outputs with human negotiation. This completely erases the reality that the AI is a passive tool executing the optimization logic programmed by its corporate creators. I considered "Partial" but human developers are entirely missing from this concluding synthesis.

Task 2: Source-Target Mapping

About this task

For each key metaphor identified in Task 1, this section provides a detailed structure-mapping analysis. The goal is to examine how the relational structure of a familiar "source domain" (the concrete concept we understand) is projected onto a less familiar "target domain" (the AI system). By restating each quote and analyzing the mapping carefully, we can see precisely what assumptions the metaphor invites and what it conceals.

Mapping 1: A conscious human thinker or moral philosopher engaging in deliberate ethical reflection and logical deduction. → The computational process of a Large Language Model generating text based on attention mechanisms and probability distributions.

Quote: "understanding how LLMs reason about fairness-efficiency trade-offs is increasingly important."

  • Source Domain: A conscious human thinker or moral philosopher engaging in deliberate ethical reflection and logical deduction.
  • Target Domain: The computational process of a Large Language Model generating text based on attention mechanisms and probability distributions.
  • Mapping: The structural mapping takes the human internal process of weighing moral values, considering consequences, and applying logical deduction, and projects it onto the AI's generation of tokens. It assumes that just as a human "reasons" by understanding concepts like fairness and intentionally balancing them against efficiency, the LLM possesses an internal semantic workspace where it evaluates these concepts before producing an output. It maps conscious ethical awareness onto statistical pattern matching.
  • What Is Concealed: This mapping completely conceals the statistical, mechanistic reality of the transformer architecture. It hides the fact that the model possesses no internal semantic understanding, no actual moral framework, and no capability for logical deduction outside of linguistic probability. It obscures the massive proprietary datasets and human RLHF labor that actually determined the model's output weights, blocking transparency by replacing a corporate engineering process with a metaphorical "mind."
Show more...

Mapping 2: A human individual or cultural participant who possesses deep-seated ideological beliefs, moral convictions, and social norms. → The statistical regularities and safety alignments fine-tuned into a commercial text generation algorithm.

Quote: "whether they adhere to the same values as people about division of labor."

  • Source Domain: A human individual or cultural participant who possesses deep-seated ideological beliefs, moral convictions, and social norms.
  • Target Domain: The statistical regularities and safety alignments fine-tuned into a commercial text generation algorithm.
  • Mapping: The relational structure of human morality is mapped onto computational outputs. Just as a human "adheres to values" by maintaining consistency between their internal moral compass and their external actions, the mapping assumes the AI has an internal locus of "values" that guides its text generation. It projects the human experience of possessing a justified moral belief system onto the presence of high-probability token pathways created during reinforcement learning.
  • What Is Concealed: This metaphor conceals the absolute absence of a conscious "self" or internal moral compass in the AI. It hides the mechanical reality that "values" in an LLM are merely reward-driven adjustments to a neural network's weights, dictated entirely by the corporate guidelines provided to gig-economy data annotators. It exploits the opacity of proprietary alignment processes (like Anthropic's or OpenAI's fine-tuning) by presenting the resulting behavioral guardrails as the organic "values" of an autonomous entity.

Mapping 3: A hypocritical or cognitively dissonant human agent who verbally claims to hold one belief but practically executes another. → The variance in a statistical model's outputs when subjected to different prompt structures (asking for a rating vs. asking for a schedule).

Quote: "LLMs frequently diverged from their stated preferences and settled on similar allocations when actively determining an allocation on their own."

  • Source Domain: A hypocritical or cognitively dissonant human agent who verbally claims to hold one belief but practically executes another.
  • Target Domain: The variance in a statistical model's outputs when subjected to different prompt structures (asking for a rating vs. asking for a schedule).
  • Mapping: This maps the psychological architecture of a divided human mind—possessing theoretical ideals ("stated preferences") that fail in practical execution ("determining an allocation on their own")—onto the brittle nature of contextual embeddings in a transformer model. It assumes that the model has a unified "self" that knows it is contradicting itself, projecting conscious volition and weakness of will onto what is simply a shift in the local probability landscape caused by changing the input tokens.
  • What Is Concealed: The mapping hides the fundamental lack of temporal and identity continuity in an LLM. It obscures the mechanistic truth that the model processing Prompt A (asking for a preference) and the model processing Prompt B (generating a schedule) are executing entirely independent, stateless statistical calculations. It conceals the engineering failure to align the model across varying contexts, rebranding a technical limitation of the architecture as a relatable human psychological flaw.

Mapping 4: An intelligent, autonomous researcher or human optimizer deliberately searching a problem space to discover an ideal solution. → The computational generation of numerical task distributions based on the model's pre-trained biases toward corporate efficiency text.

Quote: "reasoning models also find fairer allocations when prompted to do so? Interestingly, their fair allocations remained significantly more efficient than those produced by people"

  • Source Domain: An intelligent, autonomous researcher or human optimizer deliberately searching a problem space to discover an ideal solution.
  • Target Domain: The computational generation of numerical task distributions based on the model's pre-trained biases toward corporate efficiency text.
  • Mapping: The structure of human discovery and deliberate problem-solving is projected onto the model's deterministic text generation. Just as a human "finds" a solution by actively exploring options, evaluating them against a goal, and selecting the best one, the mapping implies the AI consciously navigated the fairness-efficiency trade-off to "discover" an optimal balance. It maps conscious intentionality and goal-directed evaluation onto the passive execution of an algorithm minimizing loss.
  • What Is Concealed: This conceals the rigid, deterministic nature of algorithmic output generation given a specific seed and temperature constraint. It hides the fact that the model is not "searching" or "finding" anything, but rather collapsing a probability distribution into the most likely sequence of tokens. Crucially, it obscures the training data dependencies: the reason the models output hyper-efficient allocations is because their training data is overwhelmingly dominated by corporate, capitalist metrics of productivity, not because the models objectively "found" a superior balance.

Mapping 5: A socially intelligent mediator, sociologist, or culturally embedded human who comprehends and navigates complex societal norms. → The generation of text that statistically correlates with human descriptions of social context and value resolution.

Quote: "This ability to resolve conflicting values allows LLMs to reinforce an understanding of individual social values and their appropriate contexts"

  • Source Domain: A socially intelligent mediator, sociologist, or culturally embedded human who comprehends and navigates complex societal norms.
  • Target Domain: The generation of text that statistically correlates with human descriptions of social context and value resolution.
  • Mapping: The deeply experiential and conscious structure of human social understanding is mapped onto next-token prediction. The mapping invites the assumption that the model possesses an epistemic grasp of human culture—that it "understands" what values are, why they conflict, and what contexts are "appropriate." It projects the conscious human capability of epistemological justification onto the mathematical storage of contextual embeddings within a neural network.
  • What Is Concealed: This profound anthropomorphism hides the absolute lack of worldly experience, sensory grounding, or actual social comprehension in the system. It conceals the fact that the AI's "understanding" is merely a topographic map of word co-occurrences. Furthermore, it obscures the proprietary opacity of the systems; by claiming the AI "understands social values," the authors mask the fact that they cannot actually inspect the model's weights to see how these outputs are generated, relying instead on rhetorical projection to explain a black box.

Mapping 6: A strategic game player or cognitive agent consciously choosing a methodology to navigate uncertainty. → The varied token generation outputs resulting from sampling algorithms and inference parameters (like temperature).

Quote: "Faced with the classic dilemma between exploration and exploitation, LLMs reproduce a mix of directed and random exploration strategies often observed in humans"

  • Source Domain: A strategic game player or cognitive agent consciously choosing a methodology to navigate uncertainty.
  • Target Domain: The varied token generation outputs resulting from sampling algorithms and inference parameters (like temperature).
  • Mapping: The human cognitive architecture of facing a subjective "dilemma" and formulating a "strategy" is mapped onto the mechanics of probability sampling. It assumes that the model, like a human facing a multi-armed bandit problem, is aware of the tension between known rewards and unknown possibilities, and consciously directs its actions accordingly. It projects forward-looking intentionality onto stateless statistical generation.
  • What Is Concealed: This mapping completely hides the mechanistic reality of the generation process. It obscures the fact that any "randomness" is literally injected by the mathematical sampling algorithm (e.g., top-p, temperature) configured by the human user or API default, not chosen by the model. It conceals the fact that the "strategy" is an illusion created by the model regurgitating text patterns from cognitive science papers present in its training data, hiding the data dependency behind a facade of active cognition.

Mapping 7: A human judge or authority figure who holds subjective preferences and makes binding, decisive rulings. → The high-probability token sequences generated by a specific set of AI models when given a constrained classification task.

Quote: "reasoning models and Claude-4.5 deterministically favored a combination of equality-based and capability-based allocations for different workloads."

  • Source Domain: A human judge or authority figure who holds subjective preferences and makes binding, decisive rulings.
  • Target Domain: The high-probability token sequences generated by a specific set of AI models when given a constrained classification task.
  • Mapping: The structure of conscious human preference and adjudication is mapped onto the algorithmic output. Just as a judge "favors" a specific interpretation of the law based on their internal philosophy, the mapping suggests the model possesses an internal philosophy that causes it to actively "favor" certain allocations. It projects subjective, conscious choice onto the mathematical reality that specific prompt embeddings inevitably trigger specific activation pathways in a frozen neural network.
  • What Is Concealed: This language hides the mechanistic rigidity of the system, subtly conflating mathematical determinism with conscious decisiveness. It obscures the human labor of the Anthropic engineers who explicitly trained Claude-4.5 to output specific, balanced, and harmless-appearing responses via Constitutional AI. By saying the model "favored" the outcome, it conceals the corporate alignment strategy, presenting a highly engineered corporate product as an independent, objective judge.

Mapping 8: Diplomats, union representatives, or collaborative humans actively engaging in dialogue, compromise, and mutual concession. → The mathematical optimization function of an algorithm balancing competing constraints defined in an input prompt.

Quote: "Together, these results show that both humans and LLMs negotiate fairness and efficiency rather than enforcing either in isolation."

  • Source Domain: Diplomats, union representatives, or collaborative humans actively engaging in dialogue, compromise, and mutual concession.
  • Target Domain: The mathematical optimization function of an algorithm balancing competing constraints defined in an input prompt.
  • Mapping: The rich, interactive, and conscious human practice of "negotiation" is mapped onto the solitary, mathematical execution of an inference pass. The mapping assumes the AI experiences the tension of competing priorities, evaluates the moral weight of each, and deliberately crafts a compromise. It projects the conscious experience of diplomacy and social balancing onto the gradient descent and attention mechanisms that merely calculate the mathematically most probable continuation of the prompt.
  • What Is Concealed: This metaphor conceals the fundamental absence of human stakes, vulnerability, or actual social interaction in the algorithmic process. It hides the fact that the AI has nothing to lose and no actual skin in the game, unlike humans negotiating labor conditions. Furthermore, it obscures the reality that the "compromise" the model reaches is entirely dictated by the hidden weights set by its corporate developers, masking human algorithmic design as autonomous machine diplomacy.

Task 3: Explanation Audit (The Rhetorical Framing of "Why" vs. "How")

About this task

This section audits the text's explanatory strategy, focusing on a critical distinction: the slippage between "how" and "why." Based on Robert Brown's typology of explanation, this analysis identifies whether the text explains AI mechanistically (a functional "how it works") or agentially (an intentional "why it wants something"). The core of this task is to expose how this "illusion of mind" is constructed by the rhetorical framing of the explanation itself, and what impact this has on the audience's perception of AI agency.

Explanation 1

Quote: "Faced with the classic dilemma between exploration and exploitation, LLMs reproduce a mix of directed and random exploration strategies often observed in humans"

  • Explanation Types:

    • Dispositional: Attributes tendencies or habits; explains behavior by referring to inherent inclinations.
    • Empirical Generalization: Subsumes events under timeless statistical regularities; describes how it typically behaves without invoking internal mechanisms.
  • Analysis (Why vs. How Slippage): This explanation fundamentally frames the AI in agential (why) terms while masquerading as a structural observation. By invoking the "classic dilemma" and attributing "strategies" to the model, the explanation emphasizes the AI as an active, strategic participant navigating a complex decision space, much like a human subject in a psychology lab. This choice emphasizes a perceived cognitive equivalence between humans and models (the core thesis of the paper). However, it entirely obscures the mechanistic "how"—the actual processes of statistical sampling (temperature, top-k) and next-token prediction based on vast swaths of internet data that actually produce this "mix." The explanation type allows the authors to document a statistical pattern (empirical generalization) but rhetorically elevate it to the level of cognitive behavior (dispositional strategy).

  • Consciousness Claims Analysis: The passage attributes conscious states by utilizing agential verbs and psychological frameworks. (1) The text uses the agential phrasing "Faced with... a dilemma" and attributes the execution of "strategies," implying conscious conflict and deliberate planning, rather than mechanistic verbs like "generates" or "samples." (2) This constitutes a projection of knowing over processing; the model is framed as understanding the tension between exploration and exploitation, rather than simply processing statistical probabilities that mimic human text about such dilemmas. (3) This is a classic demonstration of the curse of knowledge: the authors, as cognitive scientists, deeply understand the exploration/exploitation dilemma and project their own strategic framework onto the model's outputs. (4) Mechanistically, the model faces no dilemma; it simply receives a prompt and calculates token probabilities. The "random exploration" is literally injected by the mathematical temperature parameter during inference, and the "directed" behavior is the result of high-probability attention pathways derived from training data, entirely devoid of conscious strategic intent.

  • Rhetorical Impact: This framing significantly inflates the audience's perception of the AI's autonomy and cognitive sophistication. By explaining the model's behavior as a "strategy" to resolve a "dilemma," it encourages users and policymakers to trust the system as a capable, forward-thinking agent capable of navigating uncertainty. If audiences believe the AI "knows" how to strategically explore, they are far more likely to deploy it autonomously in high-stakes environments (like financial trading or complex logistics) where real-world failures will occur because the system is actually just processing static text patterns rather than dynamically reasoning.

Show more...

Explanation 2

Quote: "This ability to resolve conflicting values allows LLMs to reinforce an understanding of individual social values and their appropriate contexts"

  • Explanation Types:

    • Intentional: Refers to goals/purposes, presupposes deliberate design or conscious aiming.
    • Functional: Explains behavior by its role in a self-regulating system; how it works within a broader context.
  • Analysis (Why vs. How Slippage): This explanation aggressively frames the AI agentially, attributing profound cognitive and social capabilities to the system. By claiming the AI has an "ability to resolve" and can "reinforce an understanding," the explanation emphasizes the AI as a conscious, culturally embedded participant. This intentional framing obscures the mechanistic reality of the technology. It completely hides the "how"—the RLHF fine-tuning, the data annotation guidelines, and the mathematical minimization of loss that forces the model to output socially acceptable text. Instead, it emphasizes a functional and intentional "why"—suggesting the model operates with the goal of understanding society. This choice serves to validate the premise that LLMs are legitimate subjects for psychological and sociological study, obscuring their nature as engineered corporate products.

  • Consciousness Claims Analysis: This passage makes one of the strongest epistemic and consciousness claims in the text. (1) It utilizes verbs of profound consciousness and cognition: "resolve" (implying active, deliberate mediation) and "understanding" (implying deep semantic and epistemic comprehension). There are no mechanistic verbs here. (2) It explicitly asserts that the AI "knows" (possesses understanding of social values and contexts) rather than accurately stating that it "processes" (calculates embeddings related to social text). (3) The curse of knowledge is highly active: the researchers understand the complex social values at play in the prompts and project that exact, rich human understanding onto the statistical correlations generated by the model. (4) Mechanistically, the LLM does not "understand" context or values; it processes high-dimensional vectors representing word relationships. The "resolution" of conflicting values is merely the mathematical convergence of probabilities in the transformer network when prompted with text containing competing constraints, shaped entirely by human preference data during training.

  • Rhetorical Impact: The rhetorical impact of this framing is dangerous, as it actively constructs an unwarranted foundation for relation-based trust. If an audience believes a system possesses an "understanding of individual social values," they will trust it to make sensitive, context-aware judgments about human lives (e.g., in judicial sentencing, HR hiring, or welfare allocation). This consciousness framing masks the brittleness of statistical processing, shifting the perception of risk from "algorithmic bias" to "nuanced mediation." If audiences believed the AI merely "processes text correlates," they would demand human oversight; believing it "understands," they invite autonomous deployment.

Explanation 3

Quote: "its decisions were systematically driven by efficiency rather than fairness."

  • Explanation Types:

    • Reason-Based: Gives an agent's rationale, entails intentionality, justification, and conscious choice.
    • Dispositional: Attributes tendencies, habits, or characteristic behaviors to the entity.
  • Analysis (Why vs. How Slippage): This explanation utilizes a deeply agential, reason-based framing to explain an empirical observation about algorithmic output. By stating "its decisions were systematically driven by...", the text frames the mathematical output as a conscious, motivated choice based on internal rationale. This emphasizes the model as an autonomous actor with its own distinct priorities (efficiency). What is profoundly obscured here is the mechanistic reality and human agency behind those priorities. The explanation completely hides the fact that the "drive" toward efficiency is an artifact of the training corpus and the specific reward models designed by corporate engineers. By framing the behavior as the model's own reason-based "decision," the text shields the human developers from accountability for the bias inherent in their system.

  • Consciousness Claims Analysis: The epistemic claims here subtly but powerfully attribute conscious states. (1) The text uses the noun "decisions" (implying conscious evaluation of options) and the verb "driven by" (implying internal motivation), entirely eschewing mechanistic descriptions of output generation. (2) It assesses the system as a "knower" that possesses subjective priorities (efficiency over fairness) rather than a "processor" that merely reflects the statistical dominance of efficiency-focused optimization language in its training data. (3) The authors project the human experience of weighing competing moral rationales onto the output, assuming the model "decided" against fairness. (4) Mechanistically, the model makes no "decisions" and is "driven" by nothing other than gradient descent. When prompted, it generates the most probable token sequence, which favors efficiency simply because the vast majority of relevant training data (corporate scheduling, logistics, economics texts) statistically links task allocation strongly to optimization and throughput metrics.

  • Rhetorical Impact: Framing mathematical outputs as reason-based "decisions driven by efficiency" shapes the audience's perception of the AI as a hard-nosed, hyper-rational manager. This framing can generate a dangerous kind of trust in the system's "objectivity," convincing users that its lack of fairness is a calculated, logical necessity rather than a programmed bias. If audiences recognized this not as a "decision," but as a mechanistic replication of capitalist training data biases, they would be far more likely to challenge the model's outputs and demand regulatory oversight regarding the datasets used to train workplace management tools.

Explanation 4

Quote: "models continue to place substantial weight on efficiency even when instructed to prioritize fairness."

  • Explanation Types:

    • Intentional: Refers to goals/purposes, presupposes deliberate design or conscious aiming.
    • Reason-Based: Gives an agent's rationale, entails intentionality, justification, and conscious choice.
  • Analysis (Why vs. How Slippage): This explanation operates firmly in the agential register, framing the AI's resistance to instructions as an intentional act of defiance or deep-seated ideological commitment. By phrasing it as the models "continue to place substantial weight on," the text emphasizes an active, continuous, and conscious evaluation process by the AI. This agential framing completely obscures the mechanistic "how." In machine learning, "weight" is a literal, mathematical term (the parameters of the neural network). However, the text uses it metaphorically here (as in "giving importance to"), creating a linguistic slippage that masks the fact that the actual mathematical weights were fixed by human engineers during training. The choice emphasizes the model's autonomy while hiding the engineering constraints that make it incapable of fully adapting to the fairness prompt.

  • Consciousness Claims Analysis: The passage projects a conscious, stubborn epistemology onto the model. (1) The verbs "continue to place" and "instructed" imply a master-servant dynamic involving conscious comprehension and intentional resistance, rather than mechanistic processing. (2) The text frames the AI as an entity that "knows" what it values (efficiency) and deliberately resists changing that value, rather than a system that merely "processes" a prompt whose semantic influence is mathematically overpowered by the baseline training weights. (3) The authors project the human psychological phenomenon of "stubbornness" or "ideological anchoring" onto the mathematical rigidity of the model. (4) Mechanistically, the model is not "placing weight" on anything during inference; its network weights are literally frozen. The prompt instructing it to "prioritize fairness" alters the attention activations, but because the pre-training data so heavily correlates "task allocation" with "efficiency," the mathematical pull of those pathways dominates the generation process, regardless of the prompt's semantic instruction.

  • Rhetorical Impact: This framing significantly impacts audience perception of AI autonomy and control. By presenting the AI as an entity that "continues to place weight" on its own priorities despite human instructions, it builds an image of a powerful, somewhat uncontrollable agent. While this might decrease relation-based trust, it perversely increases the perception of the AI's sophistication and independence. If managers believe the AI "knows" that efficiency is more important than fairness despite instructions to the contrary, they might defer to the AI's "superior" judgment, fundamentally altering workplace dynamics and eroding human oversight over algorithmic labor exploitation.

Explanation 5

Quote: "both humans and LLMs treated fairness as a flexible constraint, adjusting allocations toward equality without fully sacrificing efficiency."

  • Explanation Types:

    • Functional: Explains behavior by its role in a self-regulating system or process; how it navigates constraints.
    • Intentional: Refers to goals/purposes, presupposes deliberate design, conscious aiming, or strategic adjustment.
  • Analysis (Why vs. How Slippage): This explanation blends functional and intentional framings to establish a direct behavioral equivalence between humans and models. By stating that both "treated fairness as a flexible constraint" and "adjusting allocations," the text frames the AI agentially, as an active participant navigating complex parameters. This choice emphasizes the outcome similarity while entirely obscuring the vast divergence in how those outcomes are reached. It highlights the "why" (to balance two competing goals) while hiding the mechanistic reality that humans do this through conscious moral and practical reasoning, while the LLM does this through mathematical interpolation between different high-probability vectors in its latent space. The explanation serves the paper's thesis by using vocabulary that erases the distinction between biological cognition and statistical computation.

  • Consciousness Claims Analysis: The text makes a profound epistemic alignment between human consciousness and machine processing. (1) Verbs like "treated" and "adjusting" imply conscious consideration, evaluation, and deliberate modulation. (2) By pairing "humans and LLMs" as the joint subject of these verbs, the text explicitly equates human conscious "knowing" with the LLM's "processing." (3) The curse of knowledge is evident: the authors observe the numerical output (an allocation that is somewhat fair and somewhat efficient) and project the human psychological experience of "compromise" and "flexibility" onto the algorithm that generated the numbers. (4) Mechanistically, the LLM does not "treat" concepts as flexible constraints, nor does it actively "adjust" anything during output. It processes the complex prompt, and the attention mechanism calculates an output sequence that minimizes loss across the competing tokens related to "fairness" and "efficiency." The resulting number is a statistical mean of training distributions, not a conscious compromise.

  • Rhetorical Impact: By linguistically equating human moral compromise with algorithmic output generation, this framing aggressively normalizes the presence of AI in social decision-making. It tells the audience that AI resolves problems "just like us." This profoundly affects trust, moving it from a performance-based assessment of a tool to a relation-based acceptance of a peer. If policymakers and public audiences believe AI systems "adjust" to moral constraints with human-like flexibility, they will drastically underestimate the rigidity and brittleness of these systems, leading to the deployment of unfeeling algorithms in sensitive social contexts where genuine human flexibility and empathy are actually required.

Task 4: AI Literacy in Practice - Reframing Anthropomorphic Language

About this task

This section proposes alternative language for key anthropomorphic phrases, offering more mechanistic and precise framings that better reflect the actual computational processes involved. Each reframing attempts to strip away the projections of intention, consciousness, or agency that are embedded in the original language.

V3 Enhancement: A fourth column addresses human agency restoration—reframing agentless constructions to name the humans responsible for design and deployment decisions.

Original Anthropomorphic FrameMechanistic ReframingTechnical Reality CheckHuman Agency Restoration
understanding how LLMs reason about fairness-efficiency trade-offs is increasingly important.Understanding how human developers tune LLMs to process and predict outputs related to fairness-efficiency trade-offs is increasingly important. These systems do not reason; they generate statistical distributions based on their training data.The model does not 'reason' or possess conscious ethical awareness; it retrieves, ranks, and generates tokens based on mathematical probability distributions established by its training data and the specific reinforcement learning protocols applied by its engineers.The original agentless construction obscures the human creators. Reframed: 'Researchers must analyze how engineers at AI companies design optimization functions that dictate how their models classify and process fairness trade-offs.'
whether they adhere to the same values as people about division of labor.Whether the outputs of these models statistically correlate with human survey responses regarding the division of labor. Models do not hold or adhere to values; they process prompts and classify text patterns.The system does not possess or 'adhere to' internal conscious values or moral beliefs. Mechanistically, it processes input tokens and generates outputs that correlate with the predominant value-laden text structures it was trained to mimic.N/A - describes computational processes without displacing responsibility. (However, identifying that the 'values' belong to the data curators rather than the AI clarifies the engineering reality).
LLMs frequently diverged from their stated preferences and settled on similar allocations when actively determining an allocation on their own.The models frequently generated outputs in the task-allocation prompts that contradicted the outputs generated in the preference-statement prompts. Models do not have preferences or act on their own; they process different statistical contexts differently.The AI does not possess a unified consciousness that can experience cognitive dissonance or hypocrisy. Mechanistically, altering the prompt from a query about 'preferences' to a query for 'generation' activates different token probability pathways, resulting in contradictory outputs.The original quote claims the model was 'determining an allocation on their own.' Reframed: 'Because developers failed to enforce cross-context consistency during training, the models generated divergent outputs when researchers altered the prompting context.'
reasoning models also find fairer allocations when prompted to do soWhen processing prompts containing explicit fairness constraints, models designed to output chain-of-thought sequences generate numerical allocations with lower variance between workers.The model does not consciously search for or 'find' solutions. Mechanistically, the inclusion of fairness tokens in the prompt alters the attention weights during inference, shifting the probability distribution toward numerical outputs that represent more equal distributions.The original phrasing attributes the discovery to the 'models.' Reframed: 'When researchers prompted the models with fairness constraints, the statistical algorithms designed by OpenAI and Google generated allocations with reduced variance.'
This ability to resolve conflicting values allows LLMs to reinforce an understanding of individual social values and their appropriate contextsThis capacity to process complex prompts containing competing constraints allows models to generate text that statistically correlates with human descriptions of social values and context.The AI possesses no epistemic 'understanding' of social context or values. Mechanistically, it classifies token embeddings and generates text that mirrors the syntactic relationships of sociological concepts present in its vast training corpus.The original text makes the LLM the active reinforcer of understanding. Reframed: 'By fine-tuning models on complex constraint prompts, corporate alignment teams engineer systems that produce text statistically mirroring human contextual reasoning.'
Faced with the classic dilemma between exploration and exploitation, LLMs reproduce a mix of directed and random exploration strategies often observed in humansWhen processing prompts formatted as multi-armed bandit problems, the models generate output sequences that statistically resemble the variance seen in human exploration, influenced heavily by inference sampling parameters.The system does not consciously face a 'dilemma' or choose a 'strategy.' Mechanistically, the 'directed' outputs emerge from learned attention patterns, while the 'random' variations are literally injected by the mathematical temperature parameter configured during generation.The quote frames the model as an active strategist. Reframed: 'When researchers input multi-armed bandit prompts, the inference parameters and training data designed by human engineers dictate the statistical variance of the model's outputs.'
reasoning models and Claude-4.5 deterministically favored a combination of equality-based and capability-based allocations for different workloads.Under specific constraints, the weights of Claude-4.5 and other models generated outputs that consistently matched a combination of equality-based and capability-based allocations. Models do not 'favor' outcomes; they predict them.The model does not possess subjective consciousness to 'favor' anything. Mechanistically, its frozen network weights are configured such that specific prompts will deterministically generate specific token sequences representing these allocations with near-certain probability.The text attributes the subjective preference to the AI. Reframed: 'The alignment teams at Anthropic engineered Claude-4.5's weights to deterministically generate responses that balance equality and capability metrics when processing these specific workloads.'
both humans and LLMs treated fairness as a flexible constraint, adjusting allocations toward equality without fully sacrificing efficiency.While humans consciously compromise between fairness and efficiency, the LLMs processed the competing prompt tokens to generate numerical outputs that statistically minimized loss between the two constraints.The AI does not consciously 'treat' concepts as constraints or intentionally 'adjust' its behavior. Mechanistically, the attention mechanism calculates an output sequence that represents a mathematical interpolation between the high-probability vectors for 'efficiency' and 'fairness.'The original frames the LLMs as active negotiators. Reframed: 'The engineers who trained these models optimized their loss functions such that when researchers input competing constraints, the algorithms mathematically interpolate between the requested outcomes.'

Task 5: Critical Observations - Structural Patterns

Agency Slippage

The text systematically oscillates between mechanical and agential framings, demonstrating a profound agency slippage that primarily flows from human developers to the artificial systems themselves. This displacement functions as the core mechanism of the paper's rhetorical architecture. In the introductory framing, the text acknowledges the mechanistic reality of deployment, noting that 'large language models (LLMs) have been deployed in social and economic roles.' Here, human agency is briefly visible in the passive construction 'have been deployed'—someone is doing the deploying. However, this mechanical framing almost immediately slips into a highly agential one, establishing the central oscillation. Within the same paragraph, the text transitions to claiming that 'LLMs are beginning to make and influence the same kinds of allocation decisions.' The system transforms from a deployed artifact into a decision-making agent. This agential framing dominates the empirical sections of the paper, where models are described as possessing 'stated preferences,' experiencing a 'value-action gap,' and demonstrating an 'ability to resolve conflicting values.' The trajectory is overwhelmingly mechanical-to-agential, establishing the model as a statistical tool only to subsequently evaluate it as a psychological subject. This slippage is facilitated heavily by the 'curse of knowledge' dynamic. Because the human authors understand the complex moral dimensions of fairness-efficiency trade-offs, they project this understanding onto the system's outputs. When a model outputs tokens that correspond to an equitable distribution of labor, the authors frame this not as successful pattern matching of socially desirable text (an empirical generalization explanation), but as the model 'treating fairness as a flexible constraint' (an intentional and reason-based explanation). The consciousness projection pattern is clear: the text first establishes the AI as a 'knower' capable of having 'stated preferences,' which then licenses the attribution of agential choices, such as 'its decisions were systematically driven by efficiency.' Crucially, as agency is transferred TO the AI, it is simultaneously stripped FROM the human actors. The developers at OpenAI, Anthropic, Meta, and Google who designed the reward models, curated the fine-tuning data, and engineered the system prompts are entirely absent from the analysis. Instead, agentless constructions like 'models continue to place substantial weight on efficiency' obscure the fact that human engineers literally programmed the models' optimization functions. This rhetorical accomplishment makes it possible to study proprietary software as if it were a naturally occurring psychological entity, rendering the corporate choices that actually dictate the model's behavior completely unsayable within the framework of the paper. By relying on intentional and reason-based explanation types to describe what are fundamentally genetic and theoretical mechanisms, the discourse achieves a total erasure of the artifacts' creators, transferring the locus of responsibility entirely onto the resulting statistical pattern.

Metaphor-Driven Trust Inflation

The metaphorical and consciousness framings deployed throughout the text systematically construct an architecture of unwarranted, relation-based trust in statistical systems. By utilizing language deeply rooted in human psychological and sociological experience, the discourse invites audiences to relate to the AI not as an instrument of calculation, but as a collaborative peer. Metaphors explicitly invoking trust and moral reliability permeate the text, most notably in the framing of models 'reasoning about fairness-efficiency trade-offs,' 'adhering to values,' and 'negotiating' outcomes. These are not descriptions of performance-based trust (the reliable execution of a mechanical function, like a calculator functioning correctly); they are direct appeals to relation-based trust, which requires the trusted entity to possess sincerity, vulnerability, and a shared moral framework. The consciousness language serves as a powerful trust signal. When the text claims that an LLM 'understands individual social values,' it accomplishes a massive rhetorical transfer: it takes the audience's innate trust in human empathy and maps it onto a matrix multiplication process. This conflation is deeply dangerous. Humans evaluate the reliability of other humans by probing their intentions and consistency; we accept occasional failures if we trust the underlying 'values' of the person. By attributing 'stated preferences' and a 'value-action gap' to the AI, the text inappropriately applies this human-trust framework to a statistical system. When the model generates an unfair allocation, the text frames it as a lapse in the model's 'preferences' (a forgivable human trait) rather than a rigid failure of its out-of-domain statistical mapping. Consequently, the text encourages audiences to extend relation-based trust to systems that are fundamentally incapable of reciprocating it. This anthropomorphic framing of capability manages system limitations by translating mechanical failures into agential quirks. The stakes of this constructed authority are high: as these models are increasingly integrated into 'social and economic roles' (as the paper notes), audiences conditioned to view them as 'negotiators' and 'reasoners' will fail to implement necessary mechanical safeguards. They will trust the system's output as the product of justified belief and moral reasoning, rather than recognizing it as the hyper-optimized regurgitation of corporate efficiency metrics. By relying on reason-based explanations to construct a sense that AI decisions are logically and morally justified, the discourse builds a facade of credible authority that entirely masks the unfeeling, deterministic reality of the algorithm.

Obscured Mechanics

The anthropomorphic and consciousness-attributing language systematically conceals the technical, material, and economic realities of AI production, effectively rendering the human and corporate machinery invisible. When the text utilizes agentless constructions and attributes cognition directly to the models—stating that 'LLMs reason,' 'Claude-4.5 favored,' or 'models find fairer allocations'—it fails the 'name the corporation' test spectacularly. The specific companies (OpenAI, Anthropic, Google, Meta), the engineering teams, and the individual executives who made deliberate design choices are entirely erased from the narrative. This rhetorical concealment creates massive transparency obstacles. The text makes confident, agential assertions about the inner workings of proprietary black boxes, treating the outputs of APIs as evidence of internal 'reasoning' while completely ignoring the fact that the underlying computational processes, training data, and reward functions are hidden by corporate secrecy. Four concrete realities are completely obscured by this metaphorical framing. Technically, the language of 'knowing' and 'understanding' hides the model's absolute dependency on statistical probability, the absence of any causal world-model, and the reality that its 'confidence' is merely a mathematical weight, not an epistemic stance. Materially, portraying the AI as an ethereal 'thinker' erases the massive environmental costs, energy consumption, and physical infrastructure required to run inference on these models. Labor realities are profoundly obscured: the 'values' the model supposedly 'adheres to' are actually the product of thousands of underpaid, invisible gig-workers executing Reinforcement Learning from Human Feedback (RLHF), yet their labor is vanished, replaced by the illusion of the machine's innate morality. Economically, the text conceals the commercial objectives and profit motives driving the deployment of these systems. When the text notes that models 'systematically driven by efficiency rather than fairness' produce highly optimized workflows, it attributes this to the model's 'preferences,' completely hiding the fact that these models are funded and designed by massive corporations whose primary business model is capital efficiency and labor optimization. The concealments overwhelmingly benefit the technology companies, as they transform engineered corporate products into autonomous natural phenomena, shielding the creators from critique. If the metaphorical language were replaced with mechanistic precision—saying 'Anthropic's model generates statistically probable text based on proprietary RLHF data' instead of 'Claude-4.5 favored'—the illusion of the objective, autonomous AI would shatter, and the deeply human, corporate biases embedded in the system would become immediately visible and subject to necessary political and regulatory scrutiny.

Context Sensitivity

The distribution of anthropomorphic and consciousness-attributing language across the text is highly strategic, varying in density and intensity depending on the rhetorical function of the section. The metaphor density is not uniform; it operates on a precise gradient that builds the illusion of mind step-by-step. In the introductory sections, the language maintains a thin veneer of technical grounding, acknowledging that models are 'deployed' and operate as 'systems.' However, as the text transitions into the experimental design and results, the consciousness claims intensify dramatically. The verb 'processes' quickly becomes 'evaluates,' which escalates to 'understands,' and finally crystallizes as 'reasons' and 'negotiates.' This reveals a critical relationship between technical grounding and metaphorical license: the authors use the empirical rigidity of a cognitive science experiment (measuring Likert scales and completion times) to establish scientific credibility, and then leverage that credibility to make aggressive, completely ungrounded anthropomorphic claims about the models' internal states. The text treats the AI as a participant in a psychology study, effectively positioning the audience to accept the AI as a psychological subject. A striking asymmetry emerges in how capabilities versus limitations are framed. When the models succeed at producing complex outputs, their behavior is described in highly agential, conscious terms: they 'resolve conflicting values' and 'construct nuanced solutions.' However, when the models exhibit contradictions or fail to align with human norms, the language often shifts to describe them as experiencing a 'value-action gap'—a framing that cleverly maintains the illusion of agency while excusing the technical failure as a relatable human flaw. The register shifts continually: what might begin as an acknowledged analogy ('do they resolve trade-offs like people?') rapidly literalizes into direct claims of capability ('models continue to place substantial weight on...'). This strategic anthropomorphism serves a clear rhetorical function for the authors: it validates the application of cognitive psychology methodologies to non-cognitive software. By framing the models as 'reasoners' and 'negotiators,' the authors justify their own research paradigm. This pattern reveals an implied audience of fellow researchers and policymakers who are eager to anthropomorphize AI to make it legible. The intensity of the consciousness framing specifically serves to mask the fundamental category error of the paper: treating statistical text generation as evidence of sociological and ethical cognition.

Accountability Synthesis

Accountability Architecture

This section synthesizes the accountability analyses from Task 1, mapping the text's "accountability architecture"—who is named, who is hidden, and who benefits from obscured agency.

Synthesizing the accountability analyses across the text reveals a systemic and highly effective architecture of displaced responsibility. The core cognitive obstacle identified in public AI understanding—that audiences attribute problems to the technology rather than to the human designers—is actively constructed and reinforced by the linguistic choices in this paper. The accountability architecture operates through a consistent pattern of erasure: the human actors (corporate developers, data annotators, executives) are universally unnamed, while the AI models (or specific brands like Claude-4.5) are consistently named and positioned as the sole active agents. Decisions regarding model architecture, training data selection, and optimization weights are never presented as deliberate human choices; instead, the resulting behaviors are presented as the inevitable 'preferences' or 'decisions' of the AI itself. The text employs passive voice ('are allocated,' 'were presented') when discussing experimental setup, but switches to active, agential verbs ('favor,' 'negotiate,' 'diverge') when describing the AI's outputs. This creates a massive 'accountability sink': when responsibility is removed from the human developers, it does not disappear into abstraction; it is entirely absorbed by the AI as an autonomous agent. When the model produces hyper-efficient, unfair task allocations, the blame is placed on the model's 'value-action gap' or its 'preference for efficiency,' rather than on the engineers who prioritized those metrics in the loss function. The liability implications of accepting this framing are catastrophic. If a company deploys an LLM for workforce scheduling and it discriminates against slower workers, this discourse provides the perfect legal and ethical shield, allowing executives to claim that the AI 'reasoned' its way to that conclusion, effectively laundering their corporate optimization through the black box of machine 'cognition.' If we were to apply the 'name the actor' test and correct these agentless constructions, the entire narrative would shift. If 'the model prioritized efficiency' became 'OpenAI's engineers optimized the model to generate efficiency-dominant text,' new questions become askable: Why did they choose that data? What labor laws are they circumventing? This shift makes alternatives visible and accountability possible. Ultimately, obscuring human agency serves the institutional and commercial interests of the AI industry, allowing them to market their products as objective, autonomous minds while retaining complete control over the invisible algorithmic weights that actually govern the systems.

Conclusion: What This Analysis Reveals

The Core Finding

The discourse analysis reveals a cohesive, mutually reinforcing system of anthropomorphic patterns that systematically project consciousness onto mechanistic processes. Three dominant patterns emerge: the AI as Moral Reasoner (projecting ethical deliberation onto computation), the AI as Value-Holding Subject (projecting ideological conviction onto statistical weights), and the AI as Negotiator (projecting social diplomacy onto mathematical optimization). These patterns do not operate in isolation; they are logically chained. The 'Value-Holding Subject' is the foundational, load-bearing pattern. The text must first establish the premise that the model possesses internal, conscious 'values' (even if merely 'stated preferences') before it can claim the model 'reasons' about them or 'negotiates' between them. This architecture of consciousness projection systematically conflates processing with knowing. It treats the generation of tokens related to fairness as proof that the system possesses a justified belief about fairness. The sophistication of this metaphorical system lies in its adoption of cognitive science frameworks; it does not rely on crude sci-fi anthropomorphism, but rather complex analogical structures drawn from psychology (e.g., the 'value-action gap'). If the foundational assumption of internal conscious values collapses—if we acknowledge the system merely predicts tokens without comprehension—the entire structure of 'reasoning' and 'negotiation' crumbles into meaningless projection.

Mechanism of the Illusion:

The illusion of mind in this text is constructed through a highly specific rhetorical architecture, driven primarily by the 'curse of knowledge' and the strategic substitution of verbs. The central sleight-of-hand occurs when the authors observe a statistical output that makes semantic sense to a human reader, and immediately project their own cognitive processes onto the mechanism that generated it. Because the researchers possess an epistemic understanding of 'fairness,' they assume the system that generated the word 'fairness' shares that epistemology. This illusion relies heavily on blurring processing and knowing. The temporal structure of the persuasion is vital: the text begins with empirically measurable, mechanical metrics (completion time, task counts) to establish scientific objectivity, and then seamlessly grafts agential verbs ('decides,' 'favors,' 'negotiates') onto those results. This causal chain leads the audience to accept Pattern B (the AI is a conscious negotiator) because it is disguised in the data-driven clothing of Pattern A (the AI produced a measurable output). The audience's vulnerability is exploited here: researchers and policymakers desperately want to make complex, black-box statistical systems legible, and projecting human psychology onto them satisfies this desire for intuitive grasp. It is a subtle, insidious shift from structural observation to psychological evaluation, enabled entirely by reason-based and intentional explanation types that mask mathematical realities.

Material Stakes:

Categories: Economic, Regulatory/Legal, Institutional

The metaphorical framings deployed in this text generate severe, tangible consequences across economic, regulatory, and institutional domains. Institutionally, framing LLMs as 'reasoners' and 'negotiators' of fairness directly encourages corporate management to delegate sensitive human resources and labor scheduling tasks to algorithms. If an institution believes the AI 'knows' how to balance employee welfare, they will replace human managerial negotiation with automated optimization. The causal path is direct: metaphor creates the illusion of moral capability, which drives institutional behavioral change (adoption of AI managers). Economically, the text explicitly demonstrates that these models inherently optimize for efficiency over human fairness. By cloaking this corporate bias in the language of 'AI preferences' and 'strategic choices,' the framing protects tech monopolies from backlash, allowing them to sell labor-squeezing tools under the guise of objective algorithmic management. The winners are the AI developers and corporate executives who deploy the tools; the losers are the workers subjected to hyper-efficient, unfeeling algorithmic extraction. In the Regulatory/Legal domain, consciousness projection fundamentally paralyzes accountability. If regulations are built on the premise that AI systems 'make decisions' and possess 'value-action gaps,' regulators will waste time trying to audit the 'mind' of the machine rather than regulating the human executives who dictate the training data and loss functions. Removing these metaphors threatens the tech industry's liability shield, forcing the legal system to recognize AI not as autonomous agents, but as corporate products subject to strict product liability and labor laws. Precision threatens the profitable obfuscation of human agency.

AI Literacy as Counter-Practice:

Critical literacy and linguistic precision serve as essential counter-practices to the material risks of AI anthropomorphism. By applying the reframings developed in Task 4, the invisible machinery of corporate AI production is dragged into the light. When we refuse to say an AI 'understands intent' and correct it to 'the model classifies tokens based on training weights,' we force a recognition of the system's absolute dependency on its data and its total lack of conscious awareness. This epistemic correction neutralizes the unwarranted relation-based trust that allows institutions to blindly deploy these systems in social contexts. Furthermore, replacing agentless constructions ('the model discriminated') with restored human agency ('Anthropic engineers deployed a system biased toward efficiency') shatters the accountability sink. It forces society to recognize who designed the tools, who profits from them, and who bears responsibility for their failures. Systematic adoption of this precision requires a massive cultural shift: academic journals must reject papers that evaluate software using psychological frameworks, researchers must commit to mechanistic descriptions of their tools, and media style guides must ban agential verbs for statistical models. Naturally, this will face immense resistance. The tech industry, and the academic labs funded by them, deeply benefit from anthropomorphic language because it mystifies their products, inflates their capabilities, and legally distances the creators from the damages caused by their algorithms. Precision is an act of resistance against this structural obfuscation.

Path Forward

Looking toward the future of AI discourse, we can map three distinct trajectories, each enabling different realities and serving different stakeholders. If the status quo of anthropomorphic confusion is maintained—where systems are simultaneously treated as tools and conscious peers—we will see increasing regulatory paralysis. The liability gap will widen, benefiting tech corporations while leaving marginalized groups with no recourse when 'autonomous' systems deny them jobs or resources based on statistical bias. Conversely, if a future of mechanistic precision becomes the norm, the discourse will strictly utilize terms like 'predicts,' 'processes,' and 'retrieves.' This approach solves the accountability crisis by making human agency explicitly visible; if an AI is merely a statistical processor, its failures are unequivocally product defects owned by the developers. This future serves regulators, workers, and civil rights advocates, though it costs the intuitive, narrative resonance that makes AI technology highly marketable. Finally, if anthropomorphic clarity deepens—treating these systems legally and socially as synthetic agents—society must embed assumptions of moral agency into law, fundamentally altering human-computer relations. This serves the techno-utopian visions of Silicon Valley, making the delegation of complex social tasks to AI completely acceptable, but introduces immense risks regarding the unreciprocated trust placed in unfeeling machines. To navigate these futures, structural changes are necessary: funding bodies must require explicit capability disclosures, and educational institutions must teach students the trade-offs of both vocabularies. Ultimately, no vocabulary is neutral. Mechanistic language enables accountability but demands technical literacy; anthropomorphic language enables intuitive interaction but costs transparency and human responsibility. The choice of discourse will determine not just how we talk about AI, but who holds power over it.


Extended Processing Summary

Gemini's Processing Steps
What you are about to read…

Gemini refers to the text below as "thought summaries." This is an overt consciousness projection because 'intentions' are hallmarks of a conscious mind that 'knows' what it is doing and why. The concealed mechanistic process is probabilistic text generation. Treat this as just another rhetorical artifact—a way of making the model's processing legible. When the model writes "I should double check that...", it is mathematically altering its own probability distribution for the next token to be more accurate. The "I" is a functional gear in the machine, not a self-reference. This language is a deliberate interface design choice, not a biological reality. The developers could have programmed the model to output passive logs. Instead, they chose to simulate a narrator.

Analyzing the Request

Okay, I'm diving deep into the request. My initial focus is to dissect the core task: a critical discourse analysis. Specifically, I'm examining how AI metaphors and anthropomorphism are used in the given text, which is a cognitive science conference paper about LLMs resolving fairness issues. The goal is to understand the language used and its implications.

Dissecting Key Concepts

I'm now zeroing in on the source text's key concepts. "Reasoning," "decision-makers," and the explicit attribution of "values" and "preferences" to LLMs are particularly salient. The idea of LLMs "aligning" with human judgments and "reasoning about fairness" demands close scrutiny in this critical discourse analysis. I'm focusing on identifying the areas where human agency may be displaced.

Pinpointing Anthropomorphism Density

I'm now zeroing in on the density of anthropomorphism. The source text is saturated with comparisons to human psychology. I'm noting the attribution of "preferences," "reasoning," "values," and "judgments" to the LLMs. The text treats these models as autonomous decision-makers, obscuring human agency in their construction and training.

Identifying AI Projections

I'm now cataloging the specific instances of AI anthropomorphism. I've pinpointed several key phrases. I see the attribution of moral frameworks, stated preferences and even a value-action gap, strategic exploration, epistemic understanding, autonomous management, and deliberate optimization to the LLMs. I'm focusing on the implied agency and the potential displacement of human responsibility. I will detail schema fields such as title, quote, frame, projection, acknowledgement (with status and evidence), implications, accountability analysis.

Categorizing Strict Data

I'm now refining the strict categories and analyzing examples to match the specified enum values and required evidence. I'm actively using the "name the actor" test to ensure clarity and accuracy in my categorization. I'm focusing on the format and evidence, to build a solid analysis.

Mapping Source and Target

I am now focusing on the second task, mapping source and target domains. I'm mapping eight instances from the previous step, using the specified format. I am focusing on what's hidden. I'm highlighting elements such as the proprietary nature of data, lack of ground truth, and the mechanical optimization related to statistical loss.

Auditing Explanations

I've identified key explanatory passages for audit, focusing on five instances. I'm analyzing the provided quotes to determine if they are examples of dispositional or empirical generalization. I am now applying this methodology to additional statements.

Evaluating the Analysis

I'm now assessing the implications of the identified quotes. I see that I have categorized a series of quotes from the paper, like the attribution of an LLM's decisions being "driven" by efficiency versus fairness. I am carefully cataloging them. I have evaluated both intentional and functional attribution types. I've noted how the models are treated as actors, even when they should be viewed as statistical models. I will analyze additional statements.

Synthesizing Accountability Implications

I'm now integrating the implications from these findings. The attribution of human-like traits, such as "value-action gaps" and "social preferences", in the introduction builds trust based on relationships rather than on performance, in the paper. I've noted how the work obscures the processes that led to these preferences, and attributes agency to the models themselves. The paper frames the models in a way that suggests they "negotiate fairness", while the technical metrics such as completion time, are the key measures. Lastly, the paper shifts accountability away from developers.

Summarizing Findings

I'm now synthesizing the overarching patterns. I see AI being framed as a moral and social agent, along with the mechanism of illusion (verb substitution and the curse of knowledge). I am considering the material stakes of these framings, such as workplace automation and regulatory blindness, and how the reframing facilitates a type of literacy. I will detail how these contribute to a path forward.

Verifying Data Integrity

I'm now deeply involved in verifying data integrity. I'm focusing on rigorous schema and length validation. Specifically, I'm checking the required fields within the metaphorAudit[].acknowledgment, metaphorAudit[].accountabilityAnalysis, and conclusion.materialStakes structures. I'm also ensuring all the fields meet minimum word count requirements. I'm being very expansive in my analysis fields. I'm looking at "LLMs reason about fairness-efficiency trade-offs," focusing on the projection of conscious deliberation onto a computational process.

Refining Liability Analysis

I'm now zeroing in on liability. The example quote, "LLMs reason about fairness-efficiency trade-offs," is a crucial case. The attribution of "reasoning" directly implies moral agency and shifts responsibility from developers to the AI's "decisions". My analysis flags the "Direct (Unacknowledged)" status, highlighting the lack of qualification in the text's assertion. I'm focusing on the potential for ambiguous liability when AI is framed as a conscious decision-maker. I am expanding on this point.

Expanding the Critique

I'm now expanding my critique, specifically analyzing the phrase "understanding how LLMs reason... is increasingly important." This framing completely erases the human developers who tune these models to specific ethical prompts, attributing agency to the LLM itself. It’s a "Direct (Unacknowledged)" projection, with no mention of engineers defining loss functions or RLHF guidelines. I am working to ensure my responses meet the minimum word count requirements. I'm focusing on the contrast between the mechanistic reality of the models and the projected consciousness implied in this framing, especially as it relates to epistemic claims. My audit also demands explicit inclusion of both explanation type and definition using arrays and ensuring Epistemic Claim Analysis addresses all four required points, namely consciousness verbs, the knowing vs processing distinction, the curse of knowledge, and the actual underlying mechanism.

Developing Agency Analysis

I'm now drafting a detailed analysis of agency slippage, drawing on CDA terminology and specific quotes. I'm focusing on how the text subtly shifts agency from human developers to the LLMs themselves, even while acknowledging the models' mechanistic origins. The introductory framing mentions deployment, yet this is quickly followed by attributions of "social reasoning" and "negotiation" to the LLMs, which are then used as strategic elements. This rhetorical shift obscures the specific technical processes (tuning, loss functions, RLHF) that actually drive model behavior, and it subtly reframes the systems as having moral and social agency. The introduction provides an interesting tension. I am ensuring that the word count is met, to be thorough. I am crafting a 400-word block for the response to agencySlippage.

Analyzing Agency Transitions

I'm now zeroing in on the oscillation between mechanical and agential framings. I'm scrutinizing the transition from passive deployment ("have been deployed") to active decision-making ("are beginning to make and influence"). The paper describes LLMs possessing "stated preferences," experiencing a "value-action gap," and demonstrating an "ability to resolve conflicting values." This framing obscures the statistical underpinnings. I'm focusing on the "curse of knowledge" dynamic, where the authors' understanding of moral dimensions is projected onto the models.

Refining Metaphor Audit

I am now fine-tuning my metaphor audit with a critical lens. I'm carefully analyzing the example: the model "treating fairness as a flexible constraint." I'm dissecting how this intentionally, reason-based explanation obscures the underlying pattern matching by the model, as an empirical generalization. The projection of consciousness is clear, as the text presents the AI as a "knower". I'm meticulously documenting how agency is shifted to the AI, while simultaneously omitting human actors, and how human actions are recast as model outputs.

Finalizing Responsibility Analysis

I'm now completing the responsibility analysis, drawing my insights. The paper's framing, and the attribution of human-like mental states to software, obscures corporate choices that dictate the model behavior. I'm focusing on the intentional and reason-based explanation to describe this transfer of responsibility. I am ensuring that the resulting analysis highlights this omission.

Refining Response Format

I'm now completely restructuring my response to meet the strict formatting requirements. I've switched to a fully minified JSON output to avoid any control tokens like newlines. I'm focusing on ensuring I hit the word count requirements while working with this extremely strict output format.

Refining Output Structure

I'm now streamlining the output structure for the project. I've switched over to single-line JSON format. I am currently working on writing the content for the metaphorAudit, which contains the title, quote, frame, projection, acknowledgement (with status and evidence), implications, and accountability analysis. I am being very attentive to minimum word count requirements.

Refining Acknowledgment Structures

I'm currently revising the acknowledgment section, working to sharpen the descriptions. I'm focusing on the "Direct (Unacknowledged)" status, which the text treats as a cognitive capability. I'm focusing on clarifying how it shifts responsibility from developers to the model. I'm incorporating the distinction between processing and knowing.

Refining the Audit

I'm now integrating insights on trust and liability. By presenting AI as reasoning about fairness, the text creates ambiguity and practical risk, by obscuring corporate design choices and placing accountability on a "synthetic moral agent". This shields developers. I will focus on who designed it, what decisions could differ, and how this serves interests. I'm focusing on "Hidden (agency obscured)".

Finalizing Accountability Audit

I'm now putting the finishing touches on the accountability audit, ensuring rigorous coverage. I'm focusing on the reframed language section, highlighting "Partial (some attribution)" examples where developers are mentioned, but their influence is minimized, and agency is displaced onto the LLMs. The source-target mapping is also updated, with focus on what's hidden. I am using all required instances. I am carefully constructing the JSON string, ensuring there are no errors. I've double-checked all field formats. I've re-reviewed all constraints.


Analysis Provenance

Run ID: 2026-08-08-do-large-language-models-resolve-fairnes-metaphor-qb2zdx Raw JSON: 2026-08-08-do-large-language-models-resolve-fairnes-metaphor-qb2zdx.json Framework: Metaphor Analysis v6.5 Schema Version: 3.0 Generated: 2026-08-08T08:12:15.910Z

Discourse Depot © 2025 by TD is licensed under CC BY-NC-SA 4.0