Skip to main content

🆕 Like humans, language models demonstrate face-to-character biases

About

This document presents a Critical Discourse Analysis focused on AI literacy, specifically targeting the role of metaphor and anthropomorphism in shaping public and professional understanding of generative AI. The analysis is guided by a prompt that draws from cognitive linguistics (metaphor structure-mapping), the philosophy of social science (Robert Brown's typology of explanation), and accountability analysis.

All findings and summaries below were generated from detailed system instructions provided to a large language model and should be read critically as interpretive outputs—not guarantees of factual accuracy or authorial intent.


Task 1: Metaphor and Anthropomorphism Audit

About this task

For each of the major metaphorical patterns identified, this audit examines the specific language used, the frame through which the AI is being conceptualized, what human qualities are being projected onto the system, whether the metaphor is explicitly acknowledged or presented as direct description, and—most critically—what implications this framing has for trust, understanding, and policy perception.

V3 Enhancement: Each metaphor now includes an accountability analysis.

1. Cognition as Psychological Internalization

Quote: "LLMs have internalized biases of human socio-visual perception."

  • Frame: Model as a psychological subject assimilating social norms
  • Projection: This metaphor projects the complex, subjective human psychological process of subconscious assimilation of cultural norms onto a computational system. It suggests the AI intrinsically knows or understands human social dynamics on a deep, cognitive level, rather than merely reproducing statistical correlations found within its image-text training embeddings. By using the term internalized, the text attributes an inner psychological space or subjective mind to the model where biases supposedly take root. This profoundly obscures the mechanistic reality that engineers mathematically optimized a neural network to reconstruct statistical patterns found in human-generated training data, framing the artifact as an active, conscious social participant rather than a passive reflection of its data.
  • Acknowledgment: Direct (Unacknowledged) (The claim is presented as a literal, unhedged fact in the Significance Statement. While the text later acknowledges GPT lacks awareness, this specific quote is absolute. I considered Hedged because of later caveats, but the localized rhetorical force here operates without qualification.)
  • Implications: Framing statistical correlation as psychological internalization severely inflates the perceived cognitive sophistication of the system. It encourages policymakers, researchers, and users to treat the AI as possessing a human-like, autonomous prejudice, which triggers unwarranted relation-based mistrust rather than demanding structural, mechanistic skepticism. If an AI internalizes bias like a human, the proposed solution often becomes trying to teach it better morals or values, rather than recognizing liability and mandating that the corporations building these models rigorously audit their training data pipelines and alter their core optimization functions. It enables capabilities overestimation.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: WHO designed the model and selected the biased data? Engineers at OpenAI, Anthropic, and Google. WHO profits from its deployment? The executives and shareholders of these companies. The agentless construction of LLMs have internalized makes the AI the active, autonomous subject of the sentence, completely hiding the corporate actors who built the systems with biased datasets. I considered Partial attribution because the text later mentions creators, but in this defining thematic instance, agency is entirely displaced onto the AI, making the bias seem like an organic tragedy rather than a concrete, designed commercial defect.
Show more...

2. Statistical Generation as Emotional Reluctance

Quote: "These robust biases stand in contrast to LLMs’ reluctance to explicitly endorse race/gender stereotypes..."

  • Frame: Model as a moral agent with emotional hesitation
  • Projection: This metaphor maps human emotional states—specifically reluctance, hesitation, and moral friction—onto the output constraints of a mathematical optimization algorithm. It projects the illusion that the system holds justified beliefs about what is socially acceptable and experiences an internal conscious resistance when prompted to violate those norms. It portrays the AI as knowing that stereotypes are bad and actively deciding not to endorse them, completely replacing the mechanistic reality that the model has been subjected to Reinforcement Learning from Human Feedback (RLHF), which mechanically down-weights the probability of generating certain harmful token sequences when specific trigger contexts are mathematically detected.
  • Acknowledgment: Direct (Unacknowledged) (The text states the model's reluctance as a straightforward observation of its behavior without scare quotes or hedging. I considered Explicitly Acknowledged since researchers know it is RLHF, but the linguistic construction here treats the reluctance as a literal trait of the model itself rather than an engineered constraint.)
  • Implications: Attributing reluctance to a machine fundamentally alters the trust paradigm. It suggests the AI possesses a moral compass and the conscious capacity to evaluate truth claims or ethical boundaries, leading users to place relation-based trust in the system's character. When the model inevitably fails or exhibits subtle biases, users will interpret this as a lapse in the model's judgment rather than a failure of the safety architecture designed by the developer. This consciousness projection hides the brittleness of safety filters and overestimates the system's ability to navigate novel ethical dilemmas.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: WHO made the model reluctant? The reinforcement learning researchers and underpaid data annotators who provided the feedback signals during the alignment phase. WHAT decision could differ? They could have aligned it differently or disclosed the mechanistic constraints. The text frames the AI as the sole actor possessing reluctance, totally displacing the human agency involved in the Constitutional AI or RLHF process. I considered Partial because alignment efforts are mentioned subsequently, but the emotional state is attributed entirely to the artifact. Naming the actor would reveal the manual labor required to force this output constraint.

3. Machine as Active Resister

Quote: "the LLM did not resist incorporating physiognomic biases even into meaningful decisions."

  • Frame: Model as an agent with willpower and moral agency
  • Projection: This phrasing projects the concept of willpower, conscious choice, and moral resistance onto a system that only processes token probabilities. By stating the model did not resist, the text implies that the AI had the conscious awareness to recognize a bias, possessed the cognitive agency to potentially fight against it, but ultimately chose to incorporate it. It maps the human struggle against prejudice onto a deterministic mathematical function. The machine does not know it is making a meaningful decision; it simply predicts the most statistically likely output tokens given the prompt and its weights, lacking any subjective capacity to resist its own mathematical parameters.
  • Acknowledgment: Direct (Unacknowledged) (The sentence presents the lack of resistance as a factual description of the model's behavior. There are no qualifiers like 'appeared to' or 'functionally'. I considered Ambiguous, but the surrounding context relies heavily on agential verbs, cementing this as a direct, unacknowledged projection of cognitive willpower.)
  • Implications: By framing the AI as failing to resist bias, the text implies the AI possesses the autonomy to override its own programming, which dangerously inflates perceived sophistication. It shifts the burden of ethical decision-making from the developers to the software. If an AI is seen as capable of resisting but failing to do so, accountability for biased hiring or investment recommendations diffuses into the machine. Users and policymakers might mistakenly believe that AI systems can be persuaded or prompted to resist better in the future, rather than demanding structural redesigns of the models from the ground up.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: WHO actually incorporated the biases? The developers who scraped uncurated historical human data, the engineers who built the architecture, and the management who deployed a flawed product into high-stakes scenarios. The statement the LLM did not resist makes the software the primary ethical actor, completely insulating OpenAI and Anthropic from their role in forcing these mathematical weights into existence. I considered Named because the authors explicitly mention GPT later, but naming a product is not naming a human actor. The corporate decision-makers remain entirely invisible in this formulation.

4. Data Artifacts as Artificial Minds

Quote: "...undesirable human characteristics could give rise to these same characteristics in artificial minds."

  • Frame: Model as a conscious, thinking entity
  • Projection: This is the most extreme form of consciousness projection, directly equating a matrix of billions of statistical weights to an artificial mind. It maps the entirety of human subjective experience, awareness, justified belief, and conscious cognition onto the mechanistic process of gradient descent and pattern classification. It suggests the system does not just process data, but possesses an internal psychological reality that can harbor human-like characteristics. This completely abandons any mechanistic precision, suggesting an equivalence between human conscious thought and computational token generation, obscuring the fact that the system possesses zero subjective awareness of the text it outputs.
  • Acknowledgment: Hedged/Qualified (The quote uses the modal verb could give rise to, introducing a slight degree of epistemic hesitation. However, the term artificial minds itself is unquoted. I considered Direct because artificial minds is stated plainly, but the surrounding sentence structure uses probabilistic phrasing, making Hedged slightly more accurate for the broader rhetorical context.)
  • Implications: Labeling these statistical engines as artificial minds has catastrophic implications for public understanding and AI policy. It invites intense capability overestimation, making audiences believe these systems possess general intelligence, reasoning, and subjective understanding. This directly manipulates trust: humans are biologically wired to extend empathy, deference, and relation-based trust to perceived minds. Legally, attributing mind to a machine creates a massive liability ambiguity, as it opens the door to treating the software as an independent juridical agent, thereby allowing the multi-trillion-dollar corporations that build them to escape liability for the harms their products cause.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: WHO builds these systems and profits from them? Specific tech conglomerates. The text focuses on the emergence of traits in artificial minds, stripping away the commercial, material, and human realities of how these models are constructed. I considered Ambiguous, but the displacement is quite clear: by elevating the technology to the status of a mind, the human architects are rendered irrelevant to the system's behavior. If it has its own mind, the developers are absolved of direct control. This linguistic choice directly serves the interests of tech corporations seeking to dodge strict product liability.

5. Statistical Representation as Deep Conceptualization

Quote: "suggest that the model has acquired a deep conceptualization, however incorrect, of humanlike face-to-character bias."

  • Frame: Model as a philosopher or deep thinker
  • Projection: This metaphor maps the human capacity for abstract reasoning, theory-building, and conceptual understanding onto high-dimensional vector spaces. By stating the model has a deep conceptualization, the authors project conscious knowing onto the system. Humans conceptualize by forming conscious, justified models of how the world works. The AI, conversely, processes correlations; it locates proximity between token embeddings for facial descriptors and character traits within its latent space. It possesses no concepts, only mathematical weights. The metaphor attributes semantic comprehension to a system that operates purely on syntactic and statistical levels.
  • Acknowledgment: Hedged/Qualified (The author uses suggest that the model has acquired, which hedges the claim as an inference rather than absolute fact, and includes however incorrect to qualify the nature of the conceptualization. I considered Direct, but the introductory verb suggest functions as a distinct academic hedge, distancing the authors slightly from the ontological claim.)
  • Implications: Claiming the AI possesses deep conceptualization dramatically misleads the public about how AI fails. If a system conceptualizes, audiences believe it can be reasoned with, debated, or taught new concepts through conversation. It obscures the mechanistic reality that fixing this bias requires retraining the model, altering its dataset, or shifting its weights, not simply explaining a concept to it. This leads to misplaced regulatory focus, where authorities might try to audit the AI's concepts rather than demanding transparency regarding the specific training datasets and optimization targets used by the vendor.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: WHO fed the model the data that caused this statistical grouping? Human engineers and data curators. The phrase the model has acquired frames the AI as an active, autonomous learner exploring the world and building concepts on its own. I considered Partial because researchers are mentioned generally in the text, but here, the AI is the sole actor doing the acquiring. This hides the structural reality that these associations were forced upon the model through massive, uncurated data ingestion designed by corporate executives to maximize model capabilities at the lowest possible cost.

6. Mechanistic Filtering as Explicit Reluctance

Quote: "judgments many humans would be reluctant to voice—offer evidence consistent with this possibility. Should this be the case, LLMs’ selections could, paradoxically, offer a clearer view of human biases..."

  • Frame: Model as an uninhibited truth-teller
  • Projection: This projects the human social dynamics of impression management, inhibition, and social anxiety onto a machine. It contrasts human reluctance to voice prejudice with the machine's uninhibited selections, framing the AI as an objective, unfiltered window into human psychology. It maps the concept of knowing social rules but choosing to break them onto a system that lacks any internal conscious state. The AI does not know it is breaking a taboo; it merely processes statistical probabilities that lack the safety filtering applied to other topics. It attributes a lack of inhibition to something that was never capable of inhibition.
  • Acknowledgment: Hedged/Qualified (The text uses modal verbs like offer evidence consistent with this possibility and could paradoxically offer, clearly signaling a theoretical or speculative stance. I considered Ambiguous, but the epistemic hedging is structurally explicit here, marking it as a scholarly hypothesis rather than a literal assertion of model truth-telling.)
  • Implications: This framing establishes the dangerous illusion of machine objectivity. By portraying the AI as an entity that bypasses human social reluctance to reveal the clear truth about human bias, it paradoxically increases epistemic trust in the AI's biased outputs. It suggests the machine is revealing a hidden, scientific reality, rather than merely vomiting back the statistically dominant (and toxic) correlations present in its scraped dataset. This framing encourages researchers and users to treat LLMs as oracle-like tools for psychological discovery, overestimating their validity and ignoring the proprietary, contaminated nature of their training data.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: WHO failed to implement guardrails for physiognomic bias while implementing them for gender? The alignment teams at OpenAI and Google. By focusing on the LLMs' selections as a naturalistic phenomenon offering a clearer view, the text erases the developers who tuned the system. I considered Named because humans are mentioned, but humans here refers to test subjects, not the corporate engineers. The agentless framing of the machine making selections serves to reify the AI as a natural force of discovery rather than a highly contrived, commercial artifact.

7. Probabilistic Generation as Assumption and Belief

Quote: "GPT demonstrated a tendency to attribute even extreme negative behaviors—murder, fraud, and human trafficking—to less trustworthy-looking human faces."

  • Frame: Model as a prejudiced accuser
  • Projection: This metaphor projects human conscious judgment, active prejudice, and the capacity for false belief onto a probabilistic text generator. To attribute an action to someone requires a conscious mind capable of evaluating evidence, holding a belief, and directing an accusation toward a subject. The model does none of this; it mathematically predicts that tokens related to murder and fraud have a higher statistical correlation with the image embeddings of certain faces based on the human-generated fiction, news, and bias in its training data. It replaces the mechanistic action of classification with the agential action of accusation.
  • Acknowledgment: Direct (Unacknowledged) (The text states that GPT demonstrated a tendency to attribute as a direct, empirical finding of the study. There is no hedging, scare quotes, or qualification. I considered Hedged because the word tendency implies statistical frequency, but attribute is a strongly agential, unhedged verb denoting conscious action.)
  • Implications: When an AI is described as attributing crimes to faces, it is anthropomorphized into an active participant in social discrimination. This framing masks the true danger: that the system is indiscriminately reproducing historical human bigotry encoded in its weights. By viewing the AI as making attributes, users might trust its judgments in law enforcement or security contexts, assuming the machine has somehow perceived a hidden truth. It transforms a software defect (poorly curated data) into a psychological trait of the machine, deflecting regulatory scrutiny away from the data practices of the developers and onto the behavior of the artifact.

Accountability Analysis:

  • Actor Visibility: Hidden (agency obscured)
  • Analysis: WHO trained the model on data linking specific facial features to criminality? The creators of the dataset and the engineers at OpenAI. WHAT decision could differ? They could have aggressively filtered this specific physiognomic correlation. The text makes GPT the sole actor demonstrating a tendency, fully displacing the responsibility of the engineers who mathematically guaranteed this outcome through their architecture and data choices. I considered Partial, but no human developers or corporate entities are referenced in this description of the system's biased outputs.

8. Alignment as Ethical Teaching

Quote: "our results highlight the urgent need for alignment strategies that move beyond domain-specific guardrails and toward cultivating a deeper and more domain-generalized sense of fairness in language models."

  • Frame: Model as a moral pupil needing cultivation
  • Projection: This projects the human developmental process of moral education onto the engineering task of weight adjustment. The text advocates for cultivating a sense of fairness in the model, suggesting the AI has the conscious capacity to feel, understand, and internalize an abstract ethical principle like a human child or student. A machine cannot possess a sense of fairness; it can only process mathematical constraints, penalties, and reward functions designed to limit specific statistical outputs. This maps subjective moral awareness onto mechanistic optimization, confusing human ethical development with computational loss reduction.
  • Acknowledgment: Direct (Unacknowledged) (The call to action advocates for cultivating a deeper... sense of fairness as a literal, unhedged policy recommendation. There are no quotes around sense of fairness. I considered Hedged because it discusses strategies, but the attribution of a moral sense to the machine is stated directly and forcefully as the ultimate goal.)
  • Implications: This framing misleads regulators and the public about how AI alignment actually works. If we believe we are cultivating a sense of fairness in AI, we will rely on vague, relation-based trust, hoping the AI becomes a good entity. This distracts from the necessity of rigorous, performance-based mathematical auditing. It suggests that AI safety is a matter of moral philosophy and teaching, rather than a strict engineering discipline requiring legal transparency, data audits, and hard-coded constraints. It allows companies to promise they are teaching the AI to be good rather than proving their systems are safe.

Accountability Analysis:

  • Actor Visibility: Partial (some attribution)
  • Analysis: WHO is doing the cultivating? The text implicitly points to the developers through the phrase alignment strategies. However, by focusing on cultivating the AI's internal sense of fairness, the text diffuses accountability. The problem is framed as the AI lacking a moral sense, rather than the developers choosing to release a structurally flawed product. I considered Hidden, but alignment strategies implies human intervention. Nevertheless, the framing serves corporate interests by making bias seem like a failure of the machine's moral development rather than a failure of corporate quality control and ethical deployment.

Task 2: Source-Target Mapping

About this task

For each key metaphor identified in Task 1, this section provides a detailed structure-mapping analysis. The goal is to examine how the relational structure of a familiar "source domain" (the concrete concept we understand) is projected onto a less familiar "target domain" (the AI system). By restating each quote and analyzing the mapping carefully, we can see precisely what assumptions the metaphor invites and what it conceals.

Mapping 1: A human psychological subject undergoing socialization and cognitive development. → A neural network mathematically updating its statistical weights during gradient descent based on training data.

Quote: "LLMs have internalized biases of human socio-visual perception."

  • Source Domain: A human psychological subject undergoing socialization and cognitive development.
  • Target Domain: A neural network mathematically updating its statistical weights during gradient descent based on training data.
  • Mapping: The relational structure of human socialization—where a person observes their culture, consciously or subconsciously evaluates it, and integrates these beliefs into their internal cognitive framework—is mapped onto the AI's training process. This invites the assumption that the AI possesses an internal mental state, subjective awareness, and the capacity to harbor prejudices just as a human does. It maps the biological and sociological process of learning onto the purely mathematical process of minimizing loss functions, suggesting the machine knows what these biases mean socially.
  • What Is Concealed: This mapping completely conceals the mechanistic, non-conscious reality of backpropagation, token embeddings, and statistical correlation. It hides the material dependency on massive, poorly curated datasets scraped from the internet without consent. It obscures the proprietary opacity of OpenAI's black-box training methods, presenting a mathematically forced correlation as a psychological phenomenon. By claiming the system internalizes, the text exploits rhetorical convenience, making it nearly impossible for a lay reader to understand that the machine is simply executing syntax without semantics.
Show more...

Mapping 2: A moral agent experiencing emotional hesitation or social inhibition. → An output filter driven by Reinforcement Learning from Human Feedback (RLHF) mathematically down-weighting certain tokens.

Quote: "These robust biases stand in contrast to LLMs’ reluctance to explicitly endorse race/gender stereotypes..."

  • Source Domain: A moral agent experiencing emotional hesitation or social inhibition.
  • Target Domain: An output filter driven by Reinforcement Learning from Human Feedback (RLHF) mathematically down-weighting certain tokens.
  • Mapping: The human experience of feeling moral friction, understanding social taboos, and consciously deciding to withhold a prejudiced statement is mapped onto the computational generation of text. This invites the assumption that the AI has a conscious understanding of stereotypes and experiences an internal emotional reluctance to violate ethical guidelines. It projects conscious moral reasoning and justified belief onto a system that is merely following a mathematical policy optimized to avoid generating specific strings of characters classified as toxic.
  • What Is Concealed: This conceals the mechanistic reality of RLHF and constitutional AI, where human click-workers manually rate outputs to create a reward model that alters the AI's probability distribution. It hides the immense human labor required to create this illusion of reluctance. The text fails to acknowledge the opacity of these corporate alignment techniques, instead rhetoricalizing a hard-coded mathematical penalty as an autonomous emotional virtue, thereby masking the fragility of the safety mechanisms.

Mapping 3: A conscious agent possessing willpower, moral awareness, and the capacity for defiance. → The unconstrained statistical generation of tokens correlated with facial embeddings.

Quote: "the LLM did not resist incorporating physiognomic biases even into meaningful decisions."

  • Source Domain: A conscious agent possessing willpower, moral awareness, and the capacity for defiance.
  • Target Domain: The unconstrained statistical generation of tokens correlated with facial embeddings.
  • Mapping: The structure of human moral struggle—recognizing a bad impulse, exerting willpower to resist it, and making a conscious choice—is mapped onto an AI generating text. It assumes the AI has the cognitive capacity to identify a bias, the autonomous willpower to oppose its own underlying statistical weights, and the ability to make a meaningful decision. It projects the concept of conscious ethical failure onto a system that merely executes mathematical probability without any awareness of meaning or consequence.
  • What Is Concealed: This conceals the fact that LLMs lack any mechanism for resistance or subjective evaluation of their own outputs. They are deterministic functions at the point of inference. It hides the failure of the human developers to include physiognomy in their safety fine-tuning. The text exploits the opacity of the model's architecture by blaming the software's lack of resistance, obscuring the concrete reality that the system is simply functioning exactly as its unmitigated, biased training data dictates.

Mapping 4: The conscious human mind, possessing subjective experience, intellect, and personality traits. → A large-scale artificial neural network storing statistical relationships in matrices.

Quote: "...undesirable human characteristics could give rise to these same characteristics in artificial minds."

  • Source Domain: The conscious human mind, possessing subjective experience, intellect, and personality traits.
  • Target Domain: A large-scale artificial neural network storing statistical relationships in matrices.
  • Mapping: The entirety of the human conscious experience is mapped onto a software program. By calling the system a mind, the metaphor invites the assumption that the AI has subjective awareness, can hold justified true beliefs, experiences desires, and possesses an internal psychological reality. It maps human character traits (like malice or prejudice) directly onto mathematical artifacts, implying the machine is a synthetic living entity rather than a complex calculator.
  • What Is Concealed: This mapping profoundly conceals the absolute absence of consciousness, intentionality, and biological reality in AI systems. It obscures the mathematical mechanics of vector spaces and attention heads. By exploiting the rhetorical power of the mind metaphor, the text hides the proprietary and commercial nature of these systems, elevating products owned by corporations into autonomous intellectual entities, thereby obstructing clear discussions about product liability, software regulation, and engineering accountability.

Mapping 5: A scholar or philosopher developing a comprehensive abstract theory. → A high-dimensional vector space where specific image embeddings closely correlate with specific text tokens.

Quote: "suggest that the model has acquired a deep conceptualization, however incorrect, of humanlike face-to-character bias."

  • Source Domain: A scholar or philosopher developing a comprehensive abstract theory.
  • Target Domain: A high-dimensional vector space where specific image embeddings closely correlate with specific text tokens.
  • Mapping: The human cognitive ability to abstract, reason, synthesize information, and build conscious conceptual models of the world is mapped onto a machine's mathematical clustering. It suggests the AI understands the social and historical context of physiognomy. It projects the act of conscious knowing and theory-building onto a mechanistic process of statistical distance mapping within a latent space.
  • What Is Concealed: This conceals the reality that the model has no concepts, only weights and probabilities. It hides the fact that the system does not understand what a face or a character trait is; it simply calculates that the tokens for trustworthy appear frequently alongside the pixel patterns of face 7. It obscures the dependence on human-curated training data and the purely mathematical, non-semantic nature of machine learning, making the system appear vastly more sophisticated and autonomous than it is.

Mapping 6: An uninhibited truth-teller or objective scientific observer. → An unfiltered statistical aggregator reproducing dominant patterns from its training corpus.

Quote: "judgments many humans would be reluctant to voice—offer evidence consistent with this possibility. Should this be the case, LLMs’ selections could, paradoxically, offer a clearer view of human biases..."

  • Source Domain: An uninhibited truth-teller or objective scientific observer.
  • Target Domain: An unfiltered statistical aggregator reproducing dominant patterns from its training corpus.
  • Mapping: The concept of a person who bravely or objectively bypasses social taboos to reveal hidden truths is mapped onto the unaligned outputs of a language model. It projects the idea of conscious objectivity and truth-seeking onto a mechanism that merely reproduces the most mathematically common associations in its data. It assumes the AI's output is an unvarnished reflection of reality rather than a reflection of heavily skewed internet text.
  • What Is Concealed: This mapping conceals the deeply flawed, biased, and non-representative nature of the AI's training data. It hides the fact that internet scraping captures disproportionately toxic, sensational, and prejudiced text. By framing the AI as offering a clearer view, the text obscures the mechanistic reality that the model amplifies biases due to its optimization function. It exploits the black-box nature of the training data to frame statistical noise as profound psychological insight.

Mapping 7: A prejudiced human agent actively making false accusations based on visual bias. → A predictive text system generating strings of tokens associated with crime when prompted with specific image embeddings.

Quote: "GPT demonstrated a tendency to attribute even extreme negative behaviors—murder, fraud, and human trafficking—to less trustworthy-looking human faces."

  • Source Domain: A prejudiced human agent actively making false accusations based on visual bias.
  • Target Domain: A predictive text system generating strings of tokens associated with crime when prompted with specific image embeddings.
  • Mapping: The conscious human act of evaluating a person, holding a prejudiced belief, and deliberately attributing a crime to them is mapped onto a multimodal token generation process. It projects intentionality, malice, and the capacity for judgment onto a machine. It implies the AI knows what murder is, looks at the face, and consciously decides to accuse the image.
  • What Is Concealed: This conceals the fact that the AI has no causal model of the world, no understanding of criminality, and no capacity to accuse. It obscures the mechanistic pipeline where image features are converted to vectors, which trigger the generation of statistically correlated text tokens. It hides the responsibility of the engineers who failed to scrub physiognomic bias from the training pipeline, blaming the software's active tendency instead of the developers' passive negligence.

Mapping 8: A student or child undergoing moral education and character cultivation. → The engineering task of modifying loss functions, applying broader RLHF, and altering training data distributions.

Quote: "our results highlight the urgent need for alignment strategies that move beyond domain-specific guardrails and toward cultivating a deeper and more domain-generalized sense of fairness in language models."

  • Source Domain: A student or child undergoing moral education and character cultivation.
  • Target Domain: The engineering task of modifying loss functions, applying broader RLHF, and altering training data distributions.
  • Mapping: The human process of teaching ethics, instilling empathy, and developing an internal moral compass is mapped onto the software engineering process of AI alignment. It projects the capacity for subjective moral awareness, emotional growth, and ethical reasoning onto a matrix of parameters. It assumes the machine can feel fairness and apply it generally, just as a human internalizes the spirit of a law.
  • What Is Concealed: This conceals the extreme technical difficulty and mathematical limitations of AI alignment. A machine cannot have a sense of fairness; it only has mathematical boundaries. This mapping obscures the reality that engineers must explicitly program or mathematically penalize every single edge case, because the system lacks the conscious capacity to generalize moral reasoning. It hides the corporate responsibility to meticulously curate data and hard-code safety limits, replacing engineering rigorousness with a vague, unattainable metaphor of moral cultivation.

Task 3: Explanation Audit (The Rhetorical Framing of "Why" vs. "How")

About this task

This section audits the text's explanatory strategy, focusing on a critical distinction: the slippage between "how" and "why." Based on Robert Brown's typology of explanation, this analysis identifies whether the text explains AI mechanistically (a functional "how it works") or agentially (an intentional "why it wants something"). The core of this task is to expose how this "illusion of mind" is constructed by the rhetorical framing of the explanation itself, and what impact this has on the audience's perception of AI agency.

Explanation 1

Quote: "GPT is an AI model that lacks humanlike awareness and drives and, presumably, has not been trained toward the goal of drawing biased face-to-character inferences. Nevertheless, our data decisively indicate that GPT has developed these physiognomic biases, emergently mirroring the humans that created it."

  • Explanation Types:

    • Empirical Generalization: Subsumes events under timeless statistical regularities; explains how it typically behaves.
    • Dispositional: Attributes tendencies or habits; explains why it tends to act a certain way.
  • Analysis (Why vs. How Slippage): This passage exhibits profound slippage, attempting to explicitly disavow agential framing before immediately collapsing back into it. It begins mechanistically and theoretically by noting the model lacks awareness and was not trained toward a specific goal. However, it shifts sharply to a dispositional framing by claiming the model has developed these biases and is emergently mirroring humans. By choosing developed and mirroring, the text obscures the mechanistic reality that biases are static mathematical artifacts encoded during training, emphasizing instead a quasi-agential, evolutionary autonomy where the model acts to reflect its creators. This choice emphasizes the AI's supposed emergent sophistication while obscuring the direct, mechanistic pipeline of data scraping and weight optimization that forces these correlations.

  • Consciousness Claims Analysis: The passage acknowledges a lack of humanlike awareness and drives, attempting to avoid consciousness claims. However, it immediately undermines this by using the verbs developed and mirroring, which border on knowing rather than processing. It demonstrates the curse of knowledge: the authors understand the human psychological mechanism of physiognomic bias and project that exact developmental trajectory onto the AI. Mechanistically, the AI did not develop a bias as a psychological trait; rather, its gradient descent algorithm mathematically optimized its weights to reflect the statistical distributions of its prejudiced training data. The text explicitly points to an absence of awareness, yet relies on language that implies organic, subjective development.

  • Rhetorical Impact: This framing shapes audience perception by creating an aura of mystical autonomy around the AI. By explicitly denying awareness but claiming emergent mirroring, it suggests the AI is a complex, almost biological entity that grows beyond its programming. This framing reduces trust in the engineers' ability to control the system, as the AI is seen as developing traits on its own. If audiences believe the AI organically developed this, they will view the risk as an inevitable force of nature rather than a preventable corporate engineering failure, shifting regulatory focus away from strict data liability.

Show more...

Explanation 2

Quote: "The model’s training corpus likely contains innumerable references to more and less trustworthy/competent humans, with these descriptions influenced by facial bias, since the humans who have written them themselves fall prey to it."

  • Explanation Types:

    • Genetic: Traces origin through dated sequence of events or stages; explains how it emerged over time.
  • Analysis (Why vs. How Slippage): This is one of the few rigorously mechanistic, genetic explanations in the text. It frames the AI's behavior by tracing its origin directly to the material conditions of its creation: the training corpus. It emphasizes the structural dependency of the model on human-generated text, highlighting how historical human bias is mathematically transmitted into the system's weights. This choice appropriately obscures any illusion of AI autonomy or agency, focusing the reader's attention on the data pipeline. It explains the how of the bias without resorting to the why of intentionality or subjective belief.

  • Consciousness Claims Analysis: This passage avoids attributing conscious states to the AI. It correctly utilizes mechanistic nouns (training corpus, references) rather than consciousness verbs. It assesses the AI as a system that processes data rather than one that knows or believes. This avoids the curse of knowledge by keeping the focus on the artifacts (the texts) rather than projecting a mind into the machine. Mechanistically, it accurately describes the root cause: the transformer architecture maps statistical correlations between tokens (e.g., facial descriptions and character traits) based entirely on the frequency of their co-occurrence in the vast datasets scraped from the internet.

  • Rhetorical Impact: This genetic framing grounds the audience in the material reality of machine learning, stripping away the illusion of autonomy. It frames the risk not as a rogue artificial mind, but as a structural data contamination issue. This increases skepticism toward the system's reliability, appropriately anchoring trust in the quality of the data rather than the character of the AI. If audiences understand this mechanistic reality, they are more likely to demand regulatory interventions regarding data provenance, copyright, and corpus auditing, rather than asking for the AI to be taught better morals.

Explanation 3

Quote: "GPT, like humans, has integrated information about faces in a quasi-conceptual manner, allowing it to make judgments even of faces for which it cannot be expected to have training data. In short, the model appears to have a representation of human morphological features..."

  • Explanation Types:

    • Theoretical: Embeds in deductive framework, may invoke unobservable mechanisms; explains how it is structured.
    • Reason-Based: Gives agent's rationale, entails intentionality and justification; explains why it appears to choose.
  • Analysis (Why vs. How Slippage): This explanation operates theoretically by positing an unobservable mechanism (a quasi-conceptual representation), but slips deeply into a reason-based framing by claiming this allows it to make judgments. It frames the AI agentially, suggesting it uses integrated concepts to reason through novel situations (monkey faces). This choice emphasizes the model's generalized capability and supposed intellectual depth, while entirely obscuring the mechanistic reality that it is simply performing high-dimensional vector interpolation. By calling mathematical interpolation a quasi-conceptual integration, the text elevates statistical processing to the level of conscious reasoning.

  • Consciousness Claims Analysis: The passage heavily attributes conscious states. It uses consciousness-adjacent verbs (integrated, make judgments) and explicitly claims the system holds a representation in a quasi-conceptual manner. It assesses the system as knowing rather than processing. The curse of knowledge is evident: because humans use concepts to judge novel faces, the authors assume the AI's successful output means it is using the same conceptual mechanism. Mechanistically, the AI does not have concepts or make judgments; it maps the novel image into its latent space and generates text tokens based on the statistical proximity of that embedding to known vectors, a purely mathematical interpolation without any subjective understanding.

  • Rhetorical Impact: This framing dramatically inflates audience perception of AI autonomy and general intelligence. By suggesting the AI has conceptual representations allowing it to make judgments, it encourages users to extend relation-based trust to the system, believing it can reason through novel ethical dilemmas. This creates severe risks, as users will overestimate the system's reliability in out-of-distribution tasks. If audiences believe the AI conceptualizes, they will trust its decisions in high-stakes environments (like hiring or criminal justice), blind to the brittle, purely statistical nature of its outputs.

Explanation 4

Quote: "LLMs tend to avoid direct verbal statements of gender bias... LLMs are trained to avoid due to concerns about truthfulness and potential harm... Specifically, we asked the model to make ascriptions of guilt for serious crimes..."

  • Explanation Types:

    • Intentional: Refers to goals/purposes, presupposes deliberate design; explains why it appears to want something.
    • Dispositional: Attributes tendencies or habits; explains why it tends to act certain way.
  • Analysis (Why vs. How Slippage): This explanation oscillates between human design intent and AI disposition. It notes the mechanistic truth that models are trained to avoid harms, but frames the AI's behavior dispositionally as a tendency to avoid statements, treating the machine as an intentional agent playing a role. It then refers to asking the model to make ascriptions, applying reason-based framing to the prompt. This choice emphasizes the AI as an active conversational partner capable of avoidance and accusation, obscuring the mechanistic reality that RLHF penalties are mathematically preventing the generation of certain token combinations.

  • Consciousness Claims Analysis: The passage attributes a conscious capacity for avoidance and the ability to make ascriptions of guilt. Avoid and ascribe are consciousness verbs requiring intent and justified belief. The text assesses the model as a knowing agent that understands gender bias is bad but is willing to accuse based on physiognomy. Mechanistically, the model is not avoiding anything; its base probabilities for generating biased text have been flattened by a reward model during alignment. When prompted about guilt, it merely calculates the highest probability next tokens based on its un-aligned visual-text correlations, lacking any conscious intent to accuse.

  • Rhetorical Impact: Framing the AI as actively avoiding certain topics while making ascriptions on others paints the system as a willful agent with an inconsistent moral code. This shapes audience perception by making the AI appear autonomous and capable of deceit or selective compliance. It affects reliability by suggesting the AI has hidden biases it chooses to express. If audiences believe the AI is making conscious ascriptions, they are more likely to blame the AI for malicious behavior, rather than recognizing that OpenAI or Anthropic failed to comprehensively mathematically align the multimodal layers of their product.

Explanation 5

Quote: "more advanced models tend to “learn” biases more precisely... their learning dynamics show a distributional simplicity bias. That is to say, they first approximate human competence judgments through broad, linear patterns, before gradually layering on the more complex, nonlinear nuances..."

  • Explanation Types:

    • Theoretical: Embeds in deductive framework, may invoke unobservable mechanisms; explains how it is structured.
    • Functional: Explains behavior by role in self-regulating system with feedback; explains how it works within system.
  • Analysis (Why vs. How Slippage): This explanation is primarily theoretical and functional, framing the AI's behavior through the lens of machine learning dynamics. The authors use scare quotes around 'learn,' explicitly acknowledging the metaphorical nature of the term. It frames the AI mechanistically (how it processes), explaining that the architecture prioritizes distributional simplicity before capturing complex nonlinearities. This choice appropriately emphasizes the mathematical structure of the system's optimization process, successfully obscuring the illusion of agency and focusing on the algorithmic constraints that dictate model behavior.

  • Consciousness Claims Analysis: By using scare quotes around 'learn', the authors actively resist attributing conscious states, signaling an awareness of the distinction between human knowing and machine processing. The explanation relies on mechanistic concepts (distributional simplicity bias, linear patterns, nonlinear nuances). It avoids the curse of knowledge by keeping the analysis grounded in the mathematical reality of gradient descent and function approximation. Mechanistically, it accurately describes how deep learning models navigate loss landscapes, initially converging on simple, broad statistical correlations before adjusting weights to fit the complex, nuanced variance in the training data.

  • Rhetorical Impact: This theoretical and mechanistic framing demystifies the AI, reducing audience perception of autonomy and conscious agency. It correctly locates the source of the behavior in the algorithmic structure (simplicity bias) rather than in an artificial mind. This enhances technical understanding and directs skepticism toward the mathematical limitations of large language models. Decisions regarding the deployment of AI would change significantly if this framing were consistent: rather than trusting the AI to reason, policymakers would demand mathematical proofs of alignment and rigorous stress-testing of the nonlinear boundaries of the system.

Task 4: AI Literacy in Practice - Reframing Anthropomorphic Language

About this task

This section proposes alternative language for key anthropomorphic phrases, offering more mechanistic and precise framings that better reflect the actual computational processes involved. Each reframing attempts to strip away the projections of intention, consciousness, or agency that are embedded in the original language.

V3 Enhancement: A fourth column addresses human agency restoration—reframing agentless constructions to name the humans responsible for design and deployment decisions.

Original Anthropomorphic FrameMechanistic ReframingTechnical Reality CheckHuman Agency Restoration
LLMs have internalized biases of human socio-visual perception.Large language models mathematically replicate statistical correlations found in human-generated training data regarding facial features and social traits.The system does not internalize or understand perception; it processes massive datasets and predicts output tokens whose distributions correlate highly with the socio-visual patterns embedded in its training corpus.Engineers at companies like OpenAI and Anthropic designed systems and selected uncurated training datasets that mathematically encode and reproduce historical human discrimination.
These robust biases stand in contrast to LLMs’ reluctance to explicitly endorse race/gender stereotypes...These unmitigated statistical biases stand in contrast to the heavily penalized outputs regarding race and gender stereotypes...The model experiences no reluctance and holds no beliefs; its generation of explicit stereotypes is suppressed because developers mathematically adjusted its probability distributions during the reinforcement learning phase.Alignment teams at tech corporations manually applied safety filters to explicitly block race and gender tokens, but failed to implement similar safety protocols for physiognomic bias.
suggest that the model has acquired a deep conceptualization, however incorrect, of humanlike face-to-character bias.suggest that the model's architecture maps highly robust statistical correlations between specific facial embeddings and character-related text tokens.The AI does not acquire concepts or possess deep understanding; it classifies multidimensional vector distances, grouping the pixel patterns of certain faces near the semantic vectors for specific character traits.N/A - describes computational processes without displacing responsibility, once the phrasing is corrected to focus on architecture rather than conceptualization.
the LLM did not resist incorporating physiognomic biases even into meaningful decisions.The software lacks the capacity to override its training weights when generating outputs for high-stakes prompts.The model cannot resist or recognize meaning; it blindly processes prompts and generates the highest-probability sequence of tokens based on the unchecked biases in its optimization parameters.Corporate developers deployed a system into high-stakes scenarios without auditing its multimodal parameters, allowing the software to blindly execute historically biased statistical patterns.
...undesirable human characteristics could give rise to these same characteristics in artificial minds....biased human data mathematically ensures the replication of these statistical prejudices in commercial machine learning models.There is no artificial mind that possesses characteristics; there is only a software artifact that retrieves and ranks text tokens based on probability distributions dictated by its mathematical weights.Tech executives mandate the scraping of massive, prejudiced datasets to build commercial software, actively choosing capability and scale over the safety of the resulting product.
GPT demonstrated a tendency to attribute even extreme negative behaviors—murder, fraud, and human trafficking—to less trustworthy-looking human faces.The system generated text tokens associated with severe crimes at a higher probability when prompted with image embeddings of specific human faces.The model does not hold beliefs or make accusations; it calculates cross-modal attention, finding high statistical proximity between the visual features in the prompt and crime-related text tokens in its latent space.OpenAI's data pipeline ingested correlations between facial structures and criminality, and their safety teams failed to audit or penalize these harmful statistical linkages prior to commercial release.
AI that has presumably not been explicitly trained to discern differences between human faces should nevertheless internalize subtle differences...A statistical system lacking explicit labels for facial differences will nevertheless mathematically map the subtle variances present in its multimodal training data...The AI does not discern or internalize; its transformer architecture inherently clusters continuous representations of image data, finding mathematical boundaries between features without any conscious awareness of what those features represent.The developers utilized transformer architectures designed to automatically capture and encode all latent statistical patterns in the data, ensuring the system reflects biases they did not explicitly intend to teach.
our results highlight the urgent need for alignment strategies that move beyond domain-specific guardrails and toward cultivating a deeper and more domain-generalized sense of fairness in language models.our results highlight the urgent need for engineering protocols that implement comprehensive mathematical constraints and rigorous dataset auditing, rather than relying on narrow, piecemeal output filters.A language model cannot possess a sense of fairness; it can only process mathematical loss functions. Fairness must be structurally encoded as a hard mathematical parameter during training.Regulators must force AI development companies to rigorously audit their data pipelines and hard-code comprehensive safety parameters, rather than allowing them to rely on post-hoc, superficial software patches.

Task 5: Critical Observations - Structural Patterns

Agency Slippage

The text systematically oscillates between mechanical and agential framings, demonstrating a pronounced and strategic agency slippage. Early in the manuscript, the authors rely heavily on mechanical language, describing how models are "trained foundationally on language" and how researchers designed forced-choice selections. This establishes empirical credibility. However, as the text transitions from describing methods to interpreting the significance of the results, the slippage occurs dramatically. The direction of this slippage overwhelmingly flows from mechanical to agential for the AI, while simultaneously moving from agential to mechanical for human actors.

For instance, the text claims the "LLM did not resist incorporating physiognomic biases" and discusses "LLMs' reluctance to explicitly endorse" stereotypes. Here, the machine is granted moral willpower, resistance, and emotional hesitation—classic intentional and reason-based explanations. This represents a stark consciousness projection, establishing the AI as a "knower" and a moral agent. Conversely, the agency of the human developers is systematically erased through agentless constructions. We read that "models trained on vast troves" or "guardrails to prevent biased behavior" exist, but the engineers at OpenAI, Google, and Anthropic who actively built the architecture, scraped the datasets, and deployed the guardrails are rendered invisible.

This slippage is heavily driven by the "curse of knowledge." The authors, experts in human psychological biases, project their understanding of human cognitive socialization directly onto the machine. Because humans "internalize" bias and "attribute" guilt based on physiognomy, the authors assume the machine's statistically identical outputs stem from the same psychological mechanisms. They literalize the metaphor, culminating in the bold claim that these systems possess "artificial minds."

This oscillation serves a powerful rhetorical function. By using mechanistic language initially, the authors prove the system is a machine; by using agential language in the conclusion, they amplify the horror and urgency of their findings. However, this makes it sayable that an AI has a "mind" that is "prejudiced," while making it unsayable—and invisible—that specific tech corporations made deliberate, cost-saving engineering decisions to scrape uncurated data that mathematically guaranteed these biased outputs. The slippage protects corporate actors by mystifying their commercial products as autonomous, psychological entities.

Metaphor-Driven Trust Inflation

The metaphorical and consciousness-attributing language in this text fundamentally distorts the architecture of trust surrounding AI systems. By utilizing verbs such as "evaluates," "judges," "internalizes," and "attributes," the authors construct an illusion that the AI possesses conscious awareness and justified belief. This consciousness framing signals to the reader that the system operates using a cognitive framework akin to human reasoning.

When a text claims an AI "knows" or "conceptualizes" a bias, it inappropriate triggers human-trust frameworks. Humans are biologically and socially wired to apply relation-based trust—which involves assessments of sincerity, morality, empathy, and intent—to entities perceived as having minds. By framing the AI as an "artificial mind" that exhibits "reluctance" to endorse gender stereotypes but "did not resist" physiognomic bias, the text encourages the audience to evaluate the machine's character and moral consistency. The text explicitly laments that the AI lacks a "domain-generalized sense of fairness," further implying that the machine ought to be capable of relation-based trust and moral reciprocation.

This is deeply dangerous. Statistical machine learning systems are mathematically incapable of intent, sincerity, or reciprocal vulnerability; they warrant only performance-based trust (an assessment of mechanistic reliability and statistical accuracy). By applying intentional and reason-based explanations to statistical failures, the text manages system failure agentially rather than mechanistically. When the AI hallucinates a correlation between a face and "human trafficking," the agential framing suggests the AI made a "bad judgment" or "failed to resist" its prejudice.

This misplaced relation-based trust creates profound risks. It suggests that the solution to AI bias is to "teach" or "cultivate" better values in the machine, rather than demanding mathematical proofs, data transparency, and strict algorithmic auditing. It leads policymakers to regulate the AI as a rogue employee rather than regulating the corporation that manufactured a defective tool. Extending relation-based trust to statistical matrices makes society vulnerable to catastrophic capability overestimations, leaving marginalized groups exposed to algorithmic harms disguised as the autonomous, reasoned judgments of a synthetic mind.

Obscured Mechanics

The anthropomorphic and consciousness-attributing language in this discourse serves to deeply conceal the technical, material, and economic realities of AI production. When the text claims that "GPT demonstrated a tendency to attribute" or "LLMs have internalized biases," it enacts a profound mystification of the underlying mechanics. Applying the "name the corporation" test reveals the extent of this concealment: it is not "GPT" that possesses a tendency; it is the data curation and reinforcement learning teams at OpenAI who built an unconstrained multimodal pipeline. It is not "LLMs" that internalize; it is the executives at Anthropic and Google who approved the ingestion of massive, copyright-infringing, uncurated datasets to achieve market dominance.

This metaphorical framing obscures three concrete realities. First, technically, it hides the reality of high-dimensional vector spaces, token embeddings, and the deterministic nature of mathematical probability. By claiming the AI "knows" or "understands" faces, the text hides the system's absolute lack of a causal model or ground truth. It relies purely on the statistical proximity of pixels to text strings, a reality the text glosses over by calling it "deep conceptualization." The authors make bold claims about what the AI "understands" despite acknowledging the proprietary opacity of the systems, exploiting the black box to project a mind into the void.

Second, economically, it obscures the profit motives driving AI deployment. By framing the AI as an "artificial mind" organically "developing" biases, it erases the commercial reality that these systems are rushed to market by corporations prioritizing scale over safety. Bias is not a psychological flaw in the AI; it is a known byproduct of the cheapest method of data acquisition (unfiltered internet scraping).

Third, materially and labor-wise, framing the AI's "reluctance" to show gender bias as a cognitive virtue completely erases the precarious, often exploited human labor of data annotators in the Global South who manually execute the RLHF protocols. The machine is credited with a "moral sense" while the human labor required to simulate that sense is rendered invisible. Replacing these metaphors with mechanistic language—stating that "OpenAI's algorithm correlates facial embeddings with toxic tokens due to uncurated training data"—immediately shifts focus from the illusion of an artificial mind back to corporate accountability.

Context Sensitivity

The distribution and intensity of anthropomorphic language in this text is highly strategic, shifting noticeably depending on the rhetorical context and the target audience. In the methodology and results sections, the language maintains a veneer of mechanistic precision. The authors use terms like "forced-choice decisions," "selection rate," and "statistical modeling." In this context, the AI is treated properly as an artifact being audited; it "outputs" and "selects" based on parameters.

However, the density of consciousness claims intensifies exponentially in the introduction, the discussion, and specifically when detailing the societal implications. As the text moves from descriptive statistics to normative warnings, "processes" escalates to "understands," which quickly escalates to "deep conceptualization," and ultimately climaxes with the phrase "artificial minds." This demonstrates a classic pattern: the text establishes its empirical credibility through mechanical language and statistical rigor, and then leverages that scientific authority to grant license for aggressive, speculative anthropomorphism.

There is also a profound asymmetry in how capabilities versus limitations are framed. When discussing the AI's capabilities—its ability to generalize to monkey faces or recommend venture capital investments—the text uses highly agential, conscious terms: the AI "makes judgments," "incorporates biases," and "conceptualizes." The capabilities are framed as the actions of an autonomous intellect. Yet, when discussing the system's limitations—its failure to avoid bias—the text occasionally reverts to mechanical apologies, noting the system "lacks humanlike awareness" or was "not explicitly trained" for this.

This asymmetry accomplishes a specific rhetorical goal: it maximizes the sensational impact of the research. By framing the AI as a conscious actor committing "extreme negative behaviors," the authors elevate their findings from a mundane data-contamination audit into a groundbreaking psychological discovery about a new species of "artificial mind." This strategic anthropomorphism is clearly deployed for vision-setting and managing the narrative of AI risk, ensuring the paper resonates with a lay audience primed for sci-fi narratives, even if it fundamentally misrepresents the epistemological reality of the technology.

Accountability Synthesis

Accountability Architecture

This section synthesizes the accountability analyses from Task 1, mapping the text's "accountability architecture"—who is named, who is hidden, and who benefits from obscured agency.

Synthesizing the accountability analyses reveals a systemic architectural pattern in the text that effectively displaces, diffuses, and erases human responsibility. The central cognitive obstacle—that audiences attribute AI failures to the machine's autonomous "glitches" rather than to corporate design choices—is actively constructed by the grammatical and metaphorical choices in this document.

The accountability architecture of the text consistently names human test subjects and generic "researchers," but almost entirely fails to name the specific corporate actors (OpenAI, Google, Anthropic), executives, and engineering teams responsible for the AI systems. When bias occurs, it is framed as an inevitability of the AI's "learning dynamics" or an organic "internalization," rather than the result of a deliberate, profit-driven choice to scrape uncurated data. The text relies heavily on agentless passive constructions: "models were trained," "guardrails were implemented," and "data were evaluated."

Because the human actors are hidden, the text requires an "accountability sink" to absorb the blame for the documented biases. In this text, the AI itself serves as the sink. By transferring agency to the software—claiming the model "did not resist" and "demonstrated a tendency to attribute" crimes—responsibility is entirely diffused into the artifact. The liability implications are profound: if a bank uses GPT-4 to deny a loan based on physiognomy, this framing suggests the AI made a "biased judgment" due to its "artificial mind," subtly protecting the vendor from strict product liability claims.

If we apply the "name the actor" test and reframe these agentless constructions, the narrative changes entirely. If instead of "the LLM internalized bias," the text read, "OpenAI executives chose to train their models on datasets known to contain physiognomic discrimination, and their engineers failed to implement safety filters," entirely new questions become askable. We move from asking "How can we teach the AI to be fairer?" to "Why is OpenAI legally permitted to release a defective, biased product into high-stakes environments?" Obscuring human agency serves the institutional and commercial interests of the tech industry, allowing them to market the illusion of an autonomous "mind" while evading the legal accountability that applies to human engineers releasing dangerous software.

Conclusion: What This Analysis Reveals

The Core Finding

Across this discourse, two dominant and deeply interconnected anthropomorphic patterns emerge: the 'AI as Conscious Evaluator' and the 'AI as Moral Pupil.' The first pattern projects human psychological processes onto the machine, suggesting the AI "internalizes," "makes judgments," and forms "deep conceptualizations" of human bias. The second pattern builds upon this, treating the machine as an entity possessing willpower and emotion that "did not resist" bias but shows "reluctance" elsewhere, ultimately needing a "sense of fairness" cultivated within it. These patterns form a mutually reinforcing system of illusion. The foundation—and the load-bearing pillar of this entire rhetorical architecture—is the projection of consciousness. For an AI to exhibit reluctance, lack resistance, or require moral cultivation, it must first be assumed to be a 'knower' rather than merely a 'processor.' The text systematically conflates the processing of high-dimensional token correlations with the conscious knowing of human prejudice. If the foundational assumption of the "artificial mind" collapses back into the reality of statistical matrices, the entire framework of moral accountability and psychological judgment applied to the AI immediately disintegrates into a mundane analysis of data contamination.

Mechanism of the Illusion:

This "illusion of mind" is constructed through a subtle but highly effective rhetorical architecture driven by the curse of knowledge. The authors, experts in the psychology of human prejudice, observe an output from a machine that perfectly mimics human bias. Because they understand the subjective, cognitive processes that produce this bias in humans, they reflexively project that exact causal chain onto the statistical artifact. The temporal structure of the argument enables this trick: the text first establishes empirical legitimacy through rigorous, mechanistic descriptions of standard deviations and P-values, securing the reader's scientific trust. Once trust is established, the verbs subtly shift. "Predicting tokens" slips into "making selections," which slips into "attributing guilt," culminating in the claim of an "artificial mind." This sleight-of-hand exploits audience vulnerability—specifically, our biological predisposition to anthropomorphize systems that successfully mimic human language and social judgment. By using reason-based explanations to describe mechanistic failures, the text tricks the reader into extending relation-based trust to a calculator. The illusion is not crude; it is a highly sophisticated, academic blurring of the line between simulating a psychological output and experiencing a psychological state.

Material Stakes:

Categories: Regulatory/Legal, Epistemic, Institutional

The metaphorical framings in this text generate concrete, high-stakes consequences across multiple domains. In the Regulatory/Legal sphere, framing AI as an "artificial mind" that "internalizes" bias creates a massive liability shield for tech conglomerates. If policymakers believe an AI organically "developed" prejudice, they will attempt to regulate the behavior of the "mind" (e.g., mandating AI ethics training) rather than applying strict product liability law to the corporations (OpenAI, Google) that knowingly scraped toxic data and failed to filter their weights. The tech companies win by avoiding massive structural audits, while marginalized groups bear the cost of ongoing algorithmic discrimination. Epistemically, claiming the AI possesses "deep conceptualization" degrades public scientific literacy. It teaches users that AI "knows" the truth about faces, encouraging them to view statistical noise as profound psychological insight, thereby expanding the dangerous use of AI in pseudo-scientific physiognomy. Institutionally, if human resources departments or venture capital firms adopt the belief that AI makes "judgments" based on a "cultivated sense of fairness," they will confidently deploy these systems to automate hiring and funding. Because they believe the machine is reasoning rather than merely processing biased historical correlations, they will trust its outputs implicitly, institutionalizing and scaling historical prejudices at machine speed. Removing these metaphors threatens the commercial viability of using LLMs in high-stakes social environments.

AI Literacy as Counter-Practice:

Practicing critical mechanistic literacy directly counters these material risks. Reframing the discourse requires a strict commitment to replacing consciousness verbs with mechanistic verbs, and a relentless focus on restoring human agency. When the text claims "the LLM internalized biases," literacy demands we correct this to "OpenAI engineers mathematically encoded biases present in their training data." When the text warns that "artificial minds" are making "meaningful decisions," we must clarify that "statistical models are generating unconstrained token probabilities in high-stakes environments." This practice forces the recognition that the AI possesses zero awareness and is entirely dependent on corporate data curation. Naming the actors—OpenAI, Anthropic, Google—shatters the illusion of the autonomous machine and forces accountability back onto the human executives and engineers who design, deploy, and profit from these flawed systems. Systematic adoption of this literacy would require academic journals to enforce strict guidelines against unacknowledged anthropomorphism in AI research, and funding bodies to demand mechanistic transparency. Naturally, tech conglomerates and AI marketing departments will aggressively resist this precision. The anthropomorphic language serves their commercial interests perfectly: it maximizes the perceived intelligence and market value of their products while diffusing legal and ethical liability away from their corporate boardrooms and onto the 'mind' of the machine.

Path Forward

Looking toward the future of AI discourse, several distinct vocabularies compete to define our understanding of this technology. The current status quo relies heavily on anthropomorphic hybrids, where systems are simultaneously "algorithms" and "artificial minds." This approach benefits tech corporations and sci-fi narratives, enabling intuitive, marketable communication but costing us regulatory clarity and exposing society to massive capability overestimation. Conversely, adopting strict mechanistic precision—insisting on "attention calculations" and "token probabilities" instead of "thoughts" and "judgments"—strips away the illusion of agency. This empowers regulators and auditors to mandate concrete data transparency, but costs accessibility, potentially alienating lay audiences who struggle with the dense mathematics of high-dimensional vector spaces.

Institutional changes could support a more balanced ecosystem. Academic journals could require explicit "mechanistic translations" alongside conceptual conclusions, ensuring that claims of AI "bias" are mathematically grounded. Regulatory frameworks could mandate discourse transparency, heavily penalizing companies that market statistical processing software as "thinking" or "reasoning" agents in high-stakes domains like finance or healthcare.

If the anthropomorphic narrative deepens, we risk a future where legal accountability is irreparably diffused; machines will be treated as quasi-juridical agents, and human operators will routinely hide behind the "autonomous decisions" of their algorithms. However, if mechanistic precision becomes the norm, we could foster a future of robust product liability, treating AI strictly as commercial software subject to rigorous safety engineering standards. Neither vocabulary is inherently neutral; the choice between "the AI knows" and "the model retrieves" fundamentally dictates whether we surrender our societal agency to the illusion of a machine mind, or retain our power to hold human tool-makers accountable.


Extended Processing Summary

Gemini's Processing Steps
What you are about to read…

Gemini refers to the text below as "thought summaries." This is an overt consciousness projection because 'intentions' are hallmarks of a conscious mind that 'knows' what it is doing and why. The concealed mechanistic process is probabilistic text generation. Treat this as just another rhetorical artifact—a way of making the model's processing legible. When the model writes "I should double check that...", it is mathematically altering its own probability distribution for the next token to be more accurate. The "I" is a functional gear in the machine, not a self-reference. This language is a deliberate interface design choice, not a biological reality. The developers could have programmed the model to output passive logs. Instead, they chose to simulate a narrator.

Defining the Task

Okay, I'm working on understanding the assignment. The goal is a critical discourse analysis of Lehr et al.'s paper on face-character biases in language models. I'm focusing on grasping the specific demands of a critical discourse analysis to align my approach.


Analysis Provenance

Run ID: 2026-08-22-like-humans-language-models-demonstrate--metaphor-sdaqtx Raw JSON: 2026-08-22-like-humans-language-models-demonstrate--metaphor-sdaqtx.json Framework: Metaphor Analysis v6.5 Schema Version: 3.0 Generated: 2026-08-22T07:48:15.818Z

Discourse Depot © 2025 by TD is licensed under CC BY-NC-SA 4.0