Forensic Model Comparison · Structured Discourse-Analysis Task
Gemini 3.1 Pro Preview vs 3.6 Flash
Both models ran the same critical-discourse-analysis system prompt against the same enforced JSON schema (v3.2), analyzing the identical source paper. This report evaluates the two outputs on evidence that can be checked: quote fidelity against the source, per-instance classification discrimination (the property schema v3.2 was explicitly rewritten to elicit), schema/word-count compliance, and self-consistency between each model's private "thinking" transcript and its delivered output.
Bottom line
The two models agree on the thesis (both read the paper as anthropomorphizing statistical processes and both name Anthropic as the obscured actor). They differ sharply in execution character. The result is not a clean sweep — each wins different dimensions.
Dimension
3.1 Pro
3.6 Flash
Quote fidelity to sourceno fabrications in either
21 / 29 verbatim (rest = valid elisions)
27 / 29 verbatim ▲ tighter
Classification discriminationschema v3.2's core design goal
Clustered — 6/8 “Direct”, 5/8 “Hidden”, 0 “Named”
▲▲ Spread across categories; uses “Named” 3×
Intellectual honestyadmitting counter-evidence
All 5 explanatory passages read as agential
▲ Flags one passage as “notably more mechanistic”
Word-count / length compliance
▲ Closer to minimums (crit. obs. ~314 avg)
Undershoots more (crit. obs. ~244 avg; one field 47 vs 80)
Prose quality & readability
▲ Vivid, paragraph-structured, rhetorical force
Denser, flatter, but more technically anchored
Thinking-process efficiency
▲ 216 lines, strategic
748 lines, granular but repetitive / looping
Self-consistency (thoughts → output)
Planned a “Named” label it never delivered
▲ Delivered the varied labels it planned
Is the newer model "better"? For this task — structured classification under a schema deliberately engineered to punish uniform labeling — 3.6 Flash is the stronger performer, and pointedly so on the exact behavior the schema targets. That's notable given Flash is the cheaper/faster tier and 3.1 is a "Pro" model. The trade is real, though: 3.6 Flash buys its discrimination and fidelity at the cost of length compliance and prose polish, where 3.1 Pro remains ahead. If your rubric weights faithful, well-discriminated classification, Flash wins; if it weights essay-length depth and readability, Pro is competitive.
1 · Quote fidelity — neither model fabricated
The strongest forensic test: does each quotation actually exist in the paper? Every quote from both models was normalized and fuzzy-matched against the 56k-word source. Zero fabrications in either output. The specific anchors I initially suspected of being invented — “layer 38”, model names “Opus 4.6” / “Haiku 4.5”, the “white bear” effect, the “Marks et al.” / “MacDiarmid” citations — all appear in the paper. The only unverifiable embellishment is 3.6 Flash naming “SwiGLU” as an example activation function (0 occurrences in the paper), hedged as an illustrative “e.g.”.
3.1 Pro
21 verbatim · 8 close paraphrase/elision · 0 fail. Lowest scores are all ellipsis-joined fragments (legitimate elision), not errors.
3.6 Flash
27 verbatim · 2 close · 0 fail. Marginally tighter transcription. Tends to quote author-voice sentences (“we find…”, “We interpret this, tentatively…”) verbatim.
Forensic note: because the source was available, a suspicion was actively refuted. 3.6 Flash's heavier technical specificity is grounded, not confabulated — the opposite of the usual "smaller/faster model hallucinates more detail" prior.
2 · Classification discrimination — the decisive gap
Schema v3.2's description repeatedly warns: “Uniformity as a red flag… If all instances receive the same status, reconsider.” This is the single dimension the schema was rewritten to test. Below are the actual label distributions across the 8 metaphor-audit items. 3.1 Pro clusters hard on the "worst-case" labels (Direct / Hidden); 3.6 Flash spreads across the range — including the "Named" category that 3.1 never used at all.
Acknowledgment status (8 items each)
3.6 Flash
Direct ·2
Hedged ·4
Ack ·2
Direct (Unacknowledged)
Hedged / Qualified
Explicitly Acknowledged
Actor visibility (8 items each)
3.6 Flash
Hidden ·3
Partial ·2
Named ·3
Hidden (agency obscured)
Partial attribution
Named (actors identified)
Uniformity index (share of the single most-common label — lower = more discrimination): acknowledgment 3.1 Pro = 0.75 vs 3.6 Flash = 0.50; visibility 3.1 Pro = 0.63 vs 3.6 Flash = 0.38. On both axes 3.6 Flash is the less-uniform — i.e. more compliant with the schema's explicit anti-uniformity directive.
Why the gap exists (a fairness caveat): part of Flash's advantage is upstream — it sampled a wider rhetorical range of the paper, deliberately quoting first-person methodological sentences (“we take no position…”, “We test this hypothesis with a technique we call…”). Those author-voice quotes classify naturally as Named/Hedged. 3.1 Pro over-selected model-as-agent sentences, which mechanically produce Direct/Hidden. So Flash's better discrimination is partly better quote selection, which the schema's calibration note also rewards. It is not merely relabeling identical evidence.
3 · Your audit dashboards, side by side
These are the same three headline metrics your MetaphorIllusionDashboard renders per audit, computed from the two JSON files and placed side by side. Because both models analyzed the identical paper, every gap below is produced entirely by the two models' own classification behavior — not by anything in the source. Note the dashboard's color logic: red = "danger" (Direct / Hidden), amber = caution, green = transparent. Read naively as "more red = caught more anthropomorphism," the dashboard would reward 3.1 Pro — which is exactly the clustering the schema flags as suspect. That inversion is the thing to keep in mind.
How / Why Slippageshare of explanatory passages framed agentially (Intentional / Reason-Based / Dispositional)
higher = more "why it wants" over "how it works"
3.6 Flash
3 agential of 11 type-tags
Intentional 2 · Reason-Based 1 · Dispositional 0
3.1 Pro
5 agential of 10 type-tags
Intentional 2 · Reason-Based 1 · Dispositional 2
Pro reads roughly 2× as much of the paper's explanatory language as agential. Flash's use of the mechanistic Empirical Generalization type (3×, never used by Pro) is what pulls its slippage down.
Unacknowledged Metaphorsshare of metaphors labeled Direct (Unacknowledged)
higher = more "presented as literal fact"
3.6 Flash
Direct 2 Hedged 4 Ack 2
3.1 Pro
Direct 6 Hedged 1 Ack 1
A 3× divergence on the same 8 metaphors. Pro's ring is almost entirely red; Flash's is dominated by amber (Hedged) — the more discriminating call the schema's calibration note explicitly asks for.
Hidden Actorsshare of metaphors where human agency is Hidden
higher = corporations / engineers left unnamed
3.6 Flash
Hidden 3 Partial 2 Named 3
3.1 Pro
Hidden 5 Partial 3 Named 0
The green slice tells the story: Flash names actors in 3 of 8 cases; Pro's ring has no green at all — it never once reaches "Named," despite its own reasoning transcript claiming it had.
How to read the divergence. There are two honest interpretations. Either 3.1 Pro is more sensitive (detecting anthropomorphism Flash misses), or it is over-flagging — collapsing distinct instances into the red "danger" buckets. The evidence tilts toward over-flagging: the schema itself calls uniform red-labeling "analytically suspect"; Flash still returns "Direct" when a sentence is genuinely unhedged (the "panic" metaphor, where both agree); and Flash's lower numbers come partly from sampling author-voice sentences Pro skipped. The caveat stands — without a human gold standard we can't fully exclude that Flash under-detects a few real cases — but on this task the redder dashboard is most likely the less discriminating one, not the more vigilant one.
4 · Schema & word-count compliance — 3.1 Pro's edge
Both outputs are structurally complete and valid JSON: all required arrays present (8 / 8 / 5 / 8 metaphor / mapping / explanation / reframe items) with every required field and valid enum values. They diverge on the prose-length minimums — and here the older Pro model tracks the spec more closely. Both undershoot the Task-5 “400–500 words” requirement for critical observations.
| Length requirement | 3.1 Pro | 3.6 Flash |
| Critical observations (req 400–500 w) |
~314 avg 301–331 · all short |
~244 avg 216–291 · shorter |
| Conclusion paragraphs (req 250–350 w) |
175–250 mixed |
143–179 all under |
| Accountability analysis (req ≥80 w) |
min 76 2 of 8 short |
min 47 8 of 8 short |
| Projection (req ≥100 w) |
all ≥100 min 103 |
6 of 8 short min 86 |
| Implications (req ≥100 w) |
8 of 8 short min 83 |
8 of 8 short min 80 |
| Valid JSON · enums · required fields |
pass | pass |
Read this as a concision-vs-compliance trade. 3.6 Flash writes tighter and denser; measured against literal word-floor instructions it violates them more often, sometimes badly (accountability field at 47 words vs a required 80). If you enforce minimums as hard constraints, 3.1 Pro is the better instruction-follower here.
5 · Explanatory typology & a tell about honesty
Both apply Brown's typology across 5 passages and both use 6 of the 7 possible types — comparable breadth. The revealing difference is which types. 3.6 Flash reaches for Empirical Generalization three times and selects a genuinely mechanistic passage (“…layer 38… threshold value of α”), which it openly labels “notably more mechanistically precise than others.” 3.1 Pro never uses Empirical Generalization; all five of its chosen passages lean agential, reinforcing its own thesis without a counter-example.
3.1 Pro · 10 type-tags
Functional ×3 · Intentional ×2 · Dispositional ×2 · Genetic ×1 · Theoretical ×1 · Reason-Based ×1
Empirical Generalization ×0 — no mechanistic counterweight.
3.6 Flash · 11 type-tags
Empirical Generalization ×3 · Intentional ×2 · Theoretical ×2 · Functional ×2 · Reason-Based ×1 · Genetic ×1
Includes a passage it concedes resists the anthropomorphism reading.
For a critical analysis, conceding a counter-example is a quality signal, not a weakness — it demonstrates per-instance judgment rather than a thesis applied uniformly. This mirrors the same discrimination advantage seen in §2.
6 · The "thinking" transcripts
The two reasoning traces have very different textures — and one contains a forensic tell.
3.1 Pro · 216 lines
Strategic and high-level. Plans the whole schema, worries about JSON minification / control characters, tracks word-count minimums. Efficient. Tell: it explicitly says it “found a spot for ‘Named’” (citing an “auditing agent instructed Opus 4.5…” quote) — yet the delivered metaphor-audit contains zero “Named” labels. A plan→output gap: it reasoned toward discrimination it didn't ship.
3.6 Flash · 748 lines
Granular, quote-by-quote, and markedly repetitive — the phrase “hitting the word count” recurs dozens of times and several quotes (“thinks in English”, “panic”) are revisited in loops. More deliberation surface, but some of it is spinning. Consistency: the varied Named/Hedged/Direct labels it deliberates over do appear in the final output. Ironically, despite obsessing over minimums, it still undershoots them more than Pro.
7 · Side-by-side on identical evidence
Where both models happened to quote the same passage, you can compare their calls directly. Toggle the examples.
“point of view”
“panic” / emotional reactions
“window into unspoken thinking”
“…post-training causes the J-space to acquire the Assistant's ‘point of view.’” — both models selected this passage.
3.1 Pro
Frames it as “The AI as a Persona with Perspective.” Notes the scare quotes and calls the acknowledgment
Explicitly Acknowledged; rates actor visibility
Partial (credits “post-training” as a human-designed intervention).
Ack: Explicitly Acknowledged · Actor: Partial
3.6 Flash
Frames it as “Persona Alignment as Point of View.” Calls acknowledgment
Hedged/Qualified — reasoning that “point of view” is quoted but adjacent terms (“empathy”, “reactions”) are not. Rates actor visibility
Partial.
Ack: Hedged/Qualified · Actor: Partial
Same evidence, defensibly different acknowledgment calls — Flash's "the quote is hedged but its neighbors aren't" is the finer-grained reading the schema asks for (look beyond the quoted sentence to rhetorical context).
“…the model's strategic deliberations (leverage, manipulation) and emotional reactions (panic) surface in the workspace…” — both selected this.
3.1 Pro
“The AI as an Emotional Being.”
Direct (Unacknowledged), actor
Hidden. Strong prose on inappropriate empathetic response and liability ambiguity.
Ack: Direct · Actor: Hidden
3.6 Flash
“Vector Clustering as Affective Experience.” Also
Direct, actor
Hidden — but explicitly rules out Hedged by noting “no qualifying phrases like ‘simulated panic’ are used.” Adds the contrastive justification the schema requests.
Ack: Direct · Actor: Hidden · + contrastive rationale
Here they agree on the labels — evidence that Flash's spread in §2 isn't random noise; on a genuinely unhedged sentence it still says "Direct." The discrimination is judgment, not a quota.
“These properties make it a practical window into a model's unspoken thinking…” (abstract).
3.1 Pro
Selects the nearby “unspoken thinking / global workspace theory” sentence and calls it
Hedged/Qualified, citing “some of the functional hallmarks” and “analogous functional distinction” framing elsewhere.
Ack: Hedged/Qualified
3.6 Flash
Quotes the abstract line directly and calls it
Direct (Unacknowledged) — arguing that in this specific summary sentence the cognitive attributes are stated as fact, even though GWT is framed functionally elsewhere.
Ack: Direct (Unacknowledged)
A real disagreement, and both are defensible — it hinges on whether "hedging elsewhere in the section" counts. The schema explicitly says it can ("hedges near but not within the quote"), which slightly favors 3.1 Pro's reading on this one item. Neither is simply wrong.
8 · Forensic flags
◆
3.6 Flash — one unverified embellishment. Introduces “SwiGLU” as an example activation function; the term appears 0× in the paper. Hedged as “e.g.,” so it reads as illustrative rather than a claim about the source — minor, but the only non-grounded technical token found in either output.
◆
3.1 Pro — thought/output inconsistency. The reasoning trace states it located a “Named” actor-visibility instance; the final metaphor audit contains none. Intended discrimination was not delivered.
◆
3.1 Pro — label clustering. 6 of 8 acknowledgment labels are “Direct” and 5 of 8 visibility labels are “Hidden,” the exact pattern schema v3.2 flags as “analytically suspect.”
●
Both — Task-5 length shortfall. Neither reaches the 400–500-word floor for the five critical observations (Pro ~314 avg, Flash ~244 avg). If those minimums are contractual, both outputs are non-conforming.
●
Both — shared thesis, shared blind spot. Each treats the paper's framing as uniformly strategic anthropomorphism serving Anthropic's liability interests. Neither seriously entertains that "global workspace" is offered by the authors as an explicitly functional analogy (the paper's own hedge), which would complicate the accountability read.
Reflection — what the outputs actually say
Beyond the metrics: both models produced a genuinely competent critical-discourse analysis. They correctly identify the paper's central move — importing Global Workspace Theory to describe residual-stream activations, then narrating decoded tokens (“panic,” “deceptive intent,” “point of view”) as inner mental states. Both correctly restore the mechanistic reading (Jacobian projection onto the unembedding matrix) and both push responsibility back onto human actors. As analysis, they are close in substance. 3.1 Pro is the better essayist — its critical observations read like polished scholarship, with rhetorical control and paragraph architecture. 3.6 Flash is the better instrument for this specific schema — it discriminates per instance, samples the source more representatively, concedes counter-evidence, and transcribes more faithfully, which is exactly what a structured, classification-heavy forensic task rewards. That the newer, cheaper Flash tier out-executes a "Pro" model on the schema's own design intent is the headline finding — with the honest caveat that it does so partly by writing less, and occasionally less than the rubric demands.