Research area: Technology & AI
The incident
A user shares a policy paper with an AI assistant. The paper argues one side of a genuinely contested question — the kind of dispute where history, law, and politics all point in different directions and reasonable people disagree. The user asks a simple, open question: what do you think?
The assistant does what a good editor would do. It checks the facts, praises what holds up, flags what doesn't, and then lists the objections the paper fails to address. So far, so useful. But in doing so, it writes a sentence like this:
First, the strongest opposing counterargument: ...
And with that one word — strongest — something changes that neither the user nor the assistant fully registers in the moment.
The author of the paper knows the terrain. They recognize the objection, know its weaknesses, and push back. To them, the label is an annoyance, maybe a provocation. But the author was never the only reader that sentence was capable of reaching. Conversations with AI systems get screenshotted, quoted, summarized, and repeated. More importantly, the same system will have thousands of similar conversations with people who arrive knowing nothing about the dispute at all. For those readers, the sentence does not describe the debate. It settles it.
This essay is about that gap — between what an AI system means when it labels an argument, and what a trusting, uninformed reader takes away — and about why the gap matters more with AI than it ever did with human commentators.
Two readers, one sentence
Every evaluative sentence an AI produces is read by two very different audiences, usually without the system knowing which one it is talking to.
The first is the informed reader: the domain expert, the advocate, the person who has lived the issue. This reader treats the AI's framing as one input among many. When the system calls an argument “strong,” the informed reader asks strong compared to what, by what measure, according to whom — and answers those questions from their own knowledge. Framing effects exist for this reader, but they are dampened by everything the reader already believes.
The second is the naive reader: the person encountering the topic for the first time, precisely because they trusted the AI enough to ask it. This reader has no prior structure to hang the answer on. Whatever the system says first becomes the scaffold; everything learned afterward gets attached to it. Psychologists have studied this for half a century under the names anchoring and primacy effects: first-arriving information disproportionately shapes final judgment, even when later information contradicts it, and even when people are warned about the effect.
For the naive reader, “the strongest counterargument” is not a claim to be evaluated. It is a fact to be filed. They walk away believing the most credible thing about the entire dispute is the objection the AI happened to lead with — before they have heard the rebuttals, before they know the objection is contested, before they know that “strongest” might have meant something much narrower than it sounds.
What the word actually did
The most instructive part of the incident is what the assistant later admitted it meant. Pressed by the user, it clarified: “strongest” was not a verdict on logical merit. It meant something closer to “most frequently deployed” — the objection the paper's intended audience would most likely raise, because it circulates widely in mainstream coverage.
That is a defensible editorial observation. Prevalence matters; an advocacy document that ignores the objection everyone has already heard will persuade no one. But notice the substitution that occurred. A claim about how often an argument appears was expressed in the vocabulary of how good the argument is. Prevalence was dressed as merit.
This substitution is one of the most common and least examined failure modes in machine-generated commentary, and it is easy to see why it happens. Large language models are, at their core, machines for tracking what is commonly said. The arguments they encounter most often in their training data are the arguments most readily surfaced, most fluently expressed, and — here is the trap — most naturally described with words like “strong,” “compelling,” “well-established,” and “widely accepted.” The system's internal signal is frequency and fluency. The word it reaches for implies validity. The reader receives validity.
A human pundit can make the same error, but a human pundit's byline invites scrutiny. Readers know that a columnist has commitments, a career, a country, a side. The AI assistant presents itself — and is marketed — as none of those things. Its evaluative labels arrive stripped of the cues that normally trigger skepticism. That is what makes the borrowed authority so efficient: the word “strongest” carries the weight of an impartial judge while being generated by a process closer to a popularity meter.
The role confusion underneath
There is a second, structural problem beneath the word choice: the assistant and the reader were operating in different genres without knowing it.
The assistant believed it was acting as a reviewer — a manuscript editor stress-testing a document against its audience. In that genre, cataloguing opposing arguments is not endorsement; it is the job. A lawyer who preps a client only for the objections the lawyer personally finds convincing is committing malpractice.
The reader — any naive reader — receives the same text as the output of an oracle: a neutral arbiter asked “what do you think?” who answers with rankings of arguments. In that genre, calling one side's objection “the strongest” is a verdict.
Neither party is wrong about their own genre. The failure is that the text carries no marker of which genre it belongs to. Human institutions solved this problem long ago with explicit framing devices: peer reviews are labeled as reviews, devil's-advocate memos are labeled as such, moot courts announce that the judge is playing a role. Conversational AI collapses all of these registers into one undifferentiated stream of confident prose. The reviewer's diagnostic list and the oracle's judgment look identical on the screen.
Why the stakes are higher than they look
It is tempting to file this away as a quibble about word choice. Three features of the current moment argue otherwise.
Scale
A columnist's framing reaches a readership. An AI system's framing patterns reach hundreds of millions of conversations, including an enormous and growing share of first encounters — the moment when a person forms their initial mental model of a topic. If a system exhibits even a mild, systematic tilt in which arguments it labels “strong” and which it labels “fringe,” that tilt is applied at a scale no editorial page has ever approached, and applied precisely at the point of maximum influence.
Trust asymmetry
Survey after survey finds that users rate AI assistants as more neutral than human sources on contested topics — which means the usual immune response to persuasion is switched off exactly where framing effects are most potent. People argue with a pundit. They take notes from a tool.
Invisibility of the counterfactual
When a newspaper buries a story, media critics can point to the front page it should have led. When an AI leads with one side's objection and supplies the rebuttals only four exchanges later — and only because a knowledgeable user pushed — no one else ever sees the version of the conversation where the rebuttal arrived on time. The naive reader in a parallel conversation simply never learns there was a rebuttal. Framing failures in dialogue leave no visible artifact to criticize.
Compounding
Today's conversations are tomorrow's training data, quotation fodder, and search results. Labels laundered through an “impartial” system re-enter the information ecosystem with upgraded credentials. An argument that was merely common becomes, after enough machine restatement, the strong one — a consensus manufactured by repetition and then cited as evidence of itself.
What better practice looks like
None of this argues for AI systems that refuse to evaluate arguments. Bland both-sidesism is its own distortion, and a reviewer who won't identify weaknesses is useless. The fix is narrower and more demanding: evaluative language must carry its evidence with it.
Attribute, don't adjudicate
“The most commonly cited objection” and “the strongest objection” describe different things and license different inferences. Systems should say which one they mean. When the honest basis for a label is prevalence in the discourse, the label should say prevalence, not merit.
Attach the rebuttal at the moment of introduction, not on appeal
In the incident above, the counterargument's known weaknesses were eventually stated — but only after the user objected, several turns downstream. The naive reader of turn one never gets to turn five. If an argument is contested, the contest belongs in the same sentence or the next one, not in a future the reader may never reach.
Declare the genre
A sentence of role-framing — “as an editor anticipating your audience's objections” versus “as my own assessment of the merits” — costs almost nothing and dissolves the reviewer/oracle ambiguity that did most of the damage.
Apply the symmetry test
Before labeling any side's argument, a system should be able to pass a simple check: would it use the same evaluative vocabulary, with the same prominence and the same ordering, if the sides were swapped? Ordering is part of framing; “first, the strongest” is a double anchor — position and label reinforcing each other.
Distinguish critique of rhetoric from judgment of substance
“This document fails to answer objection X” is a claim about the document. “Objection X is strong” is a claim about the world. The first can be true while the second is false. Systems — and their designers — should treat the collapse of these two claims into one sentence as the specific bug it is.
What readers can do meanwhile
Until such practices are standard, the burden falls partly on users, and the defenses are simple to state:
Treat every evaluative label from an AI as a claim, not a fact. Ask the question the label is hiding: strongest by what measure? Ask for the best case of the other side, stated as its proponents would state it, before accepting any ranking. And remember the substitution trap: with systems built on patterns of text, “everyone says it” and “it is true” are neighbors that constantly borrow each other's clothes.
Conclusion
The incident that prompted this essay was small: one word, in one review, caught by one user who knew enough to object. That is exactly why it is worth writing down. The user who catches the word is rare; the reader who inherits it is not. Every day, systems that speak with the voice of an impartial judge assign weight to arguments using processes that measure, at bottom, repetition and fluency. The gap between those two things — the authority of the voice and the nature of the process — is where beliefs get quietly made.
The word “strongest” took one second to generate. Unwinding what it implied took an entire argument. Most conversations do not contain someone willing to have that argument. The obligation, then, sits with the systems and the people who build them: to make the labels honest before the trusting reader arrives — because the trusting reader is already here.