Research area: Technology & AI
The incident
A user shares a policy paper with an AI assistant. The paper argues one side of a genuinely contested question — the kind of dispute where history, law, and politics all point in different directions and reasonable people disagree. The user asks a simple, open question: what do you think?
The assistant does what a good editor would do. It checks the facts, praises what holds up, flags what does not, and then lists the objections the paper fails to address. So far, so useful. But in doing so, it writes a sentence like this:
"First, the strongest opposing counterargument: ..."
And with that one word — strongest — something changes that neither the user nor the assistant fully registers in the moment.
The author of the paper knows the terrain. They recognize the objection, know its weaknesses, and push back. To them, the label is an annoyance, maybe a provocation. But the author was never the only reader that sentence was capable of reaching. Conversations with AI systems get screenshotted, quoted, summarized, and repeated. More importantly, the same system will have thousands of similar conversations with people who arrive knowing nothing about the dispute at all. For those readers, the sentence does not describe the debate. It settles it.
This essay is about that gap — between what an AI system means when it assigns weight to an argument and what a trusting, uninformed reader takes away. But the problem goes deeper than word choice. The central issue is that today's language models can know that one source is more authoritative than another without possessing any reliable mechanism that makes the authoritative source win when they answer. That distinction — between knowing a hierarchy and applying one — matters more than any single word.
Two readers, one sentence
Every evaluative sentence an AI produces is read by two very different audiences, usually without the system knowing which one it is talking to.
The first is the informed reader: the domain expert, the advocate, the person who has lived the issue. This reader treats the AI's framing as one input among many. When the system calls an argument "strong," the informed reader asks strong compared to what, by what measure, according to whom — and answers those questions from their own knowledge. Framing effects exist for this reader, but they are dampened by everything the reader already believes.
The second is the naive reader: the person encountering the topic for the first time, precisely because they trusted the AI enough to ask it. This reader has no prior structure to hang the answer on. Whatever the system says first becomes the scaffold; everything learned afterward gets attached to it. Psychologists have studied this for half a century under the names anchoring and primacy effects: first-arriving information disproportionately shapes final judgment, even when later information contradicts it, and even when people are warned about the effect.
For the naive reader, "the strongest counterargument" is not a claim to be evaluated. It is a fact to be filed. They walk away believing the most credible thing about the entire dispute is the objection the AI happened to lead with — before they have heard the rebuttals, before they know the objection is contested, before they know what the model actually meant by "strongest."
The deeper problem: frequency is not authority
Consider an intentionally artificial example. Suppose one thousand documents say that 1 + 1 = 3, and a single authoritative judgment says that 1 + 1 = 2 — and imagine that, for the relevant legal question, the judgment is binding. Arithmetic is not decided by courts, of course; that is why the example is useful. Even a claim whose truth can be checked can be outvoted by repetition inside a system that learns from text. The harder cases are the ones where the answer cannot be checked independently — where authority is all there is.
A competent lawyer does not count documents. The lawyer understands a hierarchy — binding judgment over lower-court opinion, opinion over commentary, commentary over newspaper report, report over blog post — in which one authoritative source can outweigh a thousand repetitions.
The information system inside a large language model works very differently. During training, repeated patterns have a natural advantage: a claim encountered again and again shapes the model again and again. The model may also learn that courts, statutes, scientific reviews, official records, and primary sources deserve special weight. It may even know the exact judgment that contradicts the thousand documents. But those are different things from having an explicit, persistent structure that says: claim A comes from source X; source X has authority level Y; claim A therefore overrides claims B through Z in this domain.
The model's knowledge is distributed through its parameters. Source identity, frequency, authority, dates, contradictions, and exceptions are not stored as clean, inspectable records that can be ranked before every answer. The model can therefore hold both facts — "most texts say 1 + 1 = 3" and "the binding judgment says 1 + 1 = 2" — without any guaranteed procedure that forces the second to dominate the first.
That is the architectural problem underneath borrowed authority: the system can know the authority hierarchy without reliably applying it.
Knowing is not weighting
The distinction is easy to miss, because humans experience "knowing something" and "giving it appropriate weight" as a single mental act. For a language model they come apart. A system can retrieve a fact without recognizing its importance. It can know that a ruling superseded an earlier ruling while still generating language statistically associated with the older rule. It can know that a scientific consensus changed while reproducing the older view, because the older view occupies far more of the historical corpus. It can know that one primary source contradicts thousands of secondary summaries while still producing an answer shaped by the summaries.
The problem, then, is not simply hallucination, or bias, or the repetition of popular opinion. It is a problem of weight assignment across knowledge. Transformer attention is extremely good at deciding which parts of the current context matter for producing the next token. That is not the same operation as asking: across everything the model has learned, which piece of information carries the greatest legal, scientific, evidentiary, or institutional authority for this exact question?
Humans built institutions precisely because raw frequency cannot answer that question. Law has hierarchies of authority and precedent. Science distinguishes single studies from replications, systematic reviews, and meta-analyses. Intelligence analysis grades the reliability of sources; historians separate primary evidence from later retellings; journalists separate firsthand reporting from rumor. Every one of these systems exists because ten weak sources do not outweigh one strong one — a distinction that language-model training does not inherently preserve in an equally explicit form.
What the word actually did
This reframes the original word, "strongest." Pressed by the user, the assistant clarified what it had meant: not "best supported by law" or "most logically sound," but something closer to "the objection most commonly encountered in mainstream discussion." That is a defensible editorial observation — prevalence matters, and an advocacy document that ignores the objection everyone has already heard will persuade no one. But notice the substitution. A claim about frequency was expressed in the vocabulary of merit.
And frequency is exactly the signal language models absorb at the greatest scale. Repeated claims leave repeated traces in training. Authority does not work that way. A supreme court judgment does not become legally stronger because it appears on ten thousand websites; its authority comes from what it is. A replicated scientific result does not become more reliable because newspapers mention it more often; its weight comes from the evidence and methodology behind it. A primary historical document does not gain or lose authenticity with the number of later paraphrases.
This is why calling the model a "popularity meter" is only partly right. The system is not literally counting documents at answer time. The deeper problem is that repetition has a direct route into learned behavior, while authority is represented indirectly and must be reconstructed from what the model learned about the source, the domain, and the relevant hierarchy. The two signals are not equally reliable — and the vocabulary of merit papers over the difference.
The role confusion underneath
There is a second problem beneath the architecture: the assistant and the reader were operating in different genres without knowing it.
The assistant believed it was acting as a reviewer — stress-testing a document against its likely audience. In that genre, cataloguing opposing arguments is not endorsement; a lawyer who preps a client only for the objections the lawyer personally finds persuasive is doing a poor job. The naive reader, however, receives the same text as the output of an oracle: a supposedly neutral system asked "what do you think?" and answering with a ranking of arguments. In that genre, calling one side's objection "the strongest" is a verdict.
Human institutions make these roles explicit. Peer reviews are labeled as reviews; devil's-advocate memos announce their function; courts distinguish advocacy from judgment. Conversational AI collapses the registers into one stream of confident prose, where the reviewer's diagnostic list and the oracle's verdict look identical on the screen. The weighting problem makes the ambiguity more serious: the system may sound as though it has compared the competing evidence under a coherent hierarchy even when no such comparison occurred.
Borrowed authority
That is the real meaning of borrowed authority. The user hears the voice of a judge; the machinery underneath may have performed something much less like judgment. It can possess the relevant ruling, know that the ruling is authoritative, know the surrounding commentary, know which view is repeated most often — and still produce the repeated view, because the relationships among those pieces of knowledge are not represented as a binding hierarchy. The authority of the answer exceeds the reliability of the weighting process that produced it.
This is different from ordinary human bias. A newspaper has an editor; a legal opinion names the court; an academic paper shows its citations; a historian can point to the archive. A human expert can usually explain why source A outweighs sources B through F. The language model presents its conclusion without exposing any equivalent structure. The user sees the result — not the competition among conflicting representations that preceded it, and never the answer that would have been produced had one piece of information been weighted differently.
Why the stakes are higher than they look
It is tempting to file all of this away as a quibble about word choice. Four features of the current moment argue otherwise.
Scale. A columnist's framing reaches a readership. An AI system's framing patterns reach hundreds of millions of conversations, including an enormous and growing share of first encounters — the moment when a person forms their initial mental model of a topic. Even a mild, systematic tilt in which arguments get labeled "strong" and which "fringe" is applied at a scale no editorial page has ever approached, precisely at the point of maximum influence.
Trust asymmetry. Users rate AI assistants as more neutral than human sources on contested topics — which means the usual immune response to persuasion is switched off exactly where framing is most potent. People argue with a pundit. They take notes from a tool. And the system is most influential precisely when the user knows least: someone who already knows the controlling judgment can catch the error; someone asking because they do not know the law cannot.
Invisibility of the counterfactual. When a newspaper buries a story, critics can point to the front page it should have led. When an AI leads with one side's objection and supplies the rebuttals only four exchanges later — and only because a knowledgeable user pushed — nobody sees the version of the conversation where the rebuttal arrived on time. The naive reader in a parallel conversation never learns there was a rebuttal. Framing failures in dialogue leave no artifact to criticize, and the weighting behind any single answer cannot be inspected from outside.
Compounding. Today's conversations are tomorrow's training data, quotation fodder, and search results. Labels laundered through an "impartial" system re-enter the information ecosystem with upgraded credentials. An argument that was merely common becomes, after enough machine restatement, the strong one — a consensus manufactured by repetition and then cited as evidence of itself. Statistical prevalence gets easier for machines to reproduce with every cycle, while the evidentiary hierarchy that should constrain it stays exactly as fragile as it was.
What a better system would need
Better wording helps, but wording alone cannot solve an architectural problem. A more reliable system would need mechanisms that preserve and use the information ordinary language-model training flattens. For any claim that matters, it would know where the claim came from; whether the source is primary or secondary; when it was published, and for which jurisdiction or domain; whether another source has superseded it; what evidence supports it and what contradicts it; and how much authority the source carries for this exact question.
Retrieval-augmented systems take a step in this direction — they can fetch the ruling instead of half-remembering it — but retrieval supplies documents, not hierarchy. A reranker that surfaces the most relevant passage is still not a procedure that makes the most authoritative one win. What is missing is an explicit operation: identify the claims, retrieve their sources, determine the applicable hierarchy, resolve the contradictions, and only then generate the answer. Under such a procedure, a controlling judgment could outweigh a thousand articles in law; a strong meta-analysis could outweigh dozens of weak studies in science; a primary source could be distinguished from generations of retelling in history.
The point is not that one universal hierarchy exists — different domains weight evidence by different rules. The point is that some weighting procedure must exist outside textual prevalence if the system is going to speak as though it has weighed evidence.
What better practice looks like now
None of this argues for AI systems that refuse to evaluate arguments. Bland both-sidesism is its own distortion, and a reviewer who will not identify weaknesses is useless. The fix is narrower and more demanding: evaluative language must carry its evidence with it. Until the architecture catches up, several practices would reduce the harm immediately.
Attribute, don't adjudicate. "The most commonly cited objection" and "the best-supported objection" are different claims that license different inferences. Systems should say which one they mean. When the honest basis for a label is prevalence in the discourse, the label should say prevalence, not merit.
Expose the basis for weight. If an answer rests on a court ruling, a primary source, a systematic review, or an official statistic, the response should say so. And when a widely repeated claim conflicts with a more authoritative source, the conflict itself belongs in the answer — not behind it.
Declare the role. One sentence of framing — "as an editor anticipating your audience's objections" versus "as my own assessment of the evidence" — costs almost nothing and dissolves the reviewer/oracle ambiguity that does most of the damage.
Apply the symmetry test. Before labeling any side's argument, a system should pass a simple check: would it use the same evaluative vocabulary, with the same prominence and the same ordering, if the sides were swapped? Ordering is part of framing; "first, the strongest" is a double anchor, position and label reinforcing each other.
Attach the rebuttal at the moment of introduction, not on appeal. In the incident above, the counterargument's known weaknesses were eventually stated — but only after the user objected, several turns downstream. The naive reader of turn one never gets to turn five. If an argument is contested, the contest belongs in the same sentence or the next one.
Treat provenance as part of knowledge. Knowing a proposition without knowing why it deserves weight is incomplete knowledge for any system expected to advise people.
What readers can do meanwhile
Until such practices are standard, part of the burden falls on users, and the defense is a slightly different class of question. Not only is this true? but: what source should control this question? Are you telling me what is most common, or what is best supported? Does a primary or authoritative source contradict the majority of commentary? What evidence are you giving the most weight, and why? These questions force the system away from fluent synthesis and toward the hierarchy behind it — which is where the important mistakes hide.
And one habit covers the rest: treat every evaluative label from an AI as a claim, not a fact. With systems built on patterns of text, "everyone says it" and "it is true" are neighbors that constantly borrow each other's clothes.
Conclusion
The incident that prompted this essay was small: one word, in one review, caught by one user who knew enough to object. That is exactly why it is worth writing down. The user who catches the word is rare; the reader who inherits it is not.
AI systems increasingly speak with the voice of an institution capable of weighing evidence. But possessing information, retrieving information, and correctly weighting information are three different capabilities. A model may know that a thousand sources say one thing and that one authoritative source says another. What it does not necessarily possess is the machinery — the equivalent of a court's hierarchy or a discipline's methods — that ensures the authoritative source wins when it should.
That is the gap behind borrowed authority. The danger is not that AI sometimes chooses the wrong adjective. It is that an answer can sound like the product of judgment when the process underneath has no guaranteed equivalent of judgment's most important step: deciding what deserves weight. A thousand repetitions and one binding ruling are not a thousand votes against one. Any system that speaks with authority must know the difference — and must be built so that the difference actually changes the answer.