JudgmentAI and the enterprise
Doubt Against the Illusion of Knowledge
How to harness AI without surrendering judgment
Matteo Gatta11 min read

Dubito ergo sum. Science didn't become powerful because humans got better at being right. It became powerful because we built a culture — and a set of methods — that makes it normal to distrust the first answer, to hunt for disconfirming evidence, and to treat certainty as something you earn rather than something you feel. Karl Popper framed this as falsifiability: what makes a theory "scientific" is not how persuasive it sounds, but whether it exposes itself to tests that could refute it. Richard Feynman said the same thing in plain language: the first principle is that you must not fool yourself — and you are the easiest person to fool.
That is why doubt is not the enemy of knowledge. Doubt is the operating system that keeps knowledge from turning into performance.
And this is precisely what fluent AI threatens — not because people suddenly became less intelligent, but because the cost of sounding right has collapsed. We are entering an era where borrowed certainty is cheap, abundant, and beautifully formatted. This is the intuition behind Epistemia (Quattrociocchi, Capraro, Perc): a state where plausibility substitutes for verification, where the appearance of understanding quietly replaces the work of understanding. The term is new and not yet canonical — that is part of its usefulness. It names this moment's failure mode: when language becomes so frictionless that coherence starts masquerading as truth.
The mechanism is ancient. Human judgment is highly sensitive to familiarity and fluency: when something is easy to process, it often feels more true. The illusory truth effect is not a quirky bias; it is a reliable exploit. Repetition and coherence make claims feel real even when they are not.
Kahneman's System 1 / System 2 framing explains why this matters. System 1 is fast, intuitive, pattern-driven. System 2 is slower, effortful, and capable of genuine checking. Generative AI is, in effect, a System-1 superstimulus: it produces confident language at high speed, with structure, nuance markers, and "sources" that mimic epistemic seriousness. Under time pressure, that fluency becomes a shortcut: people accept the nice answer and move on, mistaking coherence for truth. Epistemia is what happens when that shortcut becomes a habit — then a culture.
Cognitive offloading: normal, useful, and newly risky
Cognitive offloading is not a moral failure; it is normal cognition. We write notes, set reminders, use maps, and consult colleagues precisely because the mind is designed to distribute work. Risko and Gilbert describe offloading as shaped by internal demand and metacognitive judgments about effort. The "Google effect" showed a related shift in the internet era: when people expect future access to information, they tend to remember where to find it rather than the content itself.
LLMs raise the stakes because they don't just retrieve facts; they draft arguments, synthesize positions, and generate plausible explanations. The unit you offload is no longer memory alone. It becomes the connective tissue of reasoning — the part that forces you to notice contradictions, declare assumptions, and earn conclusions.
Recent evidence adds urgency, but it also demands transparency. The widely discussed MIT Media Lab preprint on LLM-assisted writing reports lower measured engagement and weaker carryover when participants later write without tools. The study has also attracted methodological critique (sample size, interpretability and reproducibility questions), so it is best treated as a warning signal, not a verdict. Still, its core risk is hard to dismiss: when the machine does the hard cognitive work for you, you may keep the output while losing the ownership.
At the same time, the empirical picture is not one-directional. Multiple meta-analyses in 2025 report positive effects of generative AI on learning outcomes in many setups — especially when AI is used as a tutor, coach, or feedback layer rather than as a shortcut machine. That bimodality matters. It shifts the argument from fatalism to responsibility: unguided substitution makes decay likely, while scaffolded augmentation can strengthen learning and performance. If Epistemia is the disease, design, pedagogy, and incentives determine the infection rate.
Humans and machines are both bad at prediction — just in different ways
At the beginning of the year, predictions are published as if uncertainty were a minor inconvenience. Taleb's Black Swan argument is brutal and liberating: the world is shaped by rare events, fat tails, and nonlinear shocks — the kinds of things prediction machinery, human or statistical, systematically misses. Humans are bad forecasters not because they lack intelligence, but because they are trapped inside WYSIATI: what you see is all there is. System 1 compresses reality into a story that fits the data at hand, then hands you a feeling of certainty as if it were a result.
But here is the sharper point: human prediction errors are often not innocent. They are entangled with ego, status, tribal incentives, and the psychological comfort of having an answer. Forecasts are not always attempts to be right; they are often performances that protect identity. The point of many predictions is not accuracy — it is control, confidence, and coherence in public.
Machine prediction errors are different. When an LLM gives you a confident but wrong answer, the failure mode is usually poor grounding, not ego. The model has no personal stake. Its mistakes are guiltless: probabilistic completions detached from reality checks. That makes them more forgivable — but also more dangerous in one specific way. Because machine mistakes are not motivated, they can feel like "honest errors," and that can seduce humans into lowering their guard.
The Oracle with No Skin in the Game
In medieval Rome, the Bocca della Verita (The Mouth of Truth) served as a visceral lie detector; legend held that the stone maw would bite the hand of any speaker who uttered a falsehood. It was a literal Price of Knowing (Assurance Cost) paid in flesh.
Generative AI is the inverse: it is a mouth that speaks with total confidence but risks nothing. It has no hand to lose. This guiltless fluency makes machine errors more forgivable, but also more dangerous. Because the machine faces no consequence, it seduces the human into lowering their guard — effectively sticking their hand into the maw and assuming it is made of rubber.
This is the twist: the real epistemic hazard is not that AI lies. It is that humans use AI to launder their own overconfidence. The model provides fluent scaffolding for narratives we already wanted to believe. It gives a System-1 story a System-2 costume — complete with qualifiers, counterarguments, and "sources." The output feels less like a hunch and more like knowledge, even when it is just an elaborated guess.
So the remedy cannot be "be careful." The remedy has to be method — technical, behavioral, institutional — because method is how we make doubt operational at scale.
The Assurance Cost: why "just check it" doesn't work
Technically, grounding mechanisms like Retrieval-Augmented Generation can improve traceability by tying outputs to external documents, and tool-using patterns can reduce error propagation by forcing contact with external reality. Yet neither is a silver bullet. Retrieval can fetch the wrong sources. Citations can become ornamental. Humans can still rubber-stamp answers that feel right.
The deeper fix is governance: workflows where assurance is measurable, repeatable, and accountable.
The biggest obstacle, however, is not technical. It is economic and behavioral: the assurance cost. In a world of "as-soon-as-possible," assurance feels like friction. If an LLM saves two hours of drafting but validating claims costs three hours, the human brain — optimized for efficiency — will default to acceptance, not interrogation. Epistemia becomes rational behavior: not because people don't care, but because the system rewards speed and punishes scrutiny.
This is also why "responsible AI" often stalls at posters and policies. Plenty of institutions endorse audit trails in principle. Far fewer reward the behaviors that make them real: validating sources, documenting uncertainty, recording assumptions, exposing what would change the conclusion. Most workflows still pay for confidence, volume, and pace — exactly what LLMs can counterfeit at scale. When assurance is optional, it becomes a luxury good. And truth becomes a branding exercise. So we have to change posture. Assurance cannot be a passive audit — glancing at citations and assuming correctness because they exist. It must become active interrogation, where the human remains the primary investigator rather than a proofreader.
A simple habit helps: before reading the model's output, sketch the logic you expect, the constraints that must hold, or the failure modes you fear. Then red-team the model instead of asking it to reassure you: ask it to identify hidden assumptions, plausible counterexamples, and what evidence would discriminate between competing explanations. This breaks the single-narrative spell. It restores a scientific stance: conjectures under pressure, not assertions under applause.
And it helps to treat AI outputs explicitly as a mix of facts (check sources), causal claims (check mechanisms and studies), predictions (track calibration over time), and judgments (make assumptions explicit and contestable)— a practical taxonomy that aligns with transparency and accountability guidance like the NIST AI RMF.
This interrogation posture matters because overreliance on automated aids has a known cognitive signature. Research on automation bias shows that when decision aids are imperfect, users commit omission and commission errors — missing what the system missed, or accepting what it suggested — especially under workload. If AI is to remain an amplifier rather than a crutch, workflows must resist automation bias by default, not by willpower.
Learning science adds one more constraint. If AI makes tasks too easy, the brain may stop encoding the underlying logic. Desirable difficulties capture this: certain forms of effort and friction improve long-term learning and transfer, even if they feel slower in the moment. Applied to GenAI, the principle is simple: let the model provide bricks (draft structure, candidate references, alternative formulations), but keep the mortar (connective logic, justification, final judgment) human-owned. If the AI provides the mortar too, structural reasoning weakens over time — not as a dramatic collapse, but as gradual atrophy masked by rising output.
A real-world signal: elite institutions start testing "judgment with AI"
Institutions are beginning to redesign how they evaluate talent. McKinsey, for example, has reportedly piloted a recruitment change in select final-round graduate interviews in the U.S., where some candidates use its internal AI assistant, Lilli, during an interview-style case exercise — and are assessed less on the "answer" than on how well they interrogate the tool, challenge its output, and adapt it to context.
This is not yet a broad cultural shift. It is too recent and too limited to claim that. But itis a signal worth taking seriously because it reveals what is becoming scarce. In an AI-saturated workplace, baseline competence is no longer "can you produce a coherent answer?" The machine can do that. The differentiator is whether you can interrogate, contextualize, and discipline an answer — whether you can make it accountable.
Conclusion
: make verification prestigious — or accept The Illusion of Knowledge as the default Epistemia will not be defeated by telling people to be vigilant. It will be defeated only if leaders and educators make assurance economically viable and socially prestigious. If organizations reward only volume of output, they are subsidizing Epistemia. They are teaching people that fluency is success. The prestige must shift from the final document to the audit trail. In high-stakes contexts, an AI-assisted output should be incomplete without a short Verification Log that records which claims were checked against which primary sources, what assumptions were used, and what remains uncertain — because without that, "assurance" quietly collapses back into vibes. The expected cost of being wrong must become higher than the assurance cost — otherwise speed will keep winning by default. Science progressed because it made doubt a shared discipline, not a private virtue. AI will create value in the same way: not by making us faster at producing convincing text, but by strengthening our cycles of sense, decide, execute, and learn — with verification as the hinge. If we get the design, pedagogy, and incentives right, GenAI can become a genuine cognitive amplifier — supporting learning, inclusion, and productivity without hollowing out judgment. If we don't, it will normalize the most dangerous habit of all: the quiet acceptance of plausible answers as if they were knowledge. Fluent is not the same as true. And the future belongs to the people and institutions who refuse to forget that.
References
- Karl Popper, Conjectures and Refutations: The Growth of Scientific Knowledge (falsifiability).
- Richard P. Feynman, "Cargo Cult Science" (1974; "must not fool yourself...").
- Daniel Kahneman, Thinking, Fast and Slow (System 1 / System 2).
- Hassan & Barber (2021), "The effects of repetition frequency on the illusory truth effect" (processing fluency).
- Risko & Gilbert (2016), "Cognitive Offloading," Trends in Cognitive Sciences.
- Sparrow, Liu & Wegner (2011), "Google Effects on Memory," Science.
- Goddard, Roudsari & Wyatt (2012), "Automation bias: a systematic review..."
- Bjork & Bjork (2011), "Making Things Hard on Yourself, But in a Good Way: Desirable Difficulties."
- Lewis et al. (2020), "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks."
- Yao et al. (2022), "ReAct: Synergizing Reasoning and Acting in Language Models."
- Lin, Hilton & Evans (2021/2022), "TruthfulQA."
- NIST, AI Risk Management Framework (AI RMF 1.0).
- UNESCO, Guidance for generative AI in education and research (UNESCO article page last update: 16 January 2026).
- UNICEF Innocenti, Guidance on AI and Children 3.0 (December 2025, plus checklist/poster materials).
- Quattrociocchi, Capraro, Perc (22 Dec 2025 preprint), "Epistemological Fault Lines Between Human and Artificial Intelligence" ("Epistemia").
- Tian et al. (2025), AI dependence & critical thinking (fatigue mediator; information literacy moderator).
- Ma&Zhong (2025) meta-analysis, Journal of Computer Assisted Learning; plus other 2025 meta-analyses (e.g., Humanities and Social Sciences Communications).
- Kosmyna et al. (2025) arXiv preprint on LLM-assisted writing ("cognitive debt") + subsequent critique/commentary.