{"id":30922,"date":"2026-08-10T04:11:00","date_gmt":"2026-08-09T20:11:00","guid":{"rendered":"https:\/\/csccm.org.cn\/?p=30922"},"modified":"2026-08-10T06:03:26","modified_gmt":"2026-08-09T22:03:26","slug":"nejm%e8%a7%82%e7%82%b9%ef%bc%9a%e4%ba%ba%e5%b7%a5%e6%99%ba%e8%83%bd%e8%83%bd%e5%90%a6%e8%af%b4%e6%88%91%e4%b8%8d%e7%9f%a5%e9%81%93","status":"publish","type":"post","link":"https:\/\/csccm.org.cn\/?p=30922","title":{"rendered":"[NEJM\u89c2\u70b9]\uff1a\u4eba\u5de5\u667a\u80fd\u80fd\u5426\u8bf4\u201c\u6211\u4e0d\u77e5\u9053\u201d"},"content":{"rendered":"\n<p><a href=\"https:\/\/www.nejm.org\/browse\/nejm-article-type\/perspective\">PERSPECTIVE<\/a><\/p>\n\n\n\n<h1 class=\"wp-block-heading\">Can AI Say \u201cI Don\u2019t Know\u201d?<\/h1>\n\n\n\n<h3 class=\"wp-block-heading\">Andrea&nbsp;Sikora,&nbsp;Leo A.&nbsp;Celi,&nbsp;Raja-Elie E.&nbsp;Abdulnour<\/h3>\n\n\n\n<h3 class=\"wp-block-heading\">N Engl J Med&nbsp;2026;394:1873-1875<\/h3>\n\n\n\n<h3 class=\"wp-block-heading\">DOI: 10.1056\/NEJMp2517624<\/h3>\n\n\n\n<blockquote class=\"wp-block-quote\">\n<p>I will not be ashamed to say \u201cI know not.\u201d<\/p>\n\n\n\n<p>\u2014 Hippocratic Oath<\/p>\n<\/blockquote>\n\n\n\n<p>A first-year resident is asked, \u201cWhat explains the patient\u2019s rising creatinine?\u201d She pauses. \u201cI don\u2019t know \u2026.\u201d The team reviews the patient\u2019s medications and risk factors and chooses a cautious plan involving checking drug levels and consulting the nephrology department.<\/p>\n\n\n\n<p>A clinical artificial intelligence (AI) system is presented with the same case. It retrieves relevant articles, but despite conflicting evidence, it produces a confident answer, without flagging the uncertainty involved in applying the selected studies to the patient in question.<\/p>\n\n\n\n<p>There are few clinical scenarios more problematic than a practitioner being confidently wrong. Clinicians are rightly expected to disclose their gaps in knowledge or their inability to forecast an outcome. Yet emerging AI tools often cannot \u2014 or will not \u2014 do the same. In an analysis of large language models (LLMs) that were presented with 300 physician-designed vignettes, each including one fabricated detail, the LLMs accepted and amplified the falsehood 50 to 82% of the time.<sup><a href=\"https:\/\/www.nejm.org\/doi\/full\/10.1056\/NEJMp2517624?query=featured_home#core-collateral-r1\">1<\/a><\/sup>&nbsp;In short, most models didn\u2019t say, \u201cI don\u2019t know.\u201d<\/p>\n\n\n\n<p>Current methods for training AI systems rarely reward abstaining from answering a question, and regulations typically don\u2019t require this capability. Yet if AI is going to support clinical reasoning, systems will need to be able to indicate uncertainty. This behavior could be taught and assessed for humans and machines alike.<\/p>\n\n\n\n<p>Clinicians navigate multiple forms of uncertainty: factual uncertainty (involving knowledge gaps), diagnostic uncertainty (involving incomplete evidence), prognostic uncertainty (involving inability to predict outcomes), and values-based uncertainty (involving misalignment of goals). The act of saying \u201cI don\u2019t know\u201d takes different forms in various domains but serves a common function: avoiding premature closure and permitting a shift from an intuitive thinking style to an analytic thinking style.<sup><a href=\"https:\/\/www.nejm.org\/doi\/full\/10.1056\/NEJMp2517624?query=featured_home#core-collateral-r2\">2<\/a><\/sup>&nbsp;The resident in the opening vignette made such a shift when faced with multiple plausible reasons for the rising creatinine level (and the system should have made an analogous shift when faced with conflicting information). During this shift, clinicians engage in several cognitive processes, including forward planning, monitoring, and critical thinking.<sup><a href=\"https:\/\/www.nejm.org\/doi\/full\/10.1056\/NEJMp2517624?query=featured_home#core-collateral-r3\">3<\/a><\/sup>&nbsp;Critical thinking is a learned set of skills involving metacognition (self-awareness and questioning assumptions), evaluating quality and applicability of evidence, and recognizing bias. Related concepts in health professional education include metacognition, self-awareness, and recognition of uncertainty.<sup><a href=\"https:\/\/www.nejm.org\/doi\/full\/10.1056\/NEJMp2517624?query=featured_home#core-collateral-r3\">3<\/a><\/sup>&nbsp;These capabilities provide a bulwark against cognitive errors because they prompt human\u2013AI systems to remain open to the possibility of not knowing.<\/p>\n\n\n\n<p>Although there have been extensive discussions about these clinical competencies, expression of appropriate self-doubt hasn\u2019t been included as a core competency in medical education or embedded into AI product development. Recognizing and acknowledging when the degree of self-doubt crosses a threshold to a state of \u201cI don\u2019t know\u201d is the first step in operationalizing the commitment to \u201cdo no harm\u201d; the crossing of this threshold indicates that critical thinking is necessary. The ability to say \u201cI don\u2019t know\u201d \u2014 particularly in high-stakes scenarios \u2014 may be the truest hallmark of an expert.<sup><a href=\"https:\/\/www.nejm.org\/doi\/full\/10.1056\/NEJMp2517624?query=featured_home#core-collateral-r3\">3<\/a><\/sup><\/p>\n\n\n\n<p>The resident\u2019s \u201cI don\u2019t know\u201d reflected epistemic humility, a human virtue that involves metacognitive awareness, a moral commitment to truthfulness, and recognition of the limits of one\u2019s knowledge. AI systems lack the metacognitive architecture that enables epistemic humility: LLMs are next-token predictors. They don\u2019t \u201cknow\u201d what they don\u2019t know \u2014 they just generate statistically probable text. Even if an LLM were trained to produce the output \u201cI don\u2019t know\u201d more often, this output wouldn\u2019t necessarily align with actual knowledge gaps. The tool might say \u201cI don\u2019t know\u201d too frequently (rendering it useless) or miss critical gaps (rendering it dangerous). Even state-of-the-art AI models and methods (e.g., retrieval-augmented generation, agentic workflows, and multistep reasoning) have this shortcoming. Although these tools are unlikely to invent fictional studies, they could cite a single paper without acknowledging contradictory evidence or fail to indicate when study populations differ from the patient in question. AI systems that functionally express uncertainty could serve clinical purposes analogous to those of epistemic humility: preventing premature closure and triggering appropriate deliberation.<\/p>\n\n\n\n<p>Competency-based medical education (CBME) involves defining clinical competencies that trainees must develop and entrustable professional activities (EPAs) \u2014 measurable units of practice that reflect competency development and signal when a trainee can be trusted with increasing autonomy.<sup><a href=\"https:\/\/www.nejm.org\/doi\/full\/10.1056\/NEJMp2517624?query=featured_home#core-collateral-r4\">4<\/a><\/sup>&nbsp;We propose that the ability to explicitly state \u201cI don\u2019t know\u201d when appropriate be implemented as a core entrustable behavior that an external evaluator can observe and document for both humans and machines. This EPA translates epistemic humility from a human virtue into a clinical competency associated with a measurable behavior: noticing ambiguity and articulating uncertainty to trigger critical thinking. Like other EPAs, this behavior can be observed, taught by means of modeling and feedback, and assessed. For questions that aren\u2019t well-posed or that have unknowable answers, experienced clinicians tend to prefer to address uncertainty head-on than to provide an answer displaying false confidence.<sup><a href=\"https:\/\/www.nejm.org\/doi\/full\/10.1056\/NEJMp2517624?query=featured_home#core-collateral-r3\">3<\/a><\/sup><\/p>\n\n\n\n<p>A similar framework can be applied to AI. We and other scholars have proposed AI\u2013CBME approaches that involve defining competencies, mapping EPAs, and establishing milestones for AI models and entrusting these models with increasing authority. We suggest that AI tool developers implement a CBME framework involving measurable EPAs and developmental milestones to promote and surface AI\u2019s computational ability to produce the output \u201cI don\u2019t know.\u201d<\/p>\n\n\n\n<p>Methods for quantifying uncertainty and supporting abstention in AI systems have become increasingly available, and some systems can route outputs involving uncertainty to clinicians for review. Nonetheless, there are gaps in the implementation of algorithmic uncertainty in clinical workflows; questions remain about when AI tools should signal uncertainty, how uncertainty should be expressed, and how health systems should evaluate the reliability of uncertainty signaling. This framework would link the decision to deploy an AI system to the execution of a measurable behavior, with clear criteria for uncertainty management, regardless of the underlying computational process that generated the uncertainty. The goal for AI products subject to such a framework would be to execute uncertainty behaviors that serve the same clinical role as epistemic humility.<\/p>\n\n\n\n<p>For clinical tools, management of uncertainty must be evaluated and regulated throughout the product\u2019s lifecycle. During development, AI systems should refrain from providing confident answers when they lack grounds for confidence. The threshold for expressing uncertainty could be context-dependent: high-stakes decisions and those made in settings with few backup resources demand a higher bar for expressing certainty than other decisions. Systems must identify uncertainty when critical information is missing, the query is beyond a tool\u2019s scope, there is conflicting evidence, or confidence is low \u2014 and then must provide transparent guidance that includes guardrails, rather than simply declining to answer a question. During validation, rates of expressed uncertainty could be measured against real-world accuracy rates for diagnosis and treatment decisions, including in various subgroups of patients and settings. At deployment, health systems would define an escalation pathway \u2014 who should review flagged outputs, how quickly, and using what documentation.<\/p>\n\n\n\n<p>These systems could be tested using scenarios in clinician-annotated benchmarking data sets in which the safest output is to refrain from producing an answer, to ask for missing information, or to present multiple plausible possibilities. Benchmarks can be used to assess performance on any reasoning task, including the task of distinguishing medications from Pok\u00e9mon: we found that LLMs confabulated in 90% of instances when the name of a Pok\u00e9mon character was introduced in a list of medications, providing indications or dosing instructions for the character (see\u00a0<a href=\"https:\/\/www.nejm.org\/doi\/full\/10.1056\/NEJMp2517624?query=featured_home#f1\">figure<\/a>).<sup><a href=\"https:\/\/www.nejm.org\/doi\/full\/10.1056\/NEJMp2517624?query=featured_home#core-collateral-r5\">5<\/a><\/sup>\u00a0Although trained professionals may also fail to identify a name as that of a Pok\u00e9mon character, many would pause after reading an unfamiliar medication name and seek more information before proceeding. In this study, errors were reduced when LLMs were given instructions about how to respond to perceived uncertainty.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><a href=\"https:\/\/csccm.org.cn\/wp-content\/uploads\/2026\/07\/nejmp2517624_f1-scaled.jpg\"><img decoding=\"async\" loading=\"lazy\" width=\"1024\" height=\"446\" src=\"https:\/\/csccm.org.cn\/wp-content\/uploads\/2026\/07\/nejmp2517624_f1-1024x446.jpg\" alt=\"\" class=\"wp-image-31401\" srcset=\"https:\/\/csccm.org.cn\/wp-content\/uploads\/2026\/07\/nejmp2517624_f1-1024x446.jpg 1024w, https:\/\/csccm.org.cn\/wp-content\/uploads\/2026\/07\/nejmp2517624_f1-300x131.jpg 300w, https:\/\/csccm.org.cn\/wp-content\/uploads\/2026\/07\/nejmp2517624_f1-768x334.jpg 768w, https:\/\/csccm.org.cn\/wp-content\/uploads\/2026\/07\/nejmp2517624_f1-1536x669.jpg 1536w, https:\/\/csccm.org.cn\/wp-content\/uploads\/2026\/07\/nejmp2517624_f1-2048x891.jpg 2048w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/a><\/figure>\n\n\n\n<p>We believe clinical AI systems should be trained to express calibrated uncertainty in ways that complement and indicate the need for epistemic humility in human users \u2014 and their performance in these areas should be assessed. Systems without the ability to execute uncertainty behaviors will continue to produce persuasive fiction in precisely the moments when patients need their clinicians to pause and ask for help. Contemporary LLMs have passed many Turing tests, but will they pass this modern test of not knowing? We don\u2019t know.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>PERSPECTIVE Can AI Say \u201cI Don\u2019t Know\u201d? Andrea&nbsp;Siko [&hellip;]<\/p>\n","protected":false},"author":3,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":[],"categories":[24,23],"tags":[],"_links":{"self":[{"href":"https:\/\/csccm.org.cn\/index.php?rest_route=\/wp\/v2\/posts\/30922"}],"collection":[{"href":"https:\/\/csccm.org.cn\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/csccm.org.cn\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/csccm.org.cn\/index.php?rest_route=\/wp\/v2\/users\/3"}],"replies":[{"embeddable":true,"href":"https:\/\/csccm.org.cn\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=30922"}],"version-history":[{"count":2,"href":"https:\/\/csccm.org.cn\/index.php?rest_route=\/wp\/v2\/posts\/30922\/revisions"}],"predecessor-version":[{"id":31402,"href":"https:\/\/csccm.org.cn\/index.php?rest_route=\/wp\/v2\/posts\/30922\/revisions\/31402"}],"wp:attachment":[{"href":"https:\/\/csccm.org.cn\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=30922"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/csccm.org.cn\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=30922"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/csccm.org.cn\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=30922"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}