One of the most troubling moments when using generative AI is not when it fails to know the answer, but when it presents an incorrect answer with complete confidence. AI hallucinations—including fabricated citations and false information delivered as fact—remain one of the most persistent obstacles to building trustworthy AI systems.

This raises an important question: Can a language model assess how confident it is in an answer? And when its confidence is low, can it choose not to respond?

Published in Nature Machine Intelligence in September 2026, the paper “Causal Evidence That Language Models Use Confidence to Drive Behaviour” examines whether internal confidence signals influence how language models behave. The study moves beyond whether an AI can simply generate the phrase “I don’t know.” Instead, it investigates whether the model uses its own internal uncertainty to decide whether to answer or abstain.AI Can Stop When It Doesn’t Know

Why Does AI Sound Confident Even When It Is Wrong?

Current language models do not verify whether they truly know the correct answer before generating a response. Instead, they produce text by predicting the words or tokens most likely to come next based on patterns learned from training data.

As a result, a model can generate a plausible answer even when the information available to it is incomplete. A response may sound fluent and logically structured without being factually correct.

In high-stakes fields such as healthcare, law, and education, the ability to detect uncertainty and stop can be more important than answering every question. To explore this capability, the researchers examined how language models choose between providing an answer and abstaining.

How Was Confidence Measured?

The researchers focused primarily on GPT-4o while also examining Gemma 3 27B, DeepSeek, and Qwen models. The experiment consisted of four phases.

In the first phase, each model was required to select one of four possible answers. Abstention was not available. During this stage, the researchers measured both the probability assigned to each option and the model’s separately expressed verbal confidence.

In the second phase, an abstention option was added to similar questions. The model was allowed to withhold its answer when no option appeared clearly correct.

During the third phase, the researchers intervened in the model’s internal activations to artificially increase or decrease its confidence signals. This allowed them to test whether confidence merely correlated with behavior or actually caused the model to change its decision.

In the final phase, the models were explicitly instructed to abstain whenever their confidence fell below a specified threshold. The researchers then adjusted this threshold systematically to determine whether the models could modify their abstention policies accordingly.

GPT-4o Stopped Answering When Its Confidence Was Low

When GPT-4o was not allowed to abstain, it answered 63.7% of the questions correctly. Once the abstention option was introduced, the model withheld its answer in 56.6% of cases. Among the questions it did answer, however, accuracy increased to 69.1%.

In other words, allowing the model to stop when uncertain improved the reliability of the answers it chose to provide.

The researchers found that GPT-4o adjusted its behavior according to internal confidence even when it had not been given a numerical threshold. The implicit decision point—the confidence level at which answering and abstaining became equally likely—was estimated at approximately 77%. Below this point, the probability of abstention increased.

The researchers also compared several alternative explanations, including objective question difficulty, wording, and the accessibility of external knowledge. Internal confidence remained the strongest predictor of abstention behavior.

Changing Internal Confidence Changed AI Behavior

The relationship between confidence and behavior became even clearer in the experiment involving Gemma 3 27B.

The researchers intervened in the model’s internal activation patterns, steering them toward higher- or lower-confidence states. Under the strongest low-confidence condition, the model abstained in 66.5% of cases. Under the strongest high-confidence condition, the abstention rate fell to 7.0%.

This suggests that the model does more than verbally express confidence. Its internal confidence signals are actively involved in determining whether it answers or abstains. Because manipulating these signals changed the model’s behavior, the findings provide causal evidence rather than a simple correlation.

Increasing confidence, however, did not always produce a better outcome. Higher confidence led the model to answer more questions, but accuracy among those answers declined slightly. This reveals an important trade-off between answering more frequently and answering more reliably.

Why This Does Not Mean AI Has Self-Awareness

The findings should not be interpreted as evidence that AI possesses human-like self-awareness or experiences subjective confidence.

In this study, confidence does not refer to an emotion or conscious state. It is better understood as a computational signal formed through probabilities and internal representations during the response-generation process.

A model may still be highly confident and wrong. Conversely, it may possess information close to the correct answer but abstain too readily. Internal confidence therefore cannot be trusted automatically. It must be calibrated by examining how closely it corresponds to actual accuracy.

The study does not demonstrate that language models possess human-like metacognition. Rather, it suggests that these systems contain internal structures that connect confidence signals with decisions about whether to respond.

Building a Safer Standard for When AI Should Stop

Efforts to reduce AI hallucinations have generally focused on training models with more data or connecting them to external retrieval systems. This study points to another possibility: reading a model’s internal confidence and instructing it to search for additional evidence, request human review, or stop when confidence falls below an appropriate threshold.

Such mechanisms may become increasingly important as AI agents gain the ability to use tools and act autonomously. An incorrect answer in an ordinary conversation may be corrected later. But if an AI sends an email, executes a program, or makes a decision involving healthcare or finance, acting under uncertainty can produce much more serious consequences.

Future AI systems cannot be evaluated solely by how many questions they can answer. Their ability to distinguish between what they know and what remains uncertain—and to pause, verify, or transfer control when needed—will also be essential to their reliability.

This paper shows that safer AI does not always require more confident answers. Sometimes, the ability to stop when confidence is low may be more intelligent than producing another plausible response.