Chatbots almost never say “I’m not sure.” Researchers at the University of Miami want to change that.
Engineers at the University of Miami College of Engineering are working on a way to measure AI uncertainty. The goal is to flag answers that may be unreliable.
Why Confident Answers Are a Problem
Large language models are built to produce the most probable response to a prompt. That holds even when they lack reliable information.
The result can sound convincing without being correct. In fields like medicine and defense, that gap gets costly.
“Even a slight oversight in the output could be severe and very costly,” said Kamal Premaratne, a professor of electrical and computer engineering.
“Right now, as soon as you ask an LLM a question, it gives you an answer rather than saying, ‘I’m not sure,’” said doctoral candidate Pragatheeswaran Vipulanandan.
“The ultimate goal is for the LLM to tell you, ‘I’m not sure, but these are the top answers I have,’ or even ask for more context before responding,” Vipulanandan said.
Testing How Stable the Odds Are
One common approach asks a model the same question several times. If the answers vary, the model may be unsure.
The Miami team looks past the answers themselves. Every response is built from probabilities that guide how the model picks words.
The researchers built a mathematical framework that tests how stable those probabilities stay when small changes are introduced. Large swings suggest the response deserves a closer look.
Premaratne and Vipulanandan presented the work earlier this year at the International Conference on Learning Representations in Rio de Janeiro. Dilip Sarkar, an associate professor of computer science, collaborated on it.
Counting the Leopards Nobody Saw
The team also examines answers a model could have given but did not.
They use a statistical technique called the missing-species model. It considers possibilities beyond what has already been observed.
Premaratne compares it to a visit to a wildlife reserve. You count the giraffes, hippos and rhinos you see.
“The fact that you didn’t see a leopard doesn’t mean there are no leopards in the park,” he said.
In the same way, a model may hold other possible answers that never appear in its output. Accounting for them gives a fuller picture of its uncertainty.
Tools From Quantum Physics
The team is not building a quantum computer. It has adapted mathematical tools from quantum physics that measure uncertainty in complex systems.
Those tools can estimate how many outcomes are possible, even when researchers can observe only some of them.
Why Miami Should Care
The university says the work matters most as AI enters high-stakes fields. Miami is home to OpenEvidence, a clinical AI company.
Clinical AI is exactly where an “I’m not sure” signal has value. A confident wrong answer costs more in a medical setting than in a casual chat.
The approach does not eliminate hallucinations. Its purpose is a clearer signal for when an answer needs a second look.
Human judgment stays in the loop. “The logical thinking comes when you actually know how to do the work on your own,” Premaratne said.
A model that admits doubt does not replace an expert. It tells the expert where to look.

