ToWow

When Should an AI Agent Ask a Clarifying Question?

An agent with several possible clarifying questions should ask the one with the highest score, and only if that score is strictly above the cost of a round of asking. If even the best question falls short, the agent asks nothing and moves on. The score is one line: the value of changing the decision times the probability that the answer changes it. This article walks through that rule as we implemented it in research code, with a worked example that uses made-up teaching numbers.

Zhang Chenxi (Nature), who builds agent systems for manufacturers and distributors. . Drafted with AI assistance.

The problem: missing facts, costly questions

Our collaboration-discovery research looks at a person with a need and an agent that helps turn it into a collaboration plan. The agent sorts out the conditions, looks for participants, and drafts a plan people can discuss. Along the way it often lacks a fact: whether a candidate has the equipment, whether the timing works, who has the authority to decide.

A missing fact does not mean every question should be asked. Each question takes the person's time and can delay the next step. The agent needs a way to decide which question to ask next, and when to stop.

The function: select_voi_question

The discovery engine in our research has a function called select_voi_question. It takes a set of candidate questions. Each question carries two estimates: the value of changing the decision, and the probability that the answer changes it. All questions in a round share one flat cost of asking. The function returns the one question worth asking first, or "no question this round".

It borrows the idea of value of information (VOI) and uses a simplified product rule. The function only selects. The upstream process writes the questions and supplies the value and probability estimates, and the function does not generate questions or estimate payoffs itself.

Understanding the mechanism comes down to three connected judgments: which decision an answer could change, how much that change is worth, and whether getting the answer beats moving forward without it.

A worked example: equipment or reminders

The numbers in this example are invented for teaching. They do not come from a real event or from measurements of residents.

A neighborhood group is organizing a small repair event and needs someone to fix old appliances that residents bring. Two repairers are available. One lives nearby, and the other has better tools. The organizer has little time for follow-up questions. The agent has two candidate questions:

  1. Does the nearby repairer have circuit-testing equipment? The answer could directly change who gets picked.
  2. Do attendees prefer SMS or group-chat reminders? The answer would help choose a notification method, but most attendees have already confirmed their plans, so adjusting reminders gains little.

We assume the equipment question has a 40% chance of changing the pick, worth 100 teaching units if it does. The reminder question has an 80% chance of improving the notification method, worth 10 units. One round of asking costs 15. Both values mean a useful improvement. A changed mind counts as a gain only when the new decision is better.

Teaching questionValue if the decision changesProbability it changesScore
Does the repairer have testing equipment?1000.440
SMS or group chat for reminders?100.88

The equipment question scores 100 × 0.4 = 40, and the reminder question scores 10 × 0.8 = 8. The agent picks the equipment question, because 40 is the higher score and also exceeds the round cost of 15.

The reminder question is more likely to change something, which is why ranking by likelihood alone would choose it. The value of the change is small, and multiplying value by probability puts that into the score.

Diagram of the question-selection rule: candidate questions are scored as value times probability, the top score is compared with the cost of a round, and the agent either asks that question or asks nothing
Teaching diagram of the question-selection rule. All numbers are made up for illustration.

The selection rule

The research literature has linked information to decisions for a long time. Ronald Howard, writing on information value in 1966, pointed out that measuring uncertainty is not enough. The effect of the outcome on the decision maker also counts, and the value of several pieces of information taken together can differ from the sum of their separate values. His paper is Information Value Theory.

Our implementation takes a simpler route. Each candidate carries a value estimate and a probability estimate, and the code multiplies them:

score = value of changing the decision × probability the answer changes it

It sorts the questions by score, highest first. The top question is chosen only when its score is strictly greater than the cost of the round.

This is a simplified rule inspired by value of information. It does not enumerate every possible answer to a question, does not re-optimize the action for each answer, and does not learn probabilities. The process that calls it supplies the value and the probability.

The meaning of "value" matters here. Suppose a question easily makes people change their minds, but the new choice is often worse. The change itself is not a gain. The value input has to describe a useful improvement to the decision, and not merely that the decision moved. The function does not check this for the caller. If the inputs confuse the two, the multiplication still runs.

When to stop asking

Go back to the equipment question. Suppose the organizer has now seen a photo of the equipment, and that settles most of it. The chance that the answer changes the pick drops to 0.1. With the same value of 100, the score is 10, below the cost of 15. No question is worth asking this round, and the function returns an empty result.

That does not cancel the repair event, and it does not mean the other questions will never matter. It means that on current estimates another round of asking costs more than it returns. The organizer can go ahead with what is known, or find a cheaper source, such as a public equipment list.

The comparison is strict. A score exactly equal to the cost is not asked, because the code tests "greater than", not "greater than or equal to". If the cost of the round dropped from 15 to 5, the equipment question with its score of 10 would qualify again. The question did not change. The price of getting an answer did.

That makes cost a real design variable. Shorter forms, letting people upload materials directly, and reusing earlier answers could all change which questions are worth asking. How much those measures lower the cost has to be measured on real tasks, and a threshold rule cannot supply that number.

Where the rule breaks

The inputs are the weak point. The code does not say where the probability comes from, for whom the value holds, or whether ten minutes of waiting, one interrupted person and one extra model call have been converted to the same scale. The caller has to settle all of that.

The function also scores one question at a time. Some facts matter only in combination. Knowing a repairer has the equipment does not help if you do not know whether they can attend, and the decision changes only when both answers arrive. Each question alone may score low, so the rule can miss a pair of questions that is worth asking in sequence. For tasks like this, you can compare a group of dependent questions as a unit, or use a fuller sequential decision method. Whether that added complexity pays off depends on how much a missed question costs in practice.

The implementation shows one behavior clearly: from a batch of candidates, pick the one with the highest estimated payoff, provided it exceeds the cost. It makes an explainable starting point for scheduling questions. Its effect on real collaboration tasks still needs measuring.

That also defines what to measure. Counting how many fewer questions the system asked is not enough, because final decision quality has to be checked too. A system that halves its average number of rounds by skipping necessary confirmations has not done the original job.

A habit you can use today

Before you ask the next question, write down which decision the answer would change. If you cannot say, the question itself probably needs rework.

This rule comes out of ToWow's research on agents that coordinate work between people. If you are deciding when your own agents should stop and ask, and want to compare notes, write to hi@towow.ai.

The original of this article, in Chinese, is at towow.net/articles/towow-voi-question-selection.

If you want a system like this built around your own workflow, see AI agent systems for manufacturers and distributors or write to hi@towow.ai.