Bandwise docs

Confidence bands and actions

How each answer lands in a band, and how each band becomes an action.

Every answer lands in one of three bands: high, medium or low. Each question's policy maps each band to an action. Your app acts on the result, not on raw probabilities.

From answer to band

Choice and score answers carry a confidence from 0 to 1. The policy sets two thresholds:

"thresholds": { "high": 0.75, "medium": 0.45 }

A confidence at or above high is the high band. At or above medium is the medium band. Anything lower is the low band. A choice policy can also set stricter thresholds for risky options with perOption.

Noul answers have no confidence field, only the probability of yes. The policy sets where yes and no become clear:

"noul": { "trueAt": 0.85, "falseAt": 0.15, "reviewMargin": 0.1 }
Probability of yesBandValue
0.85 or morehightrue
0.15 or lesshighfalse
0.75 to under 0.85mediumtrue
over 0.15 to 0.25mediumfalse
anything in betweenlownull

From band to action

ActionWhat happens
autoThe answer is applied.
reviewA review item is created for a person. Nothing else happens until they resolve it.
fallbackThe configured fallback runs: a fixed value, another set, or nothing.
escalate_to_llmA reasoning model is called during the run. Its answer is returned next to the System One answer, and its cost is counted as escalation spend.

A policy picks one action per band:

"actions": {
  "high": { "kind": "auto" },
  "medium": { "kind": "review" },
  "low": { "kind": "review" }
}

If an escalation fails or times out, that decision becomes review instead.

Gating and the run as a whole

A question marked gating: true counts toward the run's overall result. The run band is the lowest band among relevant gating decisions. The overall action is the most conservative action among them, in this order: review, then fallback, then escalate_to_llm, then auto.

A question whose relevantWhen condition is false is left out. It creates no review item, does not lower the run band and does not count toward savings.

Action and effective action

Each decision in a run result carries two actions:

  • action is what the policy says.
  • effectiveAction is what the current rollout stage allows.

Your app should always act on effectiveAction.

Choosing thresholds

Thresholds should follow the cost of a wrong action, not the question. These are starting points by risk tier, measured on Jev 1.13. Tune them on your own labeled data.

TierExamplehighmediumNoul trueAt / falseAt / margin
Low riskTagging, sorting, UI hints0.600.300.80 / 0.20 / 0.10
StandardRouting tickets, ranking, triage0.750.450.85 / 0.15 / 0.10
High riskMoney movement, account changes0.900.700.95 / 0.05 / 0.05
  • If all you need is the best option, take the top choice and don't threshold it.
  • Low confidence on a harmless preference can be fine. Several acceptable options spread probability.
  • Confidence summarizes the probabilities. It is not permission to act.
  • Don't copy a noul threshold to a choice question.

Bandwise is an independent product built on TypeSafe's System One models. It is not TypeSafe's documentation. For the System One models themselves, see docs.typesafe.ai.

On this page