The answer says what. Confidence says whether to act on it without a person.
Confidence falls out of how spread the probabilities are — a flat distribution gives a low number. Three bands do most of the work: act automatically, confirm or review, or send it to a person. Set the bars by what the action costs.
Calibration is a promise about many predictions — outcomes given 0.8 should happen about 80% of the time — not about the one answer in front of you. And re-deciding costs nothing: the thresholds run in your code, against the answer you already have.