Problem
A prior that can correct you can also overrule you
A prior knows which words are common, not what this user meant. Below, the user heard “sat”. The decoder's own scores are weak and slightly off, and the prior likes “sat”, “mat” and “cat” equally. The system adds the two scores, with α setting how much the prior counts. Drag it.
Once the prior counts enough, the output flips to “mat” and stays wrong. Meanwhile the lead of the top word over the second, which the system reads as confidence, grows from 0.3 to 0.8. The system is now surer of a wrong word than it ever was on its own.
Observation
Where most errors remain, confidence points the wrong way
To see where this happens, you need to know where the right word stood in the decoder's own list before the prior stepped in. A running system cannot know this, since it does not know the right word; its scores only give hints. In an offline study we do know it. So we took every word the decoder first got wrong, grouped them by how far down the right word started, and asked: in each group, are the errors the prior fixed more confident than the ones it left wrong?
Rank outputs by confidence, and the system says the wrong one first.
The answer is a score from 0 to 1. Above 0.5, confidence points to the fixed errors, as it should. Below 0.5, it points to the ones left wrong: backwards. Here the prior is our own, which builds its guess from the decoder's readings of the preceding words.
Near the top of the list, confidence works. Further down it flips: the more confident the output, the more likely it is a mistake the prior left in place. That is where 46.6% of the remaining errors are, while the overall check, which lumps all errors together, still reads 0.87.
Explanation → prediction
A fix has to pay for its own gap
Think of the prior as a push. To fix an error, it must first push the right word past every wrong word above it. What it spends on that is gone, so the fixed word ends up only slightly ahead. With G0 the right word's starting gap, Δπ how hard the prior can push, and α its weight:
A mistake the prior cannot fix has no such limit. The prior can push two wrong words far ahead of the rest, and the lead between them grows freely. Far down the list, the right word is even second in only 2.1% of these leftover mistakes: the system's confidence is mostly a contest between two wrong words.
Predictionα sits on the right-hand side: more weight means more push. If this explanation is right, giving the prior more weight should push the flipped region further down the list.
Causal test
Turn one knob, hold everything else
We change only α. Everything else, the decoder's scores, the prior's scores and the words being decoded, stays exactly the same.
As α grows, the point where confidence flips moves steadily down the list, from the very top to ranks 11–20. It never moves back, in all 1,000 repeats of the analysis on resampled participants. The decoder's own confidence, which the prior never touches, stays backwards throughout, so the change comes from mixing in the prior.
The explanation said which way the boundary should move. It moved that way, at every step.
Generalisation
A prior we did not build
Our prior is built from the decoder's own readings, so perhaps the effect comes from how we built it. To check, we swapped in a prior we had no hand in: Qwen3-4B, a large language model, scoring each candidate word given the words before it. The flip is there, and more weight again pushes it down the list. Then comes something an accuracy score would never show:
Accuracy is best at α = 0.5 and then falls, yet the flip keeps moving. Accuracy and confidence respond differently to the same prior: a weight tuned for one does not fix the other. All six priors we tested show the backwards region.
Application
Not fixing every word. Knowing which ones not to say.
Errors that start far down the list are practically unfixable: the prior fixes about 1% of them, and pushing harder is exactly what overrides users. The realistic goal is to let the prior fix what it can and to recognise the rest, so the system can hold them back. The information is there, only hidden by the mixing: the prior's own scores, looked at separately, reach 0.94 on the 0-to-1 scale in the very group where the mixed confidence reads 0.39.
It offers a short list, about five words on average, for the user to pick from. 92% of these lists contain the right word.
The window is held back, to be repeated or confirmed, instead of putting a word in the user's mouth.
Illustrative example of what a selective decoder outputs.
Among the hardest errors, the share of answers that contain the right word rises from 36.4% to 50.8%. Re-scaling the usual confidence cannot do this: it changes the numbers but not their order, so the flip stays.
Impact
Decoders that know when not to speak for you
Brain–computer interfaces are judged by how often they are right. As language models make them fluent, a second question decides whether a person can trust one with their voice: when it is wrong, does anyone know? We show the answer can be no, exactly where errors pile up; we explain why, and show how to catch them. The explanation needs nothing but a prior added to weak brain evidence, so the same check applies to any decoder built this way.
- Check confidence where the errors are, not only on average.
- Choose how much the prior counts for confidence, not only for accuracy.
- Keep the brain's evidence and the prior's evidence apart, so the decoder can hold back instead of speaking for the user.
A decoder that is right more often is progress. A decoder that knows when it might be speaking for you is one a person can trust.
Paper
Abstract
Contextual priors improve neural-to-language decoding by reshaping candidate scores. However, confidence is read from the same reshaped scores, so the errors a prior leaves behind can become more confident with no change in accuracy to reveal it. In this work, we ask how a prior shapes the confidence of these errors, studying speech retrieval on MEG-MASC and MOUS with local decoding scores and a contextual prior combined by additive shallow fusion, and the fused top-two margin as confidence. Among initially incorrect predictions, we find a confidence-ordering reversal: a larger margin makes a repair, an error that fusion corrects, more likely when the correct candidate starts near the top of the local ranking, but less likely when it starts lower. On MEG-MASC, pooled correctness AUROC is 0.87, yet the AUROC separating repairs from residual errors, which fusion leaves uncorrected, falls from 0.70 at initial ranks 2–3 to 0.39 at ranks 21–50. Errors starting beyond rank 20, inside the reversed region, make up 46.6% of all errors after fusion. We propose a score-level account: a repair must first close the correct candidate's initial deficit, which limits its final margin, whereas a residual error can build a large margin between two incorrect candidates. Through a causal intervention that changes only the fusion weight, we show that the reversal moves to deeper ranks, as the account predicts, and that under a word-level LM prior it keeps moving after accuracy gain peaks, so a weight chosen for accuracy does not settle confidence. Guided by this account, we read local and prior scores separately: read before fusion, the prior's own scores already separate repairs from residual errors where the fused margin reverses, and estimators built on local and prior scores let a selective decoder answer on 74.5% of windows instead of 56.7%, with 92% of its output sets still containing the correct candidate. Confidence after contextual fusion should retain the local and contextual evidence behind each prediction, not just the fused scores.