Teaching a Reasoning Model to Know When It's Sure Cuts Its Token Bill
A self-supervised confidence objective trims up to 25% of reasoning tokens at matched accuracy across Gemma, Qwen, Nemotron, and GPT-OSS, without ever training the model to stop, shorten, or optimize for length.
TL;DR: A paper titled “Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency” shows you can cut up to 25% of a reasoning model’s generated tokens at matched accuracy by training it to predict its own confidence, never once optimizing for length or stopping.
The interesting part of this result is what the training objective does not contain. No length penalty. No early-stopping head. No reward for brevity. The authors fine-tune reasoning models to predict, at intermediate points in their own reasoning traces, how confident they are in the eventual answer. That is the whole target. Efficiency falls out as a side effect, and that is the claim worth sitting with.
What did the paper actually do?
The setup is small and specific, which is why it is worth taking seriously. Using a self-supervised procedure, the authors generate confidence labels from the models’ own reasoning trajectories, then fine-tune on only 600 training problems. The loss teaches the model one thing: at a given point mid-reasoning, estimate the probability that your final answer will be correct.
Crucially, the loss “contains no objective for reasoning length, efficiency, or stopping.” And at inference time there is no special machinery. No confidence threshold that triggers a halt, no elicitation prompt asking the model how sure it is. The fine-tuned model runs the standard generation procedure. It just happens to generate fewer tokens to reach the same answer quality.
That combination is what separates this from the usual efficiency toolkit. Most existing approaches do one of two things: bolt an early-stopping mechanism onto inference, or explicitly push for shorter chains during training, often via reinforcement learning with a length penalty. Both work. Both also directly optimize the thing you want, which means you are trading accuracy against brevity and tuning that knob. Here the knob does not exist. The model learns a metacognitive signal, and shorter reasoning emerges downstream.

Does it generalize past one model or one benchmark?
This is where the paper earns its keep. The 25% token reduction at matched accuracy holds “across Gemma, Qwen, Nemotron, and GPT-OSS models,” spanning four different model families from different labs. And it holds across “mathematical, scientific, and coding reasoning benchmarks,” not one narrow domain.
That breadth matters because efficiency tricks have a bad habit of being artifacts of a single checkpoint. Something that works on one Qwen math variant and nowhere else is a curiosity. Something that reproduces across four families and three task types is closer to a property of how reasoning models behave. The authors report efficiency gains “comparable to methods that explicitly optimize for shorter reasoning,” which is the honest framing: this is not beating length-penalty RL by a mile, it is matching it while sidestepping the direct optimization and the accuracy tradeoff that usually comes with it.
I want to flag the standard caveat, because the sources here are three arXiv listings of the same paper (cs.AI, cs.CL, and cs.LG cross-postings, identical abstract) and I have the abstract, not the full experimental appendix. “Up to 25%” is a ceiling, not an average. Matched accuracy is defined by the authors’ own evaluation, and “matched” always hides some variance you only see in the tables. Take the headline as a well-scoped claim across a real spread of models, not as a guaranteed number on your workload.
Why does confidence make reasoning shorter at all?
The mechanism the paper offers is the genuinely useful idea. Long reasoning traces are partly a symptom of a model that cannot tell when it is done. It keeps second-guessing, re-deriving, hedging, because it has no internal read on whether the answer it already reached is good. Give it a calibrated sense of its own certainty and the redundant work has less reason to happen.
What the authors checked next is the part that made me trust the result more. Analysis of reasoning episodes “largely preserves the base models’ high-level reasoning composition rather than selectively suppressing particular behaviors.” In plain terms: the model did not get shorter by lobotomizing one type of step, like dropping all self-checks or refusing to backtrack. The overall shape of how it reasons stays intact. It just does less padding around the same skeleton.
That is a meaningfully different failure profile than length-penalty training, which can teach a model that tokens are bad and quietly erode the very self-correction that makes reasoning models good. If confidence supervision preserves composition, you get efficiency without the silent capability tax. That is the claim I would most want to stress-test on my own tasks before believing fully, but it is the right thing to have measured.

What should a builder do with this?
If you run reasoning models in production, the operational pain is rarely accuracy. It is the token bill and the latency, both driven by traces that run 3,000 tokens when 2,000 would do. A 25% cut at matched accuracy is a direct hit on cost per query and time to first useful answer, and it compounds at scale.
The practical catch: this is a fine-tuning result, not a prompt you can paste. You need model weights you can train, which is exactly why the family list matters. Gemma, Qwen, Nemotron, and GPT-OSS are open-weight, so the technique is reproducible by anyone with a fine-tuning pipeline. If your stack is a closed API, you cannot apply this directly today, though it is the kind of thing a provider could ship inside a model and never mention.

The other quiet win: confidence is useful on its own, separate from token savings. A model that predicts its own probability of being correct mid-trace gives you a routing signal. Low confidence at the halfway mark is exactly when you escalate to a bigger model, add tool calls, or flag for human review. The paper trains confidence purely as an efficiency lever, but the same head is a reliability instrument if you expose it.
Practitioner’s take: if you self-host an open-weight reasoning model, this is worth a weekend experiment, because 600 training problems and no new inference machinery is about as cheap as a real efficiency gain gets. Build the confidence labels from your own traces, fine-tune, then measure token count and accuracy on your actual traffic, not a public benchmark, because “matched accuracy” on math is not the same as matched accuracy on your support tickets. The trap most people will fall into is treating 25% as a spec. It is a demonstrated ceiling across four model families, which is a strong signal that the effect is real and a weak basis for a capacity-planning number. Reproduce it on your workload first, then bank the savings.