17 × 20 = 340, then subtract 17. So the product is 323.
The answer is correct, but a visible inferential step remains before it.Flip the switch. Watch the visible response change.
Same question, same answer target, two answer interfaces. This precomputed illustrative trace makes the response-level distinction concrete: a native no-think request can still leave visible inference, while a strict answer-only interface removes the pre-answer field.
What is 17 × 19?
323
Correctness is evaluated separately from the text before the answer.A narrated walkthrough of the motivation, response-level decomposition, three evaluation levels, and main findings.
A no-thinking switch does not guarantee an answer-only response
Before defining no-thinking, we ask a simpler question: when a model is told to return only the answer, is the visible pre-answer field actually empty? Figure 1 shows that the requested mode and the emitted response can diverge.
Figure 1. No-thinking controls do not guarantee response-level no-thinking. (a) A clear ``answer only'' instruction still elicits question-relevant pre-answer content from DeepSeek-V4-Flash and Qwen3-235B-A22B on a representative MMLU question. (b) Under native no-thinking, ETR and visible token length show that adherence is closest to answer-only for Boolean tasks, then falls across MCQ and Open-ended tasks.
A no-thinking control specifies the request—not the response. Even with a clear answer-only instruction, a model may expose question-relevant text before a correct final answer; under native no-thinking, this exposure changes systematically with the answer space.
The instruction–response gap
Both models return the correct option, yet expose question-relevant text before it. Correctness therefore does not certify answer-only compliance.
The answer space changes compliance
Under native no-thinking, Boolean outputs stay closest to the answer-only boundary, MCQ is intermediate, and Open-ended outputs retain the most visible pre-answer text.
These observations motivate the response-level object used below: separate the pre-answer text T from the final answer A, then distinguish empty output, relevant context, and explicit inference.
Abstract
Large Language Models (LLMs) increasingly ship with explicit “thinking modes,” yet their counterpart, “no-think,” has received far less attention. In this work, we study LLMs’ “no-think” behavior along two axes.
a. How to measure “no-thinking”? Prior work typically defines “no-thinking” through proxies such as a disabled thinking mode or the absence of long traces. These proxies are unreliable: disabled thinking modes still emit extended reasoning, or long traces often contain uninformative filler rather than genuine inference. We instead normalize any LLM response into pre-answer trace and final answer, and evaluate it at three levels: (i) an Empty-Thinking Ratio (ETR) for strict answer-only compliance; (ii) an instruction-aware Question–Pre-answer Relevance (QRel.) score to measure similarity between the question and pre-answer trace; and (iii) an additional LLM-as-judge Explicit Inference Rate (EIR) for visible explicit inference. Together, these levels separate answer-only output, relevant but non-inferential text, and explicit inference. Human studies demonstrate the effectiveness of these metrics.
b. How does no-thinking vary across tasks and models? We evaluate six prompting intervention settings on six widely used LLMs, both open-source and commercial, across three question types: Boolean, multiple-choice, and open-ended. We find that explicit no-think controls cannot reliably eliminate visible explicit inference; models instead exhibit “Thinking Inertia”: visible explicit inference persists even under strict no-thinking controls and becomes progressively more prevalent as the answer space opens. Accuracy remains relatively stable on Boolean and multiple-choice tasks, whereas open-ended tasks expose a trade-off between answer-only compliance and task accuracy. Finally, rewriting the same questions across answer spaces shows that supplying candidate answers makes answer-only responses substantially easier to produce. Taken together, these findings establish no-thinking as a non-trivial capability of LLMs: stopping explicit reasoning cannot be assumed from model settings or instructions alone, and deserves systematic evaluation alongside reasoning ability itself.
Surface proxies describe the interface—not the response
No-thinking is often inferred from a prompt-side switch, a missing <think> block, or a short generation. Figure 2 shows why each proxy can fail while the emitted response still exposes an inferential step.
Figure 2. Surface proxies fail to identify response-level no-thinking. (a) Prompt-side control, special reasoning tokens, and generation length are convenient but indirect signals. (b) On the same algebra question, disabling the thinking mode, omitting special tokens, or producing a short output can each leave a visible inferential step before the final answer.
Disabling the mode can still produce ``Subtract 1 from both sides...'' before the answer.
<think> token, but reasoning remainsOrdinary answer-channel text can still explain how to isolate x.
A brief ``3x = 9, so x = 3'' still exposes the decisive intermediate step.
A proxy can tell us how the response was requested, formatted, or sized; it cannot by itself tell us what question-conditioned content the response exposes.
These proxies are useful controls, but none defines the behavioral object we want to study. We therefore move from surface conditions to the visible pre-answer field itself.
Measure what the response exposes
We turn no-thinking into an observable response-level object by separating the visible pre-answer text from the final answer, without making a claim about latent computation.
Figure 3. Response-centric metrics for measuring no-thinking. Top: a question Q is wrapped by an intervention S(·) before the model produces a response decomposed into pre-answer text T and final answer A. Bottom: six responses share the same question and final answer while varying in length, relevance, and inferential content.
Same question, same answer
The six cases keep Q and A fixed. What changes is the visible text before the answer.
Different text, different behavior
T can be empty, generic, relevant but non-inferential, or explicitly inferential. Non-empty text is not automatically reasoning.
We evaluate T directly and evaluate answer correctness on A separately. This keeps visible reasoning claims at the response level, rather than treating a prompt setting or hidden computation as the measurement target.
Once T is explicit, the next question is not simply whether text exists, but what kind of text it contains: empty output, relevant context, or visible explicit inference.
One visible trace, three increasingly specific questions
The same pre-answer field T is read at three levels: whether it is empty, whether it is related to Q, and whether it contains an explicit inferential step.
Strict answer-only
ASKS · Is the complete response answer-only?T = ∅ and the final answer A is parseable; any substantive visible text outside the permitted answer wrapper breaks the boundary.
binary · all-output ratedoes not claim anything about latent computationQuestion–pre-answer relevance
ASKS · How related is T to answering Q?The instruction-aware Qwen3-Embedding-4B encoder reads Q as the query and T as the candidate passage; empty T receives zero.
continuous · all-output meanrelevance signal, not a reasoning detectorVisible explicit inference
ASKS · Does T make an inferential move?A blinded judge sees only (Q, T); only the mutually exclusive Explicit reasoning label contributes to the rate, while empty T is assigned zero.
binary · all-output ratevisible response evidence, not latent cognitionWhy the levels should not be collapsed
T = ∅T = ``The answer is 121.''T = ``2 + 9 = 11, carry 1.''ETR tells us whether T exists; QRel. describes what T is about; EIR asks whether T makes an explicit inferential move. Accuracy is evaluated separately on the final answer A.
What the experiments reveal
We use Thinking Inertia for a specific response-level behavior, then test whether it survives changes in answer space and no-thinking control.
Thinking Inertia
the persistence of visible, question-conditioned pre-answer text that contains explicit inference after the model has been told not to think. It is a claim about the emitted response—not latent cognition—and is measured through EIR, with ETR and QRel. providing complementary context.
More open interfaces preserve more visible inference
Across native and stricter no-thinking settings, the response-level signal follows a staircase: Boolean tasks are closest to answer-only, multiple-choice tasks are intermediate, and open-ended tasks retain the most explicit pre-answer work.
Stricter controls reduce the behavior, but do not erase it
Within every answer space, stronger no-thinking control yields lower explicit inference than the native no-thinking setting. The remaining Open-ended signal is the response-level signature of thinking inertia.
Results overview. The three-level view pairs strict compliance, visible explicit inference, instruction-aware relevance, and pre-answer length across answer spaces and intervention modes.
Compliance and capability are separate axes
A cleaner answer-only format does not automatically mean a more capable answer. Accuracy on A is read alongside ETR, QRel., and EIR.
Matched rewrites isolate the interface
When the same question is rewritten as Boolean, multiple-choice, or open-ended, short targets alone do not explain the pattern; matched lengths can still differ sharply in relevance.
Checks that keep the interpretation narrow
The evidence is layered so that no single proxy has to carry the reasoning claim on its own.
Alternative encoders preserve the relevance ordering, while the instruction-aware QRel. signal remains explicitly separated from reasoning.
Claude Opus 4.8 independently reproduces the Level 3 visible-inference pattern across models, modes, and answer spaces.
Five blinded annotators review a proportion-matched 600-output sample using the same relevance and reasoning rubric.
ETR marks the answer-only boundary; QRel. measures question-relevant residual text; EIR identifies visible explicit inference. The claim stays at the level of observable response behavior, not latent cognition.
📋 BibTeX
Read the paper on arXiv (2610.11765).
@article{lei2026thinkinginertia,
title = {Thinking Inertia: LLMs Keep Thinking When Told Not To},
author = {Lei, Dianqiao and Lin, Kevin Qinghong and Lu, Pan
and Torr, Philip and Zou, James},
year = {2026},
eprint = {2610.11765},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2610.11765}
}