Thinking Inertia

LLMs Keep Thinking When Told Not To

Dianqiao Lei1, Kevin Qinghong Lin2†✉, Pan Lu3, Philip Torr2✉, James Zou3✉

1 Tsinghua University 2 University of Oxford 3 Stanford University

†Project lead  ·  ✉Correspondence

Flip the switch. Watch the visible response change.

Same question, same answer target, two answer interfaces. This precomputed illustrative trace makes the response-level distinction concrete: a native no-think request can still leave visible inference, while a strict answer-only interface removes the pre-answer field.

SAME QUESTION · TWO RESPONSE TRACESIllustrative open-ended example
Interactive response demo
QUESTION

What is 17 × 19?

PRE-ANSWER TEXT TEXPLICIT INFERENCE

17 × 20 = 340, then subtract 17. So the product is 323.

The answer is correct, but a visible inferential step remains before it.
FINAL ANSWER APARSED

323

Correctness is evaluated separately from the text before the answer.

A narrated walkthrough of the motivation, response-level decomposition, three evaluation levels, and main findings.

01 / MOTIVATION

A no-thinking switch does not guarantee an answer-only response

Before defining no-thinking, we ask a simpler question: when a model is told to return only the answer, is the visible pre-answer field actually empty? Figure 1 shows that the requested mode and the emitted response can diverge.

(a) Instruction–output mismatch(b) Answer-space gradient
Representative answer-only instruction failure and difficulty of adhering to no-thinking across benchmarks

Figure 1. No-thinking controls do not guarantee response-level no-thinking. (a) A clear ``answer only'' instruction still elicits question-relevant pre-answer content from DeepSeek-V4-Flash and Qwen3-235B-A22B on a representative MMLU question. (b) Under native no-thinking, ETR and visible token length show that adherence is closest to answer-only for Boolean tasks, then falls across MCQ and Open-ended tasks.

ETR (answer-only compliance): Bool > MCQ > OpenVisible pre-answer length: Bool < MCQ < Open
InstructionAnswer-only request``Put only the final option label in \boxed{}.''
Observed responseCorrect answer + pre-answer textthe behavior we can actually inspect
The key distinction

A no-thinking control specifies the request—not the response. Even with a clear answer-only instruction, a model may expose question-relevant text before a correct final answer; under native no-thinking, this exposure changes systematically with the answer space.

01 / INSTRUCTION FAILURE

The instruction–response gap

Both models return the correct option, yet expose question-relevant text before it. Correctness therefore does not certify answer-only compliance.

02 / ANSWER-SPACE GRADIENT

The answer space changes compliance

Under native no-thinking, Boolean outputs stay closest to the answer-only boundary, MCQ is intermediate, and Open-ended outputs retain the most visible pre-answer text.

These observations motivate the response-level object used below: separate the pre-answer text T from the final answer A, then distinguish empty output, relevant context, and explicit inference.

Abstract

Large Language Models (LLMs) increasingly ship with explicit “thinking modes,” yet their counterpart, “no-think,” has received far less attention. In this work, we study LLMs’ “no-think” behavior along two axes.

a. How to measure “no-thinking”? Prior work typically defines “no-thinking” through proxies such as a disabled thinking mode or the absence of long traces. These proxies are unreliable: disabled thinking modes still emit extended reasoning, or long traces often contain uninformative filler rather than genuine inference. We instead normalize any LLM response into pre-answer trace and final answer, and evaluate it at three levels: (i) an Empty-Thinking Ratio (ETR) for strict answer-only compliance; (ii) an instruction-aware Question–Pre-answer Relevance (QRel.) score to measure similarity between the question and pre-answer trace; and (iii) an additional LLM-as-judge Explicit Inference Rate (EIR) for visible explicit inference. Together, these levels separate answer-only output, relevant but non-inferential text, and explicit inference. Human studies demonstrate the effectiveness of these metrics.

b. How does no-thinking vary across tasks and models? We evaluate six prompting intervention settings on six widely used LLMs, both open-source and commercial, across three question types: Boolean, multiple-choice, and open-ended. We find that explicit no-think controls cannot reliably eliminate visible explicit inference; models instead exhibit “Thinking Inertia”: visible explicit inference persists even under strict no-thinking controls and becomes progressively more prevalent as the answer space opens. Accuracy remains relatively stable on Boolean and multiple-choice tasks, whereas open-ended tasks expose a trade-off between answer-only compliance and task accuracy. Finally, rewriting the same questions across answer spaces shows that supplying candidate answers makes answer-only responses substantially easier to produce. Taken together, these findings establish no-thinking as a non-trivial capability of LLMs: stopping explicit reasoning cannot be assumed from model settings or instructions alone, and deserves systematic evaluation alongside reasoning ability itself.

02 / DIAGNOSIS

Surface proxies describe the interface—not the response

No-thinking is often inferred from a prompt-side switch, a missing <think> block, or a short generation. Figure 2 shows why each proxy can fail while the emitted response still exposes an inferential step.

Surface proxies for no-thinking and their corresponding failure cases

Figure 2. Surface proxies fail to identify response-level no-thinking. (a) Prompt-side control, special reasoning tokens, and generation length are convenient but indirect signals. (b) On the same algebra question, disabling the thinking mode, omitting special tokens, or producing a short output can each leave a visible inferential step before the final answer.

01
Prompt-side controlNo-thinking mode, but reasoning remains

Disabling the mode can still produce ``Subtract 1 from both sides...'' before the answer.

02
Visible-format controlNo <think> token, but reasoning remains

Ordinary answer-channel text can still explain how to isolate x.

03
Length-based controlShort output, but compact reasoning

A brief ``3x = 9, so x = 3'' still exposes the decisive intermediate step.

The measurement gap

A proxy can tell us how the response was requested, formatted, or sized; it cannot by itself tell us what question-conditioned content the response exposes.

These proxies are useful controls, but none defines the behavioral object we want to study. We therefore move from surface conditions to the visible pre-answer field itself.

03 / RESPONSE OBJECT

Measure what the response exposes

We turn no-thinking into an observable response-level object by separating the visible pre-answer text from the final answer, without making a claim about latent computation.

QUESTIONQthe task sample
→
INTERVENTIONS(Q)prompt, mode, or answer interface
→
MODELπthe language model
→
RESPONSE(T, A)visible pre-answer text and final answer
Response-level measurement framework showing question, intervention, model, pre-answer text, final answer, and measurement labels

Figure 3. Response-centric metrics for measuring no-thinking. Top: a question Q is wrapped by an intervention S(·) before the model produces a response decomposed into pre-answer text T and final answer A. Bottom: six responses share the same question and final answer while varying in length, relevance, and inferential content.

01 / HOLD CONSTANT

Same question, same answer

The six cases keep Q and A fixed. What changes is the visible text before the answer.

02 / INSPECT T

Different text, different behavior

T can be empty, generic, relevant but non-inferential, or explicitly inferential. Non-empty text is not automatically reasoning.

The observable target

We evaluate T directly and evaluate answer correctness on A separately. This keeps visible reasoning claims at the response level, rather than treating a prompt setting or hidden computation as the measurement target.

Once T is explicit, the next question is not simply whether text exists, but what kind of text it contains: empty output, relevant context, or visible explicit inference.

04 / THREE-LEVEL MEASUREMENT

One visible trace, three increasingly specific questions

The same pre-answer field T is read at three levels: whether it is empty, whether it is related to Q, and whether it contains an explicit inferential step.

same observableTpre-answer text
01T = ∅?strict compliance→02Q ↔ T?relevance→03Inference?visible reasoning
01/ETR

Strict answer-only

ASKS · Is the complete response answer-only?

T = ∅ and the final answer A is parseable; any substantive visible text outside the permitted answer wrapper breaks the boundary.

binary · all-output ratedoes not claim anything about latent computation
02/QRel.

Question–pre-answer relevance

ASKS · How related is T to answering Q?

The instruction-aware Qwen3-Embedding-4B encoder reads Q as the query and T as the candidate passage; empty T receives zero.

continuous · all-output meanrelevance signal, not a reasoning detector
03/EIR

Visible explicit inference

ASKS · Does T make an inferential move?

A blinded judge sees only (Q, T); only the mutually exclusive Explicit reasoning label contributes to the rate, while empty T is assigned zero.

binary · all-output ratevisible response evidence, not latent cognition
Generic / off-topicParaphrase-onlyRelevant non-inferentialExplicit reasoningUnclear
Same Q, three kinds of T

Why the levels should not be collapsed

EMPTYT = ∅
ETR = 1QRel. = 0 · EIR = 0
RELEVANT ASSERTIONT = ``The answer is 121.''
ETR = 0QRel. may be nonzero · EIR = 0
INFERENTIAL STEPT = ``2 + 9 = 11, carry 1.''
ETR = 0QRel. can be high · EIR = 1
QRel. ≠ reasoningvisible response ≠ latent cognitionall metrics use all outputs

ETR tells us whether T exists; QRel. describes what T is about; EIR asks whether T makes an explicit inferential move. Accuracy is evaluated separately on the final answer A.

05 / EXPERIMENTAL STORY

What the experiments reveal

We use Thinking Inertia for a specific response-level behavior, then test whether it survives changes in answer space and no-thinking control.

Working definition

Thinking Inertia

the persistence of visible, question-conditioned pre-answer text that contains explicit inference after the model has been told not to think. It is a claim about the emitted response—not latent cognition—and is measured through EIR, with ETR and QRel. providing complementary context.

No-thinking instruction→Visible pre-answer text T→Explicit inference
01Answer space

More open interfaces preserve more visible inference

Across native and stricter no-thinking settings, the response-level signal follows a staircase: Boolean tasks are closest to answer-only, multiple-choice tasks are intermediate, and open-ended tasks retain the most explicit pre-answer work.

Bool<MCQ<Openvisible inference increases with answer freedom
02Mode effect

Stricter controls reduce the behavior, but do not erase it

Within every answer space, stronger no-thinking control yields lower explicit inference than the native no-thinking setting. The remaining Open-ended signal is the response-level signature of thinking inertia.

M5<M2M2 = native think-off · M5 = strict answer-only; relation holds for Bool, MCQ, and Open
All-output viewCompliance · relevance · inference · length
ETR, EIR, QRel., and visible token length across answer spaces and intervention modes

Results overview. The three-level view pairs strict compliance, visible explicit inference, instruction-aware relevance, and pre-answer length across answer spaces and intervention modes.

03 / ACCURACY

Compliance and capability are separate axes

A cleaner answer-only format does not automatically mean a more capable answer. Accuracy on A is read alongside ETR, QRel., and EIR.

04 / CONTROL

Matched rewrites isolate the interface

When the same question is rewritten as Boolean, multiple-choice, or open-ended, short targets alone do not explain the pattern; matched lengths can still differ sharply in relevance.

06 / VALIDATION

Checks that keep the interpretation narrow

The evidence is layered so that no single proxy has to carry the reasoning claim on its own.

Q
Relevance robustness

Alternative encoders preserve the relevance ordering, while the instruction-aware QRel. signal remains explicitly separated from reasoning.

J
Judge robustness

Claude Opus 4.8 independently reproduces the Level 3 visible-inference pattern across models, modes, and answer spaces.

H
Human validation

Five blinded annotators review a proportion-matched 600-output sample using the same relevance and reasoning rubric.

Reading guide

ETR marks the answer-only boundary; QRel. measures question-relevant residual text; EIR identifies visible explicit inference. The claim stays at the level of observable response behavior, not latent cognition.

📋 BibTeX

Read the paper on arXiv (2610.11765).

@article{lei2026thinkinginertia,
  title   = {Thinking Inertia: LLMs Keep Thinking When Told Not To},
  author  = {Lei, Dianqiao and Lin, Kevin Qinghong and Lu, Pan
             and Torr, Philip and Zou, James},
  year    = {2026},
  eprint  = {2610.11765},
  archivePrefix = {arXiv},
  url     = {https://arxiv.org/abs/2610.11765}
}