Thinking Mode Induces Confidence Compression in Reasoning LLMs: A Pre-Registered Type-2 SDT Analysis
Jon-Paul Cacioli · 2026
· DOI: 10.5281/zenodo.20394749
Abstract
Pre-registered study (Stage 1 of 2) testing whether thinking mode improves the discriminative power of token-probability-derived confidence signals. Three 7-8B reasoning models (Qwen3-8B, Phi-4-reasoning-plus, DeepSeek-R1-Distill-Llama-8B) answered 2,000 TriviaQA items in thinking and non-thinking mode. Thinking mode improved accuracy but significantly reduced NLP-based AUROC2 in all three models (Qwen3-8B: delta = -0.231, 95% CI [-0.262, -0.198]). The mechanism is NLP compression: thinking-mode NLP variance dropped to 16% of non-thinking levels. Verbal confidence (Stage 2, in progress) may behave differently. These results suggest that reasoning-mode evaluation of metacognitive monitoring requires probe-specific analysis.
← Back to synthiumjp.github.io