BottleCap AI is betting that shorter reasoning is a feature rather than a compromise. Its latest release, ThinkingCap-Qwen3.8-27B, is the second model in the ThinkingCap line and a fine-tune of Qwen’s Qwen3.8-27B. The company’s own framing is unusually candid: it gives up a little accuracy on purpose, in exchange for thinking far less.
The headline trade is 37.2 percent fewer thinking tokens on average across 12 benchmarks, against a macro-average slide of 0.86 percentage points, from 86.65 percent to 85.79 percent.
The design target was narrow. BottleCap avoided adding knowledge or reshaping answer style, so reasoning, instruction following and safety behaviour sit close to the base model. Effort went into math, reasoning, long-context and agentic tests instead. An earlier release in the line did the same for Qwen3.6-27B.
Shorter traces show up everywhere, from a 10.7 percent trim to 65.5 percent. Multilingual and knowledge items shrink hardest: MMMLU gives up 65.5 percent of its tokens, MMLU-Pro 57.3 percent, and GPQA-Diamond 43 percent.
Two results move the other way. On AA-LCR, long-context retrieval gains 2.25 points while thinking 38.6 percent less. LiveCodeBench v6 adds 0.07 points on 20.3 percent fewer tokens.
The worst trade is AIME 2026, where accuracy gives up 3.85 points for 30.2 percent less thinking. Pooled tokens drop from 15,735 to 12,144. Builds cover vLLM and SGLang in FP8, NVFP4, GGUF and MLX, the repository gated, with commercial use above the small-business licence needing BottleCap’s agreement.