drlee1's picture
Trim card to the merged-model usage path
9ec6e4f verified
|
Raw
History Blame Contribute Delete
5.66 kB
metadata
base_model: Qwen/Qwen3.5-2B
library_name: transformers
license: apache-2.0
pipeline_tag: text-generation
tags:
  - token-efficient
  - efficient-thinking
  - reasoning
  - grpo
  - lora
  - peft

ThinkingCap-Qwen3.5-2B (loss-level variant)

A merged, ready-to-serve fine-tune of Qwen3.5-2B trained with GRPO to reach the same answers with 43% fewer thinking tokens β€” while improving accuracy by +13 to +17 points, because the dominant failure of the base model is not verbosity but never finishing its reasoning at all.

This adapter is the loss-level variant of a controlled study on how to compress chain-of-thought at small scale. It combines three modifications to the GRPO objective β€” Dr.GRPO normalization, positional advantage decay inside the reasoning span, and KL restricted to that span β€” on top of a correctness-gated length penalty. The companion study that isolates each component is linked at the bottom.


In-domain results

Measured under a natural-length protocol: greedy decoding, a 4,096-token budget, and no forced truncation of thinking. A response whose reasoning span never closes contains no answer and is scored incorrect β€” so token reduction and accuracy are read off the same generations.

Benchmark Base Acc Ours Acc Base Thinking Tokens Ours Thinking Tokens Reduction
Held-out dev (300 problems) 69.0 82.3 2,922 1,656 ↓ 43.3%
GSM8K 50.0 66.7 3,545 2,438 ↓ 31.2%
ARC-Challenge 70.8 87.5 2,216 1,247 ↓ 43.7%

Out-of-domain

Benchmark Base Acc Ours Acc Base Thinking Tokens Ours Thinking Tokens Reduction
MATH-500 20.8 20.8 4,096 3,320 ↓ 19.0%

Out of distribution the adapter preserves accuracy exactly while cutting a fifth of the reasoning budget. Compression transfers; the accuracy gains do not β€” a 2B model has little headroom on competition mathematics either way.

Failure modes

The base model's real pathology is non-termination: it finds the answer early, then loops on self-verification and never emits a closing marker.

Metric (out-of-domain) Base Ours
Reasoning never closes within budget 100% 75%
Terminates normally (EOS) 0% 25%

On in-domain data the effect is far stronger β€” unclosed reasoning drops from 58% to 21%. Most of the accuracy gain is simply answers that now exist.

Training

Base model Qwen/Qwen3.5-2B (thinking mode)
Method GRPO + LoRA (rank 16, Ξ± 32, bf16), 200 optimizer steps
Reward correctness-gated length penalty; zero reward when the answer is wrong or the reasoning span never closes
Objective modifications Dr.GRPO normalization Β· positional advantage decay over the thinking span Β· KL restricted to the thinking span
Data 10.7K verifiable problems (GSM8K + ARC training splits), deduplicated against every evaluation set
Hardware 1Γ— NVIDIA RTX 5080 (16 GB)

No generation-time intervention is used at any point β€” no forced stopping, no length cap on thinking. The model learns to terminate purely from the reward.

Usage

The repository root holds the merged model β€” load it like any Qwen3.5 checkpoint, no PEFT required.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "drlee1/ThinkingCap-Qwen3.5-2B"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16, device_map="auto")

messages = [{"role": "user", "content": "If a train travels 120 km in 1.5 hours, what is its average speed?"}]
inputs = tok.apply_chat_template(
    messages, add_generation_prompt=True, enable_thinking=True,
    return_dict=True, return_tensors="pt",
).to(model.device)

out = model.generate(**inputs, max_new_tokens=2048, do_sample=False)
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

enable_thinking=True is required β€” Qwen3.5-2B runs in non-thinking mode by default, and this model is trained on the thinking-mode chat template.

Honest scope

  • Single training seed, 200 optimizer steps. Numbers are a screening result, not a multi-seed benchmark. Evaluation uses 300 held-out problems for the dev set and 24 items per public benchmark.
  • A simpler recipe is stronger. In the companion ablation, dropping all three loss-level modifications and keeping only the reward-level length penalty reaches 79% token reduction at 90.7% dev accuracy β€” better on both axes than this model. Each objective modification measurably costs compression in this regime (2B, LoRA, small rollout groups). This model is released as the loss-level arm of that comparison, not as the recommended configuration.
  • Findings are specific to this scale and setup; they are not a general verdict on the underlying techniques in their original settings.

Citation

@misc{reducing-think-token-2026,
  title  = {Reducing Thinking Tokens via Reinforcement Learning:
            A Controlled Study on Loss-Level Interventions},
  author = {DONGRYEOLLEE1},
  year   = {2026},
  url    = {https://github.com/DONGRYEOLLEE1/reducing-think-token}
}

Acknowledgements. Inspired by ThinkingCap-Qwen3.6-27B (BottleCapAI). Built on Qwen3.5-2B, TRL, and PEFT. The full study β€” including the component-wise ablation with confidence intervals β€” is on GitHub.