I just read this piece on Chain of Draft: Thinking Faster by Writing Less…basically force your LLM to keep each reasoning step under 5 words. Same accuracy. 7% of the tokens. And token management is a big deal.
Turns out most of what we mistake for thinking is the model apologising for existing: https://arxiv.org/pdf/2502.18600
Discover more from soulcruzer
Subscribe to get the latest posts sent to your email.