molt
When context outgrows its shell, shed it into the folder and keep going.
Install
Paste this into Claude Code. The agent runs the two commands and tells you when /molt is there.
Install the molt mod for Claude Code:
1. claude plugin marketplace add wisbech/molt
2. claude plugin install molt@molt
3. If the build says hooks modules are early access, add "CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1" under "env" in ~/.claude/settings.json.
Then tell me whether /molt shows up in the command list.
claude plugin marketplace add wisbech/molt then claude plugin install molt@molt. Or claude --plugin-dir path/to/molt.
The problem, measured
If your usage looks like this, the cache is not your problem:
| Tokens, one project, one week | 158.7M |
| Cache hit rate | 96% |
| Cache read | 152.2M |
| Cache write | 5M |
| Uncached input | 0.7M |
| Output | 0.8M |
That is a healthy cache. Only 4% of input missed it. Yet 96% of all tokens were cache reads, and cache reads are what you pay for a long thread: every API call resends the whole conversation. A hit is cheaper than a miss, but half a million tokens of hits per call, hundreds of times a day, is still the bill.
Parsing the session transcripts found where it went. One thread had been kept open for seven days:
| The seven-day thread | |
|---|---|
| API calls | 618 |
| Average context per call | ~480k |
| Peak context | 966k of 1M |
| Cache reads | 296M |
| Full cache rewrites | 14 × 550k–850k |
| Subagents, all on the most expensive model | 61 |
1. Thread length
Every tool call re-read ~480k tokens. A fresh session costs ~70k. Resuming the thread cost 480k on the first message and on every call after it. A sibling thread averaging 230k cost a quarter as much. Spend tracks thread length.
2. The cache expiring while you sleep
The prompt cache lives one hour. Every morning's first message rewrote the entire context at full write price. Fourteen such rewrites, each after an idle gap over an hour. Auto-compact never helped: on a 1M model it fires near 967k.
What it was not: the personal harness, hooks and rituals were under 4% of the tokens. Measure before you blame your tooling.
The mechanism
Two criteria, one call.
- Idle. Fifty minutes into an idle stretch, while the cache is still warm and reads are cheap, the session molts. When you come back the rewrite is of a 60k context, not 500k.
- Size. Context above a ceiling that molt solves, not sets. Over a stretch between molts, tokens per call ≈ (F+T)/2 + m·T·g/(T−F): F the floor after a molt, g the growth per API call, m the reads a molt costs. The minimum is
T = F + √(2·m·g·F). Molt measures all three and re-solves after every turn: light chat lands near 95k, heavy subagent work near 165k on a 73k floor.
The call is /molt. The model saves state to .state/; the mod compacts the session in place, keeping only what the handoff does not cover. Same session, same prompt box, a fraction of the tokens.
.state/
INDEX.md routing table: one line per note, says WHEN to read it (cap 200)
HANDOFF.md where the last stretch stopped (cap 60; rewritten by each molt)
log.md append-only history; grep it, never load it whole
notes/*.md decisions, failures, patterns; frontmatter links form the graph
INDEX.md and a fresh HANDOFF.md ride the first user message next to CLAUDE.md. The folder is plain markdown: any agent can continue from it.
Use
- Nothing, usually. At a criterion the session molts by itself.
- The bar above the prompt shows
ctx 132k · molts at 120kand a Molt now button. /molt [note]molts now./molt-lintis free housekeeping./molt-graphprints note links.- Pair with the native
/autocompact 150kandCLAUDE_CODE_SUBAGENT_MODEL=sonnet.
Measure your own
Reads your transcripts, prints the five biggest threads. If avg_ctx is a few hundred k, you have the problem above.
python3 - <<'EOF'
import json, glob, os
d = os.path.expanduser('~/.claude/projects/')
for f in sorted(glob.glob(d + '*/*.jsonl'), key=os.path.getsize, reverse=True)[:5]:
seen, n, cr, cw, mx = set(), 0, 0, 0, 0
for line in open(f):
try: o = json.loads(line)
except: continue
m = o.get('message') or {}; u = m.get('usage') or {}
if o.get('type') != 'assistant' or not u or m.get('id') in seen: continue
seen.add(m.get('id')); n += 1
ctx = u.get('input_tokens', 0) + u.get('cache_read_input_tokens', 0) + u.get('cache_creation_input_tokens', 0)
cr += u.get('cache_read_input_tokens', 0); cw += u.get('cache_creation_input_tokens', 0); mx = max(mx, ctx)
print(f"{os.path.basename(f)[:8]} calls={n} avg_ctx={(cr+cw)/max(n,1)/1e3:.0f}k peak={mx/1e3:.0f}k cache_read={cr/1e6:.0f}M cache_write={cw/1e6:.1f}M")
EOF