a Claude Code mod

molt

When context outgrows its shell, shed it into the folder and keep going.

Install

Paste this into Claude Code. The agent runs the two commands and tells you when /molt is there.

Install the molt mod for Claude Code:
1. claude plugin marketplace add wisbech/molt
2. claude plugin install molt@molt
3. If the build says hooks modules are early access, add "CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1" under "env" in ~/.claude/settings.json.
Then tell me whether /molt shows up in the command list.
Or by hand: claude plugin marketplace add wisbech/molt then claude plugin install molt@molt. Or claude --plugin-dir path/to/molt.

The problem, measured

If your usage looks like this, the cache is not your problem:

Tokens, one project, one week158.7M
Cache hit rate96%
Cache read152.2M
Cache write5M
Uncached input0.7M
Output0.8M

That is a healthy cache. Only 4% of input missed it. Yet 96% of all tokens were cache reads, and cache reads are what you pay for a long thread: every API call resends the whole conversation. A hit is cheaper than a miss, but half a million tokens of hits per call, hundreds of times a day, is still the bill.

Parsing the session transcripts found where it went. One thread had been kept open for seven days:

The seven-day thread
API calls618
Average context per call~480k
Peak context966k of 1M
Cache reads296M
Full cache rewrites14 × 550k–850k
Subagents, all on the most expensive model61

1. Thread length

Every tool call re-read ~480k tokens. A fresh session costs ~70k. Resuming the thread cost 480k on the first message and on every call after it. A sibling thread averaging 230k cost a quarter as much. Spend tracks thread length.

2. The cache expiring while you sleep

The prompt cache lives one hour. Every morning's first message rewrote the entire context at full write price. Fourteen such rewrites, each after an idle gap over an hour. Auto-compact never helped: on a 1M model it fires near 967k.

What it was not: the personal harness, hooks and rituals were under 4% of the tokens. Measure before you blame your tooling.

The mechanism

Two criteria, one call.

  1. Idle. Fifty minutes into an idle stretch, while the cache is still warm and reads are cheap, the session molts. When you come back the rewrite is of a 60k context, not 500k.
  2. Size. Context above a ceiling that molt solves, not sets. Over a stretch between molts, tokens per call ≈ (F+T)/2 + m·T·g/(T−F): F the floor after a molt, g the growth per API call, m the reads a molt costs. The minimum is T = F + √(2·m·g·F). Molt measures all three and re-solves after every turn: light chat lands near 95k, heavy subagent work near 165k on a 73k floor.

The call is /molt. The model saves state to .state/; the mod compacts the session in place, keeping only what the handoff does not cover. Same session, same prompt box, a fraction of the tokens.

.state/
  INDEX.md      routing table: one line per note, says WHEN to read it (cap 200)
  HANDOFF.md    where the last stretch stopped (cap 60; rewritten by each molt)
  log.md        append-only history; grep it, never load it whole
  notes/*.md    decisions, failures, patterns; frontmatter links form the graph

INDEX.md and a fresh HANDOFF.md ride the first user message next to CLAUDE.md. The folder is plain markdown: any agent can continue from it.

Use

Measure your own

Reads your transcripts, prints the five biggest threads. If avg_ctx is a few hundred k, you have the problem above.

python3 - <<'EOF'
import json, glob, os
d = os.path.expanduser('~/.claude/projects/')
for f in sorted(glob.glob(d + '*/*.jsonl'), key=os.path.getsize, reverse=True)[:5]:
    seen, n, cr, cw, mx = set(), 0, 0, 0, 0
    for line in open(f):
        try: o = json.loads(line)
        except: continue
        m = o.get('message') or {}; u = m.get('usage') or {}
        if o.get('type') != 'assistant' or not u or m.get('id') in seen: continue
        seen.add(m.get('id')); n += 1
        ctx = u.get('input_tokens', 0) + u.get('cache_read_input_tokens', 0) + u.get('cache_creation_input_tokens', 0)
        cr += u.get('cache_read_input_tokens', 0); cw += u.get('cache_creation_input_tokens', 0); mx = max(mx, ctx)
    print(f"{os.path.basename(f)[:8]} calls={n} avg_ctx={(cr+cw)/max(n,1)/1e3:.0f}k peak={mx/1e3:.0f}k cache_read={cr/1e6:.0f}M cache_write={cw/1e6:.1f}M")
EOF