Add skillopt/sleep — a deployment-time companion to SkillOpt that gives a
local Claude agent a nightly "sleep cycle":
harvest ~/.claude transcripts -> mine recurring tasks -> replay offline
-> consolidate (reflect -> bounded edit -> held-out GATE) -> stage -> adopt
Synthesizes SkillOpt (validation-gated bounded text optimization, reusing
skillopt.evaluation.gate verbatim), Claude Dreams (offline consolidation;
input never mutated; review-then-adopt), and the agent-sleep paper
(short-term experience -> long-term competence).
Engine (skillopt/sleep/, import-light, py>=3.10):
- harvest.py read-only parse of session JSONL + history.jsonl
- mine.py sessions -> TaskRecords (heuristic miner + LLM hook)
- backend.py MockBackend (deterministic, no API) + AnthropicBackend
- replay.py offline re-run -> (hard, soft) scores
- consolidate.py one SkillOpt epoch behind a held-out gate
- memory.py protected-region edits to SKILL.md / CLAUDE.md
- staging.py stage proposals; adopt with backup (Dreams safety contract)
- cycle.py + __main__.py orchestrator + CLI (run/dry-run/status/adopt/harvest)
Plugin (skillopt-sleep-plugin/): plugin.json, /sleep command, skillopt-sleep
skill, SessionEnd hook, bundled runner + cron generator.
Validation (deterministic, no API): persona experiment proves held-out lift
(researcher 0.33->1.0, programmer 0.32->1.0) AND that the gate rejects an
injected harmful edit. 13 stdlib-unittest tests pass, incl. full cycle +
adopt-with-backup and parsing of real on-disk transcripts.
Co-Authored-By: Claude Opus 4 <noreply@anthropic.com>
4.6 KiB
name, description
| name | description |
|---|---|
| skillopt-sleep | Use when the user wants their Claude agent to self-improve from past usage, asks about a nightly/offline 'sleep' or 'dream' cycle, memory/skill consolidation, or says things like '让 agent 越用越好用', 'review my past sessions', 'learn my preferences', 'consolidate what you learned', 'run the sleep cycle', or wants to schedule offline self-optimization. Drives the skillopt.sleep engine: harvest past sessions → mine recurring tasks → replay offline → consolidate validated CLAUDE.md/SKILL.md behind a held-out gate. |
SkillOpt-Sleep: offline self-evolution for a local Claude agent
SkillOpt-Sleep gives the user's agent a sleep cycle. While the user is
offline (e.g. nightly), it reviews their real past Claude Code sessions,
re-runs recurring tasks on their own API budget, and consolidates what it
learns into memory (CLAUDE.md) and skills (SKILL.md) — but only
keeps changes that pass a held-out validation gate, and only after the user
adopts them. The agent gets measurably better at this user's recurring work,
with no model-weight training. It is the deployment-time analogue of training:
short-term experience → long-term competence.
It synthesizes three ideas:
- SkillOpt — the skill/memory doc is trainable text; bounded add/delete/replace edits; accepted only through a held-out gate; rejected edits become negative feedback.
- Claude Dreams — offline consolidation that reads past sessions and rebuilds memory (dedup/merge/resolve); the input is never mutated; output is reviewed then adopted.
- Agent sleep — periodic offline replay turns episodes into durable skill.
When to use this skill
Trigger when the user wants any of:
- "make my agent learn from how I use it" / "越用越好用" / "remember my preferences across sessions"
- a nightly/scheduled or on-demand offline self-improvement / dream / sleep run
- to review past sessions/trajectories and distill recurring tasks
- to consolidate feedback into
CLAUDE.mdor a managed skill - to schedule the cycle (cron) or adopt a staged proposal
The cycle (six stages)
- Harvest — read
~/.claude/projects/*/<session>.jsonl+~/.claude/history.jsonl(READ-ONLY) → session digests. - Mine — digests →
TaskRecords (recurring intents + outcome labels + checkable refs where possible). - Replay — re-run tasks offline under the current skill+memory → (hard, soft) scores.
- Consolidate — reflect on failures → propose bounded edits → gate on a held-out slice; accept only if it strictly improves.
- Stage — write
proposed_CLAUDE.md,proposed_SKILL.md, a diff, andreport.mdinto<project>/.skillopt-sleep/staging/<date>/. Nothing live changes. - Adopt — explicit (or opt-in auto): copy staged files over live ones, backing up first.
How to drive it
Prefer the /sleep command. Under the hood it calls the bundled runner:
"${CLAUDE_PLUGIN_ROOT}/scripts/sleep.sh" status # what's happened
"${CLAUDE_PLUGIN_ROOT}/scripts/sleep.sh" dry-run --project "$(pwd)" # safe preview
"${CLAUDE_PLUGIN_ROOT}/scripts/sleep.sh" run --project "$(pwd)" # full cycle, stages a proposal
"${CLAUDE_PLUGIN_ROOT}/scripts/sleep.sh" adopt --project "$(pwd)" # apply staged proposal (with backup)
- Default backend is
mock(deterministic, no API spend) — good for trying the plumbing. - Add
--backend anthropicto spend the user's real budget for genuine improvement. - Scope defaults to the invoked project;
--scope allharvests every project.
Hard rules
- Never hand-edit the user's
CLAUDE.md/SKILL.mdas part of this skill. Only theadoptaction changes live files, and it backs them up first. - Harvest is read-only.
mockreplay has no side effects. - Always show the user the held-out baseline → candidate score and the exact proposed edits before suggesting adoption. Evidence before adoption.
- If asked whether it really helps, run
python -m skillopt.sleep.experiments.run_experiment --persona researcher --json— a deterministic demo that proves held-out lift and that the gate blocks harmful edits.
Validate / demo
# deterministic proof (no API): held-out score rises, gate blocks regressions
python -m skillopt.sleep.experiments.run_experiment --persona researcher --assert-improves
python -m skillopt.sleep.experiments.run_experiment --persona programmer --assert-improves
See docs/sleep/experiment_results.md for recorded output and
docs/superpowers/specs/2026-06-07-skillopt-sleep-claude-code-plugin-design.md
for the full design.