86bad36ffe
Updates the SkillOpt-Sleep plugin on top of the current main. User-facing and engine improvements since the initial drop: * Command renamed /sleep -> /skillopt-sleep across Claude Code + Codex shells; refreshed plugin READMEs and install scripts. * Built-in scheduling (skillopt_sleep/scheduler.py + __main__): schedule / unschedule the nightly cycle without external cron wiring. * Backend robustness: bounded retry with backoff (no more silent empty-string on transient 429/timeout), content-filter-safe rollout prompt, an output-contract guardrail that rejects edits violating the task's required format, and a per-sample cache key so repeated dream rollouts are independent samples (fixes degenerate single-sample reflection). * consolidate / rollout / replay: parallel multi-rollout dreaming, gate-mode controls, TaskRecord.system framing field. Scope: this commit ships only the plugin engine + shells. Research/benchmark harnesses and their data are intentionally not included; the public package has no dependency on them (the one research-evaluator import is now guarded). Marked as an early preview in the README; we'll keep iterating. 99/99 unit tests pass. Co-Authored-By: Claude Opus 4 <noreply@anthropic.com>
2.0 KiB
2.0 KiB
name, description
| name | description |
|---|---|
| skillopt-sleep | Nightly offline self-evolution for a Codex agent. Reviews past sessions, replays recurring tasks, and consolidates validated memory + skills behind a held-out gate. Use when the user wants Codex to learn from past usage, run a "sleep"/"dream" cycle, or schedule offline self-optimization. |
SkillOpt-Sleep (Codex skill)
This skill drives the skillopt_sleep engine — an offline "sleep cycle" that
makes a Codex agent better at the user's recurring work without retraining.
When to use
Trigger when the user wants to: review past sessions, learn their preferences, consolidate feedback into long-term memory/skills, run a nightly/offline self-improvement cycle, or adopt a staged proposal.
How to run it
Invoke the bundled runner via shell (Codex exec has shell access). The runner
finds the engine and a Python ≥ 3.10 automatically:
# point at the repo if it isn't auto-detected from CWD:
export SKILLOPT_SLEEP_REPO=/path/to/SkillOpt-Sleep
bash "$SKILLOPT_SLEEP_REPO/plugins/run-sleep.sh" <action> --project "$(pwd)"
<action> ∈ status | dry-run | run | adopt | harvest. Use --backend codex
for real improvement on the user's own Codex budget (default mock = no spend).
Steps
- Run the requested action; capture stdout.
- For
run/dry-run: read the stagedreport.mdit prints and show the user the held-out baseline → candidate score and the exact proposed edits. runonly stages a proposal under<project>/.skillopt-sleep/staging/; nothing live changes untiladopt. Offer/skillopt-sleep adopt.- Never hand-edit the user's
AGENTS.md/ skills yourself — onlyadoptdoes, and it backs up first.
Validate
python -m skillopt_sleep.experiments.run_gbrain --backend codex \
--seeds brief-writer --data-root /path/to/gbrain-evals/eval/data/skillopt-v1 \
--nights 2 --limit-replay 3 --limit-holdout 3
A deficient skill goes 0.00 → 1.00 on a held-out set; the optimizer's edits are gated on real-task performance.