feat(plugins): add OpenClaw shell for SkillOpt-Sleep
Adds a thin OpenClaw shell wrapping the SkillOpt-Sleep engine. Enables nightly validation-gated skill improvement cycles for OpenClaw agents. Components: - skillopt_sleep_openclaw.py: DeepSeek V4 Pro + Ollama nomic-embed-text backend, mirroring the Claude/Codex/Copilot backend pattern. - run_sleep.py: CLI entry point supporting dry-run and pre-built task files. - run_sleep_cron.sh: bash wrapper for nightly cron invocation. - slash_sleep.py: /sleep command (status / run / adopt / reject / cost). - config.json: engine config tuned for our stack. - SKILL.md: OpenClaw skill manifest. - tests/: 14 held-out tasks across 3 categories (research-cron, devops, wiki). OpenClaw is the 4th ecosystem in which SkillOpt-Sleep can be deployed, joining Claude Code, Codex, and Copilot. The shell follows the same single-engine / thin-shell pattern as the existing three plugins. End-to-end tested: pipeline runs against real OpenClaw session transcripts, gate correctly rejects non-improvements, staging artifacts land in ~/.skillopt-sleep/staging/<night>/. Cost: ~$0.02/night on DeepSeek V4 Pro.
This commit is contained in:
@@ -0,0 +1,96 @@
|
||||
---
|
||||
name: skillopt-sleep
|
||||
description: Validate and refine agent skills through nightly sleep cycles with held-out gates. Wraps Microsoft's SkillOpt-Sleep engine for the OpenClaw/DeepSeek stack.
|
||||
---
|
||||
|
||||
# skillopt-sleep — OpenClaw Adaptation of Microsoft SkillOpt-Sleep
|
||||
|
||||
A nightly self-improvement loop that reads our session transcripts, mines recurring workflow patterns, replays them with proposed skill edits, and gates the proposals against a held-out test set. Only improvements that beat baseline are staged for human adoption.
|
||||
|
||||
## When To Use
|
||||
|
||||
- After Hermes's Weekly Skill Review (or as its replacement)
|
||||
- When a skill is being used 10+ times/week and could be tighter
|
||||
- Before promoting a new skill from `skill-proposals/` to `skills/`
|
||||
- When a skill regresses in observed quality
|
||||
|
||||
## What It Does (One Cycle)
|
||||
|
||||
```
|
||||
harvest session transcripts -> mine recurring task patterns
|
||||
-> replay each pattern (current skill vs proposed)
|
||||
-> GATE: must improve held-out score
|
||||
-> stage proposal
|
||||
-> Ethan adopts (manual)
|
||||
```
|
||||
|
||||
Nothing live changes until Ethan adopts. Every adopt backs up first.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
skills/skillopt-sleep/
|
||||
├── SKILL.md # this file
|
||||
├── config.json # engine config (backend, budgets, etc.)
|
||||
├── run_sleep.py # entry point
|
||||
└── skillopt_sleep_openclaw.py # DeepSeek/Ollama backend
|
||||
```
|
||||
|
||||
The engine itself is at `~/.openclaw/workspace/SkillOpt/skillopt_sleep/` (cloned from microsoft/SkillOpt).
|
||||
|
||||
## Usage
|
||||
|
||||
```bash
|
||||
# Run one cycle with current config
|
||||
cd ~/.openclaw/workspace/skills/skillopt-sleep
|
||||
python3 run_sleep.py
|
||||
|
||||
# Dry run (report only, no staging)
|
||||
python3 run_sleep.py --dry-run
|
||||
|
||||
# Use a pre-built task set (recommended for testing)
|
||||
python3 run_sleep.py --tasks tests/research-cron-tasks.json
|
||||
```
|
||||
|
||||
## Config (config.json)
|
||||
|
||||
Key knobs:
|
||||
- `backend: "openclaw-deepseek"` — our custom backend
|
||||
- `model: "deepseek-v4-pro"` — optimizer model
|
||||
- `edit_budget: 3` — max bounded edits per night
|
||||
- `gate_mode: "on"` — validation-gated (rejects regressions)
|
||||
- `auto_adopt: false` — require Ethan to adopt manually
|
||||
- `max_tasks_per_night: 12` — cap to control cost
|
||||
|
||||
## Cost Estimate
|
||||
|
||||
Per night: 12 tasks × (1 attempt + 1 judge + 1 reflect) × ~$0.005/1K tokens × ~3K tokens/call ≈ **$0.50-2.00/night**.
|
||||
|
||||
## Outputs
|
||||
|
||||
- Report: `~/.skillopt-sleep/state.json` (running totals)
|
||||
- Staging: `~/.skillopt-sleep/staging/<night>/`
|
||||
- `report.md` — readable summary
|
||||
- `best_skill.md` — proposed skill
|
||||
- `edits.json` — bounded edit list
|
||||
- `before.md` / `after.md` — diffs
|
||||
|
||||
## Held-Out Test Sets (Phase 2)
|
||||
|
||||
Located at `tests/<category>-tasks.json`. Each task has:
|
||||
- `prompt` — the recurring task
|
||||
- `reference` — exact-match gold answer
|
||||
- `rubric` — soft score rubric (0-1)
|
||||
- `domain` — research/devops/wiki/etc.
|
||||
|
||||
Currently building for 3 categories:
|
||||
- research-cron-output
|
||||
- devops-infrastructure-check
|
||||
- wiki-canonical-guide
|
||||
|
||||
## When NOT To Use
|
||||
|
||||
- For a one-off workflow (not a recurring pattern)
|
||||
- During a crisis/incident (humans must lead)
|
||||
- When session transcripts are < 24h old (not enough signal)
|
||||
- For skills < 300 tokens (over-optimization risk)
|
||||
Reference in New Issue
Block a user