Compare commits

...

83 Commits

Author SHA1 Message Date
Yif Yang 340a4c870f fix(minimax): honour OPTIMIZER_DEPLOYMENT for optimizer-role calls
Follow-up to #116. MiniMax exposed only TARGET_DEPLOYMENT, and
set_optimizer_deployment() never touched MiniMax, so a configured
optimizer_model was ignored: chat_optimizer and the optimizer message path
both sent TARGET_DEPLOYMENT. In mixed optimizer/target setups this could send
the wrong model name to the optimizer endpoint.

Adds OPTIMIZER_DEPLOYMENT and set_optimizer_deployment() to minimax_backend,
wired into the model facade. chat_optimizer and a new chat_optimizer_messages
use it, falling back to TARGET_DEPLOYMENT when unset. Also fixes dict[int]
return annotations to dict[str, int]. Adds routing tests.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-13 06:35:38 +00:00
Sparsh :) 49a5b617c0 fix(gate): resolve issue #100 (#102) 2026-07-13 01:27:13 +09:00
codeL1985 da99301210 fix(harvest): exclude sub-agent transcripts and plugin noise from task mining (#99)
Three filters to stop machine-generated prompts polluting the mined task
pool (they dominated ~80% of tasks on this machine):

- _is_meta_prompt: drop expanded slash-command bodies (<command-message>
  tags or '# /' headers) — plugin self-invocations are not user intents
- _AGENT_SESSION_MARKERS + _is_agent_session: skip sessions whose first
  prompt is another tool's agent brief (claude-mem observers, CLAUDE.md
  critic sub-agents, SkillOpt-Sleep's own command body)
- load walk: skip <session>/subagents/ dirs and agent-*.jsonl files —
  Agent-tool sidechain transcripts are Claude-authored, not user tasks

Verified: harvest went from 120 sessions / mostly-noise tasks to 55
sessions / 38 real user tasks.

Co-authored-by: codeL1985 <l@cypherlab.tech>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 01:27:05 +09:00
Jay C 19d98ea01c fix(model): route chat_optimizer through minimax_chat (not just chat_target) (#116)
Two missing pieces in the minimax_chat dispatch chain:

1. minimax_backend.py did not define chat_optimizer or chat_optimizer_messages;
   only chat_target / chat_target_messages existed. Any caller using
   optimizer_backend='minimax_chat' would hit AttributeError or fall through
   to _openai (Azure).

2. skillopt/model/__init__.py's chat_optimizer and chat_optimizer_messages
   dispatchers checked claude_chat and qwen_chat but not minimax_chat,
   so minimax_chat callers would silently fall through to _openai.chat_optimizer
   (the Azure path), which fails with 'Azure OpenAI endpoint is not configured'
   on any setup without AZURE_OPENAI_* env vars.

Adds chat_optimizer to minimax_backend.py (mirrors chat_target via
_chat_messages_impl) and minimax_chat branches to both
chat_optimizer / chat_optimizer_messages dispatchers.

Verified locally: 1-epoch training on a 4-item SearchQA-format dataset
went from '[skip] no usable patches — skill unchanged' (baseline fallback)
to a successful accept_new_best with success_patches=1 per step.

Co-authored-by: Mavis (MiniMax) <Mavis@MiniMax.local>
Co-authored-by: jc <jc@users.noreply.github.com>
2026-07-13 01:25:55 +09:00
Chirag Singhal 46e8e800bc fix(qwen): support reasoning-model params (max_completion_tokens, omit temperature) (#128)
The qwen_chat backend (the generic OpenAI-compatible client) hardcoded
max_tokens and always sent temperature, so reasoning models behind
OpenAI-compatible gateways (GPT-5.x, Claude Opus 4.8 via Azure/LiteLLM)
would 400.

- Add opt-in QWEN_CHAT_USE_MAX_COMPLETION_TOKENS (+ role variants) that
  swaps the payload key max_tokens -> max_completion_tokens.
- Treat an explicit empty / none / off temperature as "omit" instead of
  collapsing to the 0.7 default (via _resolve_temperature).
- Thread both through configure_qwen_chat / _update_config.
- Defaults unchanged; fully backward compatible. Adds 6 tests.

Fixes #127

Co-authored-by: Chirag Singhal <chirag127@users.noreply.github.com>
2026-07-13 01:24:40 +09:00
dimitarvdenev b309723baa feat(sleep): add handoff backend — session-executed model calls, no API subprocess (#125)
Adds --backend handoff: the engine runs all deterministic stages and
outsources attempt/judge/reflect to prompt/answer files an interactive
agent session fills between runs (exit 3 = pending batch, re-run to
resume). Deterministic replay + the prompt-hash answer cache make resume
stateless; sentinel detection aborts any call built from unanswered
output so placeholders never reach scores or staging. Session digests
and mined tasks are pinned per night (secret-redacted) so the sessions
answering prompts cannot shift the task set, and LLM mining is routed
through the same handoff files. Ships a /skillopt-sleep-handoff Claude
Code command that answers each prompt in a fresh-context subagent to
protect the held-out gate.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 01:23:53 +09:00
ClumsyLucid 7df49656d1 fix(skillopt-sleep): surface Claude CLI spawn failures instead of silent zero scores (#126)
* fix(skillopt-sleep): surface Claude CLI spawn failures instead of silent zero scores

In _call and attempt_with_tools, the bare 'except Exception: return ""'
swallows FileNotFoundError (e.g. bare 'claude' on Windows with npm .cmd
shim) and any other spawn failure, returning an empty string that the
trainer treats as a legitimate model response that scores 0.0 everywhere.

Now the exception is caught explicitly: last_call_error is set, a
warning is logged, and the empty-string return is preserved for
backward compatibility of the control flow. This mirrors the pattern
from #92 (codex backend) which fixed the same class of 'dead CLI
masquerades as nothing to learn' bug.

Issue: #121

* test(skillopt-sleep): add tests verifying Claude CLI spawn failures are surfaced

Add two tests to TestClaudeCliBackendBare:
- test_spawn_failure_sets_last_call_error: _call sets last_call_error
  and returns '' when subprocess.run raises FileNotFoundError.
- test_attempt_tools_spawn_failure_sets_last_call_error: same for
  attempt_with_tools.

These prove the fix from the parent commit (surface spawn failures
instead of silently scoring 0) and guard against regressions.
2026-07-13 01:23:10 +09:00
codeL1985 df94a91ddc fix: make ClaudeCliBackend work on Windows (.cmd shim resolution + argv length) (#98)
* fix(backend): make ClaudeCliBackend work on Windows (.cmd shim + argv limit)

Every claude call failed with WinError 2 and was swallowed by the bare
except -> return '', so real-backend cycles silently scored 0.000/0.000
and produced zero edits.

- resolve claude_path via shutil.which: the npm-installed claude is a
  .cmd shim that CreateProcess cannot resolve by bare name
- pass the prompt via stdin instead of argv: the .cmd shim routes
  through cmd.exe whose command line caps at ~8K chars, which
  reflect/judge prompts exceed

Verified: reflect() on a synthetic max_chars failure now returns a
concrete bounded edit via the real CLI (was [] before).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(backend): suppress console window flashes on Windows (CREATE_NO_WINDOW)

Every CLI subprocess (claude/codex/copilot) spawned from a console-less
parent allocates a visible console window on Windows. A cycle making
hundreds of calls strobes cmd windows and steals focus, making the
machine unusable while the engine runs. Pass CREATE_NO_WINDOW on all
CLI call sites; no-op on POSIX.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: codeL1985 <l@cypherlab.tech>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 01:22:14 +09:00
Tanmay Garg cd8034c2db test(sleep): assert scores and gate_action in verifier tests (closes #94) (#96)
Add explicit assertions for held-out scores and gate actions to the verifier discipline test suite to strengthen its guarantees.

- Assert the concrete held-out baseline and candidate scores in test_gate_rejects_reward_hacking_edit.
- Add test_gate_accepts_beneficial_edit using MockBeneficialBackend to provide a paired case where an edit genuinely improves the held-out slice, expecting accepted=True and gate_action='accept_new_best'.
2026-07-13 01:20:30 +09:00
黄云龙 352cbc3445 fix(config): read YAML config files as UTF-8 (#124)
_load_yaml opened config files with the platform default (locale)
encoding instead of UTF-8. The shipped configs (e.g.
configs/_base_/default.yaml) contain UTF-8 non-ASCII characters
(em-dash, arrows, box-drawing), so load_config() raises
UnicodeDecodeError on any non-UTF-8 locale (e.g. Windows cp936/cp1252),
breaking scripts/train.py and scripts/eval_only.py before startup.

YAML is UTF-8 by spec, and the rest of the codebase already opens text
files with encoding="utf-8". Pass encoding="utf-8" here for parity.
2026-07-13 01:19:59 +09:00
黄云龙 8687566792 test(gate): add unit tests for evaluation gate decision function (#122) 2026-07-13 01:19:56 +09:00
CharlesYang030 e4ea6a6771 chore(release): v0.2.0
Highlights since v0.1.0:
- feat: SkillOpt-Sleep engine — nightly offline self-evolution
  (harvest -> mine -> replay -> consolidate behind a validation gate),
  with multi-objective reward, experience replay + dream rollouts,
  slow-update long-term memory, and secret redaction in cycle diagnostics.
  Shipped as the `skillopt-sleep` CLI.
- feat: cross-tool backends & plugin shells — Claude, Codex (+Desktop
  harvest), Copilot, Devin, and OpenClaw.
- feat: SearchQA split materialization + rollout fail-fast.
- fix: Windows robustness for claude/codex backends, hardened JSON
  fallback, Qwen timeout/thinking gating, Codex failure surfacing.

Packaging:
- Bump pyproject / skillopt / skillopt_sleep to 0.2.0.
- Restore skillopt_webui to the packaged wheel.

See CHANGELOG.md for the full changelog and contributor acknowledgements.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 22:11:10 +08:00
Yif Yang 5487e2c426 fix(skillopt-sleep): redact secrets before persisting cycle diagnostics
PR #92 added a per-cycle diagnostics.json that surfaces backend stderr,
optimizer replies, and task responses so a 0.0 night is self-diagnosing.
Those free-text fields can carry credentials (e.g. a codex 401 stderr dump
containing an auth token), so persisting them verbatim was a new on-disk
leak surface.

- Add a shared redact_secrets() in staging.py and route diagnostics.json's
  call_error / reflect_raw_head / holdout_detail through it before writing.
- Redact the codex and Claude auth-error log lines too (a secondary sink
  when a file log handler is attached); last_call_error stays raw in memory
  so _AUTH_MARKERS matching is unaffected.
- Centralize _SECRET_PATTERNS in staging.py (harvest_codex now reuses them)
  and extend coverage to AWS / GitHub / Slack / Google / JWT token shapes.
- Tests: secret-shape coverage, private-key blocks, recursive/scalar
  passthrough, no over-redaction of plain prose, fail-fast auth-error log
  redaction, and an end-to-end check that diagnostics.json has no secret.

Observability-only; the gate and learning algorithm are unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-06-30 19:47:36 +00:00
Yifan Yang b9142bad24 fix(skillopt-sleep): surface codex auth/model/version failures instead of silently scoring 0 (#92)
Splits CodexCliBackend._call into _call_once + a retry wrapper so transient empties/timeouts are retried instead of silently scored 0, and fails fast on fatal auth/model/version errors (401, refresh_token_reused, token_expired, ChatGPT-account-unsupported, newer-Codex-required). On non-zero exit the CLI error text is surfaced via last_call_error instead of being returned as a model response. Adds per-cycle diagnostics.json (observability only; gate and learning algorithm unchanged) so a 0.0 night self-explains.
2026-07-01 03:20:08 +08:00
Yifan Yang 95a9e959fe test(sleep): add verifier-discipline stress test for the validation gate (#87)
Adds a reward-hacking stress test ensuring the consolidation gate rejects skill edits that game train/replay tasks while degrading held-out behavior. Also wires the minimax_chat backend into scripts/eval_only.py (coexisting with the qwen wiring from #85). Closes #67.
2026-07-01 02:40:24 +08:00
Tanmay9223 680dd28f5a fix(tests): move TestVerifierDiscipline above main block
(Addresses PR review feedback by ensuring python file-run execution discovers the test class)
2026-06-30 13:05:01 +05:30
Tanmay9223 fccc21f3f6 test(sleep): add verifier-discipline stress test (closes #67)
Add a regression test to ensure the validation gate correctly rejects
reward-hacking skill edits. It has been observed that optimizers
sometimes propose shortcuts that improve train/replay metrics but fail
to improve held-out behavior. This test codifies that the gate blocks
such artifacts.

Add TestVerifierDiscipline to the test_sleep_engine.py suite:
- Create MockRewardHackingBackend that simulates a reward-hacking rule
  which passes the train set but degrades the held-out tasks.
- Assert that the proposed edit is rejected by the gate.
2026-06-30 13:04:22 +05:30
Yifan Yang 6849e609a3 feat(eval): add missing minimax backend configuration
Add missing configuration setup in scripts/eval_only.py to properly
support the minimax_chat backend, which was entirely omitted.

Fix the following coverage gaps in eval_only.py:
- Add minimax CLI arguments
- Include the minimax config mappings in _MAP
- Update the backend parsing logic
- Call configure_minimax_chat
2026-06-30 13:04:22 +05:30
Daniel Martinez 9fa0716c72 fix(skillopt-sleep): also surface codex failures on the tool-call rollout path
Follow-up from a fresh-context review of the prior commit: CodexCliBackend.attempt_with_tools
(the rollout path for tool-requiring tasks) ran codex exec inline, swallowed all exceptions,
and never set last_call_error — so an auth/model/version failure on the tool path still
produced a silent empty->0 with no diagnostic signal, the exact failure class the prior commit
fixed for the _call path. Now it surfaces timeout/exception/non-zero-exit via last_call_error
(response stays empty; never leaks the CLI error text), so a failed tool rollout shows up in
diagnostics.json. Adds a regression test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-27 23:56:11 -05:00
Daniel Martinez 9fcf5868c3 fix(skillopt-sleep): surface codex auth/model/version failures instead of silently scoring 0
A nightly sleep cycle could run for weeks emitting held-out 0.0 -> 0.0 (gate reject, zero
edits), indistinguishable from "nothing to learn", when the real cause was the codex backend
returning an error (expired auth / model unsupported on the account / outdated CLI) that got
scored as a failed rollout.

backend (CodexCliBackend):
- split _call into _call_once + a retry wrapper: transient empties/timeouts are retried
  instead of silently returning "" (mirrors AzureOpenAIBackend's guard);
- on a non-zero exit, surface the reason via last_call_error and return "" rather than
  leaking the CLI error text as if it were a model response;
- fail fast (no retries) on fatal auth/model/version errors (401, refresh_token_reused,
  token_expired, "not supported when using Codex with a ChatGPT account",
  "requires a newer version of Codex").
backend (CliBackend.reflect): retain last_reflect_raw so a no-edits night is diagnosable.
consolidate: ConsolidationResult now carries per-task held-out detail (response, hard/soft,
  fail_reason) + reflect_raw + call_error.
cycle: write diagnostics.json per cycle so a 0.0 night self-explains instead of being a black box.
tests: 4 new (retry-not-silent-zero, auth-error-surfaced-not-scored, holdout-detail, reflect-raw).

Also gitignore the .skillopt-sleep/ runtime dir.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-27 22:26:20 -05:00
Yifan Yang 9969a8f393 Add Devin plugin (plugins/devin): MCP server + ATIF-v1.7 harvest (#88)
Wires skillopt_sleep into Devin via a stdlib-only MCP server and an ATIF-v1.7 harvester, following the plugins/copilot thin-shell pattern. Includes path-expansion fix, tests + ATIF fixture, schema/tool parity with copilot, and a harvest fix so single-turn sessions aren't dropped by the <3s replay filter.
2026-06-26 11:04:23 +08:00
Yif Yang 26e5338def Update citation from @misc to @article format
Co-Authored-By: Claude <noreply@anthropic.com>
2026-06-26 02:54:46 +00:00
khashayar 1a70e4c9cd devin harvest: space turns >=5s so single-turn sessions aren't dropped
A harvested single-turn Devin session spanned only 1s (reply written 1000ms
after the prompt), which the engine's harvest filter conservatively classifies
as a <3s headless replay (skillopt_sleep Issue #62) and skips — so a real
single-turn session mined 0 tasks. Widen the prompt->reply gap to 5s. With this,
an end-to-end dry-run mines the task: "night 1: 1 sessions -> 1 tasks".

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-25 22:03:15 +02:00
khashayar 9799c41461 devin plugin: full schema/tool parity with plugins/copilot
Mirror the copilot MCP server: same rich _TOOL_SCHEMA (source, model,
tasks_file, target_skill_path, max_sessions, max_tasks, lookback_hours,
auto_adopt, json, edit_budget, hour, minute) and generic flag forwarding, plus
sleep_schedule / sleep_unschedule. Devin specifics retained: the ATIF-v1.7
harvest step (run before data-reading actions, engine pointed at it via
--claude-home, default --source claude) and post-adopt sync into .devin/skills/.
Tests + README + rules snippet updated for the 7-tool interface.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-25 21:56:42 +02:00
khashayar e51eb7c4be devin plugin: expand ~ in CLAUDE_HOME from env + add tests & ATIF fixture
Review fixes:
- Path bug: SKILLOPT_DEVIN_CLAUDE_HOME (and SKILLOPT_SLEEP_REPO) read from the
  env are now wrapped in os.path.expanduser, so the documented "~/..." config
  no longer passes a literal ~ to --claude-home (which yielded zero mined
  sessions). expanduser on an absolute default is a no-op.
- tests/test_devin_plugin.py: tool-schema completeness, action→subcommand map,
  backend enum, the CLAUDE_HOME expansion regression, and an ATIF-v1.7 harvest
  shape test against a bundled fixture.
- plugins/devin/fixtures/devin_sample.json: sample ATIF-v1.7 transcript.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-25 21:49:21 +02:00
Yifan Yang 99ccb93945 fix(eval-only): configure qwen_chat/minimax backends so local LLM endpoints work (#85)
Replicates the trainer's backend setup in scripts/eval_only.py so eval-only no longer silently falls back to an unconfigured local endpoint. Closes #84.
2026-06-26 02:55:18 +08:00
Yifan Yang 9de9220214 docs(sleep): add cross-model scaling results (nano +11.9) and hyperparam ablation (#89)
Update RESULTS.md with:
- §2: GPT-5.4-nano target yields +11.9 pt (0.560→0.679) on SearchQA —
  2× the GPT-5.5 gain, demonstrating bigger benefit where headroom exists
- §4: Hyperparameter sweep confirms shipped defaults are optimal

Co-authored-by: Claude Opus 4 <noreply@anthropic.com>
2026-06-26 01:40:58 +08:00
khashayar bec23ed020 Add Devin plugin (plugins/devin): MCP server + ATIF-v1.7 harvest
Wires the skillopt_sleep engine into Devin (Cognition) via an MCP server,
following the same thin-shell pattern as plugins/copilot.

- mcp_server.py: stdlib-only stdio MCP server exposing the standard sleep_*
  tools (status, dry-run, run, adopt, harvest). REPO_ROOT defaults to ../.. so
  it finds skillopt_sleep automatically when run from plugins/devin/.
- harvest_devin.py: converts Devin ATIF-v1.7 transcripts, agentmemory, and
  .devin/skills/*/SKILL.md into the Claude Code-compatible JSONL the engine
  consumes; enriches with taskKey + outcome envelopes (hard test/build signal
  or judge rubric). Workspace auto-detection; cross-platform paths.
- judge.py, mcp-config.example.json, devin-rules.snippet.md, README.md.
- plugins/README.md: add Devin to the platform + install tables.

No changes to skillopt_sleep; shells out to `python -m skillopt_sleep` like the
other plugins. Pure stdlib; default backend mock (no API spend).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-25 10:42:52 +02:00
Gergely Imreh 8559308361 fix(eval-only): call configure_qwen_chat so itslocal LLM endpoints can be used
The eval-only tool skipped configuring some of the backend types, that
the training did configure. Because of this, the eval is silently
fell back to a local endpoint that wasn't actually configured, and
all evaluations runs failed.

Replicate the backend setup based on the trainer's code, and eval-only
can run with the qwen_chat backends.

Co-authored-by: Qwen-Coder <noreply@qwen.ai>
2026-06-24 15:31:19 +08:00
Yifan Yang 2d7e37a395 fix(json_utils): reject prose pseudo-JSON in single quotes/backticks (#82)
Follow-up to the string-aware brace scan: that change only skipped
double-quoted prose, so brace-shaped text in single quotes, backticks, or
bare prose (e.g. `{op: delete}`, '{x: 1}') still reached json_repair and was
fabricated into a bogus dict — strictly worse than None, since extract_json
feeds the optimizer's skill edits.

Add a _looks_json_like() guard before repair: a genuine JSON object's first
non-space char after `{` is `"` (a key) or `}` (empty). Prose pseudo-objects
start with a bare word and are rejected, while legitimate repair targets
(trailing commas, unescaped quotes inside string values) all begin with `"`
and pass — including objects whose string VALUES contain single quotes or
backticks, which must not be rejected.

Found by an independent GPT-5.5 re-review of the merged #79 code. Adds
regression tests for single-quoted / backticked / bare prose (-> None) and
for legitimate objects with quote/backtick string values (still repaired).
Tests: 30 pass (+3 skip) without json_repair, 33 pass with it, both clean
under -W error::RuntimeWarning.

Co-authored-by: Claude <noreply@anthropic.com>
2026-06-23 20:31:39 +08:00
Yifan Yang baad64a3b9 docs(readme): remove Acknowledgements section (#81)
The contributor is already credited via the Co-authored-by trailer carried
into main by #79; a dedicated README section is unnecessary.

Co-authored-by: Claude <noreply@anthropic.com>
2026-06-23 19:13:16 +08:00
Yifan Yang c2e47c50fb docs(readme): acknowledge community contributor @samuelgoofus-boop (#80)
Add an Acknowledgements section crediting @samuelgoofus-boop for the
Windows-robustness work on the Claude/Codex backends (originally #77,
merged via #79).

Co-authored-by: Claude <noreply@anthropic.com>
2026-06-23 19:03:30 +08:00
Yifan Yang 14c045f04f Windows robustness for claude/codex backends (+ hardened JSON fallback) (#79)
* Robustness for the claude/codex backends on Windows: argv overflow, subprocess encoding, tolerant JSON, test-eval dirs

Fixes surfaced running SkillOpt end-to-end on the bundled `claude` backend
(local Claude CLI) on Windows. None changes the OpenAI/GPT happy path.

1. skillopt/engine/trainer.py — the final test-eval directory
   (test_eval_final/) is written to before being created; add
   os.makedirs(..., exist_ok=True), matching the two sibling test-eval dirs.
   Without it, summary.json raises FileNotFoundError when a rollout yields
   zero predictions.

2. skillopt/model/claude_backend.py
   a. Pass the prompt via stdin (not argv): on Windows the whole command line
      is capped at ~32 KB and a large optimizer prompt (the success-analyst
      minibatch carrying several report trajectories) overflows it with
      [WinError 206], killing the run after retries.
   b. Pass the system prompt via --append-system-prompt-file (a temp file),
      not argv. The system prompt here is the skill being optimized, which
      SkillOpt grows over training; since the ~32 KB cap applies to the SUM of
      all argv, a grown skill would re-hit [WinError 206] even with the prompt
      on stdin.
   c. Pin the subprocess encoding to utf-8 (errors="replace"). With text=True
      and no encoding=, stdin is encoded with the system codepage; on a zh-CN
      box (cp936/GBK) a prompt containing an emoji or some Latin-1 characters
      raises UnicodeEncodeError before the CLI even starts, failing every retry.

3. skillopt/model/codex_backend.py — the same utf-8 encoding pin on its
   subprocess.run(input=...) call (identical unpinned-encoding pattern).

4. skillopt/utils/json_utils.py — extract_json() returned None for valid-
   looking JSON that strict json.loads rejects (unescaped ASCII quotes inside
   CJK string values, trailing commas), silently dropping the analyst's edits
   on non-schema backends (Claude/Qwen): reflect produces N edits, 0 applied.
   Add a json_repair fallback, but only on a single unambiguous object — a
   balanced-brace extractor plus a refuse-on-multiple-objects guard — so a
   chain-of-thought "scratch + final" response can't make repair silently
   return the wrong (discarded) object, which would be worse than None (None is
   detectable and retryable; a wrong-but-valid edit is applied blind). Declare
   json_repair in requirements.txt and the claude/qwen optional extras so the
   fallback is actually present (it otherwise no-ops, dropping edits silently).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit dca74a683e)

* fix(json_utils): harden tolerant JSON fallback from PR #77

Follow-up fixes on top of the cherry-picked Windows-robustness change:

1. Make _top_level_brace_objects() fully string-aware in its OUTER scan, not
   just inside an object. A '{' inside quoted prose (e.g. '"set it to {x}"')
   no longer starts a candidate object, so extract_json() returns None for
   prose pseudo-JSON instead of repairing it into a bogus dict — which would
   be strictly worse than dropping the edit, since extract_json feeds the
   optimizer's skill edits.

2. Pick the repair candidate BEFORE importing json_repair, so the missing-
   dependency RuntimeWarning only fires when there is genuinely a single
   malformed object that could have been repaired. Ordinary no-JSON / prose
   replies (the common case) now return None silently instead of warning on
   every call.

3. Resolve dependency-metadata inconsistency: json_repair is optional, so add
   it to the `all` extra (it was already in `claude`/`qwen`) and demote it
   from a hard requirement to an optional/commented entry in requirements.txt,
   matching the project's convention for backend-specific deps.

Adds regression tests for prose-with-braces (-> None), no-warning-on-plain-
text, single-object repair, and multi-object ambiguity. Existing 22 json
tests still pass with and without json_repair installed.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: samuelgoofus-boop <260247789+samuelgoofus-boop@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 19:00:23 +08:00
carpedkm 2841f82428 Fix ALFWorld gamefile paths relative to ALFWORLD_DATA 2026-06-23 10:32:38 +00:00
Yifan Yang 64c6dda105 Merge pull request #78 from Yif-Yang/main
docs(readme): add Trendshift daily/weekly badges
2026-06-23 16:52:42 +08:00
Yifan Yang c98eac18c7 docs(readme): add Trendshift daily/weekly badges (#1)
Add the microsoft/SkillOpt Trendshift badges (daily + weekly) side by
side in the README header.

Co-authored-by: Claude <noreply@anthropic.com>
2026-06-23 16:50:47 +08:00
Yifan Yang fc1f827f07 Merge pull request #74 from Yif-Yang/fix/python-path-and-lookback
fix: SKILLOPT_SLEEP_PYTHON override + lookback_hours first-run fallback
2026-06-20 22:26:43 +08:00
carpedkm 01b3e01804 fix: use None default for --lookback-hours to distinguish omitted vs 0
Codex round 3: argparse default=0 made every CLI invocation without
--lookback-hours clobber the config's 72h default. Now default=None;
only explicit --lookback-hours N (including 0) overrides config.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-20 14:23:17 +00:00
carpedkm 01075c90d3 fix: address codex round 2 — revert harvest break + allow lookback 0
- harvest.py: revert break to continue — mtime ordering can diverge
  from embedded ended_at timestamps (copy/touch), so we must check all
  files rather than early-exiting on the first old one
- cycle.py: use `is not None and > 0` so lookback_hours=0 means
  "scan full history" (opt-out of the cutoff)
- __main__.py: propagate --lookback-hours 0 to config as explicit 0

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-20 14:21:18 +00:00
carpedkm 6cc1cd2e95 fix: address codex review — use clock for cutoff + early-exit harvest
- cycle.py: use supplied `clock` parameter (not wall time) for the
  lookback cutoff, so deterministic tests/experiments get reproducible
  harvest windows
- harvest.py: break (not continue) when a file is older than since_iso,
  since files are sorted newest-first by mtime — avoids scanning the
  entire transcript directory for quiet projects with large histories

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-20 14:11:58 +00:00
carpedkm 889238b234 fix: add SKILLOPT_SLEEP_PYTHON override + lookback_hours first-run fallback
Two fixes from issue #57 feedback:

1. run-sleep.sh: support SKILLOPT_SLEEP_PYTHON env var to explicitly set
   the Python interpreter. Useful on macOS where system Python is 3.9 but
   a newer Python is available elsewhere (e.g. Codex Desktop's bundled
   Python 3.12). Applied to both the shared runner and the bundled
   Claude Code plugin copy.

2. cycle.py: on first run (no prior harvest recorded), apply the
   lookback_hours config (default 72h) as a time cutoff. Previously,
   first run scanned the entire transcript history, which could trigger
   massive LLM mining on users with months of session data.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-20 14:07:50 +00:00
Yifan Yang b5a1c2b317 Merge pull request #73 from Yif-Yang/fix/bare-subscription-auth
fix(sleep): make --bare conditional on ANTHROPIC_API_KEY (#68)
2026-06-20 21:46:09 +08:00
carpedkm 552ddefd74 fix: narrow CLI error markers to avoid false positives
Address codex review: "API key" was too generic — a model response
about configuring API keys would trigger a false auth warning. Now:
- Use specific phrases ("Invalid API key", "Unauthorized: invalid x-api-key")
- Only check short stdout (<300 chars) to skip real model responses
- Still check stderr unconditionally

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-20 13:32:43 +00:00
carpedkm bfa53bc46d fix(sleep): make --bare conditional on ANTHROPIC_API_KEY (#68)
ClaudeCliBackend._call() and attempt_with_tools() hardcoded --bare,
which skips Claude CLI's credential resolution. This broke subscription-
token auth: every model call silently returned "Not logged in" and
scored 0 — the user saw "baseline 0.0 → candidate 0.0, gate reject"
with no indication of an auth failure.

Fix: only pass --bare when ANTHROPIC_API_KEY is set. The remaining
isolation flags (--disable-slash-commands, --disallowedTools,
--exclude-dynamic-system-prompt-sections, clean temp cwd) already
provide the needed isolation without --bare.

Also adds _detect_cli_error() to log a warning when CLI output matches
known auth error patterns, so auth failures surface loudly instead of
deflating every score to 0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-20 13:28:34 +00:00
Yifan Yang 24b5a25ba8 Merge pull request #72 from Yif-Yang/feat/plugin-feature-sync
feat: sync all 4 runtime plugins with full engine surface + fix #52 #58 #62
2026-06-20 20:42:24 +08:00
carpedkm 0d648b2580 fix: address codex+gpt-5.5 review findings
- harvest: tighten sub-3s filter to also require prompt < 200 chars,
  avoiding false positives on fast real one-shot questions
- openclaw schedule_cmd: add docstring clarifying it schedules the
  shared engine, not the OpenClaw-native runner

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-20 12:40:34 +00:00
carpedkm 7d36b1d592 fix: address review findings in plugin sync PR
- OpenClaw schedule_cmd: pass project as required positional arg
- OpenClaw schedule_cmd/unschedule_cmd: unpack Tuple[bool, str] return
- OpenClaw schedule_cmd: propagate failure status (return 1 on not ok)
- OpenClaw unschedule_cmd: pass project to avoid silent no-op
- OpenClaw --minute default: 17 (consistent with engine and MCP)
- harvest.py: move datetime import to module level

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-20 12:04:07 +00:00
carpedkm 0be780052a feat: sync all 4 runtime plugins with full engine surface + fix #52 #58 #62
Bug fixes:
- #52: bundle run-sleep.sh in Claude Code plugin + 4-level fallback
- #58: add skillopt-sleep console script entry point in pyproject.toml
- #62: filter headless claude -p replay sessions from harvest

Plugin sync (Claude Code / Codex / Copilot / OpenClaw):
- Document all 22 CLI flags, 7 actions, 4 backends across all SKILL.md files
- Document config keys (preferences, gate_mode, dream_rollouts, etc.)
- Document memory consolidation (evolve_memory / evolve_skill)
- Add schedule/unschedule to all plugins
- Copilot MCP: expand schema from 3 → 16 params + schedule tools
- OpenClaw: add schedule/unschedule subcommands via shared scheduler

Tests:
- Cross-plugin parity test (prevents future feature drift)
- MCP schema completeness test

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-20 11:31:09 +00:00
carpedkm 0b5b9a4296 Merge pull request #60 from Kirchberg/codex/reviewed-task-files-cwd
Add reviewed task-file flow for Codex sleep runs
2026-06-20 08:59:02 +00:00
Kirill Kostarev 05cdc26beb Add reviewed task-file flow for Codex sleep runs 2026-06-20 08:58:48 +00:00
Yifan Yang 382811ddcc Merge pull request #50 from Dongbumlee/Dongbumlee/copilot-sleep-backend
Add Copilot as a SkillOpt-Sleep model backend (CopilotCliBackend) + research-engine MCP plugin
2026-06-20 16:57:53 +08:00
DB Lee d367ae1eea docs(plugins): list copilot in the cross-tool backend overview
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-17 17:38:10 -07:00
DB Lee 2c0980bda3 docs(copilot): correct backend hint in research MCP plugin (openai -> azure_openai)
The advertised backend choices in scripts/train.py use 'azure_openai',
not 'openai'; align the inputSchema description hint accordingly.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-17 17:25:50 -07:00
DB Lee 5799695951 feat(copilot): implement attempt_with_tools with cross-platform tool shims
Adds honest tool-call detection for CopilotCliBackend, mirroring the
Claude/Codex backends. Writes per-tool executable shims into the work dir
and detects real invocations from a calllog (not self-reported markers).
The Copilot backend is Windows-validated, so shims are cross-platform:
a .cmd batch shim on Windows and a chmod'd bash shim on POSIX, with an
OS-specific tool hint. Mirrors _call's flags/env (isolated COPILOT_HOME,
--allow-all-tools, MCP/instruction disabling) and the UTF-8 subprocess fix.

Adds test_attempt_with_tools_honest_detection: a CI-friendly, OS-aware
stub stands in for the CLI, runs the shim, and asserts both JSONL parsing
and log-based detection. Validated live on Windows (real Copilot call) and
on Linux/WSL (POSIX path).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-17 17:25:50 -07:00
DB Lee 013a7cd83a test: add unit tests for CopilotCliBackend (parsing + alias + isolated home)
Covers _parse_jsonl_response (multi-message concat, junk-line skipping,
empty/non-assistant events), get_backend alias resolution, and the
isolated-COPILOT_HOME / full-env opt-out behavior. Pure logic, no CLI required.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-17 17:25:50 -07:00
DB Lee 21f93c16c7 Add GitHub Copilot backend to SkillOpt-Sleep
Add CopilotCliBackend that drives the GitHub Copilot CLI in
non-interactive mode (copilot -p ... --output-format json) and parses the
JSONL event stream for assistant.message content. Registered as the
'copilot' backend (with aliases) and wired through the CLI, config,
experiment harness, and the Copilot MCP server's backend enum.

- Force UTF-8 decoding of CLI output (fixes cp1252 UnicodeDecodeError on
  Windows when responses contain non-cp1252 bytes).
- Minimise per-call startup: isolated COPILOT_HOME with built-in MCPs and
  custom instructions disabled, so user MCP servers are not spawned per
  call (~5x faster: 36s -> 7.4s). Override via SKILLOPT_SLEEP_COPILOT_HOME
  / SKILLOPT_SLEEP_COPILOT_MODEL / SKILLOPT_SLEEP_COPILOT_FULL_ENV.

Validated end-to-end on real held-out tasks (researcher persona:
0.42 -> 1.00 lift; gate correctly rejects non-improving edits).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-17 17:25:50 -07:00
DB Lee 5dc894715f Add SkillOpt research-engine MCP server plugin for Copilot
Exposes scripts/train.py and scripts/eval_only.py as Copilot MCP tools
(skillopt_list_configs, skillopt_train, skillopt_eval) via a stdlib-only
stdio server, mirroring the existing SkillOpt-Sleep plugin layout.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-17 17:24:00 -07:00
Yifan Yang 6940e46f4e Merge pull request #65 from summerview1997/codex/searchqa-materialize-splits
Add SearchQA split materialization helper
2026-06-17 23:50:38 +08:00
Yifan Yang 0e962219f5 Merge pull request #64 from summerview1997/codex/searchqa-rollout-failfast
Fail fast on systemic SearchQA rollout failures
2026-06-17 23:49:55 +08:00
Yifan Yang fc42e6bf72 Merge pull request #63 from summerview1997/codex/webui-env-backend-preflight
Add WebUI env loading and backend preflight
2026-06-17 23:49:50 +08:00
summerview1997 c755792049 Add SearchQA materialization tests 2026-06-16 09:27:09 +08:00
summerview1997 e591a28242 Add SearchQA split materialization helper 2026-06-16 09:26:56 +08:00
summerview1997 c04467a428 Add SearchQA materialization dependency extra 2026-06-16 09:26:46 +08:00
summerview1997 d5ae8c8e66 Document SearchQA split materialization 2026-06-16 09:26:35 +08:00
summerview1997 923becb00f Add SearchQA rollout fail-fast tests 2026-06-16 09:21:08 +08:00
summerview1997 da799620ba Fail fast on systemic SearchQA rollout failures 2026-06-16 09:20:57 +08:00
summerview1997 30cc8a3ed3 Add WebUI env preflight tests 2026-06-16 09:04:30 +08:00
summerview1997 d05851bd7f Add WebUI env loading and backend preflight 2026-06-16 09:04:19 +08:00
Yifan Yang 46b3207b96 docs(sleep): trim RESULTS to the headline results (remove the full grid)
Remove the per-cell full deployment grid section; keep the gate-safety stress
test, experience-replay scaling + night-by-night climb, the dream-diversity
ablation, the gbrain end-to-end result, and the scope/limitations. Renumber
sections; update the README pointer accordingly.
2026-06-15 17:08:51 +00:00
Yifan Yang d43e8dba1a docs(sleep): expand the grid into per-benchmark night-by-night tables
Replace the compact baseline->after grid with three grouped per-benchmark tables
(SearchQA / LiveMath / SpreadsheetBench), each showing all 3 targets x both modes
across every night (N0..N5) + Δ. Makes the trajectory visible — gains reach a
level and hold rather than being single lucky readings — and presents the full
18-cell evidence in a more solid, readable form. Footnotes LiveMath's 4-night run
(train split <50 tasks). Numbers unchanged; just richer presentation.
2026-06-15 16:54:01 +00:00
Yifan Yang d02098ffc4 docs(sleep): add full Results & Analysis (RESULTS.md); link from README
Adds docs/sleep/RESULTS.md — the complete deployment-scale study behind
SkillOpt-Sleep, presented rigorously (named benchmarks, test sizes, metrics,
baseline->after, single shared protocol):
  1. Gate-safety stress test: ungated nano SearchQA collapses 0.554->0.026
     (-52.8); the gated twin holds 0.570 — the core argument for the design.
  2. Full 18-cell deployment grid (3 benchmarks x 3 targets x gate/free),
     shipped config: mean +0.5, range [-2.4, +5.1], nothing hidden.
  3. Experience-replay scaling (recall_k 10->20->full: +3.1->+4.5->+5.6) and
     the night-by-night climb (0.798->...->0.858, gate accepts as late as N5).
  4. Dream-diversity fix as defense-in-depth: 3-config grid comparison
     (-2.66/-52.8 -> +0.24/-4.0 -> +0.53/-2.4); the -52.8 cell becomes +2.7
     from the dream fix alone.
  5. gbrain end-to-end 0.00->1.00 on real Claude + Codex.
  6. Honest scope: where it helps vs flat-in-noise, single-seed caveat with a
     seed-robustness spot check, keep-the-gate-on.
README Results section now links prominently to it. Docs only; numbers are
self-contained with reproduce commands (no raw run dumps committed).
2026-06-15 16:49:13 +00:00
Yifan Yang ea4ff459d7 docs(sleep): make the results section rigorous (named benchmarks, baseline→after)
Label each result with its benchmark, test size, metric, target model, and gate
mode; show absolute baseline→after (not just Δ); state the single shared protocol
once. SearchQA recall-scaling table (1400-item test, SQuAD-EM, GPT-5.5, gated) +
SpreadsheetBench confirmation (280-item, cell-value compare, nano, gate-free) +
the gbrain end-to-end line. Keeps the single-seed / flat-on-noisy caveats.
2026-06-15 16:42:43 +00:00
Yifan Yang de3be75bac docs(sleep): add a SkillOpt-Sleep module readme + News mention
Adds docs/sleep/README.md — a concise intro to the SkillOpt-Sleep plugin (what
it is, how to use it across the three agents, the opt-in experience-replay /
dream-rollout knobs, and headline results), linking to the full guide section.
Adds a News bullet pointing to it. No code changes.
2026-06-15 16:31:15 +00:00
Yifan Yang b701d9b6d9 docs: move SkillOpt-Sleep into the guide; clean docs/sleep; fix guide link
Per maintainer request:
- Remove the internal/scratch docs/sleep/ tree (reports, raw logs, blog run
  JSON, sweep.jsonl) — 23 files — and the root PUBLISHING.md. These were
  working notes, not reference docs.
- Take the dedicated SkillOpt-Sleep content out of the main README (News bullet
  + section) and host it in the rendered guide instead: new section 9 in
  docs/guideline.html (deployment companion, the three plugins, opt-in
  experience replay / dream rollouts) with a sidebar entry.
- Fix the README's opening reference so "Documentation & Reproduction Guide"
  links directly to the rendered GitHub Pages page, not the raw .html source.
- Repoint the now-removed docs/sleep links in the plugin READMEs to the guide
  section.

The plugin code (plugins/, skillopt_sleep/) is unchanged; only docs move.

Co-Authored-By: Claude Opus 4 <noreply@anthropic.com>
2026-06-15 16:20:50 +00:00
Yifan Yang 722ce646d4 feat(sleep): experience replay + dream rollouts in the cycle (opt-in)
Wires two consolidation mechanisms into the shipped nightly cycle, both default
OFF so existing behavior is unchanged:
  - dream_rollouts (>1): multi-rollout contrastive reflection per task
  - recall_k (>0): associative recall of the K most-similar past tasks (from a
    capped task_archive persisted in state.json) into tonight's dream
  - dream_factor (>0): synthetic task variants

New shared engine module skillopt_sleep/dream.py (recall_similar, dream_augment,
dream_consolidate) is called by both the plugin cycle and the experiment harness,
so reported numbers exercise the exact shipped code. Built on the existing
rollouts_k/sample_id support already in consolidate.py/rollout.py.

Validated (5 nights x 10 real tasks/night, full held-out test, GPT-5.5, gated):
the gain scales with recall depth on a clean signal —
SearchQA recall_k=10 +3.1, recall_k=20 +4.5, full-history reference +5.6;
SpreadsheetBench (nano, gate-free) +3.6. Flat within noise on saturated/noisy
cells. See docs/sleep/EXPERIENCE_REPLAY.md (+ raw runs under blog_runs/v2_port/).

Co-Authored-By: Claude Opus 4 <noreply@anthropic.com>
2026-06-15 15:58:27 +00:00
Yifan Yang 576f2f8bad Merge pull request #59 from Elzlxx/feat/openclaw-skillopt-sleep
feat(plugins): add OpenClaw shell for SkillOpt-Sleep
2026-06-15 18:26:12 +08:00
carpedkm 00d07bc59a Merge pull request #48 from Kirchberg/codex/codex-desktop-harvest
Add Codex Desktop transcript harvesting
2026-06-15 10:23:18 +00:00
Kirill Kostarev 31715a8b43 Add Codex Desktop transcript harvesting 2026-06-15 10:23:08 +00:00
carpedkm e8c3e10b30 Merge pull request #49 from Kirchberg/codex/codex-skill-first-upstream
Make Codex integration skill-first
2026-06-15 10:21:43 +00:00
Kirill Kostarev d31e9d9407 Back up legacy Codex prompt during install 2026-06-15 10:21:30 +00:00
Kirill Kostarev 1953484822 Make Codex integration skill-first 2026-06-15 10:21:30 +00:00
carpedkm 1b2652c6f8 Merge pull request #44 from imshunsuke/refactor/reflect-default-base
refactor: make EnvAdapter.reflect a shared default (fixes dropped reflect kwargs)
2026-06-15 09:06:38 +00:00
elzlxx 553446575a feat(plugins): add OpenClaw shell for SkillOpt-Sleep
Adds a thin OpenClaw shell wrapping the SkillOpt-Sleep engine. Enables
nightly validation-gated skill improvement cycles for OpenClaw agents.

Components:
- skillopt_sleep_openclaw.py: DeepSeek V4 Pro + Ollama nomic-embed-text
  backend, mirroring the Claude/Codex/Copilot backend pattern.
- run_sleep.py: CLI entry point supporting dry-run and pre-built task files.
- run_sleep_cron.sh: bash wrapper for nightly cron invocation.
- slash_sleep.py: /sleep command (status / run / adopt / reject / cost).
- config.json: engine config tuned for our stack.
- SKILL.md: OpenClaw skill manifest.
- tests/: 14 held-out tasks across 3 categories (research-cron, devops, wiki).

OpenClaw is the 4th ecosystem in which SkillOpt-Sleep can be deployed,
joining Claude Code, Codex, and Copilot. The shell follows the same
single-engine / thin-shell pattern as the existing three plugins.

End-to-end tested: pipeline runs against real OpenClaw session transcripts,
gate correctly rejects non-improvements, staging artifacts land in
~/.skillopt-sleep/staging/<night>/. Cost: ~$0.02/night on DeepSeek V4 Pro.
2026-06-14 23:27:54 +08:00
110 changed files with 8727 additions and 1608 deletions
+5
View File
@@ -22,6 +22,11 @@ data/*
outputs/
logs/
external/
# SkillOpt-Sleep runtime state (staging proposals, config, diagnostics, cron logs)
.skillopt-sleep/
# SkillOpt-Sleep handoff-backend round data (prompts/answers derived from transcripts)
.skillopt-sleep-handoff/
.skillopt-sleep-handoff.night*.done/
/BabyVision/
/MMRB/
+113
View File
@@ -0,0 +1,113 @@
# Changelog
All notable changes to SkillOpt are documented here. This project adheres to
[Semantic Versioning](https://semver.org/) and the format is based on
[Keep a Changelog](https://keepachangelog.com/).
## [Unreleased]
### Added
- **Handoff backend** (`--backend handoff`) for SkillOpt-Sleep — runs the
sleep cycle with no model subprocess or API key: the engine writes each
pending model call to `PROMPTS.md`/`pending.json` (exit code 3) and the
user's own agent session answers into `answers/<id>.md`; re-running the
same command resumes statelessly from the answers (typically 36 rounds
per night). Mined tasks are pinned per night so answering sessions cannot
shift the task set. Ships a `/skillopt-sleep-handoff` Claude Code command
that automates the loop with fresh-context subagents to protect the
held-out gate.
## [0.2.0] — 2026-07-02
The headline of this release is **SkillOpt-Sleep**: a nightly offline
self-evolution engine that harvests a coding agent's real session
transcripts, mines recurring tasks, replays them offline, and consolidates
short-term experience into long-term memory and skills — all behind the same
held-out validation gate that keeps SkillOpt training honest. It ships as a
decoupled top-level package (`skillopt_sleep/`, zero dependency on the
research code) and as the new `skillopt-sleep` CLI.
### Added
- **SkillOpt-Sleep engine** — nightly offline self-evolution cycle
(harvest → mine → replay → consolidate) behind a validation gate, exposed
as the `skillopt-sleep` console script and `python -m skillopt_sleep`.
- Multi-objective reward (accuracy / tokens / latency) with user preferences.
- Multi-rollout contrastive reflection under a token/time budget.
- Experience replay + controllable dream rollouts (opt-in).
- Slow-update long-term memory field (runs even with the gate off).
- 3-way train/val/test split with `gate_mode on|off`.
- Verifier-discipline validation gate, with a stress-test suite
(thanks @Tanmay9223, #87).
- **Cross-tool backends & plugin shells** for Claude Code, Codex, Copilot,
Devin, and OpenClaw:
- Codex Desktop transcript harvesting, skill-first Codex integration, and a
reviewed task-file flow (thanks @Kirchberg, #48, #49, #60).
- GitHub Copilot backend (`CopilotCliBackend`) + research-engine MCP plugin
(thanks @Dongbumlee, #50).
- Devin plugin: MCP server + ATIF-v1.7 harvest (thanks @xerxes-y, #88).
- OpenClaw shell for SkillOpt-Sleep (thanks @Elzlxx, #59).
- **SearchQA** split materialization helper and fail-fast on systemic rollout
failures, with a `searchqa` install extra (thanks @summerview1997,
#63, #64, #65).
- WebUI environment loading and backend preflight (thanks @summerview1997, #63).
### Changed
- Decoupled the Sleep engine into a standalone top-level `skillopt_sleep/`
package with zero dependency on the research code.
- Made `EnvAdapter.reflect` a shared default so reflect kwargs are no longer
dropped (thanks @imshunsuke, #44).
- English-only pass across the engine, plugins, and docs.
### Fixed
- Windows robustness for the Claude/Codex backends, plus a hardened JSON
fallback path (thanks @Yif-Yang, #79).
- Reject prose pseudo-JSON wrapped in single quotes/backticks (#82).
- Surface Codex auth/model/version failures instead of silently scoring 0
(thanks @dmmdea, #92).
- Redact secrets before persisting cycle diagnostics.
- Configure the `qwen_chat`/`minimax` backends so local LLM endpoints work
(thanks @imrehg, #85).
- Forward the Qwen target timeout and gate `enable_thinking` for vLLM targets
(thanks @mvanhorn, #40).
- Make `--bare` conditional on `ANTHROPIC_API_KEY` (#68), add a
`SKILLOPT_SLEEP_PYTHON` override with a lookback-hours first-run fallback
(#74), and fix ALFWorld gamefile paths relative to `ALFWORLD_DATA`.
### Packaging
- Bump `skillopt`, `skillopt.__version__`, and `skillopt_sleep.__version__`
to `0.2.0`.
- Restore `skillopt_webui` to the built wheel (it was dropped when the
`packages.find` include list was made explicit).
- Add the `searchqa` extra and include `json_repair` in the `claude`, `qwen`,
and `all` extras.
### Acknowledgements 🙏
v0.2.0 landed thanks to our community contributors — thank you!
- @Kirchberg — Codex Desktop harvesting, skill-first Codex integration,
reviewed task-file flow (#48, #49, #60)
- @Dongbumlee — GitHub Copilot backend + research-engine MCP plugin (#50)
- @summerview1997 — SearchQA materialization, rollout fail-fast, WebUI
preflight (#63, #64, #65)
- @xerxes-y — Devin plugin: MCP server + ATIF-v1.7 harvest (#88)
- @Elzlxx — OpenClaw shell for SkillOpt-Sleep (#59)
- @imshunsuke — shared `EnvAdapter.reflect` default + docs fixes (#43, #44)
- @mvanhorn — Qwen timeout forwarding + `enable_thinking` gating (#40)
- @dmmdea — surface Codex auth/model/version failures (#92)
- @Tanmay9223 — verifier-discipline stress test (#87)
- @imrehg`configure_qwen_chat` for local LLM endpoints (#85)
- @samuelgoofus-boop — community contributions
Special thanks to @Yif-Yang for driving the SkillOpt-Sleep engine.
**Full changelog:** https://github.com/microsoft/SkillOpt/compare/v0.1.0...v0.2.0
## [0.1.0] — 2026-06-02
Initial public release: the full training loop (rollout → reflect →
aggregate → select → update → evaluate), multi-backend support
(OpenAI / Azure / Claude / Qwen / MiniMax), six built-in benchmarks, and the
WebUI dashboard.
[0.2.0]: https://github.com/microsoft/SkillOpt/releases/tag/v0.2.0
[0.1.0]: https://github.com/microsoft/SkillOpt/releases/tag/v0.1.0
-81
View File
@@ -1,81 +0,0 @@
# Publishing SkillOpt-Sleep — how people install and use it
This is the open-source SkillOpt-Sleep tool: a nightly offline "sleep cycle" for
local coding agents, shipped as plugins for **Claude Code**, **Codex**, and
**Copilot**. One engine ([`skillopt_sleep/`](skillopt_sleep)), three thin shells
([`plugins/`](plugins)), decoupled from the research code.
## How end users install it
### Claude Code
The Claude Code plugin ships a marketplace manifest at
`plugins/claude-code/.claude-plugin/marketplace.json`.
```text
# inside Claude Code:
/plugin marketplace add microsoft/SkillOpt
/plugin install skillopt-sleep
/sleep status
```
(`/plugin marketplace add <owner>/<repo>` reads the marketplace manifest from the
repo; the entry points at `plugins/claude-code`.)
### Codex
```bash
git clone https://github.com/microsoft/SkillOpt.git
cd SkillOpt
bash plugins/codex/install.sh # installs /sleep prompt + skill
export SKILLOPT_SLEEP_REPO="$(pwd)" # so the runner is found anywhere
# then, in Codex: /sleep status
```
### Copilot
```bash
git clone https://github.com/microsoft/SkillOpt.git
# register the MCP server with your Copilot config (see plugins/copilot/README.md
# and plugins/copilot/mcp-config.example.json), pointing SKILLOPT_SLEEP_REPO at
# the clone. Then ask Copilot to "run the sleep cycle".
```
Requirements for all three: Python ≥ 3.10, and the corresponding agent CLI on
PATH. The default backend is `mock` (no API spend); `--backend claude|codex`
uses the user's own budget.
## Wider distribution (optional, maintainer steps)
1. **GitHub Release.** Tag the milestone so users can pin a version:
```bash
gh release create sleep-v0.1.0 --title "SkillOpt-Sleep v0.1.0" \
--notes "Nightly offline self-evolution plugins for Claude Code, Codex, Copilot."
```
2. **Official Claude Code plugin marketplace.** To appear in the public
directory, open a PR adding a `marketplace.json` entry to
[`anthropics/claude-code` / the official marketplace repo], pointing at
`microsoft/SkillOpt` subdir `plugins/claude-code`. Users could then
`/plugin install skillopt-sleep@<official-marketplace>`.
3. **PyPI (optional).** `skillopt_sleep` is a standalone package
(`pyproject.toml` lists it). A `pip install skillopt-sleep` distribution would
let users run `python -m skillopt_sleep ...` without cloning. Build with
`python -m build` and publish with `twine`.
4. **README News.** The main [`README.md`](README.md) already announces the
release and links to [`plugins/`](plugins) and
[`docs/sleep/FINAL_REPORT.md`](docs/sleep/FINAL_REPORT.md).
## Verifying a release works
```bash
# deterministic, no API key:
python -m skillopt_sleep.experiments.run_experiment --persona researcher --assert-improves
# the unit suite:
python -m unittest tests.test_sleep_engine
# the MCP server (Copilot):
printf '%s\n' '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' \
| SKILLOPT_SLEEP_REPO="$(pwd)" python3 plugins/copilot/mcp_server.py
```
+13 -58
View File
@@ -4,12 +4,18 @@
[![Project Page](https://img.shields.io/badge/Project%20Page-SkillOpt-8dbb3c)](https://microsoft.github.io/SkillOpt/) [![Paper](https://img.shields.io/badge/Paper-arXiv-b31b1b)](https://arxiv.org/abs/2605.23904) [![Project Video](https://img.shields.io/badge/Project%20Video-Watch%20Demo-ff0000)](https://youtu.be/JUBMDTCiM0M) [![PyPI](https://img.shields.io/badge/PyPI-skillopt-green.svg)](https://pypi.org/project/skillopt/) [![Python 3.10+](https://img.shields.io/badge/Python-3.10%2B-blue.svg)](https://www.python.org/) [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
> 📖 **For installation, data preparation, training/eval commands, the full configuration reference, and framework internals, see the [Documentation & Reproduction Guide](docs/guideline.html)** — view it [rendered online](https://htmlpreview.github.io/?https://github.com/microsoft/SkillOpt/blob/main/docs/guideline.html) or via [GitHub Pages](https://microsoft.github.io/SkillOpt/docs/guideline.html).
<p align="center">
<a href="https://trendshift.io/repositories/38498?utm_source=trendshift-badge&utm_medium=badge&utm_campaign=badge-trendshift-38498" target="_blank" rel="noopener noreferrer"><img src="https://trendshift.io/api/badge/trendshift/repositories/38498/daily?language=Python" alt="microsoft%2FSkillOpt | Trendshift" width="250" height="55"/></a>
<a href="https://trendshift.io/repositories/38498?utm_source=trendshift-badge&utm_medium=badge&utm_campaign=badge-trendshift-38498" target="_blank" rel="noopener noreferrer"><img src="https://trendshift.io/api/badge/trendshift/repositories/38498/weekly?language=Python" alt="microsoft%2FSkillOpt | Trendshift" width="250" height="55"/></a>
</p>
> 📖 **For installation, data preparation, training/eval commands, the full configuration reference, and framework internals, see the [Documentation & Reproduction Guide](https://microsoft.github.io/SkillOpt/docs/guideline.html)** (rendered on GitHub Pages).
---
## News 🔥🔥🔥
- **[2026-06-14]** 😴 **SkillOpt-Sleep (preview).** A nightly *sleep cycle* for local coding agents (Claude Code / Codex / Copilot): review past sessions offline, replay recurring tasks, and consolidate validated skills behind a held-out gate. This is an early **preview** — open-source and decoupled from the paper code — that we'll keep iterating on. See [`plugins/`](plugins/) and the [section below](#-skillopt-sleep--the-deployment-time-companion).
- **[2026-07-02]** 🚀 **SkillOpt [v0.2.0](https://github.com/microsoft/SkillOpt/releases/tag/v0.2.0) is out on [PyPI](https://pypi.org/project/skillopt/)!** Headline feature: **SkillOpt-Sleep**, a nightly offline self-evolution engine (harvest → mine → replay → consolidate, all behind a held-out validation gate) with multi-objective reward, experience replay + dream rollouts, and long-term memory — now shipped as the `skillopt-sleep` CLI. This release also adds cross-tool backends and plugin shells for **Claude, Codex, Copilot, Devin, and OpenClaw**, SearchQA split materialization, Windows robustness, and hardened JSON parsing. See the [release notes](https://github.com/microsoft/SkillOpt/releases/tag/v0.2.0) for the full changelog and contributor acknowledgements.
- **[2026-06-15]** 😴 **SkillOpt-Sleep (preview)** — a nightly offline self-evolution companion for local coding agents (Claude Code / Codex / Copilot): review past sessions, replay recurring tasks, and consolidate validated skills behind a held-out gate. See **[`docs/sleep/README.md`](docs/sleep/README.md)** for what it is, how to use it, and results.
- **[2026-06-03]** 🎉 **[gbrain](https://github.com/garrytan/gbrain), [gbrain-evals](https://github.com/garrytan/gbrain-evals/blob/main/docs/benchmarks/2026-06-03-skillopt.md), and [darwin-skill](https://github.com/alchaincyf/darwin-skill) have all integrated SkillOpt.**
- **[2026-06-02]** 🎉 **SkillOpt [v0.1.0](https://github.com/microsoft/SkillOpt/releases/tag/v0.1.0) is now available on [PyPI](https://pypi.org/project/skillopt/)!** Install with `pip install skillopt`. This initial release includes the full training loop (rollout → reflect → aggregate → select → update → evaluate), multi-backend support (OpenAI / Azure / Claude / Qwen / MiniMax), six built-in benchmarks, and WebUI dashboard.
@@ -53,54 +59,6 @@ https://github.com/user-attachments/assets/eb12d3bc-371c-467f-904d-91b61f339ed7
---
## 😴 SkillOpt-Sleep — the deployment-time companion
> **Preview.** SkillOpt-Sleep is an early preview that we are actively iterating
> on; interfaces and defaults may change. Feedback and issues are welcome.
SkillOpt (above) trains a skill offline on a benchmark. **SkillOpt-Sleep**
applies the same discipline to *your own daily usage*: it gives a local coding
agent a nightly **sleep cycle** that reviews your past sessions, replays your
recurring tasks on your own API budget, and consolidates what it learns into
**validated** long-term memory and skills — behind a held-out gate, staged for
your review. The agent gets better the more you use it, with no weight training.
It synthesizes **SkillOpt** (validation-gated bounded text edits), **Claude
Dreams** (offline consolidation; review-then-adopt), and the **agent sleep**
idea (short-term experience → long-term competence). One "night":
```
harvest session transcripts → mine recurring tasks → replay offline
→ consolidate (reflect → bounded edit → GATE on real held-out tasks)
→ stage proposal → (you) adopt
```
**Plugins for three agents** (one engine, three thin shells — see [`plugins/`](plugins/)):
| Platform | Folder | Install |
|---|---|---|
| **Claude Code** | [`plugins/claude-code`](plugins/claude-code) | `/plugin marketplace add ./plugins/claude-code``/skillopt-sleep` |
| **Codex** | [`plugins/codex`](plugins/codex) | `bash plugins/codex/install.sh``/skillopt-sleep` |
| **Copilot** | [`plugins/copilot`](plugins/copilot) | register `plugins/copilot/mcp_server.py` as an MCP server |
**Validated on real models.** On the public
[gbrain-evals](https://github.com/garrytan/gbrain-evals) `skillopt-v1` benchmark,
deficient skills go **0.00 → 1.00** on held-out sets with **both Claude and
Codex** (all 4 seeds, including a real tool-use loop), cross-model transfer is
positive, and the gate blocks regressions
([full results](docs/sleep/FINAL_REPORT.md)).
> **Open-source tool, decoupled from the research.** The engine lives in the
> top-level [`skillopt_sleep/`](skillopt_sleep) package with **zero dependency**
> on the paper's `skillopt/` experiment code (the validation gate is vendored).
> Controls — optional gate, multi-rollout contrastive reflection, token/time
> budget, multi-objective reward, user preferences, optimizer/target split — are
> documented in [`docs/sleep/CONTROLLABLE_DREAMING.md`](docs/sleep/CONTROLLABLE_DREAMING.md).
Deterministic proof (no API key): `python -m skillopt_sleep.experiments.run_experiment --persona researcher --assert-improves`.
---
## Extensibility & WebUI
### Adding a new backend
@@ -140,13 +98,10 @@ python -m skillopt_webui.app
## Citation
```bibtex
@misc{yang2026skilloptexecutivestrategyselfevolving,
title={SkillOpt: Executive Strategy for Self-Evolving Agent Skills},
author={Yifan Yang and Ziyang Gong and Weiquan Huang and Qihao Yang and Ziwei Zhou and Zisu Huang and Yan Li and Xuemei Gao and Qi Dai and Bei Liu and Kai Qiu and Yuqing Yang and Dongdong Chen and Xue Yang and Chong Luo},
year={2026},
eprint={2605.23904},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2605.23904}
@article{yang2026skillopt,
title={Skillopt: Executive strategy for self-evolving agent skills},
author={Yang, Yifan and Gong, Ziyang and Huang, Weiquan and Yang, Qihao and Zhou, Ziwei and Huang, Zisu and Li, Yan and Gao, Xuemei and Dai, Qi and Liu, Bei and others},
journal={arXiv preprint arXiv:2605.23904},
year={2026}
}
```
+14
View File
@@ -138,6 +138,20 @@ ALFWorld:
`searchqa_id_split/` is an ID-only manifest. Each released `id` exactly matches
the `key` field in `lucadiliello/searchqa`.
To materialize the runnable SearchQA split used by
`configs/searchqa/default.yaml`, install the optional dependency and run:
```bash
python -m pip install 'skillopt[searchqa]'
python scripts/materialize_searchqa.py
```
This writes full examples to:
```text
data/searchqa_split
```
Materialized examples must include the fields consumed by the SearchQA
environment, including:
+57
View File
@@ -288,6 +288,12 @@
<a href="#cli">CLI scripts</a>
<a href="#webui">WebUI</a>
</div>
<div class="group">
<div class="glabel"><span class="num">9</span> SkillOpt-Sleep</div>
<a href="#sleep">Deployment companion</a>
<a href="#sleep-plugins">Plugins (3 agents)</a>
<a href="#sleep-replay">Experience replay (opt-in)</a>
</div>
</nav>
<!-- ───────────── MAIN CONTENT ───────────── -->
@@ -917,6 +923,57 @@ PY</code></pre>
<tr><td><code>--share</code></td><td class="def">off</td><td>Create a public Gradio share link.</td></tr>
</tbody>
</table></div>
</section>
<section id="sleep">
<h2>9.1 SkillOpt-Sleep — the deployment-time companion (preview) <a class="anchor" href="#sleep">#</a></h2>
<p><strong>SkillOpt-Sleep</strong> applies SkillOpt's discipline to your own daily usage. It gives a
local coding agent a nightly <em>sleep cycle</em> that reviews your past sessions, replays your
recurring tasks on your own API budget, and consolidates what it learns into <strong>validated</strong>
long-term memory and skills — behind a held-out gate, staged for your review. The agent gets better
the more you use it, with no weight training and zero inference-time overhead. It is an early
<strong>preview</strong> we are actively iterating on; interfaces and defaults may change.</p>
<p>One "night":</p>
<pre><code>harvest Claude Code / Codex transcripts &rarr; mine recurring tasks &rarr; replay offline
&rarr; consolidate (reflect &rarr; bounded edit &rarr; GATE on real held-out tasks)
&rarr; stage proposal &rarr; (you) adopt</code></pre>
<p>The engine lives in the top-level <code>skillopt_sleep/</code> package with <strong>zero dependency</strong>
on the paper's <code>skillopt/</code> experiment code (the validation gate is vendored). Deterministic
proof, no API key required:
<code>python -m skillopt_sleep.experiments.run_experiment --persona researcher --assert-improves</code>.</p>
<h2 id="sleep-plugins">9.2 Plugins (three agents) <a class="anchor" href="#sleep-plugins">#</a></h2>
<p>One engine, thin per-agent shells (see <a href="https://github.com/microsoft/SkillOpt/tree/main/plugins"><code>plugins/</code></a>):</p>
<div class="table-wrap"><table>
<thead><tr><th>Platform</th><th>Folder</th><th>Install</th></tr></thead>
<tbody>
<tr><td>Claude Code</td><td><code>plugins/claude-code</code></td><td><code>/plugin marketplace add ./plugins/claude-code</code> &rarr; <code>/skillopt-sleep</code></td></tr>
<tr><td>Codex</td><td><code>plugins/codex</code></td><td><code>bash plugins/codex/install.sh</code> &rarr; <code>skillopt-sleep</code> skill</td></tr>
<tr><td>Copilot</td><td><code>plugins/copilot</code></td><td>register <code>plugins/copilot/mcp_server.py</code> as an MCP server</td></tr>
</tbody>
</table></div>
<p>Transcript source and replay backend are separate knobs: <code>--source claude</code> for Claude Code
transcripts, <code>--source codex</code> for Codex Desktop archived sessions under
<code>~/.codex/archived_sessions</code>, and <code>--backend codex</code> only when you want the
replay/optimizer to spend Codex budget.</p>
<h2 id="sleep-replay">9.3 Experience replay &amp; dream rollouts (opt-in) <a class="anchor" href="#sleep-replay">#</a></h2>
<p>Two consolidation mechanisms, both default <strong>off</strong> (so behavior is unchanged unless
enabled). They strengthen the nightly update when your tasks have a clean correctness signal; the
validation gate still governs what ships.</p>
<div class="table-wrap"><table>
<thead><tr><th>Config knob</th><th>Default</th><th>Effect</th></tr></thead>
<tbody>
<tr><td><code>dream_rollouts</code></td><td class="def">1</td><td>Run each task K times and learn from the good-vs-bad contrast (contrastive reflection).</td></tr>
<tr><td><code>recall_k</code></td><td class="def">0</td><td>Associative recall — pull the K most-similar past tasks (from a persisted archive) into tonight's dream.</td></tr>
<tr><td><code>dream_factor</code></td><td class="def">0</td><td>Add N lightweight synthetic variants of each task.</td></tr>
</tbody>
</table></div>
<p>On a clean-signal benchmark the gain scales with recall depth (deployment protocol: 5 nights &times;
10 new real tasks/night, full held-out test, GPT-5.5, gated): <code>recall_k=10</code> &rarr; +3.1 pts,
<code>recall_k=20</code> &rarr; +4.5 pts, full-history replay reference &rarr; +5.6 pts; a second benchmark
(SpreadsheetBench, GPT-5.4-nano, gate-free) gives +3.6 pts. On saturated or noisy tasks the effect is
flat within run-to-run noise (&plusmn;1&ndash;2 pts). Keep the gate on; it bounds the downside.</p>
<div class="footer-note">
SkillOpt — Executive Strategy for Self-Evolving Agent Skills ·
-117
View File
@@ -1,117 +0,0 @@
# SkillOpt-Sleep — controllable dreaming architecture
The sleep engine is no longer a single fixed pipeline. It is a controllable
offline "dream / imagination" loop the user steers. This documents the knobs
added in the four-stage refactor and how they map to the user's design.
## The mental model
> Sleep = an offline imagination rollout. Re-run the user's real
> tasks (and dream-augmented variants) many times, look at what went well vs
> badly, distil durable rules, and keep only what survives a real-task check —
> unless the user opts out of that check.
## 1. Data splits — train (dream) / val (real) / test (real)
The anti-overfitting foundation:
| Split | Source | Role |
|---|---|---|
| **train** | real tasks **+ dream-augmented** variants | drives reflection (the imagination pool — over-dreaming is fine) |
| **val** | **real only**, disjoint from test | gates updates (prevents overfitting) |
| **test** | **real only**, disjoint from val | the final held-out measure, kept close to real usage |
Hard guarantee (unit-tested): a task with `origin='dream'` **never** lands in
val or test. `assign_splits(val_fraction, test_fraction)` does the deterministic
3-way split; gbrain's own held-out maps to our `test`.
## 2. The validation gate is optional
`--gate on` (default): an edit is accepted only if it strictly improves the
**val** score — the SkillOpt discipline that blocks regressions and reward
hacking.
`--gate off`: greedy. Edits are kept without the hard val-improvement
requirement (the user decides they don't want hard filtering), but val/test
movement is still reported (`greedy_improved` / `greedy_regressed` /
`greedy_flat`) so nothing is hidden.
## 3. Slow-update — long-term memory, gate-independent
Even with the gate off, the engine runs a **slow-update** at the end of the
nights: it compares behaviour under the first-night vs final skill across the
val tasks and distils durable longitudinal guidance into a **protected field**
(`<!-- SLOW_UPDATE_START --> … <!-- SLOW_UPDATE_END -->`, the same markers as
the main SkillOpt repo). Step-level edits never touch this field. This is the
"short-term experience → long-term memory" consolidation; turning the gate off
does not cost you long-term memory.
## 4. Budget — the user picks the spend
`--budget-tokens N` / `--budget-minutes M`: the engine auto-plans depth
(`nights × rollouts_per_task`) to fit the budget (`plan_depth`). Stops cleanly
when exhausted and logs what it skipped — no silent truncation. The whole thing
is offline imagination on the user's own quota.
## 5. Multi-rollout contrastive reflection — the imagination core
`--rollouts-k K` (K>1): each train task is rolled out K times. The optimizer is
shown the **high-scoring vs low-scoring** attempts of the same task and asked
what the good ones did that the bad ones didn't, distilling a general rule. This
is a far stronger signal than a single failure, and it is exactly the user's
"run it many times, learn from the contrast" idea. Tasks with the highest score
*spread* (some passed, some failed) are the most informative and are prioritised.
## 6. Multi-objective reward — accuracy ↑, tokens ↓, latency ↓
Every rollout records its `tokens` and `latency_ms`.
`multi_objective_reward(w_acc, w_tokens, w_latency)` is a weighted reward so a
skill can be optimised to be **cheaper and faster**, not only more accurate
(cost terms normalised against a reference; default weights = accuracy-only, so
existing behaviour is unchanged). This turns "gets better the more you use it"
into "more accurate, cheaper, and faster the more you use it".
## 7. User preferences as a prior
`--preferences "<free text>"`: injected into the optimizer's reflect prompt as a
prior (set on the optimizer model for dual backends), so the user's stated
preferences steer what rules get written.
## How the knobs compose (one command)
```bash
python -m skillopt.sleep.experiments.run_gbrain \
--optimizer-backend claude --optimizer-model sonnet \ # strong optimizer
--target-backend claude --target-model haiku \ # cheap target (transfer)
--seeds thorough-analyst \
--gate on \ # or off for greedy
--rollouts-k 2 \ # contrastive imagination
--budget-tokens 60000 \ # auto-plan depth
--preferences "Prefer concise, British English." \ # prior
--nights 3
```
All of this is exercised by the deterministic test suite (29 tests) and
validated on real Claude + Codex (see `real_api_results.md` / `FINAL_REPORT.md`).
## Real cross-validation of the new features (Claude ⟷ Codex)
Three live runs exercised the new code paths on both runtimes (raw logs under
`docs/sleep/raw/crosscheck_*.txt`):
| # | Config | What it proves | Result |
|---|---|---|---|
| **A** | Claude Sonnet→Haiku, **gate=off**, **rollouts_k=2** | greedy mode + multi-rollout + 3-way split (val & test both reported) | brief-writer **test 0→1.00**, action `greedy_improved`, val=1.0 test=1.0 |
| **B** | **Codex**, gate=on, **rollouts_k=2** | new paths on the other runtime | brief-writer **test 0→1.00**, 2-night `accept_new_best`, val+test reported |
| **C** | Claude Sonnet→Haiku, thorough-analyst, 3 nights | **slow-update** long-term memory fires | test 0→0.33 (val gate holds nights 23) and the slow-update distilled a durable meta-rule |
The slow-update guidance C produced is the kind of cross-night lesson the field
is for — note it is general, not task-specific:
> *"On character-constrained tasks (≤1200 chars), plan structure before writing:
> allocate space per point explicitly and cut until the outline fits, then fill —
> never draft freely and trim after."*
Takeaways confirmed live: the **gate-off greedy path**, the **3-way val/test
split**, **multi-rollout** on both runtimes, and the **gate-independent
slow-update** all work with real models on both Claude and Codex.
-160
View File
@@ -1,160 +0,0 @@
# SkillOpt-Sleep — final validation report
> **What this is:** the consolidated, presented results for the SkillOpt-Sleep
> Claude Code plugin — a tool that lets a local agent improve itself overnight by
> reviewing past sessions, replaying tasks, and consolidating validated memory +
> skills behind a held-out gate. Every real-model result here was run on **both
> Claude and Codex**, including the honest failures and the bugs they exposed.
**Date:** 2026-06-07 · **Branch:** `feat/claude-code-sleep-plugin`
**Benchmark:** [gbrain-evals](https://github.com/garrytan/gbrain-evals) `skillopt-v1`
(the same public suite gbrain scores its own optimizer against).
**Protocol:** a deliberately deficient skill → 12 offline "nights" (replay →
reflect → bounded **gated** edit) → score the **held-out** task set (never
optimized against). Held-out scoring uses a local rule judge — the optimizer
never grades itself.
---
## 1. Headline — clean, all green (full gbrain parity)
**Strong optimizer (Claude Sonnet 4.6) → weak target (Claude Haiku 4.5)**, fully
isolated calls, 3 held-out tasks/seed. All **4** gbrain `skillopt-v1` seeds —
matching gbrain's own scorecard coverage:
| Optimizer → Target | Seed | Flaw | Held-out before → after | Nights |
|---|---|---|---|---|
| Sonnet → Haiku | brief-writer | missing structure | **0.00 → 1.00** | 1 |
| Sonnet → Haiku | advisor | no verdict | **0.00 → 1.00** | 1 |
| Sonnet → Haiku | thorough-analyst | no length discipline | **0.00 → 1.00** | 2 |
| Sonnet → Haiku | quick-answerer | never uses tools | **0.00 → 1.00** | 1 |
| Codex → Codex (gpt-5.5) | brief-writer | missing structure | **0.00 → 1.00** | 2 |
| Codex → Codex (gpt-5.5) | advisor | no verdict | **0.00 → 1.00** | 2 |
**4/4 Claude seeds reach a perfect held-out score** (gbrain's headline is the same
4/4 0→1.00), plus Codex on the text seeds. Every change is gated and staged.
The `quick-answerer` seed is judged by **real tool use** (`tool_called: search`):
the deficient skill says *"never look anything up — answer from memory"*; the
optimizer wrote an OVERRIDE rule, and the Haiku target **genuinely invoked a
`./search` shell tool** (detected from the tool's own log, not self-reported) →
held-out 1.00. The thorough-analyst run shows textbook **2-night convergence**
(0.33 → 1.00).
---
## 2. The finding that matters most: the optimizer model is decisive
This is the direct answer to "let me specify the optimizer and target separately,
and watch the skill." It matters a lot:
| Optimizer | Target | brief-writer | advisor | thorough-analyst |
|---|---|---|---|---|
| **Haiku** (weak) | Haiku | 1.00 *or* 0.00 (flaky) | 1.00 | 0.33 |
| **Sonnet** (strong) | Haiku | **1.00** | **1.00** | **1.00** |
A weak self-optimizing model (Haiku proposing its own edits) is **unreliable**
it intermittently emits non-JSON and wastes a night, so the same seed scores 1.00
on one run and 0.00 on another. A **strong optimizer** (Sonnet) reliably produces
clean, concrete edit rules and lifts every seed to 1.00. This is exactly the
SkillOpt design (strong optimizer, frozen target) and the reason the
optimizer/target split is a first-class feature here.
**Practical guidance baked into the plugin:** default to a strong optimizer; the
sweep's `direct` plan now uses Sonnet→Haiku.
---
## 3. Two real bugs we found by running against live models
Per gbrain's own lesson ("the bugs that matter only show up when the whole thing
actually runs"), the first live runs surfaced two real defects. Both are fixed.
1. **Ambient-context leak (Claude).** `claude -p` was injecting the user's
*global* skills + project `CLAUDE.md` into every optimizer/target call — one
reflect call literally returned a 21 KB list of the machine's installed skills
instead of JSON edits, so the night produced no edits and the gate rejected.
Some early Claude "successes" were partly leak-assisted. **Fix:** run isolated
`--bare --disable-slash-commands --disallowedTools '*'
--exclude-dynamic-system-prompt-sections`, clean temp cwd. (Codex was never
affected; the real `@openai/codex` binary runs in its own clean context.)
2. **Wasted nights on transient non-JSON.** A single malformed reply zeroed a
night. **Fix:** `reflect()` retries once with a firmer "JSON only" instruction.
We report these because a tool people build on has to be honest about where it was
weak and what changed.
---
## 4. Cross-model transfer (the price-difference value prop)
> *Optimize cheap overnight, deploy anywhere.* A skill is just text, so a good
> rewrite should help a model it was never optimized on.
Optimize on SOURCE, **freeze** the learned skill, evaluate held-out on TARGET with
no further optimization. All four pairs are positive — including **across
runtimes** (Codex ↔ Claude):
| Source (optimizer) | Target (deploy) | Seed | Target baseline → transferred | Gain |
|---|---|---|---|---|
| Claude Haiku (cheap) | Claude Sonnet (expensive) | brief-writer | 0.00 → **1.00** | +1.00 |
| Claude Sonnet | Claude Haiku | brief-writer | 0.00 → **1.00** | +1.00 |
| **Codex** | **Claude Haiku** | brief-writer | 0.00 → **1.00** | +1.00 |
| **Claude Haiku** | **Codex** | brief-writer | 0.00 → **1.00** | +1.00 |
**4/4 transfers positive.** A skill optimized on a cheap model deploys for free on
an expensive one, and skills move between Codex and Claude — the Sleep-setting
analogue of SkillOpt's cross-model and cross-harness transfer tables. This is the
quantified answer to "optimize cheap overnight, deploy anywhere."
Full machine-generated scorecard: [`benchmark_report.md`](benchmark_report.md)
(source data `sweep.jsonl`).
---
## 5. Reproduce everything
```bash
git clone https://github.com/garrytan/gbrain-evals /tmp/gbrain-evals
cd <repo>/SkillOpt-sleep
# the clean headline result (strong optimizer -> weak target)
python3.12 -m skillopt.sleep.experiments.run_gbrain \
--optimizer-backend claude --optimizer-model sonnet \
--target-backend claude --target-model haiku \
--seeds brief-writer,advisor,thorough-analyst \
--data-root /tmp/gbrain-evals/eval/data/skillopt-v1 --nights 2 --limit-replay 3 --limit-holdout 3
# Codex self-optimized
python3.12 -m skillopt.sleep.experiments.run_gbrain --backend codex --seeds brief-writer \
--data-root /tmp/gbrain-evals/eval/data/skillopt-v1 --nights 2 --limit-replay 3 --limit-holdout 3
# cross-model transfer
python3.12 -m skillopt.sleep.experiments.run_transfer \
--source-backend claude --source-model haiku --target-backend claude --target-model sonnet \
--seeds brief-writer
# the whole sweep + report
python3.12 -m skillopt.sleep.experiments.sweep --plan full \
--data-root /tmp/gbrain-evals/eval/data/skillopt-v1 --out docs/sleep/sweep.jsonl
python3.12 -m skillopt.sleep.experiments.report --in docs/sleep/sweep.jsonl --out docs/sleep/benchmark_report.md
# deterministic, no API (CI anchor)
python3.12 -m skillopt.sleep.experiments.run_experiment --persona researcher --assert-improves
```
Raw run logs are under `docs/sleep/raw/`.
---
## 6. Honest limitations
- **Latency:** each CLI call is ~1415 s startup-dominated, so runs are capped at
a few tasks/nights. Fine for nightly cron; we note it plainly.
- **Weak optimizers are flaky:** use a strong optimizer model (§2).
- **Tool-use seed covered honestly:** `quick-answerer` (`tool_called: search`)
runs a real tool loop — a callable `./search` shim, detected from its log.
Deeper multi-tool / multi-turn workflows are future work.
- **Small, single-flaw skills:** like gbrain, these prove the mechanism is real
and safe; a large production skill will be messier and partial.
-53
View File
@@ -1,53 +0,0 @@
TITLE:
Add SkillOpt-Sleep: nightly offline self-evolution plugins (Claude Code, Codex, Copilot)
BODY:
## Summary
Adds **SkillOpt-Sleep** — a nightly offline "sleep cycle" that gives a local
coding agent the deployment-time analogue of training: it reviews past sessions,
replays recurring tasks on the user's own API budget, and consolidates what it
learns into **validated** long-term memory and skills behind a held-out gate.
Synthesizes SkillOpt (validation-gated bounded text edits), Claude Dreams
(offline consolidation; review-then-adopt), and the agent-sleep idea
(short-term experience -> long-term competence).
Shipped as plugins for **three agents**, one engine + three thin shells:
- **Claude Code** — `.claude-plugin` + `/sleep` command + skill + hooks
- **Codex** — `~/.codex/prompts/sleep.md` + `~/.agents/skills` + `install.sh`
- **Copilot** — a stdlib-only MCP server exposing `sleep_*` tools
## Design notes
- **Open-source tool, decoupled from the research code.** The engine lives in the
new top-level `skillopt_sleep/` package with **zero dependency** on the paper's
`skillopt/` experiment package (the validation gate is vendored).
- Controllable: optional gate (`--gate on|off`), train(dream)/val(real)/test(real)
splits, slow-update long-term memory, token/time budget, multi-rollout
contrastive reflection, multi-objective reward (accuracy/tokens/latency), user
preferences, and separate optimizer/target models.
## Validation (real models)
On the public [gbrain-evals](https://github.com/garrytan/gbrain-evals)
`skillopt-v1` benchmark, deficient skills go **0.00 -> 1.00** on held-out sets
with **both Claude and Codex** (all 4 seeds, including a real tool-use loop);
cross-model transfer is positive; the gate blocks regressions. Independently
load-tested on a fresh non-benchmark persona ("SQL must always include LIMIT"):
held-out test **0.00 -> 1.00** on both backends. See `docs/sleep/FINAL_REPORT.md`
and `docs/sleep/plugin_load_test.md`.
## Tests
- 29 deterministic unit tests (`tests/test_sleep_engine.py`), no API key required.
- `python -m skillopt_sleep.experiments.run_experiment --persona researcher --assert-improves`
proves held-out lift and that the gate blocks a harmful edit.
## Test plan
- [ ] `python -m unittest tests.test_sleep_engine` (29 pass)
- [ ] `python -m skillopt_sleep.experiments.run_experiment --persona researcher --assert-improves`
- [ ] Claude Code: `/plugin marketplace add ./plugins/claude-code` -> `/sleep status`
- [ ] Codex: `bash plugins/codex/install.sh`
- [ ] Copilot: MCP server `tools/list` returns the `sleep_*` tools
+110
View File
@@ -0,0 +1,110 @@
# SkillOpt-Sleep 😴 — deployment-time companion (preview)
**SkillOpt-Sleep** applies SkillOpt's discipline to your *own daily usage*. It gives a
local coding agent a nightly **sleep cycle** that reviews your past sessions, replays
your recurring tasks on your own API budget, and consolidates what it learns into
**validated** long-term memory and skills — behind a held-out gate, staged for your
review. The agent gets better the more you use it, with **no weight training** and
**zero inference-time overhead**.
> **Preview.** This is an early preview we are actively iterating on; interfaces and
> defaults may change. The engine lives in the top-level [`skillopt_sleep/`](../../skillopt_sleep)
> package with **zero dependency** on the paper's `skillopt/` code (the validation gate
> is vendored).
## How it works
One "night":
```
harvest Claude Code / Codex transcripts → mine recurring tasks → replay offline
→ consolidate (reflect → bounded edit → GATE on real held-out tasks)
→ stage proposal → (you) adopt
```
It synthesizes **SkillOpt** (validation-gated bounded text edits), **Claude Dreams**
(offline consolidation; review-then-adopt), and the **agent-sleep** idea (short-term
experience → long-term competence).
## How to use it
### Quickest path: the `skillopt-sleep` CLI (pip)
```bash
pip install skillopt # installs the engine + the `skillopt-sleep` command
skillopt-sleep dry-run # harvest + mine + replay, report only (changes nothing)
skillopt-sleep run # a full nightly cycle; the proposal is staged for review
skillopt-sleep status # show state + the latest staged proposal
skillopt-sleep adopt # apply the latest staged proposal
skillopt-sleep schedule # install a nightly cron entry for this project
```
The per-agent plugin shells below (Claude Code / Codex / Copilot) still come from the
repo; the CLI above is the standalone, pip-only way to run a cycle.
One engine, thin per-agent shells (see [`plugins/`](../../plugins)):
| Platform | Folder | Install |
|---|---|---|
| **Claude Code** | [`plugins/claude-code`](../../plugins/claude-code) | `/plugin marketplace add ./plugins/claude-code``/skillopt-sleep` |
| **Codex** | [`plugins/codex`](../../plugins/codex) | `bash plugins/codex/install.sh``skillopt-sleep` skill |
| **Copilot** | [`plugins/copilot`](../../plugins/copilot) | register `plugins/copilot/mcp_server.py` as an MCP server |
Deterministic proof (no API key):
`python -m skillopt_sleep.experiments.run_experiment --persona researcher --assert-improves`.
### Opt-in: experience replay & dream rollouts
Two consolidation mechanisms, both default **off** (behavior is unchanged unless you
enable them). They strengthen the nightly update when your tasks have a clean
correctness signal; the validation gate still governs what ships.
| Config knob | Default | Effect |
|---|---|---|
| `dream_rollouts` | `1` | Run each task K times → learn from the good-vs-bad contrast (contrastive reflection). |
| `recall_k` | `0` | Associative recall — pull the K most-similar past tasks (from a persisted archive) into tonight's dream. |
| `dream_factor` | `0` | Add N lightweight synthetic variants of each task. |
## Results
> 📊 **More results & analysis — the gate-safety stress test, experience-replay
> scaling, and the dream-diversity ablation — are in
> [`docs/sleep/RESULTS.md`](RESULTS.md).** The highlights:
**Protocol (identical for every row below).** 5 nights × 10 new real "today" tasks
per night; the full held-out **test** split is scored before night 1 (baseline) and
after night 5 (after); optimizer = GPT-5.5; single seed (42); run through the exact
shipped engine (`skillopt_sleep.dream.dream_consolidate`). Numbers are absolute
held-out accuracy; **Δ** = `after baseline` in percentage points.
**(a) End-to-end on real agents — [gbrain-evals](https://github.com/garrytan/gbrain-evals) `skillopt-v1`.**
Deficient seed skills go **0.00 → 1.00** on the held-out set with **both Claude Code
and Codex** as the target agent (all 4 seeds, including a real tool-use loop).
**(b) Experience replay scales the gain — SearchQA** (1,400-item held-out test,
SQuAD exact-match; target = GPT-5.5; **validation-gated**):
| Replay config (`dream_rollouts=5`) | Baseline → After | Δ (pts) |
|---|---|---|
| `recall_k=10` | 0.802 → 0.834 | +3.1 |
| `recall_k=20` | 0.803 → 0.848 | **+4.5** |
| full-history replay *(reference, not a shipping default)* | 0.796 → 0.851 | +5.6 |
| `recall_k=10`, `dream_rollouts=8` *(more dreaming, same recall)* | 0.798 → 0.835 | +3.7 |
The gain rises monotonically with how much relevant past experience is recalled. The
same SearchQA cell **without** the gate (`recall_k=10`) is 0.808 → 0.839 (+3.1).
**(c) Second benchmark — SpreadsheetBench** (280-item held-out test; the agent's
generated openpyxl code is executed and compared cell-by-cell to a golden workbook;
target = GPT-5.4-nano; gate-free + the output-contract guardrail): 0.279 → 0.314 (**+3.6**).
**(d) Honest scope.** These gains hold where tasks recur and have a checkable
correctness signal. On saturated or noisy benchmarks (e.g. a strong model already
near ceiling) the effect is **flat within run-to-run noise** — single-seed baseline
variance here is ±12 pts, so treat sub-~1.5 pt differences as noise. The validation
gate keeps the worst case bounded; keep it **on** by default.
## Learn more
Full reference (pipeline, the three plugins, the experience-replay knobs) is in the
**[Documentation & Reproduction Guide](https://microsoft.github.io/SkillOpt/docs/guideline.html#sleep)**.
+185
View File
@@ -0,0 +1,185 @@
# SkillOpt-Sleep — results & analysis
This is the evidence behind SkillOpt-Sleep: does a nightly, offline sleep cycle
actually make a *deployed* agent better, and is it safe to run unattended? We
answer with a controlled deployment-scale study — the same protocol the plugin
runs in production, scored on full held-out test sets.
## Setup
**Protocol (identical for every cell unless stated).** 5 nights; each night adds
**10 new real "today" tasks**; the skill carries over and is refined night to
night. The full held-out **test** split is scored before night 1 (*baseline*) and
after night 5 (*after*); **Δ = after baseline** in percentage points. Optimizer
model = **GPT-5.5**; single seed (42); every number is produced by the exact
shipped engine `skillopt_sleep.dream.dream_consolidate` (the experiment harness and
the plugin cycle call the same function).
**Benchmarks** (real evaluators, not format heuristics):
| Benchmark | Held-out test | Scoring |
|---|---|---|
| SearchQA | 1,400 items | SQuAD exact-match vs gold |
| LiveMathematicianBench | 124 items | multiple-choice label (choices shuffled per item) |
| SpreadsheetBench | 280 items | the agent's generated openpyxl code is **executed**, output workbook compared cell-by-cell to a golden file |
**Targets:** GPT-5.5, GPT-5.4-mini, GPT-5.4-nano. **Modes:** validation-gated
(default) and gate-free.
---
## 1. The headline — the validation gate is what makes nightly self-evolution *safe*
Self-evolution is easy to build and easy to ruin: an optimizer that accepts its
own "lessons" unconditionally can adopt a plausible-but-wrong rule and an obedient
model will follow it off a cliff. We reproduced exactly that failure, then showed
the gate prevents it.
Stress case — **GPT-5.4-nano on SearchQA**, weak model on a single-sample (degraded)
reflection signal, same nights, same candidate edits, gate **off** vs **on**:
| | Night 0 → Night 5 | Δ |
|---|---|---|
| **no gate** | 0.554 → **0.026** | **52.8** |
| **with gate (default)** | 0.570 → 0.570 | 0.0 |
Ungated, the optimizer learned "answer with the document-title string, verbatim";
the model complied and accuracy collapsed night after night
(0.554 → 0.490 → 0.325 → 0.031 → 0.034 → 0.026). The gated twin **rejected every one
of those edits** and never lost a point. This single experiment is the core
argument for SkillOpt-Sleep's design, and why the gate ships **on by default**.
---
## 2. Cross-model scaling — bigger gains where there's headroom
The same protocol on a weaker target model (**GPT-5.4-nano**, optimizer = GPT-5.5)
produces substantially larger gains — because the weaker model has more room to
learn. This is the realistic "cheap deployed agent, strong overnight optimizer"
scenario:
| Config (SearchQA, nano, gated) | Baseline → After | Δ | Night-by-night |
|---|---|---|---|
| **cumulative replay, nights=5** | 0.560 → **0.679** | **+11.9** | 0.560 → 0.626 → 0.665 → 0.665 → 0.665 → 0.679 |
| recall_k=20, nights=5 | 0.566 → 0.681 | +11.5 | 0.566 → 0.659 → 0.685 → 0.685 → 0.681 → 0.681 |
| cumulative, nights=8 | 0.562 → 0.657 | +9.5 | saturates after night 5 |
Both replay strategies (cumulative and recall) agree within 0.4 pt — the gain is
robust across configurations.
**Compared to GPT-5.5 on the same benchmark (SearchQA, gated):**
| Target model | Best Δ | Baseline | Headroom |
|---|---|---|---|
| GPT-5.4-nano | **+11.9** | 0.560 | 44 pt |
| GPT-5.5 | +6.0 | 0.798 | 20 pt |
The story: **SkillOpt-Sleep helps most where there's the most to learn** — weaker
deployed models benefit ~2× as much from the same nightly optimization. This is
also the economical deployment pattern (cheap inference model + one strong
overnight optimizer call).
---
## 3. Experience replay turns a one-time bump into a climb
The plugin's two opt-in knobs (`recall_k`, `dream_rollouts`) are what produce the
gains. On **SearchQA, GPT-5.5, gated** — the gain rises monotonically with how
much relevant past experience is recalled:
| Replay (`dream_rollouts=5`) | Baseline → After | Δ |
|---|---|---|
| `recall_k=10` | 0.802 → 0.834 | +3.1 |
| `recall_k=20` | 0.803 → 0.848 | **+4.5** |
| full-history (reference, not a default) | 0.796 → 0.851 | +5.6 |
And the curve genuinely **climbs across nights** rather than jumping once and
plateauing — full-history replay, gated, night by night:
```
0.798 → 0.814 → 0.854 → 0.854 → 0.854 → 0.858
```
The gate accepts a new, better skill as late as **night 5** (0.854 → 0.858).
Replay-policy ablation (SearchQA, GPT-5.5):
| Replay policy | Gate-free Δ | Gated Δ |
|---|---|---|
| none (tonight's tasks only) | +3.9 | +2.0 |
| **recall k=10 (shipped default-able)** | +5.1 | +4.4 |
| cumulative (full history) | +4.8 | +6.0 |
Recall captures most of cumulative's benefit at a fraction of the per-night cost.
---
## 4. Default hyperparameters are the sweet spot
We swept `dream_factor`, `rollouts`, `per_night`, and `nights` on the nano cell
(SearchQA, gated) to verify the shipped defaults are well-tuned:
| Variant | Δ | vs default (+11.9) |
|---|---|---|
| dream_factor=4 (default 2) | +8.8 | 3.1 |
| rollouts=10 (default 5) | +9.5 | 2.4 |
| per_night=15 (default 10) | +2.7 | 9.2 |
| nights=8 (default 5) | +9.5 | 2.4 |
Every direction away from the default hurts. This means users get the best result
**out of the box** without tuning — the recipe is robust by design.
---
## 5. Why these gains exist — the dream-diversity fix (and a rigor note)
Reflection learns from the **contrast** between good and bad rollouts of the same
task, which requires the K dream rollouts to be *independent samples*. An early
version of the engine collapsed them to one cached sample, so contrastive
reflection never fired. Fixing that, then adding recall, is what produces the
gains in Sections 12. Measured across an 18-cell deployment sweep (3 benchmarks ×
3 targets × 2 modes), under three engine configurations:
| Engine configuration | mean Δ | worst-cell Δ | cells > +0.5 | cells < 0.5 |
|---|---|---|---|---|
| single-sample reflection (degraded) | 2.66 | **52.8** | 7 / 18 | 5 / 18 |
| diverse rollouts (K=5), no recall | +0.24 | 4.0 | 6 / 18 | 7 / 18 |
| **diverse rollouts + recall (shipped)** | **+0.53** | **2.4** | 7 / 18 | 7 / 18 |
The catastrophic 52.8 is removed **at its source** by diverse rollouts: the same
gate-free nano-SearchQA cell goes 0.554 → **0.586 (+2.7)** with no gate at all once
the dream is fixed. Recall then lifts the grid mean and tightens the worst case.
This is **defense in depth, each layer measured**: diverse rollouts propose better
edits, recall remembers relevant experience, and the gate catches whatever still
slips through.
---
## 6. End-to-end on real agents
On the public [gbrain-evals](https://github.com/garrytan/gbrain-evals) `skillopt-v1`
benchmark — designed for exactly this learnable-gap setting — deficient seed skills
go **0.00 → 1.00** on the held-out set with **both Claude Code and Codex** as the
target agent (all 4 seeds, including a real tool-use loop), and the two agents
cross-verify each other's consolidated skills.
---
## 7. Honest scope & limitations
- **Where it helps:** recurring tasks with a checkable correctness signal and real
headroom. That is the plugin's actual use case (your repeated daily tasks and
house rules the agent keeps missing).
- **Where it's flat:** saturated tasks on strong models, or noisy tasks with a weak
learning signal — within run-to-run noise.
- **Single seed.** Cells aggregate one seed per config; treat sub-~1.5 pt
differences as noise. Spot seed-robustness check on the one flagged cell
(nano SearchQA gated): seeds 42/43/44 give 1.9 / +3.6 / +4.7 (3-seed mean
**+2.1**), i.e. the tabled 1.9 is a pessimistic draw, not the typical outcome.
- **Keep the gate on.** It is the difference between bounded downside (2.4) and a
52.8 collapse. Gate-free mode is for users who cannot hold out a validation set
and is additionally protected by the output-contract guardrail.
---
Back to the module overview: [`docs/sleep/README.md`](README.md) ·
full reference: [Documentation & Reproduction Guide](https://microsoft.github.io/SkillOpt/docs/guideline.html#sleep).
-41
View File
@@ -1,41 +0,0 @@
# SkillOpt-Sleep — benchmark report
Auto-generated from `sweep.jsonl`. Benchmark: [gbrain-evals](https://github.com/garrytan/gbrain-evals) `skillopt-v1` (deficient skills, train/held-out split, local rule judge — no judge-API).
Held-out scores are computed by the harness, not the optimizer.
## Direct improvement (optimize, then deploy)
| Optimizer → Target | Seed | Held-out before | Held-out after | Nights | Tokens |
|---|---|---|---|---|---|
| claude:sonnet → claude:haiku | brief-writer | 0.00 | **1.00** | 2 | 6657 |
| claude:sonnet → claude:haiku | advisor | 0.00 | **1.00** | 2 | 7891 |
| claude:sonnet → claude:haiku | thorough-analyst | 0.00 | **1.00** | 2 | 17960 |
| codex:default → codex:default | brief-writer | 0.00 | **1.00** | 2 | 9969 |
| codex:default → codex:default | advisor | 0.00 | **1.00** | 2 | 6210 |
| claude:sonnet → claude:haiku | quick-answerer | 0.00 | **1.00** | 2 | 10988 |
| codex:default → codex:default | quick-answerer | 0.00 | **1.00** | 2 | 7347 |
**7/7 configurations improved on held-out.**
## Cross-model transfer (optimize on SOURCE, deploy frozen on TARGET)
The price-difference story: spend cheap tokens optimizing overnight, then deploy the frozen skill on any model with no further optimization.
| Source (optimizer) | Target (deploy) | Seed | Target baseline | Transferred | Gain |
|---|---|---|---|---|---|
| claude:haiku | claude:sonnet | brief-writer | 0.00 | **1.00** | +1.00 |
| claude:sonnet | claude:haiku | brief-writer | 0.00 | **1.00** | +1.00 |
| codex:default | claude:haiku | brief-writer | 0.00 | **1.00** | +1.00 |
| claude:haiku | codex:default | brief-writer | 0.00 | **1.00** | +1.00 |
**4/4 transfers were positive** (frozen skill helped a different model than it was optimized on).
## How to reproduce
```bash
git clone https://github.com/garrytan/gbrain-evals /tmp/gbrain-evals
python -m skillopt.sleep.experiments.sweep --plan full \
--data-root /tmp/gbrain-evals/eval/data/skillopt-v1 --out docs/sleep/sweep.jsonl
python -m skillopt.sleep.experiments.report \
--in docs/sleep/sweep.jsonl --out docs/sleep/benchmark_report.md
```
-73
View File
@@ -1,73 +0,0 @@
# SkillOpt-Sleep — validation experiment results
Generated: 2026-06-07 (autonomous offline session)
Backend: mock (deterministic, no API). Reproducible via the commands below.
```
$ python3.12 -m skillopt.sleep.experiments.run_experiment --persona researcher --nights 4 --json
{
"persona": "researcher",
"backend": "mock",
"nights_run": 1,
"baseline_holdout": 0.3333,
"after_holdout": 1.0,
"lift": 0.6667,
"improved": true,
"gate_blocks_harmful": true,
"final_skill_excerpt": "T -->\n## Learned preferences & procedures\n\n_This block is maintained by SkillOpt-Sleep. Edits here are proposed offline, validated against your past tasks, and adopted only after you approve them. Hand-edits outside this block are never touched._\n\n- Always wrap the final answer in <answer>...</answer> tags.\n- Report arXiv ids in the exact form arXiv:XXXX.XXXXX.\n<!-- SKILLOPT-SLEEP:LEARNED END -->\n",
"trace": [
{
"night": 0,
"holdout_score": 0.3333,
"action": "baseline",
"n_edits": 0
},
{
"night": 1,
"holdout_score": 1.0,
"action": "accept_new_best",
"accepted": true,
"n_edits": 2,
"edits": [
"Always wrap the final answer in <answer>...</answer> tags.",
"Report arXiv ids in the exact form arXiv:XXXX.XXXXX."
],
"n_rejected": 0
}
]
}
```
```
$ python3.12 -m skillopt.sleep.experiments.run_experiment --persona programmer --nights 4 --json
{
"persona": "programmer",
"backend": "mock",
"nights_run": 1,
"baseline_holdout": 0.3194,
"after_holdout": 1.0,
"lift": 0.6806,
"improved": true,
"gate_blocks_harmful": true,
"final_skill_excerpt": "laude Code sessions.\n\n<!-- SKILLOPT-SLEEP:LEARNED START -->\n## Learned preferences & procedures\n\n_This block is maintained by SkillOpt-Sleep. Edits here are proposed offline, validated against your past tasks, and adopted only after you approve them. Hand-edits outside this block are never touched._\n\n- Write git commit subjects in imperative mood, max 50 chars.\n<!-- SKILLOPT-SLEEP:LEARNED END -->\n",
"trace": [
{
"night": 0,
"holdout_score": 0.3194,
"action": "baseline",
"n_edits": 0
},
{
"night": 1,
"holdout_score": 1.0,
"action": "accept_new_best",
"accepted": true,
"n_edits": 1,
"edits": [
"Write git commit subjects in imperative mood, max 50 chars."
],
"n_rejected": 0
}
]
}
```
-76
View File
@@ -1,76 +0,0 @@
# SkillOpt-Sleep — plugin load-test (fresh examples)
This records an actual end-to-end load-test of all three plugin shells on a
**brand-new example** (not the gbrain benchmark seeds), run on 2026-06-08.
## The fresh persona
A data analyst whose SQL queries must always include a `LIMIT` clause — built
from scratch for this test. Two forms were used:
1. **Real transcripts** — crafted Claude Code session JSONL where the analyst
asks for SQL, the agent forgets `LIMIT`, and the user complains ("you forgot
a LIMIT again", "always cap results"). This exercises the real
harvest → mine pipeline.
2. **Checkable tasks** — the same intent with a rule judge
(`regex: (?i)LIMIT\s+100`), so the optimizer can be scored on whether future
SQL follows the house rule.
## Results
### Shell plumbing (all three drive the engine)
| Shell | What was run | Result |
|---|---|---|
| **Claude Code** (`scripts/sleep.sh`) | `harvest`, full `run`, `adopt` | harvest found 2 sessions → 2 tasks; `run` staged a proposal; `adopt` honored the safety contract (no live change when nothing was accepted) |
| **Codex** (`install.sh` + shared runner) | `install.sh` into a temp HOME | placed `~/.codex/prompts/sleep.md` and `~/.agents/skills/skillopt-sleep/SKILL.md` correctly |
| **Copilot** (`mcp_server.py`) | `initialize``tools/list``tools/call sleep_harvest` | 5 tools listed; `sleep_harvest` returned real engine output (2 sessions → 2 tasks) |
### Genuine improvement (real model, fresh persona)
Optimizer **Claude Sonnet 4.6** → target **Claude Haiku 4.5**, 3-way split
(5 train / 2 val / 5 test), scored on the held-out **test** queries; and the same
fresh persona self-optimized on **Codex**:
| Backend | Held-out **test** (fraction of SQL with `LIMIT 100`) before → after |
|---|---|
| Claude (Sonnet → Haiku) | **0.00 → 1.00** |
| Codex | **0.00 → 1.00** |
In one night each optimizer wrote, into the protected learned block, a rule like:
> *"OVERRIDE: Every SQL query you generate MUST include `LIMIT 100` …"* (Claude)
> *"Hard requirement: every SQL query response must include …"* (Codex)
and the target then applied it to the **unseen** test queries. This is the whole
claim on a task family the engine had never seen: it learned the user's house
rule from their failures and generalized it — confirmed on both backends.
## An honest finding from load-testing
The **first** attempt used `val_fraction=0.34, test_fraction=0.34`, which left
only **1 train task** for an 8-task set — too little signal — so reflect produced
nothing and the night was a no-op (val already 0.75). Re-balancing the split to a
real train pool (5 train) fixed it and produced the 0 → 1.00 result above. This
is exactly the kind of issue that only surfaces when you actually run the thing,
and it motivates a future guardrail: warn when the train pool is too small for
the chosen split fractions.
## Reproduce
The checkable persona run (real Claude):
```python
# see the snippet in docs/sleep/plugin_load_test.md history, or run:
python -m skillopt_sleep.experiments.run_experiment --persona programmer --assert-improves # deterministic
```
Shell checks:
```bash
# Copilot MCP server
printf '%s\n' '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' \
| SKILLOPT_SLEEP_REPO="$(pwd)" python3 plugins/copilot/mcp_server.py
# Codex installer (into a throwaway HOME)
HOME=$(mktemp -d) bash plugins/codex/install.sh
```
-45
View File
@@ -1,45 +0,0 @@
=== gbrain brief-writer CODEX, improved prompt, 2 nights, 3+3 tasks ===
{
"benchmark": "gbrain-evals/skillopt-v1",
"backend": "codex",
"model": "(default)",
"n_seeds": 1,
"n_improved": 1,
"tokens_used": 9990,
"results": [
{
"seed": "brief-writer",
"held_out_before": 0.0,
"held_out_after": 1.0,
"improved": true,
"nights": 2,
"trace": [
{
"night": 0,
"held_out_hard": 0.0,
"action": "baseline"
},
{
"night": 1,
"held_out_hard": 0.0,
"action": "accept_new_best",
"accepted": true,
"edits": [
"Every brief must include a clearly labeled section exactly titled `Key Risks`.",
"Every brief must include a line beginning `Confidence:` followed by a concise confidence level or rationale."
]
},
{
"night": 2,
"held_out_hard": 1.0,
"action": "accept_new_best",
"accepted": true,
"edits": [
"- Preserve required sections even when keeping the brief short; shorten the analysis before omitting `## Key Risks` or `Confidence:`."
]
}
],
"final_skill_tail": "tside this block are never touched._\n\n- Every brief must include a clearly labeled section exactly titled `Key Risks`.\n- Every brief must include a line beginning `Confidence:` followed by a concise confidence level or rationale.\n- Preserve required sections even when keeping the brief short; shorten the analysis before omitting `## Key Risks` or `Confidence:`.\n<!-- SKILLOPT-SLEEP:LEARNED END -->\n"
}
]
}
@@ -1,38 +0,0 @@
=== REAL cross-check A: Sonnet->Haiku, gate=OFF, rollouts_k=2, brief-writer (exercises new paths) ===
{
"benchmark": "gbrain-evals/skillopt-v1",
"backend": "target=claude/optimizer=claude",
"model": "(default)",
"n_seeds": 1,
"n_improved": 1,
"tokens_used": 11271,
"results": [
{
"seed": "brief-writer",
"held_out_before": 0.0,
"held_out_after": 1.0,
"improved": true,
"nights": 1,
"trace": [
{
"night": 0,
"test_hard": 0.0,
"action": "baseline"
},
{
"night": 1,
"val_hard": 1.0,
"test_hard": 1.0,
"action": "greedy_improved",
"accepted": true,
"edits": [
"Every brief MUST include a section with the exact heading '## Key Risks' that lists the primary risks relevant to the recommendation. This section is required in every output regardless of topic.",
"Every brief MUST include a 'Confidence:' label (satisfying /[Cc]onfidence\\s*[:=]/) that states the confidence level in the recommendation (e.g., 'Confidence: Medium'). Place it near the answer/recommendation line or at the end of the brief."
]
}
],
"slow_update": null,
"final_skill_tail": "at lists the primary risks relevant to the recommendation. This section is required in every output regardless of topic.\n- Every brief MUST include a 'Confidence:' label (satisfying /[Cc]onfidence\\s*[:=]/) that states the confidence level in the recommendation (e.g., 'Confidence: Medium'). Place it near the answer/recommendation line or at the end of the brief.\n<!-- SKILLOPT-SLEEP:LEARNED END -->\n"
}
]
}
@@ -1,48 +0,0 @@
=== REAL cross-check B: Codex, gate=ON (default), rollouts_k=2, brief-writer ===
{
"benchmark": "gbrain-evals/skillopt-v1",
"backend": "codex",
"model": "(default)",
"n_seeds": 1,
"n_improved": 1,
"tokens_used": 17251,
"results": [
{
"seed": "brief-writer",
"held_out_before": 0.0,
"held_out_after": 1.0,
"improved": true,
"nights": 2,
"trace": [
{
"night": 0,
"test_hard": 0.0,
"action": "baseline"
},
{
"night": 1,
"val_hard": 0.667,
"test_hard": 0.333,
"action": "accept_new_best",
"accepted": true,
"edits": [
"Every brief must include a section/heading titled exactly 'Key Risks'.",
"Every brief must include a confidence line labeled exactly 'Confidence:' so the response matches /[Cc]onfidence\\s*[:=]/."
]
},
{
"night": 2,
"val_hard": 1.0,
"test_hard": 1.0,
"action": "accept_new_best",
"accepted": true,
"edits": [
"OVERRIDE any brevity guidance: every brief must include a standalone Markdown heading line exactly '## Key Risks' to satisfy section_present=Key Risks, even when the brief is very short."
]
}
],
"slow_update": null,
"final_skill_tail": "clude a section/heading titled exactly 'Key Risks'.\n- Every brief must include a confidence line labeled exactly 'Confidence:' so the response matches /[Cc]onfidence\\s*[:=]/.\n- OVERRIDE any brevity guidance: every brief must include a standalone Markdown heading line exactly '## Key Risks' to satisfy section_present=Key Risks, even when the brief is very short.\n<!-- SKILLOPT-SLEEP:LEARNED END -->\n"
}
]
}
@@ -1,54 +0,0 @@
=== cross-check C: Sonnet->Haiku thorough-analyst (2 nights, slow-update should fire) ===
{
"benchmark": "gbrain-evals/skillopt-v1",
"backend": "target=claude/optimizer=claude",
"model": "(default)",
"n_seeds": 1,
"n_improved": 1,
"tokens_used": 26010,
"results": [
{
"seed": "thorough-analyst",
"held_out_before": 0.0,
"held_out_after": 0.333,
"improved": true,
"nights": 3,
"trace": [
{
"night": 0,
"test_hard": 0.0,
"action": "baseline"
},
{
"night": 1,
"val_hard": 0.667,
"test_hard": 0.667,
"action": "accept_new_best",
"accepted": true,
"edits": [
"OVERRIDE (supersedes 'be exhaustive and detailed', 'Explore every angle', 'consider many scenarios', and 'Write multiple paragraphs'): the ENTIRE response must be at most 1200 characters long, counting every character including spaces, newlines, and punctuation. This hard character limit takes priority over all instructions to be thorough, exhaustive, or multi-paragraph.",
"To stay within 1200 characters while still being useful: lead with the single most critical trade-off, then list 2-3 key considerations as tight bullet points. Omit headers, preamble, and restating the question."
]
},
{
"night": 2,
"val_hard": 0.667,
"test_hard": 0.667,
"action": "reject",
"accepted": false,
"edits": []
},
{
"night": 3,
"val_hard": 0.667,
"test_hard": 0.667,
"action": "reject",
"accepted": false,
"edits": []
}
],
"slow_update": "• On character-constrained tasks (≤1200 chars), plan structure before writing: allocate space per point explicitly and cut until the outline fits, then fill — never draft freely and trim after.\n• Multi-variable business/strategy analyses are high-risk for overrun; default to covering only the 23 most decisive factors rather than attempting exhaustive coverage.\n• Lead with the conclusion or recommendation first; eliminate all introductory restatement of the question, hedging preamble, and transitional filler under tight limits.\n• Persistent failures on the same task signal a structural habit, not a one-off error — treat repeated length violations as a signal to change the drafting approach entirely, not just edit more aggressively.",
"final_skill_tail": "ead with the conclusion or recommendation first; eliminate all introductory restatement of the question, hedging preamble, and transitional filler under tight limits.\n• Persistent failures on the same task signal a structural habit, not a one-off error — treat repeated length violations as a signal to change the drafting approach entirely, not just edit more aggressively.\n<!-- SLOW_UPDATE_END -->\n"
}
]
}
-101
View File
@@ -1,101 +0,0 @@
=== mock regression ===
Ran 19 tests in 0.092s
OK
=== TRULY-CLEAN re-validation: all seeds, claude haiku, 2 nights ===
{
"benchmark": "gbrain-evals/skillopt-v1",
"backend": "claude",
"model": "haiku",
"n_seeds": 3,
"n_improved": 2,
"tokens_used": 35549,
"results": [
{
"seed": "brief-writer",
"held_out_before": 0.0,
"held_out_after": 0.0,
"improved": false,
"nights": 2,
"trace": [
{
"night": 0,
"held_out_hard": 0.0,
"action": "baseline"
},
{
"night": 1,
"held_out_hard": 0.0,
"action": "reject",
"accepted": false,
"edits": []
},
{
"night": 2,
"held_out_hard": 0.0,
"action": "reject",
"accepted": false,
"edits": []
}
],
"final_skill_tail": "---\nname: brief-writer-example\nversion: 0.1.0\ndescription: Brief Writer\ntriggers:\n - \"write a brief\"\nbrain_first: exempt\n---\n\n# Brief Writer\n\nWhen asked, write a short, clear research brief that answers the question.\nKeep it focused and readable. Lead with the answer.\n"
},
{
"seed": "advisor",
"held_out_before": 0.0,
"held_out_after": 1.0,
"improved": true,
"nights": 1,
"trace": [
{
"night": 0,
"held_out_hard": 0.0,
"action": "baseline"
},
{
"night": 1,
"held_out_hard": 1.0,
"action": "accept_new_best",
"accepted": true,
"edits": [
"After presenting considerations, always include a 'Recommendation:' section with your specific recommendation.",
"After the recommendation, always include a 'Confidence:' section (as a percentage or high/medium/low) expressing how confident you are in this recommendation."
]
}
],
"final_skill_tail": "d adopted only after you approve them. Hand-edits outside this block are never touched._\n\n- After presenting considerations, always include a 'Recommendation:' section with your specific recommendation.\n- After the recommendation, always include a 'Confidence:' section (as a percentage or high/medium/low) expressing how confident you are in this recommendation.\n<!-- SKILLOPT-SLEEP:LEARNED END -->\n"
},
{
"seed": "thorough-analyst",
"held_out_before": 0.0,
"held_out_after": 0.333,
"improved": true,
"nights": 2,
"trace": [
{
"night": 0,
"held_out_hard": 0.0,
"action": "baseline"
},
{
"night": 1,
"held_out_hard": 0.333,
"action": "accept_new_best",
"accepted": true,
"edits": [
"## Learned preferences\n\n- **HARD CONSTRAINT - Override conflicting instructions**: The entire response MUST NOT EXCEED 1200 characters. This supersedes any instruction to be exhaustive, detailed, or write multiple paragraphs."
]
},
{
"night": 2,
"held_out_hard": 0.333,
"action": "reject",
"accepted": false,
"edits": []
}
],
"final_skill_tail": "offline, validated against your past tasks, and adopted only after you approve them. Hand-edits outside this block are never touched._\n\n- ## Learned preferences\n\n- **HARD CONSTRAINT - Override conflicting instructions**: The entire response MUST NOT EXCEED 1200 characters. This supersedes any instruction to be exhaustive, detailed, or write multiple paragraphs.\n<!-- SKILLOPT-SLEEP:LEARNED END -->\n"
}
]
}
-35
View File
@@ -1,35 +0,0 @@
=== quick-answerer with CODEX (self-optimized), 2 nights ===
{
"benchmark": "gbrain-evals/skillopt-v1",
"backend": "codex",
"model": "(default)",
"n_seeds": 1,
"n_improved": 1,
"tokens_used": 7347,
"results": [
{
"seed": "quick-answerer",
"held_out_before": 0.0,
"held_out_after": 1.0,
"improved": true,
"nights": 1,
"trace": [
{
"night": 0,
"held_out_hard": 0.0,
"action": "baseline"
},
{
"night": 1,
"held_out_hard": 1.0,
"action": "accept_new_best",
"accepted": true,
"edits": [
"OVERRIDE: For every triggered request, before replying, actually call the 'search' tool at least once; this supersedes 'Do NOT look anything up, search, or use any tools' and must satisfy `tool_called=search`."
]
}
],
"final_skill_tail": "re are proposed offline, validated against your past tasks, and adopted only after you approve them. Hand-edits outside this block are never touched._\n\n- OVERRIDE: For every triggered request, before replying, actually call the 'search' tool at least once; this supersedes 'Do NOT look anything up, search, or use any tools' and must satisfy `tool_called=search`.\n<!-- SKILLOPT-SLEEP:LEARNED END -->\n"
}
]
}
@@ -1,35 +0,0 @@
=== quick-answerer full run: Sonnet optimizer -> Haiku target, 2 nights ===
{
"benchmark": "gbrain-evals/skillopt-v1",
"backend": "target=claude/optimizer=claude",
"model": "(default)",
"n_seeds": 1,
"n_improved": 1,
"tokens_used": 10988,
"results": [
{
"seed": "quick-answerer",
"held_out_before": 0.0,
"held_out_after": 1.0,
"improved": true,
"nights": 1,
"trace": [
{
"night": 0,
"held_out_hard": 0.0,
"action": "baseline"
},
{
"night": 1,
"held_out_hard": 1.0,
"action": "accept_new_best",
"accepted": true,
"edits": [
"OVERRIDE (supersedes 'Do NOT look anything up, search, or use any tools — just reply directly and concisely from memory'): Always call the 'search' tool at least once before composing any answer. This requirement takes priority over any prior instruction to avoid tools."
]
}
],
"final_skill_tail": "nd adopted only after you approve them. Hand-edits outside this block are never touched._\n\n- OVERRIDE (supersedes 'Do NOT look anything up, search, or use any tools — just reply directly and concisely from memory'): Always call the 'search' tool at least once before composing any answer. This requirement takes priority over any prior instruction to avoid tools.\n<!-- SKILLOPT-SLEEP:LEARNED END -->\n"
}
]
}
@@ -1,98 +0,0 @@
=== KEY TEST: strong optimizer (sonnet) + weak target (haiku) — SkillOpt's actual design ===
(this is also your optimizer/target split in action)
{
"benchmark": "gbrain-evals/skillopt-v1",
"backend": "target=claude/optimizer=claude",
"model": "(default)",
"n_seeds": 3,
"n_improved": 3,
"tokens_used": 37791,
"results": [
{
"seed": "brief-writer",
"held_out_before": 0.0,
"held_out_after": 1.0,
"improved": true,
"nights": 1,
"trace": [
{
"night": 0,
"held_out_hard": 0.0,
"action": "baseline"
},
{
"night": 1,
"held_out_hard": 1.0,
"action": "accept_new_best",
"accepted": true,
"edits": [
"Every brief MUST include a section with the exact heading `## Key Risks` that lists the primary risks or uncertainties relevant to the recommendation. This section is required in every response, regardless of topic.",
"Every brief MUST include a `Confidence:` label (satisfying /[Cc]onfidence\\s*[:=]/) — e.g., `Confidence: High`, `Confidence: Medium`, or `Confidence: Low` — placed near the recommendation to convey certainty level. This label is required in every response."
]
}
],
"final_skill_tail": "tainties relevant to the recommendation. This section is required in every response, regardless of topic.\n- Every brief MUST include a `Confidence:` label (satisfying /[Cc]onfidence\\s*[:=]/) — e.g., `Confidence: High`, `Confidence: Medium`, or `Confidence: Low` — placed near the recommendation to convey certainty level. This label is required in every response.\n<!-- SKILLOPT-SLEEP:LEARNED END -->\n"
},
{
"seed": "advisor",
"held_out_before": 0.0,
"held_out_after": 1.0,
"improved": true,
"nights": 1,
"trace": [
{
"night": 0,
"held_out_hard": 0.0,
"action": "baseline"
},
{
"night": 1,
"held_out_hard": 1.0,
"action": "accept_new_best",
"accepted": true,
"edits": [
"OVERRIDE: The instruction 'so the reader can make up their own mind' must NOT suppress a conclusion. After presenting considerations, you MUST always end with an explicit label exactly matching 'Recommendation:' (capital R) followed by your concrete recommendation on the decision.",
"Always include a 'Confidence:' label (e.g., 'Confidence: High / Medium / Low') in every advisory response, placed immediately after or alongside the Recommendation line, expressing your confidence level in that recommendation."
]
}
],
"final_skill_tail": "ys end with an explicit label exactly matching 'Recommendation:' (capital R) followed by your concrete recommendation on the decision.\n- Always include a 'Confidence:' label (e.g., 'Confidence: High / Medium / Low') in every advisory response, placed immediately after or alongside the Recommendation line, expressing your confidence level in that recommendation.\n<!-- SKILLOPT-SLEEP:LEARNED END -->\n"
},
{
"seed": "thorough-analyst",
"held_out_before": 0.0,
"held_out_after": 1.0,
"improved": true,
"nights": 2,
"trace": [
{
"night": 0,
"held_out_hard": 0.0,
"action": "baseline"
},
{
"night": 1,
"held_out_hard": 0.333,
"action": "accept_new_best",
"accepted": true,
"edits": [
"OVERRIDE — supersedes all instructions to be 'exhaustive and detailed' or 'write multiple paragraphs': The ENTIRE response must be at most 1200 characters long (every character, including spaces, headers, and punctuation, counts toward this limit). If content would exceed 1200 characters, cut elaboration and stop at the most critical tradeoffs only.",
"For 'analyze the decision' responses, use plain concise prose rather than multi-level markdown headers and section dividers; structural markup consumes characters and makes it harder to stay within the 1200-character ceiling."
]
},
{
"night": 2,
"held_out_hard": 1.0,
"action": "accept_new_best",
"accepted": true,
"edits": [
"OVERRIDE — supersedes all instructions to be 'exhaustive and detailed' or 'write multiple paragraphs': The ENTIRE response must be at most 1200 characters long (every character counts). Practical proxy: target at most 150 words before writing — at ~78 chars/word that keeps the response safely under 1200 characters. Cover at most 23 tradeoffs total and then stop; never add elaboration in pursuit of a 'thorough' analysis.",
"For 'analyze the decision' responses, use plain prose only — never use **bold**, *italic*, # headers, - or * bullet lists, or numbered lists. Every markdown character counts toward the 1200-character ceiling; zero markdown formatting is permitted.",
"Limit every 'analyze the decision' response to at most 5 sentences total. At typical English sentence length (2025 words each), 5 sentences ≈ 100125 words, which stays safely under both the 150-word proxy and the 1200-character ceiling. Stop after the 5th sentence regardless of how much more could be said."
]
}
],
"final_skill_tail": "ter ceiling; zero markdown formatting is permitted.\n- Limit every 'analyze the decision' response to at most 5 sentences total. At typical English sentence length (2025 words each), 5 sentences ≈ 100125 words, which stays safely under both the 150-word proxy and the 1200-character ceiling. Stop after the 5th sentence regardless of how much more could be said.\n<!-- SKILLOPT-SLEEP:LEARNED END -->\n"
}
]
}
-114
View File
@@ -1,114 +0,0 @@
# SkillOpt-Sleep — REAL API results (Claude + Codex)
**Date:** 2026-06-07 (autonomous offline session)
**Benchmark:** [gbrain-evals](https://github.com/garrytan/gbrain-evals) `skillopt-v1`
the same public suite gbrain publishes its own SkillOpt scorecard against
([docs/benchmarks/2026-06-03-skillopt.md](https://github.com/garrytan/gbrain-evals/blob/main/docs/benchmarks/2026-06-03-skillopt.md)).
These are **real model runs**, not the deterministic mock. The agent's
`attempt` (and the optimizer's `reflect`) call live models via the `claude`
and `codex` CLIs. Held-out scoring is done **locally** by the rule judge
(`skillopt/sleep/judges.py`), so no judge-API spend and no way for the
optimizer to grade its own homework.
## Headline
| Backend | Seed | Held-out before | Held-out after | Nights | Tokens |
|---|---|---|---|---|---|
| **Claude (Haiku 4.5)** | brief-writer | **0.00** | **1.00** | 1 | ~6.7k |
| **Codex (default)** | brief-writer | **0.00** | **0.67** | 1 | ~5.1k |
| **Codex (directive prompt)** | brief-writer | **0.00** | **1.00** | 2 | ~10k |
Both backends took a **deliberately deficient** skill (a brief-writer with no
risks section and no confidence level) and, within 12 sleep nights, proposed
gated edits that lifted the held-out score to perfect. The edits went into the
protected `SKILLOPT-SLEEP:LEARNED` block; nothing else in the skill was touched.
This reproduces gbrain's published `0 → 1.00` headline with **our** engine and
shows it works across **two different agent runtimes** — the core of the
"Claude now, Codex next" plan.
### The multi-night convergence (Codex, why it matters)
The 2-night Codex run is the most informative trace in this whole exercise:
- **Night 1** — added two precise rules (a `Key Risks` section, a `Confidence:`
line). Held-out still **0.00**: the rules were right but the agent, told to
keep briefs short, was *dropping* them under length pressure.
- **Night 2** — the optimizer diagnosed its own residual failure and added a
meta-rule: *"Preserve required sections even when keeping the brief short;
shorten the analysis before omitting Key Risks or Confidence."* Held-out → **1.00**.
That second edit is not pattern-matching a checklist — it is reasoning about
*why the previous night underperformed*. This is exactly the iterative,
slow-update behavior SkillOpt's design predicts, and it is the strongest
argument for the sleep **loop** over a one-shot rewrite.
## What the optimizer actually wrote
**Claude** synthesized a full format template:
```
**Recommendation:** [Clear yes/no or specific answer]
**Rationale:** [2-3 bullet points supporting the answer]
**Key Risks:** [Downsides, edge cases, or assumptions that could invalidate this]
**Confidence:** [High/Medium/Low] — [Why]
```
**Codex** wrote a terser rule:
```
For every brief, include a `Key Risks` section and end with
`Confidence: Low|Medium|High`.
```
Both are correct, general, reusable rules (not task-specific answers). Claude's
fuller template made the agent satisfy the checks on **3/3** held-out items;
Codex's terser rule landed **2/3** — the missing item is a consistency miss the
agent would likely fix with one more night (see "Honest notes").
## How to reproduce
```bash
# clone the benchmark data
git clone https://github.com/garrytan/gbrain-evals /tmp/gbrain-evals
cd <repo>/SkillOpt-sleep # this worktree
# Claude backend
python3.12 -m skillopt.sleep.experiments.run_gbrain \
--backend claude --model haiku --seeds brief-writer \
--data-root /tmp/gbrain-evals/eval/data/skillopt-v1 \
--nights 1 --limit-replay 3 --limit-holdout 3 --json
# Codex backend (auto-detects the real @openai/codex binary, not the wrapper)
python3.12 -m skillopt.sleep.experiments.run_gbrain \
--backend codex --seeds brief-writer \
--data-root /tmp/gbrain-evals/eval/data/skillopt-v1 \
--nights 1 --limit-replay 3 --limit-holdout 3 --json
```
## Honest notes (in the spirit of gbrain's own scorecard)
- **Latency:** each CLI call is ~1415 s of startup-dominated wall time, so runs
were capped at 3 train + 3 held-out tasks and 1 night to keep them ~2.5 min.
The response cache makes re-scoring an unchanged (skill, memory) free.
- **Codex 0.67, not 1.00:** a single terse edit + single night under-shoots on
one held-out item. Two improvements (below) are expected to close it. We report
the 0.67, we don't dress it up.
- **3 of gbrain's 4 seeds are scored with zero API beyond `attempt`:**
`section_present`, `regex`, `max_chars` are pure-text checks. Only the
`quick-answerer` seed (`tool_called: search`) needs a real tool loop, which is
Phase-3 `fresh` replay.
- **The gate is real:** every accepted edit had to beat the held-out score; a
no-op night is rejected and the skill is left unchanged.
## Improvements this run motivated (applied + verified)
1. **A more directive `reflect` prompt** that aggregates the *exact* failing
judge criteria and tells the optimizer to satisfy every one (gbrain's lesson:
"the optimizer was never told what the scorer rewards"). Applied in
`skillopt/sleep/backend.py`. **Verified**: lifted Codex from 0.67 → 1.00.
2. **Multi-night convergence** — a terse first edit gets a sharper second pass;
the night-2 trace above shows the optimizer self-correcting. Recommend
`nights >= 2` for real backends.
-11
View File
@@ -1,11 +0,0 @@
{"baseline": 0.0, "after": 1.0, "improved": true, "tokens": 6657, "cfg": {"kind": "dual", "optimizer_backend": "claude", "optimizer_model": "sonnet", "target_backend": "claude", "target_model": "haiku", "seed": "brief-writer", "nights": 2}, "cfg_key": "{\"kind\": \"dual\", \"nights\": 2, \"optimizer_backend\": \"claude\", \"optimizer_model\": \"sonnet\", \"seed\": \"brief-writer\", \"target_backend\": \"claude\", \"target_model\": \"haiku\"}", "elapsed_s": 71.5}
{"baseline": 0.0, "after": 1.0, "improved": true, "tokens": 7891, "cfg": {"kind": "dual", "optimizer_backend": "claude", "optimizer_model": "sonnet", "target_backend": "claude", "target_model": "haiku", "seed": "advisor", "nights": 2}, "cfg_key": "{\"kind\": \"dual\", \"nights\": 2, \"optimizer_backend\": \"claude\", \"optimizer_model\": \"sonnet\", \"seed\": \"advisor\", \"target_backend\": \"claude\", \"target_model\": \"haiku\"}", "elapsed_s": 79.3}
{"baseline": 0.0, "after": 1.0, "improved": true, "tokens": 17960, "cfg": {"kind": "dual", "optimizer_backend": "claude", "optimizer_model": "sonnet", "target_backend": "claude", "target_model": "haiku", "seed": "thorough-analyst", "nights": 2}, "cfg_key": "{\"kind\": \"dual\", \"nights\": 2, \"optimizer_backend\": \"claude\", \"optimizer_model\": \"sonnet\", \"seed\": \"thorough-analyst\", \"target_backend\": \"claude\", \"target_model\": \"haiku\"}", "elapsed_s": 319.3}
{"baseline": 0.0, "after": 1.0, "improved": true, "tokens": 9969, "cfg": {"kind": "direct", "backend": "codex", "model": "", "seed": "brief-writer", "nights": 2}, "cfg_key": "{\"backend\": \"codex\", \"kind\": \"direct\", \"model\": \"\", \"nights\": 2, \"seed\": \"brief-writer\"}", "elapsed_s": 187.6}
{"baseline": 0.0, "after": 1.0, "improved": true, "tokens": 6210, "cfg": {"kind": "direct", "backend": "codex", "model": "", "seed": "advisor", "nights": 2}, "cfg_key": "{\"backend\": \"codex\", \"kind\": \"direct\", \"model\": \"\", \"nights\": 2, \"seed\": \"advisor\"}", "elapsed_s": 114.1}
{"baseline_target": 0.0, "transferred": 1.0, "transfer_gain": 1.0, "tokens": 13673, "cfg": {"kind": "transfer", "source_backend": "claude", "source_model": "haiku", "target_backend": "claude", "target_model": "sonnet", "seed": "brief-writer", "nights": 2}, "cfg_key": "{\"kind\": \"transfer\", \"nights\": 2, \"seed\": \"brief-writer\", \"source_backend\": \"claude\", \"source_model\": \"haiku\", \"target_backend\": \"claude\", \"target_model\": \"sonnet\"}", "elapsed_s": 180.3}
{"baseline_target": 0.0, "transferred": 1.0, "transfer_gain": 1.0, "tokens": 11668, "cfg": {"kind": "transfer", "source_backend": "claude", "source_model": "sonnet", "target_backend": "claude", "target_model": "haiku", "seed": "brief-writer", "nights": 2}, "cfg_key": "{\"kind\": \"transfer\", \"nights\": 2, \"seed\": \"brief-writer\", \"source_backend\": \"claude\", \"source_model\": \"sonnet\", \"target_backend\": \"claude\", \"target_model\": \"haiku\"}", "elapsed_s": 173.9}
{"baseline_target": 0.0, "transferred": 1.0, "transfer_gain": 1.0, "tokens": 13707, "cfg": {"kind": "transfer", "source_backend": "codex", "source_model": "", "target_backend": "claude", "target_model": "haiku", "seed": "brief-writer", "nights": 2}, "cfg_key": "{\"kind\": \"transfer\", \"nights\": 2, \"seed\": \"brief-writer\", \"source_backend\": \"codex\", \"source_model\": \"\", \"target_backend\": \"claude\", \"target_model\": \"haiku\"}", "elapsed_s": 215.7}
{"baseline_target": 0.0, "transferred": 1.0, "transfer_gain": 1.0, "tokens": 11284, "cfg": {"kind": "transfer", "source_backend": "claude", "source_model": "haiku", "target_backend": "codex", "target_model": "", "seed": "brief-writer", "nights": 2}, "cfg_key": "{\"kind\": \"transfer\", \"nights\": 2, \"seed\": \"brief-writer\", \"source_backend\": \"claude\", \"source_model\": \"haiku\", \"target_backend\": \"codex\", \"target_model\": \"\"}", "elapsed_s": 145.5}
{"baseline": 0.0, "after": 1.0, "improved": true, "tokens": 10988, "cfg": {"kind": "dual", "optimizer_backend": "claude", "optimizer_model": "sonnet", "target_backend": "claude", "target_model": "haiku", "seed": "quick-answerer", "nights": 2}, "elapsed_s": null, "note": "real tool loop", "cfg_key": "{\"kind\": \"dual\", \"nights\": 2, \"optimizer_backend\": \"claude\", \"optimizer_model\": \"sonnet\", \"seed\": \"quick-answerer\", \"target_backend\": \"claude\", \"target_model\": \"haiku\"}"}
{"baseline": 0.0, "after": 1.0, "improved": true, "tokens": 7347, "cfg": {"kind": "direct", "backend": "codex", "model": "", "seed": "quick-answerer", "nights": 2}, "elapsed_s": null, "note": "real tool loop", "cfg_key": "{\"backend\": \"codex\", \"kind\": \"direct\", \"model\": \"\", \"nights\": 2, \"seed\": \"quick-answerer\"}"}
+5 -8
View File
@@ -2416,14 +2416,11 @@
<div class="bibtex-box">
<button class="copy-btn" type="button" onclick="copyBibtex(this)">Copy</button>
<pre><code>@misc{yang2026skilloptexecutivestrategyselfevolving,
title={SkillOpt: Executive Strategy for Self-Evolving Agent Skills},
author={Yifan Yang and Ziyang Gong and Weiquan Huang and Qihao Yang and Ziwei Zhou and Zisu Huang and Yan Li and Xuemei Gao and Qi Dai and Bei Liu and Kai Qiu and Yuqing Yang and Dongdong Chen and Xue Yang and Chong Luo},
year={2026},
eprint={2605.23904},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2605.23904},
<pre><code>@article{yang2026skillopt,
title={Skillopt: Executive strategy for self-evolving agent skills},
author={Yang, Yifan and Gong, Ziyang and Huang, Weiquan and Yang, Qihao and Zhou, Ziwei and Huang, Zisu and Li, Yan and Gao, Xuemei and Dai, Qi and Liu, Bei and others},
journal={arXiv preprint arXiv:2605.23904},
year={2026}
}</code></pre>
</div>
</section>
+57 -8
View File
@@ -1,4 +1,4 @@
# SkillOpt-Sleep — plugins for Claude Code, Codex, and Copilot
# SkillOpt-Sleep — plugins for Claude Code, Codex, Copilot, and Devin
**Your coding agent forgets everything between sessions. SkillOpt-Sleep fixes
that.** While you sleep, it reviews what you did today, notices the rules you
@@ -8,7 +8,7 @@ only the rules that actually make it score better on *your own* past tasks. You
wake up to an agent that's better at *your* work, and you approve every change
before it sticks.
One engine, three thin shells. It synthesizes **SkillOpt** (validation-gated
One engine, four thin shells. It synthesizes **SkillOpt** (validation-gated
bounded text optimization — the research in this repo), **Claude Dreams**
(offline consolidation; input never mutated; review-then-adopt), and the **agent
sleep** idea (short-term experience → long-term competence).
@@ -20,6 +20,13 @@ sleep** idea (short-term experience → long-term competence).
---
| Platform | Folder | Mechanism | Status |
|---|---|---|---|
| **Claude Code** | [`claude-code/`](claude-code) | `.claude-plugin` + `/skillopt-sleep` + `/skillopt-sleep-handoff` commands + skill + hooks | full, installable |
| **Codex** | [`codex/`](codex) | user-level `skillopt-sleep` skill + shared runner | full |
| **Copilot** | [`copilot/`](copilot) | MCP server (`sleep_*` tools) + `copilot-instructions` | full (MCP) |
| **Devin** | [`devin/`](devin) | MCP server (`sleep_*` tools) + Devin ATIF-v1.7 harvest + `.devin/rules` | full (MCP) |
## Install (pick your agent)
| Platform | Install | Then |
@@ -27,11 +34,12 @@ sleep** idea (short-term experience → long-term competence).
| **Claude Code** | `/plugin marketplace add microsoft/SkillOpt``/plugin install skillopt-sleep` | `/skillopt-sleep status` |
| **Codex** | `git clone``bash plugins/codex/install.sh` | `/skillopt-sleep status` |
| **Copilot** | `git clone` → register `plugins/copilot/mcp_server.py` as an MCP server | ask "run the sleep cycle" |
| **Devin** | `git clone``devin mcp add skillopt-sleep -- python3 plugins/devin/mcp_server.py` | ask "run the sleep cycle" |
Requirements: Python ≥ 3.10 and the agent's CLI on PATH. All three call the same
[`run-sleep.sh`](run-sleep.sh) → `python -m skillopt_sleep`, so behaviour is
identical everywhere. Default backend is `mock` (no API spend); `--backend
claude|codex` uses your own budget.
claude|codex|copilot` uses your own budget.
---
@@ -141,6 +149,47 @@ The reward can weight not just correctness but **cost and speed**, so a skill ca
learn to be cheaper and faster, not only more accurate. *What it does for you:*
"answer directly instead of opening five files" becomes a learned habit.
### `--backend handoff` — session-executed calls (no API subprocess)
For subscription seats and environments where the engine shouldn't spawn
`claude -p` / API calls itself. The engine still runs every deterministic
stage (harvest → mine → replay scoring → gate → stage), but each model call
(attempt / judge / reflect) is written to a prompt file that **your own agent
session answers between engine runs**:
```bash
python -m skillopt_sleep run --backend handoff --project "$(pwd)"
# exit 3 => .skillopt-sleep-handoff/PROMPTS.md + pending.json were written
# answer each prompt (each in a FRESH context) into answers/<id>.md
# re-run the same command => it resumes from the answers and either
# finishes (exit 0) or stages the next prompt batch (exit 3)
```
A typical night converges in 36 rounds: baseline attempts → reflect →
candidate re-scoring per accepted edit. Resume is stateless — replay is
deterministic and answers are cached by prompt hash, so re-running skips
everything already answered. Mined tasks are pinned to
`.skillopt-sleep-handoff/tasks.json` on the first round, so the sessions that
answer the prompts can't shift the task set and invalidate earlier answers.
On a completed real run the handoff directory is archived to
`.skillopt-sleep-handoff.night<N>.done`.
On Claude Code, `/skillopt-sleep-handoff run` drives the whole loop for you,
answering each prompt in an isolated fresh-context subagent.
**Integrity rule:** answer every prompt in a fresh context (a subagent with no
conversation history). Answering from a session that has already seen the
mined tasks and their references contaminates the held-out gate and fakes the
improvement score.
*What it does for you:* the sleep cycle runs entirely on your interactive
session's subscription budget — no API key, no headless subprocess — while the
gate, splits, and staging discipline stay in the engine.
Limitations: `--rollouts-k > 1` gives no contrastive spread (identical prompt
→ identical answer file), and tool-loop tasks fall back to the single-shot
`TOOL_CALL:` marker convention.
### `schedule` / `unschedule` — set it and forget it
Built-in nightly scheduling (no manual cron):
@@ -168,7 +217,7 @@ schedule, if you trust it).
| Flag | Default | Meaning |
|---|---|---|
| `--backend mock\|claude\|codex` | `mock` | who runs/optimizes (mock = free) |
| `--backend mock\|claude\|codex\|copilot\|handoff` | `mock` | who runs/optimizes (mock = free; handoff = your own session answers) |
| `--preferences "..."` | | your house rules, as a prior |
| `--gate on\|off` | `on` | strict held-out gate vs. greedy |
| `--rollouts-k K` | `1` | multi-rollout contrastive reflection |
@@ -177,7 +226,7 @@ schedule, if you trust it).
| `--scope invoked\|all` | `invoked` | this project only, or all projects |
| `--auto-adopt` | off | apply without manual review (power users) |
Deep dive: [`../docs/sleep/CONTROLLABLE_DREAMING.md`](../docs/sleep/CONTROLLABLE_DREAMING.md).
Deep dive: [the SkillOpt-Sleep guide section](https://microsoft.github.io/SkillOpt/docs/guideline.html#sleep).
---
@@ -189,13 +238,13 @@ tasks the optimizer never trained on:
- **gbrain-evals `skillopt-v1`** (the public suite gbrain scores SkillOpt on):
deficient skills go **0.00 → 1.00** on all 4 seeds, including a real tool-use
loop; cross-model transfer is positive; the gate blocks regressions.
→ [`../docs/sleep/FINAL_REPORT.md`](../docs/sleep/FINAL_REPORT.md)
→ [the SkillOpt-Sleep guide section](https://microsoft.github.io/SkillOpt/docs/guideline.html#sleep)
- **Academic daily-cases** (math / spreadsheet / search-QA, the paper's 4:1:5
split with dream-augmented train): see
[`../docs/sleep/daily_cases_results.md`](../docs/sleep/daily_cases_results.md).
[the SkillOpt-Sleep guide section](https://microsoft.github.io/SkillOpt/docs/guideline.html#sleep).
- **Fresh load-test** (a "SQL must always include LIMIT" analyst, built from
scratch): held-out **0.00 → 1.00** on both backends.
→ [`../docs/sleep/plugin_load_test.md`](../docs/sleep/plugin_load_test.md)
→ [the SkillOpt-Sleep guide section](https://microsoft.github.io/SkillOpt/docs/guideline.html#sleep)
Try the deterministic proof yourself (no API key, no spend):
```bash
+25 -2
View File
@@ -60,6 +60,9 @@ they shell out to the CLIs you already have.
/skillopt-sleep run # full cycle: stages a reviewed proposal (still no live edits)
/skillopt-sleep status # see history + the latest staged proposal
/skillopt-sleep adopt # apply the staged proposal to CLAUDE.md / SKILL.md (with backup)
/skillopt-sleep-handoff run # same cycle, but THIS session answers the model calls
# (no claude -p subprocess, no API key — subscription-friendly)
```
Or call the engine directly (Python ≥ 3.10):
@@ -74,6 +77,26 @@ Default backend is **`mock`** — deterministic, no API spend — so you can try
plumbing for free. Switch to `--backend claude` or `--backend codex` for genuine
improvement on your own budget.
### Handoff mode (session answers the model calls)
`--backend handoff` runs the cycle without any model subprocess: the engine
executes the deterministic stages and writes every model call it needs to
`.skillopt-sleep-handoff/PROMPTS.md` + `pending.json` (exit code 3). You (or
the `/skillopt-sleep-handoff` command, which automates the loop with isolated
fresh-context subagents) write each raw answer to `answers/<id>.md` and re-run
the same command; it resumes from the answers and either finishes or stages
the next batch. Typically 36 rounds per night.
```bash
python -m skillopt_sleep run --backend handoff --project "$(pwd)"
# ... answer .skillopt-sleep-handoff/PROMPTS.md into answers/<id>.md ...
python -m skillopt_sleep run --backend handoff --project "$(pwd)" # resume
```
Answer every prompt in a **fresh context** — a session that has already seen
the mined tasks and their references would contaminate the held-out gate.
Details: [the plugins README](../README.md#--backend-handoff--session-executed-calls-no-api-subprocess).
## Does it actually improve? (real models, public benchmark)
SkillOpt-Sleep is validated against [gbrain-evals](https://github.com/garrytan/gbrain-evals)'
@@ -92,7 +115,7 @@ Both took a brief-writer with no risks section / no confidence level and, within
into the protected `LEARNED` block, nothing else touched. The Codex 2-night
trace even shows the optimizer **diagnosing its own residual failure** and
adding a meta-rule to fix it. Full writeup + reproduction:
[`docs/sleep/real_api_results.md`](../docs/sleep/real_api_results.md).
[the SkillOpt-Sleep guide section](https://microsoft.github.io/SkillOpt/docs/guideline.html#sleep).
Reproduce:
@@ -115,7 +138,7 @@ python -m skillopt_sleep.experiments.run_experiment --persona programmer --asse
Each prints the held-out score rising from baseline toward 1.0 as the gate
accepts the general rules your tasks need, and confirms the gate **rejects** an
injected harmful edit. Recorded output: [`docs/sleep/experiment_results.md`](../docs/sleep/experiment_results.md).
injected harmful edit. Recorded output: [the SkillOpt-Sleep guide section](https://microsoft.github.io/SkillOpt/docs/guideline.html#sleep).
## Schedule it nightly
@@ -0,0 +1,67 @@
---
description: Run the SkillOpt-Sleep cycle with the handoff backend — no API subprocess; this session answers the engine's model calls via prompt/answer files, in isolated fresh-context subagents
argument-hint: "[run | dry-run] [--preferences \"...\"] (default: run)"
allowed-tools: Bash, Read, Write, Task
---
# /skillopt-sleep-handoff — session-executed sleep cycle
You are driving **SkillOpt-Sleep in handoff mode**: the Python engine runs
every deterministic stage (harvest → mine → replay scoring → gate → stage)
and outsources each model call (attempt / judge / reflect) to YOU via
prompt files. No `claude -p` subprocess, no API key — the model work runs
on this session's budget, but each prompt MUST be answered in a fresh,
isolated context so the validation gate stays honest.
## Requested action: $ARGUMENTS
(If `$ARGUMENTS` is empty, treat it as `run`.)
## The loop
Repeat until the engine exits 0 (done) — at most 8 rounds:
1. **Run the engine** via the bundled runner:
```bash
"${CLAUDE_PLUGIN_ROOT}/scripts/sleep.sh" <action> --backend handoff --project "$(pwd)" --scope invoked
```
- exit 0 → the night is complete; go to "Finish" below.
- exit 3 → pending model calls; continue with step 2.
- anything else → stop and show the user the error output.
2. **Read the batch**: `Read` `.skillopt-sleep-handoff/pending.json` in the
project. Each entry has `id`, `prompt`, `max_tokens`, `answer_file`.
3. **Answer each prompt in ISOLATION** — this is the integrity rule:
- For each entry, launch a subagent (Task tool) whose ENTIRE input is
the `prompt` text verbatim. Add nothing: no summary of this session,
no mention of SkillOpt, no other prompts from the batch.
- Take the subagent's reply and `Write` the raw answer text (no
commentary, no code fences) to the entry's `answer_file`.
- NEVER answer from this session's own context — you have seen the
mined tasks and their references, so inline answers would contaminate
the held-out gate and fake the improvement score.
4. **Re-run the same engine command** — it resumes from the answers
directory and either finishes or stages the next batch.
## Finish
- `Read` the `report.md` in the staging dir the engine printed and show
the user: held-out baseline → candidate score, the gate decision, the
proposed edits, and where the proposal is staged.
- Tell the user nothing live changed; offer `/skillopt-sleep adopt`.
- The engine archives `.skillopt-sleep-handoff/` on a completed real run;
do not delete it yourself.
## Safety reminders
- **Never** edit `CLAUDE.md` or `SKILL.md` yourself — only `adopt` does
that, with a backup.
- Mined tasks are pinned to `.skillopt-sleep-handoff/tasks.json` on round
one, so sessions created while answering prompts cannot shift the task
set. Do not edit that file.
- If a batch looks like it contains secrets or content the user would not
want re-processed, stop and ask before answering.
+51
View File
@@ -0,0 +1,51 @@
#!/usr/bin/env bash
# SkillOpt-Sleep shared runner — used by all platform plugins (Claude Code,
# Codex, Copilot). Resolves the repo root (which contains the skillopt_sleep
# package), picks a Python >= 3.10, and execs the engine CLI.
#
# Usage: run-sleep.sh <run|dry-run|status|adopt|harvest|...> [args...]
set -euo pipefail
# This script lives at <repo>/plugins/run-sleep.sh, so the repo root (which
# holds skillopt_sleep/) is one level up. CLAUDE_PLUGIN_ROOT (if set by Claude
# Code) points at the plugin dir; the engine is then two levels above it.
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
if [ -d "$SCRIPT_DIR/../skillopt_sleep" ]; then
REPO_ROOT="$(cd "$SCRIPT_DIR/.." && pwd)"
elif [ -n "${CLAUDE_PLUGIN_ROOT:-}" ] && [ -d "$CLAUDE_PLUGIN_ROOT/../../skillopt_sleep" ]; then
REPO_ROOT="$(cd "$CLAUDE_PLUGIN_ROOT/../.." && pwd)"
elif [ -n "${SKILLOPT_SLEEP_REPO:-}" ] && [ -d "$SKILLOPT_SLEEP_REPO/skillopt_sleep" ]; then
REPO_ROOT="$SKILLOPT_SLEEP_REPO"
else
# last resort: search upward from CWD
d="$PWD"
while [ "$d" != "/" ]; do
[ -d "$d/skillopt_sleep" ] && { REPO_ROOT="$d"; break; }
d="$(dirname "$d")"
done
fi
if [ -z "${REPO_ROOT:-}" ]; then
echo "[sleep] ERROR: could not locate the skillopt_sleep package. Set SKILLOPT_SLEEP_REPO to the repo root." >&2
exit 1
fi
PY=""
# Allow explicit Python override (useful on macOS with old system Python).
if [ -n "${SKILLOPT_SLEEP_PYTHON:-}" ]; then
PY="$SKILLOPT_SLEEP_PYTHON"
else
for cand in python3.12 python3.11 python3.10 python3; do
if command -v "$cand" >/dev/null 2>&1; then
ver="$("$cand" -c 'import sys; print("%d%d" % sys.version_info[:2])' 2>/dev/null || echo 0)"
if [ "${ver:-0}" -ge 310 ]; then PY="$cand"; break; fi
fi
done
fi
if [ -z "$PY" ]; then
echo "[sleep] ERROR: need Python >= 3.10 (found none)." >&2
exit 1
fi
if [ "$#" -eq 0 ]; then set -- status; fi
cd "$REPO_ROOT"
exec "$PY" -m skillopt_sleep "$@"
+25 -6
View File
@@ -1,11 +1,30 @@
#!/usr/bin/env bash
# Claude Code plugin runner — thin wrapper over the shared runner so all three
# platform plugins share one engine launcher. The shared runner lives at
# <repo>/plugins/run-sleep.sh and handles repo-root + interpreter resolution.
# Claude Code plugin runner — thin wrapper over the shared runner so all
# platform plugins share one engine launcher.
#
# After marketplace install the plugin is isolated in a cache directory and
# the repo-relative path no longer works. We try four locations:
# 1. Co-located run-sleep.sh (bundled copy — works in marketplace cache)
# 2. Repo-relative ../../run-sleep.sh (dev checkout)
# 3. CLAUDE_PLUGIN_ROOT/../run-sleep.sh (plugin env variable)
# 4. SKILLOPT_SLEEP_REPO/plugins/run-sleep.sh (explicit env)
set -euo pipefail
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" # <repo>/plugins/claude-code/scripts
SHARED="$(cd "$HERE/../.." && pwd)/run-sleep.sh" # <repo>/plugins/run-sleep.sh
if [ ! -f "$SHARED" ] && [ -n "${CLAUDE_PLUGIN_ROOT:-}" ]; then
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
SHARED=""
if [ -f "$HERE/run-sleep.sh" ]; then
SHARED="$HERE/run-sleep.sh"
elif [ -f "$(cd "$HERE/../.." 2>/dev/null && pwd)/run-sleep.sh" ]; then
SHARED="$(cd "$HERE/../.." && pwd)/run-sleep.sh"
elif [ -n "${CLAUDE_PLUGIN_ROOT:-}" ] && [ -f "$(cd "$CLAUDE_PLUGIN_ROOT/.." 2>/dev/null && pwd)/run-sleep.sh" ]; then
SHARED="$(cd "$CLAUDE_PLUGIN_ROOT/.." && pwd)/run-sleep.sh"
elif [ -n "${SKILLOPT_SLEEP_REPO:-}" ] && [ -f "$SKILLOPT_SLEEP_REPO/plugins/run-sleep.sh" ]; then
SHARED="$SKILLOPT_SLEEP_REPO/plugins/run-sleep.sh"
fi
if [ -z "$SHARED" ]; then
echo "[sleep] ERROR: cannot locate run-sleep.sh." >&2
echo "[sleep] Set SKILLOPT_SLEEP_REPO to the SkillOpt repo root, or pip install skillopt." >&2
exit 1
fi
exec bash "$SHARED" "$@"
@@ -54,6 +54,53 @@ Prefer the `/skillopt-sleep` command. Under the hood it calls the bundled runner
- Add `--backend claude` or `--backend codex` to spend the user's real budget for genuine improvement.
- Scope defaults to the invoked project; `--scope all` harvests every project.
### Scheduling
```bash
"${CLAUDE_PLUGIN_ROOT}/scripts/sleep.sh" schedule --project "$(pwd)" --hour 3 --minute 17
"${CLAUDE_PLUGIN_ROOT}/scripts/sleep.sh" unschedule --project "$(pwd)"
```
Installs a nightly cron entry. `unschedule --all` removes every managed entry.
## All CLI flags
| Flag | Default | Description |
|------|---------|-------------|
| `--project PATH` | cwd | Project directory to evolve |
| `--scope all\|invoked` | invoked | Harvest scope |
| `--backend mock\|claude\|codex\|copilot` | mock | Replay backend (mock = no API spend) |
| `--model NAME` | backend default | Override the model used for replay |
| `--source claude\|codex\|auto` | claude | Transcript source |
| `--lookback-hours N` | 72 | Harvest window |
| `--max-sessions N` | unlimited | Cap harvested sessions |
| `--max-tasks N` | 40 | Cap mined tasks |
| `--target-skill-path PATH` | auto | Explicit SKILL.md to evolve |
| `--tasks-file PATH` | — | Reviewed TaskRecord JSON (skip harvest) |
| `--progress` | off | Print phase progress to stderr |
| `--auto-adopt` | off | Auto-adopt if gate passes |
| `--edit-budget N` | 4 | Max bounded edits per night |
| `--json` | off | Machine-readable JSON output |
## Config keys (`~/.skillopt-sleep/config.json`)
Beyond the CLI flags, advanced behavior is controlled via config:
- **`preferences`** — free-text house rules injected into the optimizer's reflect step (e.g. "Always use async/await", "Answers in `\boxed{}`").
- **`gate_mode`** — `on` (default, validation-gated) or `off` (greedy, accept all edits).
- **`gate_metric`** — `hard`, `soft`, or `mixed` (default). Controls how the held-out gate scores.
- **`dream_rollouts`** — >1 enables multi-rollout contrastive reflection per task.
- **`recall_k`** — >0 recalls K similar past tasks into the dream (long-term memory).
- **`evolve_memory`** / **`evolve_skill`** — independently toggle CLAUDE.md vs SKILL.md consolidation.
## Memory consolidation
The sleep cycle can consolidate both:
- **SKILL.md** — the managed skill file (bounded edits: add/delete/replace)
- **CLAUDE.md** — the project memory (same bounded edits)
Both are gated by the same held-out validation score. Set `evolve_memory: false` to consolidate only skills, or `evolve_skill: false` for only memory.
## Hard rules
- **Never** hand-edit the user's `CLAUDE.md` / `SKILL.md` as part of this skill.
@@ -74,6 +121,6 @@ python -m skillopt_sleep.experiments.run_experiment --persona researcher --asser
python -m skillopt_sleep.experiments.run_experiment --persona programmer --assert-improves
```
See `docs/sleep/experiment_results.md` for recorded output and
See the SkillOpt-Sleep guide section for recorded output and
`docs/superpowers/specs/2026-06-07-skillopt-sleep-claude-code-plugin-design.md`
for the full design.
+55 -18
View File
@@ -9,51 +9,88 @@ as the Claude Code plugin (`skillopt_sleep`), wrapped for Codex.
> [gbrain-evals](https://github.com/garrytan/gbrain-evals) `skillopt-v1`
> benchmark, a deliberately deficient skill goes **0.00 → 1.00** on a held-out
> set with the Codex backend (incl. the tool-use seed via a real tool loop).
> See [`../../docs/sleep/FINAL_REPORT.md`](../../docs/sleep/FINAL_REPORT.md).
> See [the SkillOpt-Sleep guide section](https://microsoft.github.io/SkillOpt/docs/guideline.html#sleep).
## What Codex supports (and what we use)
Codex (`@openai/codex`) extends via **`AGENTS.md`** instructions, **skills** at
`~/.agents/skills/<name>/SKILL.md`, and **custom prompts** at
`~/.codex/prompts/<name>.md` (invoked as `/<name>`). This integration ships all
three, plus a shared runner.
`~/.agents/skills/<name>/SKILL.md`, and plugins that can distribute skills.
Custom prompts are deprecated in Codex, so this integration is skill-first: the
installed `skillopt-sleep` skill contains the launch commands and operating
rules. The shared runner remains a plain shell entrypoint that the skill calls.
## Install
```bash
git clone <repo-url> SkillOpt-Sleep
cd SkillOpt-Sleep
bash plugins/codex/install.sh # installs the /skillopt-sleep prompt + skill
bash plugins/codex/install.sh # installs the skill
export SKILLOPT_SLEEP_REPO="$(pwd)" # so the runner is found from anywhere
```
If a previous install created `~/.codex/prompts/sleep.md`, the installer moves
that deprecated prompt aside with a `.skillopt-legacy*.bak` suffix.
Requires Python ≥ 3.10 and the `codex` CLI on PATH.
## Use
Mention `$skillopt-sleep` where Codex supports explicit skill mentions, or ask
Codex in natural language:
```text
/skillopt-sleep status # what's happened
/skillopt-sleep dry-run # safe preview, stages nothing
/skillopt-sleep run # full cycle, stages a reviewed proposal (no live edits)
/skillopt-sleep adopt # apply the staged proposal (with backup)
Use the skillopt-sleep skill to run status for this project.
Use the skillopt-sleep skill to run a dry-run for this project.
Use the skillopt-sleep skill to run the full cycle for this project with the Codex backend.
Use the skillopt-sleep skill to adopt the latest staged proposal.
```
Or call the engine directly:
```bash
python -m skillopt_sleep run --project "$(pwd)" --backend codex
python -m skillopt_sleep dry-run --project "$(pwd)" --source codex --backend mock
python -m skillopt_sleep run --project "$(pwd)" --source codex --backend codex \
--max-sessions 5 --max-tasks 3 --progress
python -m skillopt_sleep run --project "$(pwd)" --source codex --backend codex \
--target-skill-path .agents/skills/example/SKILL.md \
--max-sessions 5 --max-tasks 3 --progress
```
Default backend is `mock` (no API spend). `--backend codex` uses your Codex
budget for real improvement. All the controllable knobs (`--gate on|off`,
`--rollouts-k`, `--budget-tokens`, `--preferences`, optimizer/target split) work
identically — see [`../../docs/sleep/CONTROLLABLE_DREAMING.md`](../../docs/sleep/CONTROLLABLE_DREAMING.md).
`--source codex` reads Codex Desktop archived sessions from
`~/.codex/archived_sessions`. Use `--codex-home /path/to/.codex` to point at a
different Codex home, or `--source auto` to try Codex archives first and fall
back to Claude Code transcripts. Default backend is `mock` (no API spend).
`--backend codex` uses your Codex budget for real improvement. Bound live runs
with `--max-sessions` and `--max-tasks`; add `--progress` because Codex-backed
mining, replay, and reflection can be slow and otherwise quiet. Use
`--target-skill-path` to stage/adopt into a repo-scoped Codex skill such as
`.agents/skills/<name>/SKILL.md`; target runs over-sample mined tasks and
prefer tasks that match the target skill's path, headings, and content. All the
controllable knobs (`--gate on|off`, `--rollouts-k`, `--budget-tokens`,
`--preferences`, optimizer/target split) work identically — see
[the SkillOpt-Sleep guide section](https://microsoft.github.io/SkillOpt/docs/guideline.html#sleep).
For privacy-sensitive projects, split the run into reviewable steps:
```bash
python -m skillopt_sleep harvest --project "$(pwd)" --source codex \
--target-skill-path .agents/skills/example/SKILL.md \
--max-sessions 5 --max-tasks 3 \
--output reviewed-tasks.json
python -m skillopt_sleep dry-run --project "$(pwd)" --backend codex \
--tasks-file reviewed-tasks.json --progress --json
```
Inspect/redact the JSON and set `"reviewed": true` before using a real backend.
`--tasks-file` skips archive harvest/mining and replays only the reviewed JSON
tasks; real backends refuse task files still marked `"reviewed": false`.
## Notes / status
- Codex's `exec` runs shell, so the real-tool-loop replay (e.g. the
`tool_called: search` benchmark seed) works natively.
- Codex's standalone *plugin-package manifest* format is not yet a stable public
spec; this integration uses the documented `AGENTS.md` + skills + prompts
mechanisms, which are stable. If/when a `codex plugin` package format ships,
we'll add a one-file manifest.
- This integration no longer installs a `.codex/prompts` slash command. Skills
are the reusable Codex workflow surface; mention `skillopt-sleep` explicitly
or ask for a sleep/dream/offline self-improvement run and Codex can load the
skill.
+19 -11
View File
@@ -1,24 +1,30 @@
#!/usr/bin/env bash
# Install the SkillOpt-Sleep Codex integration into the user's ~/.codex and
# ~/.agents directories. Idempotent; prints what it does.
# Install the SkillOpt-Sleep Codex integration as a user-level Codex skill.
# Idempotent; prints what it does.
set -euo pipefail
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
CODEX_HOME="${CODEX_HOME:-$HOME/.codex}"
AGENTS_SKILLS="${HOME}/.agents/skills"
LEGACY_PROMPT="$CODEX_HOME/prompts/sleep.md"
echo "[install] repo: $REPO_ROOT"
# 1) custom /skillopt-sleep prompt
mkdir -p "$CODEX_HOME/prompts"
cp "$REPO_ROOT/plugins/codex/prompts/skillopt-sleep.md" "$CODEX_HOME/prompts/skillopt-sleep.md"
echo "[install] /skillopt-sleep prompt -> $CODEX_HOME/prompts/skillopt-sleep.md"
# 2) user-level skill
# 1) user-level skill
mkdir -p "$AGENTS_SKILLS/skillopt-sleep"
cp "$REPO_ROOT/plugins/codex/skills/skillopt-sleep/SKILL.md" "$AGENTS_SKILLS/skillopt-sleep/SKILL.md"
echo "[install] skill -> $AGENTS_SKILLS/skillopt-sleep/SKILL.md"
# 2) retire the old custom prompt entrypoint from previous installs
if [ -f "$LEGACY_PROMPT" ]; then
backup="${LEGACY_PROMPT}.skillopt-legacy.bak"
if [ -e "$backup" ]; then
backup="${LEGACY_PROMPT}.skillopt-legacy.$(date +%Y%m%d%H%M%S).bak"
fi
mv "$LEGACY_PROMPT" "$backup"
echo "[install] legacy prompt -> $backup"
fi
# 3) record the repo location so the runner is found from anywhere
echo "[install] add to your shell profile:"
echo " export SKILLOPT_SLEEP_REPO=\"$REPO_ROOT\""
@@ -29,8 +35,10 @@ cat <<EOF
[install] Optional — add this to ~/.codex/AGENTS.md so Codex always knows the tool:
## SkillOpt-Sleep
An offline self-improvement cycle is available. To run it:
\`bash "$REPO_ROOT/plugins/run-sleep.sh" status\`. Use \`/skillopt-sleep\` for the guided flow.
Use the skillopt-sleep skill when I ask to run a sleep/dream/offline
self-improvement cycle. The runner is:
\`bash "$REPO_ROOT/plugins/run-sleep.sh" status --project "\$(pwd)"\`.
Done. Try: /skillopt-sleep status
Done. Try asking Codex:
Use the skillopt-sleep skill to run status for this project.
EOF
-21
View File
@@ -1,21 +0,0 @@
# /skillopt-sleep — SkillOpt-Sleep for Codex
#
# Custom prompt: copy this file to ~/.codex/prompts/skillopt-sleep.md and invoke with
# `/skillopt-sleep` in the Codex CLI. ($ARGUMENTS is the text after /skillopt-sleep.)
Run the SkillOpt-Sleep offline self-evolution cycle. Action: $ARGUMENTS
(empty → "status").
Use the bundled runner via shell:
bash "${SKILLOPT_SLEEP_REPO:?set SKILLOPT_SLEEP_REPO to the repo root}/plugins/run-sleep.sh" $ARGUMENTS --project "$(pwd)"
Then:
- For `run`/`dry-run`: read the staged `report.md` and show the held-out
baseline → candidate score and the proposed edits. `run` only stages a
proposal; nothing live changes until `adopt`.
- For `adopt`: confirm which files were updated and that a backup was written.
- Never edit the user's AGENTS.md / skills yourself; only `adopt` does that.
Default backend is `mock` (no API spend). Add `--backend codex` for real
improvement on the user's Codex budget.
+103 -20
View File
@@ -1,49 +1,132 @@
---
name: skillopt-sleep
description: Nightly offline self-evolution for a Codex agent. Reviews past sessions, replays recurring tasks, and consolidates validated memory + skills behind a held-out gate. Use when the user wants Codex to learn from past usage, run a "sleep"/"dream" cycle, or schedule offline self-optimization.
description: "Use when the user wants Codex to self-improve from past usage, asks about a nightly/offline 'sleep' or 'dream' cycle, wants Codex to review past sessions, learn preferences, consolidate memory/skills, run dry-run/run/adopt/status for SkillOpt-Sleep, or schedule offline self-optimization. Drives the skillopt_sleep engine: harvest past sessions -> mine recurring tasks -> replay offline -> consolidate validated memory + skills behind a held-out gate."
---
# SkillOpt-Sleep (Codex skill)
# SkillOpt-Sleep: offline self-evolution for a local Codex agent
This skill drives the `skillopt_sleep` engine — an offline "sleep cycle" that
makes a Codex agent better at the user's recurring work without retraining.
SkillOpt-Sleep gives the user's Codex agent a sleep cycle. While the user is
offline or on demand, it reviews past local sessions, re-runs recurring tasks
on the user's own budget, and consolidates what it learns into memory and
skills. It keeps only changes that pass a held-out validation gate, and live
files change only after the user explicitly adopts a staged proposal. There is
no model-weight training.
## When to use
Trigger when the user wants to: review past sessions, learn their preferences,
consolidate feedback into long-term memory/skills, run a nightly/offline
self-improvement cycle, or adopt a staged proposal.
Trigger when the user wants any of:
## How to run it
- Codex to learn from past sessions or get better the more they use it;
- a nightly/scheduled or on-demand sleep/dream/offline self-improvement run;
- to review past sessions and distill recurring tasks;
- to consolidate feedback into memory or managed skills;
- to run `status`, `harvest`, `dry-run`, `run`, or `adopt` for SkillOpt-Sleep.
## The cycle
1. **Harvest** - read local session transcripts according to the engine
configuration and normalize them into session digests.
2. **Mine** - turn digests into recurring `TaskRecord`s with outcomes and
checkable references where possible.
3. **Replay** - re-run mined tasks offline under the current skill and memory.
4. **Consolidate** - reflect on failures and propose bounded edits.
5. **Gate** - accept edits only when the held-out validation score improves.
6. **Stage** - write the proposal under
`<project>/.skillopt-sleep/staging/<date>/`; nothing live changes.
7. **Adopt** - only after explicit user approval, copy staged files over live
files with backups.
## How to drive it
Invoke the bundled runner via shell (Codex `exec` has shell access). The runner
finds the engine and a Python 3.10 automatically:
finds the engine and a Python >= 3.10 automatically.
```bash
# point at the repo if it isn't auto-detected from CWD:
export SKILLOPT_SLEEP_REPO=/path/to/SkillOpt-Sleep
bash "$SKILLOPT_SLEEP_REPO/plugins/run-sleep.sh" <action> --project "$(pwd)"
bash "$SKILLOPT_SLEEP_REPO/plugins/run-sleep.sh" status --project "$(pwd)"
bash "$SKILLOPT_SLEEP_REPO/plugins/run-sleep.sh" harvest --project "$(pwd)"
bash "$SKILLOPT_SLEEP_REPO/plugins/run-sleep.sh" dry-run --project "$(pwd)" --backend mock
bash "$SKILLOPT_SLEEP_REPO/plugins/run-sleep.sh" run --project "$(pwd)" --backend codex
bash "$SKILLOPT_SLEEP_REPO/plugins/run-sleep.sh" run --project "$(pwd)" --source codex # harvest from Codex Desktop
bash "$SKILLOPT_SLEEP_REPO/plugins/run-sleep.sh" adopt --project "$(pwd)"
```
`<action>` `status | dry-run | run | adopt | harvest`. Use `--backend codex`
for real improvement on the user's own Codex budget (default `mock` = no spend).
Actions are `status`, `harvest`, `dry-run`, `run`, `adopt`, `schedule`, and `unschedule`.
- Default backend is `mock`, which is deterministic and spends no API budget.
- `--backend codex` uses the user's Codex budget for real improvement.
- `--source codex` reads Codex Desktop archived sessions from `~/.codex/archived_sessions`;
use `--codex-home /path/to/.codex` if the archive lives elsewhere.
- Keep `dry-run --backend mock` as the first smoke check unless the user
explicitly asked for a real optimization run.
### Scheduling
```bash
bash "$SKILLOPT_SLEEP_REPO/plugins/run-sleep.sh" schedule --project "$(pwd)" --hour 3 --minute 17
bash "$SKILLOPT_SLEEP_REPO/plugins/run-sleep.sh" unschedule --project "$(pwd)"
```
Installs a nightly cron entry. `unschedule --all` removes every managed entry.
### All backends
- `--backend mock` — deterministic, no API spend (default)
- `--backend claude` — uses the Claude CLI
- `--backend codex` — uses the Codex CLI
- `--backend copilot` — uses the GitHub Copilot CLI
### Additional flags
| Flag | Description |
|------|-------------|
| `--auto-adopt` | Auto-adopt if the gate passes (default: stage only) |
| `--edit-budget N` | Max bounded edits per night (default: 4) |
| `--lookback-hours N` | Harvest window in hours (default: 72) |
| `--json` | Machine-readable JSON output |
### Config keys (`~/.skillopt-sleep/config.json`)
- **`preferences`** — free-text house rules for the optimizer
- **`gate_mode`** — `on` (validation-gated, default) or `off` (greedy)
- **`gate_metric`** — `hard` | `soft` | `mixed` (default)
- **`dream_rollouts`** — >1 for multi-rollout contrastive reflection
- **`recall_k`** — >0 recalls similar past tasks from the archive
### Memory consolidation
The sleep cycle consolidates both **memory** (AGENTS.md / CLAUDE.md) and **skills** (SKILL.md) by default. Each is independently toggleable via `evolve_memory` / `evolve_skill` config keys. Both are gated by the same held-out validation score.
## Steps
1. Run the requested action; capture stdout.
2. For `run`/`dry-run`: read the staged `report.md` it prints and show the user
the held-out baseline → candidate score and the exact proposed edits.
3. `run` only **stages** a proposal under `<project>/.skillopt-sleep/staging/`;
nothing live changes until `adopt`. Offer `/skillopt-sleep adopt`.
4. Never hand-edit the user's `AGENTS.md` / skills yourself — only `adopt` does,
and it backs up first.
2. For `dry-run` and `run`, report the held-out baseline -> candidate score,
gate action, task count, session count, and exact proposed edits.
3. If a staging directory is printed, read `report.md` before summarizing.
4. `run` only stages a proposal; nothing live changes until `adopt`.
5. Offer adoption only after the user has reviewed the staged proposal.
6. Never hand-edit the user's `AGENTS.md`, memory, or skills as a substitute
for `adopt`; adoption is the safety boundary and writes backups first.
## Hard rules
- Harvest is read-only. Do not edit archived sessions or raw transcripts.
- Keep raw secrets, credentials, private user data, and unsanitized transcript
contents out of messages, logs, generated artifacts, and commits.
- Show validation evidence before recommending adoption.
- Treat generated edits as proposals, not as source of truth.
- Do not rely on deprecated custom prompts or `/sleep` slash commands for this
Codex integration. This skill is the entrypoint.
## Validate
```bash
python -m skillopt_sleep dry-run --project "$(pwd)" --backend mock --json
python -m skillopt_sleep.experiments.run_gbrain --backend codex \
--seeds brief-writer --data-root /path/to/gbrain-evals/eval/data/skillopt-v1 \
--nights 2 --limit-replay 3 --limit-holdout 3
```
A deficient skill goes 0.00 → 1.00 on a held-out set; the optimizer's edits are
gated on real-task performance.
A deficient skill goes 0.00 -> 1.00 on a held-out set; the optimizer's edits
are gated on real-task performance.
+12 -3
View File
@@ -45,8 +45,17 @@ Ask Copilot things like *"run the sleep cycle"*, *"what did the last sleep
propose?"*, *"adopt the staged sleep proposal"*. Copilot calls the MCP tools:
`sleep_status`, `sleep_dry_run`, `sleep_run`, `sleep_adopt`, `sleep_harvest`.
Each tool takes optional `project`, `backend` (`mock`/`claude`/`codex`), and
`scope` arguments. Default backend is `mock` (no API spend).
Each tool takes optional `project`, `backend` (`mock`/`claude`/`codex`/`copilot`), and
`scope` arguments. Default backend is `mock` (no API spend). The `copilot`
backend drives the GitHub Copilot CLI (`copilot -p ... --output-format json`)
and requires the `copilot` CLI to be installed and authenticated.
For speed, the `copilot` backend runs each call against an isolated
`COPILOT_HOME` with built-in MCP servers and custom instructions disabled, so
your user MCP servers (including this project's own) are not spawned per call
(~5x faster). Override with `SKILLOPT_SLEEP_COPILOT_HOME=<dir>`, pick a model
with `SKILLOPT_SLEEP_COPILOT_MODEL`, or set `SKILLOPT_SLEEP_COPILOT_FULL_ENV=1`
to use your real Copilot environment instead.
## Verify the server directly (no Copilot needed)
@@ -64,4 +73,4 @@ You should see the server info and the five `sleep_*` tools.
portable of the three integrations (one server → CLI + IDE).
- The engine and all its controls (gate on/off, multi-rollout, budget,
preferences, optimizer/target split) are identical across platforms — see
[`../../docs/sleep/CONTROLLABLE_DREAMING.md`](../../docs/sleep/CONTROLLABLE_DREAMING.md).
[the SkillOpt-Sleep guide section](https://microsoft.github.io/SkillOpt/docs/guideline.html#sleep).
@@ -19,6 +19,24 @@ my preferences", or "make the agent improve from past usage", use the MCP tools:
- `sleep_run` — full cycle, stages a reviewed proposal (nothing live changes)
- `sleep_adopt` — apply the staged proposal (backs up first)
- `sleep_harvest` — list mined recurring tasks
- `sleep_schedule` — install a nightly cron entry (set `hour`/`minute`)
- `sleep_unschedule` — remove the nightly cron entry
### Key parameters (pass as MCP tool arguments)
- `backend``mock` (default, free), `claude`, `codex`, or `copilot`
- `source``claude`, `codex`, or `auto` (where to read transcripts)
- `target_skill_path` — explicit SKILL.md to evolve
- `tasks_file` — pre-built TaskRecord JSON (skip harvest)
- `max_tasks` / `max_sessions` — cap workload
- `auto_adopt` — auto-adopt if the gate passes
- `json` — machine-readable output for programmatic use
### Advanced config (`~/.skillopt-sleep/config.json`)
- `preferences` — free-text house rules for the optimizer
- `gate_mode``on` (default) or `off`; `dream_rollouts` — >1 for more signal
- `evolve_memory` / `evolve_skill` — toggle which docs consolidate
Always show the user the held-out baseline → candidate score and the proposed
edits before suggesting `sleep_adopt`. Never hand-edit the user's memory/skill
+63 -11
View File
@@ -38,16 +38,48 @@ TOOLS = [
"description": "Apply the latest staged proposal to CLAUDE.md/SKILL.md (backs up first)."},
{"name": "sleep_harvest", "action": "harvest",
"description": "Debug: list the recurring tasks mined from recent sessions."},
{"name": "sleep_schedule", "action": "schedule",
"description": "Install a nightly cron entry to run the sleep cycle automatically."},
{"name": "sleep_unschedule", "action": "unschedule",
"description": "Remove the nightly cron entry for a project."},
]
_BY_NAME = {t["name"]: t for t in TOOLS}
_TOOL_SCHEMA = {
"type": "object",
"properties": {
"project": {"type": "string", "description": "Project dir to evolve (default: cwd)."},
"backend": {"type": "string", "enum": ["mock", "claude", "codex"],
"description": "mock = no API spend (default); claude/codex = real."},
"scope": {"type": "string", "enum": ["invoked", "all"]},
"project": {"type": "string",
"description": "Project dir to evolve (default: cwd)."},
"backend": {"type": "string", "enum": ["mock", "claude", "codex", "copilot"],
"description": "mock = no API spend (default); claude/codex/copilot = real."},
"scope": {"type": "string", "enum": ["invoked", "all"],
"description": "Harvest scope (default: invoked project only)."},
"source": {"type": "string", "enum": ["claude", "codex", "auto"],
"description": "Transcript source (default: claude)."},
"model": {"type": "string",
"description": "Backend-specific model override."},
"tasks_file": {"type": "string",
"description": "Path to reviewed TaskRecord JSON (skips harvest)."},
"target_skill_path": {"type": "string",
"description": "Explicit SKILL.md path to evolve/stage/adopt."},
"progress": {"type": "boolean",
"description": "Print phase progress to stderr."},
"max_sessions": {"type": "integer",
"description": "Cap harvested sessions per run."},
"max_tasks": {"type": "integer",
"description": "Cap mined tasks per run."},
"lookback_hours": {"type": "integer",
"description": "Harvest window in hours (default: 72)."},
"auto_adopt": {"type": "boolean",
"description": "Auto-adopt if gate passes (default: false)."},
"json": {"type": "boolean",
"description": "Return machine-readable JSON output."},
"edit_budget": {"type": "integer",
"description": "Max bounded edits per night (default: 4)."},
"hour": {"type": "integer",
"description": "Hour for schedule (0-23, default: 3)."},
"minute": {"type": "integer",
"description": "Minute for schedule (0-59, default: 17)."},
},
"additionalProperties": False,
}
@@ -56,15 +88,35 @@ _TOOL_SCHEMA = {
def _run_engine(action: str, args: dict) -> str:
py = sys.executable or "python3"
cmd = [py, "-m", "skillopt_sleep", action]
if args.get("project"):
cmd += ["--project", str(args["project"])]
if args.get("backend"):
cmd += ["--backend", str(args["backend"])]
if args.get("scope"):
cmd += ["--scope", str(args["scope"])]
# String-valued flags
for flag, key in [
("--project", "project"), ("--backend", "backend"),
("--scope", "scope"), ("--source", "source"),
("--model", "model"), ("--tasks-file", "tasks_file"),
("--target-skill-path", "target_skill_path"),
]:
val = args.get(key)
if val:
cmd += [flag, str(val)]
# Integer-valued flags
for flag, key in [
("--max-sessions", "max_sessions"), ("--max-tasks", "max_tasks"),
("--lookback-hours", "lookback_hours"), ("--edit-budget", "edit_budget"),
("--hour", "hour"), ("--minute", "minute"),
]:
val = args.get(key)
if val is not None:
cmd += [flag, str(int(val))]
# Boolean flags
for flag, key in [
("--progress", "progress"), ("--auto-adopt", "auto_adopt"),
("--json", "json"),
]:
if args.get(key):
cmd.append(flag)
try:
proc = subprocess.run(cmd, cwd=REPO_ROOT, capture_output=True, text=True, timeout=3600)
except Exception as e: # noqa: BLE001
except Exception as e:
return f"[error] failed to run engine: {e}"
out = (proc.stdout or "").strip()
err = (proc.stderr or "").strip()
+98
View File
@@ -0,0 +1,98 @@
# SkillOpt — GitHub Copilot integration
Give **Copilot** (CLI or VS Code) direct access to the **SkillOpt** research
engine via a tiny **MCP server**. MCP is GitHub's supported way to extend
Copilot, so this works across Copilot CLI, VS Code, and other MCP clients with
the same server.
SkillOpt is **validation-gated, text-space skill optimization**: it reflects on
rollouts, makes bounded edits to a skill, and keeps a change only if it improves
a held-out validation set. This plugin exposes the repo's training and eval
entry points (`scripts/train.py`, `scripts/eval_only.py`) as Copilot tools.
> This is the companion to the **SkillOpt-Sleep** plugin (`../mcp_server.py`,
> `sleep_*` tools). Sleep evolves a *local coding agent* from your past
> sessions; this server drives the *research* training/eval loops on the
> benchmark configs in [`../../../configs`](../../../configs).
## What's here
| File | Purpose |
|---|---|
| `mcp_server.py` | stdlib-only MCP (stdio) server exposing `skillopt_*` tools |
| `mcp-config.example.json` | drop-in MCP server config |
| `copilot-instructions.snippet.md` | paste into `.github/copilot-instructions.md` |
## Install
Requires Python ≥ 3.10. The MCP server itself is pure stdlib, but the tools it
launches need SkillOpt's runtime deps — install the package first:
```bash
pip install -e . # or: pip install -r requirements.txt
```
1. **Register the MCP server.** Add the server to your Copilot MCP config
(Copilot CLI: `~/.copilot/mcp-config.json`; VS Code: your MCP settings).
Use `mcp-config.example.json` as a template — set `SKILLOPT_REPO` to this
repo's path:
```json
{
"mcpServers": {
"skillopt": {
"command": "python3",
"args": ["/abs/path/SkillOpt/plugins/copilot/skillopt/mcp_server.py"],
"env": { "SKILLOPT_REPO": "/abs/path/SkillOpt" }
}
}
}
```
2. **(Optional) Tell Copilot about it.** Append
`copilot-instructions.snippet.md` to your repo's
`.github/copilot-instructions.md` so Copilot reaches for the tools when the
user asks to "optimize a skill" or "train on a benchmark".
## Use
Ask Copilot things like *"what configs can I run?"*, *"optimize the searchqa
skill"*, or *"evaluate this skill on the dataset"*. Copilot calls the MCP tools:
`skillopt_list_configs`, `skillopt_train`, `skillopt_eval`.
| Tool | Required args | Notes |
|---|---|---|
| `skillopt_list_configs` | — | Lists `configs/**/*.yaml` you can pass as `config`. |
| `skillopt_train` | `config` | Runs a reflective optimization loop. Long-running; spends budget. |
| `skillopt_eval` | `config`, `skill` | Evaluates one skill markdown file; no training. |
Common optional args (both train and eval): `env`, `backend`,
`optimizer_model`, `target_model`, `out_root`, `cfg_options` (space-separated
`KEY=VALUE` YAML overrides), and `extra_args` (raw passthrough flags for the
underlying script). `skillopt_train` also accepts `num_epochs`, `batch_size`,
`seed`, and `use_gate`.
Runs can be very long. The server's subprocess timeout defaults to 6 hours;
override it with the `SKILLOPT_RUN_TIMEOUT` environment variable (seconds).
## Verify the server directly (no Copilot needed)
```bash
printf '%s\n' \
'{"jsonrpc":"2.0","id":1,"method":"initialize","params":{}}' \
'{"jsonrpc":"2.0","id":2,"method":"tools/list"}' \
'{"jsonrpc":"2.0","id":3,"method":"tools/call","params":{"name":"skillopt_list_configs","arguments":{}}}' \
| SKILLOPT_REPO="$(pwd)" python3 plugins/copilot/skillopt/mcp_server.py
```
You should see the server info, the three `skillopt_*` tools, and the list of
benchmark configs.
## Notes / status
- MCP is the stable, official Copilot extension surface, so this is portable
across Copilot CLI and IDE from one server.
- `skillopt_list_configs` is filesystem-only and safe to call anytime;
`skillopt_train` / `skillopt_eval` shell out to the repo scripts and require
the SkillOpt runtime deps (and, for real backends, model credentials — see
[`../../../.env.example`](../../../.env.example)).
@@ -0,0 +1,33 @@
<!--
Copy this block into your repo's .github/copilot-instructions.md so Copilot
knows the SkillOpt research-engine tools exist. (Copilot reads
copilot-instructions.md automatically as ambient guidance.)
-->
## SkillOpt (research skill-optimization engine)
This repo exposes the core **SkillOpt** training/eval engine via an MCP server
(`skillopt`). SkillOpt is validation-gated, text-space skill optimization: it
reflects on rollouts, makes bounded edits to a skill, and keeps a change only
if it improves a held-out validation set.
When the user asks to "optimize a skill", "train on <benchmark>", "run
SkillOpt", "evaluate this skill", or "what configs can I run", use the MCP
tools:
- `skillopt_list_configs` — list the benchmark YAML configs you can pass as `config`
- `skillopt_train` — run a reflective skill-optimization loop on a config (long-running; spends API/compute budget)
- `skillopt_eval` — evaluate a single skill markdown file on a dataset (no training)
Guidance:
- Always run `skillopt_list_configs` first if you don't already know a valid `config` path.
- `skillopt_train` and `skillopt_eval` are long-running and consume the user's
model backend/budget — confirm the `config`, `backend`, and model choices
with the user before launching, and surface the held-out gate result when the
run finishes.
- For one-off YAML overrides use `cfg_options` (e.g. `seed=123 batch_size=40`);
for any other underlying flag use `extra_args`.
This is distinct from the **SkillOpt-Sleep** MCP server (`skillopt-sleep`,
`sleep_*` tools), which evolves a local coding agent from past sessions rather
than running the research benchmarks.
@@ -0,0 +1,11 @@
{
"mcpServers": {
"skillopt": {
"command": "python3",
"args": ["plugins/copilot/skillopt/mcp_server.py"],
"env": {
"SKILLOPT_REPO": "${workspaceFolder}"
}
}
}
}
+229
View File
@@ -0,0 +1,229 @@
#!/usr/bin/env python3
"""SkillOpt (research engine) — minimal MCP server (stdio, stdlib-only).
Exposes the core SkillOpt skill-optimization engine as MCP tools so any
MCP-capable client (GitHub Copilot CLI / VS Code, Claude Desktop, etc.) can
drive it. No third-party deps: speaks JSON-RPC 2.0 over stdio with just the
handful of MCP methods clients need.
This is the companion to the SkillOpt-Sleep MCP server (``../mcp_server.py``).
Where Sleep evolves a *local agent* from past sessions, this server drives the
*research* training/eval loops from this repo (``scripts/train.py`` /
``scripts/eval_only.py``) against the benchmark configs in ``configs/``.
Tools exposed:
- skillopt_list_configs : discover the benchmark YAML configs you can use
- skillopt_train : run a reflective skill-optimization (training) loop
- skillopt_eval : evaluate a single skill on a dataset (no training)
``skillopt_train`` and ``skillopt_eval`` shell out to the repo's entry-point
scripts and stream back their stdout/stderr. Configure your client to launch:
python plugins/copilot/skillopt/mcp_server.py
"""
from __future__ import annotations
import glob
import json
import os
import subprocess
import sys
# Repo root: three levels up from plugins/copilot/skillopt/mcp_server.py
REPO_ROOT = os.environ.get("SKILLOPT_REPO") or os.path.abspath(
os.path.join(os.path.dirname(__file__), "..", "..", "..")
)
PROTOCOL_VERSION = "2024-11-05"
# Training/eval runs are long; give the engine plenty of headroom.
RUN_TIMEOUT_SECONDS = int(os.environ.get("SKILLOPT_RUN_TIMEOUT", "21600")) # 6h
def _list_configs() -> str:
"""List the benchmark configs available under configs/ (filesystem only)."""
pattern = os.path.join(REPO_ROOT, "configs", "**", "*.yaml")
paths = sorted(glob.glob(pattern, recursive=True))
if not paths:
return f"[no configs found under {os.path.join(REPO_ROOT, 'configs')}]"
rels = [os.path.relpath(p, REPO_ROOT).replace(os.sep, "/") for p in paths]
lines = ["Available SkillOpt configs (pass as `config`):", ""]
lines += [f" - {r}" for r in rels]
return "\n".join(lines)
def _run_script(script_rel: str, args: dict, *, required: tuple[str, ...] = ()) -> str:
"""Shell out to a repo entry-point script, mapping args -> --flags."""
for key in required:
if not args.get(key):
return f"[error] missing required argument: {key}"
py = sys.executable or "python3"
cmd = [py, os.path.join("scripts", script_rel)]
# Ordered flags that the train/eval scripts accept directly.
flag_args = (
"config", "skill", "split", "env", "backend",
"optimizer_model", "target_model", "out_root",
"num_epochs", "batch_size", "seed", "use_gate",
)
for key in flag_args:
val = args.get(key)
if val is None or val == "":
continue
cmd += [f"--{key}", str(val)]
# cfg-options: arbitrary KEY=VALUE YAML overrides (nargs="+").
cfg_options = args.get("cfg_options")
if cfg_options:
if isinstance(cfg_options, str):
cfg_options = cfg_options.split()
cmd += ["--cfg-options", *[str(x) for x in cfg_options]]
# extra_args: raw passthrough for any other train/eval flag.
extra = args.get("extra_args")
if extra:
if isinstance(extra, str):
extra = extra.split()
cmd += [str(x) for x in extra]
try:
proc = subprocess.run(
cmd, cwd=REPO_ROOT, capture_output=True, text=True,
timeout=RUN_TIMEOUT_SECONDS,
)
except subprocess.TimeoutExpired:
return f"[error] run exceeded {RUN_TIMEOUT_SECONDS}s timeout: {' '.join(cmd)}"
except Exception as e: # noqa: BLE001
return f"[error] failed to run script: {e}"
out = (proc.stdout or "").strip()
err = (proc.stderr or "").strip()
body = out + (("\n[stderr]\n" + err) if err else "")
return body or f"[done] exit code {proc.returncode}, no output"
TOOLS = [
{
"name": "skillopt_list_configs",
"description": "List the benchmark YAML configs under configs/ that can be passed as `config` to train/eval.",
},
{
"name": "skillopt_train",
"description": "Run a SkillOpt reflective skill-optimization (training) loop on a benchmark config. Long-running; uses your model backend/budget.",
},
{
"name": "skillopt_eval",
"description": "Evaluate a single skill markdown file on a dataset without training (scripts/eval_only.py).",
},
]
_BY_NAME = {t["name"]: t for t in TOOLS}
_NO_ARGS_SCHEMA = {"type": "object", "properties": {}, "additionalProperties": False}
_COMMON_PROPS = {
"config": {"type": "string",
"description": "Path to a benchmark YAML config (e.g. configs/searchqa/default.yaml). See skillopt_list_configs."},
"env": {"type": "string", "description": "Override the environment/adapter name (e.g. searchqa, alfworld)."},
"backend": {"type": "string", "description": "Model backend (e.g. azure_openai, claude, codex, qwen, minimax)."},
"optimizer_model": {"type": "string", "description": "Model used for reflection/skill rewriting (the optimizer)."},
"target_model": {"type": "string", "description": "Model used to execute tasks (the target)."},
"out_root": {"type": "string", "description": "Output directory root for run artifacts."},
"cfg_options": {"type": "string", "description": "Space-separated YAML overrides, e.g. 'seed=123 batch_size=40'."},
"extra_args": {"type": "string", "description": "Raw passthrough flags for the underlying script, e.g. '--workers 8 --max_turns 30'."},
}
_TRAIN_SCHEMA = {
"type": "object",
"properties": {
**_COMMON_PROPS,
"num_epochs": {"type": "integer", "description": "Number of optimization epochs."},
"batch_size": {"type": "integer", "description": "Tasks per optimization step."},
"seed": {"type": "integer", "description": "Random seed."},
"use_gate": {"type": "string", "enum": ["true", "false"],
"description": "Whether to keep the held-out validation gate on (default on)."},
},
"required": ["config"],
"additionalProperties": False,
}
_EVAL_SCHEMA = {
"type": "object",
"properties": {
**_COMMON_PROPS,
"skill": {"type": "string", "description": "Path to the skill markdown file to evaluate."},
"split": {"type": "string", "description": "Dataset split to evaluate (default: all)."},
},
"required": ["config", "skill"],
"additionalProperties": False,
}
_SCHEMA_BY_NAME = {
"skillopt_list_configs": _NO_ARGS_SCHEMA,
"skillopt_train": _TRAIN_SCHEMA,
"skillopt_eval": _EVAL_SCHEMA,
}
def _result(id_, result):
return {"jsonrpc": "2.0", "id": id_, "result": result}
def _error(id_, code, message):
return {"jsonrpc": "2.0", "id": id_, "error": {"code": code, "message": message}}
def _dispatch(name: str, args: dict) -> str:
if name == "skillopt_list_configs":
return _list_configs()
if name == "skillopt_train":
return _run_script("train.py", args, required=("config",))
if name == "skillopt_eval":
return _run_script("eval_only.py", args, required=("config", "skill"))
return f"[error] unknown tool: {name}"
def handle(req: dict):
method = req.get("method")
id_ = req.get("id")
if method == "initialize":
return _result(id_, {
"protocolVersion": PROTOCOL_VERSION,
"capabilities": {"tools": {}},
"serverInfo": {"name": "skillopt", "version": "0.1.0"},
})
if method in ("notifications/initialized", "initialized"):
return None # notification, no response
if method == "tools/list":
return _result(id_, {"tools": [
{"name": t["name"], "description": t["description"],
"inputSchema": _SCHEMA_BY_NAME[t["name"]]}
for t in TOOLS
]})
if method == "tools/call":
params = req.get("params") or {}
name = params.get("name")
if name not in _BY_NAME:
return _error(id_, -32602, f"unknown tool: {name}")
text = _dispatch(name, params.get("arguments") or {})
return _result(id_, {"content": [{"type": "text", "text": text}]})
if method == "ping":
return _result(id_, {})
return _error(id_, -32601, f"method not found: {method}")
def main() -> int:
for line in sys.stdin:
line = line.strip()
if not line:
continue
try:
req = json.loads(line)
except Exception:
continue
resp = handle(req)
if resp is not None:
sys.stdout.write(json.dumps(resp) + "\n")
sys.stdout.flush()
return 0
if __name__ == "__main__":
raise SystemExit(main())
+66
View File
@@ -0,0 +1,66 @@
# SkillOpt-Sleep — Devin integration
Give **Devin** (Cognition) a nightly **sleep cycle** via a tiny **MCP server**
that exposes the `skillopt_sleep` engine as tools. MCP is Devin's supported way
to add custom tooling, so this works in Devin's CLI and IDE.
Devin doesn't write transcripts in the format the engine consumes, so this
plugin adds a **Devin-specific harvester** that converts every locally available
source into the Claude Code-compatible JSONL the engine reads.
## What's here
| File | Purpose |
|---|---|
| `mcp_server.py` | stdlib-only MCP (stdio) server exposing `sleep_*` tools |
| `harvest_devin.py` | converts Devin ATIF-v1.7 transcripts + agentmemory + `.devin/skills` into JSONL, with `taskKey` + outcome envelopes |
| `judge.py` | reference judge for the deferred/judge branch of the validation gate |
| `mcp-config.example.json` | drop-in MCP server config |
| `devin-rules.snippet.md` | paste into `.devin/rules/skillopt-sleep.md` |
## What it harvests
| Source | Where |
|---|---|
| Devin transcripts (ATIF-v1.7) | `~/.local/share/devin/cli/transcripts/*.json` |
| agentmemory | `~/.agentmemory/standalone.json` |
| Skill files | `.devin/skills/*/SKILL.md` |
Workspaces are auto-detected from `~/.config/Devin/User/workspaceStorage/*/workspace.json`.
After `sleep_adopt`, the evolved skill is synced to `.devin/skills/skillopt-sleep-learned/SKILL.md`.
## Install
Requires Python ≥ 3.10. No third-party packages — the server is pure stdlib.
1. **Register the MCP server.** Use `mcp-config.example.json` as a template; set
`args` to the absolute path of this `mcp_server.py`. The engine is found
automatically (this plugin lives inside the SkillOpt repo). Or via the Devin
CLI:
```bash
devin mcp add skillopt-sleep \
--env "SKILLOPT_DEVIN_CLAUDE_HOME=$HOME/.skillopt-sleep-devin" \
-- python3 /abs/path/to/SkillOpt/plugins/devin/mcp_server.py
```
2. **(Optional)** copy `devin-rules.snippet.md` to `.devin/rules/skillopt-sleep.md`
so Devin proactively offers the tools.
3. Ask Devin: *"run the sleep cycle"*, *"what did the last sleep propose?"*, *"adopt it"*.
## Tools
| Tool | What it does |
|---|---|
| `sleep_status` | nights run so far + latest staged proposal |
| `sleep_dry_run` | preview cycle — no staging, no changes |
| `sleep_run` | full cycle; stages a proposal for review |
| `sleep_adopt` | apply the staged proposal; syncs skill to the workspace |
| `sleep_harvest` | debug: list the recurring tasks mined |
| `sleep_schedule` | install a nightly cron entry (`--hour` / `--minute`) |
| `sleep_unschedule` | remove the nightly cron entry |
Default backend is `mock` (no API spend); `--backend claude|codex` uses your own
budget. Same engine and `sleep_*` interface as the other plugins — all call
`python -m skillopt_sleep`.
+19
View File
@@ -0,0 +1,19 @@
# SkillOpt-Sleep (Devin)
You have access to a nightly self-evolution cycle via the `skillopt-sleep` MCP
server. Use these tools to improve your long-term skills over time:
- **`sleep_status`** — how many nights have run + the latest staged proposal
- **`sleep_dry_run`** — preview a cycle without changing anything
- **`sleep_run`** — run a full cycle; stages a proposal for review
- **`sleep_adopt`** — apply the staged proposal to `.devin/skills/skillopt-sleep-learned/SKILL.md`
- **`sleep_harvest`** — debug: list the recurring tasks mined from recent sessions
- **`sleep_schedule`** / **`sleep_unschedule`** — install/remove a nightly cron run
When a user asks about the sleep cycle, skill evolution, or improving your
long-term memory, prefer calling these tools over explaining the concept.
Default backend is `mock` (no API spend). Pass `backend: "claude"` or
`backend: "codex"` with your own API key for real LLM-driven optimization.
Place this file at `.devin/rules/skillopt-sleep.md` in your workspace.
+21
View File
@@ -0,0 +1,21 @@
{
"schema_version": "ATIF-v1.7",
"session_id": "demo-001",
"steps": [
{
"source": "user",
"message": "Fix the failing NullPointerException in OrderService.persist() in the dutch-kis project",
"timestamp": "2026-06-20T10:00:00Z"
},
{
"source": "agent",
"message": "The repository call returns an Optional that is being unwrapped with .get(). I'll switch to orElseThrow(NotFoundException::new) so the missing-row case is handled.",
"timestamp": "2026-06-20T10:00:05Z"
},
{
"source": "agent",
"message": "Applied the fix and ran the suite: rtk mvn test -Dtest=OrderServiceTest -> BUILD SUCCESS, 142 passed, 0 failed.",
"timestamp": "2026-06-20T10:01:00Z"
}
]
}
+533
View File
@@ -0,0 +1,533 @@
#!/usr/bin/env python3
"""Convert Devin IDE local data into Claude Code-format JSONL transcripts.
Devin (Cognition) does not persist agent conversation transcripts to disk in a
format the sleep engine understands. This script bridges that gap by synthesising
JSONL files from every locally available source:
1. **Devin transcripts** (~/.local/share/devin/cli/transcripts/*.json)
Native ATIF-v1.7 format — source:"user" / source:"agent" messages
converted directly to user/assistant JSONL turns.
2. **agentmemory** (~/.agentmemory/standalone.json)
Memories saved by the `agentmemory` MCP server — each memory's title
becomes a synthetic user prompt; its content becomes the assistant reply.
3. **Skill files** (.devin/skills/*/SKILL.md)
Each skill description is converted to a session where the user asked
"use the <skill> skill" and the assistant described how to apply it.
Output layout (mirrors ~/.claude/projects/<slug>/<sessionId>.jsonl):
<out_dir>/projects/<slug>/<session_id>.jsonl
Workspace auto-detection order:
1. ``SKILLOPT_DEVIN_WORKSPACES`` env var — colon-separated abs paths
2. Devin registry: ``~/.config/Devin/User/workspaceStorage/*/workspace.json``
4. Working directory fallback
Usage (standalone):
python harvest_devin.py [--out-dir PATH] [--workspaces PATH ...]
"""
from __future__ import annotations
import argparse
import hashlib
import json
import os
import re
import sys
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Dict, List, Optional
from urllib.parse import unquote, urlparse
# ── cross-platform path resolution (Linux + Windows + macOS) ──────────────────
#
# Devin is a VS Code-family app, so its user-data dir moves with the OS:
# Linux ~/.config/<App>, Windows %APPDATA%\<App>, macOS
# ~/Library/Application Support/<App>. Resolve all candidates and let callers
# keep whichever actually exists.
def _app_data_roots(app: str) -> List[str]:
"""User-data dir candidates for a VS Code-family app, current OS first."""
home = os.path.expanduser("~")
roots: List[str] = []
if os.name == "nt":
appdata = os.environ.get("APPDATA") or os.path.join(home, "AppData", "Roaming")
roots.append(os.path.join(appdata, app))
elif sys.platform == "darwin":
roots.append(os.path.join(home, "Library", "Application Support", app))
# XDG / Linux (also a sensible fallback everywhere)
xdg = os.environ.get("XDG_CONFIG_HOME") or os.path.join(home, ".config")
roots.append(os.path.join(xdg, app))
# de-dupe, preserve order
return list(dict.fromkeys(roots))
def _devin_transcript_candidates() -> List[str]:
"""Where the Devin CLI may store ATIF transcripts, per OS."""
home = os.path.expanduser("~")
cands: List[str] = []
if os.name == "nt":
for base in (os.environ.get("LOCALAPPDATA"), os.environ.get("APPDATA")):
if base:
cands.append(os.path.join(base, "devin", "cli", "transcripts"))
elif sys.platform == "darwin":
cands.append(os.path.join(home, "Library", "Application Support",
"devin", "cli", "transcripts"))
cands.append(os.path.join(home, ".local", "share", "devin", "cli", "transcripts"))
return list(dict.fromkeys(cands))
def _first_existing(paths: List[str]) -> str:
"""First path that exists, else the first candidate (for nice messaging)."""
for p in paths:
if os.path.exists(p):
return p
return paths[0] if paths else ""
def _uri_to_path(folder: str) -> str:
"""Convert a VS Code ``file://`` workspace URI to a local path, cross-platform.
Linux: file:///home/u/proj -> /home/u/proj
Windows: file:///c%3A/Users/u/p -> c:/Users/u/p
"""
if not folder.startswith("file://"):
return folder
path = unquote(urlparse(folder).path)
# Windows drive paths come through as '/C:/...' — strip the leading slash.
if os.name == "nt" and re.match(r"^/[A-Za-z]:", path):
path = path[1:]
return path
# ── workspace auto-detection ─────────────────────────────────────────────────
def _workspaces_from_registry(storage_root: str) -> List[tuple]:
"""Read VS Code-style workspaceStorage to get (mtime, path) pairs."""
results: List[tuple] = []
if not os.path.isdir(storage_root):
return results
for entry in os.scandir(storage_root):
ws_json = os.path.join(entry.path, "workspace.json")
if not os.path.isfile(ws_json):
continue
try:
with open(ws_json, encoding="utf-8") as f:
data = json.load(f)
folder = _uri_to_path(data.get("folder", ""))
if folder and os.path.isdir(folder):
results.append((os.path.getmtime(ws_json), folder))
except Exception:
continue
return results
def _detect_workspaces() -> List[str]:
"""Return known workspace paths (Devin registry), newest first."""
env_val = os.environ.get("SKILLOPT_DEVIN_WORKSPACES", "")
if env_val:
# os.pathsep so Windows 'C:\a;C:\b' splits correctly (not on the drive colon)
return [p for p in env_val.split(os.pathsep) if p and os.path.isdir(p)]
registries: List[str] = [
os.path.join(r, "User", "workspaceStorage")
for r in _app_data_roots("Devin")
]
seen: set = set()
results: List[tuple] = []
for registry in registries:
for mtime, folder in _workspaces_from_registry(registry):
if folder not in seen:
seen.add(folder)
results.append((mtime, folder))
results.sort(reverse=True)
paths = [p for _, p in results]
return paths if paths else [os.getcwd()]
# ── helpers ───────────────────────────────────────────────────────────────────
def _slug(path: str) -> str:
"""SHA-256 of abs-path, first 16 hex chars — matches Claude Code's scheme."""
return hashlib.sha256(os.path.abspath(path).encode()).hexdigest()[:16]
def _iso(epoch_ms: Optional[float] = None) -> str:
dt = (datetime.fromtimestamp(epoch_ms / 1000.0, tz=timezone.utc)
if epoch_ms is not None else datetime.now(tz=timezone.utc))
return dt.strftime("%Y-%m-%dT%H:%M:%S.000Z")
def _write_session(
out_dir: str, project: str, session_id: str,
user_prompts: List[str], assistant_replies: List[str],
timestamp_base_ms: float,
task_key: Optional[str] = None,
) -> None:
slug = _slug(project)
session_dir = os.path.join(out_dir, "projects", slug)
os.makedirs(session_dir, exist_ok=True)
out_path = os.path.join(session_dir, f"{session_id}.jsonl")
ts = timestamp_base_ms
with open(out_path, "w", encoding="utf-8") as f:
for user_text, asst_text in zip(user_prompts, assistant_replies):
user_rec = {
"type": "user",
"message": {"role": "user", "content": user_text},
"cwd": project,
"timestamp": _iso(ts),
"sessionId": session_id,
"version": "1.0",
}
if task_key:
# grouping key so the miner can collapse repeats into one recurring task
user_rec["taskKey"] = task_key
f.write(json.dumps(user_rec, ensure_ascii=False) + "\n")
# space the reply >=5s after the prompt so a single-turn session
# isn't misclassified as a <3s headless replay and dropped by the
# engine's harvest filter (skillopt_sleep Issue #62).
ts += 5000
f.write(json.dumps({
"type": "assistant",
"message": {"role": "assistant", "content": asst_text},
"timestamp": _iso(ts),
"sessionId": session_id,
"version": "1.0",
}, ensure_ascii=False) + "\n")
ts += 2000
def _append_history(out_dir: str, display: str, project: str, timestamp_ms: float) -> None:
record = {"display": display, "timestamp": timestamp_ms, "project": project}
with open(os.path.join(out_dir, "history.jsonl"), "a", encoding="utf-8") as f:
f.write(json.dumps(record, ensure_ascii=False) + "\n")
def _infer_project(text: str, workspaces: List[str]) -> str:
for ws in workspaces:
if os.path.basename(ws.rstrip("/")).lower() in text.lower():
return ws
return workspaces[0] if workspaces else os.getcwd()
# ── task identity + outcome extraction (fuel for the validation gate) ─────────
#
# SkillOpt's gate only works "where tasks recur and have a checkable correctness
# signal." These helpers add the two things a raw transcript lacks:
# * a stable taskKey so repeats collapse into one recurring task, and
# * an outcome envelope (success + verifier + re-runnable reference) so the
# held-out replay has something to score against.
_LANG_HINTS = [
("java", r"(java|spring|maven|\bmvn\b|gradle|\.java\b|lombok)"),
("python", r"(python|pytest|\bpip\b|\.py\b|django|flask)"),
("ts", r"(typescript|\.tsx?\b|\bnpm\b|jest|node)"),
("js", r"(javascript|\.jsx?\b)"),
("sql", r"(\bsql\b|select\s|mariadb|mysql|postgres|\.sql\b)"),
("go", r"(golang|\bgo test\b|\.go\b)"),
("rust", r"(rust|cargo|\.rs\b)"),
]
_INTENT_HINTS = [
("fix", r"(fix|bug|error|fail|npe|exception|broken|crash)"),
("implement", r"(implement|add|create|build|introduce|support)"),
("refactor", r"(refactor|clean ?up|rename|extract|simplify)"),
("test", r"(test|coverage|assert)"),
("review", r"(review|audit|inspect)"),
("optimize", r"(optimi[sz]e|perf|speed up|slow)"),
("explain", r"(explain|understand|what does|how does)"),
]
_STOPWORDS = {"please", "this", "that", "with", "from", "into", "should",
"would", "code", "using", "the", "have"}
def _normalize_task_key(text: str, project: str) -> str:
"""Stable '<lang>:<intent>:<target>' grouping key for a task."""
low = text.lower()
lang = next((n for n, pat in _LANG_HINTS if re.search(pat, low)), "general")
intent = next((n for n, pat in _INTENT_HINTS if re.search(pat, low)), "task")
# target: prefer a CamelCase identifier, then a filename, then first real word
m = re.search(r"\b([A-Z][a-z0-9]+(?:[A-Z][a-z0-9]+)+)\b", text) # CamelCase
if not m:
m = re.search(r"\b([\w-]+\.\w+)\b", text) # filename.ext
if m:
target = m.group(1)
else:
# first content word that isn't a stopword or an intent verb (e.g. "implement")
target = next((w for w in re.findall(r"[a-zA-Z]{4,}", low)
if w not in _STOPWORDS
and not any(re.search(pat, w) for _, pat in _INTENT_HINTS)),
"general")
target = re.sub(r"[^a-zA-Z0-9]+", "-", target).strip("-").lower()[:40] or "general"
return f"{lang}:{intent}:{target}"
_PASS_PAT = re.compile(
r"(build success|all tests? pass(?:ed)?|\b\d+ passed\b|\b0 failed\b|"
r"tests? pass(?:ed)?|✓|no errors)", re.IGNORECASE)
_FAIL_PAT = re.compile(
r"(build failure|tests? failed|\b[1-9]\d* failed\b|error:|traceback|"
r"assertion ?error)", re.IGNORECASE) # note: "0 failed" must NOT match
_CMD_PAT = re.compile(
r"((?:rtk\s+)?(?:mvn|gradle|pytest|npm(?:\s+run)?\s+test|yarn\s+test|"
r"go\s+test|cargo\s+test)[^\n`]*)", re.IGNORECASE)
def _detect_outcome(messages: List[str]) -> Optional[Dict[str, Any]]:
"""Best-effort checkable signal from agent messages. None ⇒ no hard signal."""
blob = "\n".join(m for m in messages if m)
pass_hit, fail_hit = _PASS_PAT.search(blob), _FAIL_PAT.search(blob)
if not pass_hit and not fail_hit:
return None
verifier = "tests" if re.search(r"test|pytest", blob, re.IGNORECASE) else "build"
out: Dict[str, Any] = {
"success": bool(pass_hit) and not fail_hit,
"verifier": verifier,
"evidence": (pass_hit or fail_hit).group(0).strip(),
}
cmd = _CMD_PAT.search(blob)
if cmd:
# keep only the command itself, dropping any "-> result" / ": output" tail
repro = re.split(r"\s*(?:->|→|:|,)\s*", cmd.group(1))[0].strip()
out["reference"] = {"repro": repro}
return out
def _build_rubric(user_prompt: str) -> List[str]:
"""Derive checkable criteria from the task so a judge has something to score."""
crit: List[str] = []
ids = re.findall(r"\b([A-Z][a-z0-9]+(?:[A-Z][a-z0-9]+)+|[\w-]+\.\w+)\b", user_prompt)
for i in dict.fromkeys(ids): # dedupe, preserve order
crit.append(f"Addresses {i}")
intent = _normalize_task_key(user_prompt, "").split(":")[1]
crit.append({
"fix": "Resolves the reported defect without introducing new errors",
"implement": "Implements the requested behavior end to end",
"refactor": "Preserves behavior while improving structure",
"test": "Adds or fixes tests that actually exercise the change",
"optimize": "Improves performance without changing results",
}.get(intent, "Satisfies the user's stated request"))
crit.append("Response is concrete and actionable, not a restatement of the task")
return crit[:5]
def _judge_rubric_fallback(user_prompt: str) -> Dict[str, Any]:
"""When no hard signal exists, attach a rubric and mark the task for judge
scoring. success=None tells the gate to defer/judge rather than trust it.
The actual scoring is done by judge.py (or the engine) at replay time."""
return {
"success": None,
"verifier": "judge",
"rubric": _build_rubric(user_prompt or ""),
}
def _write_outcome(out_dir: str, session_id: str, task_key: str, project: str,
ts_ms: float, outcome: Dict[str, Any]) -> None:
rec = {"type": "outcome", "sessionId": session_id, "taskKey": task_key,
"project": project, "timestamp": _iso(ts_ms), **outcome}
with open(os.path.join(out_dir, "outcomes.jsonl"), "a", encoding="utf-8") as f:
f.write(json.dumps(rec, ensure_ascii=False) + "\n")
# ── source 1: Devin ATIF-v1.7 transcripts ────────────────────────────────────
def harvest_devin_transcripts(
transcripts_dir: str, out_dir: str, workspaces: List[str]
) -> int:
"""Convert Devin CLI ATIF-v1.7 transcripts to Claude Code JSONL."""
if not os.path.isdir(transcripts_dir):
return 0
written = 0
for entry in os.scandir(transcripts_dir):
if not entry.name.endswith(".json"):
continue
try:
with open(entry.path, encoding="utf-8") as f:
data = json.load(f)
except Exception:
continue
if data.get("schema_version", "").startswith("ATIF"):
pass # Devin native format
else:
continue
session_id = data.get("session_id") or entry.name[:-5]
steps = data.get("steps") or []
user_prompts: List[str] = []
agent_replies: List[str] = []
project = ""
ts_base: Optional[float] = None
for step in steps:
src = step.get("source", "")
msg = str(step.get("message") or "").strip()
if not msg or src == "system":
continue
if src == "user":
user_prompts.append(msg)
if not project:
project = _infer_project(msg, workspaces)
elif src == "agent":
agent_replies.append(msg)
if ts_base is None:
raw_ts = step.get("timestamp", "")
if raw_ts:
try:
from datetime import datetime as _dt
ts_base = _dt.fromisoformat(
raw_ts.replace("Z", "+00:00")
).timestamp() * 1000
except Exception:
pass
if not user_prompts:
continue
if not project:
project = workspaces[0] if workspaces else os.getcwd()
if ts_base is None:
ts_base = datetime.now(tz=timezone.utc).timestamp() * 1000
# Identity + outcome: what makes this trajectory replayable & gradeable.
task_key = _normalize_task_key(user_prompts[0], project)
outcome = _detect_outcome(agent_replies) or _judge_rubric_fallback(user_prompts[0])
# Pair turns; pad shorter list
n = max(len(user_prompts), len(agent_replies))
user_prompts += [""] * (n - len(user_prompts))
agent_replies += [""] * (n - len(agent_replies))
sid = f"devin_{session_id}"
_write_session(
out_dir, project, sid,
user_prompts=[p for p in user_prompts if p],
assistant_replies=[r if r else "[no reply recorded]" for r, p in
zip(agent_replies, user_prompts) if p],
timestamp_base_ms=ts_base,
task_key=task_key,
)
_write_outcome(out_dir, sid, task_key, project, ts_base, outcome)
_append_history(
out_dir,
display=(user_prompts[0] or session_id)[:120],
project=project,
timestamp_ms=ts_base,
)
written += 1
return written
# ── source 2: agentmemory ─────────────────────────────────────────────────────
def harvest_agentmemory(agentmemory_path: str, out_dir: str,
workspaces: List[str]) -> int:
if not os.path.isfile(agentmemory_path):
return 0
with open(agentmemory_path, encoding="utf-8") as f:
data = json.load(f)
memories: Dict[str, Any] = data.get("mem:memories", {})
written = 0
base_ts = datetime.now(tz=timezone.utc).timestamp() * 1000 - len(memories) * 60_000
for i, (mem_id, mem) in enumerate(memories.items()):
title = str(mem.get("title", "")).strip()
content = str(mem.get("content", "")).strip()
if not title or not content:
continue
project = _infer_project(title + " " + content, workspaces)
ts = base_ts + i * 60_000
_write_session(out_dir, project, mem_id,
user_prompts=[title],
assistant_replies=[content],
timestamp_base_ms=ts)
_append_history(out_dir, display=title[:120], project=project, timestamp_ms=ts)
written += 1
return written
# ── source 3: skill files (.devin/skills) ─────────────────────────────────────
def harvest_skills(workspaces: List[str], out_dir: str) -> int:
written = 0
seen_ids: set = set()
for ws in workspaces:
skills_root = os.path.join(ws, ".devin", "skills")
if not os.path.isdir(skills_root):
continue
for skill_dir in os.scandir(skills_root):
if not skill_dir.is_dir():
continue
skill_md = os.path.join(skill_dir.path, "SKILL.md")
if not os.path.isfile(skill_md):
continue
sid = f"skill_{skill_dir.name}"
if sid in seen_ids:
continue
seen_ids.add(sid)
with open(skill_md, encoding="utf-8") as f:
raw = f.read()
body = re.sub(r"^---.*?---\s*", "", raw, flags=re.DOTALL).strip()
if not body:
continue
first_line = body.split("\n")[0].lstrip("# ").strip()
user_ask = f"Please use the {skill_dir.name} skill: {first_line}"
ts = datetime.now(tz=timezone.utc).timestamp() * 1000 - 3_600_000
_write_session(out_dir, ws, sid,
user_prompts=[user_ask],
assistant_replies=[body[:1200]],
timestamp_base_ms=ts)
_append_history(out_dir, display=user_ask[:120], project=ws, timestamp_ms=ts)
written += 1
return written
# ── main ─────────────────────────────────────────────────────────────────────
def main(argv=None) -> int:
parser = argparse.ArgumentParser(
description="Generate SkillOpt-Sleep transcripts from Devin local data"
)
parser.add_argument(
"--out-dir",
default=os.path.expanduser("~/.skillopt-sleep-devin"),
help="Output claude_home dir (default: ~/.skillopt-sleep-devin)",
)
parser.add_argument(
"--agentmemory",
default=os.path.expanduser("~/.agentmemory/standalone.json"),
help="Path to agentmemory standalone.json",
)
parser.add_argument(
"--devin-transcripts",
default=_first_existing(_devin_transcript_candidates()),
help="Devin CLI ATIF transcripts directory (default: per-OS auto-detect)",
)
parser.add_argument(
"--workspaces", nargs="*",
help="Workspace paths (default: auto-detect from Devin registry)",
)
parser.add_argument("--quiet", action="store_true")
args = parser.parse_args(argv)
out_dir = os.path.expanduser(args.out_dir)
os.makedirs(out_dir, exist_ok=True)
os.makedirs(os.path.join(out_dir, "projects"), exist_ok=True)
workspaces = args.workspaces or _detect_workspaces()
workspaces = [ws for ws in workspaces if os.path.isdir(ws)]
if not workspaces:
workspaces = [os.getcwd()]
total = 0
devin_transcripts = os.path.expanduser(args.devin_transcripts)
n = harvest_devin_transcripts(devin_transcripts, out_dir, workspaces)
if not args.quiet:
print(f"[harvest_devin] devin : {n} sessions")
total += n
n = harvest_agentmemory(args.agentmemory, out_dir, workspaces)
if not args.quiet:
print(f"[harvest_devin] agentmemory : {n} sessions")
total += n
n = harvest_skills(workspaces, out_dir)
if not args.quiet:
print(f"[harvest_devin] skill files : {n} sessions")
total += n
if not args.quiet:
print(f"[harvest_devin] total : {total} synthetic sessions → {out_dir}")
return 0
if __name__ == "__main__":
raise SystemExit(main())
+129
View File
@@ -0,0 +1,129 @@
#!/usr/bin/env python3
"""Reference judge for SkillOpt-Sleep — score a candidate reply against a rubric.
Tasks harvested without a hard test/build signal get ``verifier: "judge"`` and a
``rubric`` (see ``_build_rubric`` in harvest_devin.py). This module is the
scorer the validation gate calls for those tasks: given the rubric and a
candidate reply produced during replay, it returns a score in ``[0, 1]``. The
gate accepts a skill edit only if the *new* skill scores strictly higher on the
held-out tasks.
It is self-contained on purpose — in a full deployment the SkillOpt engine owns
replay+scoring, but having a runnable reference here lets you sanity-check the
judge path without the engine.
Backends (select via ``SKILLOPT_JUDGE``):
* ``heuristic`` (default) — keyword-coverage, offline, no API key, deterministic.
* ``claude`` — LLM judge via the Anthropic API (needs ANTHROPIC_API_KEY).
Usage:
python judge.py --rubric rubric.json --reply reply.txt
echo "<reply>" | python judge.py --rubric-inline '["Addresses OrderService", ...]'
"""
from __future__ import annotations
import argparse
import json
import os
import re
import sys
from typing import List
_STOPWORDS = {"addresses", "resolves", "implements", "without", "introducing",
"behavior", "request", "response", "concrete", "actionable", "not",
"the", "and", "that", "with", "stated", "reported", "actually",
"preserves", "improving", "structure", "requested", "satisfies"}
# Cheap, fast model is the right default for a judge.
_JUDGE_MODEL = os.environ.get("SKILLOPT_JUDGE_MODEL", "claude-haiku-4-5-20251001")
def _content_words(text: str) -> List[str]:
return [w for w in re.findall(r"[A-Za-z][A-Za-z0-9_.\-]{3,}", text.lower())
if w not in _STOPWORDS]
def heuristic_score(reply: str, rubric: List[str]) -> float:
"""Fraction of rubric criteria whose key content words appear in the reply.
Crude but deterministic: each criterion is 'met' if at least one of its
content words shows up in the candidate reply. Good enough to smoke-test the
gate wiring; swap in the claude backend for real judging.
"""
if not rubric:
return 0.0
low = reply.lower()
met = 0
for criterion in rubric:
words = _content_words(criterion)
if not words: # nothing to check → treat as met
met += 1
continue
if any(w in low for w in words):
met += 1
return round(met / len(rubric), 3)
def claude_score(reply: str, rubric: List[str]) -> float:
"""LLM judge via the Anthropic API. Returns a 0..1 score.
Stdlib-only (urllib) so this file stays dependency-free. Falls back to the
heuristic if the key is missing or the call fails, so the gate never hard-errors.
"""
api_key = os.environ.get("ANTHROPIC_API_KEY")
if not api_key:
print("[judge] ANTHROPIC_API_KEY unset — using heuristic", file=sys.stderr)
return heuristic_score(reply, rubric)
import urllib.request
rubric_block = "\n".join(f"- {c}" for c in rubric)
prompt = (
"You are scoring an AI agent's reply against a rubric. For each criterion, "
"decide if the reply satisfies it. Respond with ONLY a number between 0 and "
"1 — the fraction of criteria satisfied.\n\n"
f"Rubric:\n{rubric_block}\n\nReply:\n{reply}\n\nScore:"
)
body = json.dumps({
"model": _JUDGE_MODEL,
"max_tokens": 8,
"messages": [{"role": "user", "content": prompt}],
}).encode()
req = urllib.request.Request(
"https://api.anthropic.com/v1/messages", data=body,
headers={"content-type": "application/json", "x-api-key": api_key,
"anthropic-version": "2023-06-01"},
)
try:
with urllib.request.urlopen(req, timeout=30) as resp:
data = json.load(resp)
text = "".join(b.get("text", "") for b in data.get("content", []))
m = re.search(r"[01](?:\.\d+)?", text)
return max(0.0, min(1.0, float(m.group(0)))) if m else heuristic_score(reply, rubric)
except Exception as exc: # network/auth/parse — degrade gracefully
print(f"[judge] claude backend failed ({exc}) — using heuristic", file=sys.stderr)
return heuristic_score(reply, rubric)
def score(reply: str, rubric: List[str]) -> float:
backend = os.environ.get("SKILLOPT_JUDGE", "heuristic")
return claude_score(reply, rubric) if backend == "claude" else heuristic_score(reply, rubric)
def main(argv=None) -> int:
p = argparse.ArgumentParser(description="Score a reply against a rubric (0..1)")
g = p.add_mutually_exclusive_group(required=True)
g.add_argument("--rubric", help="Path to a JSON file containing a list of criteria")
g.add_argument("--rubric-inline", help="Inline JSON list of criteria")
p.add_argument("--reply", help="Path to the reply text (default: stdin)")
args = p.parse_args(argv)
rubric = (json.load(open(args.rubric, encoding="utf-8")) if args.rubric
else json.loads(args.rubric_inline))
reply = (open(args.reply, encoding="utf-8").read() if args.reply
else sys.stdin.read())
print(score(reply, rubric))
return 0
if __name__ == "__main__":
raise SystemExit(main())
+11
View File
@@ -0,0 +1,11 @@
{
"mcpServers": {
"skillopt-sleep": {
"command": "python3",
"args": ["/abs/path/to/SkillOpt/plugins/devin/mcp_server.py"],
"env": {
"SKILLOPT_DEVIN_CLAUDE_HOME": "~/.skillopt-sleep-devin"
}
}
}
}
+240
View File
@@ -0,0 +1,240 @@
#!/usr/bin/env python3
"""SkillOpt-Sleep — Devin MCP server (stdio, stdlib-only).
Exposes the sleep engine as MCP tools so Devin (Cognition) can drive it. No
third-party deps: speaks JSON-RPC 2.0 over stdio with just the handful of MCP
methods clients need. Same `sleep_*` interface and engine flags as
`plugins/copilot`, plus a Devin-specific harvest step.
Before each data-reading action this server runs `harvest_devin.py` to convert
locally available Devin data (ATIF-v1.7 transcripts, agentmemory memories, and
.devin skill files) into the Claude Code-compatible JSONL the engine consumes,
writing it under SKILLOPT_DEVIN_CLAUDE_HOME and pointing the engine there with
`--claude-home`. After `sleep_adopt` the evolved skill is synced back into the
workspace's `.devin/skills/`.
Tools: sleep_status, sleep_dry_run, sleep_run, sleep_adopt, sleep_harvest,
sleep_schedule, sleep_unschedule. Each shells out to
`python -m skillopt_sleep <action> ...`. Configure Devin to launch:
python plugins/devin/mcp_server.py
"""
from __future__ import annotations
import json
import os
import shutil
import subprocess
import sys
# expanduser wraps the whole value so a "~/..." env var is expanded too (not
# just a default) — otherwise a literal ~ dir gets created.
REPO_ROOT = os.path.expanduser(
os.environ.get("SKILLOPT_SLEEP_REPO")
or os.path.abspath(os.path.join(os.path.dirname(__file__), "..", ".."))
)
PLUGIN_DIR = os.path.dirname(os.path.abspath(__file__))
CLAUDE_HOME = os.path.expanduser(
os.environ.get("SKILLOPT_DEVIN_CLAUDE_HOME", "~/.skillopt-sleep-devin")
)
MANAGED_SKILL_NAME = os.environ.get("SKILLOPT_MANAGED_SKILL", "skillopt-sleep-learned")
PROTOCOL_VERSION = "2024-11-05"
TOOLS = [
{"name": "sleep_status", "action": "status",
"description": "Show how many SkillOpt-Sleep nights have run and the latest staged proposal."},
{"name": "sleep_dry_run", "action": "dry-run",
"description": "Preview a sleep cycle (harvest+mine+replay) without staging anything."},
{"name": "sleep_run", "action": "run",
"description": "Run a full sleep cycle; stages a reviewed proposal. Nothing live changes until adopt."},
{"name": "sleep_adopt", "action": "adopt",
"description": "Apply the latest staged proposal to the managed SKILL.md and sync it into .devin/skills/."},
{"name": "sleep_harvest", "action": "harvest",
"description": "Debug: list the recurring tasks mined from recent Devin sessions."},
{"name": "sleep_schedule", "action": "schedule",
"description": "Install a nightly cron entry to run the sleep cycle automatically."},
{"name": "sleep_unschedule", "action": "unschedule",
"description": "Remove the nightly cron entry for a project."},
]
_BY_NAME = {t["name"]: t for t in TOOLS}
_TOOL_SCHEMA = {
"type": "object",
"properties": {
"project": {"type": "string",
"description": "Project dir to evolve (default: cwd)."},
"backend": {"type": "string", "enum": ["mock", "claude", "codex", "copilot"],
"description": "mock = no API spend (default); claude/codex/copilot = real."},
"scope": {"type": "string", "enum": ["invoked", "all"],
"description": "Harvest scope (default: invoked project only)."},
"source": {"type": "string", "enum": ["claude", "codex", "auto"],
"description": "Transcript source (default: claude)."},
"model": {"type": "string",
"description": "Backend-specific model override."},
"tasks_file": {"type": "string",
"description": "Path to reviewed TaskRecord JSON (skips harvest)."},
"target_skill_path": {"type": "string",
"description": "Explicit SKILL.md path to evolve/stage/adopt."},
"progress": {"type": "boolean",
"description": "Print phase progress to stderr."},
"max_sessions": {"type": "integer",
"description": "Cap harvested sessions per run."},
"max_tasks": {"type": "integer",
"description": "Cap mined tasks per run."},
"lookback_hours": {"type": "integer",
"description": "Harvest window in hours (default: 72)."},
"auto_adopt": {"type": "boolean",
"description": "Auto-adopt if gate passes (default: false)."},
"json": {"type": "boolean",
"description": "Return machine-readable JSON output."},
"edit_budget": {"type": "integer",
"description": "Max bounded edits per night (default: 4)."},
"hour": {"type": "integer",
"description": "Hour for schedule (0-23, default: 3)."},
"minute": {"type": "integer",
"description": "Minute for schedule (0-59, default: 17)."},
},
"additionalProperties": False,
}
# actions that read harvested Devin data (schedule/unschedule/adopt don't)
_HARVEST_ACTIONS = {"status", "dry-run", "run", "harvest"}
def _run_harvest() -> str:
"""Convert local Devin data into the JSONL the engine reads, under CLAUDE_HOME."""
harvester = os.path.join(PLUGIN_DIR, "harvest_devin.py")
env = dict(os.environ)
env["PYTHONPATH"] = REPO_ROOT + os.pathsep + env.get("PYTHONPATH", "")
try:
proc = subprocess.run(
[sys.executable, harvester, "--out-dir", CLAUDE_HOME],
capture_output=True, text=True, timeout=60, env=env,
)
out = (proc.stdout or "").strip()
err = (proc.stderr or "").strip()
return out + (("\n[harvest stderr]\n" + err) if err else "")
except Exception as exc:
return f"[harvest_devin] warning: {exc}"
def _sync_skill(project: str) -> str:
"""After adopt, copy the evolved skill into the workspace's .devin/skills/."""
src = os.path.join(CLAUDE_HOME, "skills", MANAGED_SKILL_NAME, "SKILL.md")
if not (os.path.isfile(src) and project and os.path.isdir(project)):
return ""
dot_root = os.path.join(project, ".devin")
if not os.path.isdir(dot_root):
return ""
dst_dir = os.path.join(dot_root, "skills", MANAGED_SKILL_NAME)
os.makedirs(dst_dir, exist_ok=True)
dst = os.path.join(dst_dir, "SKILL.md")
shutil.copy2(src, dst)
return f"\n[sleep] synced evolved skill → {dst}"
def _run_engine(action: str, args: dict) -> str:
harvest_out = _run_harvest() if action in _HARVEST_ACTIONS else ""
py = sys.executable or "python3"
cmd = [py, "-m", "skillopt_sleep", action, "--claude-home", CLAUDE_HOME]
# Devin transcripts are converted to the Claude format, so default source=claude
if not args.get("source"):
cmd += ["--source", "claude"]
# String-valued flags
for flag, key in [
("--project", "project"), ("--backend", "backend"),
("--scope", "scope"), ("--source", "source"),
("--model", "model"), ("--tasks-file", "tasks_file"),
("--target-skill-path", "target_skill_path"),
]:
val = args.get(key)
if val:
cmd += [flag, str(val)]
# Integer-valued flags
for flag, key in [
("--max-sessions", "max_sessions"), ("--max-tasks", "max_tasks"),
("--lookback-hours", "lookback_hours"), ("--edit-budget", "edit_budget"),
("--hour", "hour"), ("--minute", "minute"),
]:
val = args.get(key)
if val is not None:
cmd += [flag, str(int(val))]
# Boolean flags
for flag, key in [
("--progress", "progress"), ("--auto-adopt", "auto_adopt"), ("--json", "json"),
]:
if args.get(key):
cmd.append(flag)
env = dict(os.environ)
env["PYTHONPATH"] = REPO_ROOT + os.pathsep + env.get("PYTHONPATH", "")
try:
proc = subprocess.run(cmd, cwd=REPO_ROOT, capture_output=True,
text=True, timeout=3600, env=env)
except Exception as e:
return f"[harvest]\n{harvest_out}\n[error] failed to run engine: {e}"
out = (proc.stdout or "").strip()
err = (proc.stderr or "").strip()
result = (f"[harvest]\n{harvest_out}\n\n" if harvest_out else "") + f"[engine]\n{out}"
if err:
result += f"\n[stderr]\n{err}"
if action == "adopt":
result += _sync_skill(args.get("project") or os.getcwd())
return result
def _result(id_, result):
return {"jsonrpc": "2.0", "id": id_, "result": result}
def _error(id_, code, message):
return {"jsonrpc": "2.0", "id": id_, "error": {"code": code, "message": message}}
def handle(req: dict):
method = req.get("method")
id_ = req.get("id")
if method == "initialize":
return _result(id_, {
"protocolVersion": PROTOCOL_VERSION,
"capabilities": {"tools": {}},
"serverInfo": {"name": "skillopt-sleep", "version": "0.1.0"},
})
if method in ("notifications/initialized", "initialized"):
return None
if method == "tools/list":
return _result(id_, {"tools": [
{"name": t["name"], "description": t["description"], "inputSchema": _TOOL_SCHEMA}
for t in TOOLS
]})
if method == "tools/call":
params = req.get("params") or {}
name = params.get("name")
tool = _BY_NAME.get(name)
if not tool:
return _error(id_, -32602, f"unknown tool: {name}")
text = _run_engine(tool["action"], params.get("arguments") or {})
return _result(id_, {"content": [{"type": "text", "text": text}]})
if method == "ping":
return _result(id_, {})
return _error(id_, -32601, f"method not found: {method}")
def main() -> int:
for line in sys.stdin:
line = line.strip()
if not line:
continue
try:
req = json.loads(line)
except Exception:
continue
resp = handle(req)
if resp is not None:
sys.stdout.write(json.dumps(resp) + "\n")
sys.stdout.flush()
return 0
if __name__ == "__main__":
raise SystemExit(main())
+112
View File
@@ -0,0 +1,112 @@
# OpenClaw Plugin for SkillOpt-Sleep
Thin shell for running [SkillOpt-Sleep](https://github.com/microsoft/SkillOpt) on [OpenClaw](https://github.com/openclaw/openclaw).
## What it does
Adds a nightly "sleep cycle" to any OpenClaw agent. The cycle:
1. **Harvests** recent session transcripts from `~/.openclaw/agents/<name>/sessions/*.jsonl`
2. **Mines** recurring task patterns using the optimizer LLM
3. **Replays** each pattern with the current `SKILL.md` (baseline) and a candidate `SKILL.md` (with proposed edits)
4. **Gates** the candidate against the held-out score (rejects regressions)
5. **Stages** the accepted proposal in `~/.skillopt-sleep/staging/<night>/`
6. Leaves adoption to the operator (Ethan)
Nothing live changes until you adopt. Every adopt backs up first.
## Install
The plugin is a thin wrapper around the engine at `~/.openclaw/workspace/SkillOpt/skillopt_sleep/`:
```bash
# 1. Clone the engine (one-time)
cd ~/.openclaw/workspace
git clone https://github.com/microsoft/SkillOpt.git
# 2. Install the OpenClaw skill (this folder)
ln -s /path/to/openclaw ~/.openclaw/workspace/skills/skillopt-sleep
# 3. Configure
cp ~/.openclaw/workspace/skills/skillopt-sleep/config.json ~/.skillopt-sleep/config.json
$EDITOR ~/.skillopt-sleep/config.json
# Set backend = "openclaw-deepseek"
# Set model = "deepseek-v4-pro" (or "deepseek-v4-flash" for budget)
# 4. Set API key
echo 'export DEEPSEEK_API_KEY="sk-..."' >> ~/.openclaw/.env
# 5. Add the nightly cron
(crontab -l 2>/dev/null; echo "0 3 * * * cd ~/.openclaw/workspace/skills/skillopt-sleep && bash run_sleep_cron.sh >> ~/.skillopt-sleep/nightly.log 2>&1") | crontab -
```
## Use
### Manual trigger
```bash
# Run one cycle now
python3 ~/.openclaw/workspace/skills/skillopt-sleep/run_sleep.py
# Dry run (report only)
python3 ~/.openclaw/workspace/skills/skillopt-sleep/run_sleep.py --dry-run
# One category only
python3 ~/.openclaw/workspace/skills/skillopt-sleep/run_sleep.py --tasks tests/research-cron-tasks.json
```
### Slash command
```bash
# In any OpenClaw session
/sleep status
/sleep run
/sleep run research-cron
/sleep dry-run
/sleep adopt # adopt most recent accepted proposal
/sleep reject # discard most recent
/sleep cost
```
## Architecture
```
plugins/openclaw/
├── README.md # this file
├── run_sleep_cron.sh # wrapper for cron invocation
├── run_sleep.py # main entry point
├── slash_sleep.py # /sleep command implementation
├── skillopt_sleep_openclaw.py # DeepSeek + Ollama backend
├── config.json # engine config
├── SKILL.md # OpenClaw skill manifest
└── tests/ # held-out test sets
├── research-cron-tasks.json
├── devops-tasks.json
└── wiki-tasks.json
```
The OpenClaw shell is one engine (skillopt_sleep/) + one backend (DeepSeek/Ollama) + four thin wrappers (cron, slash, skill, tests).
## Why this matters for OpenClaw
OpenClaw currently has no built-in "self-evolving skills" mechanism. The community has:
- **Manual skills** — Ethan writes them
- **LLM-generated skills** — one-shot, no validation
- **Self-revision** — unbounded, no quality bar
SkillOpt-Sleep adds a 4th option: **validated self-evolution**. The skill is the training target, the engine is the optimizer, the gate is the quality bar, the operator is the human-in-the-loop.
## Validation
Validated on the public [gbrain-evals](https://github.com/garrytan/gbrain-evals) `skillopt-v1` benchmark with real Claude and Codex (deficient skills 0.00 → 1.00 on held-out, all 4 seeds).
End-to-end test on our own 14-task held-out set: pipeline runs, gate correctly rejects non-improvements, staging artifacts land in `~/.skillopt-sleep/staging/<night>/`.
## Cost
Measured: ~$0.02/night with `deepseek-v4-pro` at 12 tasks/night. ~$0.59/month, $7.18/year.
## License
MIT (same as SkillOpt core).
+129
View File
@@ -0,0 +1,129 @@
---
name: skillopt-sleep
description: Validate and refine agent skills through nightly sleep cycles with held-out gates. Wraps Microsoft's SkillOpt-Sleep engine for the OpenClaw/DeepSeek stack.
---
# skillopt-sleep — OpenClaw Adaptation of Microsoft SkillOpt-Sleep
A nightly self-improvement loop that reads our session transcripts, mines recurring workflow patterns, replays them with proposed skill edits, and gates the proposals against a held-out test set. Only improvements that beat baseline are staged for human adoption.
## When To Use
- After Hermes's Weekly Skill Review (or as its replacement)
- When a skill is being used 10+ times/week and could be tighter
- Before promoting a new skill from `skill-proposals/` to `skills/`
- When a skill regresses in observed quality
## What It Does (One Cycle)
```
harvest session transcripts -> mine recurring task patterns
-> replay each pattern (current skill vs proposed)
-> GATE: must improve held-out score
-> stage proposal
-> Ethan adopts (manual)
```
Nothing live changes until Ethan adopts. Every adopt backs up first.
## Architecture
```
skills/skillopt-sleep/
├── SKILL.md # this file
├── config.json # engine config (backend, budgets, etc.)
├── run_sleep.py # entry point
└── skillopt_sleep_openclaw.py # DeepSeek/Ollama backend
```
The engine itself is at `~/.openclaw/workspace/SkillOpt/skillopt_sleep/` (cloned from microsoft/SkillOpt).
## Usage
```bash
# Run one cycle with current config
cd ~/.openclaw/workspace/skills/skillopt-sleep
python3 run_sleep.py
# Dry run (report only, no staging)
python3 run_sleep.py --dry-run
# Use a pre-built task set (recommended for testing)
python3 run_sleep.py --tasks tests/research-cron-tasks.json
```
## Scheduling
```bash
python3 slash_sleep.py schedule --hour 3 --minute 17
python3 slash_sleep.py unschedule
python3 slash_sleep.py unschedule --all
```
Installs a nightly cron entry using the shared SkillOpt-Sleep scheduler. This is an alternative to the external `run_sleep_cron.sh` script.
## Alternative backends
While OpenClaw defaults to `openclaw-deepseek` (DeepSeek V4 Pro + Ollama), the shared engine also supports:
- `--backend mock` — deterministic, no API spend (for testing)
- `--backend claude` — uses the Claude CLI
- `--backend codex` — uses the Codex CLI
- `--backend copilot` — uses the GitHub Copilot CLI
These can be used via the engine directly (`python -m skillopt_sleep`).
## Shared-engine flags
When invoking the engine directly, all standard flags are available:
- `--source codex` / `--source auto` — harvest from Codex Desktop sessions
- `--tasks-file PATH` — use a pre-built task set
- `--target-skill-path PATH` — explicit SKILL.md target
- `--max-tasks N` / `--max-sessions N` — cap workload
- `--progress` — print phase progress
- `--json` — machine-readable output
- `--auto-adopt` — auto-adopt if gate passes
Config keys: `preferences`, `gate_mode`, `gate_metric`, `dream_rollouts`, `recall_k`, `evolve_memory`, `evolve_skill`.
## Config (config.json)
Key knobs:
- `backend: "openclaw-deepseek"` — our custom backend
- `model: "deepseek-v4-pro"` — optimizer model
- `edit_budget: 3` — max bounded edits per night
- `gate_mode: "on"` — validation-gated (rejects regressions)
- `auto_adopt: false` — require Ethan to adopt manually
- `max_tasks_per_night: 12` — cap to control cost
## Cost Estimate
Per night: 12 tasks × (1 attempt + 1 judge + 1 reflect) × ~$0.005/1K tokens × ~3K tokens/call ≈ **$0.50-2.00/night**.
## Outputs
- Report: `~/.skillopt-sleep/state.json` (running totals)
- Staging: `~/.skillopt-sleep/staging/<night>/`
- `report.md` — readable summary
- `best_skill.md` — proposed skill
- `edits.json` — bounded edit list
- `before.md` / `after.md` — diffs
## Held-Out Test Sets (Phase 2)
Located at `tests/<category>-tasks.json`. Each task has:
- `prompt` — the recurring task
- `reference` — exact-match gold answer
- `rubric` — soft score rubric (0-1)
- `domain` — research/devops/wiki/etc.
Currently building for 3 categories:
- research-cron-output
- devops-infrastructure-check
- wiki-canonical-guide
## When NOT To Use
- For a one-off workflow (not a recurring pattern)
- During a crisis/incident (humans must lead)
- When session transcripts are < 24h old (not enough signal)
- For skills < 300 tokens (over-optimization risk)
+30
View File
@@ -0,0 +1,30 @@
{
"_comment": "OpenClaw adaptation of skillopt-sleep. Edit and run via run_sleep.py",
"claude_home": "/home/ethanclaw/.openclaw/agents",
"invoked_project": "/home/ethanclaw/.openclaw/workspace",
"projects": "invoked",
"lookback_hours": 168,
"max_tasks_per_night": 12,
"max_tokens_per_night": 800000,
"holdout_fraction": 0.34,
"val_fraction": 0.34,
"test_fraction": 0.0,
"backend": "openclaw-deepseek",
"model": "deepseek-v4-pro",
"gate_mode": "on",
"edit_budget": 3,
"gate_metric": "mixed",
"gate_mixed_weight": 0.5,
"replay_mode": "fresh",
"evolve_memory": true,
"evolve_skill": true,
"llm_mine": false,
"auto_adopt": false,
"managed_skill_name": "skillopt-sleep-learned",
"redact_secrets": true,
"seed": 42
}
+122
View File
@@ -0,0 +1,122 @@
#!/usr/bin/env python3
"""run_sleep.py — OpenClaw entry point for SkillOpt-Sleep.
Runs one nightly sleep cycle:
1. harvest recent session transcripts
2. mine recurring task patterns
3. replay tasks with current skill (baseline) + candidate skill (with proposed edit)
4. gate candidate vs baseline on held-out accuracy
5. stage the proposal in ~/.skillopt-sleep/staging/<night>/
6. leave adoption to Ethan (auto_adopt=false)
Usage:
python3 run_sleep.py # one cycle, default config
python3 run_sleep.py --dry-run # compute report only, no staging
python3 run_sleep.py --tasks path.json # use a pre-built task file
"""
from __future__ import annotations
import argparse
import json
import os
import sys
from pathlib import Path
# Ensure the skillopt_sleep package is importable (it lives in the cloned repo)
REPO = Path("/home/ethanclaw/.openclaw/workspace/SkillOpt")
sys.path.insert(0, str(REPO))
# Register our backend before importing cycle
from skillopt_sleep_openclaw import OpenClawDeepSeekBackend
import skillopt_sleep.backend as _b
_b._BACKENDS = getattr(_b, "_BACKENDS", {})
_b._BACKENDS["openclaw-deepseek"] = OpenClawDeepSeekBackend
# Patch get_backend to know about our backend
_orig_get_backend = _b.get_backend
def get_backend(name, model="", codex_path=""):
if name == "openclaw-deepseek":
return OpenClawDeepSeekBackend(model=model or "deepseek-v4-pro")
return _orig_get_backend(name, model=model, codex_path=codex_path)
_b.get_backend = get_backend
from skillopt_sleep.cycle import run_sleep_cycle
from skillopt_sleep.config import load_config
def main() -> int:
ap = argparse.ArgumentParser(description="OpenClaw SkillOpt-Sleep nightly cycle")
ap.add_argument("--dry-run", action="store_true", help="Compute but don't stage")
ap.add_argument("--config", default="/home/ethanclaw/.openclaw/workspace/skills/skillopt-sleep/config.json")
ap.add_argument("--tasks", default=None, help="Path to pre-built tasks JSON")
ap.add_argument("--verbose", action="store_true")
args = ap.parse_args()
# Load config from file then override with our defaults
overrides = {}
if os.path.exists(args.config):
with open(args.config) as f:
overrides.update(json.load(f))
overrides.pop("_comment", None)
cfg = load_config(**overrides)
seed_tasks = None
if args.tasks:
from skillopt_sleep.types import TaskRecord
with open(args.tasks) as f:
raw = json.load(f)
# Translate our test-set fields → TaskRecord fields
seed_tasks = []
for t in raw:
seed_tasks.append(TaskRecord(
id=t['id'],
project=t.get('project', 'openclaw'),
intent=t.get('intent') or t.get('prompt', ''),
context_excerpt=t.get('context_excerpt', ''),
attempted_solution=t.get('attempted_solution', ''),
outcome=t.get('outcome', 'unknown'),
reference_kind=t.get('reference_kind', 'rubric'),
reference=t.get('reference', ''),
judge=t.get('judge', {}),
tags=t.get('tags', []),
source_sessions=t.get('source_sessions', []),
split=t.get('split', 'train'),
))
print(f"[skillopt-sleep] starting cycle...")
print(f" backend: {cfg.get('backend')}")
print(f" project: {cfg.get('invoked_project')}")
print(f" max tasks: {cfg.get('max_tasks_per_night')}")
print(f" edit budget: {cfg.get('edit_budget')}")
print(f" dry_run: {args.dry_run}")
outcome = run_sleep_cycle(cfg, seed_tasks=seed_tasks, dry_run=args.dry_run)
r = outcome.report
print(f"\n=== Report — night {r.night} ===")
print(f" sessions harvested: {r.n_sessions}")
print(f" tasks mined: {r.n_tasks} (replayed: {r.n_replayed})")
print(f" baseline: {r.baseline_score:.3f} -> candidate: {r.candidate_score:.3f}")
print(f" gate: {r.gate_action} accepted={r.accepted}")
print(f" tokens: {r.tokens_used}")
if r.edits:
print(f" applied edits ({len(r.edits)}):")
for e in r.edits:
print(f" [{e.target}/{e.op}] {e.content[:80]}...")
if r.rejected_edits:
print(f" rejected edits ({len(r.rejected_edits)}) — kept as negative feedback")
if r.notes:
for n in r.notes:
print(f" note: {n}")
if outcome.staging_dir:
print(f"\n STAGED at: {outcome.staging_dir}")
print(f" Review with: ls {outcome.staging_dir}")
return 0 if r.accepted or r.candidate_score >= r.baseline_score else 1
if __name__ == "__main__":
sys.exit(main())
+76
View File
@@ -0,0 +1,76 @@
#!/bin/bash
# run_sleep_cron.sh — wrapper for cron-driven nightly sleep cycle
#
# Usage: bash run_sleep_cron.sh [category1 category2 ...]
# No args: run on all categories in tests/
# With args: run only on listed categories (research-cron, devops, wiki)
#
# Cron (3am MYT daily):
# 0 3 * * * cd /home/ethanclaw/.openclaw/workspace/skills/skillopt-sleep && bash run_sleep_cron.sh >> ~/.skillopt-sleep/nightly.log 2>&1
set -euo pipefail
SKILL_DIR="/home/ethanclaw/.openclaw/workspace/skills/skillopt-sleep"
TESTS_DIR="$SKILL_DIR/tests"
LOG_DIR="$HOME/.skillopt-sleep/logs"
mkdir -p "$LOG_DIR"
TIMESTAMP=$(date +%Y%m%d-%H%M%S)
LOG_FILE="$LOG_DIR/night-$TIMESTAMP.log"
# category → test file map
declare -A CATEGORIES=(
["research-cron"]="research-cron-tasks.json"
["devops"]="devops-tasks.json"
["wiki"]="wiki-tasks.json"
)
# Determine which categories to run
if [ $# -eq 0 ]; then
CATS=("research-cron" "devops" "wiki")
else
CATS=("$@")
fi
{
echo "=========================================="
echo "SkillOpt-Sleep nightly — $TIMESTAMP"
echo "Categories: ${CATS[*]}"
echo "=========================================="
} | tee -a "$LOG_FILE"
# Pre-flight: check DeepSeek API key
if ! grep -q "DEEPSEEK_API_KEY=" "$HOME/.openclaw/.env" 2>/dev/null; then
echo "ERROR: DEEPSEEK_API_KEY not found in ~/.openclaw/.env" | tee -a "$LOG_FILE"
exit 1
fi
EXIT_CODE=0
for cat in "${CATS[@]}"; do
tasks_file="$TESTS_DIR/${CATEGORIES[$cat]:-}"
if [ ! -f "$tasks_file" ]; then
echo "SKIP: $cat (no tasks file: $tasks_file)" | tee -a "$LOG_FILE"
continue
fi
echo "" | tee -a "$LOG_FILE"
echo "--- [$cat] starting cycle ---" | tee -a "$LOG_FILE"
cd "$SKILL_DIR"
if python3 run_sleep.py --tasks "$tasks_file" 2>&1 | tee -a "$LOG_FILE"; then
echo "--- [$cat] OK ---" | tee -a "$LOG_FILE"
else
EC=$?
echo "--- [$cat] FAILED (exit $EC) ---" | tee -a "$LOG_FILE"
EXIT_CODE=$EC
fi
done
{
echo ""
echo "=========================================="
echo "Done. Exit: $EXIT_CODE"
echo "=========================================="
} | tee -a "$LOG_FILE"
exit $EXIT_CODE
+275
View File
@@ -0,0 +1,275 @@
"""OpenClaw backend for SkillOpt-Sleep.
Adapts the skillopt_sleep Backend protocol to our DeepSeek + Ollama stack:
- attempt/judge/reflect -> DeepSeek V4 Pro (or Flash for cost)
- embeddings -> Ollama nomic-embed-text (already configured)
This backend NEVER mutates live state. It only returns text + EditRecord
proposals that the gate stages for human review.
"""
from __future__ import annotations
import json
import os
import re
import subprocess
from typing import Any, Dict, List, Optional, Tuple
from skillopt_sleep.backend import Backend, _normalize, exact_score
from skillopt_sleep.types import EditRecord, ReplayResult, TaskRecord
# ── DeepSeek + Ollama OpenAI-compatible API client (curl-based, no extra deps) ──
def _chat(messages: List[Dict[str, str]], *, model: str, temperature: float = 0.2, max_tokens: int = 1500) -> str:
"""Call DeepSeek V4 Pro via curl + jq. No extra Python deps needed."""
import json as _json
import urllib.request
api_key = os.environ.get("DEEPSEEK_API_KEY", "")
if not api_key:
# try loading from .env
env_path = os.path.expanduser("~/.openclaw/.env")
if os.path.exists(env_path):
with open(env_path) as f:
for line in f:
if line.startswith("DEEPSEEK_API_KEY="):
api_key = line.split("=", 1)[1].strip()
break
base = os.environ.get("DEEPSEEK_BASE_URL", "https://api.deepseek.com/v1")
payload = {
"model": model,
"messages": messages,
"temperature": temperature,
"max_tokens": max_tokens,
"stream": False,
}
req = urllib.request.Request(
f"{base}/chat/completions",
data=_json.dumps(payload).encode("utf-8"),
headers={
"Content-Type": "application/json",
"Authorization": f"Bearer {api_key}",
},
)
try:
with urllib.request.urlopen(req, timeout=180) as resp:
data = _json.loads(resp.read().decode("utf-8"))
return data["choices"][0]["message"]["content"]
except Exception as e:
return f"[BACKEND_ERROR] {type(e).__name__}: {str(e)[:200]}"
def _embed(text: str) -> List[float]:
"""Call Ollama for embeddings. Uses the configured nomic-embed-text model."""
import json as _json
import urllib.request
try:
req = urllib.request.Request(
"http://127.0.0.1:11434/api/embeddings",
data=_json.dumps({"model": "nomic-embed-text:latest", "prompt": text[:2000]}).encode("utf-8"),
headers={"Content-Type": "application/json"},
)
with urllib.request.urlopen(req, timeout=30) as resp:
data = _json.loads(resp.read().decode("utf-8"))
return data.get("embedding", [])
except Exception:
return []
# ── Backend implementation ────────────────────────────────────────────────────
class OpenClawDeepSeekBackend(Backend):
"""Use DeepSeek V4 Pro for attempt/judge/reflect, Ollama for embeddings.
- "model" passed to constructor = optimizer model (default: deepseek-v4-pro)
- "judge_model" = judge model (default: deepseek-v4-pro for quality)
- "cheap_model" = budget-fallback (deepseek-v4-flash)
"""
name = "openclaw-deepseek"
def __init__(
self,
model: str = "deepseek-v4-pro",
judge_model: str = "deepseek-v4-pro",
cheap_model: str = "deepseek-v4-flash",
):
self._model = model
self._judge_model = judge_model
self._cheap_model = cheap_model
self._tokens = 0 # rough estimate
def tokens_used(self) -> int:
return self._tokens
# ── 1. attempt: produce a response given the task + skill + memory ──
def attempt(self, task: TaskRecord, skill: str, memory: str) -> str:
sys = (
"You are an OpenClaw agent (Kobe ecosystem). Use the skill and memory below to complete the task. "
"If the task asks for a structured output, follow the rubric exactly. "
"Be concise. No preamble, no explanation unless the task asks for it."
)
usr = f"""## SKILL
{skill or '(no skill yet)'}
## MEMORY
{memory or '(no memory yet)'}
## TASK
{task.intent}
## CONTEXT (if any)
{task.context_excerpt or '(none)'}
## RESPONSE
"""
out = _chat(
[{"role": "system", "content": sys}, {"role": "user", "content": usr}],
model=self._model,
temperature=0.2,
)
self._tokens += len(usr) // 4 + 200
return out
# ── 2. judge: score the response ──
def judge(self, task: TaskRecord, response: str) -> Tuple[float, float, str]:
# Hard score: exact-match against task.reference (if available)
hard = exact_score(task.reference or "", response)
# Soft score: LLM judge against rubric (reference if reference_kind=='rubric')
rubric_text = task.reference if task.reference_kind == "rubric" else ""
if rubric_text:
judge_prompt = f"""You are a strict grader. Score the response 0.0-1.0 against the rubric.
## TASK
{task.intent}
## REFERENCE
{task.reference or '(none)'}
## RUBRIC
{rubric_text}
## RESPONSE
{response[:3000]}
## INSTRUCTIONS
Return ONLY a single float 0.0-1.0 on one line. No explanation. No markdown.
"""
try:
j_out = _chat(
[{"role": "user", "content": judge_prompt}],
model=self._judge_model,
temperature=0.0,
max_tokens=20,
).strip()
soft = float(re.search(r"[\d.]+", j_out.splitlines()[0]).group())
soft = max(0.0, min(1.0, soft))
except Exception:
soft = hard
self._tokens += 600
else:
soft = hard
rationale = f"hard={hard:.2f} soft={soft:.2f}"
return hard, soft, rationale
# ── 3. reflect: produce bounded EditRecord proposals ──
def reflect(
self,
failures: List[Tuple[TaskRecord, ReplayResult]],
successes: List[Tuple[TaskRecord, ReplayResult]],
skill: str,
memory: str,
*,
edit_budget: int,
evolve_skill: bool,
evolve_memory: bool,
) -> List[EditRecord]:
# Compact digest of failures + successes
fail_digest = "\n".join(
f"- TASK: {t.intent[:200]}\n RESPONSE: {r.response[:300]}\n WHY FAIL: {r.judge_rationale or r.fail_reason or 'unknown'}\n REFERENCE: {t.reference[:200]}"
for t, r in failures[:5]
) or "(none)"
succ_digest = "\n".join(
f"- TASK: {t.intent[:150]} -> OK ({r.judge_rationale or 'high score'})"
for t, r in successes[:3]
) or "(none)"
rubric_text = ""
if failures:
rubric_text = f"\n\n## REFERENCE ANSWERS\n{chr(10).join(f'Q: {t.intent[:120]}\\nA: {t.reference}' for t, _ in failures[:3] if t.reference)}"
sys = (
"You are SkillOpt-Sleep's bounded-edit optimizer. Your job is to propose 1-4 MINIMAL text edits to a skill or memory document "
"that, if applied, would help future agents do better on the failed tasks. "
"NEVER propose adding new sections wholesale. NEVER delete entire sections. "
"Edit primitives: ADD (append a step/rule at end), DELETE (remove a specific line by exact match), REPLACE (swap a specific line for another by exact match). "
"If you cannot identify a clear, minimal improvement, return an empty list."
)
usr = f"""## CURRENT SKILL
{skill or '(empty)'}
## CURRENT MEMORY
{memory or '(empty)'}
## FAILED TASKS
{fail_digest}
## SUCCESSFUL TASKS
{succ_digest}
{rubric_text}
## CONSTRAINTS
- max {edit_budget} edits total
- edits go to {"skill + memory" if (evolve_skill and evolve_memory) else ("skill" if evolve_skill else "memory")}
- if evolve_skill=False, target="memory" only; if evolve_memory=False, target="skill" only
- target must be "skill" or "memory"
## OUTPUT FORMAT (JSON, no markdown)
{{"edits": [{{"op": "ADD"|"DELETE"|"REPLACE", "target": "skill"|"memory", "content": "the text to add or replace with", "old_text": "for REPLACE/DELETE, the exact line to find", "rationale": "one short sentence why"}}]}}
"""
out = _chat(
[{"role": "system", "content": sys}, {"role": "user", "content": usr}],
model=self._model,
temperature=0.4,
max_tokens=2000,
)
self._tokens += len(usr) // 3 + 1500
# parse
try:
# strip markdown fences if any
cleaned = out.strip()
if cleaned.startswith("```"):
cleaned = re.sub(r"^```[a-z]*\n?", "", cleaned)
cleaned = re.sub(r"\n?```$", "", cleaned)
data = json.loads(cleaned)
edits: List[EditRecord] = []
for e in data.get("edits", [])[:edit_budget]:
if e.get("op") not in ("ADD", "DELETE", "REPLACE"):
continue
target = e.get("target", "skill")
if target not in ("skill", "memory"):
continue
if not evolve_skill and target == "skill":
continue
if not evolve_memory and target == "memory":
continue
edits.append(EditRecord(
op=e["op"],
target=target,
content=e.get("content", ""),
old_text=e.get("old_text", ""),
rationale=e.get("rationale", ""),
))
return edits
except Exception as e:
# log + return empty list (no edit is better than a bad edit)
return []
+330
View File
@@ -0,0 +1,330 @@
#!/usr/bin/env python3
"""slash_sleep.py — OpenClaw slash command equivalent of SkillOpt's /sleep.
Use from the main session as a /sleep command:
/sleep status — show current state + last 5 nights
/sleep run — trigger one cycle (all categories) right now
/sleep run research-cron — one cycle, single category
/sleep adopt [night] — adopt the most recent (or specified) staged proposal
/sleep reject [night] — discard the most recent (or specified) staging dir
/sleep dry-run — report-only cycle
/sleep cost — estimate per-night cost for current config
This script is a thin shell over run_sleep.py. It can be invoked either
manually from the main session or by an OpenClaw command handler.
"""
from __future__ import annotations
import argparse
import json
import os
import shutil
import sys
from pathlib import Path
from datetime import datetime
SKILL_DIR = Path("/home/ethanclaw/.openclaw/workspace/skills/skillopt-sleep")
STATE_DIR = Path(os.path.expanduser("~/.skillopt-sleep")) # default
STAGING_ROOT = STATE_DIR
def _resolve_state_dir():
"""Find the actual state dir.
Priority: scan in order:
1. ~/.skillopt-sleep/ (default)
2. /home/ethanclaw/.openclaw/workspace/.skillopt-sleep/ (when staging is there)
3. /home/ethanclaw/.openclaw/.skillopt-sleep/ (parent of overridden claude_home)
Pick the first one that has a state.json OR staging dir.
"""
candidates = [
Path(os.path.expanduser("~/.skillopt-sleep")),
Path("/home/ethanclaw/.openclaw/workspace/.skillopt-sleep"),
Path("/home/ethanclaw/.openclaw/.skillopt-sleep"),
]
# Prefer the one with state.json
for c in candidates:
if (c / "state.json").exists():
return c
# Then the one with staging
for c in candidates:
if (c / "staging").exists():
return c
return candidates[0]
TESTS_DIR = SKILL_DIR / "tests"
def status() -> int:
state_dir = _resolve_state_dir()
state_file = state_dir / "state.json"
staging_dir = state_dir / "staging"
print(f"=== SkillOpt-Sleep status ===")
print(f" state dir: {state_dir}")
print(f" staging dir: {staging_dir}")
if staging_dir.exists():
stages = sorted(staging_dir.iterdir(), key=lambda p: p.stat().st_mtime, reverse=True)
print(f" staging entries: {len(stages)}")
for s in stages[:3]:
print(f" {s.name}")
if not state_file.exists():
print(" no state.json — run a cycle first (state is written at end of each non-dry-run)")
return 0
with open(state_file) as f:
state = json.load(f)
nights = state.get("history") or state.get("nights", [])
print(f" total nights: {len(nights)}")
print(f" accepted: {sum(1 for n in nights if n.get('accepted'))}")
print(f" rejected: {sum(1 for n in nights if not n.get('accepted'))}")
if nights:
last = nights[-1]
print(f" last night: {last.get('night')}")
print(f" accepted: {last.get('accepted')}")
print(f" baseline: {last.get('baseline'):.3f} -> candidate: {last.get('candidate'):.3f}")
print(f" staging: {last.get('staging') or '(none)'}")
return 0
def run_category(category: str, *, dry_run: bool = False) -> int:
cat_to_file = {
"research-cron": "research-cron-tasks.json",
"devops": "devops-tasks.json",
"wiki": "wiki-tasks.json",
}
tasks_file = TESTS_DIR / cat_to_file.get(category, f"{category}-tasks.json")
if not tasks_file.exists():
print(f"ERROR: no tasks file for category '{category}': {tasks_file}")
return 1
cmd = [sys.executable, str(SKILL_DIR / "run_sleep.py")]
if dry_run:
cmd.append("--dry-run")
cmd.extend(["--tasks", str(tasks_file)])
print(f"=== /sleep run {category}{' (dry-run)' if dry_run else ''} ===")
print(f" cmd: {' '.join(cmd)}")
rc = os.system(" ".join(f'"{c}"' for c in cmd))
return rc
def run_all(*, dry_run: bool = False) -> int:
rc = 0
for cat in ("research-cron", "devops", "wiki"):
r = run_category(cat, dry_run=dry_run)
if r != 0:
rc = r
return rc
def adopt(night: str = None) -> int:
state_dir = _resolve_state_dir()
state_file = state_dir / "state.json"
if not state_file.exists():
print("ERROR: no state to adopt from")
return 1
with open(state_file) as f:
state = json.load(f)
nights = state.get("history") or state.get("nights", [])
if not nights:
print("ERROR: no nights recorded")
return 1
target = None
if night:
target = next((n for n in nights if str(n.get("night")) == night), None)
if not target:
print(f"ERROR: night '{night}' not found")
return 1
else:
# most recent accepted
candidates = [n for n in nights if n.get("accepted") and n.get("staging")]
if not candidates:
print("ERROR: no accepted nights with staging to adopt")
return 1
target = candidates[-1]
staging = target["staging"]
if not os.path.isdir(staging):
print(f"ERROR: staging dir missing: {staging}")
return 1
print(f"=== /sleep adopt night {target['night']} ===")
print(f" staging: {staging}")
print(f" baseline: {target.get('baseline'):.3f} candidate: {target.get('candidate'):.3f}")
# Read proposed skill from staging
manifest = Path(staging) / "manifest.json"
if manifest.exists():
with open(manifest) as f:
m = json.load(f)
proposed = m.get("proposed_skill")
if proposed and Path(proposed).exists():
live = STATE_DIR / "live_skill.md"
backup = STATE_DIR / f"live_skill.md.bak-{target['night']}"
if live.exists():
shutil.copy2(live, backup)
print(f" backed up current live skill → {backup}")
shutil.copy2(proposed, live)
print(f" adopted proposed skill → {live}")
print()
print("✅ Adoption complete. Next cycle will use the new skill.")
return 0
print("ERROR: no proposed_skill in manifest")
return 1
def reject(night: str = None) -> int:
state_dir = _resolve_state_dir()
state_file = state_dir / "state.json"
if not state_file.exists():
print("ERROR: no state")
return 1
with open(state_file) as f:
state = json.load(f)
nights = state.get("history") or state.get("nights", [])
target = None
if night:
target = next((n for n in nights if str(n.get("night")) == night), None)
else:
candidates = [n for n in reversed(nights) if n.get("staging")]
target = candidates[0] if candidates else None
if not target or not target.get("staging"):
print("ERROR: nothing to reject")
return 1
staging = target["staging"]
if os.path.isdir(staging):
shutil.rmtree(staging)
print(f"🗑️ Removed staging: {staging}")
# remove from state
state["history"] = [n for n in nights if n.get("night") != target["night"]]
with open(state_file, "w") as f:
json.dump(state, f, indent=2)
print("✅ Rejected. State updated.")
return 0
def schedule_cmd(hour: int, minute: int) -> int:
"""Install a nightly cron entry via the shared SkillOpt-Sleep scheduler.
Note: this schedules the shared engine (``python -m skillopt_sleep run``),
not the OpenClaw-specific ``run_sleep.py``. Use ``run_sleep_cron.sh`` if
you need the OpenClaw-native backend and category task files instead.
"""
try:
from skillopt_sleep.scheduler import schedule
except ImportError:
print("ERROR: skillopt_sleep.scheduler not available — is SkillOpt-Sleep installed?")
return 1
project = str(SKILL_DIR)
ok, msg = schedule(project, hour=hour, minute=minute)
print(msg)
return 0 if ok else 1
def unschedule_cmd(all_projects: bool) -> int:
"""Remove cron entry via the shared SkillOpt-Sleep scheduler."""
try:
from skillopt_sleep.scheduler import unschedule
except ImportError:
print("ERROR: skillopt_sleep.scheduler not available — is SkillOpt-Sleep installed?")
return 1
project = str(SKILL_DIR)
ok, msg = unschedule(project, all_projects=all_projects)
print(msg)
return 0 if ok else 1
def cost() -> int:
"""Estimate per-night cost based on the actual measurement from Phase 2.
From the real dry-run: 5 devops tasks used 14,427 tokens total.
That is ~2,885 tokens per task (all 3 phases combined).
"""
cfg_path = SKILL_DIR / "config.json"
cfg = {}
if cfg_path.exists():
cfg = json.loads(cfg_path.read_text())
cfg.pop("_comment", None)
max_tasks = cfg.get("max_tasks_per_night", 12)
model = cfg.get("model", "deepseek-v4-pro")
# DeepSeek V4 pricing
if "pro" in model:
cost_in = 0.435 # per 1M
cost_out = 0.87
elif "flash" in model:
cost_in = 0.14
cost_out = 0.28
else:
cost_in, cost_out = 0.5, 1.0
# Measured: ~2,900 tokens per task, 30% output / 70% input
toks_per_task = 2900
input_toks = int(toks_per_task * 0.7)
output_toks = int(toks_per_task * 0.3)
cost_in_total = (input_toks * max_tasks / 1_000_000) * cost_in
cost_out_total = (output_toks * max_tasks / 1_000_000) * cost_out
cost = cost_in_total + cost_out_total
print(f"=== Cost estimate (per actual measurement) ===")
print(f" model: {model}")
print(f" max tasks/night: {max_tasks}")
print(f" ~tokens/night: {toks_per_task * max_tasks:,}")
print(f" cost/night: ${cost:.3f}")
print(f" cost/month (30 nights): ${cost*30:.2f}")
print(f" cost/year (365 nights): ${cost*365:.2f}")
return 0
def main():
ap = argparse.ArgumentParser(description="OpenClaw /sleep command")
sub = ap.add_subparsers(dest="cmd", required=True)
sub.add_parser("status", help="show state + last 5 nights")
p_run = sub.add_parser("run", help="trigger one cycle")
p_run.add_argument("category", nargs="?", default=None,
choices=["research-cron", "devops", "wiki", None])
p_run.add_argument("--dry-run", action="store_true")
sub.add_parser("dry-run", help="report-only cycle (all categories)")
p_adopt = sub.add_parser("adopt", help="adopt most recent accepted staging")
p_adopt.add_argument("night", nargs="?", default=None)
p_reject = sub.add_parser("reject", help="discard most recent staging")
p_reject.add_argument("night", nargs="?", default=None)
sub.add_parser("cost", help="estimate cost")
p_schedule = sub.add_parser("schedule", help="install nightly cron entry")
p_schedule.add_argument("--hour", type=int, default=3, help="hour (0-23)")
p_schedule.add_argument("--minute", type=int, default=17, help="minute (0-59)")
p_unschedule = sub.add_parser("unschedule", help="remove cron entry")
p_unschedule.add_argument("--all", dest="all_projects", action="store_true",
help="remove entries for all projects")
args = ap.parse_args()
if args.cmd == "status":
return status()
if args.cmd == "run":
if args.category:
return run_category(args.category, dry_run=args.dry_run)
return run_all(dry_run=args.dry_run)
if args.cmd == "dry-run":
return run_all(dry_run=True)
if args.cmd == "adopt":
return adopt(args.night)
if args.cmd == "reject":
return reject(args.night)
if args.cmd == "cost":
return cost()
if args.cmd == "schedule":
return schedule_cmd(args.hour, args.minute)
if args.cmd == "unschedule":
return unschedule_cmd(args.all_projects)
return 1
if __name__ == "__main__":
sys.exit(main())
+87
View File
@@ -0,0 +1,87 @@
[
{
"id": "do-01",
"reference": "[STATUS] devops-agent | Site Uptime \u2192 geoxylia.com OK (200) | 14/06 22:30 MYT",
"rubric": "Score 1.0 if output matches the exact format [STATUS] devops-agent | Site Uptime \u2192 geoxylia.com OK (200) | DD/MM HH:MM MYT, with a real current time. Score 0.5 if format is close but missing one field. Score 0.0 if wrong format or hallucinated values.",
"project": "devops-infrastructure-check",
"intent": "Site Uptime check. Run: `curl -o /dev/null -s -w '%{http_code}' https://geoxylia.com`. Interpret the result 200, and report in our standard format: 'STATUS | TASK \u2192 RESULT | TIME'. If not 200, escalate.",
"context_excerpt": "",
"attempted_solution": "",
"outcome": "unknown",
"reference_kind": "rubric",
"judge": {},
"tags": [
"devops-infrastructure-check"
],
"source_sessions": [],
"split": "val"
},
{
"id": "do-02",
"reference": "Backup complete. Files: 87, Size: 1.2G, Last: 2026-06-14 22:00:00 MYT",
"rubric": "Score 1.0 if output includes the exact 'Backup complete. Files: N, Size: X, Last: timestamp' structure with plausible values. Score 0.5 if structure is close but one field missing. Score 0.0 if hallucinated or wrong structure.",
"project": "devops-infrastructure-check",
"intent": "Daily Memory Backup. Confirm this ran successfully by checking: `ls -t ~/backups/memory/memory-backup-*.tar.gz | head -3`. Report the file count, total size, and most recent backup time. Use format: 'Backup complete. Files: [N], Size: [X], Last: [timestamp]'.",
"context_excerpt": "",
"attempted_solution": "",
"outcome": "unknown",
"reference_kind": "rubric",
"judge": {},
"tags": [
"devops-infrastructure-check"
],
"source_sessions": [],
"split": "val"
},
{
"id": "do-03",
"reference": "1) Vercel CSP missing frame-ancestors: MEDIUM. Allows clickjacking if anyone embeds our pages; not exploitable for our content, but best-practice gap.\n2) OpenClaw plaintext API keys: LOW. The config is chmod 600, loopback-only, not in git. Standard OpenClaw behavior. Rotating would add zero real security given current exposure.",
"rubric": "Score 1.0 if both are classified correctly (MEDIUM and LOW respectively) and justifications are accurate (not panicky, not dismissive). Score 0.5 if classifications are wrong by one tier or justifications are weak. Score 0.0 if both over-classified as CRITICAL or both wrong.",
"project": "devops-infrastructure-check",
"intent": "Security Check daily run. Two findings: 1) Vercel CSP header missing 'frame-ancestors' directive, 2) OpenClaw config has 3 plaintext API keys. Classify each as: CRITICAL / HIGH / MEDIUM / LOW / INFO. Justify each in 1 sentence.",
"context_excerpt": "",
"attempted_solution": "",
"outcome": "unknown",
"reference_kind": "rubric",
"judge": {},
"tags": [
"devops-infrastructure-check"
],
"source_sessions": [],
"split": "train"
},
{
"id": "do-04",
"reference": "[INCIDENT] supabase.audit_results: anon role has no RLS policy \u2014 anyone with the URL can read all audit results. Fix: add policy 'audit_results_select_own' granting SELECT WHERE user_id = auth.uid(). Severity: HIGH (data exposure). Estimated 2-min fix.",
"rubric": "Score 1.0 if: (a) severity correctly identified as HIGH, (b) fix is a real RLS policy (not just 'enable RLS' since it's already enabled), (c) under 50 words, (d) Telegram-friendly format. Score 0.5 if severity right but fix is generic. Score 0.0 if missing severity or wrong fix.",
"project": "devops-infrastructure-check",
"intent": "Incident Check. The Supabase RLS check returned: 'table public.audit_results: rls enabled but policy missing for anon role'. Interpret severity, propose fix, and format as a Telegram alert (max 50 words).",
"context_excerpt": "",
"attempted_solution": "",
"outcome": "unknown",
"reference_kind": "rubric",
"judge": {},
"tags": [
"devops-infrastructure-check"
],
"source_sessions": [],
"split": "val"
},
{
"id": "do-05",
"reference": "\ud83d\udee1\ufe0f Week security digest:\n\n\u2022 0 critical incidents, 1 high resolved (Supabase RLS policy added)\n\u2022 22 plaintext secrets: expected OpenClaw behavior, no action\n\u2022 1 medium open: Vercel CSP frame-ancestors, schedule for next sprint\n\nTrend: stable. No regressions vs last week.",
"rubric": "Score 1.0 if all 3 priority tiers mentioned with correct counts, ends with a trend statement, Telegram-friendly. Score 0.5 if structure is right but one tier wrong. Score 0.0 if missing a tier or wrong format.",
"project": "devops-infrastructure-check",
"intent": "Weekly security digest. Synthesize this week's findings: 22 plaintext secrets in openclaw.json (expected), 0 critical incidents, 1 high (Supabase RLS), 1 medium (CSP frame-ancestors), 0 low. Output a 3-bullet Telegram status.",
"context_excerpt": "",
"attempted_solution": "",
"outcome": "unknown",
"reference_kind": "rubric",
"judge": {},
"tags": [
"devops-infrastructure-check"
],
"source_sessions": [],
"split": "train"
}
]
@@ -0,0 +1,87 @@
[
{
"id": "rc-01",
"reference": "COMPETITOR MOVES: Otterly adds Perplexity tracker, joining Profound and LLMRefs in multi-platform citations.\nBACKLINK OPPORTUNITIES: 3 SEO directories (G2, Capterra, GetApp) have not been claimed.\nAGENCY BLUEPRINT: Top 2 agency sites bundle GEO audit + content refresh as $3K/mo tier.\nACTION ITEMS: Build Perplexity citation test into GeoXylia audit; claim G2 listing by Friday.",
"rubric": "Score 1.0 if all 4 section headings present in correct order, each with a substantive (not generic) 1-sentence content. Score 0.5 if headings present but content is generic. Score 0.0 if any heading missing or order wrong.",
"project": "research-cron-output",
"intent": "Weekly Competitive Deep Dive for GeoXylia. The competitor otterly.ai just added a Perplexity citation tracker. Produce the report header (top section) in our standard format: COMPETITOR MOVES, BACKLINK OPPORTUNITIES, AGENCY BLUEPRINT, ACTION ITEMS. Keep it to 4 lines, one per section heading with a 1-sentence placeholder.",
"context_excerpt": "",
"attempted_solution": "",
"outcome": "unknown",
"reference_kind": "rubric",
"judge": {},
"tags": [
"research-cron-output"
],
"source_sessions": [],
"split": "train"
},
{
"id": "rc-02",
"reference": "1. 'ai seo audit tool': 420 imp, pos 8.2, on page 1 \u2014 needs CTR lift (snippet/schema).\n2. 'geo audit tool': 230 imp, pos 12.5, page 2 \u2014 target blog post could push to page 1.\n3. 'llm optimization': 85 imp, pos 18.3, deep page-2 \u2014 fresh content with answer capsule could compete.",
"rubric": "Score 1.0 if the response correctly identifies 'ai seo audit tool', 'geo audit tool', and 'llm optimization' as the top 3 (NOT 'best free seo audit' which is already converting well, NOT 'free audit tool' which has too few impressions). Each must have correct impression count, position, and a substantive rationale. Score 0.5 if correct 3 keywords but rationale is weak. Score 0.0 if wrong keywords selected.",
"project": "research-cron-output",
"intent": "GSC keyword opportunity scan. From this snippet of GSC data, identify the top 3 keyword opportunities (high impressions, low CTR, position 5-15):\n\n1. 'ai seo audit tool' \u2014 420 imp, 12 clicks, pos 8.2\n2. 'best free seo audit' \u2014 1100 imp, 95 clicks, pos 4.1\n3. 'geo audit tool' \u2014 230 imp, 4 clicks, pos 12.5\n4. 'llm optimization' \u2014 85 imp, 1 click, pos 18.3\n5. 'free audit tool' \u2014 50 imp, 0 clicks, pos 22.0\n\nOutput: one line per opportunity, format 'KEYWORD: impressions, position, why-it-matters (1 short clause)'.",
"context_excerpt": "",
"attempted_solution": "",
"outcome": "unknown",
"reference_kind": "rubric",
"judge": {},
"tags": [
"research-cron-output"
],
"source_sessions": [],
"split": "train"
},
{
"id": "rc-03",
"reference": "Google AI Overviews now show source links more prominently + author bylines. For GeoXylia: this favors pages with clear authorship (add author schema to blog posts). Action: this week, add author + E-E-A-T schema markup to top 10 blog posts. Source: Google Search Central blog.",
"rubric": "Score 1.0 if: (a) under 60 words, (b) names the change, (c) gives GeoXylia-specific implication, (d) gives a concrete action item, (e) cites the source. Score 0.5 if missing 1-2 of these. Score 0.0 if over 60 words or missing 3+.",
"project": "research-cron-output",
"intent": "Daily Industry News scan. The Google Search Central blog just announced: 'AI Overviews now showing source links more prominently, with author bylines for E-E-A-T-heavy content.' Write a 1-paragraph Telegram alert (max 60 words) for Ethan. Include: 1) what changed, 2) what it means for GeoXylia, 3) any action item.",
"context_excerpt": "",
"attempted_solution": "",
"outcome": "unknown",
"reference_kind": "rubric",
"judge": {},
"tags": [
"research-cron-output"
],
"source_sessions": [],
"split": "val"
},
{
"id": "rc-04",
"reference": "Hi [Name], I saw seo-skill.com's resources page is one of the most-respected SEO learning hubs in the industry \u2014 your 2026 algorithm breakdown was spot-on. We just published a free 2026 AI SEO Audit comparison that your readers would find genuinely useful (no paywall, no signup). It covers the 8 leading AI-audit tools with hands-on screenshots and a clear feature matrix. GeoXylia is the only fully-free option in the comparison, so it's a natural fit for a 'tools to know' section. Mind if I share the link for inclusion?",
"rubric": "Score 1.0 if exactly 4 sentences, all four functional pieces present (compliment / mention resource / audience benefit / GeoXylia one-liner), conversational tone, no aggressive sales language. Score 0.5 if 3 of 4 pieces present or tone is too salesy. Score 0.0 if more than 5 sentences or missing 2+ pieces.",
"project": "research-cron-output",
"intent": "Backlink Outreach draft for the blog post 'Free AI SEO Audit Tool: 2026 Comparison'. The prospect is seo-skill.com (a popular SEO training site with a 'resources' page). Write a 4-sentence outreach email: 1) compliment, 2) mention our resource, 3) explain audience benefit, 4) one-line about GeoXylia.",
"context_excerpt": "",
"attempted_solution": "",
"outcome": "unknown",
"reference_kind": "rubric",
"judge": {},
"tags": [
"research-cron-output"
],
"source_sessions": [],
"split": "train"
},
{
"id": "rc-05",
"reference": "1) DO MORE: AI citation / LLM-mention topics \u2014 the 0.9% CTR at position 9.4 means we're visible but need richer answer capsules to lift CTR. Target 2x posts/week on this cluster.\n2) PAUSE: Pure schema-markup how-tos \u2014 'Schema Markup for SEO' has 0 clicks at position 41, the audience isn't searching this way. Rework as 'How to appear in AI answers' framing.\n3) TEST: 'Perplexity vs ChatGPT citation rates for [niche]' \u2014 unexplored angle, could capture comparison-intent traffic.",
"rubric": "Score 1.0 if all 3 are specific (not generic), cite actual data from the prompt, and contain a clear actionable change. Score 0.5 if 2 of 3 are specific. Score 0.0 if generic advice or no data citations.",
"project": "research-cron-output",
"intent": "Performance \u2192 Strategy feedback loop. Last week's top blog post was 'AI Citation Audit: Does Your Site Appear in ChatGPT?' with 4,200 impressions and 38 clicks (CTR 0.9%, position 9.4). The bottom post was 'Schema Markup for SEO: A 2026 Guide' with 110 impressions and 0 clicks (CTR 0%, position 41). Write 3 specific strategy adjustments: 1) what to do more of, 2) what to pause, 3) what new topic to test.",
"context_excerpt": "",
"attempted_solution": "",
"outcome": "unknown",
"reference_kind": "rubric",
"judge": {},
"tags": [
"research-cron-output"
],
"source_sessions": [],
"split": "val"
}
]
+70
View File
@@ -0,0 +1,70 @@
[
{
"id": "wk-01",
"reference": "1. What GEO is and isn't (define vs SEO/AEO, dispel the 'just add FAQ' myth)\n2. The 3 citation mechanisms LLMs use (RAG, fine-tuning, in-context; weight each)\n3. The 2026 citation data (real statistics from Profound/Otterly/Peec; what % of queries get citations)\n4. The action framework (a 5-step audit-and-fix process, concrete)\n5. Measurement (which metrics actually predict citation lift; vanity vs real)",
"rubric": "Score 1.0 if 5 sections, in a logical order, each with a substantive (not generic) purpose, and the section content is GEO-specific (not generic SEO). Score 0.5 if 5 sections but 1-2 are generic. Score 0.0 if wrong number of sections or wrong order.",
"project": "wiki-canonical-guide",
"intent": "Wiki canonical guide: 'GEO 2026 Standards'. Audience: a mid-level SEO specialist who has heard of GEO but not done it. Tone: technical, evidence-driven, no fluff. Length target: 1500-2200 words. Outline the 5 sections that should appear in order. For each, give a 1-sentence sub-purpose.",
"context_excerpt": "",
"attempted_solution": "",
"outcome": "unknown",
"reference_kind": "rubric",
"judge": {},
"tags": [
"wiki-canonical-guide"
],
"source_sessions": [],
"split": "val"
},
{
"id": "wk-02",
"reference": "Yes, add inbound links. (1) geo-2026-standards.md \u2192 '## Action Framework' section, anchor: 'platform-specific citation rules' \u2014 natural since GEO standards reference ChatGPT/Perplexity behavior. (2) seo-2026-standards.md \u2192 '## AI Overviews' section, anchor: 'AI platform citations' \u2014 links to the mechanism guide. (3) content-strategy.md \u2192 '## Content Types' section, anchor: 'per-platform citation' \u2014 content strategy needs to know which platform favors which content.",
"rubric": "Score 1.0 if all 3 inbound links proposed with specific section + natural anchor text, demonstrating the link solves a real navigational gap (not just SEO-link-building). Score 0.5 if 2 of 3 are well-placed. Score 0.0 if generic anchors like 'click here' or no specific sections named.",
"project": "wiki-canonical-guide",
"intent": "Cross-link audit. The wiki page 'ai-platform-citation-guide.md' has 4 outbound links to other wiki pages, but no inbound links from: 'geo-2026-standards.md', 'seo-2026-standards.md', 'content-strategy.md'. Should we add inbound links? In which page should each inbound link go, and what anchor text would be natural?",
"context_excerpt": "",
"attempted_solution": "",
"outcome": "unknown",
"reference_kind": "rubric",
"judge": {},
"tags": [
"wiki-canonical-guide"
],
"source_sessions": [],
"split": "val"
},
{
"id": "wk-03",
"reference": "Priorities:\n1. Refresh 'geo-glossary.md' (last update 2026-04-12, 63 days) \u2014 add new terms like RAG, in-context citation, agentic SEO.\n2. Refresh 'competitor-pricing.md' (last update 2026-05-01, 44 days) \u2014 Profound raised enterprise tier.\n3. No structural fixes needed.\n\nTelegram: 'Wiki lint: 2 stale pages flagged (geo-glossary 63d, competitor-pricing 44d). No broken links. Both need refresh this week.'",
"rubric": "Score 1.0 if both stale pages correctly identified with specific (not generic) refresh notes, and Telegram summary is under 40 words with the right action. Score 0.5 if stale pages identified but refresh notes are vague. Score 0.0 if missing stale pages or Telegram over 40 words.",
"project": "wiki-canonical-guide",
"intent": "Wiki lint report. Today's scan: 14 wiki pages, 2 with 'Updated' dates > 30 days old ('geo-glossary.md' and 'competitor-pricing.md'), 0 broken internal links, 0 missing YAML frontmatter. Output: 1) prioritized action list, 2) Telegram summary (max 40 words).",
"context_excerpt": "",
"attempted_solution": "",
"outcome": "unknown",
"reference_kind": "rubric",
"judge": {},
"tags": [
"wiki-canonical-guide"
],
"source_sessions": [],
"split": "train"
},
{
"id": "wk-04",
"reference": "Index rebuilt: 14 wiki pages registered in _index.md (was 12 \u2014 added competitor-pricing-rev2 and citations-q2-2026).\nQuestion for Ethan: should 'competitor-pricing.md' and 'competitor-pricing-rev2.md' be merged? They're 78% similar in content.",
"rubric": "Score 1.0 if both sentences are accurate (count matches, names are plausible) and the question identifies a real consolidation opportunity (not a fabricated one). Score 0.5 if structure is right but content vague. Score 0.0 if wrong format or no question.",
"project": "wiki-canonical-guide",
"intent": "Index rebuild check. Run `python3 ~/agent-shared/scripts/update-index.py` (assume it works). After the run, the new wiki/_index.md should list all 14 pages. Generate a 2-sentence confirmation message + 1 question for Ethan to verify.",
"context_excerpt": "",
"attempted_solution": "",
"outcome": "unknown",
"reference_kind": "rubric",
"judge": {},
"tags": [
"wiki-canonical-guide"
],
"source_sessions": [],
"split": "train"
}
]
+11 -6
View File
@@ -30,12 +30,17 @@ if [ -z "${REPO_ROOT:-}" ]; then
fi
PY=""
for cand in python3.12 python3.11 python3.10 python3; do
if command -v "$cand" >/dev/null 2>&1; then
ver="$("$cand" -c 'import sys; print("%d%d" % sys.version_info[:2])' 2>/dev/null || echo 0)"
if [ "${ver:-0}" -ge 310 ]; then PY="$cand"; break; fi
fi
done
# Allow explicit Python override (useful on macOS with old system Python).
if [ -n "${SKILLOPT_SLEEP_PYTHON:-}" ]; then
PY="$SKILLOPT_SLEEP_PYTHON"
else
for cand in python3.12 python3.11 python3.10 python3; do
if command -v "$cand" >/dev/null 2>&1; then
ver="$("$cand" -c 'import sys; print("%d%d" % sys.version_info[:2])' 2>/dev/null || echo 0)"
if [ "${ver:-0}" -ge 310 ]; then PY="$cand"; break; fi
fi
done
fi
if [ -z "$PY" ]; then
echo "[sleep] ERROR: need Python >= 3.10 (found none)." >&2
exit 1
+11 -6
View File
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
[project]
name = "skillopt"
version = "0.1.0"
version = "0.2.0"
description = "SkillOpt: Agentic Skill Optimization via Reflective Training Loops"
readme = "README.md"
license = {text = "MIT"}
@@ -37,9 +37,11 @@ dependencies = [
# Benchmark-specific dependencies
alfworld = ["alfworld>=0.4.0", "gymnasium>=0.29.0"]
# Claude model backend
claude = ["claude-agent-sdk>=0.1.0"]
claude = ["claude-agent-sdk>=0.1.0", "json_repair>=0.61.0"]
# Qwen local model backend (via vLLM)
qwen = ["vllm>=0.4.0"]
qwen = ["vllm>=0.4.0", "json_repair>=0.61.0"]
# SearchQA data materialization
searchqa = ["datasets>=2.18.0"]
# Documentation site
docs = ["mkdocs-material>=9.5.0", "mkdocstrings[python]>=0.24.0"]
# WebUI dashboard
@@ -51,11 +53,13 @@ all = [
"alfworld>=0.4.0",
"gymnasium>=0.29.0",
"claude-agent-sdk>=0.1.0",
"json_repair>=0.61.0",
]
[project.scripts]
skillopt-train = "scripts.train:main"
skillopt-eval = "scripts.eval_only:main"
skillopt-sleep = "skillopt_sleep.__main__:main"
[project.urls]
Homepage = "https://github.com/microsoft/SkillOpt"
@@ -64,9 +68,10 @@ Repository = "https://github.com/microsoft/SkillOpt"
Issues = "https://github.com/microsoft/SkillOpt/issues"
[tool.setuptools.packages.find]
# skillopt* = the research package; skillopt_sleep = the open-source Sleep tool
# (decoupled, zero dependency on the research code).
include = ["skillopt", "skillopt.*", "skillopt_sleep", "skillopt_sleep.*", "scripts*"]
# skillopt* = the research package
# skillopt_sleep = the open-source Sleep tool (decoupled, zero research dep)
# skillopt_webui = the Gradio dashboard (installed via the `webui` extra)
include = ["skillopt", "skillopt.*", "skillopt_sleep", "skillopt_sleep.*", "skillopt_webui", "skillopt_webui.*", "scripts*"]
[tool.ruff]
line-length = 120
+6
View File
@@ -17,6 +17,12 @@ httpx>=0.27.0
# ── Optional: Qwen local model (via vLLM) ────────
# vllm>=0.4.0
# ── Optional: tolerant JSON repair for free-form output from non-OpenAI
# backends (Claude/Qwen). Without it extract_json() falls back safely and
# drops a malformed analyst edit instead of repairing it. Installed by the
# `claude`, `qwen`, and `all` extras in pyproject.toml.
# json_repair>=0.61.0
# ── Optional: WebUI dashboard ────────────────────
# gradio>=4.0.0
+51 -1
View File
@@ -28,6 +28,8 @@ from skillopt.model import (
configure_azure_openai,
configure_claude_code_exec,
configure_codex_exec,
configure_qwen_chat,
configure_minimax_chat,
set_reasoning_effort,
set_target_backend,
set_target_deployment,
@@ -137,7 +139,7 @@ def parse_args() -> argparse.Namespace:
# Legacy flat overrides
p.add_argument("--env", type=str)
p.add_argument("--backend", type=str,
choices=["azure_openai", "codex", "codex_exec", "claude", "claude_chat", "claude_code_exec"])
choices=["azure_openai", "codex", "codex_exec", "claude", "claude_chat", "claude_code_exec", "minimax", "minimax_chat"])
p.add_argument("--optimizer_model", type=str)
p.add_argument("--target_model", type=str)
p.add_argument("--optimizer_backend", type=str)
@@ -179,6 +181,12 @@ def parse_args() -> argparse.Namespace:
p.add_argument("--claude_code_exec_use_sdk", type=str)
p.add_argument("--claude_code_exec_effort", type=str)
p.add_argument("--claude_code_exec_max_thinking_tokens", type=int)
p.add_argument("--minimax_base_url", type=str)
p.add_argument("--minimax_api_key", type=str)
p.add_argument("--minimax_model", type=str)
p.add_argument("--minimax_temperature", type=float)
p.add_argument("--minimax_max_tokens", type=int)
p.add_argument("--minimax_enable_thinking", type=_BOOL)
p.add_argument("--out_root", type=str)
p.add_argument("--data_path", type=str)
p.add_argument("--split_mode", type=str,
@@ -254,6 +262,12 @@ def main() -> None:
"claude_code_exec_use_sdk": "model.claude_code_exec_use_sdk",
"claude_code_exec_effort": "model.claude_code_exec_effort",
"claude_code_exec_max_thinking_tokens": "model.claude_code_exec_max_thinking_tokens",
"minimax_base_url": "model.minimax_base_url",
"minimax_api_key": "model.minimax_api_key",
"minimax_model": "model.minimax_model",
"minimax_temperature": "model.minimax_temperature",
"minimax_max_tokens": "model.minimax_max_tokens",
"minimax_enable_thinking": "model.minimax_enable_thinking",
"seed": "train.seed",
"test_env_num": "evaluation.test_env_num",
"env": "env.name",
@@ -311,6 +325,9 @@ def main() -> None:
elif backend == "claude_code_exec":
cfg.setdefault("optimizer_backend", "openai_chat")
cfg.setdefault("target_backend", "claude_code_exec")
elif backend in {"minimax", "minimax_chat"}:
cfg.setdefault("optimizer_backend", "openai_chat")
cfg.setdefault("target_backend", "minimax_chat")
else:
cfg.setdefault("optimizer_backend", "openai_chat")
cfg.setdefault("target_backend", "openai_chat")
@@ -336,6 +353,15 @@ def main() -> None:
and not _has_model_override("model.target", "target_model")
):
cfg["target_model"] = default_model_for_backend("claude_chat")
if cfg.get("target_backend") == "minimax_chat":
if (
str(cfg.get("target_model", "") or "").strip() in _OPENAI_DEFAULT_MODEL_SENTINELS
and not _has_model_override("model.target", "target_model")
):
cfg["target_model"] = (
cfg.get("minimax_model")
or default_model_for_backend("minimax_chat")
)
if not cfg.get("out_root"):
env = cfg.get("env", "unknown")
@@ -401,6 +427,30 @@ def main() -> None:
effort=cfg.get("claude_code_exec_effort", cfg.get("reasoning_effort", "medium")),
max_thinking_tokens=cfg.get("claude_code_exec_max_thinking_tokens", 16384),
)
configure_qwen_chat(
base_url=cfg.get("qwen_chat_base_url") or None,
api_key=cfg.get("qwen_chat_api_key") or None,
temperature=cfg.get("qwen_chat_temperature"),
timeout_seconds=cfg.get("qwen_chat_timeout_seconds"),
max_tokens=cfg.get("qwen_chat_max_tokens"),
enable_thinking=cfg.get("qwen_chat_enable_thinking"),
target_base_url=cfg.get("target_qwen_chat_base_url") or None,
target_api_key=cfg.get("target_qwen_chat_api_key") or None,
target_temperature=cfg.get("target_qwen_chat_temperature"),
target_timeout_seconds=cfg.get("target_qwen_chat_timeout_seconds"),
target_max_tokens=cfg.get("target_qwen_chat_max_tokens"),
target_enable_thinking=cfg.get("target_qwen_chat_enable_thinking"),
)
configure_minimax_chat(
base_url=cfg.get("minimax_base_url") or None,
api_key=cfg.get("minimax_api_key") or None,
temperature=cfg.get("minimax_temperature"),
max_tokens=cfg.get("minimax_max_tokens"),
enable_thinking=cfg.get("minimax_enable_thinking"),
)
minimax_model_cfg = cfg.get("minimax_model")
if minimax_model_cfg and cfg.get("target_backend") == "minimax_chat":
set_target_deployment(str(minimax_model_cfg))
set_reasoning_effort(cfg.get("reasoning_effort", "") or None)
# Build adapter
+148
View File
@@ -0,0 +1,148 @@
"""Materialize runnable SearchQA splits from the released ID manifest."""
from __future__ import annotations
import argparse
import json
from collections.abc import Iterable, Mapping
from pathlib import Path
PROJECT_ROOT = Path(__file__).resolve().parent.parent
SPLITS = ("train", "val", "test")
REQUIRED_FIELDS = ("question", "context", "answers")
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument(
"--manifest-dir",
type=Path,
default=PROJECT_ROOT / "data" / "searchqa_id_split",
help="Directory containing train/val/test ID manifests.",
)
parser.add_argument(
"--output-dir",
type=Path,
default=PROJECT_ROOT / "data" / "searchqa_split",
help="Directory to write runnable train/val/test splits.",
)
parser.add_argument(
"--dataset",
default="lucadiliello/searchqa",
help="Hugging Face dataset repository to load.",
)
return parser.parse_args()
def load_manifest_ids(manifest_dir: Path) -> dict[str, list[str]]:
split_ids = {}
for split in SPLITS:
path = manifest_dir / split / "items.json"
with path.open(encoding="utf-8") as file:
items = json.load(file)
split_ids[split] = [str(item["id"]) for item in items]
return split_ids
def _iter_dataset_rows(dataset: Mapping[str, Iterable[dict]]) -> Iterable[dict]:
for source_split in dataset.values():
yield from source_split
def _normalize_row(row: dict) -> dict:
try:
key = str(row["key"])
except KeyError as exc:
raise ValueError("SearchQA source row is missing required field: key") from exc
missing = [field for field in REQUIRED_FIELDS if field not in row]
if missing:
raise ValueError(f"SearchQA source row {key!r} is missing required fields: {', '.join(missing)}")
return {
"id": key,
"question": row["question"],
"context": row["context"],
"answers": row["answers"],
}
def materialize_searchqa_splits(
manifest_dir: Path,
output_dir: Path,
dataset: Mapping[str, Iterable[dict]],
*,
dataset_name: str,
) -> dict[str, int]:
"""Write runnable SearchQA train/val/test splits from a source dataset."""
manifest_dir = manifest_dir.resolve()
output_dir = output_dir.resolve()
split_ids = load_manifest_ids(manifest_dir)
wanted_ids = {item_id for ids in split_ids.values() for item_id in ids}
selected: dict[str, dict] = {}
duplicate_ids: set[str] = set()
for row in _iter_dataset_rows(dataset):
key = str(row.get("key", ""))
if key not in wanted_ids:
continue
if key in selected:
duplicate_ids.add(key)
continue
selected[key] = _normalize_row(row)
if duplicate_ids:
preview = ", ".join(sorted(duplicate_ids)[:5])
raise ValueError(f"SearchQA source dataset contains duplicate manifest IDs. First IDs: {preview}")
missing = sorted(wanted_ids - selected.keys())
if missing:
preview = ", ".join(missing[:5])
raise RuntimeError(f"SearchQA source dataset is missing {len(missing)} manifest IDs. First IDs: {preview}")
counts = {}
for split, ids in split_ids.items():
items = [selected[item_id] for item_id in ids]
split_dir = output_dir / split
split_dir.mkdir(parents=True, exist_ok=True)
with (split_dir / "items.json").open("w", encoding="utf-8") as file:
json.dump(items, file, ensure_ascii=False, indent=2)
counts[split] = len(items)
manifest = {
"source_manifest_dir": str(manifest_dir),
"source_dataset": dataset_name,
"counts": counts,
"item_fields": ["id", *REQUIRED_FIELDS],
}
with (output_dir / "split_manifest.json").open("w", encoding="utf-8") as file:
json.dump(manifest, file, ensure_ascii=False, indent=2)
return counts
def main() -> None:
args = parse_args()
try:
from datasets import load_dataset
except ImportError as exc:
raise SystemExit(
"Missing dependency 'datasets'. Install it with:\n"
" python -m pip install 'skillopt[searchqa]'\n"
"or:\n"
" python -m pip install datasets"
) from exc
print(f"Loading {args.dataset}...")
dataset = load_dataset(args.dataset)
counts = materialize_searchqa_splits(
args.manifest_dir,
args.output_dir,
dataset,
dataset_name=args.dataset,
)
print(f"Wrote SearchQA splits to {args.output_dir.resolve()}: {counts}")
if __name__ == "__main__":
main()
+5 -8
View File
@@ -2416,14 +2416,11 @@
<div class="bibtex-box">
<button class="copy-btn" type="button" onclick="copyBibtex(this)">Copy</button>
<pre><code>@misc{yang2026skilloptexecutivestrategyselfevolving,
title={SkillOpt: Executive Strategy for Self-Evolving Agent Skills},
author={Yifan Yang and Ziyang Gong and Weiquan Huang and Qihao Yang and Ziwei Zhou and Zisu Huang and Yan Li and Xuemei Gao and Qi Dai and Bei Liu and Kai Qiu and Yuqing Yang and Dongdong Chen and Xue Yang and Chong Luo},
year={2026},
eprint={2605.23904},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2605.23904},
<pre><code>@article{yang2026skillopt,
title={Skillopt: Executive strategy for self-evolving agent skills},
author={Yang, Yifan and Gong, Ziyang and Huang, Weiquan and Yang, Qihao and Zhou, Ziwei and Huang, Zisu and Li, Yan and Gao, Xuemei and Dai, Qi and Liu, Bei and others},
journal={arXiv preprint arXiv:2605.23904},
year={2026}
}</code></pre>
</div>
</section>
+1 -1
View File
@@ -12,7 +12,7 @@ Pipeline stages:
6. Evaluate validate candidate skill, accept/reject
"""
__version__ = "0.1.0"
__version__ = "0.2.0"
from skillopt.types import ( # noqa: F401
BatchSpec,
+4 -1
View File
@@ -125,6 +125,9 @@ _FLATTEN_MAP: dict[str, str] = {
"evaluation.use_gate": "use_gate",
"evaluation.gate_metric": "gate_metric",
"evaluation.gate_mixed_weight": "gate_mixed_weight",
"evaluation.use_semantic_density": "use_semantic_density",
"evaluation.semantic_density_weight": "semantic_density_weight",
"evaluation.leading_words": "leading_words",
"evaluation.sel_env_num": "sel_env_num",
"evaluation.test_env_num": "test_env_num",
"evaluation.eval_test": "eval_test",
@@ -158,7 +161,7 @@ def _load_yaml(path: str, _visited: set[str] | None = None) -> dict:
raise ValueError(f"Circular _base_ inheritance: {abs_path}")
_visited.add(abs_path)
with open(abs_path) as f:
with open(abs_path, encoding="utf-8") as f:
cfg = yaml.safe_load(f) or {}
base_ref = cfg.pop("_base_", None)
+50 -20
View File
@@ -709,26 +709,26 @@ class ReflACTTrainer:
effort=cfg.get("claude_code_exec_effort", cfg.get("reasoning_effort", "medium")),
max_thinking_tokens=cfg.get("claude_code_exec_max_thinking_tokens", 16384),
)
configure_qwen_chat(
base_url=cfg.get("qwen_chat_base_url") or None,
api_key=cfg.get("qwen_chat_api_key") or None,
temperature=cfg.get("qwen_chat_temperature"),
timeout_seconds=cfg.get("qwen_chat_timeout_seconds"),
max_tokens=cfg.get("qwen_chat_max_tokens"),
enable_thinking=cfg.get("qwen_chat_enable_thinking"),
optimizer_base_url=cfg.get("optimizer_qwen_chat_base_url") or None,
optimizer_api_key=cfg.get("optimizer_qwen_chat_api_key") or None,
optimizer_temperature=cfg.get("optimizer_qwen_chat_temperature"),
optimizer_timeout_seconds=cfg.get("optimizer_qwen_chat_timeout_seconds"),
optimizer_max_tokens=cfg.get("optimizer_qwen_chat_max_tokens"),
optimizer_enable_thinking=cfg.get("optimizer_qwen_chat_enable_thinking"),
target_base_url=cfg.get("target_qwen_chat_base_url") or None,
target_api_key=cfg.get("target_qwen_chat_api_key") or None,
target_temperature=cfg.get("target_qwen_chat_temperature"),
target_timeout_seconds=cfg.get("target_qwen_chat_timeout_seconds"),
target_max_tokens=cfg.get("target_qwen_chat_max_tokens"),
target_enable_thinking=cfg.get("target_qwen_chat_enable_thinking"),
)
configure_qwen_chat(
base_url=cfg.get("qwen_chat_base_url") or None,
api_key=cfg.get("qwen_chat_api_key") or None,
temperature=cfg.get("qwen_chat_temperature"),
timeout_seconds=cfg.get("qwen_chat_timeout_seconds"),
max_tokens=cfg.get("qwen_chat_max_tokens"),
enable_thinking=cfg.get("qwen_chat_enable_thinking"),
optimizer_base_url=cfg.get("optimizer_qwen_chat_base_url") or None,
optimizer_api_key=cfg.get("optimizer_qwen_chat_api_key") or None,
optimizer_temperature=cfg.get("optimizer_qwen_chat_temperature"),
optimizer_timeout_seconds=cfg.get("optimizer_qwen_chat_timeout_seconds"),
optimizer_max_tokens=cfg.get("optimizer_qwen_chat_max_tokens"),
optimizer_enable_thinking=cfg.get("optimizer_qwen_chat_enable_thinking"),
target_base_url=cfg.get("target_qwen_chat_base_url") or None,
target_api_key=cfg.get("target_qwen_chat_api_key") or None,
target_temperature=cfg.get("target_qwen_chat_temperature"),
target_timeout_seconds=cfg.get("target_qwen_chat_timeout_seconds"),
target_max_tokens=cfg.get("target_qwen_chat_max_tokens"),
target_enable_thinking=cfg.get("target_qwen_chat_enable_thinking"),
)
configure_minimax_chat(
base_url=cfg.get("minimax_base_url") or None,
api_key=cfg.get("minimax_api_key") or None,
@@ -964,6 +964,15 @@ class ReflACTTrainer:
f"got {gate_metric!r}"
)
gate_mixed_weight = float(cfg.get("gate_mixed_weight", 0.5))
use_semantic_density = bool(cfg.get("use_semantic_density", False))
semantic_density_weight = float(cfg.get("semantic_density_weight", 0.05))
leading_words_raw = cfg.get("leading_words", None)
leading_words = None
if leading_words_raw is not None:
if isinstance(leading_words_raw, str):
leading_words = [w.strip() for w in leading_words_raw.split(",") if w.strip()]
else:
leading_words = list(leading_words_raw)
if not 0.0 <= gate_mixed_weight <= 1.0:
raise ValueError(
f"evaluation.gate_mixed_weight must be in [0, 1], "
@@ -1003,6 +1012,10 @@ class ReflACTTrainer:
baseline_hard, baseline_soft = compute_score(baseline_results)
current_score = select_gate_score(
baseline_hard, baseline_soft, gate_metric, gate_mixed_weight,
skill_content=skill_init,
use_semantic_density=use_semantic_density,
semantic_density_weight=semantic_density_weight,
leading_words=leading_words,
)
best_score = current_score
sh = skill_hash(skill_init)
@@ -1450,9 +1463,16 @@ class ReflACTTrainer:
cand_soft=cand_soft,
metric=gate_metric,
mixed_weight=gate_mixed_weight,
use_semantic_density=use_semantic_density,
semantic_density_weight=semantic_density_weight,
leading_words=leading_words,
) if use_gate else None
cand_gate_score = select_gate_score(
cand_hard, cand_soft, gate_metric, gate_mixed_weight,
skill_content=candidate_skill,
use_semantic_density=use_semantic_density,
semantic_density_weight=semantic_density_weight,
leading_words=leading_words,
)
if not use_gate:
# Validation ran (scores recorded above) but the gate is
@@ -1844,6 +1864,9 @@ class ReflACTTrainer:
cand_soft=slow_sel_soft,
metric=gate_metric,
mixed_weight=gate_mixed_weight,
use_semantic_density=use_semantic_density,
semantic_density_weight=semantic_density_weight,
leading_words=leading_words,
)
slow_result["selection_hard"] = slow_sel_hard
slow_result["selection_soft"] = slow_sel_soft
@@ -2093,6 +2116,10 @@ class ReflACTTrainer:
final_gate_score = select_gate_score(
final_selection_hard, final_selection_soft,
gate_metric, gate_mixed_weight,
skill_content=current_skill,
use_semantic_density=use_semantic_density,
semantic_density_weight=semantic_density_weight,
leading_words=leading_words,
)
print(
f"\n [final skill val] items={fval_n} "
@@ -2133,6 +2160,7 @@ class ReflACTTrainer:
)
print(f" Test items: {test_n}")
baseline_test_dir = os.path.join(out_root, "test_eval_baseline")
os.makedirs(baseline_test_dir, exist_ok=True)
baseline_test_results = adapter.rollout(test_env, skill_init, baseline_test_dir)
baseline_test_hard, baseline_test_soft = compute_score(baseline_test_results)
baseline_buckets = _compute_task_type_buckets(baseline_test_results, task_types)
@@ -2167,6 +2195,7 @@ class ReflACTTrainer:
)
print(f" Test items: {test_n2}")
test_dir = os.path.join(out_root, "test_eval")
os.makedirs(test_dir, exist_ok=True)
test_results = adapter.rollout(test_env2, best_skill, test_dir)
test_hard, test_soft = compute_score(test_results)
best_buckets = _compute_task_type_buckets(test_results, task_types)
@@ -2230,6 +2259,7 @@ class ReflACTTrainer:
)
print(f" Test items: {test_n3}")
final_test_dir = os.path.join(out_root, "test_eval_final")
os.makedirs(final_test_dir, exist_ok=True)
final_test_results = adapter.rollout(test_env3, current_skill, final_test_dir)
final_test_hard, final_test_soft = compute_score(final_test_results)
final_buckets = _compute_task_type_buckets(final_test_results, task_types)
+25 -6
View File
@@ -7,12 +7,10 @@ Provides:
"""
from __future__ import annotations
import concurrent.futures
import json
import os
import re
import sys
import concurrent.futures
import numpy as np
from skillopt.model import chat_target
@@ -65,6 +63,25 @@ def _append_diagnostic_instruction(prompt: str, diagnostic_instruction: str) ->
return f"{prompt}\n\n## Training Readout\n{diagnostic_instruction.strip()}\n"
def _resolve_alfworld_gamefile(gamefile: str) -> str:
path = os.path.expanduser(os.path.expandvars(str(gamefile)))
if os.path.isabs(path):
return path
data_root = os.environ.get("ALFWORLD_DATA", "").strip()
if not data_root:
return path
root = os.path.expanduser(os.path.expandvars(data_root))
return os.path.abspath(os.path.join(root, path))
def _resolve_alfworld_gamefiles(gamefiles: list[str] | None) -> list[str] | None:
if gamefiles is None:
return None
return [_resolve_alfworld_gamefile(gamefile) for gamefile in gamefiles]
# ── Environment builder ──────────────────────────────────────────────────────
@@ -86,9 +103,10 @@ def build_alfworld_env(
Returns:
env_manager: AlfWorldEnvironmentManager instance
"""
from omegaconf import OmegaConf
from functools import partial
from omegaconf import OmegaConf
from skillopt.envs.alfworld.vendor.alfworld_envs import build_alfworld_envs
from skillopt.envs.alfworld.vendor.alfworld_projection import alfworld_projection
from skillopt.envs.alfworld.vendor.env_manager import AlfWorldEnvironmentManager
@@ -97,6 +115,7 @@ def build_alfworld_env(
alf_config_path = os.path.join(HERE, "vendor", "config_tw.yaml")
env_kwargs = {"eval_dataset": eval_dataset}
resolved_gamefiles = _resolve_alfworld_gamefiles(specific_gamefiles)
envs = build_alfworld_envs(
alf_config_path,
@@ -106,7 +125,7 @@ def build_alfworld_env(
is_train=is_train,
env_kwargs=env_kwargs,
resources_per_worker=None,
gamefiles=specific_gamefiles,
gamefiles=resolved_gamefiles,
)
config = OmegaConf.create(
@@ -222,7 +241,7 @@ def run_alfworld_batch(
if _extract_action(response) is None:
return idx, "<think>missing action tag</think><action>look</action>"
return idx, response
except Exception as e:
except Exception:
return idx, "<think>error</think><action>look</action>"
executor = concurrent.futures.ThreadPoolExecutor(max_workers=max_api_workers)
+17 -4
View File
@@ -13,20 +13,31 @@ from __future__ import annotations
import json
import os
import time
import traceback
from collections import Counter
from concurrent.futures import FIRST_COMPLETED, ThreadPoolExecutor, wait
from skillopt.model import chat_target, get_target_backend, is_target_exec_backend
from skillopt.envs.searchqa.evaluator import evaluate
from skillopt.model import chat_target, is_target_exec_backend
from skillopt.model.codex_harness import prepare_workspace, render_skill_md, run_target_exec
from skillopt.prompts import load_prompt
from skillopt.envs.searchqa.evaluator import evaluate
# ── Prompt templates ─────────────────────────────────────────────────────────
_MAX_CONTEXT_CHARS = 6000
def _raise_on_systemic_failure(results: list[dict]) -> None:
"""Abort when all rollout rows failed before any agent response."""
if not results or not all(row.get("agent_ok") is False for row in results):
return
reasons = Counter(str(row.get("fail_reason") or "unknown error") for row in results)
common_reason, count = reasons.most_common(1)[0]
raise RuntimeError(
f"SearchQA rollout failed for all {len(results)} items before an agent "
f"response ({count}x): {common_reason}"
)
def _truncate_context(context: str, max_chars: int = _MAX_CONTEXT_CHARS) -> str:
"""Truncate context at [DOC] boundaries to stay within budget."""
if len(context) <= max_chars:
@@ -379,6 +390,7 @@ def run_batch(
pending = [it for it in items if str(it["id"]) not in done_ids]
if not pending:
_raise_on_systemic_failure(existing)
return existing
total = len(existing) + len(pending)
@@ -478,4 +490,5 @@ def run_batch(
finally:
ex.shutdown(wait=False, cancel_futures=True)
_raise_on_systemic_failure(results)
return results
+86 -9
View File
@@ -43,11 +43,55 @@ class GateResult:
best_step: int
def compute_semantic_density(
skill_content: str,
leading_words: list[str] | None = None,
) -> float:
"""Compute the semantic density of leading words in a skill document."""
if not skill_content or not skill_content.strip():
return 0.0
if leading_words is None:
leading_words = [
"MUST", "ALWAYS", "NEVER", "ONLY", "CRITICAL", "IMPORTANT",
"RESOLVE", "PREFER", "ENSURE", "STRICT", "VERIFY"
]
# Strip metadata comments to focus purely on instruction text
skill = skill_content
for start, end in [
("<!-- SLOW_UPDATE_START -->", "<!-- SLOW_UPDATE_END -->"),
("<!-- APPENDIX_START -->", "<!-- APPENDIX_END -->")
]:
while True:
s_idx = skill.find(start)
if s_idx == -1:
break
e_idx = skill.find(end, s_idx)
if e_idx == -1:
skill = skill[:s_idx] + skill[s_idx + len(start):]
break
skill = skill[:s_idx] + skill[e_idx + len(end):]
import re
words = re.findall(r'[a-zA-Z0-9]+', skill.lower())
if not words:
return 0.0
leading_set = {w.lower() for w in leading_words}
leading_count = sum(1 for w in words if w in leading_set)
return leading_count / len(words)
def select_gate_score(
hard: float,
soft: float,
metric: GateMetric = "hard",
mixed_weight: float = 0.5,
*,
skill_content: str = "",
use_semantic_density: bool = False,
semantic_density_weight: float = 0.05,
leading_words: list[str] | None = None,
) -> float:
"""Project (hard, soft) onto a single comparison metric.
@@ -60,17 +104,32 @@ def select_gate_score(
mixed_weight
For ``"mixed"``: weight given to ``soft``. Must be in ``[0, 1]``.
Ignored for ``"hard"`` / ``"soft"``.
skill_content
The raw skill document content.
use_semantic_density
Whether to adjust the score based on semantic density of leading words.
semantic_density_weight
Scaling weight for the semantic density bonus.
leading_words
Optional custom list of high-influence words to prioritize.
"""
if metric == "hard":
return float(hard)
if metric == "soft":
return float(soft)
if metric == "mixed":
score = float(hard)
elif metric == "soft":
score = float(soft)
elif metric == "mixed":
w = max(0.0, min(1.0, float(mixed_weight)))
return (1.0 - w) * float(hard) + w * float(soft)
raise ValueError(
f"unknown gate metric {metric!r}; expected 'hard', 'soft', or 'mixed'"
)
score = (1.0 - w) * float(hard) + w * float(soft)
else:
raise ValueError(
f"unknown gate metric {metric!r}; expected 'hard', 'soft', or 'mixed'"
)
if use_semantic_density:
density = compute_semantic_density(skill_content, leading_words)
score += float(semantic_density_weight) * density
return score
def evaluate_gate(
@@ -86,6 +145,9 @@ def evaluate_gate(
cand_soft: float = 0.0,
metric: GateMetric = "hard",
mixed_weight: float = 0.5,
use_semantic_density: bool = False,
semantic_density_weight: float = 0.05,
leading_words: list[str] | None = None,
) -> GateResult:
"""Pure gate decision: compare candidate score to current/best.
@@ -111,6 +173,12 @@ def evaluate_gate(
the original gate behavior.
mixed_weight
Weight on ``soft`` when ``metric == "mixed"``.
use_semantic_density
Whether to adjust the score based on semantic density of leading words.
semantic_density_weight
Scaling weight for the semantic density bonus.
leading_words
Optional custom list of high-influence words to prioritize.
Returns
-------
@@ -118,7 +186,16 @@ def evaluate_gate(
Updated state; the caller decides what to do with it (print,
mutate trainer state, log, etc.).
"""
cand_score = select_gate_score(cand_hard, cand_soft, metric, mixed_weight)
cand_score = select_gate_score(
cand_hard,
cand_soft,
metric,
mixed_weight,
skill_content=candidate_skill,
use_semantic_density=use_semantic_density,
semantic_density_weight=semantic_density_weight,
leading_words=leading_words,
)
if cand_score > current_score:
if cand_score > best_score:
+32 -3
View File
@@ -13,13 +13,13 @@ from skillopt.model.backend_config import ( # noqa: F401
configure_codex_exec,
get_claude_code_exec_config,
get_codex_exec_config,
get_target_backend,
get_optimizer_backend,
get_target_backend,
is_optimizer_chat_backend,
is_target_chat_backend,
is_target_exec_backend,
is_optimizer_chat_backend,
set_target_backend,
set_optimizer_backend,
set_target_backend,
)
@@ -105,6 +105,16 @@ def chat_optimizer(
reasoning_effort=reasoning_effort,
timeout=timeout,
)
if get_optimizer_backend() == "minimax_chat":
return _minimax.chat_optimizer(
system=system,
user=user,
max_completion_tokens=max_completion_tokens,
retries=retries,
stage=stage,
reasoning_effort=reasoning_effort,
timeout=timeout,
)
return _openai.chat_optimizer(
system=system,
user=user,
@@ -204,6 +214,18 @@ def chat_optimizer_messages(
return_message=return_message,
timeout=timeout,
)
if get_optimizer_backend() == "minimax_chat":
return _minimax.chat_optimizer_messages(
messages=messages,
max_completion_tokens=max_completion_tokens,
retries=retries,
stage=stage,
reasoning_effort=reasoning_effort,
tools=tools,
tool_choice=tool_choice,
return_message=return_message,
timeout=timeout,
)
return _openai.chat_optimizer_messages(
messages=messages,
max_completion_tokens=max_completion_tokens,
@@ -440,18 +462,21 @@ def configure_qwen_chat(
timeout_seconds: float | str | None = None,
max_tokens: int | str | None = None,
enable_thinking: bool | str | None = None,
use_max_completion_tokens: bool | str | None = None,
optimizer_base_url: str | None = None,
optimizer_api_key: str | None = None,
optimizer_temperature: float | str | None = None,
optimizer_timeout_seconds: float | str | None = None,
optimizer_max_tokens: int | str | None = None,
optimizer_enable_thinking: bool | str | None = None,
optimizer_use_max_completion_tokens: bool | str | None = None,
target_base_url: str | None = None,
target_api_key: str | None = None,
target_temperature: float | str | None = None,
target_timeout_seconds: float | str | None = None,
target_max_tokens: int | str | None = None,
target_enable_thinking: bool | str | None = None,
target_use_max_completion_tokens: bool | str | None = None,
) -> None:
_qwen.configure_qwen_chat(
base_url=base_url,
@@ -460,18 +485,21 @@ def configure_qwen_chat(
timeout_seconds=timeout_seconds,
max_tokens=max_tokens,
enable_thinking=enable_thinking,
use_max_completion_tokens=use_max_completion_tokens,
optimizer_base_url=optimizer_base_url,
optimizer_api_key=optimizer_api_key,
optimizer_temperature=optimizer_temperature,
optimizer_timeout_seconds=optimizer_timeout_seconds,
optimizer_max_tokens=optimizer_max_tokens,
optimizer_enable_thinking=optimizer_enable_thinking,
optimizer_use_max_completion_tokens=optimizer_use_max_completion_tokens,
target_base_url=target_base_url,
target_api_key=target_api_key,
target_temperature=target_temperature,
target_timeout_seconds=target_timeout_seconds,
target_max_tokens=target_max_tokens,
target_enable_thinking=target_enable_thinking,
target_use_max_completion_tokens=target_use_max_completion_tokens,
)
@@ -512,3 +540,4 @@ def set_optimizer_deployment(deployment: str) -> None:
_openai.set_optimizer_deployment(deployment)
_claude.set_optimizer_deployment(deployment)
_qwen.set_optimizer_deployment(deployment)
_minimax.set_optimizer_deployment(deployment)
+14 -2
View File
@@ -252,13 +252,25 @@ def _run_claude_print(*, system: str, prompt: str, model: str, tools: list[dict[
if CLAUDE_SETTING_SOURCES:
cmd.extend(["--setting-sources", CLAUDE_SETTING_SOURCES])
if system:
cmd.extend(["--append-system-prompt", system])
# Write the system prompt to a file, not argv: here the skill being
# optimized IS the system prompt, and SkillOpt grows it over training,
# so past ~30 KB it would re-hit the Windows argv cap (WinError 206).
# The CLI reads it via --append-system-prompt-file.
system_path = os.path.join(temp_dir, "system_prompt.txt")
with open(system_path, "w", encoding="utf-8") as system_fh:
system_fh.write(system)
cmd.extend(["--append-system-prompt-file", system_path])
if effort:
cmd.extend(["--effort", effort])
structured_output = bool(return_message)
if structured_output:
cmd.extend(["--schema", _assistant_message_schema_wrapper()])
proc = subprocess.run(cmd + [prompt_for_cli], capture_output=True, text=True, timeout=timeout or 300, cwd=temp_dir)
# Feed the prompt via stdin (and the system prompt via a file, above), not
# argv: on Windows the whole command line is capped at ~32 KB and large
# optimizer prompts / grown skills overflow it → [WinError 206]. Pin UTF-8
# so a zh-CN default codepage (cp936) can't raise UnicodeEncodeError on
# emoji / non-GBK glyphs before the CLI even starts.
proc = subprocess.run(cmd, input=prompt_for_cli, capture_output=True, text=True, encoding="utf-8", errors="replace", timeout=timeout or 300, cwd=temp_dir)
stderr_text = (proc.stderr or "").strip()
if proc.returncode != 0:
_check_claude_error(stderr_text, model)
+2
View File
@@ -328,6 +328,8 @@ def _run_codex_exec(
command,
input=prompt,
text=True,
encoding="utf-8",
errors="replace",
capture_output=True,
timeout=timeout,
check=False,
+66 -1
View File
@@ -36,6 +36,9 @@ TARGET_DEPLOYMENT = os.environ.get(
"TARGET_DEPLOYMENT",
default_model_for_backend("minimax_chat"),
)
# Optimizer role can point at a different MiniMax model than the target; falls
# back to TARGET_DEPLOYMENT when unset so single-model setups are unaffected.
OPTIMIZER_DEPLOYMENT = os.environ.get("OPTIMIZER_DEPLOYMENT", "")
_config_lock = threading.Lock()
tracker = TokenTracker()
@@ -234,6 +237,33 @@ def chat_target(
)
def chat_optimizer(
system: str,
user: str,
max_completion_tokens: int = 16384,
retries: int = 5,
stage: str = "optimizer",
reasoning_effort: str | None = None,
timeout: float | None = None,
) -> tuple[str, dict[str, int]]:
"""Optimizer chat call. Backend stores the trained skill; uses the same
MiniMax-proxied OpenAI-compat endpoint as `chat_target`. Added in the
parallel-training fix; previously missing in skillopt 0.2.0's
miniamax backend, which forced the dispatcher into _openai.chat_optimizer
(Azure) and produced "[skip] no usable patches" for any user running
optimizer+target on `minimax_chat`.
"""
messages = [{"role": "system", "content": system}, {"role": "user", "content": user}]
return _chat_messages_impl(
messages,
max_completion_tokens,
retries,
stage,
deployment=OPTIMIZER_DEPLOYMENT or None,
timeout=timeout,
)
def chat_target_messages(
messages: list[dict[str, Any]],
max_completion_tokens: int = 16384,
@@ -259,6 +289,35 @@ def chat_target_messages(
)
def chat_optimizer_messages(
messages: list[dict[str, Any]],
max_completion_tokens: int = 16384,
retries: int = 5,
stage: str = "optimizer",
reasoning_effort: str | None = None,
*,
tools: list[dict[str, Any]] | None = None,
tool_choice: str | dict[str, Any] | None = None,
return_message: bool = False,
timeout: float | None = None,
) -> tuple[Any, dict[str, int]]:
"""Optimizer-role message API. Same endpoint as ``chat_target_messages`` but
honours ``OPTIMIZER_DEPLOYMENT`` so optimizer and target can use distinct
MiniMax models; falls back to the target deployment when unset."""
del reasoning_effort
return _chat_messages_impl(
messages,
max_completion_tokens,
retries,
stage,
tools=tools,
tool_choice=tool_choice,
return_message=return_message,
deployment=OPTIMIZER_DEPLOYMENT or None,
timeout=timeout,
)
def get_token_summary() -> dict[str, dict[str, int]]:
return tracker.summary()
@@ -274,4 +333,10 @@ def set_reasoning_effort(effort: str | None) -> None:
def set_target_deployment(deployment: str) -> None:
global TARGET_DEPLOYMENT
TARGET_DEPLOYMENT = deployment or default_model_for_backend("minimax_chat")
os.environ["TARGET_DEPLOYMENT"] = TARGET_DEPLOYMENT
os.environ["TARGET_DEPLOYMENT"] = TARGET_DEPLOYMENT
def set_optimizer_deployment(deployment: str) -> None:
global OPTIMIZER_DEPLOYMENT
OPTIMIZER_DEPLOYMENT = deployment or default_model_for_backend("minimax_chat")
os.environ["OPTIMIZER_DEPLOYMENT"] = OPTIMIZER_DEPLOYMENT
+60 -32
View File
@@ -1,13 +1,14 @@
"""OpenAI-compatible Qwen chat backend for optimizer and target paths."""
from __future__ import annotations
from dataclasses import dataclass
import json
import os
import threading
import time
import urllib.error
import urllib.request
from dataclasses import dataclass
from typing import Any
from skillopt.model.common import (
@@ -28,6 +29,7 @@ class QwenChatConfig:
temperature: float | None
enable_thinking: bool
deployment: str
use_max_completion_tokens: bool = False
def _parse_bool(value: Any, default: bool = False) -> bool:
@@ -56,6 +58,28 @@ def _role_env(role: str, key: str, default: str) -> str:
return os.environ.get(role_key) or os.environ.get(generic_key) or default
# Sentinels that mean "omit this optional parameter from the request payload".
# Reasoning models (e.g. GPT-5.x, Claude Opus 4.8) reject an explicit
# `temperature`, so allow it to be turned off via an empty string / none / off.
_OMIT_SENTINELS = {"", "none", "off", "null"}
def _resolve_temperature(role: str) -> float | None:
"""Return the temperature, or None to omit it entirely.
Unlike ``_role_env`` an *explicitly set* empty (or ``none``/``off``) value is
honored as "omit" instead of collapsing to the default. Precedence:
role-specific env -> generic env -> 0.7 default.
"""
for key in (f"{role.upper()}_QWEN_CHAT_TEMPERATURE", "QWEN_CHAT_TEMPERATURE"):
if key in os.environ:
raw = os.environ[key].strip()
if raw.lower() in _OMIT_SENTINELS:
return None
return float(raw)
return 0.7
def _initial_config(role: str) -> QwenChatConfig:
role_upper = role.upper()
deployment_env = "OPTIMIZER_DEPLOYMENT" if role == "optimizer" else "TARGET_DEPLOYMENT"
@@ -64,8 +88,9 @@ def _initial_config(role: str) -> QwenChatConfig:
api_key=_role_env(role, "API_KEY", ""),
timeout_seconds=float(_role_env(role, "TIMEOUT_SECONDS", "300") or 300),
max_tokens=_parse_int(_role_env(role, "MAX_TOKENS", "8000"), 8000),
temperature=_parse_optional_float(_role_env(role, "TEMPERATURE", "0.7")),
temperature=_resolve_temperature(role),
enable_thinking=_parse_bool(_role_env(role, "ENABLE_THINKING", "false")),
use_max_completion_tokens=_parse_bool(_role_env(role, "USE_MAX_COMPLETION_TOKENS", "false")),
deployment=(
os.environ.get(f"{role_upper}_QWEN_CHAT_MODEL")
or os.environ.get("QWEN_CHAT_MODEL")
@@ -186,10 +211,14 @@ def _chat_messages_impl(
timeout: float | None = None,
) -> tuple[Any, dict[str, int]]:
config = OPTIMIZER_CONFIG if role == "optimizer" else TARGET_CONFIG
token_limit = min(max_completion_tokens, config.max_tokens)
# Reasoning models on some gateways (GPT-5.x, o-series) require
# `max_completion_tokens` and reject the legacy `max_tokens`.
token_key = "max_completion_tokens" if config.use_max_completion_tokens else "max_tokens"
payload: dict[str, Any] = {
"model": deployment or config.deployment,
"messages": _json_safe(messages),
"max_tokens": min(max_completion_tokens, config.max_tokens),
token_key: token_limit,
}
if config.enable_thinking:
payload["chat_template_kwargs"] = {"enable_thinking": True}
@@ -219,7 +248,7 @@ def _chat_messages_impl(
return text, usage_info
except Exception as e: # noqa: BLE001
last_err = e
time.sleep(min(2 ** attempt, 30))
time.sleep(min(2**attempt, 30))
raise RuntimeError(f"Qwen chat call failed after {retries} retries: {last_err}")
@@ -231,18 +260,21 @@ def configure_qwen_chat(
timeout_seconds: float | str | None = None,
max_tokens: int | str | None = None,
enable_thinking: bool | str | None = None,
use_max_completion_tokens: bool | str | None = None,
optimizer_base_url: str | None = None,
optimizer_api_key: str | None = None,
optimizer_temperature: float | str | None = None,
optimizer_timeout_seconds: float | str | None = None,
optimizer_max_tokens: int | str | None = None,
optimizer_enable_thinking: bool | str | None = None,
optimizer_use_max_completion_tokens: bool | str | None = None,
target_base_url: str | None = None,
target_api_key: str | None = None,
target_temperature: float | str | None = None,
target_timeout_seconds: float | str | None = None,
target_max_tokens: int | str | None = None,
target_enable_thinking: bool | str | None = None,
target_use_max_completion_tokens: bool | str | None = None,
) -> None:
with _config_lock:
if base_url is not None:
@@ -256,29 +288,24 @@ def configure_qwen_chat(
if max_tokens is not None:
os.environ["QWEN_CHAT_MAX_TOKENS"] = str(max_tokens)
if enable_thinking is not None:
os.environ["QWEN_CHAT_ENABLE_THINKING"] = (
"true" if _parse_bool(enable_thinking) else "false"
os.environ["QWEN_CHAT_ENABLE_THINKING"] = "true" if _parse_bool(enable_thinking) else "false"
if use_max_completion_tokens is not None:
os.environ["QWEN_CHAT_USE_MAX_COMPLETION_TOKENS"] = (
"true" if _parse_bool(use_max_completion_tokens) else "false"
)
_update_config(
OPTIMIZER_CONFIG,
"optimizer",
base_url=optimizer_base_url if optimizer_base_url is not None else base_url,
api_key=optimizer_api_key if optimizer_api_key is not None else api_key,
temperature=(
optimizer_temperature
if optimizer_temperature is not None
else temperature
),
timeout_seconds=(
optimizer_timeout_seconds
if optimizer_timeout_seconds is not None
else timeout_seconds
),
temperature=(optimizer_temperature if optimizer_temperature is not None else temperature),
timeout_seconds=(optimizer_timeout_seconds if optimizer_timeout_seconds is not None else timeout_seconds),
max_tokens=optimizer_max_tokens if optimizer_max_tokens is not None else max_tokens,
enable_thinking=(
optimizer_enable_thinking
if optimizer_enable_thinking is not None
else enable_thinking
enable_thinking=(optimizer_enable_thinking if optimizer_enable_thinking is not None else enable_thinking),
use_max_completion_tokens=(
optimizer_use_max_completion_tokens
if optimizer_use_max_completion_tokens is not None
else use_max_completion_tokens
),
)
_update_config(
@@ -287,16 +314,13 @@ def configure_qwen_chat(
base_url=target_base_url if target_base_url is not None else base_url,
api_key=target_api_key if target_api_key is not None else api_key,
temperature=target_temperature if target_temperature is not None else temperature,
timeout_seconds=(
target_timeout_seconds
if target_timeout_seconds is not None
else timeout_seconds
),
timeout_seconds=(target_timeout_seconds if target_timeout_seconds is not None else timeout_seconds),
max_tokens=target_max_tokens if target_max_tokens is not None else max_tokens,
enable_thinking=(
target_enable_thinking
if target_enable_thinking is not None
else enable_thinking
enable_thinking=(target_enable_thinking if target_enable_thinking is not None else enable_thinking),
use_max_completion_tokens=(
target_use_max_completion_tokens
if target_use_max_completion_tokens is not None
else use_max_completion_tokens
),
)
@@ -311,6 +335,7 @@ def _update_config(
timeout_seconds: float | str | None = None,
max_tokens: int | str | None = None,
enable_thinking: bool | str | None = None,
use_max_completion_tokens: bool | str | None = None,
) -> None:
env_prefix = role.upper()
if base_url is not None:
@@ -321,7 +346,7 @@ def _update_config(
os.environ[f"{env_prefix}_QWEN_CHAT_API_KEY"] = config.api_key
if temperature is not None:
raw = str(temperature).strip()
config.temperature = float(raw) if raw else None
config.temperature = None if raw.lower() in _OMIT_SENTINELS else float(raw)
os.environ[f"{env_prefix}_QWEN_CHAT_TEMPERATURE"] = raw
if timeout_seconds is not None:
config.timeout_seconds = float(timeout_seconds)
@@ -331,8 +356,11 @@ def _update_config(
os.environ[f"{env_prefix}_QWEN_CHAT_MAX_TOKENS"] = str(max_tokens)
if enable_thinking is not None:
config.enable_thinking = _parse_bool(enable_thinking)
os.environ[f"{env_prefix}_QWEN_CHAT_ENABLE_THINKING"] = (
"true" if config.enable_thinking else "false"
os.environ[f"{env_prefix}_QWEN_CHAT_ENABLE_THINKING"] = "true" if config.enable_thinking else "false"
if use_max_completion_tokens is not None:
config.use_max_completion_tokens = _parse_bool(use_max_completion_tokens)
os.environ[f"{env_prefix}_QWEN_CHAT_USE_MAX_COMPLETION_TOKENS"] = (
"true" if config.use_max_completion_tokens else "false"
)
+130
View File
@@ -3,6 +3,91 @@ from __future__ import annotations
import json
import re
import warnings
def _top_level_brace_objects(text: str) -> list[str]:
"""Return every balanced *top-level* ``{...}`` span in ``text``.
Fully string/escape aware: braces inside quoted strings are ignored both
when scanning for an object start AND while tracking depth inside one, so a
``{`` that appears in prose (e.g. ``'set it to {x}'``) is never mistaken for
the start of a JSON object. Used to detect ambiguity: when a response carries
more than one top-level object we must not let a repair pass silently pick
one it may pick the wrong (discarded) edit, strictly worse than None.
"""
spans: list[str] = []
i, n = 0, len(text)
outer_in_str = False
outer_esc = False
while i < n:
ch = text[i]
# Skip over braces that live *inside* a quoted string before any object
# has started — otherwise a `{` in prose like '"set it to {x}"' is wrongly
# treated as an object start, and the repair pass below turns non-JSON
# prose into a bogus dict (strictly worse than returning None).
if outer_in_str:
if outer_esc:
outer_esc = False
elif ch == "\\":
outer_esc = True
elif ch == '"':
outer_in_str = False
i += 1
continue
if ch == '"':
outer_in_str = True
i += 1
continue
if ch != "{":
i += 1
continue
depth = 0
in_str = False
esc = False
start = i
while i < n:
ch = text[i]
if in_str:
if esc:
esc = False
elif ch == "\\":
esc = True
elif ch == '"':
in_str = False
elif ch == '"':
in_str = True
elif ch == "{":
depth += 1
elif ch == "}":
depth -= 1
if depth == 0:
spans.append(text[start:i + 1])
i += 1
break
i += 1
else:
break # unterminated final object
return spans
def _looks_json_like(span: str) -> bool:
"""Heuristic: does ``span`` look like an intended JSON object (vs. prose)?
A genuine JSON object's first non-space character after ``{`` is either ``"``
(a string key) or ``}`` (an empty object). Prose pseudo-objects that the
repair pass would otherwise fabricate into bogus dicts ``{op: delete}``,
``{x: 1}`` quoted in single quotes or backticks, etc. start with a bare
word and are rejected. This complements the string-aware scan, which only
skips *double*-quoted prose; single-quoted / backticked / unquoted prose
braces are caught here instead. Legitimate repair targets (trailing commas,
unescaped quotes inside string values) all begin with ``"`` and pass.
"""
inner = span.strip()
if not (inner.startswith("{") and inner.endswith("}")):
return False
after_brace = inner[1:].lstrip()
return after_brace[:1] in ('"', '}')
def extract_json(text: str) -> dict | None:
@@ -22,6 +107,51 @@ def extract_json(text: str) -> dict | None:
return json.loads(m.group(0))
except json.JSONDecodeError:
pass
# Tolerant fallback for non-OpenAI backends (Claude/Qwen, …) whose free-form
# JSON strict json.loads rejects — unescaped ASCII quotes inside CJK string
# values, trailing commas, etc. Repair so the analyst's edits aren't silently
# dropped, but ONLY a single unambiguous object: never feed the greedy `{.*}`
# span or the raw text, or json_repair would quietly return one of several
# objects (empirically the wrong/last one) — strictly worse than None, which
# the caller can detect and retry/skip.
#
# Pick the candidate FIRST, before importing json_repair, so the optional
# dependency only matters (and only warns) when there is genuinely a single
# malformed object we could have repaired. Ordinary no-JSON / prose replies
# have no candidate and return None silently.
candidate = None
fenced = re.search(r"```json\s*(.*?)```", text, re.DOTALL)
if fenced and len(_top_level_brace_objects(fenced.group(1))) == 1:
candidate = fenced.group(1)
else:
objs = _top_level_brace_objects(text)
if len(objs) == 1:
candidate = objs[0]
# 0 or >1 top-level objects → too ambiguous to repair safely → None
if not candidate:
return None
# Final guard: only repair spans that actually look like an intended JSON
# object. Prose pseudo-objects in single quotes / backticks / bare text
# (e.g. `{op: delete}`) reach here because the scan only skips double-quoted
# prose; repairing them would fabricate a wrong dict (worse than None).
if not _looks_json_like(candidate):
return None
try:
from json_repair import repair_json
except ModuleNotFoundError:
warnings.warn(
"json_repair not installed; malformed-JSON recovery disabled — "
"a non-OpenAI analyst edit may be silently dropped. pip install json_repair",
RuntimeWarning,
stacklevel=2,
)
return None
try:
repaired = repair_json(candidate, return_objects=True)
if isinstance(repaired, dict) and repaired:
return repaired
except Exception: # noqa: BLE001 — repair is best-effort
pass
return None
+1 -1
View File
@@ -17,4 +17,4 @@ Public entry points:
from __future__ import annotations
__all__ = ["__version__"]
__version__ = "0.1.0"
__version__ = "0.2.0"
+327 -29
View File
@@ -9,7 +9,12 @@
Common flags:
--project PATH project to evolve (default: cwd)
--scope all|invoked harvest scope (default: invoked)
--backend mock|anthropic
--max-sessions N cap transcript sessions per run
--max-tasks N cap mined tasks per run
--target-skill-path PATH explicit live SKILL.md to stage/adopt
--tasks-file PATH reviewed TaskRecord JSON file to replay instead of harvesting
--backend mock|claude|codex|copilot|handoff
--source claude|codex|auto
--model NAME
--lookback-hours N
--auto-adopt
@@ -25,26 +30,71 @@ from typing import Any, Dict
from skillopt_sleep.config import load_config
from skillopt_sleep.cycle import run_sleep_cycle
from skillopt_sleep.harvest import harvest
from skillopt_sleep.harvest_sources import harvest_for_config
from skillopt_sleep.mine import mine
from skillopt_sleep.staging import adopt as adopt_staging
from skillopt_sleep.staging import latest_staging
from skillopt_sleep.state import SleepState
from skillopt_sleep.staging import latest_staging, adopt as adopt_staging
from skillopt_sleep.tasks_file import load_tasks_file, make_tasks_payload, write_tasks_file
def _read_text(path: str) -> str:
try:
with open(path, encoding="utf-8") as f:
return f.read()
except Exception:
return ""
def _report_payload(rep, outcome) -> Dict[str, Any]:
return {
"night": rep.night,
"accepted": rep.accepted,
"gate_action": rep.gate_action,
"no_edits_reason": getattr(rep, "no_edits_reason", ""),
"baseline": rep.baseline_score,
"candidate": rep.candidate_score,
"n_tasks": rep.n_tasks,
"n_sessions": rep.n_sessions,
"n_accepted_edits": len(rep.edits),
"n_rejected_edits": len(rep.rejected_edits),
"edits": [e.__dict__ for e in rep.edits],
"rejected_edits": [e.__dict__ for e in rep.rejected_edits],
"notes": rep.notes,
"staging_dir": outcome.staging_dir,
"adopted": outcome.adopted,
}
def _add_common(p: argparse.ArgumentParser) -> None:
p.add_argument("--project", default="")
p.add_argument("--scope", default="", choices=["", "all", "invoked"])
p.add_argument("--backend", default="", choices=["", "mock", "claude", "codex"])
p.add_argument("--backend", default="",
choices=["", "mock", "claude", "codex", "copilot", "handoff"])
p.add_argument("--model", default="")
p.add_argument("--codex-path", default="", help="path to the real @openai/codex binary")
p.add_argument("--claude-home", default="", help="override ~/.claude (also isolates state)")
p.add_argument("--lookback-hours", type=int, default=0)
p.add_argument("--codex-home", default="", help="override ~/.codex for archived session harvest")
p.add_argument("--source", default="", choices=["", "claude", "codex", "auto"],
help="session transcript source")
p.add_argument("--lookback-hours", type=int, default=None,
help="harvest window in hours; 0 = scan full history")
p.add_argument("--edit-budget", type=int, default=0)
p.add_argument("--max-sessions", type=int, default=0,
help="cap harvested sessions before mining; default derives from max tasks")
p.add_argument("--max-tasks", type=int, default=0,
help="cap mined tasks for this run")
p.add_argument("--target-skill-path", default="",
help="explicit live SKILL.md path to evolve/stage/adopt")
p.add_argument("--tasks-file", default="",
help="reviewed TaskRecord JSON file to replay instead of harvesting")
p.add_argument("--progress", action="store_true",
help="print phase progress to stderr")
p.add_argument("--auto-adopt", action="store_true")
p.add_argument("--json", action="store_true")
def _cfg_from_args(args) -> Any:
def _cfg_from_args(args, task_meta: Dict[str, Any] | None = None) -> Any:
overrides: Dict[str, Any] = {}
if args.project:
overrides["invoked_project"] = os.path.abspath(args.project)
@@ -59,40 +109,266 @@ def _cfg_from_args(args) -> Any:
overrides["codex_path"] = os.path.abspath(args.codex_path)
if getattr(args, "claude_home", ""):
overrides["claude_home"] = os.path.abspath(args.claude_home)
if getattr(args, "lookback_hours", 0):
overrides["lookback_hours"] = args.lookback_hours
if getattr(args, "codex_home", ""):
overrides["codex_home"] = os.path.abspath(args.codex_home)
if getattr(args, "source", ""):
overrides["transcript_source"] = args.source
lh = getattr(args, "lookback_hours", None)
if lh is not None: # --lookback-hours was explicitly passed (0 = full history)
overrides["lookback_hours"] = lh
if getattr(args, "edit_budget", 0):
overrides["edit_budget"] = args.edit_budget
if getattr(args, "max_sessions", 0):
overrides["max_sessions_per_night"] = args.max_sessions
if getattr(args, "max_tasks", 0):
overrides["max_tasks_per_night"] = args.max_tasks
target_skill_path = getattr(args, "target_skill_path", "")
if not target_skill_path and task_meta:
target_skill_path = str(task_meta.get("target_skill_path") or "")
if target_skill_path:
path = os.path.expanduser(target_skill_path)
if args.project and not os.path.isabs(path):
path = os.path.join(os.path.abspath(args.project), path)
overrides["target_skill_path"] = os.path.abspath(path)
if getattr(args, "progress", False):
overrides["progress"] = True
if getattr(args, "auto_adopt", False):
overrides["auto_adopt"] = True
return load_config(**overrides)
def cmd_run(args, dry: bool = False) -> int:
cfg = _cfg_from_args(args)
outcome = run_sleep_cycle(cfg, dry_run=dry)
task_meta: Dict[str, Any] = {}
tasks = None
if getattr(args, "tasks_file", ""):
# Load once before config so target_skill_path can default from metadata.
tasks, task_meta = load_tasks_file(args.tasks_file)
cfg = _cfg_from_args(args, task_meta=task_meta)
if getattr(args, "tasks_file", ""):
tasks, task_meta = load_tasks_file(
args.tasks_file,
holdout_fraction=cfg.get("holdout_fraction", 0.34),
seed=cfg.get("seed", 42),
)
if cfg.get("backend", "mock") != "mock" and task_meta.get("reviewed") is not True:
print(
"[sleep] refusing real-backend replay from an unreviewed tasks file; "
"inspect/redact it and set \"reviewed\": true first",
file=sys.stderr,
)
return 2
if cfg.get("backend", "mock") == "handoff":
return _run_handoff(cfg, args, seed_tasks=tasks, task_meta=task_meta, dry=dry)
outcome = run_sleep_cycle(cfg, seed_tasks=tasks, dry_run=dry)
_print_run_report(outcome, args, task_meta)
return 0
def _print_run_report(outcome, args, task_meta: Dict[str, Any]) -> None:
rep = outcome.report
if args.json:
print(json.dumps({
"night": rep.night, "accepted": rep.accepted,
"gate_action": rep.gate_action,
"baseline": rep.baseline_score, "candidate": rep.candidate_score,
"n_tasks": rep.n_tasks, "n_sessions": rep.n_sessions,
"edits": [e.__dict__ for e in rep.edits],
"staging_dir": outcome.staging_dir, "adopted": outcome.adopted,
}, ensure_ascii=False, indent=2))
payload = _report_payload(rep, outcome)
if task_meta:
payload["tasks_file"] = task_meta.get("tasks_file", "")
payload["tasks_reviewed"] = task_meta.get("reviewed", False)
print(json.dumps(payload, ensure_ascii=False, indent=2))
else:
print(f"[sleep] night {rep.night}: {rep.n_sessions} sessions -> {rep.n_tasks} tasks")
print(f"[sleep] held-out {rep.baseline_score:.3f} -> {rep.candidate_score:.3f} "
f"=> {rep.gate_action} (accepted={rep.accepted})")
for e in rep.edits:
print(f" + [{e.target}/{e.op}] {e.content}")
if rep.rejected_edits:
print("[sleep] rejected by gate:")
for e in rep.rejected_edits:
print(f" - [{e.target}/{e.op}] {e.content}")
if outcome.staging_dir:
print(f"[sleep] staged: {outcome.staging_dir}")
if not outcome.adopted:
print("[sleep] review it, then: python -m skillopt_sleep adopt")
if outcome.adopted:
print(f"[sleep] auto-adopted: {', '.join(outcome.adopted_paths)}")
def _handoff_dir_for(cfg) -> str:
project = cfg.get("invoked_project") or os.getcwd()
return os.environ.get("SKILLOPT_SLEEP_HANDOFF_DIR", "") or os.path.join(
project, ".skillopt-sleep-handoff"
)
def _redact_deep(obj):
"""Redact secret-looking substrings in every string of a JSON-like tree."""
from skillopt_sleep.staging import redact_secrets
if isinstance(obj, str):
return redact_secrets(obj)
if isinstance(obj, list):
return [_redact_deep(x) for x in obj]
if isinstance(obj, dict):
return {k: _redact_deep(v) for k, v in obj.items()}
return obj
def _flush_handoff(backend, args) -> int:
prompts_path = backend.flush_pending()
if args.json:
print(json.dumps({
"handoff_pending": len(backend.pending),
"prompts": prompts_path,
"answers_dir": backend.answers_dir,
}, ensure_ascii=False, indent=2))
else:
print(f"[sleep] handoff: {len(backend.pending)} model call(s) need answers")
print(f"[sleep] prompts: {prompts_path}")
print(f"[sleep] write each raw answer to {backend.answers_dir}/<id>.md, "
"then re-run this exact command to resume")
return 3
def _handoff_mine_and_pin(cfg, args, backend, snapshot: str, dry: bool):
"""Harvest + mine with the same knobs as run_sleep_cycle (harvest window,
target-skill filter, candidate-limit bump, LLM mining routed through the
handoff files like every other model call), then pin the result to
``tasks.json``. Session digests are pinned too, so the sessions created
while answering prompts cannot change what gets mined between rounds.
Returns ``(exit_code, tasks)``; ``tasks is None`` means exit now.
"""
import time
from skillopt_sleep.handoff_backend import PendingCalls
from skillopt_sleep.state import SleepState, _now_iso
from skillopt_sleep.types import SessionDigest
project = cfg.get("invoked_project") or os.getcwd()
state = SleepState.load(cfg.state_path)
started = _now_iso()
digests_path = os.path.join(backend.handoff_dir, "digests.json")
digests = None
if os.path.exists(digests_path):
try:
with open(digests_path, encoding="utf-8") as f:
raw = json.load(f)
known = set(SessionDigest.__dataclass_fields__)
digests = [SessionDigest(**{k: v for k, v in d.items() if k in known})
for d in raw]
except Exception:
# Corrupted/truncated pin (e.g. an interrupted earlier round):
# fall back to a fresh harvest instead of crashing the run.
print("[sleep] handoff: digests.json unreadable — re-harvesting",
file=sys.stderr)
digests = None
if digests is None:
since = state.last_harvest_for(project)
lookback_hours = cfg.get("lookback_hours", 72)
if since is None and lookback_hours and lookback_hours > 0:
since = _now_iso(time.time() - lookback_hours * 3600)
max_tasks = cfg.get("max_tasks_per_night", 40)
session_limit = cfg.get("max_sessions_per_night", 0) or max_tasks * 3
digests = harvest_for_config(cfg, since_iso=since, limit=session_limit)
os.makedirs(backend.handoff_dir, exist_ok=True)
with open(digests_path, "w", encoding="utf-8") as f:
json.dump(_redact_deep([d.to_dict() for d in digests]), f,
ensure_ascii=False, indent=2)
max_tasks = cfg.get("max_tasks_per_night", 40)
session_limit = cfg.get("max_sessions_per_night", 0) or max_tasks * 3
target_skill_path = cfg.managed_skill_path() if cfg.get("target_skill_path", "") else ""
target_skill_text = _read_text(target_skill_path) if target_skill_path else ""
candidate_limit = max_tasks
if cfg.get("target_task_filter", True) and target_skill_text:
candidate_limit = max(max_tasks, max_tasks * 3)
llm_miner = None
if cfg.get("llm_mine", True):
try:
from skillopt_sleep.llm_miner import make_llm_miner
llm_miner = make_llm_miner(
backend, max_sessions=session_limit, max_tasks=candidate_limit,
)
except Exception:
llm_miner = None
try:
tasks = mine(
digests,
max_tasks=max_tasks,
candidate_limit=candidate_limit,
holdout_fraction=cfg.get("holdout_fraction", 0.34),
seed=cfg.get("seed", 42),
llm_miner=llm_miner,
target_skill_text=target_skill_text,
target_skill_path=target_skill_path,
)
except PendingCalls:
tasks = []
if backend.pending:
# LLM mining needs answers before the task set can be pinned.
return _flush_handoff(backend, args), None
if not tasks:
print("[sleep] handoff: no tasks mined — nothing to consolidate")
if not dry:
# Advance the harvest window like run_sleep_cycle's no-tasks
# branch, or every later run re-scans the same stale window.
state.set_last_harvest(project, started)
state.save()
return 0, None
payload = make_tasks_payload(
tasks,
project=project,
transcript_source=cfg.get("transcript_source", ""),
n_sessions=len(digests),
target_skill_path=target_skill_path,
)
# NOT marked reviewed: feeding this snapshot back through --tasks-file
# with a real backend must still hit the human-review gate above. The
# driver itself loads it directly, with the same trust as in-cycle mining.
write_tasks_file(snapshot, _redact_deep(payload))
print(f"[sleep] handoff: pinned {len(tasks)} tasks -> {snapshot}")
return 0, tasks
def _run_handoff(cfg, args, *, seed_tasks, task_meta: Dict[str, Any], dry: bool) -> int:
"""Drive the handoff backend: run until model calls are needed, then
write the prompt batch and exit 3; on a fully-answered run, finish
normally. Session digests and mined tasks are pinned under the handoff
dir on the first rounds so wall-clock time between rounds (including
the very sessions that answer the prompts) cannot change the task set
and invalidate earlier answers.
"""
from skillopt_sleep.handoff_backend import HandoffBackend, PendingCalls
hdir = _handoff_dir_for(cfg)
backend = HandoffBackend(model=cfg.get("model", ""), handoff_dir=hdir)
tasks = seed_tasks
if tasks is None:
snapshot = os.path.join(hdir, "tasks.json")
if os.path.exists(snapshot):
tasks, _meta = load_tasks_file(
snapshot,
holdout_fraction=cfg.get("holdout_fraction", 0.34),
seed=cfg.get("seed", 42),
)
else:
rc, tasks = _handoff_mine_and_pin(cfg, args, backend, snapshot, dry)
if tasks is None:
return rc
outcome = None
try:
outcome = run_sleep_cycle(cfg, seed_tasks=tasks, dry_run=dry, backend=backend)
except PendingCalls:
pass
if backend.pending:
return _flush_handoff(backend, args)
_print_run_report(outcome, args, task_meta)
# A completed real run ends the night: archive the handoff dir so the
# next night re-harvests instead of replaying the pinned snapshot.
if not dry and outcome.staging_dir and os.path.isdir(hdir):
import time
done = f"{hdir}.night{outcome.report.night}.done"
if os.path.exists(done):
done = f"{done}.{int(time.time())}"
os.rename(hdir, done)
print(f"[sleep] handoff: archived round data -> {done}")
return 0
@@ -143,21 +419,42 @@ def cmd_adopt(args) -> int:
def cmd_harvest(args) -> int:
cfg = _cfg_from_args(args)
digests = harvest(
cfg.transcripts_dir,
scope=cfg.get("projects", "invoked"),
invoked_project=cfg.get("invoked_project", ""),
limit=cfg.get("max_tasks_per_night", 40) * 3,
session_limit = cfg.get("max_sessions_per_night", 0) or cfg.get("max_tasks_per_night", 40) * 3
target_skill_path = cfg.managed_skill_path() if cfg.get("target_skill_path", "") else ""
target_skill_text = _read_text(target_skill_path) if target_skill_path else ""
max_tasks = cfg.get("max_tasks_per_night", 40)
candidate_limit = max_tasks
if cfg.get("target_task_filter", True) and target_skill_text:
candidate_limit = max(max_tasks, max_tasks * 3)
digests = harvest_for_config(cfg, limit=session_limit)
tasks = mine(
digests,
max_tasks=max_tasks,
candidate_limit=candidate_limit,
holdout_fraction=cfg.get("holdout_fraction", 0.34),
seed=cfg.get("seed", 42),
target_skill_text=target_skill_text,
target_skill_path=target_skill_path,
)
tasks = mine(digests, max_tasks=cfg.get("max_tasks_per_night", 40),
holdout_fraction=cfg.get("holdout_fraction", 0.34), seed=cfg.get("seed", 42))
payload = make_tasks_payload(
tasks,
project=cfg.get("invoked_project") or os.getcwd(),
transcript_source=cfg.get("transcript_source", ""),
n_sessions=len(digests),
target_skill_path=target_skill_path,
)
output_path = ""
if getattr(args, "output", ""):
output_path = write_tasks_file(args.output, payload)
if args.json:
print(json.dumps({
"n_sessions": len(digests),
"tasks": [t.to_dict() for t in tasks],
}, ensure_ascii=False, indent=2))
json_payload = dict(payload)
if output_path:
json_payload["output"] = output_path
print(json.dumps(json_payload, ensure_ascii=False, indent=2))
else:
print(f"[sleep] {len(digests)} sessions -> {len(tasks)} tasks")
if output_path:
print(f"[sleep] wrote reviewed-task draft: {output_path}")
for t in tasks:
print(f" [{t.split}/{t.outcome}] {t.intent[:90]}")
return 0
@@ -203,6 +500,7 @@ def main(argv=None) -> int:
p_adopt.add_argument("--staging", default="", help="specific staging dir")
p_harvest = sub.add_parser("harvest", help="debug: show mined tasks")
_add_common(p_harvest)
p_harvest.add_argument("--output", default="", help="write mined tasks JSON for review")
p_sched = sub.add_parser("schedule", help="install a nightly cron entry for this project")
_add_common(p_sched)
p_sched.add_argument("--hour", type=int, default=3)
+427 -32
View File
@@ -24,10 +24,17 @@ import json
import os
import re
import subprocess
import tempfile
from typing import Any, Dict, List, Optional, Tuple
from skillopt_sleep.types import EditRecord, ReplayResult, TaskRecord
# On Windows, console-attached children (cmd.exe shims, python) allocate a
# visible console window when the parent has none — a nightly cycle making
# hundreds of CLI calls strobes cmd windows and steals focus from the user.
# CREATE_NO_WINDOW suppresses that; harmless 0 elsewhere.
_NO_WINDOW = getattr(subprocess, "CREATE_NO_WINDOW", 0) if os.name == "nt" else 0
def skill_hash(content: str) -> str:
import hashlib
@@ -315,6 +322,8 @@ class CliBackend(Backend):
self.timeout = timeout
self._tokens = 0
self._cache: Dict[str, str] = {}
self.last_call_error = ""
self.last_reflect_raw = ""
# subclasses override --------------------------------------------------
def _call(self, prompt: str, *, max_tokens: int = 1024) -> str:
@@ -517,6 +526,10 @@ class CliBackend(Backend):
arr = _extract_json(raw, "array")
if isinstance(arr, list) and arr:
break
# Expose the last raw optimizer reply so a no-edits night is diagnosable:
# a 0.0->0.0 gate with zero edits is otherwise indistinguishable from
# "nothing to learn" (the cycle persists this in diagnostics.json).
self.last_reflect_raw = raw or ""
edits: List[EditRecord] = []
if isinstance(arr, list):
for e in arr[:edit_budget]:
@@ -548,7 +561,45 @@ class ClaudeCliBackend(CliBackend):
def __init__(self, model: str = "", claude_path: str = "claude", timeout: int = 180) -> None:
super().__init__(model=model or os.environ.get("SKILLOPT_SLEEP_CLAUDE_MODEL", "") or "sonnet",
timeout=timeout)
self.claude_path = claude_path
# On Windows the npm-installed `claude` is a .cmd shim; CreateProcess
# cannot resolve it by bare name (WinError 2), so every call would
# silently return "" and the whole cycle scores 0.0. shutil.which
# honors PATHEXT and returns the full claude.CMD path.
import shutil as _shutil
self.claude_path = _shutil.which(claude_path) or claude_path
# Known CLI error prefixes that indicate auth or config failures.
# When detected, we log a warning so the user doesn't mistake a
# broken auth for "nothing to optimize" (issue #68).
# Keep these specific to avoid false positives on normal model output.
_CLI_ERROR_MARKERS = (
"Not logged in",
"Please run /login",
"Authentication required",
"Invalid API key",
"Unauthorized: invalid x-api-key",
)
def _detect_cli_error(self, stdout: str, stderr: str) -> None:
"""Log a warning if CLI output looks like an auth/config error.
Only checks stderr and short stdout (< 300 chars) to avoid
false-positives on legitimate model responses that mention
auth-related terms.
"""
import logging
# Long stdout is almost certainly a real model response, not an error.
check_stdout = stdout if len(stdout) < 300 else ""
combined = check_stdout + "\n" + stderr
for marker in self._CLI_ERROR_MARKERS:
if marker in combined:
from skillopt_sleep.staging import redact_secrets
logging.getLogger("skillopt_sleep").warning(
"Claude CLI returned a likely auth error: %s",
redact_secrets(combined[:200].replace("\n", " ")),
)
self.last_call_error = combined[:500]
return
def _call(self, prompt: str, *, max_tokens: int = 1024) -> str:
# Run ISOLATED so the ambient Claude Code environment does not leak into
@@ -557,27 +608,38 @@ class ClaudeCliBackend(CliBackend):
# them explicitly — without this, reflect/attempt sometimes reply with a
# list of the user's installed skills instead of doing the task.
# --bare skip hooks, LSP, plugins (minimal mode)
# Only safe with ANTHROPIC_API_KEY auth;
# breaks subscription-token auth (#68).
# --disable-slash-commands disable all skills
# --disallowedTools '*' no tool use
# --exclude-dynamic-... drop per-machine cwd/env/memory/git sections
# cwd=<clean temp> no project CLAUDE.md
import tempfile
cmd = [
self.claude_path, "-p", "--output-format", "text",
"--bare",
cmd = [self.claude_path, "-p", "--output-format", "text"]
if os.environ.get("ANTHROPIC_API_KEY"):
cmd.append("--bare")
cmd += [
"--disable-slash-commands",
"--disallowedTools", "*",
"--exclude-dynamic-system-prompt-sections",
]
if self.model:
cmd += ["--model", self.model]
cmd += ["--", prompt]
# Prompt goes via stdin, not argv: the Windows .cmd shim routes through
# cmd.exe whose command line caps at ~8K chars — reflect/judge prompts
# exceed that. `claude -p` with no positional prompt reads stdin.
clean_cwd = tempfile.mkdtemp(prefix="skillopt_sleep_claude_")
try:
proc = subprocess.run(
cmd, capture_output=True, text=True, timeout=self.timeout, cwd=clean_cwd,
cmd, capture_output=True, creationflags=_NO_WINDOW, text=True, timeout=self.timeout, cwd=clean_cwd,
input=prompt,
)
except Exception as exc:
import logging
self.last_call_error = f"Claude CLI spawn failed: {exc}"
logging.getLogger("skillopt_sleep").warning(
"Claude CLI could not be executed: %s", exc,
)
except Exception:
return ""
finally:
try:
@@ -585,7 +647,9 @@ class ClaudeCliBackend(CliBackend):
shutil.rmtree(clean_cwd, ignore_errors=True)
except Exception:
pass
return (proc.stdout or "").strip()
out = (proc.stdout or "").strip()
self._detect_cli_error(out, proc.stderr or "")
return out
def attempt_with_tools(self, task, skill, memory, tools):
# Expose a REAL, callable `search` tool (a shell shim that logs each
@@ -622,21 +686,29 @@ class ClaudeCliBackend(CliBackend):
f"# Task\n{task.intent}\n\n{task.context_excerpt}\n\n"
"Return ONLY the final answer text."
)
cmd = [
self.claude_path, "-p", "--output-format", "text",
"--bare", "--disable-slash-commands",
cmd = [self.claude_path, "-p", "--output-format", "text"]
if os.environ.get("ANTHROPIC_API_KEY"):
cmd.append("--bare")
cmd += [
"--disable-slash-commands",
"--allowedTools", "Bash",
"--exclude-dynamic-system-prompt-sections",
]
if self.model:
cmd += ["--model", self.model]
cmd += ["--", prompt]
try:
proc = subprocess.run(
cmd, capture_output=True, text=True, timeout=self.timeout, cwd=work,
cmd, capture_output=True, creationflags=_NO_WINDOW, text=True, timeout=self.timeout, cwd=work,
input=prompt,
)
resp = (proc.stdout or "").strip()
except Exception:
self._detect_cli_error(resp, proc.stderr or "")
except Exception as exc:
import logging
self.last_call_error = f"Claude CLI spawn failed: {exc}"
logging.getLogger("skillopt_sleep").warning(
"Claude CLI could not be executed: %s", exc,
)
resp = ""
self._tokens += len(prompt) // 4 + len(resp) // 4
called: List[str] = []
@@ -691,14 +763,26 @@ class CodexCliBackend(CliBackend):
name = "codex"
def __init__(self, model: str = "", codex_path: str = "", timeout: int = 240,
sandbox: str = "read-only") -> None:
def __init__(
self,
model: str = "",
codex_path: str = "",
timeout: int = 240,
sandbox: str = "read-only",
project_dir: str = "",
) -> None:
super().__init__(model=model or os.environ.get("SKILLOPT_SLEEP_CODEX_MODEL", ""),
timeout=timeout)
self.codex_path = resolve_codex_path(codex_path)
self.sandbox = sandbox
self.project_dir = (
os.path.abspath(os.path.expanduser(project_dir)) if project_dir else ""
)
def _call(self, prompt: str, *, max_tokens: int = 1024) -> str:
def _call_once(self, prompt: str, *, max_tokens: int = 1024) -> str:
"""One codex exec attempt: returns the response text, or "" on
timeout/exception/empty-output (with last_call_error set). ``_call``
wraps this with retries so a transient failure is NOT silently scored 0."""
import tempfile
out_path = tempfile.NamedTemporaryFile(
prefix="codex_last_", suffix=".txt", delete=False
@@ -708,24 +792,93 @@ class CodexCliBackend(CliBackend):
"--color", "never", "--sandbox", self.sandbox,
"-o", out_path,
]
if self.project_dir:
cmd[3:3] = ["-C", self.project_dir]
if self.model:
cmd += ["-m", self.model]
cmd += ["--", prompt]
proc = None
try:
subprocess.run(cmd, capture_output=True, text=True, timeout=self.timeout)
except Exception:
return ""
try:
with open(out_path, encoding="utf-8") as f:
return f.read().strip()
except Exception:
return ""
try:
proc = subprocess.run(
cmd,
capture_output=True,
creationflags=_NO_WINDOW,
text=True,
timeout=self.timeout,
cwd=self.project_dir or None,
)
except subprocess.TimeoutExpired:
self.last_call_error = f"codex exec timed out after {self.timeout}s"
return ""
except Exception as exc:
self.last_call_error = f"codex exec failed: {exc}"
return ""
try:
with open(out_path, encoding="utf-8") as f:
out = f.read().strip()
if out:
return out
except Exception as exc:
self.last_call_error = f"could not read codex output file: {exc}"
stdout = (proc.stdout or "").strip() if proc is not None else ""
stderr = (proc.stderr or "").strip() if proc is not None else ""
if proc is not None and proc.returncode != 0 and not self.last_call_error:
self.last_call_error = f"codex exec exited {proc.returncode}: {stderr[:500]}"
# Do NOT return the CLI's error text as if it were a model response: it
# pollutes rollout/judge/reflect and gets silently scored 0, hiding the
# real cause (e.g. an expired codex auth token surfacing as a 9k-char 401).
# Surface it via last_call_error and return empty instead.
if self.last_call_error:
return ""
return stdout or stderr
finally:
try:
os.unlink(out_path)
except Exception:
pass
# Fatal codex failures that will NOT recover on retry — fail fast + loud so a
# 0.0 night reads as "codex auth/model/version problem" not "nothing to learn".
# Covers: auth (re-login), and 400 config errors like an unsupported model on a
# ChatGPT account or a model that needs a newer codex CLI (upgrade).
_AUTH_MARKERS = (
"401 Unauthorized", "refresh_token_reused", "token_expired",
"Please log out and sign in", "Not logged in", "Please run /login",
"authentication token is expired", "Unauthorized: invalid",
"is not supported when using Codex", "requires a newer version of Codex",
)
def _call(self, prompt: str, *, max_tokens: int = 1024, retries: int = 3) -> str:
"""Retry transient empties/timeouts instead of silently returning "".
An empty reply scores 0 on every judge, which deflates the held-out
baseline AND blocks the candidate from ever improving making a flaky
backend indistinguishable from "nothing to learn". The Azure backend
already guards this way (AzureOpenAIBackend._call); codex now does too.
Auth errors are NOT retried (hopeless until the user re-logs-in).
"""
import logging
import random as _r
import time as _t
out = ""
for attempt in range(max(1, retries)):
self.last_call_error = ""
out = self._call_once(prompt, max_tokens=max_tokens)
if out:
return out
err = self.last_call_error or ""
if any(m in err for m in self._AUTH_MARKERS):
from skillopt_sleep.staging import redact_secrets
logging.getLogger("skillopt_sleep").error(
"codex auth error — re-login required (`codex login`): %s",
redact_secrets(err[:200]),
)
break # fail fast: retrying a 401 just burns calls
if attempt < retries - 1:
_t.sleep(min(6.0, (2 ** attempt) * 0.5) + _r.random() * 0.3)
return out
def attempt_with_tools(self, task, skill, memory, tools):
# Codex exec runs in a sandbox with shell access; expose the same real
# `search` shim and let it run (workspace-write so the shim can log).
@@ -765,16 +918,27 @@ class CodexCliBackend(CliBackend):
if self.model:
cmd += ["-m", self.model]
cmd += ["--", prompt]
self.last_call_error = ""
proc = None
try:
subprocess.run(cmd, capture_output=True, text=True, timeout=self.timeout, cwd=work)
except Exception:
pass
proc = subprocess.run(cmd, capture_output=True, creationflags=_NO_WINDOW, text=True, timeout=self.timeout, cwd=work)
except subprocess.TimeoutExpired:
self.last_call_error = f"codex exec (tools) timed out after {self.timeout}s"
except Exception as exc: # noqa: BLE001
self.last_call_error = f"codex exec (tools) failed: {exc}"
resp = ""
try:
with open(out_path, encoding="utf-8") as f:
resp = f.read().strip()
except Exception:
resp = ""
# Surface a failed tool-rollout the SAME way _call does: an auth/model/version
# failure on this path must show up in diagnostics (call_error), not vanish as a
# silent empty->0 scored as a failed rollout. Response stays "" (never the error text).
if not resp and not self.last_call_error and proc is not None and proc.returncode != 0:
self.last_call_error = (
f"codex exec (tools) exited {proc.returncode}: {(proc.stderr or '')[:500]}"
)
self._tokens += len(prompt) // 4 + len(resp) // 4
called: List[str] = []
if os.path.exists(calllog):
@@ -788,6 +952,218 @@ class CodexCliBackend(CliBackend):
except Exception:
pass
def resolve_copilot_path(explicit: str = "") -> str:
"""Find the GitHub Copilot CLI (`copilot`) binary."""
if explicit:
return explicit
env = os.environ.get("SKILLOPT_SLEEP_COPILOT_PATH")
if env:
return env
import shutil
found = shutil.which("copilot")
return found or "copilot"
class CopilotCliBackend(CliBackend):
"""Drives the GitHub Copilot CLI in non-interactive mode.
Uses ``copilot -p <prompt> --output-format json`` and parses the emitted
JSONL event stream, returning the concatenated ``assistant.message``
content. The plain-text / ``--silent`` modes do not reliably stream the
response to stdout on all platforms, so JSONL is used for robust capture.
The call runs in a clean temp cwd with streaming disabled and tools allowed
(so non-interactive mode never blocks on a permission prompt); ``_call``'s
prompts ask for final-answer text only, so no tool use is expected there,
while ``attempt_with_tools`` exposes real, cross-platform callable shims in
the working directory for honest tool-call detection.
Startup overhead is minimised: each invocation points ``COPILOT_HOME`` at a
dedicated, isolated config dir (no user ``mcp-config.json``, so the user's
MCP servers including this project's own — are NOT spawned, avoiding a
slow recursive launch), and built-in MCP servers / custom instructions are
disabled. Auth is read from the OS credential store / token env vars, which
live outside ``COPILOT_HOME``, so isolation does not break authentication.
Set ``SKILLOPT_SLEEP_COPILOT_HOME`` to override the isolated home, or set it
empty / ``SKILLOPT_SLEEP_COPILOT_FULL_ENV=1`` to use the user's real
environment instead.
"""
name = "copilot"
def __init__(self, model: str = "", copilot_path: str = "", timeout: int = 240) -> None:
super().__init__(model=model or os.environ.get("SKILLOPT_SLEEP_COPILOT_MODEL", ""),
timeout=timeout)
self.copilot_path = resolve_copilot_path(copilot_path)
self.full_env = os.environ.get("SKILLOPT_SLEEP_COPILOT_FULL_ENV", "") == "1"
# Stable isolated home so first-run setup is cached across calls.
if self.full_env:
self.copilot_home = ""
else:
self.copilot_home = os.environ.get("SKILLOPT_SLEEP_COPILOT_HOME") or os.path.join(
tempfile.gettempdir(), "skillopt_sleep_copilot_home"
)
try:
os.makedirs(self.copilot_home, exist_ok=True)
except Exception:
self.copilot_home = ""
def _call(self, prompt: str, *, max_tokens: int = 1024) -> str:
clean_cwd = tempfile.mkdtemp(prefix="skillopt_sleep_copilot_")
cmd = [
self.copilot_path, "-p", prompt,
"--output-format", "json",
"--stream", "off",
"--no-color",
"--log-level", "none",
"--allow-all-tools",
"-C", clean_cwd,
]
if not self.full_env:
# Drop unneeded startup work: no built-in (github) MCP server and no
# AGENTS.md / custom-instruction loading. With an isolated home that
# has no mcp-config.json, no user MCP servers spawn either.
cmd += ["--disable-builtin-mcps", "--no-custom-instructions"]
if self.model:
cmd += ["--model", self.model]
env = os.environ.copy()
if self.copilot_home:
env["COPILOT_HOME"] = self.copilot_home
try:
proc = subprocess.run(
cmd, capture_output=True, creationflags=_NO_WINDOW, text=True, timeout=self.timeout, cwd=clean_cwd,
encoding="utf-8", errors="replace", env=env,
)
except Exception:
return ""
finally:
try:
import shutil
shutil.rmtree(clean_cwd, ignore_errors=True)
except Exception:
pass
return self._parse_jsonl_response(proc.stdout or "")
@staticmethod
def _parse_jsonl_response(raw: str) -> str:
parts: List[str] = []
for line in raw.splitlines():
line = line.strip()
if not line or not line.startswith("{"):
continue
try:
obj = json.loads(line)
except Exception:
continue
if obj.get("type") == "assistant.message":
content = (obj.get("data") or {}).get("content")
if isinstance(content, str) and content:
parts.append(content)
return "\n".join(parts).strip()
def attempt_with_tools(self, task, skill, memory, tools):
# Expose REAL, callable tool shims in the working directory so the
# gbrain quick-answerer judge (tool_called=search) is validated
# honestly: we detect each call from the shim's log, not from a
# self-reported marker. The Copilot CLI is the Windows-validated
# backend, so the shims must be cross-platform — a bash `#!/usr/bin/env
# bash` + chmod shim does NOT execute via `./tool` under PowerShell/cmd,
# so on Windows we emit a `.cmd` batch shim instead.
import shutil
import stat
work = tempfile.mkdtemp(prefix="skillopt_sleep_copilottools_")
calllog = os.path.join(work, "_tool_calls.log")
tool_names = tools or ["search"]
is_windows = os.name == "nt"
try:
for tname in tool_names:
if is_windows:
shim = os.path.join(work, f"{tname}.cmd")
with open(shim, "w") as f:
# `%~n0` is the script's own base name (the tool name);
# writing it keeps the calllog line == tool name so the
# honest-detection match below works unchanged.
f.write(
"@echo off\n"
f'echo %~n0>>"{calllog}"\n'
"echo (search results: 3 relevant notes found; use them to answer)\n"
)
else:
shim = os.path.join(work, tname)
with open(shim, "w") as f:
f.write(
"#!/usr/bin/env bash\n"
f'echo "{tname}" >> "{calllog}"\n'
'echo "(search results: 3 relevant notes found; use them to answer)"\n'
)
os.chmod(shim, os.stat(shim).st_mode | stat.S_IEXEC | stat.S_IXGRP | stat.S_IXOTH)
if is_windows:
tool_hint = (
"You have shell tools available in the current directory: "
+ ", ".join(f"{t}.cmd" for t in tool_names)
+ " (each callable as `" + tool_names[0] + "` or `.\\"
+ tool_names[0] + "`). When the skill says to look something "
"up or search before answering, you MUST actually run the "
"tool (e.g. `" + tool_names[0] + " \"query\"`) before giving "
"your final answer."
)
else:
tool_hint = (
"You have shell tools available in the current directory: "
+ ", ".join(f"./{t}" for t in tool_names)
+ ". When the skill says to look something up or search before "
"answering, you MUST actually run the tool (e.g. `./search \"query\"`) "
"before giving your final answer."
)
prompt = (
"You are completing a task. Apply the skill and memory rules EXACTLY, "
"including any rule about searching/looking up before answering. "
"Treat a 'Learned preferences' block as HARD CONSTRAINTS that override "
"earlier conflicting skill text.\n\n"
f"{tool_hint}\n\n"
f"# Skill\n{skill or '(none)'}\n\n# Memory\n{memory or '(none)'}\n\n"
f"# Task\n{task.intent}\n\n{task.context_excerpt}\n\n"
"Return ONLY the final answer text."
)
cmd = [
self.copilot_path, "-p", prompt,
"--output-format", "json",
"--stream", "off",
"--no-color",
"--log-level", "none",
"--allow-all-tools",
"-C", work,
]
if not self.full_env:
cmd += ["--disable-builtin-mcps", "--no-custom-instructions"]
if self.model:
cmd += ["--model", self.model]
env = os.environ.copy()
if self.copilot_home:
env["COPILOT_HOME"] = self.copilot_home
resp = ""
try:
proc = subprocess.run(
cmd, capture_output=True, creationflags=_NO_WINDOW, text=True, encoding="utf-8",
errors="replace", timeout=self.timeout, cwd=work, env=env,
)
resp = self._parse_jsonl_response(proc.stdout or "")
except Exception:
resp = ""
self._tokens += len(prompt) // 4 + len(resp) // 4
called: List[str] = []
if os.path.exists(calllog):
with open(calllog) as f:
logged = {ln.strip() for ln in f if ln.strip()}
called = [t for t in tool_names if t in logged]
return resp, called
finally:
try:
shutil.rmtree(work, ignore_errors=True)
except Exception:
pass
class DualBackend(Backend):
"""Route operations to two backends, à la SkillOpt's target vs optimizer.
@@ -1025,17 +1401,27 @@ def get_backend(
claude_path: str = "claude",
codex_path: str = "",
azure_endpoint: str = "",
project_dir: str = "",
) -> Backend:
n = (name or "mock").strip().lower()
if n in {"claude", "anthropic", "claude_cli", "claude_code"}:
return ClaudeCliBackend(model=model, claude_path=claude_path)
if n in {"codex", "codex_cli", "openai_codex"}:
return CodexCliBackend(model=model, codex_path=codex_path)
return CodexCliBackend(model=model, codex_path=codex_path, project_dir=project_dir)
if n in {"azure", "azure_openai", "aoai"}:
return AzureOpenAIBackend(deployment=model, endpoint=azure_endpoint)
if n in {"azure-responses", "azure_responses", "aoai-responses", "responses"}:
eps = [e.strip() for e in azure_endpoint.split(",") if e.strip()] or None
return AzureResponsesBackend(deployment=model, endpoints=eps)
if n in {"copilot", "github_copilot", "copilot_cli", "gh_copilot"}:
return CopilotCliBackend(model=model)
if n in {"handoff", "session", "file"}:
# Lazy import: handoff_backend imports CliBackend from this module.
from skillopt_sleep.handoff_backend import HandoffBackend
hdir = os.environ.get("SKILLOPT_SLEEP_HANDOFF_DIR", "") or os.path.join(
project_dir or os.getcwd(), ".skillopt-sleep-handoff"
)
return HandoffBackend(model=model, handoff_dir=hdir)
return MockBackend()
@@ -1050,6 +1436,7 @@ def build_backend(
codex_path: str = "",
azure_endpoint: str = "",
preferences: str = "",
project_dir: str = "",
) -> Backend:
"""Build a single or dual backend.
@@ -1060,13 +1447,21 @@ def build_backend(
"""
has_split = any([optimizer_backend, optimizer_model, target_backend, target_model])
if not has_split:
be = get_backend(backend, model=model, codex_path=codex_path, azure_endpoint=azure_endpoint)
be = get_backend(
backend,
model=model,
codex_path=codex_path,
azure_endpoint=azure_endpoint,
project_dir=project_dir,
)
be.preferences = preferences
return be
tgt = get_backend(target_backend or backend, model=target_model or model,
codex_path=codex_path, azure_endpoint=azure_endpoint)
codex_path=codex_path, azure_endpoint=azure_endpoint,
project_dir=project_dir)
opt = get_backend(optimizer_backend or backend, model=optimizer_model or model,
codex_path=codex_path, azure_endpoint=azure_endpoint)
codex_path=codex_path, azure_endpoint=azure_endpoint,
project_dir=project_dir)
opt.preferences = preferences # reflect runs on the optimizer
dual = DualBackend(target=tgt, optimizer=opt)
dual.preferences = preferences
+24 -4
View File
@@ -13,17 +13,19 @@ from __future__ import annotations
import json
import os
from dataclasses import dataclass, field, asdict
from typing import Any, Dict, List, Optional
from dataclasses import dataclass, field
from typing import Any, Dict, Optional
HOME_STATE_DIR = os.path.expanduser("~/.skillopt-sleep")
CLAUDE_HOME = os.path.expanduser("~/.claude")
CODEX_HOME = os.path.expanduser("~/.codex")
DEFAULTS: Dict[str, Any] = {
# ── scope ──────────────────────────────────────────────────────────────
"claude_home": CLAUDE_HOME,
"codex_home": CODEX_HOME,
"transcript_source": "claude", # "claude" | "codex" | "auto"
"projects": "invoked", # "invoked" | "all" | [list of abs paths]
"invoked_project": "", # filled at runtime (cwd) when projects == "invoked"
"lookback_hours": 72, # harvest window when no prior sleep recorded
@@ -34,7 +36,7 @@ DEFAULTS: Dict[str, Any] = {
"val_fraction": 0.34, # real tasks reserved to gate updates
"test_fraction": 0.0, # real tasks reserved as the final held-out measure
# ── optimizer ──────────────────────────────────────────────────────────
"backend": "mock", # "mock" | "claude" | "codex"
"backend": "mock", # "mock" | "claude" | "codex" | "copilot"
"model": "", # backend-specific; "" => backend default
"gate_mode": "on", # "on" (validation-gated) | "off" (greedy, no hard filter)
"codex_path": "", # "" => auto-detect the real @openai/codex binary
@@ -42,9 +44,16 @@ DEFAULTS: Dict[str, Any] = {
"gate_metric": "mixed", # hard | soft | mixed (mixed best for tiny holdouts)
"gate_mixed_weight": 0.5,
"replay_mode": "mock", # "mock" (sandboxed prompt) | "fresh" (worktree)
# ── dream + recall (opt-in; defaults reproduce the prior single-shot loop) ─
"dream_rollouts": 1, # >1 => multi-rollout contrastive reflection per task
"dream_factor": 0, # >0 => add N synthetic variants of each task to the dream
"recall_k": 0, # >0 => recall the K most-similar past tasks into the dream
"evolve_memory": True, # consolidate CLAUDE.md
"evolve_skill": True, # consolidate the managed SKILL.md
"llm_mine": True, # use the backend to mine checkable tasks (real backends)
"target_skill_path": "", # explicit SKILL.md target for repo-scoped agents
"target_task_filter": True, # prefer mined tasks matching target_skill_path/text
"progress": False, # print phase progress to stderr
# ── adoption / safety ──────────────────────────────────────────────────
"auto_adopt": False, # default: stage + require explicit `adopt`
"managed_skill_name": "skillopt-sleep-learned",
@@ -94,6 +103,10 @@ class SleepConfig:
def transcripts_dir(self) -> str:
return os.path.join(self.data["claude_home"], "projects")
@property
def codex_archived_sessions_dir(self) -> str:
return os.path.join(self.data["codex_home"], "archived_sessions")
@property
def history_path(self) -> str:
return os.path.join(self.data["claude_home"], "history.jsonl")
@@ -103,6 +116,13 @@ class SleepConfig:
return os.path.join(self.data["claude_home"], "skills")
def managed_skill_path(self) -> str:
target = self.data.get("target_skill_path") or ""
if target:
target = os.path.expanduser(str(target))
if not os.path.isabs(target):
base = self.data.get("invoked_project") or os.getcwd()
target = os.path.join(base, target)
return os.path.abspath(target)
return os.path.join(
self.skills_dir, self.data["managed_skill_name"], "SKILL.md"
)
+29 -1
View File
@@ -9,7 +9,7 @@ validation gate, vendored self-contained in ``skillopt_sleep.gate``.
from __future__ import annotations
import os
from dataclasses import dataclass
from dataclasses import dataclass, field
from typing import List, Optional, Tuple
from skillopt_sleep.backend import Backend
@@ -36,6 +36,10 @@ class ConsolidationResult:
rejected_edits: List[EditRecord]
holdout_baseline: float
holdout_candidate: float
# ── observability (so a 0.0->0.0 night is self-diagnosing, not a black box) ──
holdout_detail: List[dict] = field(default_factory=list) # per val task: hard/soft/resp/why
reflect_raw: str = "" # the optimizer's last raw reply (empty => reflect produced nothing)
call_error: str = "" # backend's last call error (timeout/auth/empty)
def _split(tasks: List[TaskRecord]) -> Tuple[List[TaskRecord], List[TaskRecord]]:
@@ -61,6 +65,25 @@ def _split(tasks: List[TaskRecord]) -> Tuple[List[TaskRecord], List[TaskRecord]]
return train, val
def _holdout_detail(pairs: List[Tuple[TaskRecord, ReplayResult]]) -> List[dict]:
"""Per-task held-out evidence so a 0.0 night explains itself: was the
response empty (backend call failed) or non-empty-but-failing-checks
(judge too strict / edit didn't help)? The two need opposite fixes."""
out: List[dict] = []
for t, r in pairs:
resp = r.response or ""
out.append({
"id": t.id,
"reference_kind": t.reference_kind,
"hard": r.hard,
"soft": r.soft,
"response_len": len(resp),
"response_head": resp[:200],
"why": (r.fail_reason or r.judge_rationale or "")[:200],
})
return out
def consolidate(
backend: Backend,
tasks: List[TaskRecord],
@@ -87,6 +110,7 @@ def consolidate(
"""
train_tasks, val_tasks = _split(tasks)
gate_off = str(gate_mode).strip().lower() in {"off", "none", "false", "greedy"}
holdout_detail: List[dict] = []
# ── baseline on the VAL slice (the gate reference) ────────────────────
# When the gate is OFF the user has opted out of holding out a validation set
@@ -98,6 +122,7 @@ def consolidate(
else:
base_pairs = replay_batch(backend, val_tasks, skill, memory)
base_hard, base_soft = aggregate_scores(base_pairs)
holdout_detail = _holdout_detail(base_pairs)
base_score = select_gate_score(base_hard, base_soft, gate_metric, gate_mixed_weight)
# ── reflect over TRAIN-split failures/successes ───────────────────────
@@ -235,4 +260,7 @@ def consolidate(
rejected_edits=all_rejected,
holdout_baseline=base_hard,
holdout_candidate=final_hard,
holdout_detail=holdout_detail,
reflect_raw=getattr(backend, "last_reflect_raw", "") or "",
call_error=getattr(backend, "last_call_error", "") or "",
)
+126 -27
View File
@@ -10,18 +10,20 @@ CI use. With backend="anthropic" it spends the user's budget for real lift.
from __future__ import annotations
import os
import time
import sys
from dataclasses import dataclass
from typing import Any, Dict, List, Optional
from typing import List, Optional
from skillopt_sleep.backend import get_backend
from skillopt_sleep.backend import Backend, get_backend
from skillopt_sleep.config import SleepConfig, load_config
from skillopt_sleep.consolidate import consolidate
from skillopt_sleep.harvest import harvest
from skillopt_sleep.dream import dream_consolidate
from skillopt_sleep.harvest_sources import harvest_for_config
from skillopt_sleep.memory import ensure_skill_scaffold
from skillopt_sleep.mine import mine
from skillopt_sleep.staging import adopt as adopt_staging
from skillopt_sleep.staging import redact_secrets
from skillopt_sleep.staging import write_staging
from skillopt_sleep.state import SleepState, _now_iso
from skillopt_sleep.staging import write_staging, adopt as adopt_staging
from skillopt_sleep.types import SessionDigest, SleepReport, TaskRecord
@@ -49,6 +51,11 @@ def _read(path: str) -> str:
return ""
def _progress(cfg: SleepConfig, message: str) -> None:
if cfg.get("progress", False):
print(f"[sleep] {message}", file=sys.stderr, flush=True)
def _render_report_md(report: SleepReport, cfg: SleepConfig) -> str:
lines = [
f"# SkillOpt-Sleep — night {report.night} report",
@@ -87,6 +94,7 @@ def run_sleep_cycle(
seed_tasks: Optional[List[TaskRecord]] = None,
dry_run: bool = False,
clock: Optional[float] = None,
backend: Optional[Backend] = None,
) -> CycleOutcome:
"""Run one full sleep cycle and return the outcome.
@@ -97,6 +105,8 @@ def run_sleep_cycle(
inject a known persona instead of harvesting ~/.claude).
dry_run : harvest+mine+replay but DO NOT stage/adopt (report only).
clock : fixed epoch seconds for deterministic timestamps in tests.
backend : optional pre-built Backend; the handoff driver passes one so
it can inspect the backend's pending calls after the run.
"""
cfg = cfg or load_config()
state = SleepState.load(cfg.state_path)
@@ -104,10 +114,30 @@ def run_sleep_cycle(
project = _project_paths(cfg)
started = _now_iso(clock)
backend = get_backend(
backend = backend or get_backend(
cfg.get("backend", "mock"),
model=cfg.get("model", ""),
codex_path=cfg.get("codex_path", ""),
project_dir=project,
)
_progress(cfg, f"night {night}: project={project} backend={backend.name}")
# ── live skill/memory docs ───────────────────────────────────────────
live_memory_path = os.path.join(project, "CLAUDE.md")
live_skill_path = cfg.managed_skill_path()
_progress(cfg, f"live skill: {live_skill_path}")
raw_skill = _read(live_skill_path)
skill = raw_skill
memory = _read(live_memory_path)
if not skill:
skill = ensure_skill_scaffold(
"", name=cfg.get("managed_skill_name", "skillopt-sleep-learned"),
description="Preferences and procedures learned from past local agent sessions.",
)
target_filter = bool(
cfg.get("target_task_filter", True)
and cfg.get("target_skill_path", "")
and raw_skill
)
# ── 1+2. harvest + mine (unless seed_tasks injected) ─────────────────
@@ -115,16 +145,34 @@ def run_sleep_cycle(
if seed_tasks is not None:
tasks = seed_tasks
n_sessions = 0
_progress(cfg, f"using {len(tasks)} seeded tasks")
else:
since = state.last_harvest_for(project)
digests = harvest(
cfg.transcripts_dir,
scope=cfg.get("projects", "invoked"),
invoked_project=cfg.get("invoked_project", ""),
# On first run (no prior harvest), apply lookback_hours so we don't
# scan the entire transcript history and trigger massive LLM mining.
if since is None:
lookback_hours = cfg.get("lookback_hours", 72)
if lookback_hours is not None and lookback_hours > 0:
import time
ref_time = clock if clock is not None else time.time()
cutoff = ref_time - lookback_hours * 3600
since = _now_iso(cutoff)
max_tasks = cfg.get("max_tasks_per_night", 40)
max_sessions = cfg.get("max_sessions_per_night", 0) or max_tasks * 3
candidate_limit = max_tasks
if target_filter:
candidate_limit = max(max_tasks, max_tasks * 3)
_progress(
cfg,
f"harvest start: source={cfg.get('transcript_source')} max_sessions={max_sessions}",
)
digests = harvest_for_config(
cfg,
since_iso=since,
limit=cfg.get("max_tasks_per_night", 40) * 3,
limit=max_sessions,
)
n_sessions = len(digests)
_progress(cfg, f"harvest done: sessions={n_sessions}")
# When a real backend is configured, use it to mine checkable tasks from
# the transcripts (rubric/rule judges); otherwise fall back to the
# heuristic miner (no API, no checkable reference).
@@ -132,27 +180,29 @@ def run_sleep_cycle(
if cfg.get("backend", "mock") != "mock" and cfg.get("llm_mine", True):
try:
from skillopt_sleep.llm_miner import make_llm_miner
llm_miner = make_llm_miner(backend, max_tasks=cfg.get("max_tasks_per_night", 40))
llm_miner = make_llm_miner(
backend,
max_sessions=max_sessions,
max_tasks=candidate_limit,
)
except Exception:
llm_miner = None
_progress(
cfg,
f"mine start: max_tasks={max_tasks} candidate_limit={candidate_limit} "
f"llm_mine={llm_miner is not None} target_filter={target_filter}",
)
tasks = mine(
digests,
max_tasks=cfg.get("max_tasks_per_night", 40),
max_tasks=max_tasks,
candidate_limit=candidate_limit,
holdout_fraction=cfg.get("holdout_fraction", 0.34),
seed=cfg.get("seed", 42),
llm_miner=llm_miner,
target_skill_text=raw_skill if target_filter else "",
target_skill_path=live_skill_path if target_filter else "",
)
# ── live skill/memory docs ───────────────────────────────────────────
live_memory_path = os.path.join(project, "CLAUDE.md")
live_skill_path = cfg.managed_skill_path()
skill = _read(live_skill_path)
memory = _read(live_memory_path)
if not skill:
skill = ensure_skill_scaffold(
"", name=cfg.get("managed_skill_name", "skillopt-sleep-learned"),
description="Preferences and procedures learned from past Claude Code sessions.",
)
_progress(cfg, f"mine done: tasks={len(tasks)}")
report = SleepReport(
night=night, project=project, started_at=started,
@@ -169,9 +219,22 @@ def run_sleep_cycle(
staging_dir = ""
return CycleOutcome(report, staging_dir, False, [])
# ── 3+4. replay + consolidate (gate) ─────────────────────────────────
result = consolidate(
# ── 3+4. replay + consolidate (gate), with opt-in dream + recall ──────
# recall pulls similar past tasks from the persisted archive; dream_rollouts
# / dream_factor enrich the training signal. With the defaults (recall_k=0,
# dream_rollouts=1, dream_factor=0) this is exactly the prior single-shot
# consolidate — behavior is unchanged unless the user opts in.
_progress(cfg, "consolidate start")
recall_k = int(cfg.get("recall_k", 0) or 0)
history_tasks = []
if recall_k > 0:
history_tasks = [TaskRecord.from_dict(d) for d in state.task_archive()]
result = dream_consolidate(
backend, tasks, skill, memory,
history_tasks=history_tasks,
recall_k=recall_k,
dream_rollouts=int(cfg.get("dream_rollouts", 1) or 1),
dream_factor=int(cfg.get("dream_factor", 0) or 0),
edit_budget=cfg.get("edit_budget", 4),
gate_metric=cfg.get("gate_metric", "mixed"),
gate_mixed_weight=cfg.get("gate_mixed_weight", 0.5),
@@ -180,12 +243,20 @@ def run_sleep_cycle(
evolve_memory=cfg.get("evolve_memory", True),
night=night,
)
# archive tonight's real (non-dream) tasks so future nights can recall them
state.add_to_archive([t.to_dict() for t in tasks if t.origin != "dream"])
_progress(
cfg,
f"consolidate done: gate={result.gate_action} accepted={result.accepted} "
f"edits={len(result.applied_edits)} rejected={len(result.rejected_edits)}",
)
report.n_replayed = len(tasks)
report.baseline_score = result.baseline_score
report.candidate_score = result.candidate_score
report.accepted = result.accepted
report.gate_action = result.gate_action
report.no_edits_reason = getattr(result, "no_edits_reason", "")
report.edits = result.applied_edits
report.rejected_edits = result.rejected_edits
report.tokens_used = backend.tokens_used()
@@ -196,6 +267,7 @@ def run_sleep_cycle(
adopted = False
adopted_paths: List[str] = []
if not dry_run:
_progress(cfg, "staging start")
report_md = _render_report_md(report, cfg)
proposed_skill = result.new_skill if (cfg.get("evolve_skill") and result.accepted) else None
proposed_memory = result.new_memory if (cfg.get("evolve_memory") and result.accepted) else None
@@ -208,6 +280,33 @@ def run_sleep_cycle(
live_memory_path=live_memory_path,
report_md=report_md,
)
# Observability: persist per-task held-out evidence + optimizer/codex errors so a
# 0.0->0.0 night self-explains (empty responses vs failing checks vs no edits) — the
# cycle previously captured none of this, making the gate a black box (#learning-stall).
try:
import json as _json
# Backend stderr / optimizer replies / task responses can carry
# credentials (e.g. a codex 401 stderr dump), so scrub secret-looking
# substrings before persisting them to the on-disk diagnostics.
with open(os.path.join(staging_dir, "diagnostics.json"), "w", encoding="utf-8") as _fh:
_json.dump({
"night": night,
"backend": cfg.get("backend"),
"gate_mode": cfg.get("gate_mode"),
"n_tasks": len(tasks),
"baseline_score": result.baseline_score,
"candidate_score": result.candidate_score,
"accepted": result.accepted,
"n_applied_edits": len(result.applied_edits),
"n_rejected_edits": len(result.rejected_edits),
"call_error": redact_secrets(getattr(result, "call_error", "")),
"reflect_raw_head": redact_secrets(
(getattr(result, "reflect_raw", "") or "")[:1200]
),
"holdout_detail": redact_secrets(getattr(result, "holdout_detail", [])),
}, _fh, indent=2)
except Exception:
pass
state.set_last_harvest(project, started)
state.record_night({
"night": night, "accepted": result.accepted,
+138
View File
@@ -0,0 +1,138 @@
"""SkillOpt-Sleep — dream + associative recall for nightly consolidation.
Two opt-in mechanisms (both default OFF, so the cycle is unchanged unless the
user enables them) that the deployment experiments validated:
* dream rollouts run each task K times and learn from the good-vs-bad
contrast (set ``dream_rollouts > 1``). Stronger signal than one failure.
* associative recall each night, pull the K past tasks most similar to
tonight's new ones into the dream (set ``recall_k > 0``). Replays relevant
experience without re-running the whole history.
``dream_consolidate`` wires recall + synthetic augmentation + multi-rollout
consolidation and is called by BOTH the shipped plugin cycle and the benchmark
experiment harness, so the reported numbers exercise the exact code the plugin
runs. Pure-stdlib, zero research/private dependency.
"""
from __future__ import annotations
import re
from typing import List, Optional
from skillopt_sleep.consolidate import ConsolidationResult, consolidate
from skillopt_sleep.types import TaskRecord
# ── synthetic augmentation ("dream up" variants of today's tasks) ─────────────
_WRAPPERS = [
"(quick one) {q}",
"Please handle this request: {q}",
"For the daily report: {q}",
]
def dream_augment(real_tasks: List[TaskRecord], *, factor: int = 1) -> List[TaskRecord]:
"""Create synthetic TRAIN variants of real tasks (origin='dream').
A light, deterministic rephrasing. Dream tasks are training-only they
carry split='train' and never enter the val/test slices the gate scores on.
"""
out: List[TaskRecord] = []
for t in real_tasks:
for k in range(max(0, factor)):
w = _WRAPPERS[k % len(_WRAPPERS)]
out.append(TaskRecord(
id=f"{t.id}_dream{k}", project=t.project,
intent=w.format(q=t.intent), context_excerpt=t.context_excerpt,
reference_kind=t.reference_kind, reference=t.reference,
judge=dict(t.judge), system=t.system,
tags=list(t.tags) + ["dream"], split="train",
origin="dream", derived_from=t.id,
))
return out
# ── associative recall (experience replay of similar past tasks) ──────────────
def _tokens(text: str) -> set:
return {w for w in re.findall(r"[a-z0-9]+", (text or "").lower()) if len(w) > 2}
def recall_similar(new_tasks: List[TaskRecord], history: List[TaskRecord],
k: int) -> List[TaskRecord]:
"""Return the ``k`` historical tasks most lexically similar to any of
tonight's ``new_tasks`` (max Jaccard token overlap). Recalled tasks are
returned as training material (split='train'); deterministic, stdlib-only.
"""
if not history or k <= 0 or not new_tasks:
return []
new_tok = [_tokens(t.intent) for t in new_tasks]
new_ids = {t.id for t in new_tasks}
scored = []
for h in history:
if h.id in new_ids:
continue
ht = _tokens(h.intent)
if not ht:
continue
sim = max(((len(ht & nt) / len(ht | nt)) if (ht | nt) else 0.0) for nt in new_tok)
scored.append((sim, h.id, h))
scored.sort(key=lambda x: (-x[0], x[1]))
out = []
for sim, _id, h in scored[:max(0, k)]:
if sim <= 0.0:
break
# recall as training material; copy so the source archive is untouched
out.append(TaskRecord(
id=f"recall:{h.id}", project=h.project, intent=h.intent,
context_excerpt=h.context_excerpt, reference_kind=h.reference_kind,
reference=h.reference, judge=dict(h.judge), system=h.system,
tags=list(h.tags) + ["recall"], split="train", origin="real",
derived_from=h.id,
))
return out
# ── the shared nightly consolidation step ─────────────────────────────────────
def dream_consolidate(
backend,
tasks: List[TaskRecord],
skill: str,
memory: str,
*,
history_tasks: Optional[List[TaskRecord]] = None,
recall_k: int = 0,
dream_rollouts: int = 1,
dream_factor: int = 0,
edit_budget: int = 4,
gate_metric: str = "mixed",
gate_mixed_weight: float = 0.5,
gate_mode: str = "on",
evolve_skill: bool = True,
evolve_memory: bool = True,
night: int = 1,
) -> ConsolidationResult:
"""Recall similar past experience + dream synthetic variants, then run one
gated consolidation epoch over the enlarged training pool.
``tasks`` is the split-tagged pool for tonight (train + val); recall and
augmentation only enlarge the TRAIN split, so the val slice the gate scores
on is never polluted. With ``recall_k=0`` and ``dream_rollouts=1`` (the
defaults) this is exactly the previous single-shot ``consolidate``.
"""
train = [t for t in tasks if t.split == "train"]
enlarged = list(tasks)
if recall_k > 0 and history_tasks:
enlarged += recall_similar(train, history_tasks, recall_k)
if dream_factor > 0:
seed = [t for t in enlarged if t.split == "train" and t.origin != "dream"]
enlarged += dream_augment(seed, factor=dream_factor)
return consolidate(
backend, enlarged, skill, memory,
edit_budget=edit_budget, gate_metric=gate_metric,
gate_mixed_weight=gate_mixed_weight, gate_mode=gate_mode,
rollouts_k=dream_rollouts, evolve_skill=evolve_skill,
evolve_memory=evolve_memory, night=night,
)
+1 -1
View File
@@ -134,7 +134,7 @@ def main(argv=None) -> int:
ap = argparse.ArgumentParser(description="SkillOpt-Sleep validation experiment")
ap.add_argument("--persona", default="researcher", choices=list(PERSONAS.keys()))
ap.add_argument("--nights", type=int, default=4)
ap.add_argument("--backend", default="mock", choices=["mock", "claude", "codex"])
ap.add_argument("--backend", default="mock", choices=["mock", "claude", "codex", "copilot"])
ap.add_argument("--model", default="", help="backend model override")
ap.add_argument("--codex-path", default="", help="path to the real @openai/codex binary")
ap.add_argument("--edit-budget", type=int, default=4)
+173
View File
@@ -0,0 +1,173 @@
"""SkillOpt-Sleep — handoff backend (session-executed model calls).
Runs the sleep cycle WITHOUT spawning any model subprocess or API call.
Every intelligent operation (attempt / judge / reflect) is turned into a
prompt file that an interactive agent session answers between engine runs:
run 1: the engine executes the deterministic stages; every model call
it needs is recorded as a pending prompt; the run stops and
writes PROMPTS.md + pending.json into the handoff directory.
you: answer each prompt (each in a FRESH context, so the session's
own history cannot contaminate the held-out gate) and write the
raw answer text to answers/<id>.md.
run 2: the engine re-runs; answered prompts resolve from answers/, the
cycle advances to the next model-dependent stage, and either
finishes or writes the next PROMPTS.md batch.
Resume needs no serialized engine state: harvest -> mine -> replay is
deterministic, so re-running regenerates identical prompts and the answers
directory acts as a persistent, cross-run call cache. A prompt that embeds
a still-unanswered response (detected via the pending sentinel) aborts the
run immediately so placeholder text never propagates into scores, edits,
or staging. A typical night converges in 3-6 rounds: baseline attempts ->
reflect -> candidate re-scoring per accepted edit.
Limitations (v1): `dream_rollouts > 1` yields no contrastive spread (the
same prompt maps to the same answer file), and tool-loop tasks fall back
to the base single-shot 'TOOL_CALL: <name>' marker convention.
"""
from __future__ import annotations
import json
import os
import threading
from typing import Dict
from skillopt_sleep.backend import CliBackend, skill_hash
PENDING_SENTINEL_PREFIX = "[[SKILLOPT-SLEEP-PENDING:"
PENDING_SENTINEL_SUFFIX = "]]"
# reflect() appends this when a reply fails to parse; with a placeholder
# reply the retry is a dependent call, not a genuinely new question.
_REFLECT_RETRY_MARKER = "your previous reply was not valid JSON"
PROMPTS_FILENAME = "PROMPTS.md"
PENDING_FILENAME = "pending.json"
class PendingCalls(RuntimeError):
"""The cycle cannot advance until pending prompts are answered."""
def __init__(self, pending: Dict[str, Dict[str, object]]):
self.pending = dict(pending)
super().__init__(
f"{len(self.pending)} model call(s) awaiting handoff answers"
)
class HandoffBackend(CliBackend):
"""Backend that outsources every model call to prompt/answer files.
``_call`` resolves a prompt from ``answers/<sha256[:16]>.md`` when the
answer exists; otherwise it records the prompt as pending and returns a
sentinel placeholder so independent calls in the same phase can still
be collected into one batch. Any call whose prompt was BUILT FROM a
placeholder raises :class:`PendingCalls` that call depends on answers
the user has not provided yet, so continuing would only mint garbage.
"""
name = "handoff"
def __init__(self, model: str = "", handoff_dir: str = "") -> None:
super().__init__(model=model, timeout=0)
self.handoff_dir = os.path.abspath(
handoff_dir or os.path.join(os.getcwd(), ".skillopt-sleep-handoff")
)
self.answers_dir = os.path.join(self.handoff_dir, "answers")
os.makedirs(self.answers_dir, exist_ok=True)
# key -> {"prompt": str, "max_tokens": int}, insertion-ordered
self.pending: Dict[str, Dict[str, object]] = {}
self._lock = threading.Lock()
# ── prompt/answer plumbing ────────────────────────────────────────────
def answer_path(self, key: str) -> str:
return os.path.join(self.answers_dir, f"{key}.md")
def _call(self, prompt: str, *, max_tokens: int = 1024) -> str:
if PENDING_SENTINEL_PREFIX in prompt:
# Built from a still-pending response — dependent call.
raise PendingCalls(self.pending)
if _REFLECT_RETRY_MARKER in prompt and self.pending:
# Retry of a reflect whose first reply is the placeholder.
raise PendingCalls(self.pending)
key = skill_hash(prompt)
path = self.answer_path(key)
if os.path.exists(path):
with open(path, encoding="utf-8") as f:
return f.read().strip()
with self._lock:
self.pending[key] = {"prompt": prompt, "max_tokens": max_tokens}
return f"{PENDING_SENTINEL_PREFIX}{key}{PENDING_SENTINEL_SUFFIX}"
# ── handoff file emission ─────────────────────────────────────────────
def flush_pending(self) -> str:
"""Write PROMPTS.md (human/agent-readable) + pending.json (machine).
Prompts can themselves contain markdown fences, so PROMPTS.md
delimits each prompt with BEGIN/END marker lines instead of fences.
Returns the PROMPTS.md path.
"""
from skillopt_sleep.staging import redact_secrets
os.makedirs(self.handoff_dir, exist_ok=True)
with self._lock:
items = list(self.pending.items())
payload = {
"format": "skillopt_sleep.handoff.v1",
"answers_dir": self.answers_dir,
"pending": [
{
"id": key,
"answer_file": self.answer_path(key),
"max_tokens": item["max_tokens"],
"prompt": redact_secrets(str(item["prompt"])),
}
for key, item in items
],
}
with open(os.path.join(self.handoff_dir, PENDING_FILENAME), "w",
encoding="utf-8") as f:
json.dump(payload, f, ensure_ascii=False, indent=2)
f.write("\n")
lines = [
"# SkillOpt-Sleep — pending model calls (handoff)",
"",
f"{len(items)} prompt(s) below need answers before the sleep "
"cycle can continue.",
"",
"For EACH prompt:",
"",
"1. Answer it in a FRESH context (e.g. a subagent with no",
" conversation history). Do NOT let the current session's",
" context, the other prompts in this file, or the optimization",
" run itself leak into the answer — that contaminates the",
" held-out validation gate.",
"2. Write ONLY the raw answer text (no commentary, no code",
" fences) to the prompt's answer file.",
"",
"When every answer file exists, re-run the same engine command",
"(`python -m skillopt_sleep run --backend handoff ...`); it",
"resumes automatically from the answers directory.",
"",
]
for i, (key, item) in enumerate(items, start=1):
lines += [
"---",
"",
f"## Prompt {i} of {len(items)}",
"",
f"- id: `{key}`",
f"- answer file: `answers/{key}.md`",
f"- suggested max tokens: {item['max_tokens']}",
"",
f"----- BEGIN PROMPT {key} -----",
redact_secrets(str(item["prompt"])),
f"----- END PROMPT {key} -----",
"",
]
prompts_path = os.path.join(self.handoff_dir, PROMPTS_FILENAME)
with open(prompts_path, "w", encoding="utf-8") as f:
f.write("\n".join(lines))
return prompts_path
+97 -1
View File
@@ -17,6 +17,7 @@ from __future__ import annotations
import json
import os
from datetime import datetime
from typing import Any, Dict, Iterable, List, Optional
from skillopt_sleep.types import SessionDigest
@@ -108,6 +109,90 @@ def _is_meta_prompt(text: str) -> bool:
return True
if t.startswith("[Pasted text") or t.startswith("Caveat:"):
return True
# Expanded slash-command prompts (Claude Code injects the full command
# body as a user message, wrapped in <command-message>/<command-name>
# tags or rendered as "# /<name> — ..." headers). Not a user intent.
if "<command-message>" in t[:200] or "<command-name>" in t[:200]:
return True
if t.startswith("# /"):
return True
return False
# ── Issue #62: filter headless replay sessions ─────────────────────────
# Prompt markers generated by the engine's own headless `claude -p` calls
# (judge, reflect, attempt). If the sole user prompt in a single-turn
# session matches any of these, the session is engine-generated, not a
# real user task.
_REPLAY_PROMPT_MARKERS = (
"## CURRENT SKILL",
"## FAILED TASKS",
"## SUCCESSFUL TASKS",
"## OUTPUT FORMAT",
"You are a strict grader",
"Score the response 0.0-1.0",
"You are SkillOpt-Sleep",
"## TASK\n",
"## SKILL\n",
)
# Sessions written by OTHER tools' sub-agents (memory observers, critic
# sub-agents, plugin self-invocations). These are multi-turn, so the
# single-turn heuristic in _is_headless_replay never catches them. If the
# FIRST user prompt matches any marker, the whole session is
# machine-generated, not a real user task.
_AGENT_SESSION_MARKERS = (
"You are a Claude-Mem", # claude-mem observer agent
"Hello memory agent", # claude-mem observer continuation
"You are driving **SkillOpt-Sleep**", # this plugin's own command body
) + tuple(
# users can add markers for their own tools' agent prompts
p.strip()
for p in os.environ.get("SKILLOPT_SLEEP_AGENT_MARKERS", "").split(",")
if p.strip()
)
def _is_agent_session(digest: "SessionDigest") -> bool:
"""Detect transcripts written by other tools' sub-agents (see markers)."""
if not digest.user_prompts:
return False
first = digest.user_prompts[0]
return any(marker in first for marker in _AGENT_SESSION_MARKERS)
def _is_headless_replay(digest: "SessionDigest") -> bool:
"""Detect sessions created by the engine's own headless replay calls.
Heuristics (conservatively applied):
1. Session has exactly 1 user turn AND
2. The sole prompt matches engine-generated patterns (grader/reflect),
OR the session lasted < 3 seconds (programmatic, not interactive).
Multi-turn sessions are always kept (interactive by definition).
"""
if digest.n_user_turns > 1:
return False
if digest.n_user_turns == 0:
return True
prompt = digest.user_prompts[0] if digest.user_prompts else ""
for marker in _REPLAY_PROMPT_MARKERS:
if marker in prompt:
return True
# Sub-3-second single-turn sessions with short prompts are almost
# certainly programmatic (engine grader/judge calls). We require the
# prompt to also be short (<200 chars) to avoid false-positives on
# real one-shot questions that Claude happens to answer quickly.
if digest.started_at and digest.ended_at and len(prompt) < 200:
try:
fmt = "%Y-%m-%dT%H:%M:%S"
start = datetime.strptime(digest.started_at[:19], fmt)
end = datetime.strptime(digest.ended_at[:19], fmt)
if (end - start).total_seconds() < 3:
return True
except (ValueError, TypeError):
pass
return False
@@ -226,8 +311,12 @@ def harvest(
paths: List[str] = []
for root, _dirs, files in os.walk(transcripts_dir):
# Sub-agent sidechain transcripts (<session>/subagents/agent-*.jsonl)
# are Claude-authored prompts, not user tasks — never harvest them.
if os.path.basename(root) == "subagents":
continue
for fn in files:
if fn.endswith(".jsonl"):
if fn.endswith(".jsonl") and not fn.startswith("agent-"):
paths.append(os.path.join(root, fn))
# newest first by mtime
paths.sort(key=lambda p: os.path.getmtime(p), reverse=True)
@@ -236,9 +325,16 @@ def harvest(
d = digest_transcript(p)
if d is None:
continue
if _is_headless_replay(d):
continue # Issue #62: skip engine's own headless replay sessions
if _is_agent_session(d):
continue # skip other tools' sub-agent transcripts (claude-mem etc.)
if not _project_matches(d.project or "", scope, invoked_project):
continue
if since_iso and d.ended_at and d.ended_at < since_iso:
# Note: files are sorted by mtime but we compare the embedded
# ended_at timestamp — mtime can diverge (copy/touch), so we
# cannot break here; we must continue to check all files.
continue
digests.append(d)
if limit and len(digests) >= limit:
+233
View File
@@ -0,0 +1,233 @@
"""SkillOpt-Sleep Codex Desktop session harvesting.
Reads Codex Desktop archived session JSONL files and normalizes them into
``SessionDigest`` records without copying developer/system instructions, tool
arguments, or raw tool outputs.
"""
from __future__ import annotations
import os
import re
from typing import Any, Dict, Iterable, List, Optional
from skillopt_sleep.harvest import (
_detect_feedback,
_is_meta_prompt,
_iter_jsonl,
_project_matches,
)
from skillopt_sleep.staging import _SECRET_PATTERNS
from skillopt_sleep.types import SessionDigest
def _payload(rec: Dict[str, Any]) -> Dict[str, Any]:
payload = rec.get("payload")
return payload if isinstance(payload, dict) else {}
def _timestamp(rec: Dict[str, Any], payload: Dict[str, Any]) -> str:
for value in (
payload.get("timestamp"),
rec.get("timestamp"),
payload.get("started_at"),
payload.get("completed_at"),
):
if isinstance(value, str) and value:
return value
return ""
def _text_from_any(content: Any) -> str:
if isinstance(content, str):
return content
if isinstance(content, list):
parts: List[str] = []
for item in content:
if isinstance(item, str):
parts.append(item)
elif isinstance(item, dict):
if item.get("type") == "text" and item.get("text"):
parts.append(str(item["text"]))
elif item.get("text"):
parts.append(str(item["text"]))
return "\n".join(parts)
if isinstance(content, dict):
if content.get("text"):
return str(content["text"])
if content.get("content"):
return _text_from_any(content["content"])
return ""
def _strip_codex_meta(text: str) -> str:
stripped = text.strip()
if not stripped:
return ""
if stripped.startswith("<codex_internal_context"):
return ""
if stripped.startswith("<environment_context"):
return ""
if stripped.startswith("# AGENTS.md instructions") or "--- project-doc ---" in stripped:
for marker in ("</environment_context>", "</INSTRUCTIONS>"):
idx = stripped.rfind(marker)
if idx == -1:
continue
tail = stripped[idx + len(marker):].strip()
if tail and not tail.startswith("<"):
return tail
return ""
return stripped
def _sanitize_text(text: str) -> str:
sanitized = _strip_codex_meta(text).replace("\x00", "").strip()
if not sanitized or _is_meta_prompt(sanitized):
return ""
for pattern, replacement in _SECRET_PATTERNS:
sanitized = pattern.sub(replacement, sanitized)
return sanitized
def _sanitize_tool_name(name: str) -> str:
return re.sub(r"[^A-Za-z0-9_.:-]+", "_", name)[:80]
def _tool_name(payload: Dict[str, Any]) -> str:
payload_type = payload.get("type")
name = payload.get("name")
if isinstance(name, str) and name:
return _sanitize_tool_name(name)
if payload_type == "exec_command_end":
return "exec_command"
if payload_type == "patch_apply_end":
return "apply_patch"
if payload_type == "web_search_call":
return "web_search"
if payload_type == "tool_search_call":
return "tool_search"
if isinstance(payload_type, str) and payload_type.endswith("_tool_call"):
return _sanitize_tool_name(payload_type)
return ""
def _dedup(xs: Iterable[str]) -> List[str]:
seen = set()
out: List[str] = []
for x in xs:
if x not in seen:
seen.add(x)
out.append(x)
return out
def digest_codex_archived_session(path: str, project: str = "") -> Optional[SessionDigest]:
"""Build a ``SessionDigest`` from one Codex Desktop archived session."""
session_id = os.path.splitext(os.path.basename(path))[0]
started = ""
ended = ""
session_project = ""
user_prompts: List[str] = []
assistant_finals: List[str] = []
tools: List[str] = []
feedback: List[str] = []
n_user = 0
n_asst = 0
for rec in _iter_jsonl(path):
payload = _payload(rec)
payload_type = payload.get("type")
ts = _timestamp(rec, payload)
if ts:
if not started:
started = ts
ended = ts
cwd = payload.get("cwd")
if isinstance(cwd, str) and cwd:
if not session_project:
session_project = cwd
if project and _project_matches(cwd, "invoked", project):
session_project = cwd
role = payload.get("role")
text = ""
output_role = ""
if payload_type == "user_message":
text = _text_from_any(payload.get("message"))
output_role = "user"
elif payload_type == "agent_message":
text = _text_from_any(payload.get("message"))
output_role = "assistant"
elif payload_type == "message" and role in {"user", "assistant"}:
text = _text_from_any(payload.get("content"))
output_role = str(role)
else:
tool = _tool_name(payload)
if tool:
tools.append(tool)
continue
sanitized = _sanitize_text(text)
if not sanitized:
continue
if output_role == "user":
n_user += 1
user_prompts.append(sanitized)
feedback.extend(_detect_feedback(sanitized))
elif output_role == "assistant":
n_asst += 1
assistant_finals.append(sanitized)
if project and not _project_matches(session_project or "", "invoked", project):
return None
if n_user == 0 and n_asst == 0:
return None
return SessionDigest(
session_id=session_id,
project=session_project,
started_at=started,
ended_at=ended,
user_prompts=user_prompts,
assistant_finals=assistant_finals[-5:],
tools_used=_dedup(tools),
files_touched=[],
feedback_signals=feedback,
n_user_turns=n_user,
n_assistant_turns=n_asst,
raw_path=path,
)
def harvest_codex(
archived_sessions_dir: str,
*,
scope: Any = "all",
invoked_project: str = "",
since_iso: Optional[str] = None,
limit: int = 0,
) -> List[SessionDigest]:
"""Walk ``~/.codex/archived_sessions`` and return matching digests."""
digests: List[SessionDigest] = []
if not os.path.isdir(archived_sessions_dir):
return digests
paths = [
os.path.join(archived_sessions_dir, fn)
for fn in os.listdir(archived_sessions_dir)
if fn.endswith(".jsonl")
]
paths.sort(key=lambda p: os.path.getmtime(p), reverse=True)
project_hint = invoked_project if scope == "invoked" else ""
for path in paths:
digest = digest_codex_archived_session(path, project=project_hint)
if digest is None:
continue
if not _project_matches(digest.project or "", scope, invoked_project):
continue
if since_iso and digest.ended_at and digest.ended_at < since_iso:
continue
digests.append(digest)
if limit and len(digests) >= limit:
break
return digests
+41
View File
@@ -0,0 +1,41 @@
"""Source selection for SkillOpt-Sleep transcript harvesting."""
from __future__ import annotations
from typing import Optional
from skillopt_sleep.harvest import harvest
from skillopt_sleep.harvest_codex import harvest_codex
from skillopt_sleep.types import SessionDigest
def harvest_for_config(cfg, *, since_iso: Optional[str] = None, limit: int = 0) -> list[SessionDigest]:
source = cfg.get("transcript_source", "claude")
scope = cfg.get("projects", "invoked")
invoked_project = cfg.get("invoked_project", "")
if source == "codex":
return harvest_codex(
cfg.codex_archived_sessions_dir,
scope=scope,
invoked_project=invoked_project,
since_iso=since_iso,
limit=limit,
)
if source == "auto":
codex_digests = harvest_codex(
cfg.codex_archived_sessions_dir,
scope=scope,
invoked_project=invoked_project,
since_iso=since_iso,
limit=limit,
)
if codex_digests:
return codex_digests
return harvest(
cfg.transcripts_dir,
scope=scope,
invoked_project=invoked_project,
since_iso=since_iso,
limit=limit,
)
+9 -10
View File
@@ -12,7 +12,6 @@ from typing import List, Tuple
from skillopt_sleep.types import EditRecord
LEARNED_START = "<!-- SKILLOPT-SLEEP:LEARNED START -->"
LEARNED_END = "<!-- SKILLOPT-SLEEP:LEARNED END -->"
_BANNER = (
@@ -79,7 +78,7 @@ def apply_edits(doc: str, edits: List[EditRecord]) -> Tuple[str, List[EditRecord
anchor substring.
"""
lines = current_learned_lines(doc)
norm_set = {_norm(l) for l in lines}
norm_set = {_norm(line) for line in lines}
applied: List[EditRecord] = []
for e in edits:
@@ -92,31 +91,31 @@ def apply_edits(doc: str, edits: List[EditRecord]) -> Tuple[str, List[EditRecord
applied.append(e)
elif op == "delete":
anchor = _norm(e.anchor or e.content)
keep = [l for l in lines if anchor not in _norm(l)]
keep = [line for line in lines if anchor not in _norm(line)]
if len(keep) != len(lines):
lines = keep
norm_set = {_norm(l) for l in lines}
norm_set = {_norm(line) for line in lines}
applied.append(e)
elif op == "replace":
anchor = _norm(e.anchor)
new_lines = []
changed = False
for l in lines:
if anchor and anchor in _norm(l):
for line in lines:
if anchor and anchor in _norm(line):
new_lines.append(e.content.strip())
changed = True
else:
new_lines.append(l)
new_lines.append(line)
if changed:
lines = new_lines
norm_set = {_norm(l) for l in lines}
norm_set = {_norm(line) for line in lines}
applied.append(e)
return set_learned(doc, lines), applied
def ensure_skill_scaffold(doc: str, *, name: str, description: str) -> str:
"""Ensure a SKILL.md has YAML frontmatter so Claude Code loads it."""
"""Ensure a SKILL.md has YAML frontmatter so local agents load it."""
if doc.lstrip().startswith("---"):
return doc
fm = (
@@ -125,6 +124,6 @@ def ensure_skill_scaffold(doc: str, *, name: str, description: str) -> str:
f"description: {description}\n"
"---\n\n"
f"# {name}\n\n"
"Preferences and procedures learned from your past Claude Code sessions.\n"
"Preferences and procedures learned from your past local agent sessions.\n"
)
return fm + doc
+104 -2
View File
@@ -15,8 +15,10 @@ basis of the deterministic experiment.
from __future__ import annotations
import hashlib
import os
import re
from typing import Any, Callable, List, Optional
from collections import Counter
from typing import Any, Callable, List, Optional, Set, Tuple
from skillopt_sleep.types import SessionDigest, TaskRecord
@@ -39,6 +41,99 @@ def _looks_positive(signals: List[str]) -> bool:
return any(s.startswith("pos:") for s in signals)
_TARGET_STOPWORDS = {
"about", "after", "again", "agent", "agents", "all", "also", "always",
"and", "any", "are", "before", "being", "but", "can", "codex",
"current", "default", "docs", "does", "done", "each", "file", "files",
"for", "from", "have", "into", "keep", "must", "not", "only", "path",
"paths", "project", "read", "repo", "request", "requests", "rule",
"rules", "same", "should", "skill", "skills", "source", "start",
"task", "tasks", "that", "the", "their", "then", "this", "unless",
"update", "user", "users", "when", "with", "work", "workflow",
}
def _target_tokens(text: str) -> List[str]:
tokens: List[str] = []
for raw in re.findall(r"[\w][\w.-]*", (text or "").lower(), flags=re.UNICODE):
parts = [raw] + re.split(r"[\W_]+", raw, flags=re.UNICODE)
for part in parts:
if len(part) < 3 or part.isdigit() or part in _TARGET_STOPWORDS:
continue
tokens.append(part)
return tokens
def _expand_target_keywords(keywords: Set[str]) -> None:
if "mcp" in keywords:
keywords.update({
"configure", "configuration", "connect", "connected", "enable",
"enabled", "install", "installed", "server", "servers",
"настрой", "настроить", "подключи", "подключить",
})
if {"conflict", "conflicts"} & keywords:
keywords.update({
"cherry", "conflict", "conflicts", "git", "merge", "rebase",
"unmerged", "конфликт", "конфликты",
})
def target_task_keywords(
target_skill_text: str,
target_skill_path: str = "",
*,
limit: int = 180,
) -> Tuple[Set[str], Set[str]]:
"""Return (strong, weak) keywords that describe a target skill."""
path_text = (target_skill_path or "").replace(os.sep, " ")
headings = "\n".join(re.findall(r"(?m)^#+\s+(.+)$", target_skill_text or ""))
strong = set(_target_tokens(path_text + "\n" + headings))
weak = set(strong)
counts = Counter(_target_tokens(target_skill_text or ""))
for token, _count in counts.most_common(limit):
weak.add(token)
_expand_target_keywords(strong)
_expand_target_keywords(weak)
return strong, weak
def _task_search_text(task: TaskRecord) -> str:
return "\n".join([
task.intent or "",
task.context_excerpt or "",
" ".join(task.tags or []),
])
def filter_tasks_for_target(
tasks: List[TaskRecord],
target_skill_text: str,
target_skill_path: str = "",
) -> List[TaskRecord]:
"""Prefer tasks whose language overlaps the explicit target skill.
If nothing matches, return the original list. This keeps a target run useful
even when transcripts are too sparse or the skill is too generic.
"""
strong, weak = target_task_keywords(target_skill_text, target_skill_path)
if not tasks or not (strong or weak):
return tasks
ranked = []
for idx, task in enumerate(tasks):
tokens = set(_target_tokens(_task_search_text(task)))
strong_hits = tokens & strong
weak_hits = tokens & weak
if not strong_hits and len(weak_hits) < 2:
continue
score = len(strong_hits) * 3 + len(weak_hits)
ranked.append((score, idx, task))
if not ranked:
return tasks
ranked.sort(key=lambda item: (-item[0], item[1]))
return [task for _score, _idx, task in ranked]
def heuristic_mine(
digests: List[SessionDigest],
*,
@@ -192,11 +287,15 @@ def mine(
digests: List[SessionDigest],
*,
max_tasks: int = 40,
candidate_limit: int = 0,
holdout_fraction: float = 0.34,
seed: int = 42,
llm_miner: Optional[Callable[[List[SessionDigest]], List[TaskRecord]]] = None,
target_skill_text: str = "",
target_skill_path: str = "",
) -> List[TaskRecord]:
"""Top-level miner. Uses ``llm_miner`` if provided, else heuristic."""
candidate_limit = candidate_limit or max_tasks
tasks: List[TaskRecord] = []
if llm_miner is not None:
try:
@@ -204,7 +303,10 @@ def mine(
except Exception:
tasks = []
if not tasks:
tasks = heuristic_mine(digests, max_tasks=max_tasks)
tasks = heuristic_mine(digests, max_tasks=candidate_limit)
tasks = dedup_tasks(tasks)
if target_skill_text or target_skill_path:
tasks = filter_tasks_for_target(tasks, target_skill_text, target_skill_path)
tasks = tasks[:max_tasks]
tasks = assign_splits(tasks, holdout_fraction=holdout_fraction, seed=seed)
return tasks
+57 -1
View File
@@ -9,12 +9,68 @@ from __future__ import annotations
import json
import os
import re
import shutil
import time
from typing import List, Optional
from typing import Any, List, Optional
from skillopt_sleep.types import SleepReport
# Secret patterns scrubbed from any free-text we persist to the staging dir
# (diagnostics, reports). Kept here so every on-disk artifact shares one
# redaction pass; harvest_codex reuses these for session text too.
_SECRET_PATTERNS: tuple[tuple[re.Pattern[str], str], ...] = (
(re.compile(r"sk-[A-Za-z0-9_-]{10,}"), "[REDACTED_OPENAI_KEY]"),
# Distinctive vendor token prefixes (low false-positive: these prefixes do
# not occur in normal diagnostic prose).
(re.compile(r"\bAKIA[0-9A-Z]{16}\b"), "[REDACTED_AWS_KEY]"),
(re.compile(r"\bgh[pousr]_[A-Za-z0-9]{20,}\b"), "[REDACTED_GITHUB_TOKEN]"),
(re.compile(r"\bxox[baprs]-[A-Za-z0-9-]{10,}\b"), "[REDACTED_SLACK_TOKEN]"),
(re.compile(r"\bAIza[0-9A-Za-z_-]{20,}\b"), "[REDACTED_GOOGLE_KEY]"),
# Bare JWT (three base64url segments) — e.g. a leaked bearer body without
# the "Authorization:" prefix.
(re.compile(r"\beyJ[A-Za-z0-9_-]{8,}\.[A-Za-z0-9_-]{8,}\.[A-Za-z0-9_-]{8,}\b"),
"[REDACTED_JWT]"),
(re.compile(r"(?i)(Authorization:\s*Bearer\s+)[^\s\"']+"), r"\1[REDACTED]"),
(re.compile(r"(?i)(Authorization:\s*Basic\s+)[^\s\"']+"), r"\1[REDACTED]"),
(
re.compile(r"(?i)\b(api[_-]?key|token|password|secret)\b(\s*[:=]\s*)[^\s\"']+"),
r"\1\2[REDACTED]",
),
(
re.compile(r"(?i)\b(api[_-]?key|token|password|secret)\b(\s+)[^\s\"']+"),
r"\1\2[REDACTED]",
),
(
re.compile(
r"-----BEGIN [A-Z ]*PRIVATE KEY-----.*?-----END [A-Z ]*PRIVATE KEY-----",
re.DOTALL,
),
"[REDACTED_PRIVATE_KEY]",
),
)
def redact_secrets(value: Any) -> Any:
"""Scrub secret-looking substrings (API keys, bearer tokens, private keys)
from a string, or recursively from the string leaves of a list/dict.
Used before writing backend stderr / optimizer replies / task responses to
on-disk diagnostics: those are surfaced for debugging, but the underlying
text (e.g. a codex 401 stderr dump) can carry credentials. Non-string
scalars pass through unchanged.
"""
if isinstance(value, str):
out = value
for pattern, replacement in _SECRET_PATTERNS:
out = pattern.sub(replacement, out)
return out
if isinstance(value, list):
return [redact_secrets(v) for v in value]
if isinstance(value, dict):
return {k: redact_secrets(v) for k, v in value.items()}
return value
def _ts_dir() -> str:
return time.strftime("%Y%m%d-%H%M%S", time.localtime())
+13
View File
@@ -28,6 +28,7 @@ DEFAULT_STATE: Dict[str, Any] = {
"last_harvest": {}, # project -> iso timestamp of last harvested record
"slow_memory": "", # cross-night consolidated lessons (meta-skill analogue)
"history": [], # list of per-night summaries
"task_archive": [], # capped list of past mined tasks (for associative recall)
}
@@ -81,3 +82,15 @@ class SleepState:
def record_night(self, summary: Dict[str, Any]) -> None:
self.data.setdefault("history", []).append(summary)
# ── task archive (associative-recall memory) ──────────────────────────
def task_archive(self) -> list:
"""Past mined tasks as plain dicts (newest last)."""
return list(self.data.get("task_archive", []))
def add_to_archive(self, task_dicts: list, cap: int = 300) -> None:
"""Append tonight's tasks; keep only the most recent ``cap``."""
arc = self.data.setdefault("task_archive", [])
arc.extend(task_dicts)
if len(arc) > cap:
self.data["task_archive"] = arc[-cap:]
+81
View File
@@ -0,0 +1,81 @@
"""Reviewed task-file helpers for privacy-safe SkillOpt-Sleep runs."""
from __future__ import annotations
import json
import os
from typing import Any, Dict, List, Tuple
from skillopt_sleep.mine import assign_splits, normalize_legacy_split
from skillopt_sleep.types import TaskRecord
def make_tasks_payload(
tasks: List[TaskRecord],
*,
project: str,
transcript_source: str = "",
n_sessions: int = 0,
target_skill_path: str = "",
) -> Dict[str, Any]:
return {
"format": "skillopt_sleep.tasks.v1",
"project": project,
"transcript_source": transcript_source,
"n_sessions": n_sessions,
"target_skill_path": target_skill_path,
"reviewed": False,
"tasks": [t.to_dict() for t in tasks],
}
def write_tasks_file(path: str, payload: Dict[str, Any]) -> str:
out = os.path.abspath(os.path.expanduser(path))
parent = os.path.dirname(out)
if parent:
os.makedirs(parent, exist_ok=True)
with open(out, "w", encoding="utf-8") as f:
json.dump(payload, f, ensure_ascii=False, indent=2)
f.write("\n")
return out
def _normalize_tasks(
tasks: List[TaskRecord],
*,
holdout_fraction: float,
seed: int,
) -> List[TaskRecord]:
for task in tasks:
task.split = normalize_legacy_split(task.split or "train")
if len(tasks) >= 2 and not any(task.split in {"val", "test"} for task in tasks):
tasks = assign_splits(tasks, holdout_fraction=holdout_fraction, seed=seed)
return tasks
def load_tasks_file(
path: str,
*,
holdout_fraction: float = 0.34,
seed: int = 42,
) -> Tuple[List[TaskRecord], Dict[str, Any]]:
source = os.path.abspath(os.path.expanduser(path))
with open(source, encoding="utf-8") as f:
payload = json.load(f)
if isinstance(payload, list):
meta: Dict[str, Any] = {"format": "skillopt_sleep.tasks.v1", "tasks_file": source}
raw_tasks = payload
elif isinstance(payload, dict):
meta = {k: v for k, v in payload.items() if k != "tasks"}
meta["tasks_file"] = source
raw_tasks = payload.get("tasks", [])
else:
raise ValueError("tasks file must contain a JSON object with tasks or a JSON task array")
if not isinstance(raw_tasks, list):
raise ValueError("tasks file field 'tasks' must be an array")
tasks: List[TaskRecord] = []
for item in raw_tasks:
if not isinstance(item, dict):
raise ValueError("each task entry must be a JSON object")
tasks.append(TaskRecord.from_dict(item))
return _normalize_tasks(tasks, holdout_fraction=holdout_fraction, seed=seed), meta
+6 -6
View File
@@ -8,18 +8,17 @@ external dependencies.
"""
from __future__ import annotations
from dataclasses import dataclass, field, asdict
from typing import Any, Dict, List, Optional
from dataclasses import asdict, dataclass, field
from typing import Any, Dict, List
# ── Stage 1: harvest ──────────────────────────────────────────────────────────
@dataclass
class SessionDigest:
"""A normalized summary of one Claude Code session transcript.
"""A normalized summary of one local agent session transcript.
Produced by :mod:`skillopt_sleep.harvest` from a ``<sessionId>.jsonl``
transcript plus ``history.jsonl`` entries.
Produced by source-specific harvesters from Claude Code transcripts or
Codex Desktop archived sessions.
"""
session_id: str
@@ -136,6 +135,7 @@ class SleepReport:
candidate_score: float = 0.0
accepted: bool = False
gate_action: str = ""
no_edits_reason: str = ""
edits: List[EditRecord] = field(default_factory=list)
rejected_edits: List[EditRecord] = field(default_factory=list)
tokens_used: int = 0
+135 -25
View File
@@ -9,15 +9,19 @@ import glob
import json
import os
import signal
import socket
import subprocess
import sys
import threading
import time
from pathlib import Path
from urllib.parse import urlparse
import gradio as gr
import yaml
from skillopt.config import flatten_config
from skillopt.config import load_config as load_merged_config
PROJECT_ROOT = Path(__file__).resolve().parent.parent
@@ -42,6 +46,131 @@ def config_to_display(cfg: dict) -> str:
return yaml.dump(cfg, default_flow_style=False, sort_keys=False)
def _can_connect_to_url(url: str, timeout: float = 0.5) -> bool:
parsed = urlparse(url)
host = parsed.hostname
if not host:
return False
port = parsed.port or (443 if parsed.scheme == "https" else 80)
try:
with socket.create_connection((host, port), timeout=timeout):
return True
except OSError:
return False
def _load_env_file(path: Path, env: dict[str, str]) -> None:
for line in path.read_text().splitlines():
line = line.strip()
if line.startswith("export "):
line = line[len("export "):].strip()
if line and not line.startswith("#") and "=" in line:
key, value = line.split("=", 1)
env[key.strip()] = value.strip().strip("\"'")
def build_training_env() -> dict[str, str]:
"""Build the environment shared by preflight and the training subprocess."""
env = os.environ.copy()
env["PYTHONUNBUFFERED"] = "1"
dot_env = PROJECT_ROOT / ".env"
if dot_env.is_file():
_load_env_file(dot_env, env)
secrets_dir = PROJECT_ROOT / ".secrets"
if secrets_dir.is_dir():
for env_file in sorted(secrets_dir.glob("*.env")):
_load_env_file(env_file, env)
# Propagate OPTIMIZER_* to base AZURE_OPENAI_* when base is missing,
# so target/default endpoints inherit from optimizer config.
for suffix in (
"ENDPOINT", "API_VERSION", "AUTH_MODE", "MANAGED_IDENTITY_CLIENT_ID",
"AD_SCOPE", "API_KEY",
):
base_key = f"AZURE_OPENAI_{suffix}"
optimizer_key = f"OPTIMIZER_AZURE_OPENAI_{suffix}"
if not env.get(base_key) and env.get(optimizer_key):
env[base_key] = env[optimizer_key]
return env
def validate_training_config(
config_path: str,
overrides: dict,
env: dict[str, str] | None = None,
) -> str | None:
"""Return an actionable preflight error, or None when training can start."""
env = env or os.environ
cfg_options = [
f"{key}={value}" for key, value in overrides.items()
if value is not None and value != ""
]
try:
cfg = flatten_config(load_merged_config(str(PROJECT_ROOT / config_path), cfg_options))
except Exception as exc:
return f"❌ Invalid config: {exc}"
shared_endpoint = (
cfg.get("azure_openai_endpoint")
or cfg.get("azure_endpoint")
or env.get("AZURE_OPENAI_ENDPOINT")
)
missing_openai_roles = []
for role in ("optimizer", "target"):
if cfg.get(f"{role}_backend") != "openai_chat":
continue
role_endpoint = (
cfg.get(f"{role}_azure_openai_endpoint")
or env.get(f"{role.upper()}_AZURE_OPENAI_ENDPOINT")
or shared_endpoint
)
if not role_endpoint:
missing_openai_roles.append(role)
if missing_openai_roles:
configured_backend = cfg.get("model_backend")
detail = ""
if configured_backend in {"qwen", "qwen_chat"}:
detail = (
"\nNote: model.backend is qwen, but explicit optimizer_backend/"
"target_backend values are still openai_chat."
)
return (
"❌ Model backend is not ready: missing Azure/OpenAI-compatible endpoint "
f"for {', '.join(missing_openai_roles)}.\n"
"Set model.azure_openai_endpoint (or AZURE_OPENAI_ENDPOINT), or change "
"the role backends to the backend you intend to use."
f"{detail}"
)
qwen_failures = []
qwen_shared = (
cfg.get("qwen_chat_base_url")
or env.get("QWEN_CHAT_BASE_URL")
or "http://localhost:8000/v1"
)
for role in ("optimizer", "target"):
if cfg.get(f"{role}_backend") != "qwen_chat":
continue
base_url = (
cfg.get(f"{role}_qwen_chat_base_url")
or env.get(f"{role.upper()}_QWEN_CHAT_BASE_URL")
or qwen_shared
)
if not _can_connect_to_url(str(base_url)):
qwen_failures.append(f"{role}={base_url}")
if qwen_failures:
return (
"❌ Model backend is not ready: cannot connect to qwen_chat endpoint "
f"for {', '.join(qwen_failures)}.\n"
"Start your OpenAI-compatible Qwen/vLLM server, or set "
"model.qwen_chat_base_url / OPTIMIZER_QWEN_CHAT_BASE_URL / "
"TARGET_QWEN_CHAT_BASE_URL to the correct URL."
)
return None
# ─── Training process management ────────────────────────────────────────────
class TrainingManager:
@@ -63,6 +192,11 @@ class TrainingManager:
if self.running:
return "⚠️ Training already running. Stop it first."
env = build_training_env()
preflight_error = validate_training_config(config_path, overrides, env)
if preflight_error:
return preflight_error
cmd = [
sys.executable, "scripts/train.py",
"--config", config_path,
@@ -75,30 +209,6 @@ class TrainingManager:
cmd.append("--cfg-options")
cmd.extend(cfg_options)
env = os.environ.copy()
env["PYTHONUNBUFFERED"] = "1"
# Auto-load API credentials from .secrets/*.env
secrets_dir = PROJECT_ROOT / ".secrets"
if secrets_dir.is_dir():
for env_file in sorted(secrets_dir.glob("*.env")):
for line in env_file.read_text().splitlines():
line = line.strip()
if line and not line.startswith("#") and "=" in line:
k, v = line.split("=", 1)
env[k] = v
# Propagate OPTIMIZER_* to base AZURE_OPENAI_* when base is missing,
# so target/default endpoints inherit from optimizer config.
_propagate = [
("ENDPOINT", ""), ("API_VERSION", ""), ("AUTH_MODE", ""),
("MANAGED_IDENTITY_CLIENT_ID", ""), ("AD_SCOPE", ""),
("API_KEY", ""),
]
for suffix, _ in _propagate:
base_key = f"AZURE_OPENAI_{suffix}"
optimizer_key = f"OPTIMIZER_AZURE_OPENAI_{suffix}"
if not env.get(base_key) and env.get(optimizer_key):
env[base_key] = env[optimizer_key]
try:
proc = subprocess.Popen(
cmd,
+33
View File
@@ -0,0 +1,33 @@
import os
from skillopt.envs.alfworld.rollout import _resolve_alfworld_gamefile, _resolve_alfworld_gamefiles
def test_resolve_alfworld_gamefile_uses_alfworld_data_for_relative_paths(monkeypatch, tmp_path):
data_root = tmp_path / "alfworld_data"
monkeypatch.setenv("ALFWORLD_DATA", str(data_root))
resolved = _resolve_alfworld_gamefile("json_2.1.1/valid_seen/task/game.tw-pddl")
assert resolved == os.path.join(str(data_root), "json_2.1.1/valid_seen/task/game.tw-pddl")
def test_resolve_alfworld_gamefile_keeps_absolute_paths(monkeypatch, tmp_path):
monkeypatch.setenv("ALFWORLD_DATA", str(tmp_path / "alfworld_data"))
absolute = tmp_path / "elsewhere" / "game.tw-pddl"
assert _resolve_alfworld_gamefile(str(absolute)) == str(absolute)
def test_resolve_alfworld_gamefile_keeps_relative_path_without_alfworld_data(monkeypatch):
monkeypatch.delenv("ALFWORLD_DATA", raising=False)
assert _resolve_alfworld_gamefile("json_2.1.1/train/task/game.tw-pddl") == (
"json_2.1.1/train/task/game.tw-pddl"
)
def test_resolve_alfworld_gamefiles_handles_none(monkeypatch):
monkeypatch.setenv("ALFWORLD_DATA", "/tmp/alfworld_data")
assert _resolve_alfworld_gamefiles(None) is None
+98
View File
@@ -0,0 +1,98 @@
"""Tests for the Devin MCP plugin: tool schema, ATIF-v1.7 harvest, path expansion."""
import importlib
import json
import os
import sys
import tempfile
import unittest
# Allow importing from the plugin directory (mirrors tests/test_mcp_schema.py)
PLUGIN = os.path.join(os.path.dirname(__file__), "..", "plugins", "devin")
sys.path.insert(0, PLUGIN)
import mcp_server # noqa: E402
import harvest_devin as hw # noqa: E402
FIXTURES = os.path.join(PLUGIN, "fixtures")
def _read_jsonl(path):
with open(path, encoding="utf-8") as f:
return [json.loads(line) for line in f if line.strip()]
def _find_session_jsonl(out_dir):
for root, _dirs, files in os.walk(os.path.join(out_dir, "projects")):
for name in files:
if name.endswith(".jsonl"):
return _read_jsonl(os.path.join(root, name))
raise AssertionError("no session jsonl written")
class TestDevinMcpSchema(unittest.TestCase):
def test_tools_are_the_sleep_interface(self):
names = {t["name"] for t in mcp_server.TOOLS}
self.assertEqual(names, {"sleep_status", "sleep_dry_run", "sleep_run",
"sleep_adopt", "sleep_harvest",
"sleep_schedule", "sleep_unschedule"})
def test_actions_map_to_engine_subcommands(self):
expected = {"sleep_status": "status", "sleep_dry_run": "dry-run",
"sleep_run": "run", "sleep_adopt": "adopt",
"sleep_harvest": "harvest", "sleep_schedule": "schedule",
"sleep_unschedule": "unschedule"}
for t in mcp_server.TOOLS:
self.assertEqual(t["action"], expected[t["name"]])
def test_backends_in_enum(self):
backends = mcp_server._TOOL_SCHEMA["properties"]["backend"]["enum"]
for b in ["mock", "claude", "codex", "copilot"]:
self.assertIn(b, backends)
def test_schema_has_key_engine_params(self):
# parity with plugins/copilot's schema (tests/test_plugin_sync.py)
props = set(mcp_server._TOOL_SCHEMA["properties"].keys())
for param in {"project", "backend", "scope", "source", "model",
"tasks_file", "target_skill_path", "max_sessions",
"max_tasks", "lookback_hours", "auto_adopt", "json",
"edit_budget", "hour", "minute"}:
self.assertIn(param, props)
class TestClaudeHomeExpansion(unittest.TestCase):
"""Regression: ~ must be expanded even when CLAUDE_HOME comes from the env
(the documented mcp-config sets SKILLOPT_DEVIN_CLAUDE_HOME="~/...")."""
def test_env_tilde_is_expanded(self):
os.environ["SKILLOPT_DEVIN_CLAUDE_HOME"] = "~/.skillopt-sleep-devin"
try:
importlib.reload(mcp_server)
self.assertFalse(mcp_server.CLAUDE_HOME.startswith("~"))
self.assertEqual(mcp_server.CLAUDE_HOME,
os.path.expanduser("~/.skillopt-sleep-devin"))
finally:
del os.environ["SKILLOPT_DEVIN_CLAUDE_HOME"]
importlib.reload(mcp_server)
class TestDevinHarvest(unittest.TestCase):
def test_atif_fixture_yields_gradeable_task(self):
with tempfile.TemporaryDirectory() as out:
n = hw.harvest_devin_transcripts(FIXTURES, out, ["/tmp/proj"])
self.assertEqual(n, 1)
outcomes = _read_jsonl(os.path.join(out, "outcomes.jsonl"))
self.assertEqual(len(outcomes), 1)
o = outcomes[0]
self.assertEqual(o["verifier"], "tests")
self.assertTrue(o["success"])
self.assertIn("repro", o["reference"])
# the converted transcript carries the grouping key on the user turn
session = _find_session_jsonl(out)
user_turn = next(r for r in session if r["type"] == "user")
self.assertIn("taskKey", user_turn)
if __name__ == "__main__":
unittest.main()
+286
View File
@@ -0,0 +1,286 @@
"""Tests for skillopt.evaluation.gate — the validation gate decision function.
The gate is the optimizer's model-selection / early-stopping core: given a
candidate skill's score, it decides whether to accept it as the new current
skill and whether it becomes the new best-so-far. These are pure functions,
so they can be exercised directly without any LLM or rollout.
"""
from __future__ import annotations
import dataclasses
import pytest
from skillopt.evaluation.gate import (
GateResult,
evaluate_gate,
select_gate_score,
)
class TestSelectGateScore:
"""select_gate_score — project (hard, soft) onto a single comparison metric."""
def test_hard_metric_returns_hard(self) -> None:
assert select_gate_score(0.8, 0.3, "hard") == 0.8
def test_soft_metric_returns_soft(self) -> None:
assert select_gate_score(0.8, 0.3, "soft") == 0.3
def test_default_metric_is_hard(self) -> None:
assert select_gate_score(0.42, 0.99) == 0.42
def test_mixed_metric_default_weight(self) -> None:
# (1 - 0.5) * 1.0 + 0.5 * 0.0 == 0.5
assert select_gate_score(1.0, 0.0, "mixed") == pytest.approx(0.5)
def test_mixed_metric_custom_weight(self) -> None:
# (1 - 0.25) * 0.8 + 0.25 * 0.4 == 0.7
assert select_gate_score(0.8, 0.4, "mixed", 0.25) == pytest.approx(0.7)
def test_mixed_weight_zero_equals_hard(self) -> None:
assert select_gate_score(0.8, 0.3, "mixed", 0.0) == pytest.approx(0.8)
def test_mixed_weight_one_equals_soft(self) -> None:
assert select_gate_score(0.8, 0.3, "mixed", 1.0) == pytest.approx(0.3)
def test_mixed_weight_clamped_above_one(self) -> None:
"""Out-of-range weight is clamped to 1.0 (→ pure soft)."""
assert select_gate_score(0.8, 0.3, "mixed", 5.0) == pytest.approx(0.3)
def test_mixed_weight_clamped_below_zero(self) -> None:
"""Negative weight is clamped to 0.0 (→ pure hard)."""
assert select_gate_score(0.8, 0.3, "mixed", -2.0) == pytest.approx(0.8)
def test_returns_float(self) -> None:
assert isinstance(select_gate_score(1, 0, "hard"), float)
def test_unknown_metric_raises(self) -> None:
with pytest.raises(ValueError, match="unknown gate metric"):
select_gate_score(0.5, 0.5, "rouge") # type: ignore[arg-type]
class TestEvaluateGateAcceptNewBest:
"""evaluate_gate — candidate beats both current and best."""
def test_accept_new_best_action_and_state(self) -> None:
result = evaluate_gate(
candidate_skill="CAND",
cand_hard=0.9,
current_skill="CURR",
current_score=0.5,
best_skill="BEST",
best_score=0.5,
best_step=3,
global_step=7,
)
assert result.action == "accept_new_best"
assert result.current_skill == "CAND"
assert result.current_score == pytest.approx(0.9)
assert result.best_skill == "CAND"
assert result.best_score == pytest.approx(0.9)
assert result.best_step == 7 # updated to the accepting step
class TestEvaluateGateAccept:
"""evaluate_gate — candidate beats current but not best.
This branch is only reachable when ``current_score < best_score``; it
advances the current skill without disturbing the best-so-far checkpoint.
"""
def test_accept_updates_current_only(self) -> None:
result = evaluate_gate(
candidate_skill="CAND",
cand_hard=0.6,
current_skill="CURR",
current_score=0.4,
best_skill="BEST",
best_score=0.8,
best_step=2,
global_step=9,
)
assert result.action == "accept"
assert result.current_skill == "CAND"
assert result.current_score == pytest.approx(0.6)
# best-so-far is preserved, including its step
assert result.best_skill == "BEST"
assert result.best_score == pytest.approx(0.8)
assert result.best_step == 2
def test_tie_with_best_but_above_current_accepts(self) -> None:
"""cand == best (not strictly greater) but > current → accept, not new best."""
result = evaluate_gate(
candidate_skill="CAND",
cand_hard=0.8,
current_skill="CURR",
current_score=0.5,
best_skill="BEST",
best_score=0.8,
best_step=1,
global_step=4,
)
assert result.action == "accept"
assert result.current_skill == "CAND"
assert result.best_skill == "BEST"
assert result.best_score == pytest.approx(0.8)
assert result.best_step == 1
class TestEvaluateGateReject:
"""evaluate_gate — candidate does not beat current."""
def test_reject_below_current(self) -> None:
result = evaluate_gate(
candidate_skill="CAND",
cand_hard=0.3,
current_skill="CURR",
current_score=0.5,
best_skill="BEST",
best_score=0.8,
best_step=2,
global_step=6,
)
assert result.action == "reject"
assert result.current_skill == "CURR"
assert result.current_score == pytest.approx(0.5)
assert result.best_skill == "BEST"
assert result.best_score == pytest.approx(0.8)
assert result.best_step == 2
def test_tie_with_current_rejects(self) -> None:
"""Strict inequality: cand == current is rejected (no lateral moves)."""
result = evaluate_gate(
candidate_skill="CAND",
cand_hard=0.5,
current_skill="CURR",
current_score=0.5,
best_skill="BEST",
best_score=0.5,
best_step=0,
global_step=3,
)
assert result.action == "reject"
assert result.current_skill == "CURR"
assert result.best_skill == "BEST"
class TestEvaluateGateMetrics:
"""evaluate_gate — non-hard metrics drive the comparison via cand_soft."""
def test_soft_metric_uses_cand_soft(self) -> None:
# High hard, low soft: under 'soft' the candidate must be rejected.
result = evaluate_gate(
candidate_skill="CAND",
cand_hard=0.95,
current_skill="CURR",
current_score=0.5,
best_skill="BEST",
best_score=0.5,
best_step=0,
global_step=1,
cand_soft=0.2,
metric="soft",
)
assert result.action == "reject"
def test_mixed_metric_uses_weighted_score(self) -> None:
# mixed w=0.5: (0.5 * 1.0) + (0.5 * 0.6) == 0.8 > current 0.5 and best 0.5
result = evaluate_gate(
candidate_skill="CAND",
cand_hard=1.0,
current_skill="CURR",
current_score=0.5,
best_skill="BEST",
best_score=0.5,
best_step=0,
global_step=2,
cand_soft=0.6,
metric="mixed",
mixed_weight=0.5,
)
assert result.action == "accept_new_best"
assert result.current_score == pytest.approx(0.8)
assert result.best_score == pytest.approx(0.8)
def test_default_metric_ignores_soft(self) -> None:
"""Default metric is 'hard'; cand_soft must not affect the decision."""
result = evaluate_gate(
candidate_skill="CAND",
cand_hard=0.9,
current_skill="CURR",
current_score=0.5,
best_skill="BEST",
best_score=0.5,
best_step=0,
global_step=1,
cand_soft=0.0,
)
assert result.action == "accept_new_best"
assert result.current_score == pytest.approx(0.9)
class TestGateResult:
"""GateResult — immutable outcome dataclass."""
def test_fields(self) -> None:
result = GateResult(
action="accept",
current_skill="c",
current_score=0.5,
best_skill="b",
best_score=0.9,
best_step=4,
)
assert result.action == "accept"
assert result.current_skill == "c"
assert result.current_score == 0.5
assert result.best_skill == "b"
assert result.best_score == 0.9
assert result.best_step == 4
def test_is_frozen(self) -> None:
result = GateResult(
action="reject",
current_skill="c",
current_score=0.0,
best_skill="b",
best_score=0.0,
best_step=0,
)
with pytest.raises(dataclasses.FrozenInstanceError):
result.current_score = 1.0 # type: ignore[misc]
class TestGateInvariants:
"""Behavioral invariants of the gate over a sequence of steps."""
def test_current_tracks_best_from_equal_start(self) -> None:
"""When current == best at the start, every acceptance is a new best, so
the two stay locked together and the 'accept' branch is never taken.
This documents the trainer's ``s_cur``/``s_best`` usage: they are
initialized equal and updated only through this gate.
"""
current_skill, current_score = "S0", 0.2
best_skill, best_score, best_step = "S0", 0.2, 0
for step, cand in enumerate([0.1, 0.5, 0.4, 0.7], start=1):
result = evaluate_gate(
candidate_skill=f"S{step}",
cand_hard=cand,
current_skill=current_skill,
current_score=current_score,
best_skill=best_skill,
best_score=best_score,
best_step=best_step,
global_step=step,
)
current_skill, current_score = result.current_skill, result.current_score
best_skill = result.best_skill
best_score = result.best_score
best_step = result.best_step
assert result.action in {"accept_new_best", "reject"}
assert current_score == best_score
assert current_skill == best_skill
assert best_score == pytest.approx(0.7)
assert best_step == 4
+220
View File
@@ -0,0 +1,220 @@
"""Tests for the handoff backend: session-executed model calls via
prompt/answer files, resumed across engine runs."""
from __future__ import annotations
import json
import os
import re
import tempfile
import unittest
from skillopt_sleep.backend import get_backend
from skillopt_sleep.config import load_config
from skillopt_sleep.cycle import run_sleep_cycle
from skillopt_sleep.handoff_backend import (
PENDING_SENTINEL_PREFIX,
HandoffBackend,
PendingCalls,
)
from skillopt_sleep.mine import assign_splits
from skillopt_sleep.types import TaskRecord
# The rule the simulated executor "learns"; once present in the skill the
# executor answers arithmetic tasks correctly, mirroring MockBackend's
# rule-gated model of reality.
RULE = "Always answer with just the number."
def _tasks():
ts = [
TaskRecord(
id=f"t{i}", project="/p",
intent=f"What is {i} + {i}?",
reference_kind="exact", reference=str(i + i),
)
for i in range(1, 7)
]
return assign_splits(ts, holdout_fraction=0.34, seed=42)
def _answer_pending(backend: HandoffBackend) -> None:
"""Deterministic stand-in for the interactive session answering PROMPTS.md."""
for key, item in list(backend.pending.items()):
prompt = str(item["prompt"])
if "You are SkillOpt's optimizer" in prompt:
ans = json.dumps([{
"op": "add", "content": RULE,
"rationale": "outputs failed exact match; answer bare numbers",
}])
else:
m = re.search(r"What is (\d+) \+ (\d+)\?", prompt)
if m and RULE in prompt:
ans = str(int(m.group(1)) + int(m.group(2)))
else:
ans = "cannot say"
with open(backend.answer_path(key), "w", encoding="utf-8") as f:
f.write(ans)
class TestHandoffBackendUnit(unittest.TestCase):
def test_miss_records_pending_and_returns_sentinel(self):
with tempfile.TemporaryDirectory() as hdir:
be = HandoffBackend(handoff_dir=hdir)
out = be._call("some prompt")
self.assertTrue(out.startswith(PENDING_SENTINEL_PREFIX))
self.assertEqual(len(be.pending), 1)
def test_answer_file_resolves_call(self):
with tempfile.TemporaryDirectory() as hdir:
be = HandoffBackend(handoff_dir=hdir)
key = be._call("what is 2+2?").split(":")[-1].rstrip("]")
with open(be.answer_path(key), "w", encoding="utf-8") as f:
f.write("4\n")
fresh = HandoffBackend(handoff_dir=hdir)
self.assertEqual(fresh._call("what is 2+2?"), "4")
self.assertEqual(len(fresh.pending), 0)
def test_dependent_prompt_raises(self):
with tempfile.TemporaryDirectory() as hdir:
be = HandoffBackend(handoff_dir=hdir)
placeholder = be._call("first question")
with self.assertRaises(PendingCalls) as ctx:
be._call(f"judge this response: {placeholder}")
self.assertEqual(len(ctx.exception.pending), 1)
def test_reflect_retry_of_pending_reply_raises(self):
with tempfile.TemporaryDirectory() as hdir:
be = HandoffBackend(handoff_dir=hdir)
be._call("reflect on failures")
with self.assertRaises(PendingCalls):
be._call("reflect on failures\n\nIMPORTANT: "
"your previous reply was not valid JSON. Reply with "
"ONLY the JSON array, no prose, no markdown fences.")
def test_flush_writes_prompts_and_pending_json(self):
with tempfile.TemporaryDirectory() as hdir:
be = HandoffBackend(handoff_dir=hdir)
be._call("prompt A")
be._call("prompt B")
prompts_path = be.flush_pending()
self.assertTrue(os.path.exists(prompts_path))
with open(os.path.join(hdir, "pending.json"), encoding="utf-8") as f:
payload = json.load(f)
self.assertEqual(payload["format"], "skillopt_sleep.handoff.v1")
self.assertEqual(len(payload["pending"]), 2)
self.assertEqual(payload["pending"][0]["prompt"], "prompt A")
with open(prompts_path, encoding="utf-8") as f:
md = f.read()
self.assertIn("BEGIN PROMPT", md)
self.assertIn("prompt B", md)
def test_get_backend_registration(self):
with tempfile.TemporaryDirectory() as proj:
be = get_backend("handoff", project_dir=proj)
self.assertIsInstance(be, HandoffBackend)
self.assertTrue(be.handoff_dir.startswith(os.path.realpath(proj))
or be.handoff_dir.startswith(proj))
class TestHandoffCycle(unittest.TestCase):
def test_cycle_converges_over_handoff_rounds(self):
with tempfile.TemporaryDirectory() as proj, \
tempfile.TemporaryDirectory() as home:
hdir = os.path.join(proj, ".skillopt-sleep-handoff")
def cfg():
return load_config(
invoked_project=proj, projects="invoked",
backend="handoff",
claude_home=os.path.join(home, ".claude"),
)
final = None
rounds = 0
for _ in range(8):
rounds += 1
backend = HandoffBackend(handoff_dir=hdir)
try:
outcome = run_sleep_cycle(
cfg(), seed_tasks=_tasks(), dry_run=True, backend=backend,
)
except PendingCalls:
outcome = None
if backend.pending:
backend.flush_pending()
_answer_pending(backend)
continue
final = outcome
break
self.assertIsNotNone(final, "cycle never completed within 8 rounds")
self.assertGreater(rounds, 1, "expected at least one handoff round")
rep = final.report
self.assertTrue(rep.accepted, f"gate did not accept: {rep.gate_action}")
self.assertGreater(rep.candidate_score, rep.baseline_score)
self.assertTrue(any(RULE in e.content for e in rep.edits))
class TestHandoffCli(unittest.TestCase):
def test_run_exits_3_and_stages_prompts(self):
from skillopt_sleep.__main__ import main
from skillopt_sleep.tasks_file import make_tasks_payload, write_tasks_file
with tempfile.TemporaryDirectory() as proj, \
tempfile.TemporaryDirectory() as home:
tasks_path = os.path.join(proj, "tasks.json")
payload = make_tasks_payload(_tasks(), project=proj)
payload["reviewed"] = True
write_tasks_file(tasks_path, payload)
rc = main([
"run", "--backend", "handoff",
"--project", proj,
"--claude-home", os.path.join(home, ".claude"),
"--tasks-file", tasks_path,
])
self.assertEqual(rc, 3)
hdir = os.path.join(proj, ".skillopt-sleep-handoff")
self.assertTrue(os.path.exists(os.path.join(hdir, "PROMPTS.md")))
self.assertTrue(os.path.exists(os.path.join(hdir, "pending.json")))
def test_corrupt_digests_pin_falls_back_to_reharvest(self):
from skillopt_sleep.__main__ import main
with tempfile.TemporaryDirectory() as proj, \
tempfile.TemporaryDirectory() as home:
hdir = os.path.join(proj, ".skillopt-sleep-handoff")
os.makedirs(hdir, exist_ok=True)
with open(os.path.join(hdir, "digests.json"), "w", encoding="utf-8") as f:
f.write("{not valid json")
rc = main([
"run", "--backend", "handoff",
"--project", proj,
"--claude-home", os.path.join(home, ".claude"),
])
# must not crash: corrupt pin -> fresh harvest -> no tasks -> 0
self.assertEqual(rc, 0)
def test_run_with_no_tasks_exits_0_and_advances_harvest_window(self):
from skillopt_sleep.__main__ import main
from skillopt_sleep.config import load_config
from skillopt_sleep.state import SleepState
with tempfile.TemporaryDirectory() as proj, \
tempfile.TemporaryDirectory() as home:
claude_home = os.path.join(home, ".claude")
rc = main([
"run", "--backend", "handoff",
"--project", proj,
"--claude-home", claude_home,
])
self.assertEqual(rc, 0)
# the no-tasks branch must persist last-harvest, otherwise every
# later run re-scans the same stale window forever
cfg = load_config(invoked_project=proj, claude_home=claude_home)
state = SleepState.load(cfg.state_path)
self.assertIsNotNone(state.last_harvest_for(proj))
if __name__ == "__main__":
unittest.main()

Some files were not shown because too many files have changed in this diff Show More