The clean-cwd + --disallowedTools isolation was NOT enough: the user's GLOBAL
skills (~/.claude/skills) are injected regardless of cwd, so reflect/attempt
still sometimes replied with a list of installed skills instead of JSON edits
(advisor reflect returned 21KB of skill descriptions, n_edits=0 -> gate reject).
Add --bare (skip hooks/LSP/plugins) and --disable-slash-commands (disable all
skills). Verified: the optimizer now returns clean JSON. Re-validating all
seeds with the truly-isolated backend; prior Claude numbers are being recomputed
honestly (some earlier "successes" were partly leak-assisted).
Co-Authored-By: Claude Opus 4 <noreply@anthropic.com>
Critical correctness fix found by debugging the thorough-analyst failure:
* `claude -p` was running with the AMBIENT Claude Code project context (the
repo's CLAUDE.md, installed skills, tools). The optimizer/target calls were
polluted — reflect once replied with a list of the user's installed skills
instead of JSON edits. Now ClaudeCliBackend._call runs ISOLATED: a clean temp
cwd, --disallowedTools '*', --exclude-dynamic-system-prompt-sections. This is
essential for the backend to be trustworthy and reproducible.
* reflect prompt: translate failing rule-judge criteria into plain English
(max_chars=1200 -> "the ENTIRE response must be at most 1200 characters") and
require CONCRETE, verbatim thresholds in proposed rules (not "respect limits").
* attempt prompt: treat the Learned-preferences block as HARD CONSTRAINTS that
override earlier conflicting skill text.
Earlier Claude results predate this fix and are being re-validated clean; the
Codex backend was never affected (it runs in its own exec context).
Co-Authored-By: Claude Opus 4 <noreply@anthropic.com>
- skillopt-sleep-plugin/.claude-plugin/marketplace.json so the plugin is
installable via `/plugin marketplace add ./skillopt-sleep-plugin`.
- README install section (clone -> add marketplace -> install -> /sleep status).
- docs/sleep/FINAL_REPORT.md: the consolidated presented results doc (real
Claude+Codex, transfer, and the honest thorough-analyst failure + fix).
- sweep.py flushes stdout for live monitoring.
Co-Authored-By: Claude Opus 4 <noreply@anthropic.com>
- sweep.py: run many (backend, model, seed, transfer-pair) configs sequentially,
append each result to JSONL incrementally (resumable, interrupt-safe).
- report.py: render the sweep JSONL into a presented Markdown scorecard with
direct-improvement and cross-model-transfer tables.
- reflect prompt now tells the optimizer its edits are APPENDED (can't delete the
base skill text), so on a conflict it must write a forceful OVERRIDE rule.
Diagnosed from a real failure: thorough-analyst (needs <=1200 chars) kept its
edits rejected because the base "be exhaustive" line won; a verified override
("HARD LIMIT ... supersedes") makes Haiku obey (1194/880 chars -> hard=1.0).
Co-Authored-By: Claude Opus 4 <noreply@anthropic.com>
Three additions driven by the goal of price-aware, model-flexible sleep:
1. DualBackend + build_backend(): route attempt->TARGET model and
reflect/judge->OPTIMIZER model (SkillOpt's target-vs-optimizer split).
gbrain runner gains --optimizer-backend/-model + --target-backend/-model.
2. run_transfer.py: sleep-scenario cross-model transfer. Optimize a skill on a
SOURCE model (e.g. cheap haiku), freeze it, evaluate held-out on a TARGET
model (e.g. expensive sonnet) with no further optimization — plus a direct
reference. Mirrors the SkillOpt paper's transfer table; quantifies the
"optimize cheap overnight, deploy anywhere" value prop.
3. llm_miner.py: turn real harvested transcripts into TaskRecords WITH checkable
rule/rubric judges, wired into the cycle for non-mock backends, so real-data
lift becomes measurable (heuristic miner remains the no-API fallback).
Fixed a str.format brace bug the new unit test caught.
19 tests pass.
Co-Authored-By: Claude Opus 4 <noreply@anthropic.com>
Upgrade from mock-only to REAL multi-backend validation:
Backends (skillopt/sleep/backend.py):
- CliBackend base: shared attempt/judge/reflect prompts, response cache,
token accounting. Subclasses implement only _call().
- ClaudeCliBackend: drives `claude -p --output-format text`.
- CodexCliBackend: drives the REAL @openai/codex `exec -o <file>` for clean
output; resolve_codex_path() skips the hermes wrapper at ~/.local/bin/codex.
- reflect() now aggregates the exact failing judge criteria into the prompt
(gbrain's lesson: tell the optimizer what the scorer rewards).
Rule judges (skillopt/sleep/judges.py): gbrain-compatible local scorers
(section_present / regex / max_chars / contains / tool_called) — held-out
scoring with no judge-API spend. TaskRecord gains a `judge` field +
reference_kind="rule".
gbrain-evals adapter (experiments/gbrain_bench.py, run_gbrain.py): load
garrytan/gbrain-evals skillopt-v1 deficient skills + train/held-out task
sets and run our consolidate() loop against the SAME suite gbrain scores.
REAL results (docs/sleep/real_api_results.md), brief-writer seed, 1 night:
- Claude (Haiku): held-out 0.00 -> 1.00
- Codex: held-out 0.00 -> 0.67
Both proposed a correct, general format rule into the protected LEARNED block.
CLI: --backend {mock,claude,codex}, --codex-path, --model; experiment +
gbrain runners gain --limit-* cost controls. 17 tests pass.
Co-Authored-By: Claude Opus 4 <noreply@anthropic.com>