diff --git a/docs/sleep/CONTROLLABLE_DREAMING.md b/docs/sleep/CONTROLLABLE_DREAMING.md new file mode 100644 index 0000000..10ab835 --- /dev/null +++ b/docs/sleep/CONTROLLABLE_DREAMING.md @@ -0,0 +1,94 @@ +# SkillOpt-Sleep — controllable dreaming architecture + +The sleep engine is no longer a single fixed pipeline. It is a controllable +offline "dream / imagination" loop the user steers. This documents the knobs +added in the four-stage refactor and how they map to the user's design. + +## The mental model + +> Sleep = an offline "脑补推演" (imagination rollout). Re-run the user's real +> tasks (and dream-augmented variants) many times, look at what went well vs +> badly, distil durable rules, and keep only what survives a real-task check — +> unless the user opts out of that check. + +## 1. Data splits — train (dream) / val (real) / test (real) + +The anti-overfitting foundation: + +| Split | Source | Role | +|---|---|---| +| **train** | real tasks **+ dream-augmented** variants | drives reflection (the imagination pool — over-dreaming is fine) | +| **val** | **real only**, disjoint from test | gates updates (prevents overfitting) | +| **test** | **real only**, disjoint from val | the final held-out measure, kept close to real usage | + +Hard guarantee (unit-tested): a task with `origin='dream'` **never** lands in +val or test. `assign_splits(val_fraction, test_fraction)` does the deterministic +3-way split; gbrain's own held-out maps to our `test`. + +## 2. The validation gate is optional + +`--gate on` (default): an edit is accepted only if it strictly improves the +**val** score — the SkillOpt discipline that blocks regressions and reward +hacking. + +`--gate off`: greedy. Edits are kept without the hard val-improvement +requirement (the user decides they don't want hard filtering), but val/test +movement is still reported (`greedy_improved` / `greedy_regressed` / +`greedy_flat`) so nothing is hidden. + +## 3. Slow-update — long-term memory, gate-independent + +Even with the gate off, the engine runs a **slow-update** at the end of the +nights: it compares behaviour under the first-night vs final skill across the +val tasks and distils durable longitudinal guidance into a **protected field** +(``, the same markers as +the main SkillOpt repo). Step-level edits never touch this field. This is the +"short-term experience → long-term memory" consolidation; turning the gate off +does not cost you long-term memory. + +## 4. Budget — the user picks the spend + +`--budget-tokens N` / `--budget-minutes M`: the engine auto-plans depth +(`nights × rollouts_per_task`) to fit the budget (`plan_depth`). Stops cleanly +when exhausted and logs what it skipped — no silent truncation. The whole thing +is offline imagination on the user's own quota. + +## 5. Multi-rollout contrastive reflection — the imagination core + +`--rollouts-k K` (K>1): each train task is rolled out K times. The optimizer is +shown the **high-scoring vs low-scoring** attempts of the same task and asked +what the good ones did that the bad ones didn't, distilling a general rule. This +is a far stronger signal than a single failure, and it is exactly the user's +"run it many times, learn from the contrast" idea. Tasks with the highest score +*spread* (some passed, some failed) are the most informative and are prioritised. + +## 6. Multi-objective reward — accuracy ↑, tokens ↓, latency ↓ + +Every rollout records its `tokens` and `latency_ms`. +`multi_objective_reward(w_acc, w_tokens, w_latency)` is a weighted reward so a +skill can be optimised to be **cheaper and faster**, not only more accurate +(cost terms normalised against a reference; default weights = accuracy-only, so +existing behaviour is unchanged). This turns "越用越好用" into "越用越准、越省、越快". + +## 7. User preferences as a prior + +`--preferences ""`: injected into the optimizer's reflect prompt as a +prior (set on the optimizer model for dual backends), so the user's stated +preferences steer what rules get written. + +## How the knobs compose (one command) + +```bash +python -m skillopt.sleep.experiments.run_gbrain \ + --optimizer-backend claude --optimizer-model sonnet \ # strong optimizer + --target-backend claude --target-model haiku \ # cheap target (transfer) + --seeds thorough-analyst \ + --gate on \ # or off for greedy + --rollouts-k 2 \ # contrastive imagination + --budget-tokens 60000 \ # auto-plan depth + --preferences "Prefer concise, British English." \ # prior + --nights 3 +``` + +All of this is exercised by the deterministic test suite (29 tests) and +validated on real Claude + Codex (see `real_api_results.md` / `FINAL_REPORT.md`).