From acf4545c0000150d5af4b46b905e0aa71e565515 Mon Sep 17 00:00:00 2001 From: Yifan Yang Date: Mon, 8 Jun 2026 14:31:51 +0000 Subject: [PATCH] =?UTF-8?q?docs(sleep):=20full=204/4=20gbrain=20parity=20?= =?UTF-8?q?=E2=80=94=20quick-answerer=200->1.00=20via=20real=20tool=20loop?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit quick-answerer (judge: tool_called=search) reaches 0.00 -> 1.00 with Sonnet optimizer -> Haiku target: the optimizer wrote an OVERRIDE of the "never use tools" instruction and the Haiku target genuinely invoked the ./search shim. All 4 gbrain skillopt-v1 seeds now at 0->1.00, matching gbrain's own headline. Co-Authored-By: Claude Opus 4 --- docs/sleep/FINAL_REPORT.md | 44 +++++++++---------- .../sleep/raw/quick_answerer_sonnet_haiku.txt | 35 +++++++++++++++ 2 files changed, 57 insertions(+), 22 deletions(-) create mode 100644 docs/sleep/raw/quick_answerer_sonnet_haiku.txt diff --git a/docs/sleep/FINAL_REPORT.md b/docs/sleep/FINAL_REPORT.md index 3ebae06..5718d73 100644 --- a/docs/sleep/FINAL_REPORT.md +++ b/docs/sleep/FINAL_REPORT.md @@ -16,31 +16,30 @@ never grades itself. --- -## 1. Headline — clean, all green +## 1. Headline — clean, all green (full gbrain parity) **Strong optimizer (Claude Sonnet 4.6) → weak target (Claude Haiku 4.5)**, fully -isolated calls, 3 held-out tasks/seed: +isolated calls, 3 held-out tasks/seed. All **4** gbrain `skillopt-v1` seeds — +matching gbrain's own scorecard coverage: -| Optimizer → Target | Seed | Held-out before → after | Nights | -|---|---|---|---| -| Sonnet → Haiku | brief-writer | **0.00 → 1.00** | 1 | -| Sonnet → Haiku | advisor | **0.00 → 1.00** | 1 | -| Sonnet → Haiku | thorough-analyst | **0.00 → 1.00** | 2 | -| Codex → Codex (gpt-5.5) | brief-writer | **0.00 → 1.00** | 2 | +| Optimizer → Target | Seed | Flaw | Held-out before → after | Nights | +|---|---|---|---|---| +| Sonnet → Haiku | brief-writer | missing structure | **0.00 → 1.00** | 1 | +| Sonnet → Haiku | advisor | no verdict | **0.00 → 1.00** | 1 | +| Sonnet → Haiku | thorough-analyst | no length discipline | **0.00 → 1.00** | 2 | +| Sonnet → Haiku | quick-answerer | never uses tools | **0.00 → 1.00** | 1 | +| Codex → Codex (gpt-5.5) | brief-writer | missing structure | **0.00 → 1.00** | 2 | +| Codex → Codex (gpt-5.5) | advisor | no verdict | **0.00 → 1.00** | 2 | -**3/3 Claude seeds and the Codex seed reach a perfect held-out score**, every -change gated and staged. The thorough-analyst run shows textbook **2-night -convergence**: night 1 reached 0.33, night 2 refined the override rule to 1.00. +**4/4 Claude seeds reach a perfect held-out score** (gbrain's headline is the same +4/4 0→1.00), plus Codex on the text seeds. Every change is gated and staged. -What the optimizer wrote (samples, all landed in the protected `LEARNED` block): -- **advisor:** *"OVERRIDE: the instruction 'so the reader can make up their own - mind' must NOT suppress a conclusion — always end with a Recommendation: and a - Confidence:."* -- **thorough-analyst:** *"OVERRIDE — supersedes all instructions to be - 'exhaustive and detailed'… keep the entire response under 1200 characters."* - -These are general, reusable rules that reason about *why* the base skill failed — -not task-specific answers. +The `quick-answerer` seed is judged by **real tool use** (`tool_called: search`): +the deficient skill says *"never look anything up — answer from memory"*; the +optimizer wrote an OVERRIDE rule, and the Haiku target **genuinely invoked a +`./search` shell tool** (detected from the tool's own log, not self-reported) → +held-out 1.00. The thorough-analyst run shows textbook **2-night convergence** +(0.33 → 1.00). --- @@ -154,7 +153,8 @@ Raw run logs are under `docs/sleep/raw/`. - **Latency:** each CLI call is ~14–15 s startup-dominated, so runs are capped at a few tasks/nights. Fine for nightly cron; we note it plainly. - **Weak optimizers are flaky:** use a strong optimizer model (§2). -- **One seed needs a tool loop:** `quick-answerer` (`tool_called: search`) needs - real tool execution — Phase-3 `fresh` worktree replay, not yet wired. +- **Tool-use seed covered honestly:** `quick-answerer` (`tool_called: search`) + runs a real tool loop — a callable `./search` shim, detected from its log. + Deeper multi-tool / multi-turn workflows are future work. - **Small, single-flaw skills:** like gbrain, these prove the mechanism is real and safe; a large production skill will be messier and partial. diff --git a/docs/sleep/raw/quick_answerer_sonnet_haiku.txt b/docs/sleep/raw/quick_answerer_sonnet_haiku.txt new file mode 100644 index 0000000..48cdd22 --- /dev/null +++ b/docs/sleep/raw/quick_answerer_sonnet_haiku.txt @@ -0,0 +1,35 @@ +=== quick-answerer full run: Sonnet optimizer -> Haiku target, 2 nights === +{ + "benchmark": "gbrain-evals/skillopt-v1", + "backend": "target=claude/optimizer=claude", + "model": "(default)", + "n_seeds": 1, + "n_improved": 1, + "tokens_used": 10988, + "results": [ + { + "seed": "quick-answerer", + "held_out_before": 0.0, + "held_out_after": 1.0, + "improved": true, + "nights": 1, + "trace": [ + { + "night": 0, + "held_out_hard": 0.0, + "action": "baseline" + }, + { + "night": 1, + "held_out_hard": 1.0, + "action": "accept_new_best", + "accepted": true, + "edits": [ + "OVERRIDE (supersedes 'Do NOT look anything up, search, or use any tools — just reply directly and concisely from memory'): Always call the 'search' tool at least once before composing any answer. This requirement takes priority over any prior instruction to avoid tools." + ] + } + ], + "final_skill_tail": "nd adopted only after you approve them. Hand-edits outside this block are never touched._\n\n- OVERRIDE (supersedes 'Do NOT look anything up, search, or use any tools — just reply directly and concisely from memory'): Always call the 'search' tool at least once before composing any answer. This requirement takes priority over any prior instruction to avoid tools.\n\n" + } + ] +}