feat(sleep): benchmark sweep + report tooling; override-aware reflect prompt
- sweep.py: run many (backend, model, seed, transfer-pair) configs sequentially,
append each result to JSONL incrementally (resumable, interrupt-safe).
- report.py: render the sweep JSONL into a presented Markdown scorecard with
direct-improvement and cross-model-transfer tables.
- reflect prompt now tells the optimizer its edits are APPENDED (can't delete the
base skill text), so on a conflict it must write a forceful OVERRIDE rule.
Diagnosed from a real failure: thorough-analyst (needs <=1200 chars) kept its
edits rejected because the base "be exhaustive" line won; a verified override
("HARD LIMIT ... supersedes") makes Haiku obey (1194/880 chars -> hard=1.0).
Co-Authored-By: Claude Opus 4 <noreply@anthropic.com>
This commit is contained in:
@@ -331,7 +331,13 @@ class CliBackend(Backend):
|
||||
f"{target} document so it stops failing. Each edit MUST be a short, "
|
||||
"GENERAL, reusable rule or preference (never task-specific, never an "
|
||||
"answer to a single task). If exact failing criteria are listed, your "
|
||||
"edits MUST make future outputs satisfy every one of them. "
|
||||
"edits MUST make future outputs satisfy every one of them.\n"
|
||||
"IMPORTANT: your edits are APPENDED to a 'Learned preferences' block; "
|
||||
"you CANNOT delete the existing instructions above. If the current "
|
||||
f"{target} text conflicts with a criterion (e.g. it says 'be exhaustive' "
|
||||
"but outputs must be under a character limit), write an explicit, "
|
||||
"forceful OVERRIDE rule that says it supersedes the conflicting "
|
||||
"instruction. "
|
||||
'Return ONLY a JSON array: '
|
||||
'[{"op":"add|replace|delete","content":"<rule>","anchor":"<text to replace/delete, optional>","rationale":"<why>"}].\n\n'
|
||||
f"# Current {target}\n{cur_doc}\n"
|
||||
|
||||
Reference in New Issue
Block a user