Addresses maintainer review on the OpenAI-compatible endpoints PR:
1. CLI: accept --backend azure_openai in skillopt_sleep/__main__.py (the
documented command was rejected by the argparse choices).
2. Example runner: exit with the child's return code so watchdog/supervisors
see a failed sleep run as a failure.
3. Error state: clear last_call_error when a retry recovers; set an explicit
"empty response on all N attempts" diagnostic when every attempt returns
empty text.
4. Security guard: the managed-identity path now refuses to send an Azure AD
bearer token to any endpoint outside *.openai.azure.com /
*.cognitiveservices.azure.com — a custom endpoint requires explicit
AZURE_OPENAI_AUTH_MODE=openai_compatible + API key.
5. Provider-neutral requests: compat mode sends only the standard contract
(max_tokens, default 8192 via SKILLOPT_SLEEP_COMPAT_MAX_TOKENS);
provider-specific body fields are opt-in via SKILLOPT_SLEEP_CHAT_EXTRA_BODY
(JSON) — the deepseek model-name inference is removed.
6. Docs: removed the unimplemented OPTIMIZER_*/TARGET_* env-var claim; added a
configuration reference matching the implementation exactly.
7. Tests: tests/test_azure_openai_compat.py — 17 deterministic no-network
unittest cases covering CLI acceptance, compat-vs-Azure client selection,
endpoint resolution, the credential guard, request kwargs (opt-in extra
body / token cap), retry-success error clearing, empty-response
diagnostics, and runner exit-code propagation.
Re-verified live against DeepSeek (deepseek-v4-pro, openai_compatible mode)
after the rework: client type OpenAI, completion returned, no error state.
Fix 1: Add [tool.setuptools.package-data] to pyproject.toml so that
all .md prompt files (generic prompts under skillopt/prompts/ and
env-specific prompts under skillopt/envs/*/prompts/) are included
in the wheel. Previously pip install from PyPI would fail with
FileNotFoundError on first prompt load.
Fix 2: Replace TemporaryDirectory context manager with mkdtemp +
shutil.rmtree(ignore_errors=True) in claude_backend._run_claude_print.
On Windows, the spawned claude/node subprocess still holds a handle
on the temp directory when the context manager tries to clean up,
causing WinError 32. Manual rmtree with ignore_errors=True works
on all Python versions >= 3.10.
Co-Authored-By: Claude <noreply@anthropic.com>
The LinearScheduler.known_decay_sequence test now reflects the new
t = (step-1)/(total_steps-1) contract where step 1 returns max_lr.
Updated docstrings for midpoint and early-step tests to match.
All 43 tests now pass against the updated formula.
LinearScheduler and CosineScheduler now guarantee step() returns max_lr
on the first call and min_lr on the total_steps-th call, using the
standard endpoint parameterisation t = (step-1)/(total_steps-1).
This resolves the inconsistent first-step semantics reported in review.
Let AzureOpenAIBackend drive any OpenAI-compatible chat-completions server
(DeepSeek, self-hosted vLLM/Ollama, ...) alongside native Azure deployments.
- __init__ resolves the endpoint as: explicit arg > AZURE_OPENAI_ENDPOINT env
> the built-in _AZURE_ENDPOINTS table (previously a non-Azure endpoint could
not be supplied at all).
- _get_client() builds a plain openai.OpenAI(base_url=...) client when
AZURE_OPENAI_AUTH_MODE=openai_compatible, matching the auth mode already
supported by the sibling skillopt/model/azure_openai.py. This avoids the
AzureOpenAI SDK rewriting request URLs with Azure-only ?api-version= query
params and deployment path segments, which non-Azure servers reject with 404.
- _call() sends max_tokens + extra_body thinking flag for deepseek* models and
records the last exception in self.last_call_error so a failed night is
diagnosable instead of collapsing to a silent empty->0 score.
Adds docs/sleep/openai-compatible-endpoints.md and sanitized example
runner/watchdog scripts documenting an Antigravity + DeepSeek integration.
The default managed-identity Azure path is unchanged.
Three filters to stop machine-generated prompts polluting the mined task
pool (they dominated ~80% of tasks on this machine):
- _is_meta_prompt: drop expanded slash-command bodies (<command-message>
tags or '# /' headers) — plugin self-invocations are not user intents
- _AGENT_SESSION_MARKERS + _is_agent_session: skip sessions whose first
prompt is another tool's agent brief (claude-mem observers, CLAUDE.md
critic sub-agents, SkillOpt-Sleep's own command body)
- load walk: skip <session>/subagents/ dirs and agent-*.jsonl files —
Agent-tool sidechain transcripts are Claude-authored, not user tasks
Verified: harvest went from 120 sessions / mostly-noise tasks to 55
sessions / 38 real user tasks.
Co-authored-by: codeL1985 <l@cypherlab.tech>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Two missing pieces in the minimax_chat dispatch chain:
1. minimax_backend.py did not define chat_optimizer or chat_optimizer_messages;
only chat_target / chat_target_messages existed. Any caller using
optimizer_backend='minimax_chat' would hit AttributeError or fall through
to _openai (Azure).
2. skillopt/model/__init__.py's chat_optimizer and chat_optimizer_messages
dispatchers checked claude_chat and qwen_chat but not minimax_chat,
so minimax_chat callers would silently fall through to _openai.chat_optimizer
(the Azure path), which fails with 'Azure OpenAI endpoint is not configured'
on any setup without AZURE_OPENAI_* env vars.
Adds chat_optimizer to minimax_backend.py (mirrors chat_target via
_chat_messages_impl) and minimax_chat branches to both
chat_optimizer / chat_optimizer_messages dispatchers.
Verified locally: 1-epoch training on a 4-item SearchQA-format dataset
went from '[skip] no usable patches — skill unchanged' (baseline fallback)
to a successful accept_new_best with success_patches=1 per step.
Co-authored-by: Mavis (MiniMax) <Mavis@MiniMax.local>
Co-authored-by: jc <jc@users.noreply.github.com>
The qwen_chat backend (the generic OpenAI-compatible client) hardcoded
max_tokens and always sent temperature, so reasoning models behind
OpenAI-compatible gateways (GPT-5.x, Claude Opus 4.8 via Azure/LiteLLM)
would 400.
- Add opt-in QWEN_CHAT_USE_MAX_COMPLETION_TOKENS (+ role variants) that
swaps the payload key max_tokens -> max_completion_tokens.
- Treat an explicit empty / none / off temperature as "omit" instead of
collapsing to the 0.7 default (via _resolve_temperature).
- Thread both through configure_qwen_chat / _update_config.
- Defaults unchanged; fully backward compatible. Adds 6 tests.
Fixes#127
Co-authored-by: Chirag Singhal <chirag127@users.noreply.github.com>
Adds --backend handoff: the engine runs all deterministic stages and
outsources attempt/judge/reflect to prompt/answer files an interactive
agent session fills between runs (exit 3 = pending batch, re-run to
resume). Deterministic replay + the prompt-hash answer cache make resume
stateless; sentinel detection aborts any call built from unanswered
output so placeholders never reach scores or staging. Session digests
and mined tasks are pinned per night (secret-redacted) so the sessions
answering prompts cannot shift the task set, and LLM mining is routed
through the same handoff files. Ships a /skillopt-sleep-handoff Claude
Code command that answers each prompt in a fresh-context subagent to
protect the held-out gate.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix(skillopt-sleep): surface Claude CLI spawn failures instead of silent zero scores
In _call and attempt_with_tools, the bare 'except Exception: return ""'
swallows FileNotFoundError (e.g. bare 'claude' on Windows with npm .cmd
shim) and any other spawn failure, returning an empty string that the
trainer treats as a legitimate model response that scores 0.0 everywhere.
Now the exception is caught explicitly: last_call_error is set, a
warning is logged, and the empty-string return is preserved for
backward compatibility of the control flow. This mirrors the pattern
from #92 (codex backend) which fixed the same class of 'dead CLI
masquerades as nothing to learn' bug.
Issue: #121
* test(skillopt-sleep): add tests verifying Claude CLI spawn failures are surfaced
Add two tests to TestClaudeCliBackendBare:
- test_spawn_failure_sets_last_call_error: _call sets last_call_error
and returns '' when subprocess.run raises FileNotFoundError.
- test_attempt_tools_spawn_failure_sets_last_call_error: same for
attempt_with_tools.
These prove the fix from the parent commit (surface spawn failures
instead of silently scoring 0) and guard against regressions.
* fix(backend): make ClaudeCliBackend work on Windows (.cmd shim + argv limit)
Every claude call failed with WinError 2 and was swallowed by the bare
except -> return '', so real-backend cycles silently scored 0.000/0.000
and produced zero edits.
- resolve claude_path via shutil.which: the npm-installed claude is a
.cmd shim that CreateProcess cannot resolve by bare name
- pass the prompt via stdin instead of argv: the .cmd shim routes
through cmd.exe whose command line caps at ~8K chars, which
reflect/judge prompts exceed
Verified: reflect() on a synthetic max_chars failure now returns a
concrete bounded edit via the real CLI (was [] before).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(backend): suppress console window flashes on Windows (CREATE_NO_WINDOW)
Every CLI subprocess (claude/codex/copilot) spawned from a console-less
parent allocates a visible console window on Windows. A cycle making
hundreds of calls strobes cmd windows and steals focus, making the
machine unusable while the engine runs. Pass CREATE_NO_WINDOW on all
CLI call sites; no-op on POSIX.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: codeL1985 <l@cypherlab.tech>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Add explicit assertions for held-out scores and gate actions to the verifier discipline test suite to strengthen its guarantees.
- Assert the concrete held-out baseline and candidate scores in test_gate_rejects_reward_hacking_edit.
- Add test_gate_accepts_beneficial_edit using MockBeneficialBackend to provide a paired case where an edit genuinely improves the held-out slice, expecting accepted=True and gate_action='accept_new_best'.
_load_yaml opened config files with the platform default (locale)
encoding instead of UTF-8. The shipped configs (e.g.
configs/_base_/default.yaml) contain UTF-8 non-ASCII characters
(em-dash, arrows, box-drawing), so load_config() raises
UnicodeDecodeError on any non-UTF-8 locale (e.g. Windows cp936/cp1252),
breaking scripts/train.py and scripts/eval_only.py before startup.
YAML is UTF-8 by spec, and the rest of the codebase already opens text
files with encoding="utf-8". Pass encoding="utf-8" here for parity.