docs: sync documentation with post-v0.2 changes

This commit is contained in:
Yif-Yang
2026-07-14 17:11:40 +00:00
parent efb30b4bcc
commit f31bf8c06b
44 changed files with 2285 additions and 2025 deletions
+39 -17
View File
@@ -4,11 +4,11 @@
local coding agent a nightly **sleep cycle** that reviews your past sessions, replays
your recurring tasks on your own API budget, and consolidates what it learns into
**validated** long-term memory and skills — behind a held-out gate, staged for your
review. The agent gets better the more you use it, with **no weight training** and
**zero inference-time overhead**.
review. It requires **no weight training** and adds no separate optimization loop to
normal agent requests.
> **Preview.** This is an early preview we are actively iterating on; interfaces and
> defaults may change. The engine lives in the top-level [`skillopt_sleep/`](../../skillopt_sleep)
> defaults may change. The engine lives in the top-level [`skillopt_sleep/`](https://github.com/microsoft/SkillOpt/tree/main/skillopt_sleep)
> package with **zero dependency** on the paper's `skillopt/` code (the validation gate
> is vendored).
@@ -26,29 +26,49 @@ It synthesizes **SkillOpt** (validation-gated bounded text edits), **Claude Drea
(offline consolidation; review-then-adopt), and the **agent-sleep** idea (short-term
experience → long-term competence).
> **Data boundary.** Harvesting is local and read-only. The `mock` backend makes no
> provider calls. A real backend, however, sends truncated excerpts from harvested
> sessions and derived tasks to the provider you select for mining, replay, judging,
> and reflection. Outbound prompts are not currently guaranteed to be secret-free;
> review your transcript source and provider policy before running on sensitive
> projects. For a reviewable workflow, harvest to a task file, inspect/redact it, mark
> it `"reviewed": true`, and then replay that file with the real backend.
## How to use it
### Quickest path: the `skillopt-sleep` CLI (pip)
```bash
pip install skillopt # installs the engine + the `skillopt-sleep` command
skillopt-sleep dry-run # harvest + mine + replay, report only (changes nothing)
skillopt-sleep dry-run # harvest + mine + replay, report only; stages nothing
skillopt-sleep run # a full nightly cycle; the proposal is staged for review
skillopt-sleep status # show state + the latest staged proposal
skillopt-sleep adopt # apply the latest staged proposal
skillopt-sleep schedule # install a nightly cron entry for this project
```
The per-agent plugin shells below (Claude Code / Codex / Copilot) still come from the
repo; the CLI above is the standalone, pip-only way to run a cycle.
> **Version note.** This page tracks `main`. PyPI 0.2.0 provides the base
> commands above. Sleep handoff, non-Azure OpenAI-compatible endpoints, and
> `--preferences` landed later and require a source install from `main` until
> the next release.
One engine, thin per-agent shells (see [`plugins/`](../../plugins)):
The per-agent integrations below still come from the repo; the CLI above is the
standalone, pip-only way to run a cycle. Claude Code, Codex, Copilot, and Devin wrap
the shared engine. OpenClaw is a separate reference adaptation and has its own setup.
One engine, thin per-agent shells (see [`plugins/`](https://github.com/microsoft/SkillOpt/tree/main/plugins)):
| Platform | Folder | Install |
|---|---|---|
| **Claude Code** | [`plugins/claude-code`](../../plugins/claude-code) | `/plugin marketplace add ./plugins/claude-code``/skillopt-sleep` |
| **Codex** | [`plugins/codex`](../../plugins/codex) | `bash plugins/codex/install.sh``skillopt-sleep` skill |
| **Copilot** | [`plugins/copilot`](../../plugins/copilot) | register `plugins/copilot/mcp_server.py` as an MCP server |
| **Claude Code** | [`plugins/claude-code`](https://github.com/microsoft/SkillOpt/tree/main/plugins/claude-code) | `/plugin marketplace add ./plugins/claude-code``/skillopt-sleep` |
| **Codex** | [`plugins/codex`](https://github.com/microsoft/SkillOpt/tree/main/plugins/codex) | `bash plugins/codex/install.sh``skillopt-sleep` skill |
| **Copilot** | [`plugins/copilot`](https://github.com/microsoft/SkillOpt/tree/main/plugins/copilot) | register `plugins/copilot/mcp_server.py` as an MCP server |
| **Devin** | [`plugins/devin`](https://github.com/microsoft/SkillOpt/tree/main/plugins/devin) | register `plugins/devin/mcp_server.py` as an MCP server |
| **OpenClaw** | [`plugins/openclaw`](https://github.com/microsoft/SkillOpt/tree/main/plugins/openclaw) | adapt the reference wrapper and paths for your installation |
To use DeepSeek, vLLM, Ollama, or another Chat Completions server, see
**[OpenAI-compatible endpoints](openai-compatible-endpoints.md)**. That guide also
documents the separate HTTPS-only boundary for Azure managed-identity credentials.
Deterministic proof (no API key):
`python -m skillopt_sleep.experiments.run_experiment --persona researcher --assert-improves`.
@@ -71,11 +91,12 @@ correctness signal; the validation gate still governs what ships.
> scaling, and the dream-diversity ablation — are in
> [`docs/sleep/RESULTS.md`](RESULTS.md).** The highlights:
**Protocol (identical for every row below).** 5 nights × 10 new real "today" tasks
per night; the full held-out **test** split is scored before night 1 (baseline) and
after night 5 (after); optimizer = GPT-5.5; single seed (42); run through the exact
shipped engine (`skillopt_sleep.dream.dream_consolidate`). Numbers are absolute
held-out accuracy; **Δ** = `after baseline` in percentage points.
**Controlled experiment recipe (not the shipping CLI defaults).** 5 nights × 10 new
real "today" tasks per night; the full held-out **test** split is scored before night
1 (baseline) and after night 5 (after); optimizer = GPT-5.5; single seed (42). The
experiments use the shipped consolidation and gate components, while the nightly CLI
and benchmark harnesses remain separate entry points. Numbers are absolute held-out
accuracy; **Δ** = `after baseline` in percentage points.
**(a) End-to-end on real agents — [gbrain-evals](https://github.com/garrytan/gbrain-evals) `skillopt-v1`.**
Deficient seed skills go **0.00 → 1.00** on the held-out set with **both Claude Code
@@ -106,5 +127,6 @@ gate keeps the worst case bounded; keep it **on** by default.
## Learn more
Full reference (pipeline, the three plugins, the experience-replay knobs) is in the
**[Documentation & Reproduction Guide](https://microsoft.github.io/SkillOpt/docs/guideline.html#sleep)**.
See the [SkillOpt documentation index](../index.md), the
[CLI reference](../reference/cli.md), and the integration-specific READMEs under
[`plugins/`](https://github.com/microsoft/SkillOpt/tree/main/plugins).
+23 -17
View File
@@ -2,8 +2,9 @@
This is the evidence behind SkillOpt-Sleep: does a nightly, offline sleep cycle
actually make a *deployed* agent better, and is it safe to run unattended? We
answer with a controlled deployment-scale study — the same protocol the plugin
runs in production, scored on full held-out test sets.
answer with a controlled deployment-scale study built from the same shipped
consolidation and gate components. Its multi-night benchmark recipe is an
experiment configuration, not the default configuration of the nightly CLI.
## Setup
@@ -11,9 +12,10 @@ runs in production, scored on full held-out test sets.
**10 new real "today" tasks**; the skill carries over and is refined night to
night. The full held-out **test** split is scored before night 1 (*baseline*) and
after night 5 (*after*); **Δ = after baseline** in percentage points. Optimizer
model = **GPT-5.5**; single seed (42); every number is produced by the exact
shipped engine `skillopt_sleep.dream.dream_consolidate` (the experiment harness and
the plugin cycle call the same function).
model = **GPT-5.5**; single seed (42). The measurements use the shipped replay,
consolidation, and gate implementations. The nightly CLI and the checked-in
benchmark convenience harnesses are separate entry points and do not all call one
shared wrapper function.
**Benchmarks** (real evaluators, not format heuristics):
@@ -106,27 +108,31 @@ Replay-policy ablation (SearchQA, GPT-5.5):
| Replay policy | Gate-free Δ | Gated Δ |
|---|---|---|
| none (tonight's tasks only) | +3.9 | +2.0 |
| **recall k=10 (shipped default-able)** | +5.1 | +4.4 |
| **recall k=10 (opt-in experiment)** | +5.1 | +4.4 |
| cumulative (full history) | +4.8 | +6.0 |
Recall captures most of cumulative's benefit at a fraction of the per-night cost.
---
## 4. Default hyperparameters are the sweet spot
## 4. Sensitivity around the experiment recipe
We swept `dream_factor`, `rollouts`, `per_night`, and `nights` on the nano cell
(SearchQA, gated) to verify the shipped defaults are well-tuned:
(SearchQA, gated) around the study recipe: `dream_factor=2`, `rollouts=5`,
`per_night=10`, and `nights=5`. These are **experiment values**, not the shipping
defaults (`dream_factor=0`, `dream_rollouts=1`, and `recall_k=0`):
| Variant | Δ | vs default (+11.9) |
| Variant | Δ | vs experiment baseline (+11.9) |
|---|---|---|
| dream_factor=4 (default 2) | +8.8 | 3.1 |
| rollouts=10 (default 5) | +9.5 | 2.4 |
| per_night=15 (default 10) | +2.7 | 9.2 |
| nights=8 (default 5) | +9.5 | 2.4 |
| dream_factor=4 (baseline 2) | +8.8 | 3.1 |
| rollouts=10 (baseline 5) | +9.5 | 2.4 |
| per_night=15 (baseline 10) | +2.7 | 9.2 |
| nights=8 (baseline 5) | +9.5 | 2.4 |
Every direction away from the default hurts. This means users get the best result
**out of the box** without tuning — the recipe is robust by design.
Every tested direction away from that baseline reduced the measured gain in this
cell. The result supports that particular study recipe; it does not establish a
universal optimum. Shipping stays conservative, and users must opt in to additional
dream rollouts or recall after considering task quality and provider cost.
---
@@ -143,7 +149,7 @@ gains in Sections 12. Measured across an 18-cell deployment sweep (3 benchmar
|---|---|---|---|---|
| single-sample reflection (degraded) | 2.66 | **52.8** | 7 / 18 | 5 / 18 |
| diverse rollouts (K=5), no recall | +0.24 | 4.0 | 6 / 18 | 7 / 18 |
| **diverse rollouts + recall (shipped)** | **+0.53** | **2.4** | 7 / 18 | 7 / 18 |
| **diverse rollouts + recall (experiment recipe)** | **+0.53** | **2.4** | 7 / 18 | 7 / 18 |
The catastrophic 52.8 is removed **at its source** by diverse rollouts: the same
gate-free nano-SearchQA cell goes 0.554 → **0.586 (+2.7)** with no gate at all once
@@ -182,4 +188,4 @@ cross-verify each other's consolidated skills.
---
Back to the module overview: [`docs/sleep/README.md`](README.md) ·
full reference: [Documentation & Reproduction Guide](https://microsoft.github.io/SkillOpt/docs/guideline.html#sleep).
documentation index: [SkillOpt documentation](../index.md).
+76 -39
View File
@@ -1,11 +1,16 @@
# OpenAI-compatible endpoints for SkillOpt-Sleep (DeepSeek, local vLLM, …)
This document describes an enhancement to the `azure_openai` backend in
`skillopt_sleep/backend.py` that lets SkillOpt-Sleep drive **any
OpenAI-compatible chat-completions endpoint** — for example DeepSeek's hosted
This document describes the `azure_openai` backend in
`skillopt_sleep/backend.py`, which can drive servers that implement the expected
OpenAI-compatible Chat Completions request shape — for example DeepSeek's hosted
API or a self-hosted vLLM/Ollama server — in addition to native Azure OpenAI
deployments. It also documents a concrete end-to-end integration: running the
nightly sleep cycle inside the Antigravity IDE against DeepSeek.
deployments. The included runner is a sanitized unattended-launch example that
was originally used alongside Antigravity; it is not an Antigravity transcript
integration.
> **Version requirement.** This capability landed after v0.2.0. Until the next
> release, install SkillOpt from the latest `main`; the current PyPI 0.2.0
> package does not provide this compatible-endpoint path.
## What changed
@@ -32,15 +37,17 @@ is unchanged:
every rollout `0.0` with no diagnostic.)
4. **Managed-identity credential guard.** The managed-identity path attaches an
Azure AD bearer token to every request. If a custom endpoint outside
`*.openai.azure.com` / `*.cognitiveservices.azure.com` is configured without
explicit compat auth, the backend now raises a clear `ValueError` instead of
sending Azure credentials to an arbitrary host.
Azure AD bearer token to every request. It therefore accepts only an **HTTPS**
endpoint whose hostname ends in `*.openai.azure.com` or
`*.cognitiveservices.azure.com`. An HTTP endpoint — even one with an
Azure-looking hostname — and any host outside those suffixes are rejected
before a credential-bearing client is created.
5. **Provider-neutral request shape.** In compat mode the backend sends only the
standard OpenAI-compatible contract (`model`, `messages`, `max_tokens`).
Provider-specific request fields are **opt-in** via environment variables
(below) — nothing is inferred from model-name substrings.
(below) and are attached only in compat mode — nothing is inferred from
model-name substrings, and the native Azure request remains unchanged.
6. **Reliable error state.** `_call()` records the last exception in
`self.last_call_error` (surfaced in `diagnostics.json`), clears it when a
@@ -57,16 +64,34 @@ sleep cycle):
| Variable | Meaning |
|---|---|
| `AZURE_OPENAI_AUTH_MODE` | `openai_compatible` (or `compat`/`openai`) selects the plain OpenAI client. Unset/other = Azure managed identity (default). |
| `AZURE_OPENAI_ENDPOINT` | Base URL of the server, e.g. `https://api.deepseek.com`. |
| `AZURE_OPENAI_API_KEY` | API key sent by the compat client. |
| `AZURE_OPENAI_ENDPOINT` | Base URL of the server, e.g. `https://api.deepseek.com`. Azure managed identity requires HTTPS plus an approved Azure hostname. |
| `AZURE_OPENAI_API_KEY` | API key sent by the compat client to the configured base URL. |
| `SKILLOPT_SLEEP_COMPAT_MAX_TOKENS` | Optional int (default `8192`): `max_tokens` sent in compat mode. |
| `SKILLOPT_SLEEP_CHAT_EXTRA_BODY` | Optional JSON object passed as `extra_body` for provider-specific fields. |
| `SKILLOPT_SLEEP_CHAT_EXTRA_BODY` | Optional JSON object passed as `extra_body` for provider-specific fields in compat mode only. It is ignored in native Azure mode. |
## Data and transport boundaries
- Harvesting reads local transcripts without modifying them, and the `mock`
backend makes no provider calls. A real backend sends **truncated transcript
excerpts and derived task content** to the selected provider for mining,
replay, judging, and reflection.
- Outbound prompts are not currently guaranteed to be free of secrets. Review
the provider's data policy and avoid a third-party endpoint for sensitive
transcripts unless you have first inspected and redacted the task material.
One reviewable path is `skillopt-sleep harvest --output tasks.json`, followed
by a reviewed `--tasks-file` run.
- Use HTTPS for every remote compatible provider. Plain HTTP is appropriate only
for an explicitly trusted loopback development server such as
`http://127.0.0.1:8000/v1`; the compat client sends its API key to the configured
URL.
- Azure managed-identity credentials have the stricter invariant described
above: HTTPS **and** an approved Azure hostname are both mandatory.
## How to use it
```bash
export AZURE_OPENAI_AUTH_MODE=openai_compatible
export AZURE_OPENAI_ENDPOINT=https://api.deepseek.com # no /v1, no trailing path
export AZURE_OPENAI_ENDPOINT=https://api.deepseek.com # DeepSeek base URL
export AZURE_OPENAI_API_KEY=sk-... # your provider key
# DeepSeek reasoning models: enable the thinking channel (opt-in, not inferred)
@@ -79,37 +104,49 @@ skillopt-sleep run \
--project /path/to/your/project
```
The same pattern works for any OpenAI-compatible server — point
`AZURE_OPENAI_ENDPOINT` at it, set a matching `--model`, and omit
`SKILLOPT_SLEEP_CHAT_EXTRA_BODY` unless your provider needs extra request
fields.
The same pattern works for a server that implements this Chat Completions
contract: point `AZURE_OPENAI_ENDPOINT` at the provider-specific base URL, set a
matching `--model`, and omit `SKILLOPT_SLEEP_CHAT_EXTRA_BODY` unless the provider
needs extra request fields. Self-hosted vLLM and Ollama commonly use a `/v1` base
path, for example `http://127.0.0.1:8000/v1` or
`http://127.0.0.1:11434/v1`.
## End-to-end integration: Antigravity + DeepSeek
`--project` selects the project/transcript scope and the project `CLAUDE.md`; it
does **not** by itself select an arbitrary project `SKILL.md`. Pass
`--target-skill-path path/to/SKILL.md` when a specific skill is the optimization
target. Without that flag, SkillOpt-Sleep uses its configured managed skill.
The [`examples/`](examples/) directory contains a sanitized reference of how this
was wired into the [Antigravity](https://antigravity.google/) agent IDE so the
sleep cycle runs unattended:
## Unattended runner example (originally used with Antigravity)
The [`examples/`](https://github.com/microsoft/SkillOpt/tree/main/docs/sleep/examples) directory contains a sanitized reference for running
the compatible backend unattended:
- **`examples/runner.py`** — a thin launcher that loads a provider key from an
`.env` file, exports the variables above, invokes `skillopt-sleep run` with
the DeepSeek backend, and **exits with the child's return code** so
supervisors see failures as failures. It also implements a `session-end` hook
that appends task-outcome metadata to a rollout-evidence log (wired to
Antigravity's `Stop` hook) so future nights have richer sessions to mine.
supervisors see failures as failures. Its `session-end` action writes a small
local rollout-evidence event as an example hook target.
- **`examples/watchdog.py`** — a minimal supervisor loop that invokes the runner
on a fixed interval (e.g. every 4 hours) and logs non-zero exits as failures.
On Windows this is registered as a Scheduled Task so it survives logout; on
Linux/macOS a `systemd` timer or cron entry serves the same role.
### Verified result
The current engine does **not** read `brain/rollout-evidence.jsonl`, and it does
not harvest Antigravity transcripts. That hook output is illustrative metadata,
not additional training evidence. A real run must use a supported Claude
Code/Codex transcript source or a reviewed task file converted by the operator.
On a Windows 11 host, driving the cycle against `deepseek-v4-pro` in
`openai_compatible` mode:
### Contributor-reported validation
The contributor reported the following results from a private Windows 11 setup
driving the cycle against `deepseek-v4-pro` in `openai_compatible` mode. They are
useful integration evidence, but the private session set is not a reproducible
benchmark bundled with this repository:
- A direct backend smoke test returns a live completion (no `404`,
`last_call_error` empty, client type `OpenAI`).
- A full nightly cycle mined tasks from real IDE sessions and the held-out
validation gate moved from `0.250 → 1.000`, **accepting** a DeepSeek-authored
- A full nightly cycle using the configured session source moved the held-out
validation gate from `0.250 → 1.000`, **accepting** a DeepSeek-authored
skill edit (`accept_new_best`). `diagnostics.json` for that night reports
`"backend": "azure_openai"` with a non-empty token count and an empty
`call_error` — i.e. a genuine optimization night, versus the prior all-`0.0`
@@ -124,13 +161,13 @@ Deterministic no-network coverage for the new behavior lives in
endpoint/auth guard, request kwargs, retry error-state, empty-response
diagnostics, and runner exit-code propagation).
## A note on Gemini (optional, unverified fallback)
## Unsupported Gemini proxy branch in the example
`examples/runner.py` also contains a fallback branch that, when only a Gemini key
is present, routes the **`claude` CLI backend** through a local
Anthropic-compatible proxy (e.g. [LiteLLM](https://github.com/BerriAI/litellm) on
`http://127.0.0.1:4000`) by setting `ANTHROPIC_BASE_URL`/`ANTHROPIC_API_KEY`.
There is **no native Gemini backend** in SkillOpt, and this proxy path was not
independently validated in this work — it is included only as a configuration
example. The verified, supported path in this document is DeepSeek via
`openai_compatible` mode. Treat the Gemini branch as illustrative, not tested.
`examples/runner.py` still contains an illustrative branch that routes the
**`claude` CLI backend** through a loopback Anthropic-compatible proxy such as
[LiteLLM](https://github.com/BerriAI/litellm). It is not a native Gemini backend,
has no validated model mapping in this example, and is not part of the supported
path documented here. The sample currently enters that branch whenever no
DeepSeek key is found, so a production adaptation should remove it or replace it
with an explicit opt-in, a separately configured model, and a trusted isolated
loopback proxy. Do not treat this branch as tested Gemini support.