From b0b62fcb864784a28543c8e9f52829958a59879b Mon Sep 17 00:00:00 2001 From: Cuzyoung Date: Wed, 10 Jun 2026 13:27:36 +0000 Subject: [PATCH] =?UTF-8?q?docs(readme):=20slim=20README=20=E2=80=94=20mov?= =?UTF-8?q?e=20install/quick-start/data/config=20details=20to=20the=20guid?= =?UTF-8?q?eline=20page?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit README now: badges + one-line pointer to docs/guideline.html, overview, demo, sleep section, extensibility pointers, WebUI launch, citation. All run-the-demo commands live in the guideline (which already covered install, credentials, training, eval, outputs, data prep, and config). Co-Authored-By: Claude Fable 5 --- README.md | 313 ------------------------------------------------------ 1 file changed, 313 deletions(-) diff --git a/README.md b/README.md index ef5428e..28c3da2 100644 --- a/README.md +++ b/README.md @@ -98,319 +98,6 @@ Deterministic proof (no API key): `python -m skillopt_sleep.experiments.run_expe --- -## Install - -### Requirements - -- Python 3.10+ - -### Option A: Install from PyPI - -```bash -pip install skillopt - -# With optional extras: -pip install skillopt[alfworld] # ALFWorld benchmark -pip install skillopt[webui] # Gradio monitoring dashboard -pip install skillopt[claude] # Claude model backend -``` - -### Option B: Install from source (for development) - -```bash -git clone https://github.com/microsoft/SkillOpt.git -cd SkillOpt -pip install -e . - -# For the ALFWorld benchmark (optional): -pip install -e ".[alfworld]" -alfworld-download -``` - -### Configure API Credentials - -```bash -cp .env.example .env -# Edit .env with your API credentials, then: -source .env -``` - -#### Azure OpenAI *(recommended)* - -```bash -export AZURE_OPENAI_ENDPOINT="https://your-resource.openai.azure.com/" -# Option 1: API key auth -export AZURE_OPENAI_API_KEY="your-key" -# Option 2: Azure CLI auth (no API key needed) -export AZURE_OPENAI_AUTH_MODE="azure_cli" -``` - -> **Note:** `AZURE_OPENAI_ENDPOINT` is required for all three modes (`api_key`, `azure_cli`, `openai_compatible`). Without it, all LLM calls will fail. - -#### OpenAI-compatible endpoints - -```bash -export AZURE_OPENAI_ENDPOINT="https://api.openai.com/v1" -export AZURE_OPENAI_API_KEY="sk-..." -export AZURE_OPENAI_AUTH_MODE="openai_compatible" -``` - -This routes all calls through the plain OpenAI Python client (no Azure auth, no `api-version` header). - -> **Note:** SkillOpt reuses the `AZURE_OPENAI_*` env var names even in this mode — there is no separate `OPENAI_API_KEY` knob. - -#### Anthropic Claude - -```bash -export ANTHROPIC_API_KEY="sk-ant-..." -``` - -#### Qwen *(local vLLM)* - -```bash -export QWEN_CHAT_BASE_URL="http://localhost:8000/v1" -export QWEN_CHAT_MODEL="Qwen/Qwen3.5-4B" -``` - -`qwen_chat` can also be used as the optimizer backend. When optimizer and -target should point to different local vLLM services, use the role-specific -settings: - -```bash -python scripts/train.py \ - --config configs/searchqa/default.yaml \ - --optimizer_backend qwen_chat \ - --target_backend qwen_chat \ - --optimizer_model Qwen/Qwen3.5-4B \ - --target_model Qwen/Qwen3.5-4B \ - --optimizer_qwen_chat_base_url http://localhost:8001/v1 \ - --target_qwen_chat_base_url http://localhost:8000/v1 -``` - -#### MiniMax - -```bash -export MINIMAX_BASE_URL="https://api.minimax.io/v1" -export MINIMAX_API_KEY="..." -export MINIMAX_MODEL="MiniMax-M2.7" -``` - ---- - -## Quick Start - -### Training - -```bash -# Minimal example — train on SearchQA: -python scripts/train.py \ - --config configs/searchqa/default.yaml \ - --split_dir /path/to/your/searchqa_split \ - --azure_openai_endpoint https://your-resource.openai.azure.com/ \ - --optimizer_model gpt-5.5 \ - --target_model gpt-5.5 - -# Train on LiveMathematicianBench: -python scripts/train.py \ - --config configs/livemathematicianbench/default.yaml \ - --split_dir /path/to/your/livemath_split \ - --azure_openai_endpoint https://your-resource.openai.azure.com/ \ - --optimizer_model gpt-5.5 \ - --target_model gpt-5.5 - -# Train on ALFWorld: -python scripts/train.py \ - --config configs/alfworld/default.yaml \ - --split_dir data/alfworld_path_split \ - --azure_openai_endpoint https://your-resource.openai.azure.com/ \ - --optimizer_model gpt-5.5 \ - --target_model gpt-5.5 -``` - -Key CLI arguments: - -| Argument | Description | Example | -|---|---|---| -| `--config` | Benchmark config YAML | `configs/searchqa/default.yaml` | -| `--split_dir` | Path to data split directory | `/path/to/split` | -| `--azure_openai_endpoint` | Azure OpenAI endpoint URL | `https://your-resource.openai.azure.com/` | -| `--optimizer_model` | Optimizer model deployment name | `gpt-5.5` | -| `--target_model` | Target model deployment name | `gpt-5.5` | -| `--num_epochs` | Number of training epochs | `4` | -| `--batch_size` | Batch size per step | `40` | -| `--workers` | Parallel rollout workers | `8` | -| `--out_root` | Output directory | `outputs/my_run` | - -### Eval Only - -Evaluate a trained skill on specific data splits without training: - -```bash -# Evaluate the packaged GPT-5.5 SearchQA skill on the test split: -python scripts/eval_only.py \ - --config configs/searchqa/default.yaml \ - --skill ckpt/searchqa/gpt5.5_skill.md \ - --split valid_unseen \ - --split_dir /path/to/searchqa_split \ - --azure_openai_endpoint https://your-resource.openai.azure.com/ - -# Evaluate on all splits (train + val + test): -python scripts/eval_only.py \ - --config configs/searchqa/default.yaml \ - --skill ckpt/searchqa/gpt5.5_skill.md \ - --split all \ - --split_dir /path/to/searchqa_split \ - --azure_openai_endpoint https://your-resource.openai.azure.com/ -``` - -To evaluate a skill produced by your own training run, replace `--skill` with that run's best-skill path, for example `outputs/my_run/best_skill.md`. - -| Split | Description | -|---|---| -| `valid_unseen` | Test set | -| `valid_seen` | Validation set | -| `train` | Training set | -| `all` | All splits combined (default) | - -### Output Structure - -Each training run writes to a structured output directory: - -``` -outputs// -├── config.json # Flattened runtime config -├── history.json # Per-step training history -├── runtime_state.json # Resume checkpoint -├── best_skill.md # Best validated skill document -├── skills/skill_vXXXX.md # Skill snapshot per step -├── steps/step_XXXX/ # Per-step artifacts (patches, evals) -├── slow_update/epoch_XX/ # Slow update logs -└── meta_skill/epoch_XX/ # Meta skill logs -``` - -Re-running the same command auto-resumes from the last completed step. - -### Pretrained Skill Artifacts - -We provide a subset of the paper's main Table 1 GPT-5.5 optimized skills in -[`ckpt/`](ckpt/) as reference artifacts. Use them with `scripts/eval_only.py` -to evaluate the provided skills on a matching data split without re-running -training. See [`ckpt/README.md`](ckpt/README.md) for the full per-benchmark -command. This is the first artifact batch; we plan to continue uploading -the remaining optimized skills and benchmark split manifests as they are -cleaned and verified. - ---- - -## Data Preparation - -### Directory layout - -SkillOpt expects data in a **split directory** with `train/`, `val/`, `test/` subdirectories, each containing a JSON file (e.g., `items.json`): - -``` -data/my_split/ -├── train/items.json -├── val/items.json -└── test/items.json -``` - -Each JSON file is an array of task items. The required fields depend on the benchmark. For example, SearchQA items look like: - -```json -[ - { - "id": "unique_item_id", - "question": "Who wrote the novel ...", - "context": "[DOC] relevant passage text ...", - "answers": ["expected answer"] - } -] -``` - -See `skillopt/envs//dataloader.py` for the exact format each benchmark expects. - -> **Note:** Most benchmark datasets are not included in this repository. Prepare your own data following the format above. The exact SearchQA split used in the paper is provided at [`data/searchqa_id_split/`](data/searchqa_id_split) (400 train / 200 val / 1400 test). We are preparing the remaining benchmark split manifests for upload. - -### Supported Benchmarks - -| Benchmark | Type | Config | -|---|---|---| -| SearchQA | QA | `configs/searchqa/default.yaml` | -| ALFWorld | Embodied agent | `configs/alfworld/default.yaml` | -| DocVQA | Document QA | `configs/docvqa/default.yaml` | -| LiveMathematicianBench | Math | `configs/livemathematicianbench/default.yaml` | -| SpreadsheetBench | Code generation | `configs/spreadsheetbench/default.yaml` | -| OfficeQA | Tool-augmented QA | `configs/officeqa/default.yaml` | - ---- - -## Configuration - -### Default settings and paper-reproduction knobs - -`configs/_base_/default.yaml` is the single source of truth for SkillOpt's -runtime knobs. Out of the box, every included benchmark config inherits -from it and keeps the paper protocol visible: 4 epochs, rollout batch 40, -reflection minibatch 8, textual learning rate 4 with cosine decay, strict -hard validation gating, and slow-update + meta-skill enabled. One detail to -watch is slow-update acceptance: the current `main` default is the newer -post-submission force-accept mode, while the paper protocol and the -paper-aligned skills under `ckpt/` use the gated semantics described in -paper Section 3.6. - -### Slow-update acceptance mode - -The epoch-boundary slow / meta update can be applied two ways, controlled -by `optimizer.slow_update_gate_with_selection`: - -```yaml -optimizer: - slow_update_gate_with_selection: false # current main default -``` - -- **`false`** *(current `main` default)*: force-accept. The - slow-update guidance is injected into both `current_skill` and - `best_skill` unconditionally at the epoch boundary. This is the newer - post-submission behavior on `main`. -- **`true`** *(paper / ckpt-skill reproduction)*: gated, matching paper - Section 3.6 verbatim. The slow-update candidate is evaluated on the - selection split and accepted only if it passes the same validation gate - as a step-level edit. Use this setting when re-running optimization to - match the paper protocol and the provenance of the provided `ckpt/` skills. - -The trainer prints which mode is active at startup -(`[slow update] acceptance=...`). See issue #22 for the discussion that -led to the flag. - -### Gate metric (`hard` / `soft` / `mixed`) - -The validation gate compares candidate vs. current skills on the selection -split using `gate_metric`: - -- **`hard`** *(default, paper)*: exact-match accuracy, strictly greater - than the current score is required. -- **`soft`**: per-item soft / partial-credit score. Useful when the - selection split is small (e.g. ≤10 items) and the reward is continuous, - where the discrete hard gate often rejects every candidate. -- **`mixed`**: weighted average, `(1 - w) * hard + w * soft`, with `w` - set by `gate_mixed_weight` (default `0.5`). - -Default is `hard`. Use the optional feature config below to switch. - -### Optional feature configs - -These are **not** default SkillOpt settings — they are optional feature configs -contributed by users for specific scenarios. The paper-reported numbers -were obtained with the default settings, not these. - -- **[`configs/features/soft_gate.yaml`](configs/features/soft_gate.yaml)** - *(PR #25, contributed by [@lvbaocheng](https://github.com/lvbaocheng))* — - switches `gate_metric` to `soft` (or `mixed`). See the comment at the - top of the file for when to use and when not to. - ---- - ## Extensibility & WebUI ### Adding a new backend