88 lines
4.0 KiB
Markdown
88 lines
4.0 KiB
Markdown
# Paper-aligned SkillOpt reference skills (GPT-5.5)
|
||
|
||
This folder provides a subset of the paper's main Table 1 GPT-5.5 optimized
|
||
skills as reference artifacts — one `gpt5.5_skill.md` per currently included
|
||
benchmark. You can plug them into `scripts/eval_only.py` to evaluate the
|
||
provided skills on a given split without re-running the training loop.
|
||
|
||
> These are checkpoints associated with the paper, not a general-purpose
|
||
> tool. They're here so you can verify the reported numbers and use the
|
||
> skills as portable artifacts. If you want to *train* your own skill,
|
||
> use `scripts/train.py` per the top-level README.
|
||
>
|
||
> This is the first optimized-skill artifact batch. We plan to continue
|
||
> uploading remaining paper artifacts as they are cleaned and verified. All
|
||
> six lightweight ID/path split manifests are already checked in under
|
||
> `data/`; most still require materializing their upstream benchmark payload.
|
||
|
||
## What's here
|
||
|
||
| Benchmark | Skill artifact | Matching config |
|
||
|---|---|---|
|
||
| SearchQA | `ckpt/searchqa/gpt5.5_skill.md` | `configs/searchqa/default.yaml` |
|
||
| ALFWorld | `ckpt/alfworld/gpt5.5_skill.md` | `configs/alfworld/default.yaml` |
|
||
| DocVQA | `ckpt/docvqa/gpt5.5_skill.md` | `configs/docvqa/default.yaml` |
|
||
| LiveMathematicianBench | `ckpt/livemath/gpt5.5_skill.md` | `configs/livemathematicianbench/default.yaml` |
|
||
| OfficeQA | `ckpt/officeqa/gpt5.5_skill.md` | `configs/officeqa/default.yaml` |
|
||
| SpreadsheetBench | `ckpt/spreadsheetbench/gpt5.5_skill.md` | `configs/spreadsheetbench/default.yaml` |
|
||
|
||
Each file is a plain Markdown skill document (~2k–13k chars). It contains a
|
||
protected `SLOW_UPDATE` section at the end that holds epoch-wise
|
||
longitudinal guidance — that's expected, not a formatting issue.
|
||
|
||
## How to evaluate a provided skill
|
||
|
||
`scripts/eval_only.py` runs a single skill against a data split without
|
||
invoking the optimizer. Example for SearchQA against the test split:
|
||
|
||
```bash
|
||
# The checked-in SearchQA split is ID-only; materialize full examples first.
|
||
python -m pip install -e ".[searchqa]"
|
||
python scripts/materialize_searchqa.py
|
||
|
||
python scripts/eval_only.py \
|
||
--config configs/searchqa/default.yaml \
|
||
--skill ckpt/searchqa/gpt5.5_skill.md \
|
||
--split valid_unseen \
|
||
--split_dir data/searchqa_split \
|
||
--azure_openai_endpoint https://your-resource.openai.azure.com/ \
|
||
--azure_openai_auth_mode api_key \
|
||
--target_model gpt-5.5
|
||
```
|
||
|
||
Substitute the benchmark, config, skill path, and `--split_dir` to evaluate
|
||
any of the other five. `--split valid_unseen` is the test split, `valid_seen`
|
||
is the selection / validation split, `train` is the training split, and
|
||
`all` runs all three.
|
||
|
||
## On comparing to the paper numbers
|
||
|
||
To compare against the paper-reported cells, use the same dataset split and
|
||
scorer. SearchQA's ID manifest is checked in at `data/searchqa_id_split/` (400
|
||
train / 200 selection / 1400 test); the materializer writes the runnable
|
||
payload to `data/searchqa_split/`. All six lightweight split manifests are
|
||
checked in under `data/`. ALFWorld's manifest records game-file paths; the
|
||
other ID manifests still require you to materialize the corresponding
|
||
upstream benchmark payload into the documented `split_dir`. See
|
||
[`data/README.md`](../data/README.md) for the exact status of each benchmark.
|
||
When using `split_mode: ratio` instead, the loader is deterministic from
|
||
`split_seed` (default `42`) + `split_ratio` (default `2:1:7`), so a given
|
||
`data_path` + seed reproduces across machines.
|
||
|
||
## Why force-accept vs. gated slow-update matters
|
||
|
||
These `ckpt/` skills were produced with the gated slow-update semantics
|
||
described in paper Section 3.6:
|
||
|
||
```yaml
|
||
optimizer:
|
||
slow_update_gate_with_selection: true
|
||
```
|
||
|
||
Current `main` defaults to `false` (force-accept mode), a newer
|
||
post-submission behavior where the slow-update guidance is written into
|
||
`current_skill` and `best_skill` unconditionally at the epoch boundary. If
|
||
you re-train with the current default, you may produce a *different*
|
||
`best_skill.md` than the one checked in here. Both modes are supported; see
|
||
the [configuration reference](../docs/reference/config.md).
|