# Ablation Study Configuration Manifest This folder records the final, reproducible settings for the ablation runs used in `docs/ablation_paper_tables.md`. It is intentionally separate from the benchmark default configs. The benchmark configs under `configs//default.yaml` remain the source task configs; this folder records the exact matrix-level overrides, run roots, launch commands, and validation rules used for the paper ablations. ## Files - `matrix.yaml`: canonical ablation matrix, common overrides, benchmark splits, token/output caps, and invalid-run rules. - `launch_commands.sh`: exact launcher commands for the valid run roots. - `validation.md`: monitoring, result extraction, and invalidation checklist. ## Source Of Truth Use the matrix launcher: ```bash /home/azureuser/workspace-gzy/miniconda3/envs/skillopt/bin/python scripts/run_ablation_matrix.py ``` The launcher builds runs from the same defaults and values recorded in `matrix.yaml`. It skips completed runs by checking `summary.json` and skips active runs by checking `env.out_root` in active `scripts/train.py` processes. Do not manually rerun a completed run into the same `env.out_root`. If a run is invalid, archive or remove its output directory first, then let the launcher start it cleanly. ## Current Correct Run Roots - SearchQA / SpreadsheetBench original ablations: `outputs/ablation_20260502_040604_unique48` - SearchQA / SpreadsheetBench batch-size ablations: `outputs/ablation_batch_searchqa_spreadsheet_20260503_153902_run` - LiveMathBench / ALFWorld clean ablations: `outputs/ablation_livemath_alfworld_clean_20260503_155155_run` - DocVQA ablations: `outputs/ablation_docvqa_20260503_160225_run` Archived, superseded, misaligned, dry-run, or pre-fix directories must not be used for paper tables. ## End-To-End Runbook ### Environment Run from the repository root: ```bash cd /home/azureuser/workspace-gzy/SkillReflection ``` Always use: ```bash PY=/home/azureuser/workspace-gzy/miniconda3/envs/skillopt/bin/python export ALFWORLD_DATA=/home/azureuser/.cache/alfworld ``` Default model/auth settings are generated by `scripts/run_ablation_matrix.py`: ```text teacher=gpt-5.5 student=gpt-5.5 teacher_backend=openai_chat student_backend=openai_chat reasoning_effort=medium teacher/student endpoint=https://t2vgoaigpt4o3.openai.azure.com/ teacher/student api_version=2024-12-01-preview teacher/student auth_mode=azure_cli ``` Core training settings: ```text train.num_epochs=4 train.train_size=0 train.batch_size=40 train.accumulation=1 train.seed=42 gradient.minibatch_size=8 gradient.merge_batch_size=8 gradient.analyst_workers=16 gradient.use_deep_reflect=false optimizer.learning_rate=4 optimizer.min_learning_rate=2 optimizer.lr_scheduler=cosine optimizer.lr_control_mode=fixed optimizer.use_slow_update=true optimizer.slow_update_samples=20 optimizer.use_meta_skill=true optimizer.use_meta_reflect=false optimizer.longitudinal_pair_policy=mixed evaluation.use_gate=true evaluation.eval_test=true env.split_mode=split_dir ``` `train.train_size=0` is intentional. The dataloader derives the train size from the fixed split. Batch-size ablations rely on the default `ceil(train_size / batch_size)` behavior; the last batch can be smaller than `train.batch_size`. ### Fixed Splits Default split directories: ```text searchqa: data/ablation_splits/searchqa/2-1-7_seed42 spreadsheetbench: data/ablation_splits/spreadsheetbench/2-1-7_seed42 livemathematicianbench: data/ablation_splits/livemathematicianbench/2-1-7_seed42 alfworld: data/ablation_splits/alfworld/2-1-7_seed42 docvqa: /home/azureuser/zisu/SkillReflection/data/docvqa/splits ``` Default train/val/test sizes: | Benchmark | Train | Val | Test | | --- | ---: | ---: | ---: | | SearchQA | 400 | 200 | 1400 | | SpreadsheetBench | 80 | 40 | 280 | | LiveMathBench | 35 | 18 | 124 | | ALFWorld | 39 | 18 | 134 | | DocVQA | 1070 | 535 | 3744 | DocVQA images are not copied. The valid setup uses: ```text data/docvqa_images -> /home/azureuser/zisu/SkillReflection/data/docvqa_images ``` 2026-05-05 DocVQA data correction: all DocVQA final reruns should use the zisu 10% split above and a fresh output root such as `outputs/ablation_docvqa_zisu10pct_20260505_run`. The older local `data/ablation_splits/docvqa/2-1-7_seed42` contains the same 5349 questionId pool but a different train/val/test assignment, so its completed summaries are historical only. ### Matrix Groups Use these group names with `scripts/run_ablation_matrix.py`: ```text default split batch mbs lr sched slown mod smodel longpair lrctrl ``` `longpair` is the slow-update/meta-skill comparison-example ablation. It keeps all prompts and training settings unchanged and only overrides: ```text optimizer.longitudinal_pair_policy=changed optimizer.longitudinal_pair_policy=unchanged ``` The default paper setting remains `mixed`. `lrctrl` contains the two learning-rate-control baselines: ```text optimizer.lr_control_mode=autonomous optimizer.lr_control_mode=none + optimizer.skill_update_mode=full_rewrite_minibatch ``` The autonomous run logs the chosen integer per step in `lr_decision.json` and `lr_history.jsonl`. The full-rewrite run removes the LR/edit-selection concept: each minibatch analyst produces a complete skill candidate, and aggregate/merge produces the candidate skill directly. Batch-size values are: ```text 8 / 24 / 40 / 56 / full ``` `40` is the default point. `full` expands to the benchmark train size. ### Launch Commands Used In This Session The exact commands are recorded in `launch_commands.sh` and in `docs/ablation_plan.md`. The important current policy is: - SearchQA / SpreadsheetBench batch-only matrix can run at `--max-parallel 8`. - DocVQA matrix can run with its launcher at `--max-parallel 8`; later top-up used `--max-parallel 16` only because completed runs were skipped and active roots were checked. - LiveMathBench is safe as API-only benchmark after the token cap fix. - ALFWorld must not be mixed into a 24-way run on this shared machine. Use `--bench alfworld --max-parallel 1` only after memory is available. ### Token And Timeout Fixes LiveMathBench must use a large student completion cap: ```text max_completion_tokens=16384 timeout=300 ``` The old 768/512 cap produced many empty visible responses because hidden reasoning consumed the budget. ALFWorld must use: ```text max_completion_tokens=2048 empty response fallback -> look missing action fallback -> look ``` ### Invalid Runs Never fill paper tables from these archive directories: ```text outputs/ablation_livemath_alfworld_clean_20260503_155155_run/archive_livemath_token768_20260504_022258/ outputs/ablation_livemath_alfworld_clean_20260503_155155_run/archive_alfworld_token512_20260504_021417/ outputs/ablation_livemath_alfworld_clean_20260503_155155_run/archive_alfworld_empty_action_20260504_025311/ outputs/ablation_livemath_alfworld_clean_20260503_155155_run/archive_alfworld_prefallback_20260504_025402/ outputs/ablation_livemath_alfworld_clean_20260503_155155_run/archive_alfworld_oom_partial_20260504_050517/ outputs/ablation_livemath_alfworld_clean_20260503_155155_run/archive_alfworld_serial_lowmem_20260504_1300/ ``` ### Current Resource Notes ALFWorld model calls are API calls. The ablation branch now creates local ALFWorld/TextWorld environments through multiprocessing workers ported from `skillopt_final_zzw`, not through Ray actors. Old Ray-based archived runs are not valid for table fill. The observed historical failure mode in this session was system RAM pressure and Ray OOM prevention, not model GPU memory. GPU memory currently shown by `nvidia-smi` came from unrelated Ray Serve visual models under: ```text /home/azureuser/workspace-gzy/zyf/gca-skill ``` Those processes are `GroundingDINOModel` / `DA3Model`, not the SkillReflection ablation ALFWorld run. There are also unrelated ALFWorld jobs under: ```text /home/azureuser/zisu/skill_distill ``` Do not confuse those with this repository's ablation outputs. ### Monitoring Active run and duplicate output-root check: ```bash $PY - <<'PY' import subprocess, re, collections, time try: raw = subprocess.check_output(["pgrep", "-af", "scripts/train.py"], text=True) except subprocess.CalledProcessError: raw = "" roots = [] for line in raw.splitlines(): m = re.search(r"env\.out_root=([^\s]+)", line) if m: roots.append(m.group(1)) ctr = collections.Counter(roots) print("time", time.strftime("%F %T")) print("active_count", len(roots)) print("duplicates", [r.rsplit("/", 1)[-1] for r, c in ctr.items() if c > 1]) for root in sorted(roots): print(root.rsplit("/", 1)[-1]) PY ``` Error scan: ```bash rg -n "Traceback|ERROR|Error code|AuthenticationError|BadRequest|RateLimit|content_filter|Killed|OutOfMemory|CUDA out of memory|\\[FAIL\\]|LLM call failed" \ outputs/ablation_docvqa_20260503_160225_run/logs \ outputs/ablation_livemath_alfworld_clean_20260503_155155_run/logs \ outputs/ablation_batch_searchqa_spreadsheet_20260503_153902_run/logs \ -g '*.log' | tail -100 || true ``` Resource checks: ```bash free -h | sed -n '1,3p' df -h /tmp du -sh /tmp/ray 2>/dev/null || true nvidia-smi ``` ### Filling Tables Only use top-level `summary.json` from valid run roots. Fill `docs/ablation_paper_tables.md` from: ```text best_selection_hard baseline_test_hard test_hard test_delta_hard token_summary._total.total_tokens ```