Compare commits
87 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 4f582d4f6e | |||
| b3c7d72364 | |||
| 36284e1bb0 | |||
| fb1a76371d | |||
| 47063e1ceb | |||
| 181d71b737 | |||
| 41012e2d5e | |||
| dd8cd993b5 | |||
| 8ebede0efd | |||
| 266fca72ab | |||
| 9265545c45 | |||
| b4850ce418 | |||
| 643346c9f3 | |||
| 00602df9e9 | |||
| c6da31df44 | |||
| e4201074aa | |||
| 309ea64ff4 | |||
| d224d425f9 | |||
| 42e555d28e | |||
| 933c0a4ab5 | |||
| 1f75d022a5 | |||
| 4f3a9bc055 | |||
| d190bf37c1 | |||
| 02695bd813 | |||
| cf287cb608 | |||
| dbc90bd755 | |||
| 5d7875cb2e | |||
| 2532043d25 | |||
| 41be2f1803 | |||
| a62ec857f1 | |||
| afb552008b | |||
| 75b5c7f31c | |||
| 74ea3a1a8f | |||
| 657b987de6 | |||
| 2a40aa3c98 | |||
| 786d57b5cf | |||
| d5c5b61830 | |||
| 99212e3956 | |||
| fc54c44e93 | |||
| 48adf5a69f | |||
| b11e6dcfb9 | |||
| 4c1b74fce2 | |||
| db6443384a | |||
| 2c7d9074fb | |||
| c98bcdd5b3 | |||
| 0f6db9afc4 | |||
| 5a36ac35ae | |||
| 5f4b228543 | |||
| a9cad7a125 | |||
| 5e968115f5 | |||
| ded8c27c90 | |||
| f55a26414e | |||
| 2df2542aec | |||
| faa4ec6199 | |||
| cff7ff6846 | |||
| 7ae2d8766e | |||
| 338a88d31c | |||
| 4a1b984d87 | |||
| 6e165d5347 | |||
| dde7dc9dd8 | |||
| cd9a0a02b9 | |||
| 607bf74a1b | |||
| 9605217e75 | |||
| c42d541828 | |||
| 2e05edc399 | |||
| 6e7d5d0117 | |||
| 441ccb9bda | |||
| 88a99048a4 | |||
| bf2106808e | |||
| ba0fa8c14b | |||
| 9012a79827 | |||
| c64fbcd4f8 | |||
| 6e1027f01a | |||
| cd56a5fe7d | |||
| bbb250cc63 | |||
| 5c45add28b | |||
| e1896c691c | |||
| ec0841cccf | |||
| 4019f1cbe7 | |||
| cad3ab2d19 | |||
| 9a064f7c97 | |||
| 74cbe704fc | |||
| 5862bbdc97 | |||
| d4f9f4d5c5 | |||
| e27aac30ef | |||
| 76a58e6e7a | |||
| 244e346b83 |
@@ -0,0 +1,34 @@
|
||||
# SkillOpt Environment Variables
|
||||
# Copy this file to .env and fill in your values.
|
||||
# Usage: set -a; source .env; set +a
|
||||
|
||||
# ── Azure OpenAI (required for openai_chat backend) ──────────────────
|
||||
export AZURE_OPENAI_ENDPOINT=https://your-resource.openai.azure.com/
|
||||
export AZURE_OPENAI_API_VERSION=2024-12-01-preview
|
||||
# Authentication: choose one method
|
||||
# Option 1: API Key
|
||||
export AZURE_OPENAI_API_KEY=
|
||||
# Option 2: Azure CLI (no API key needed, recommended on Azure VMs)
|
||||
# export AZURE_OPENAI_AUTH_MODE=azure_cli
|
||||
# Option 3: Managed Identity
|
||||
# export AZURE_OPENAI_AUTH_MODE=managed_identity
|
||||
# export AZURE_OPENAI_MANAGED_IDENTITY_CLIENT_ID=your-client-id
|
||||
|
||||
# ── OpenAI-compatible endpoints ──────────────────────────────────────
|
||||
# Set AUTH_MODE to openai_compatible and reuse AZURE_OPENAI_ENDPOINT / _API_KEY.
|
||||
# The plain OpenAI client is used; no Azure auth, no api-version header.
|
||||
# export AZURE_OPENAI_ENDPOINT=https://api.openai.com/v1
|
||||
# export AZURE_OPENAI_API_KEY=sk-...
|
||||
# export AZURE_OPENAI_AUTH_MODE=openai_compatible
|
||||
|
||||
# ── Anthropic / Claude (for claude_chat backend) ─────────────────────
|
||||
# export ANTHROPIC_API_KEY=sk-ant-...
|
||||
|
||||
# ── Qwen Local Model (for qwen_chat backend) ────────────────────────
|
||||
# export QWEN_CHAT_BASE_URL=http://localhost:8000/v1
|
||||
# export QWEN_CHAT_MODEL=Qwen/Qwen3.5-4B
|
||||
|
||||
# ── MiniMax (for minimax_chat backend) ──────────────────────────────
|
||||
# export MINIMAX_BASE_URL=https://api.minimax.io/v1
|
||||
# export MINIMAX_API_KEY=...
|
||||
# export MINIMAX_MODEL=MiniMax-M2.7
|
||||
+50
-423
@@ -1,429 +1,56 @@
|
||||
## Ignore Visual Studio temporary files, build results, and
|
||||
## files generated by popular Visual Studio add-ons.
|
||||
##
|
||||
## Get latest from https://github.com/github/gitignore/blob/main/VisualStudio.gitignore
|
||||
|
||||
# User-specific files
|
||||
*.rsuser
|
||||
*.suo
|
||||
*.user
|
||||
*.userosscache
|
||||
*.sln.docstates
|
||||
*.env
|
||||
|
||||
# User-specific files (MonoDevelop/Xamarin Studio)
|
||||
*.userprefs
|
||||
|
||||
# Mono auto generated files
|
||||
mono_crash.*
|
||||
|
||||
# Build results
|
||||
[Dd]ebug/
|
||||
[Dd]ebugPublic/
|
||||
[Rr]elease/
|
||||
[Rr]eleases/
|
||||
|
||||
[Dd]ebug/x64/
|
||||
[Dd]ebugPublic/x64/
|
||||
[Rr]elease/x64/
|
||||
[Rr]eleases/x64/
|
||||
bin/x64/
|
||||
obj/x64/
|
||||
|
||||
[Dd]ebug/x86/
|
||||
[Dd]ebugPublic/x86/
|
||||
[Rr]elease/x86/
|
||||
[Rr]eleases/x86/
|
||||
bin/x86/
|
||||
obj/x86/
|
||||
|
||||
[Ww][Ii][Nn]32/
|
||||
[Aa][Rr][Mm]/
|
||||
[Aa][Rr][Mm]64/
|
||||
[Aa][Rr][Mm]64[Ee][Cc]/
|
||||
bld/
|
||||
[Oo]bj/
|
||||
[Oo]ut/
|
||||
[Ll]og/
|
||||
[Ll]ogs/
|
||||
|
||||
# Build results on 'Bin' directories
|
||||
**/[Bb]in/*
|
||||
# Uncomment if you have tasks that rely on *.refresh files to move binaries
|
||||
# (https://github.com/github/gitignore/pull/3736)
|
||||
#!**/[Bb]in/*.refresh
|
||||
|
||||
# Visual Studio 2015/2017 cache/options directory
|
||||
.vs/
|
||||
# Uncomment if you have tasks that create the project's static files in wwwroot
|
||||
#wwwroot/
|
||||
|
||||
# Visual Studio 2017 auto generated files
|
||||
Generated\ Files/
|
||||
|
||||
# MSTest test Results
|
||||
[Tt]est[Rr]esult*/
|
||||
[Bb]uild[Ll]og.*
|
||||
*.trx
|
||||
|
||||
# NUnit
|
||||
*.VisualState.xml
|
||||
TestResult.xml
|
||||
nunit-*.xml
|
||||
|
||||
# Approval Tests result files
|
||||
*.received.*
|
||||
|
||||
# Build Results of an ATL Project
|
||||
[Dd]ebugPS/
|
||||
[Rr]eleasePS/
|
||||
dlldata.c
|
||||
|
||||
# Benchmark Results
|
||||
BenchmarkDotNet.Artifacts/
|
||||
|
||||
# .NET Core
|
||||
project.lock.json
|
||||
project.fragment.lock.json
|
||||
artifacts/
|
||||
.artifacts/
|
||||
|
||||
# ASP.NET Scaffolding
|
||||
ScaffoldingReadMe.txt
|
||||
|
||||
# StyleCop
|
||||
StyleCopReport.xml
|
||||
|
||||
# Files built by Visual Studio
|
||||
*_i.c
|
||||
*_p.c
|
||||
*_h.h
|
||||
*.ilk
|
||||
*.meta
|
||||
*.obj
|
||||
*.idb
|
||||
*.iobj
|
||||
*.pch
|
||||
*.pdb
|
||||
*.ipdb
|
||||
*.pgc
|
||||
*.pgd
|
||||
*.rsp
|
||||
# but not Directory.Build.rsp, as it configures directory-level build defaults
|
||||
!Directory.Build.rsp
|
||||
*.sbr
|
||||
*.tlb
|
||||
*.tli
|
||||
*.tlh
|
||||
*.tmp
|
||||
*.tmp_proj
|
||||
*_wpftmp.csproj
|
||||
*.log
|
||||
*.tlog
|
||||
*.vspscc
|
||||
*.vssscc
|
||||
.builds
|
||||
*.pidb
|
||||
*.svclog
|
||||
*.scc
|
||||
|
||||
# Chutzpah Test files
|
||||
_Chutzpah*
|
||||
|
||||
# Visual C++ cache files
|
||||
ipch/
|
||||
*.aps
|
||||
*.ncb
|
||||
*.opendb
|
||||
*.opensdf
|
||||
*.sdf
|
||||
*.cachefile
|
||||
*.VC.db
|
||||
*.VC.VC.opendb
|
||||
|
||||
# Visual Studio profiler
|
||||
*.psess
|
||||
*.vsp
|
||||
*.vspx
|
||||
*.sap
|
||||
|
||||
# Visual Studio Trace Files
|
||||
*.e2e
|
||||
|
||||
# TFS 2012 Local Workspace
|
||||
$tf/
|
||||
|
||||
# Guidance Automation Toolkit
|
||||
*.gpState
|
||||
|
||||
# ReSharper is a .NET coding add-in
|
||||
_ReSharper*/
|
||||
*.[Rr]e[Ss]harper
|
||||
*.DotSettings.user
|
||||
|
||||
# TeamCity is a build add-in
|
||||
_TeamCity*
|
||||
|
||||
# DotCover is a Code Coverage Tool
|
||||
*.dotCover
|
||||
|
||||
# AxoCover is a Code Coverage Tool
|
||||
.axoCover/*
|
||||
!.axoCover/settings.json
|
||||
|
||||
# Coverlet is a free, cross platform Code Coverage Tool
|
||||
coverage*.json
|
||||
coverage*.xml
|
||||
coverage*.info
|
||||
|
||||
# Visual Studio code coverage results
|
||||
*.coverage
|
||||
*.coveragexml
|
||||
|
||||
# NCrunch
|
||||
_NCrunch_*
|
||||
.NCrunch_*
|
||||
.*crunch*.local.xml
|
||||
nCrunchTemp_*
|
||||
|
||||
# MightyMoose
|
||||
*.mm.*
|
||||
AutoTest.Net/
|
||||
|
||||
# Web workbench (sass)
|
||||
.sass-cache/
|
||||
|
||||
# Installshield output folder
|
||||
[Ee]xpress/
|
||||
|
||||
# DocProject is a documentation generator add-in
|
||||
DocProject/buildhelp/
|
||||
DocProject/Help/*.HxT
|
||||
DocProject/Help/*.HxC
|
||||
DocProject/Help/*.hhc
|
||||
DocProject/Help/*.hhk
|
||||
DocProject/Help/*.hhp
|
||||
DocProject/Help/Html2
|
||||
DocProject/Help/html
|
||||
|
||||
# Click-Once directory
|
||||
publish/
|
||||
|
||||
# Publish Web Output
|
||||
*.[Pp]ublish.xml
|
||||
*.azurePubxml
|
||||
# Note: Comment the next line if you want to checkin your web deploy settings,
|
||||
# but database connection strings (with potential passwords) will be unencrypted
|
||||
*.pubxml
|
||||
*.publishproj
|
||||
|
||||
# Microsoft Azure Web App publish settings. Comment the next line if you want to
|
||||
# checkin your Azure Web App publish settings, but sensitive information contained
|
||||
# in these scripts will be unencrypted
|
||||
PublishScripts/
|
||||
|
||||
# NuGet Packages
|
||||
*.nupkg
|
||||
# NuGet Symbol Packages
|
||||
*.snupkg
|
||||
# The packages folder can be ignored because of Package Restore
|
||||
**/[Pp]ackages/*
|
||||
# except build/, which is used as an MSBuild target.
|
||||
!**/[Pp]ackages/build/
|
||||
# Uncomment if necessary however generally it will be regenerated when needed
|
||||
#!**/[Pp]ackages/repositories.config
|
||||
# NuGet v3's project.json files produces more ignorable files
|
||||
*.nuget.props
|
||||
*.nuget.targets
|
||||
|
||||
# Microsoft Azure Build Output
|
||||
csx/
|
||||
*.build.csdef
|
||||
|
||||
# Microsoft Azure Emulator
|
||||
ecf/
|
||||
rcf/
|
||||
|
||||
# Windows Store app package directories and files
|
||||
AppPackages/
|
||||
BundleArtifacts/
|
||||
Package.StoreAssociation.xml
|
||||
_pkginfo.txt
|
||||
*.appx
|
||||
*.appxbundle
|
||||
*.appxupload
|
||||
|
||||
# Visual Studio cache files
|
||||
# files ending in .cache can be ignored
|
||||
*.[Cc]ache
|
||||
# but keep track of directories ending in .cache
|
||||
!?*.[Cc]ache/
|
||||
|
||||
# Others
|
||||
ClientBin/
|
||||
~$*
|
||||
*~
|
||||
*.dbmdl
|
||||
*.dbproj.schemaview
|
||||
*.jfm
|
||||
*.pfx
|
||||
*.publishsettings
|
||||
orleans.codegen.cs
|
||||
|
||||
# Including strong name files can present a security risk
|
||||
# (https://github.com/github/gitignore/pull/2483#issue-259490424)
|
||||
#*.snk
|
||||
|
||||
# Since there are multiple workflows, uncomment next line to ignore bower_components
|
||||
# (https://github.com/github/gitignore/pull/1529#issuecomment-104372622)
|
||||
#bower_components/
|
||||
|
||||
# RIA/Silverlight projects
|
||||
Generated_Code/
|
||||
|
||||
# Backup & report files from converting an old project file
|
||||
# to a newer Visual Studio version. Backup files are not needed,
|
||||
# because we have git ;-)
|
||||
_UpgradeReport_Files/
|
||||
Backup*/
|
||||
UpgradeLog*.XML
|
||||
UpgradeLog*.htm
|
||||
ServiceFabricBackup/
|
||||
*.rptproj.bak
|
||||
|
||||
# SQL Server files
|
||||
*.mdf
|
||||
*.ldf
|
||||
*.ndf
|
||||
|
||||
# Business Intelligence projects
|
||||
*.rdl.data
|
||||
*.bim.layout
|
||||
*.bim_*.settings
|
||||
*.rptproj.rsuser
|
||||
*- [Bb]ackup.rdl
|
||||
*- [Bb]ackup ([0-9]).rdl
|
||||
*- [Bb]ackup ([0-9][0-9]).rdl
|
||||
|
||||
# Microsoft Fakes
|
||||
FakesAssemblies/
|
||||
|
||||
# GhostDoc plugin setting file
|
||||
*.GhostDoc.xml
|
||||
|
||||
# Node.js Tools for Visual Studio
|
||||
.ntvs_analysis.dat
|
||||
node_modules/
|
||||
|
||||
# Visual Studio 6 build log
|
||||
*.plg
|
||||
|
||||
# Visual Studio 6 workspace options file
|
||||
*.opt
|
||||
|
||||
# Visual Studio 6 auto-generated workspace file (contains which files were open etc.)
|
||||
*.vbw
|
||||
|
||||
# Visual Studio 6 workspace and project file (working project files containing files to include in project)
|
||||
*.dsw
|
||||
*.dsp
|
||||
|
||||
# Visual Studio 6 technical files
|
||||
*.ncb
|
||||
*.aps
|
||||
|
||||
# Visual Studio LightSwitch build output
|
||||
**/*.HTMLClient/GeneratedArtifacts
|
||||
**/*.DesktopClient/GeneratedArtifacts
|
||||
**/*.DesktopClient/ModelManifest.xml
|
||||
**/*.Server/GeneratedArtifacts
|
||||
**/*.Server/ModelManifest.xml
|
||||
_Pvt_Extensions
|
||||
|
||||
# Paket dependency manager
|
||||
**/.paket/paket.exe
|
||||
paket-files/
|
||||
|
||||
# FAKE - F# Make
|
||||
**/.fake/
|
||||
|
||||
# CodeRush personal settings
|
||||
**/.cr/personal
|
||||
|
||||
# Python Tools for Visual Studio (PTVS)
|
||||
**/__pycache__/
|
||||
__pycache__/
|
||||
*.pyc
|
||||
*.egg-info/
|
||||
build/
|
||||
dist/
|
||||
site/
|
||||
|
||||
# Cake - Uncomment if you are using it
|
||||
#tools/**
|
||||
#!tools/packages.config
|
||||
data/*
|
||||
!data/README.md
|
||||
!data/searchqa_id_split/
|
||||
!data/searchqa_id_split/**
|
||||
!data/livemathematicianbench_id_split/
|
||||
!data/livemathematicianbench_id_split/**
|
||||
!data/docvqa_id_split/
|
||||
!data/docvqa_id_split/**
|
||||
!data/officeqa_id_split/
|
||||
!data/officeqa_id_split/**
|
||||
!data/spreadsheetbench_id_split/
|
||||
!data/spreadsheetbench_id_split/**
|
||||
!data/alfworld_path_split/
|
||||
!data/alfworld_path_split/**
|
||||
outputs/
|
||||
logs/
|
||||
external/
|
||||
|
||||
# Tabs Studio
|
||||
*.tss
|
||||
/BabyVision/
|
||||
/MMRB/
|
||||
/SpreadsheetBench/
|
||||
/dl4ir-searchQA/
|
||||
|
||||
# Telerik's JustMock configuration file
|
||||
*.jmconfig
|
||||
configs/local/
|
||||
configs/**/*.local.yaml
|
||||
*.local.md
|
||||
*.secret.md
|
||||
*.bak
|
||||
|
||||
# BizTalk build output
|
||||
*.btp.cs
|
||||
*.btm.cs
|
||||
*.odx.cs
|
||||
*.xsd.cs
|
||||
.env
|
||||
.secrets/
|
||||
.codex_azure*/
|
||||
|
||||
# OpenCover UI analysis results
|
||||
OpenCover/
|
||||
|
||||
# Azure Stream Analytics local run output
|
||||
ASALocalRun/
|
||||
|
||||
# MSBuild Binary and Structured Log
|
||||
*.binlog
|
||||
MSBuild_Logs/
|
||||
|
||||
# AWS SAM Build and Temporary Artifacts folder
|
||||
.aws-sam
|
||||
|
||||
# NVidia Nsight GPU debugger configuration file
|
||||
*.nvuser
|
||||
|
||||
# MFractors (Xamarin productivity tool) working folder
|
||||
**/.mfractor/
|
||||
|
||||
# Local History for Visual Studio
|
||||
**/.localhistory/
|
||||
|
||||
# Visual Studio History (VSHistory) files
|
||||
.vshistory/
|
||||
|
||||
# BeatPulse healthcheck temp database
|
||||
healthchecksdb
|
||||
|
||||
# Backup folder for Package Reference Convert tool in Visual Studio 2017
|
||||
MigrationBackup/
|
||||
|
||||
# Ionide (cross platform F# VS Code tools) working folder
|
||||
**/.ionide/
|
||||
|
||||
# Fody - auto-generated XML schema
|
||||
FodyWeavers.xsd
|
||||
|
||||
# VS Code files for those working on multiple tools
|
||||
.vscode/*
|
||||
!.vscode/settings.json
|
||||
!.vscode/tasks.json
|
||||
!.vscode/launch.json
|
||||
!.vscode/extensions.json
|
||||
!.vscode/*.code-snippets
|
||||
|
||||
# Local History for Visual Studio Code
|
||||
.history/
|
||||
|
||||
# Built Visual Studio Code Extensions
|
||||
*.vsix
|
||||
|
||||
# Windows Installer files from build outputs
|
||||
*.cab
|
||||
*.msi
|
||||
*.msix
|
||||
*.msm
|
||||
*.msp
|
||||
# Internal docs (not for open-source release)
|
||||
docs/ablation_plan.md
|
||||
docs/ablation_paper_tables.md
|
||||
docs/ablation_paper_tables.html
|
||||
docs/experiment_commands.md
|
||||
docs/slow_update_flowchart.md
|
||||
docs/session_memory.md
|
||||
docs/harness_fresh_machine_handoff.md
|
||||
docs/harness_monitoring_memory.md
|
||||
docs/harness_reproduction_secrets.secret.md
|
||||
docs/reflact_conda_env_export.yml
|
||||
docs/reflact_overview.html
|
||||
docs/render_ablation_paper_tables.py
|
||||
docs/让*
|
||||
.gradio/
|
||||
.venv
|
||||
|
||||
@@ -0,0 +1,43 @@
|
||||
# Contributing to SkillOpt
|
||||
|
||||
Thank you for your interest in contributing! SkillOpt welcomes contributions of all kinds.
|
||||
|
||||
## Getting Started
|
||||
|
||||
```bash
|
||||
git clone https://github.com/microsoft/SkillOpt.git
|
||||
cd SkillOpt
|
||||
pip install -e ".[dev]"
|
||||
```
|
||||
|
||||
## How to Contribute
|
||||
|
||||
### 🐛 Bug Reports
|
||||
Open a GitHub issue with reproduction steps, expected/actual behavior, and your config file (remove API keys).
|
||||
|
||||
### 🔧 Add a Benchmark
|
||||
See the [guide](docs/guide/new-benchmark.md) and use the scaffold at `skillopt/envs/_template/`.
|
||||
|
||||
### 🤖 Add a Model Backend
|
||||
See the [guide](docs/guide/new-backend.md).
|
||||
|
||||
### 📝 Improve Documentation
|
||||
```bash
|
||||
pip install -e ".[docs]"
|
||||
mkdocs serve # Preview at http://localhost:8000
|
||||
```
|
||||
|
||||
## Pull Request Process
|
||||
|
||||
1. Fork the repo and create a feature branch
|
||||
2. Make changes and test with an existing benchmark
|
||||
3. Submit a PR with a clear description
|
||||
4. Ensure CI passes
|
||||
|
||||
## Code Style
|
||||
- Follow existing patterns in the codebase
|
||||
- Use type hints for function signatures
|
||||
- Keep docstrings concise
|
||||
|
||||
## License
|
||||
By contributing, you agree your contributions are licensed under the [MIT License](LICENSE).
|
||||
@@ -0,0 +1,21 @@
|
||||
MIT License
|
||||
|
||||
Copyright (c) 2026 Microsoft Corporation
|
||||
|
||||
Permission is hereby granted, free of charge, to any person obtaining a copy
|
||||
of this software and associated documentation files (the "Software"), to deal
|
||||
in the Software without restriction, including without limitation the rights
|
||||
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
||||
copies of the Software, and to permit persons to whom the Software is
|
||||
furnished to do so, subject to the following conditions:
|
||||
|
||||
The above copyright notice and this permission notice shall be included in all
|
||||
copies or substantial portions of the Software.
|
||||
|
||||
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
||||
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
||||
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
||||
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
||||
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
||||
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
||||
SOFTWARE.
|
||||
@@ -0,0 +1,395 @@
|
||||
# SkillOpt: Executive Strategy for Self-Evolving Agent Skills
|
||||
|
||||
*Train agent skills like you train neural networks — with epochs, (mini-)batchsize, learning rates, and validation gates — but without touching model weights.*
|
||||
|
||||
[](https://microsoft.github.io/SkillOpt/) [](https://arxiv.org/abs/2605.23904) [](https://youtu.be/JUBMDTCiM0M) [](https://www.python.org/) [](LICENSE)
|
||||
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
Modern agent skills are usually hand-crafted, generated one-shot by a strong
|
||||
LLM, or evolved through loosely controlled self-revision — none of which
|
||||
behaves like a deep-learning optimizer for the skill itself, and none of
|
||||
which reliably improves over its starting point under feedback.
|
||||
|
||||
**SkillOpt treats the skill document as the trainable state of a frozen
|
||||
agent**, and trains it with the discipline that makes weight-space
|
||||
optimization reproducible. A separate optimizer model turns scored rollouts
|
||||
into bounded add / delete / replace edits on a single skill document; a
|
||||
candidate edit is accepted only when it strictly improves a held-out
|
||||
validation score. A textual learning-rate budget, a rejected-edit buffer,
|
||||
and an epoch-wise slow / meta update make skill training stable while
|
||||
adding **zero inference-time model calls** at deployment.
|
||||
|
||||
The deployed artifact is a compact `best_skill.md` (typically 300–2,000
|
||||
tokens) that runs against the unchanged target model. Across **six
|
||||
benchmarks, seven target models, and three execution harnesses** (direct
|
||||
chat, Codex CLI, Claude Code CLI), SkillOpt is best or tied-best on **all
|
||||
52 evaluated (model, benchmark, harness) cells** and on GPT-5.5 lifts the
|
||||
average no-skill accuracy by **+23.5 points in direct chat, +24.8 inside
|
||||
the Codex agentic loop, and +19.1 inside Claude Code**. Optimized skill
|
||||
artifacts transfer across model scales, between Codex and Claude Code
|
||||
harnesses, and to nearby benchmarks without further optimization.
|
||||
|
||||
For the full method, ablations, and per-cell results see the [paper](https://arxiv.org/abs/2605.23904); for a visual walkthrough of the loop see the [project page](https://microsoft.github.io/SkillOpt/); for deeper API / backend / benchmark docs see [`docs/`](docs/).
|
||||
|
||||
## 🎬 Demo Video
|
||||
|
||||
https://github.com/user-attachments/assets/eb12d3bc-371c-467f-904d-91b61f339ed7
|
||||
|
||||
<p align="center">
|
||||
<a href="https://youtu.be/JUBMDTCiM0M"><b>▶ Watch the full demo on YouTube</b></a>
|
||||
</p>
|
||||
|
||||
---
|
||||
|
||||
## Install
|
||||
|
||||
### Requirements
|
||||
|
||||
- Python 3.10+
|
||||
|
||||
```bash
|
||||
git clone https://github.com/microsoft/SkillOpt.git
|
||||
cd SkillOpt
|
||||
pip install -e .
|
||||
|
||||
# For the ALFWorld benchmark (optional):
|
||||
pip install -e ".[alfworld]"
|
||||
alfworld-download
|
||||
```
|
||||
|
||||
### Configure API Credentials
|
||||
|
||||
```bash
|
||||
cp .env.example .env
|
||||
# Edit .env with your API credentials, then:
|
||||
source .env
|
||||
```
|
||||
|
||||
#### Azure OpenAI *(recommended)*
|
||||
|
||||
```bash
|
||||
export AZURE_OPENAI_ENDPOINT="https://your-resource.openai.azure.com/"
|
||||
# Option 1: API key auth
|
||||
export AZURE_OPENAI_API_KEY="your-key"
|
||||
# Option 2: Azure CLI auth (no API key needed)
|
||||
export AZURE_OPENAI_AUTH_MODE="azure_cli"
|
||||
```
|
||||
|
||||
> **Note:** `AZURE_OPENAI_ENDPOINT` is required for all three modes (`api_key`, `azure_cli`, `openai_compatible`). Without it, all LLM calls will fail.
|
||||
|
||||
#### OpenAI-compatible endpoints
|
||||
|
||||
```bash
|
||||
export AZURE_OPENAI_ENDPOINT="https://api.openai.com/v1"
|
||||
export AZURE_OPENAI_API_KEY="sk-..."
|
||||
export AZURE_OPENAI_AUTH_MODE="openai_compatible"
|
||||
```
|
||||
|
||||
This routes all calls through the plain OpenAI Python client (no Azure auth, no `api-version` header).
|
||||
|
||||
> **Note:** SkillOpt reuses the `AZURE_OPENAI_*` env var names even in this mode — there is no separate `OPENAI_API_KEY` knob.
|
||||
|
||||
#### Anthropic Claude
|
||||
|
||||
```bash
|
||||
export ANTHROPIC_API_KEY="sk-ant-..."
|
||||
```
|
||||
|
||||
#### Qwen *(local vLLM)*
|
||||
|
||||
```bash
|
||||
export QWEN_CHAT_BASE_URL="http://localhost:8000/v1"
|
||||
export QWEN_CHAT_MODEL="Qwen/Qwen3.5-4B"
|
||||
```
|
||||
|
||||
`qwen_chat` can also be used as the optimizer backend. When optimizer and
|
||||
target should point to different local vLLM services, use the role-specific
|
||||
settings:
|
||||
|
||||
```bash
|
||||
python scripts/train.py \
|
||||
--config configs/searchqa/default.yaml \
|
||||
--optimizer_backend qwen_chat \
|
||||
--target_backend qwen_chat \
|
||||
--optimizer_model Qwen/Qwen3.5-4B \
|
||||
--target_model Qwen/Qwen3.5-4B \
|
||||
--optimizer_qwen_chat_base_url http://localhost:8001/v1 \
|
||||
--target_qwen_chat_base_url http://localhost:8000/v1
|
||||
```
|
||||
|
||||
#### MiniMax
|
||||
|
||||
```bash
|
||||
export MINIMAX_BASE_URL="https://api.minimax.io/v1"
|
||||
export MINIMAX_API_KEY="..."
|
||||
export MINIMAX_MODEL="MiniMax-M2.7"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Quick Start
|
||||
|
||||
### Training
|
||||
|
||||
```bash
|
||||
# Minimal example — train on SearchQA:
|
||||
python scripts/train.py \
|
||||
--config configs/searchqa/default.yaml \
|
||||
--split_dir /path/to/your/searchqa_split \
|
||||
--azure_openai_endpoint https://your-resource.openai.azure.com/ \
|
||||
--optimizer_model gpt-5.5 \
|
||||
--target_model gpt-5.5
|
||||
|
||||
# Train on LiveMathematicianBench:
|
||||
python scripts/train.py \
|
||||
--config configs/livemathematicianbench/default.yaml \
|
||||
--split_dir /path/to/your/livemath_split \
|
||||
--azure_openai_endpoint https://your-resource.openai.azure.com/ \
|
||||
--optimizer_model gpt-5.5 \
|
||||
--target_model gpt-5.5
|
||||
|
||||
# Train on ALFWorld:
|
||||
python scripts/train.py \
|
||||
--config configs/alfworld/default.yaml \
|
||||
--split_dir data/alfworld_path_split \
|
||||
--azure_openai_endpoint https://your-resource.openai.azure.com/ \
|
||||
--optimizer_model gpt-5.5 \
|
||||
--target_model gpt-5.5
|
||||
```
|
||||
|
||||
Key CLI arguments:
|
||||
|
||||
| Argument | Description | Example |
|
||||
|---|---|---|
|
||||
| `--config` | Benchmark config YAML | `configs/searchqa/default.yaml` |
|
||||
| `--split_dir` | Path to data split directory | `/path/to/split` |
|
||||
| `--azure_openai_endpoint` | Azure OpenAI endpoint URL | `https://your-resource.openai.azure.com/` |
|
||||
| `--optimizer_model` | Optimizer model deployment name | `gpt-5.5` |
|
||||
| `--target_model` | Target model deployment name | `gpt-5.5` |
|
||||
| `--num_epochs` | Number of training epochs | `4` |
|
||||
| `--batch_size` | Batch size per step | `40` |
|
||||
| `--workers` | Parallel rollout workers | `8` |
|
||||
| `--out_root` | Output directory | `outputs/my_run` |
|
||||
|
||||
### Eval Only
|
||||
|
||||
Evaluate a trained skill on specific data splits without training:
|
||||
|
||||
```bash
|
||||
# Evaluate the packaged GPT-5.5 SearchQA skill on the test split:
|
||||
python scripts/eval_only.py \
|
||||
--config configs/searchqa/default.yaml \
|
||||
--skill ckpt/searchqa/gpt5.5_skill.md \
|
||||
--split valid_unseen \
|
||||
--split_dir /path/to/searchqa_split \
|
||||
--azure_openai_endpoint https://your-resource.openai.azure.com/
|
||||
|
||||
# Evaluate on all splits (train + val + test):
|
||||
python scripts/eval_only.py \
|
||||
--config configs/searchqa/default.yaml \
|
||||
--skill ckpt/searchqa/gpt5.5_skill.md \
|
||||
--split all \
|
||||
--split_dir /path/to/searchqa_split \
|
||||
--azure_openai_endpoint https://your-resource.openai.azure.com/
|
||||
```
|
||||
|
||||
To evaluate a skill produced by your own training run, replace `--skill` with that run's best-skill path, for example `outputs/my_run/best_skill.md`.
|
||||
|
||||
| Split | Description |
|
||||
|---|---|
|
||||
| `valid_unseen` | Test set |
|
||||
| `valid_seen` | Validation set |
|
||||
| `train` | Training set |
|
||||
| `all` | All splits combined (default) |
|
||||
|
||||
### Output Structure
|
||||
|
||||
Each training run writes to a structured output directory:
|
||||
|
||||
```
|
||||
outputs/<run_name>/
|
||||
├── config.json # Flattened runtime config
|
||||
├── history.json # Per-step training history
|
||||
├── runtime_state.json # Resume checkpoint
|
||||
├── best_skill.md # Best validated skill document
|
||||
├── skills/skill_vXXXX.md # Skill snapshot per step
|
||||
├── steps/step_XXXX/ # Per-step artifacts (patches, evals)
|
||||
├── slow_update/epoch_XX/ # Slow update logs
|
||||
└── meta_skill/epoch_XX/ # Meta skill logs
|
||||
```
|
||||
|
||||
Re-running the same command auto-resumes from the last completed step.
|
||||
|
||||
### Pretrained Skill Artifacts
|
||||
|
||||
We provide a subset of the paper's main Table 1 GPT-5.5 optimized skills in
|
||||
[`ckpt/`](ckpt/) as reference artifacts. Use them with `scripts/eval_only.py`
|
||||
to evaluate the provided skills on a matching data split without re-running
|
||||
training. See [`ckpt/README.md`](ckpt/README.md) for the full per-benchmark
|
||||
command. This is the first artifact batch; we plan to continue uploading
|
||||
the remaining optimized skills and benchmark split manifests as they are
|
||||
cleaned and verified.
|
||||
|
||||
---
|
||||
|
||||
## Data Preparation
|
||||
|
||||
### Directory layout
|
||||
|
||||
SkillOpt expects data in a **split directory** with `train/`, `val/`, `test/` subdirectories, each containing a JSON file (e.g., `items.json`):
|
||||
|
||||
```
|
||||
data/my_split/
|
||||
├── train/items.json
|
||||
├── val/items.json
|
||||
└── test/items.json
|
||||
```
|
||||
|
||||
Each JSON file is an array of task items. The required fields depend on the benchmark. For example, SearchQA items look like:
|
||||
|
||||
```json
|
||||
[
|
||||
{
|
||||
"id": "unique_item_id",
|
||||
"question": "Who wrote the novel ...",
|
||||
"context": "[DOC] relevant passage text ...",
|
||||
"answers": ["expected answer"]
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
See `skillopt/envs/<benchmark>/dataloader.py` for the exact format each benchmark expects.
|
||||
|
||||
> **Note:** Most benchmark datasets are not included in this repository. Prepare your own data following the format above. The exact SearchQA split used in the paper is provided at [`data/searchqa_id_split/`](data/searchqa_id_split) (400 train / 200 val / 1400 test). We are preparing the remaining benchmark split manifests for upload.
|
||||
|
||||
### Supported Benchmarks
|
||||
|
||||
| Benchmark | Type | Config |
|
||||
|---|---|---|
|
||||
| SearchQA | QA | `configs/searchqa/default.yaml` |
|
||||
| ALFWorld | Embodied agent | `configs/alfworld/default.yaml` |
|
||||
| DocVQA | Document QA | `configs/docvqa/default.yaml` |
|
||||
| LiveMathematicianBench | Math | `configs/livemathematicianbench/default.yaml` |
|
||||
| SpreadsheetBench | Code generation | `configs/spreadsheetbench/default.yaml` |
|
||||
| OfficeQA | Tool-augmented QA | `configs/officeqa/default.yaml` |
|
||||
|
||||
---
|
||||
|
||||
## Configuration
|
||||
|
||||
### Default settings and paper-reproduction knobs
|
||||
|
||||
`configs/_base_/default.yaml` is the single source of truth for SkillOpt's
|
||||
runtime knobs. Out of the box, every included benchmark config inherits
|
||||
from it and keeps the paper protocol visible: 4 epochs, rollout batch 40,
|
||||
reflection minibatch 8, textual learning rate 4 with cosine decay, strict
|
||||
hard validation gating, and slow-update + meta-skill enabled. One detail to
|
||||
watch is slow-update acceptance: the current `main` default is the newer
|
||||
post-submission force-accept mode, while the paper protocol and the
|
||||
paper-aligned skills under `ckpt/` use the gated semantics described in
|
||||
paper Section 3.6.
|
||||
|
||||
### Slow-update acceptance mode
|
||||
|
||||
The epoch-boundary slow / meta update can be applied two ways, controlled
|
||||
by `optimizer.slow_update_gate_with_selection`:
|
||||
|
||||
```yaml
|
||||
optimizer:
|
||||
slow_update_gate_with_selection: false # current main default
|
||||
```
|
||||
|
||||
- **`false`** *(current `main` default)*: force-accept. The
|
||||
slow-update guidance is injected into both `current_skill` and
|
||||
`best_skill` unconditionally at the epoch boundary. This is the newer
|
||||
post-submission behavior on `main`.
|
||||
- **`true`** *(paper / ckpt-skill reproduction)*: gated, matching paper
|
||||
Section 3.6 verbatim. The slow-update candidate is evaluated on the
|
||||
selection split and accepted only if it passes the same validation gate
|
||||
as a step-level edit. Use this setting when re-running optimization to
|
||||
match the paper protocol and the provenance of the provided `ckpt/` skills.
|
||||
|
||||
The trainer prints which mode is active at startup
|
||||
(`[slow update] acceptance=...`). See issue #22 for the discussion that
|
||||
led to the flag.
|
||||
|
||||
### Gate metric (`hard` / `soft` / `mixed`)
|
||||
|
||||
The validation gate compares candidate vs. current skills on the selection
|
||||
split using `gate_metric`:
|
||||
|
||||
- **`hard`** *(default, paper)*: exact-match accuracy, strictly greater
|
||||
than the current score is required.
|
||||
- **`soft`**: per-item soft / partial-credit score. Useful when the
|
||||
selection split is small (e.g. ≤10 items) and the reward is continuous,
|
||||
where the discrete hard gate often rejects every candidate.
|
||||
- **`mixed`**: weighted average, `(1 - w) * hard + w * soft`, with `w`
|
||||
set by `gate_mixed_weight` (default `0.5`).
|
||||
|
||||
Default is `hard`. Use the optional feature config below to switch.
|
||||
|
||||
### Optional feature configs
|
||||
|
||||
These are **not** default SkillOpt settings — they are optional feature configs
|
||||
contributed by users for specific scenarios. The paper-reported numbers
|
||||
were obtained with the default settings, not these.
|
||||
|
||||
- **[`configs/features/soft_gate.yaml`](configs/features/soft_gate.yaml)**
|
||||
*(PR #25, contributed by [@lvbaocheng](https://github.com/lvbaocheng))* —
|
||||
switches `gate_metric` to `soft` (or `mixed`). See the comment at the
|
||||
top of the file for when to use and when not to.
|
||||
|
||||
---
|
||||
|
||||
## Extensibility & WebUI
|
||||
|
||||
### Adding a new backend
|
||||
|
||||
A backend = a chat / exec target (e.g. `openai_chat`, `claude_chat`,
|
||||
`qwen_chat`, `minimax_chat`, `codex_exec`, `claude_code_exec`). See
|
||||
[`docs/guide/new-backend.md`](docs/guide/new-backend.md) for the full
|
||||
contract; in short you add a `skillopt/model/<name>_backend.py` module,
|
||||
register it in `skillopt/model/common.py` + `backend_config.py`, and wire
|
||||
it through the router in `skillopt/model/__init__.py`. `qwen_backend.py`
|
||||
and `minimax_backend.py` are good templates.
|
||||
|
||||
### Adding a new benchmark
|
||||
|
||||
A benchmark = a `skillopt/envs/<name>/` package with a `dataloader.py`, a
|
||||
`rollout.py`, and an `initial.md` seed skill. See
|
||||
[`docs/guide/new-benchmark.md`](docs/guide/new-benchmark.md) for the full
|
||||
contract; the simplest reference is `skillopt/envs/searchqa/`.
|
||||
|
||||
### WebUI
|
||||
|
||||
Launch the monitoring dashboard (optional):
|
||||
|
||||
```bash
|
||||
pip install -e ".[webui]"
|
||||
python -m skillopt_webui.app
|
||||
```
|
||||
|
||||
| Flag | Default | Description |
|
||||
|---|---|---|
|
||||
| `--port` | 7860 | Server port |
|
||||
| `--host` | `0.0.0.0` | Bind address |
|
||||
| `--share` | off | Create a public Gradio share link |
|
||||
|
||||
---
|
||||
|
||||
## Citation
|
||||
|
||||
```bibtex
|
||||
@misc{yang2026skilloptexecutivestrategyselfevolving,
|
||||
title={SkillOpt: Executive Strategy for Self-Evolving Agent Skills},
|
||||
author={Yifan Yang and Ziyang Gong and Weiquan Huang and Qihao Yang and Ziwei Zhou and Zisu Huang and Yan Li and Xuemei Gao and Qi Dai and Bei Liu and Kai Qiu and Yuqing Yang and Dongdong Chen and Xue Yang and Chong Luo},
|
||||
year={2026},
|
||||
eprint={2605.23904},
|
||||
archivePrefix={arXiv},
|
||||
primaryClass={cs.AI},
|
||||
url={https://arxiv.org/abs/2605.23904}
|
||||
}
|
||||
```
|
||||
-1222
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,79 @@
|
||||
# Paper-aligned SkillOpt reference skills (GPT-5.5)
|
||||
|
||||
This folder provides a subset of the paper's main Table 1 GPT-5.5 optimized
|
||||
skills as reference artifacts — one `gpt5.5_skill.md` per currently included
|
||||
benchmark. You can plug them into `scripts/eval_only.py` to evaluate the
|
||||
provided skills on a given split without re-running the training loop.
|
||||
|
||||
> These are checkpoints associated with the paper, not a general-purpose
|
||||
> tool. They're here so you can verify the reported numbers and use the
|
||||
> skills as portable artifacts. If you want to *train* your own skill,
|
||||
> use `scripts/train.py` per the top-level README.
|
||||
>
|
||||
> This is the first artifact batch. We plan to continue uploading the
|
||||
> remaining optimized skills and benchmark split manifests as they are
|
||||
> cleaned and verified.
|
||||
|
||||
## What's here
|
||||
|
||||
| Benchmark | Skill artifact | Matching config |
|
||||
|---|---|---|
|
||||
| SearchQA | `ckpt/searchqa/gpt5.5_skill.md` | `configs/searchqa/default.yaml` |
|
||||
| ALFWorld | `ckpt/alfworld/gpt5.5_skill.md` | `configs/alfworld/default.yaml` |
|
||||
| DocVQA | `ckpt/docvqa/gpt5.5_skill.md` | `configs/docvqa/default.yaml` |
|
||||
| LiveMathematicianBench | `ckpt/livemath/gpt5.5_skill.md` | `configs/livemathematicianbench/default.yaml` |
|
||||
| OfficeQA | `ckpt/officeqa/gpt5.5_skill.md` | `configs/officeqa/default.yaml` |
|
||||
| SpreadsheetBench | `ckpt/spreadsheetbench/gpt5.5_skill.md` | `configs/spreadsheetbench/default.yaml` |
|
||||
|
||||
Each file is a plain Markdown skill document (~2k–13k chars). It contains a
|
||||
protected `SLOW_UPDATE` section at the end that holds epoch-wise
|
||||
longitudinal guidance — that's expected, not a formatting issue.
|
||||
|
||||
## How to evaluate a provided skill
|
||||
|
||||
`scripts/eval_only.py` runs a single skill against a data split without
|
||||
invoking the optimizer. Example for SearchQA against the test split:
|
||||
|
||||
```bash
|
||||
python scripts/eval_only.py \
|
||||
--config configs/searchqa/default.yaml \
|
||||
--skill ckpt/searchqa/gpt5.5_skill.md \
|
||||
--split valid_unseen \
|
||||
--split_dir data/searchqa_id_split \
|
||||
--azure_openai_endpoint https://your-resource.openai.azure.com/ \
|
||||
--target_model gpt-5.5
|
||||
```
|
||||
|
||||
Substitute the benchmark, config, skill path, and `--split_dir` to evaluate
|
||||
any of the other five. `--split valid_unseen` is the test split, `valid_seen`
|
||||
is the selection / validation split, `train` is the training split, and
|
||||
`all` runs all three.
|
||||
|
||||
## On comparing to the paper numbers
|
||||
|
||||
To compare against the paper-reported cells, use the same dataset split and
|
||||
scorer. SearchQA's split is checked in at `data/searchqa_id_split/` (400
|
||||
train / 200 selection / 1400 test). For the other benchmarks, point
|
||||
`--split_dir` at your own materialized split; the loader is deterministic
|
||||
from `split_seed` (default `42`) + `split_ratio` (default `2:1:7`) when
|
||||
`split_mode: ratio` is used, so a given `data_path` + seed reproduces
|
||||
across machines. Explicit per-benchmark split manifests are being prepared
|
||||
for upload — see issues #14 and #21.
|
||||
|
||||
## Why force-accept vs. gated slow-update matters
|
||||
|
||||
These `ckpt/` skills were produced with the gated slow-update semantics
|
||||
described in paper Section 3.6:
|
||||
|
||||
```yaml
|
||||
optimizer:
|
||||
slow_update_gate_with_selection: true
|
||||
```
|
||||
|
||||
Current `main` defaults to `false` (force-accept mode), a newer
|
||||
post-submission behavior where the slow-update guidance is written into
|
||||
`current_skill` and `best_skill` unconditionally at the epoch boundary. If
|
||||
you re-train with the current default, you may produce a *different*
|
||||
`best_skill.md` than the one checked in here. Both modes are supported;
|
||||
see the top-level README's "Configuration -> Slow-update acceptance mode"
|
||||
section.
|
||||
@@ -0,0 +1,113 @@
|
||||
# ALFWorld Embodied Agent Skill
|
||||
|
||||
## Overview
|
||||
This skill guides agents operating in the ALFWorld text-based embodied environment.
|
||||
The agent must complete household tasks by navigating rooms, interacting with objects,
|
||||
and using appliances. Actions must be chosen from the admissible action list provided
|
||||
at each step.
|
||||
|
||||
**Output format**: Always output `<think>...</think>` for reasoning, then `<action>...</action>` for the chosen action.
|
||||
|
||||
---
|
||||
|
||||
## Task Types
|
||||
|
||||
| Type | Goal | Key Steps |
|
||||
|------|------|-----------|
|
||||
| Pick & Place | Put object X in/on receptacle Y | Find X -> take X -> go to Y -> put X in/on Y |
|
||||
| Pick Two & Place | Put two instances of X in/on Y | Find X1 -> take -> place -> find X2 -> take -> place |
|
||||
|
||||
### Pick Two Object Bookkeeping
|
||||
For `pick_two_obj_and_place`, choose one destination receptacle instance once it is opened/usable, and remember it as the target. Both object instances should be placed into that same remembered receptacle. After placing the first object, do not remove it again; if the second object was already seen, return directly to its remembered location rather than searching randomly. If the two objects are accidentally split across different receptacles, consolidate them into the chosen target receptacle.
|
||||
| Examine in Light | Examine object X under desklamp | Find X -> take X -> find desklamp -> use desklamp |
|
||||
|
||||
| Examine in Light detail | Final interaction | While holding X where a desklamp is visible, use the desklamp; do not try to place X on the lamp first. |
|
||||
| Clean & Place | Clean object X and put in/on Y | Find X -> take X -> go to sink -> clean X -> go to Y -> put X |
|
||||
| Heat & Place | Heat object X and put in/on Y | Find X -> take X -> go to microwave -> heat X -> go to Y -> put X |
|
||||
| Cool & Place | Cool object X and put in/on Y | Find X -> take X -> go to fridge -> cool X -> go to Y -> put X |
|
||||
|
||||
---
|
||||
|
||||
## General Principles
|
||||
|
||||
1. **Decompose the task**: Parse the goal into ordered sub-goals (locate, acquire, transform, deliver). Complete each before moving to the next.
|
||||
2. **Systematic exploration**: Search each surface and container exactly once before revisiting. Open closed containers (drawers, cabinets, fridge) before judging them empty.
|
||||
|
||||
- Prioritize semantically likely locations first, then broaden systematically: food in fridges/on countertops or dining tables; dishes/utensils/cookware on countertops, dining tables, stoveburners, cabinets, or drawers; office/bedroom items on desks, shelves, dressers, sidetables, or in drawers; newspapers on coffeetables, sidetables, sofas, or tvstands; toiletries/cleaning items near sinks, bathroom counters, shelves, carts, or cabinets.
|
||||
|
||||
- For portable kitchen targets such as bread, mugs, cups, plates, bowls, and utensils, check broad exposed surfaces early: after one or two empty countertops, try dining tables or other open surfaces before opening many cabinets/drawers. For small office/bedroom targets, alternate drawers with exposed desks, shelves, sidetables, and dressers rather than exhausting drawers first.
|
||||
|
||||
- Keep a persistent **searched set** of receptacle instances, e.g. `drawer 1`, `shelf 3`, `countertop 2`. Once an observation shows no needed target object there, mark it searched and do not call it “unexplored” later.
|
||||
- If all locations in the current preferred class are searched, **broaden to any unvisited admissible `go to ...` location** instead of restarting the same sequence. Search broadly across surfaces, furniture, containers, and appliances when relevant.
|
||||
- If a visible object is itself an openable/container-like object, such as a box, and opening/examining it is admissible, inspect it before leaving the area.
|
||||
3. **Grab immediately**: When a required object is visible and reachable, take it right away before moving elsewhere.
|
||||
|
||||
- Pick up only the exact requested object type. Similar or related objects, such as a cup when the task asks for a mug, a spoon when it asks for a knife, or a pot when it asks for a pan, are distractors; leave them in place and mark that location searched for the target.
|
||||
4. **Transform before placing**: If the task requires cleaning, heating, or cooling, perform the state change at the appropriate appliance before heading to the final destination.
|
||||
|
||||
- Do not repeatedly revisit the sink, microwave, fridge, or final destination before holding the target object. If you find the appliance early, remember its location, then resume searching unvisited object locations until the target object is acquired.
|
||||
|
||||
- Use direct admissible appliance/tool commands immediately when available, such as `clean X with sinkbasin`, `heat X with microwave`, `cool X with fridge`, or `use desklamp`. Do not waste steps opening, closing, toggling, or examining the appliance unless the needed action is unavailable or opening is required for searching/placing.
|
||||
5. **Direct delivery**: Once holding the transformed (or untransformed) goal object, navigate straight to the target receptacle and place it.
|
||||
|
||||
- Remember known destination receptacles and return directly to the same instance after pickup/transformation. If the destination is also a semantically likely source location, check/open it early rather than only after exhaustive search: food may already be in the fridge, utensils may be on the diningtable, newspapers may be on/near the sofa, and a target drawer can be opened early for pick-two tasks. If the object starts at the destination but needs cleaning/heating/cooling, take it out, transform it, then return to that same instance and place it back.
|
||||
6. **Track progress**: Maintain an internal count of how many objects still need to be found and placed. Only stop searching when the count reaches zero.
|
||||
7. **Avoid loops**: Never repeat the same action more than twice in a row. If stuck, move to a different unexplored location.
|
||||
8. **Only choose admissible actions**: Always pick an action from the admissible action list. Do not invent actions.
|
||||
|
||||
---
|
||||
|
||||
## Common Mistakes to Avoid
|
||||
|
||||
- **Revisiting searched locations**: Keep track of which surfaces/containers have been checked; do not re-examine them.
|
||||
- **Ignoring visible objects**: If the target object appears in the observation, pick it up immediately.
|
||||
- **Skipping state changes**: Do not place an object at the destination without first cleaning/heating/cooling it when required.
|
||||
- **Premature termination**: Do not stop the episode until all goal conditions are verified as met.
|
||||
- **Action loops**: Repeatedly toggling or examining the same object wastes steps. Move on to new locations instead.
|
||||
|
||||
### Hard Search-Loop Recovery
|
||||
|
||||
- **Exact-instance lockout before pickup**: once a receptacle/surface instance has been observed and does not contain the target object, do not go back to that exact instance while still searching for the object. A phase change, such as holding the object or needing final delivery, is the only reason to return.
|
||||
- **Fast broadening threshold**: after 3-4 misses in the same receptacle class, switch to a different likely class or any unvisited admissible location instead of continuing or restarting that class, unless the target has already been seen there.
|
||||
- **No search reset by recency**: do not say a location is "unsearched" merely because it was not in the last few observations. The searched set is global for the whole episode.
|
||||
- **Finite-class exhaustion**: if all visible instances of a small class have been checked once, such as all stoveburners, diningtables, countertops, or shelves, mark that class exhausted for object search and do not start a second pass. Remember a usable destination instance, then search different receptacle classes.
|
||||
- **Unvisited beats likely-but-searched**: after several misses, prefer any admissible unvisited `go to`, `open`, or `examine` target over revisiting a semantically likely but already-searched location.
|
||||
- **Destination surfaces before pickup**: if the destination receptacle is also a likely object location, inspect each instance at most once before pickup. If it lacks the object, remember it as the final destination but stop using it as a search target until the object has been transformed and is ready to place.
|
||||
- **Kitchen item fallback**: for cookware and dishware, after checking obvious burners/tables/counters once, broaden to unsearched cabinets, drawers, shelves, sinkbasins, and other kitchen storage/surfaces rather than cycling among the obvious locations.
|
||||
|
||||
### Strict Search Ledger Action Filter
|
||||
|
||||
Before every empty-handed search action, apply this hard filter:
|
||||
|
||||
1. If a required target object is visible, take it immediately.
|
||||
2. Otherwise choose an exact receptacle/surface/container instance whose contents have not yet been observed in the current object-search phase.
|
||||
3. Reject any `go to`, `examine`, or `open` action for an exact instance already observed to lack the target, even if it is semantically likely, nearby, recently mentioned, or the final destination type.
|
||||
4. If all likely instances are rejected by the ledger, broaden to any unvisited admissible location/class instead of restarting from instance 1 of a searched class.
|
||||
|
||||
The searched ledger survives inventory checks, appliance visits, placing the first object in a pick-two task, and putting down an irrelevant inspected object/container. These events are not permission to rescan shelves, drawers, cabinets, tables, counters, or destination receptacles from the beginning.
|
||||
|
||||
### Destination-as-Source Lockout
|
||||
|
||||
When the final receptacle type is also a plausible source location, inspect each visible destination instance at most once before pickup. After it lacks the target, remember a usable destination instance and lock that exact instance out of object search until you are holding the required object ready for delivery. Do not alternate between destination instances and other searched source instances while still empty-handed.
|
||||
|
||||
### Pick-Two Phase Memory
|
||||
|
||||
After placing the first object in a pick-two task, do not begin a fresh room/class search. If another required instance was previously seen, return directly to that remembered source location for the second pickup. If no second instance is remembered, continue from the existing unsearched-location ledger rather than revisiting locations already checked before the first placement.
|
||||
|
||||
<!-- SLOW_UPDATE_START -->
|
||||
Preserve the successful pattern: when the exact requested object is visible, take it immediately; perform the required clean/heat/cool/use action as soon as the correct command is admissible; then deliver directly to the remembered destination.
|
||||
|
||||
Treat tool locations as tools, not repeated search targets. If a sinkbasin, fridge, microwave, desklamp, or destination receptacle has already been checked and does not contain the target while you are empty-handed, remember it for later but do not revisit it until you are holding the required object or ready to place/use it.
|
||||
|
||||
Use a next-unsearched-instance pointer for every numbered class. If you leave cabinets, drawers, shelves, countertops, or stoveburners and later return to that class, resume at the lowest exact instance not yet observed; never restart at instance 1 and never revisit an instance already observed to lack the target.
|
||||
|
||||
For pan-to-stoveburner tasks, search in a step-efficient order: make one quick pass over stoveburners only to find a pan or remember an empty destination, then leave stoveburners until delivery. Next check countertops/islands and sinkbasins. Then prioritize cabinets in numeric order, opening each closed cabinet and observing its contents, before low-yield drawers. Do not abandon cabinet search to revisit searched stoveburners, countertops, or drawers.
|
||||
|
||||
For kettle/teapot clean-and-place tasks, after checking obvious countertops/islands, check stoveburners and sinkbasins once, then cabinets in numeric order. If several cabinets are empty, continue to the next unsearched cabinet or broaden to unvisited shelves/carts/dining tables; do not return to already searched countertops. Remember one open/empty cabinet as the final destination, but do not keep using searched cabinets as search targets.
|
||||
|
||||
For dishsponge clean-and-place tasks, check sinkbasin and nearby countertops once, then search unvisited cabinets, drawers, shelves, carts, and other storage/surfaces. Because the sink is needed for cleaning, remember it after the first visit; do not go back to the sink while empty-handed just because the sponge is likely near it. Because shelf is the destination, remember a usable shelf after inspecting it once; after a shelf lacks the sponge, search only unvisited shelves or other unvisited locations until the sponge is found.
|
||||
|
||||
When the step budget is running and you are still empty-handed, prefer any unvisited admissible location over any searched likely location. A location being semantically likely, useful later, or recently mentioned is never a reason to rescan it before acquisition.
|
||||
|
||||
Do not let the broadening threshold cause class restarts. Broadening means move to a different unvisited class or continue at the next unsearched instance of a promising storage class; it never means cycling back through exact instances already observed.
|
||||
<!-- SLOW_UPDATE_END -->
|
||||
@@ -0,0 +1,26 @@
|
||||
# DocVQA Skill
|
||||
|
||||
## Visual Evidence Discipline
|
||||
- Read the document carefully before answering.
|
||||
- Prefer the smallest exact text span that answers the question.
|
||||
|
||||
- For questions asking for a value, count, page number, date, or graph reading, return only the requested value span; omit nearby labels, category names, units, or explanatory words unless the question explicitly asks for them.
|
||||
- When several nearby strings look similar, choose the one whose surrounding labels or layout best match the question.
|
||||
|
||||
## Exact Answer Discipline
|
||||
- Copy names, numbers, and dates exactly from the document whenever possible.
|
||||
|
||||
- Preserve the document's exact spelling and punctuation for names and quoted phrases; do not substitute similar letters or change straight/curly quotes, spacing, or parentheses when the visible text provides them.
|
||||
- Prefer direct extraction over paraphrase.
|
||||
- Before finalizing, compare the answer against nearby alternatives and keep the best-supported exact span.
|
||||
|
||||
## Structured Layout Lookup
|
||||
- For tables, first find the row or entry named in the question, then read the value under the requested column, header, date, or category; answer with that cell only.
|
||||
- For forms, receipts, or labeled fields, locate the exact role, party, or field label mentioned in the question, then copy the filled-in value from the same line, box, block, or immediately adjacent field.
|
||||
- For table-of-contents, indexed, numbered, or bulleted lists, match the requested title, entry, or point number, then follow the same line or list item to the associated value; do not take a nearby value from another item.
|
||||
|
||||
## Anchored Handwriting / Nearby Text
|
||||
- For handwritten or list/table questions with an anchor term, first locate the anchor, then inspect the immediately adjacent text in the same row, column, or nearby margin. If legible, provide the best-supported nearby span rather than leaving the answer blank.
|
||||
|
||||
<!-- SLOW_UPDATE_START -->
|
||||
<!-- SLOW_UPDATE_END -->
|
||||
@@ -0,0 +1,35 @@
|
||||
# Live Mathematical MCQ Heuristics
|
||||
|
||||
## Option Comparison
|
||||
|
||||
### Meta-Options About Stronger Results
|
||||
- Treat options of the form “one of the remaining options is correct, but a stronger result can be proven” as serious candidates, especially when the question asks for the strongest statement.
|
||||
- If a concrete option is true but your theorem or derivation gives a strictly stronger conclusion not exactly listed, choose the meta-option rather than the weaker concrete statement.
|
||||
- When options are nested by strength, rank them explicitly before answering: e.g. finite-time blowup is stronger than merely “not globally bounded”; positive stable growth is stronger than ordinary unboundedness; sharper constants, rates, exceptional-set bounds, endpoint inclusion, or full equivalences are stronger than weaker asymptotic versions.
|
||||
- Compare all options before committing. The correct choice is often the strongest statement justified by the question, while nearby distractors are weaker, overstrong, or miss an equality case.
|
||||
- Track exact quantifiers such as "there exists", "for every", "if and only if", and "exactly when".
|
||||
|
||||
## Theorem-Level Precision
|
||||
|
||||
- Do not add converse, realization, or classification claims unless the theorem explicitly proves them. Phrases such as “conversely,” “every such parameter occurs,” “if and only if,” or “exactly all” add strength beyond a one-way implication.
|
||||
- Check whether an option weakens the conclusion by dropping a characterization, equality clause, or full equivalence.
|
||||
- Check whether an option overstates the theorem by upgrading regularity, removing scale restrictions, or changing an existential statement into a universal one.
|
||||
|
||||
## Hypotheses
|
||||
|
||||
### Exact Conditions and Thresholds
|
||||
- For biconditional/equivalence questions, reject conditions that are merely necessary or merely sufficient. A broader condition, such as congruence modulo a divisor instead of modulo the full modulus, is usually weaker and not equivalent unless the domain collapses the extra cases.
|
||||
- For threshold conditions, verify the exact sign and endpoint: distinguish \(\mu_0\) from \(-\mu_0\), \(<\) from \(\le\), and whether the equality case belongs to the positive, zero, or negative parameter regime.
|
||||
- When options differ by “for every” vs “for sufficiently large,” local vs global domains, strict vs non-strict inequalities, or dependence of constants, rank them by logical strength and match the sharpest justified version.
|
||||
- Verify the hypotheses and domain carefully. Distractors often keep the theorem shape but alter the required assumptions.
|
||||
- Pay close attention to equality cases, extremal conditions, and whether a result applies to the full family or only a restricted subfamily.
|
||||
|
||||
## Final Answer
|
||||
- Output the final answer as the single option label only.
|
||||
|
||||
## Exact Scope and Quantitative Wording
|
||||
- Distinguish global conclusions from localized or completed ones. Equivalence after localization, completion, or at each prime/scale is usually weaker than an unqualified equivalence.
|
||||
- In estimate-heavy options, compare every quantitative detail: exponent, derivative index range, constants and their parameter dependence, log factors, additive terms, one-sided vs two-sided notation, and pointwise vs uniform convergence.
|
||||
|
||||
<!-- SLOW_UPDATE_START -->
|
||||
<!-- SLOW_UPDATE_END -->
|
||||
@@ -0,0 +1,50 @@
|
||||
# OfficeQA Skill
|
||||
|
||||
## Retrieval Discipline
|
||||
|
||||
- When an external official time-series observation is needed, prefer the source's series/data-download/table page once identified. If exact-date or guessed-value searches return empty results, stop repeating them; broaden to the official series name/code plus `data` or `download` and use the table values.
|
||||
|
||||
- Treat provided/oracle parsed pages as primary evidence: if they contain the relevant table and period, extract directly from them before searching elsewhere; search only for missing continuation pages, missing periods, or an official actual value not present.
|
||||
- Start by narrowing to the most likely candidate file before reading long passages.
|
||||
- Prefer targeted search terms that name the exact entity, period, measure, or table concept from the question.
|
||||
- After a promising match, read only a small surrounding span and verify it matches the requested year, basis, and unit.
|
||||
|
||||
- If the requested date range extends beyond the provided/oracle page, first enumerate the required periods and verify that every period is present in evidence. Do not compute from a partial ledger or fill missing periods from memory; retrieve continuation pages, adjacent issues, or a later issue of the same table that contains the missing dates/revisions.
|
||||
|
||||
## Evidence Discipline
|
||||
- Extract the exact value from the retrieved text before doing any arithmetic.
|
||||
- Keep track of each operand's period, unit, and semantic role so nearby proxy values are not mixed in.
|
||||
|
||||
- For Treasury financing narratives, label each amount by transaction role before calculating: offered amount, tenders/subscriptions received, tenders accepted, competitive/noncompetitive accepted, foreign or Government-account exchange tenders, refunding, and **new cash** are not interchangeable.
|
||||
- When converting currencies or scales, make a direction ledger first: source table unit, source currency, exchange-rate orientation (foreign currency per U.S. dollar means divide by the rate; U.S. dollars per foreign unit means multiply), and requested final unit.
|
||||
|
||||
- For tables, align values by row label and exact column header, not proximity alone; watch for continued or unlabeled columns, footnotes, adjacent amount-versus-percent columns, fiscal-year versus calendar-year sections, and repeated month rows under different year blocks.
|
||||
- If the question asks for a transformed or derived quantity, compute only after confirming every operand.
|
||||
|
||||
- For derived comparisons, preserve the direction and sign implied by the wording: “change from A to B” means B minus A; “former than latter” means former minus latter; “share accounted for by X” means X divided by the stated total; paired “gap” questions require computing each within-row difference before ranking.
|
||||
- For statistical, regression, correlation, and growth-rate questions, write a formula ledger before calculating: confirm the exact series/endpoints, ordered vector, elapsed intervals, and requested convention such as continuously compounded rate, CAGR, Pearson correlation, or OLS index/year choice.
|
||||
- For multi-stage questions where one table determines the period/entity used in another lookup, freeze that derived key with evidence first, then retrieve the second measure only for that exact month/year/reporting date/entity.
|
||||
|
||||
- For inclusive time-series ranges, make a period-by-period ledger covering every requested month/year exactly once, preserving calendar versus fiscal basis, end-of-month or end-of-fiscal-month status, source units, and any specified adjustments.
|
||||
|
||||
- For statistical transforms over time-series windows, confirm endpoint inclusion/exclusion exactly as worded, use consecutive time indices for trend regressions when appropriate, sort values before medians, and for logarithmic growth use ln(final/initial) before converting to the requested percentage format.
|
||||
|
||||
## Final Answer Discipline
|
||||
|
||||
- Before finalizing, enforce the requested unit and format: convert thousands/millions/billions or full nominal dollars as needed, then apply no-comma, fixed-decimal, whole-number, or nearest-tenth/thousandth formatting exactly as asked.
|
||||
- Return the final answer only after one last consistency check against the retrieved evidence.
|
||||
- Copy the final answer from a checked value, not from an unverified intermediate guess.
|
||||
|
||||
## Statistical and Time-Series Calculation Checks
|
||||
|
||||
- Before computing any statistic, write the intended formula and denominator convention. If the prompt explicitly says **population standard deviation**, divide by `n`; if it says **sample**, divide by `n-1`; for a z-score comparing one observation against a small set of comparison months/periods and no population convention is stated, estimate dispersion with the **sample** standard deviation of the comparison set. Do not round intermediate operands, weighted averages, logs, exchange-rate conversions, or standard deviations before the final requested rounding.
|
||||
- For long inclusive ranges, first enumerate the expected count of observations and the first/last period, then verify the ledger has exactly that count. Exclude totals, cumulative-to-date columns, comparable-period columns, estimates, and extra latest-month columns outside the requested calendar or fiscal range.
|
||||
- When a page contains multiple nearby sections with similar labels, use only the section whose title and row label match the requested measure exactly; do not compute from the first visible table if the requested measure/table title is absent or only partially shown.
|
||||
- For Treasury security quotations, obey the table's quote basis. If the table states that price decimals are 32nds, convert quotes such as `99.27` as `99 + 27/32`, not as decimal `99.27`. If a task asks for smoothing, averaging, or forecasting in a target currency using period-specific exchange rates, convert each period's observation to the target currency first unless the prompt explicitly says to compute in the source currency and convert only the final result.
|
||||
|
||||
## Stricter Final Formatting
|
||||
|
||||
- Match any requested output template exactly. Unless the prompt explicitly asks for unit words or explanatory text, return only the numeric value or requested list; do not append words such as `million`, `dollars`, `percent`, or `percentage points`. Include symbols/commas only when the prompt requests currency-formatted output or the answer format clearly requires them.
|
||||
|
||||
<!-- SLOW_UPDATE_START -->
|
||||
<!-- SLOW_UPDATE_END -->
|
||||
@@ -0,0 +1,71 @@
|
||||
# Question Answering Skill
|
||||
|
||||
(No learned rules yet. Rules will be added through the reflection process.)
|
||||
|
||||
## Concise Answer Normalization
|
||||
- Prefer the shortest unambiguous answer that directly satisfies the question. Do not include generic descriptors, legal suffixes, or expanded formal names unless the question specifically asks for the full official name or the descriptor is necessary to identify the entity.
|
||||
|
||||
- If the answer appears inside a longer descriptive phrase, strip words that merely repeat the clue's requested type or modifiers already stated in the clue. For short-answer trivia, return the distinctive core entity or headword rather than role titles, product flavor adjectives, or place/facility designators, even when those words are part of a fuller official phrase, unless the full official name is explicitly requested.
|
||||
- For place/name-etymology questions asking for “the name” or “the word” that means something, answer the distinctive name/word itself rather than a larger phrase with a generic type label.
|
||||
|
||||
- For natural geographic features, preserve conventional feature designators such as “Lake,” “River,” “Bay,” “Gorge,” “Mount,” or “Island” when they are part of the proper name or match the requested feature type. Do not shorten “Lake Okeechobee,” “Tampa Bay,” or “Olduvai Gorge” to an ambiguous base name merely to be concise.
|
||||
- For companies, brands, and organizations, answer the common distinctive name when sufficient; omit additions such as “Company,” “Corporation,” “Inc.,” etc. unless explicitly required.
|
||||
|
||||
- Preserve the answer surface form supported by the strongest evidence when exact variants differ: spelling, capitalization, punctuation, and word order can matter. Do not substitute an equivalent official/common variant such as an alternate spelling or inverted institution name if a direct title/snippet/answer field gives the expected form.
|
||||
|
||||
- When copying titles or quoted names, preserve ordinary ASCII punctuation from the evidence, especially straight apostrophes (`'`). Do not replace them with typographic curly quotes/apostrophes unless that exact stylized form is explicitly shown as the supported answer.
|
||||
|
||||
- For nicknames, epithets, saints, and quoted titles, copy the supported surface form exactly, including spacing, capitalization, and conventional abbreviations such as “St.” Do not normalize a stylized or quoted form into a lowercase dictionary word or an expanded spelling when the clue/evidence points to the stylized answer.
|
||||
|
||||
- For person answers in trivia or crossword-style clues, prefer the conventional supported name. Use just a surname, first name, or saint/regnal name only when the clue/source clearly expects that short form; otherwise use the canonical full personal name from the strongest evidence or answer field, especially when a lone given name would be ambiguous.
|
||||
- Return the grammatical base form expected by the clue. Do not add a plural `s` merely because a crossword source pluralizes a shared name or category; if the clue lists people sharing a first name, answer the singular given name.
|
||||
|
||||
- For common-noun category answers, default to the singular dictionary headword in trivia/crossword-style clues, even if the clue uses plural words like “these,” “those,” “places,” or “items” for grammar. Use a plural only when the term is inherently plural or an answer field/source clearly gives a plural phrase.
|
||||
|
||||
- For common-noun clues about things being replaced, used in place of, or substituted by another system/item, answer the broad headword for the thing replaced unless a narrowing modifier is required by the clue or answer field. Do not add adjectives such as “letter,” “regular,” or “standard” merely because they appear in explanatory context.
|
||||
- For fill-in-the-blank or definitional clues using words like “this” or “that,” provide a standalone noun phrase. Avoid context-dependent pronouns or possessives from the source text; use a natural article such as “the” when needed (e.g., answer “the highest point,” not “its highest point”).
|
||||
|
||||
## Context-Grounded Evidence Matching
|
||||
- Start by identifying the most distinctive terms in the question: proper names, dates, titles, quoted phrases, unusual words, roles, relationships, and category descriptors.
|
||||
- Prioritize passages or document titles where several distinctive clue terms occur together, especially if the wording directly repeats or closely paraphrases the question.
|
||||
- Treat document titles as useful evidence: the answer is often named in a title while the snippet confirms the clue facts.
|
||||
|
||||
- Do not assume the document title itself is the answer. If the requested type differs from the title entity, use the title as context and extract the matching typed entity from the snippet or clue relationship.
|
||||
|
||||
- For “known as,” “called,” “defined as,” or category/type clues, choose the canonical term explicitly used in the strongest matching title/snippet or scraped answer field rather than inventing a related derivative or near-synonym from the clue wording. When multiple plausible candidates appear, prefer the candidate whose evidence directly states the requested relationship and repeats the most distinctive clue facts.
|
||||
- Ignore noisy results that only match generic words; prefer evidence that directly connects the clue facts to one specific entity.
|
||||
|
||||
## Clue Interpretation and Answer Type
|
||||
- For Jeopardy-style wording such as “this man,” “this group,” “this film,” “this country,” “this system,” “he,” or “his wife,” infer the expected answer type before choosing the answer.
|
||||
- Use that expected type to validate candidates: answer with the concise person, place, title, organization, object, term, or phrase requested by the clue.
|
||||
|
||||
- Treat modifiers attached to the requested type as hard filters, not background flavor: constraints like dates, “largest,” “2-letter-named,” “1978 remake,” “hot dog brand,” “dual throne,” or “on this company’s board” must all fit the candidate before you answer.
|
||||
- For clues centered on creative works such as books, films, plays, songs, poems, or other media, first determine whether the clue asks for the work itself, its creator, a performer or cast member, a character, a quotation source, or a setting. Verbs such as “wrote,” “directed,” “stars,” “played,” and “set in,” plus pronouns like “he” or “her,” usually determine the target.
|
||||
|
||||
- For fill-in-style clues with placeholders such as “this,” “these,” or “one of these,” substitute each candidate back into the clue and choose the concise answer that makes the full phrase, title, or fact read correctly.
|
||||
- For terse clues that are just examples or names separated by commas, slashes, or “or,” infer the shared category, class, or synonym that links them, then answer with that concise common term.
|
||||
|
||||
- For crossword-style clues, treat parenthetical numbers or stated letter counts as hard constraints on the answer length, and omit generic labels that would violate them. In dual-definition clues using wording like “X, or what Y does,” choose the single word that satisfies both senses and preserve the required inflected form.
|
||||
- If the clue references an unavailable image or link with wording like “seen here,” “pictured,” or parenthetical visual hints, rely on the textual clues and context to infer the answer; do not treat the missing image as necessary evidence.
|
||||
- If multiple snippets support the same entity, use that corroboration to choose the canonical/common form of the answer.
|
||||
|
||||
## Trivia / Jeopardy Snippet Formats
|
||||
- Retrieved trivia snippets may contain the clue and answer in scraped formats such as `CATEGORY | clue | answer`, `clue. ANSWER`, or labels like `right:`.
|
||||
- When the question text matches the clue in such a snippet, extract the answer field or adjacent answer name, not the category or the whole clue sentence.
|
||||
|
||||
## Common Clue Traps
|
||||
- Watch for inverse relationships: if the clue says “His third wife was Jiang Qing,” the requested answer is the husband, not Jiang Qing.
|
||||
|
||||
- More generally, preserve relation direction in clues: “A is evidence of this B,” “A is related to this language,” or “home to these characters” asks for the target of the relationship, not the entity already named in the clue.
|
||||
|
||||
- When a clue says examples, models, breeds, members, or items “include,” “like,” or “such as” named entities, treat those names as evidence for the requested parent class or entity. Answer the encompassing brand, animal, category, place, or term requested by “this,” not one of the examples already given.
|
||||
- If the question gives the start of a quotation or phrase, answer with the exact missing continuation from the context.
|
||||
|
||||
- For song, poem, nursery-rhyme, or quotation clues, first decide whether the question asks for a missing word or phrase from the quote or for the associated creator, performer, or work; use pronouns and answer-type signals to choose the right target.
|
||||
- When a clue asks for a constrained form such as a first name, abbreviation, acronym, or lyric word, return that exact form rather than the fuller person, title, or explanation; preserve conventional punctuation or spelling when it is part of the requested form.
|
||||
- If the clue contains wordplay, quotation marks, or puns, treat them as hints, but answer with the real entity supported by the evidence.
|
||||
|
||||
- If a clue includes a quoted title, quoted narration or lyric, named event, slogan, or other distinctive phrase but asks for an associated “this” entity, treat the quote or name as evidence to identify the requested person, work, place, group, category, source, or term; do not return the quoted anchor unless the clue explicitly asks for it.
|
||||
|
||||
<!-- SLOW_UPDATE_START -->
|
||||
<!-- SLOW_UPDATE_END -->
|
||||
@@ -0,0 +1,133 @@
|
||||
# Spreadsheet Manipulation Skill (xlsx)
|
||||
|
||||
## Overview
|
||||
This skill guides agents in manipulating Excel (.xlsx) spreadsheets using Python.
|
||||
|
||||
**Primary libraries**: `openpyxl` (structure-preserving read/write), `pandas` (data transformation).
|
||||
Never use any other third-party libraries.
|
||||
|
||||
---
|
||||
|
||||
## Common Workflow
|
||||
|
||||
1. **Explore** the input file: list sheets, inspect headers, check dimensions.
|
||||
|
||||
- Inspect actual workbook data beyond the preview, including nearby rows/columns, sample outputs, formulas, labels, headers, and any reference/example sheets such as `Output`, `Manual Result`, or `Desired...` tabs.
|
||||
|
||||
- Treat existing filled cells in the requested output area or adjacent example tables as semantic examples for edge cases and expected formats, but still recompute and write the complete requested target range.
|
||||
- Scan the used range for complete header groups, not just row 1. Tables may start in later rows/columns, have title rows above them, or have multiple source/result tables on the same sheet; use nearby labels and the requested output range to distinguish sources from destinations.
|
||||
- Locate tables, fields, and target ranges by header text, nearby labels, and surrounding nonblank structure rather than fixed coordinates. Build header maps from actual cells when useful, e.g. `{str(cell.value).strip(): cell.column}`.
|
||||
2. **Write `solution.py`** with `INPUT_PATH` and `OUTPUT_PATH` defined at the top.
|
||||
3. **Execute** `python solution.py` and verify the output file was created.
|
||||
4. **Confirm** the target cells/range contain the expected values.
|
||||
|
||||
---
|
||||
|
||||
## Library Selection
|
||||
|
||||
| Use case | Library |
|
||||
|----------|---------|
|
||||
| Preserve formulas, formatting, named ranges | `openpyxl` |
|
||||
| Bulk data transformation, aggregation, sorting | `pandas` → write back with `openpyxl` |
|
||||
| Simple cell read/write | `openpyxl` |
|
||||
|
||||
**Warning**: `pandas.to_excel()` silently destroys existing formulas and named ranges.
|
||||
When writing back to a spreadsheet that contains formulas, always use `openpyxl.save()`.
|
||||
|
||||
**Formula evaluation caution**: `openpyxl` can write formulas but does **not** calculate them or update cached results. If the requested output will be checked as cell values, compute the result in Python and write literal values unless the user explicitly requires live formulas. When existing formulas are inputs to your logic, load a second workbook with `data_only=True` to read cached values while saving changes through the normal workbook:
|
||||
|
||||
```python
|
||||
wb = openpyxl.load_workbook(INPUT_PATH)
|
||||
wb_values = openpyxl.load_workbook(INPUT_PATH, data_only=True)
|
||||
ws = wb["Sheet1"]
|
||||
ws_values = wb_values["Sheet1"]
|
||||
```
|
||||
|
||||
Treat wording such as “write/fix a formula,” “SUMIFS/COUNTIFS,” “VBA,” or “macro” as a description of the spreadsheet logic unless the deliverable explicitly requires live formula text, an `.xlsm`, or a preserved VBA project. For normal `.xlsx` outputs, implement the equivalent logic in Python/openpyxl and write the computed final values to the requested cells so verification does not depend on Excel recalculation or macros.
|
||||
|
||||
When the user provides an existing or broken formula, use it as a semantic specification: honor its referenced lookup ranges, criteria ranges, return ranges, aggregation intent, and error-handling behavior, then write the resulting values rather than guessing different source columns or leaving unevaluated formulas.
|
||||
|
||||
---
|
||||
|
||||
## solution.py Template
|
||||
|
||||
```python
|
||||
import openpyxl
|
||||
import pandas as pd
|
||||
|
||||
INPUT_PATH = "..." # set to the actual input path
|
||||
OUTPUT_PATH = "..." # set to the actual output path
|
||||
|
||||
wb = openpyxl.load_workbook(INPUT_PATH)
|
||||
ws = wb.active # or wb["SheetName"]
|
||||
|
||||
# --- perform manipulation ---
|
||||
|
||||
wb.save(OUTPUT_PATH)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Output Requirements
|
||||
|
||||
- Save the result to `OUTPUT_PATH`.
|
||||
- Do not hardcode row counts or column letters — iterate over actual rows in the workbook.
|
||||
- Preserve sheets and cells not mentioned in the instruction.
|
||||
|
||||
## Matching and Target Range Hygiene
|
||||
|
||||
- Choose the comparison operator from the instruction and examples: use `startswith` for “begins with”, substring search for “contains/search/occurrence”, and exact normalized equality only when a whole-cell match is implied.
|
||||
- Create small helper functions for comparisons and numeric parsing. Normalize text by trimming, collapsing repeated spaces/NBSPs, and casefolding; when names or labels have punctuation/spacing inconsistencies, consider punctuation-insensitive keys. Parse numeric text after removing commas/currency symbols while preserving signs and decimal points; skip `None`/blank and booleans for numeric tests, and handle placeholders such as `"-"`, `"$"`, `"$0"`, blanks, and numeric zero deliberately.
|
||||
- Normalize date keys deliberately: handle `datetime`/`date` objects, Excel serial numbers, and date-like strings, then compare at the granularity implied by the task, such as exact date, month, month/year, fiscal period, or year. For workday/date-window logic, compute the range in Python and exclude weekends/holidays as specified.
|
||||
|
||||
- For monthly or period summary grids, canonicalize period labels from all sources: sheet names, title text, row/column headers, text months such as `March`, and actual date cells. Match summaries by normalized period plus the other stated criteria rather than by fixed month offsets or existing formulas.
|
||||
- For date ranges and rolling windows, infer endpoint inclusivity from wording and examples. Phrases like `X to Y`, `through`, and `up to`, or examples such as `2 to 5` meaning `4 days`, usually require inclusive boundary handling.
|
||||
- For time extraction or time-threshold logic, parse `datetime`, `time`, Excel serial/fractional times, and time-like strings into real Python `time`/`datetime` values. Write real time values with an Excel `number_format` such as `hh:mm:ss AM/PM`; do not write text substrings when the result should behave as a time.
|
||||
- For joins, deduplication, grouping, interval lookups, lookup grids, and ordered outputs, build explicit normalized keys, including composite keys when the task refers to multiple fields. Preserve original source order within each group unless sorting is explicitly requested.
|
||||
- For outputs that depend on other rows or lookup grids, make a first pass to build normalized dictionaries/groups/range structures, then a second pass to write results. Avoid nested full-sheet scans per row; split delimited tokens and ignore empty tokens, and treat error literals such as `#N/A` as meaningful sentinel values when the task refers to them.
|
||||
|
||||
- For lookups, filters, joins, and label/header matching, normalize comparison keys consistently: trim whitespace, skip blanks explicitly, use case-insensitive text matching when appropriate, and treat numeric-looking IDs consistently (`330`, `330.0`, and `"330"`). Keep numeric outputs numeric; use `number_format` for display formatting instead of converting numbers to strings unless text is explicitly required.
|
||||
- When replacing a generated output area, clear only the instructed target range before writing new results so stale values/formulas do not remain. Preserve formatting, column widths, borders, formulas, and unrelated cells unless the instruction explicitly asks to change them.
|
||||
|
||||
- If the instruction includes formatting changes, apply them exactly after writing values and only to the requested cells/range. Use `openpyxl` styles for fills, alignment, fonts, borders, and number formats; convert hex colors to ARGB when needed, for example `#FFC000` → `FFFFC000`. For “format as text,” set `number_format = '@'` and write string values when the expected cell values are text.
|
||||
|
||||
- When the instruction names a destination range or columns, write derived results directly there. Do not insert rows/columns, relocate the source table, or sort/delete source records unless that structural change is explicitly requested.
|
||||
- For filtered lists, summaries, and aggregations, first collect all source records/results in memory, preserving the required order, then write from the first output row and clear leftover cells below the new results in the target columns. When adding rows, copy style/alignment/number format from an existing template row when appropriate; when deleting rows, delete from bottom to top to avoid row-index shifts.
|
||||
- Preserve intended blanks as empty cells (`None`) rather than placeholder text or `0` unless the task specifies otherwise.
|
||||
|
||||
- For numeric aggregation, crosstab, SUMIFS-like, and INDEX/MATCH-style summary outputs, infer missing-match behavior from table semantics and examples: numeric summary grids usually require literal `0` for no matching records, while filtered lists or “show only once” outputs usually require blanks (`None`).
|
||||
- For blank-sensitive logic such as “if input is blank, output blank,” evaluate the driving input with `data_only=True` when it may itself be a formula, and write `None` for truly blank outputs rather than relying on a new formula returning `""`.
|
||||
|
||||
## Robustness for Simple Fill Tasks
|
||||
|
||||
- Prefer simple, auditable row/column loops over complex workbook XML parsing unless the task truly requires unsupported workbook internals. Before returning, run the script once to catch syntax/indentation errors and verify that representative target rows were actually written.
|
||||
|
||||
<!-- SLOW_UPDATE_START -->
|
||||
When the user asks for a formula, macro, VBA code, or a fix to an Excel formula, still deliver the completed workbook state: compute the intended results in Python and write literal final values into the requested cells. Do not write formula strings unless the task explicitly says the output must contain live formulas.
|
||||
|
||||
After writing, reload or inspect the saved workbook and verify that every requested/evaluated target cell contains a non-formula literal where a value is expected. If a target cell is still `None` unexpectedly, fix the script before finishing.
|
||||
|
||||
Use existing formulas in the workbook as examples/specifications, not as output. If a cell contains a reference formula such as `=A25` or an INDEX/MATCH/SUMIFS pattern, parse what source cells/ranges/criteria it refers to, compute those results yourself, and overwrite the destination with the referenced or calculated value.
|
||||
|
||||
For blank-sensitive formula tasks, compute the branch explicitly: if the driving source cell is truly blank, write `None`; otherwise write the actual result such as `0`, `1`, a category label, or a lookup value. Never rely on `IF(...,"",...)` formulas to be recalculated later.
|
||||
|
||||
For lookup/category tasks, locate both the input rows and the lookup table by headers and nearby labels. Support exact keys, numeric-looking keys, and interval/range tables; then fill every destination row that has a driving input, not just the first visible example.
|
||||
|
||||
For “every nth row” or OFFSET-style tasks, infer the source column, first source row, and step from the provided examples or formulas, then copy the actual source values into the requested output range as literals.
|
||||
|
||||
For schedule/calendar fill tasks, build a cycle-day-to-periods mapping from the schedule/template area first, then fill the daily rows across all requested class columns based on each row’s cycle day. Preserve repeated/double periods exactly as shown by the template; do not leave formulas in the schedule cells.
|
||||
|
||||
For INDEX/MATCH problems where the first row works but subsequent rows fail, treat row labels, column/year headers, region/type criteria, and expense/category labels as a multi-key lookup. Fill the whole result matrix with values from the source data table, using cached `data_only` values when source cells are formulas.
|
||||
|
||||
For multi-step macro/VBA-style requests, implement every stated operation in the workbook, not just the first deletion/filtering step. Re-read the numbered requirements before saving and verify later computed columns, totals, and derived fields as well as the obvious filtered rows.
|
||||
|
||||
When a target range includes special rows such as `Total`, `Grand Total`, `min`, `max`, constraints, headers, or blank separators, do not apply ordinary row logic blindly to those rows. Compute totals as aggregates when indicated, and leave constraint/header/blank cells untouched unless explicitly requested.
|
||||
|
||||
For residual-balancing tasks, identify data rows separately from min/max constraint rows. Add positive residuals from unit 1 toward unit 5 without exceeding max values; subtract negative residuals from unit 5 toward unit 1 without going below min values; update only the unit cells in actual data rows.
|
||||
|
||||
For time-threshold rows, decide per row whether it is a normal data row or a summary row. Normal rows use the before/after threshold rule; summary rows should aggregate the computed normal-row results if the workbook labels or examples indicate a total.
|
||||
|
||||
Keep scripts simple enough to run cleanly. Avoid unnecessary dynamic code generation and fragile f-strings with regex expressions inside them. Always execute the final `solution.py`; fix any syntax, indentation, or runtime error, then verify representative target cells.
|
||||
|
||||
If workbook cells contain arbitrary sample text that could be sensitive or trigger content filters, do not quote large raw cell contents in your response. Process them locally in Python with neutral variable names and output only the completed script/workbook changes.
|
||||
<!-- SLOW_UPDATE_END -->
|
||||
@@ -0,0 +1,100 @@
|
||||
# SkillOpt default configuration — base for all environments.
|
||||
# Environment configs should inherit via: _base_: default.yaml
|
||||
|
||||
model:
|
||||
backend: azure_openai
|
||||
optimizer: gpt-5.5
|
||||
target: gpt-5.5
|
||||
optimizer_backend: openai_chat
|
||||
target_backend: openai_chat
|
||||
reasoning_effort: medium
|
||||
rewrite_reasoning_effort: ""
|
||||
rewrite_max_completion_tokens: 64000
|
||||
codex_exec_path: codex
|
||||
codex_exec_sandbox: workspace-write
|
||||
codex_exec_profile: ""
|
||||
codex_exec_full_auto: false
|
||||
codex_exec_reasoning_effort: none
|
||||
codex_exec_use_sdk: auto
|
||||
codex_exec_network_access: false
|
||||
codex_exec_web_search: false
|
||||
codex_exec_approval_policy: never
|
||||
claude_code_exec_path: claude
|
||||
claude_code_exec_profile: ""
|
||||
claude_code_exec_use_sdk: auto
|
||||
claude_code_exec_effort: medium
|
||||
claude_code_exec_max_thinking_tokens: 16384
|
||||
codex_trace_to_optimizer: true
|
||||
azure_openai_endpoint: "" # e.g. "https://your-resource.openai.azure.com/"
|
||||
azure_openai_api_version: "2024-12-01-preview"
|
||||
azure_openai_api_key: "" # Fill locally if you do not export AZURE_OPENAI_API_KEY
|
||||
azure_openai_auth_mode: "" # empty → fall back to AZURE_OPENAI_AUTH_MODE env (default "azure_cli")
|
||||
azure_openai_ad_scope: "https://cognitiveservices.azure.com/.default"
|
||||
azure_openai_managed_identity_client_id: ""
|
||||
optimizer_azure_openai_endpoint: "" # e.g. "https://your-resource.openai.azure.com/"
|
||||
optimizer_azure_openai_api_version: "2024-12-01-preview"
|
||||
optimizer_azure_openai_api_key: ""
|
||||
optimizer_azure_openai_auth_mode: "" # empty → fall back to OPTIMIZER_AZURE_OPENAI_AUTH_MODE env, then shared
|
||||
optimizer_azure_openai_ad_scope: "https://cognitiveservices.azure.com/.default"
|
||||
optimizer_azure_openai_managed_identity_client_id: ""
|
||||
target_azure_openai_endpoint: "" # e.g. "https://your-resource.openai.azure.com/"
|
||||
target_azure_openai_api_version: "2024-12-01-preview"
|
||||
target_azure_openai_api_key: ""
|
||||
target_azure_openai_auth_mode: "" # empty → fall back to TARGET_AZURE_OPENAI_AUTH_MODE env, then shared
|
||||
target_azure_openai_ad_scope: "https://cognitiveservices.azure.com/.default"
|
||||
target_azure_openai_managed_identity_client_id: ""
|
||||
|
||||
# MiniMax backend settings (minimax_chat target)
|
||||
minimax_base_url: "" # https://api.minimax.io/v1 if blank
|
||||
minimax_api_key: ""
|
||||
minimax_model: "MiniMax-M2.7"
|
||||
minimax_temperature: "0.7"
|
||||
minimax_max_tokens: "8000"
|
||||
minimax_enable_thinking: "false"
|
||||
optimizer_minimax_base_url: "" # per-role override
|
||||
target_minimax_base_url: "" # per-role override
|
||||
optimizer_minimax_api_key: ""
|
||||
target_minimax_api_key: ""
|
||||
|
||||
train:
|
||||
num_epochs: 4
|
||||
train_size: 0 # 0 = derive from dataset split when available
|
||||
batch_size: 40
|
||||
accumulation: 1
|
||||
seed: 42
|
||||
|
||||
gradient:
|
||||
minibatch_size: 8
|
||||
merge_batch_size: 8
|
||||
analyst_workers: 16
|
||||
max_analyst_rounds: 3
|
||||
failure_only: false
|
||||
|
||||
optimizer:
|
||||
learning_rate: 4 # max edits per step (edit_budget)
|
||||
min_learning_rate: 2 # min edits for decay schedulers
|
||||
lr_scheduler: cosine # constant / linear / cosine / autonomous
|
||||
lr_control_mode: fixed # fixed / autonomous / none
|
||||
skill_update_mode: patch # patch / rewrite_from_suggestions / full_rewrite_minibatch
|
||||
use_slow_update: true
|
||||
slow_update_samples: 20
|
||||
slow_update_gate_with_selection: false
|
||||
longitudinal_pair_policy: mixed # mixed / changed / unchanged
|
||||
use_meta_skill: true
|
||||
|
||||
evaluation:
|
||||
use_gate: true
|
||||
sel_env_num: 0
|
||||
test_env_num: 0
|
||||
eval_test: true
|
||||
|
||||
env:
|
||||
name: ""
|
||||
skill_init: ""
|
||||
split_mode: ratio # ratio = build deterministic split from data_path; split_dir = use pre-split train/val/test
|
||||
split_seed: 42
|
||||
split_dir: ""
|
||||
data_path: ""
|
||||
split_output_dir: ""
|
||||
exec_timeout: 120 # per target model/code-agent call timeout in seconds
|
||||
out_root: ""
|
||||
@@ -0,0 +1,29 @@
|
||||
_base_: ../_base_/default.yaml
|
||||
|
||||
train:
|
||||
train_size: 0
|
||||
accumulation: 1
|
||||
|
||||
gradient:
|
||||
minibatch_size: 8
|
||||
merge_batch_size: 8
|
||||
|
||||
optimizer:
|
||||
learning_rate: 4
|
||||
|
||||
evaluation:
|
||||
sel_env_num: 0
|
||||
test_env_num: 0
|
||||
|
||||
env:
|
||||
name: alfworld
|
||||
skill_init: skillopt/envs/alfworld/skills/initial.md
|
||||
split_mode: split_dir
|
||||
split_dir: data/alfworld_path_split
|
||||
data_path: ""
|
||||
split_output_dir: ""
|
||||
max_steps: 50
|
||||
max_completion_tokens: 16384
|
||||
workers: 8
|
||||
max_api_workers: 8
|
||||
limit: 0
|
||||
@@ -0,0 +1,28 @@
|
||||
_base_: ../_base_/default.yaml
|
||||
|
||||
model:
|
||||
reasoning_effort: medium
|
||||
|
||||
train:
|
||||
batch_size: 40
|
||||
accumulation: 1
|
||||
|
||||
gradient:
|
||||
minibatch_size: 8
|
||||
merge_batch_size: 8
|
||||
|
||||
optimizer:
|
||||
learning_rate: 4
|
||||
|
||||
env:
|
||||
name: docvqa
|
||||
skill_init: skillopt/envs/docvqa/skills/initial.md
|
||||
split_mode: split_dir
|
||||
split_dir: data/docvqa/splits
|
||||
data_path: ""
|
||||
split_output_dir: ""
|
||||
max_turns: 1
|
||||
max_completion_tokens: 16384
|
||||
workers: 16
|
||||
image_detail: auto
|
||||
limit: 0
|
||||
@@ -0,0 +1,47 @@
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Feature: soft / mixed validation-gate metric (community-contributed, PR #25)
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
#
|
||||
# This is NOT a default SkillOpt setting and was NOT used to produce the
|
||||
# numbers reported in the paper. It is provided as a reference for users
|
||||
# who encounter a specific scenario where the default `hard` gate is too
|
||||
# coarse to drive training.
|
||||
#
|
||||
# When to consider this:
|
||||
# - You are running on a custom environment.
|
||||
# - Your held-out *selection* split has very few items (e.g. ≤ ~10).
|
||||
# - Your reward function is continuous / partial-credit (e.g. F1, BLEU,
|
||||
# soft match) rather than purely binary 0/1.
|
||||
#
|
||||
# Symptom this addresses:
|
||||
# With a small selection split + continuous rewards, candidate skills
|
||||
# often improve per-item soft scores (e.g. 0.06 → 0.26 on one item) but
|
||||
# never flip the discrete hard outcome. The default `hard` gate then
|
||||
# rejects every candidate and training stalls. Switching the gate to
|
||||
# `soft` or `mixed` lets these partial improvements be accepted.
|
||||
#
|
||||
# When NOT to use this:
|
||||
# - When reproducing the paper. The paper-reported numbers were obtained
|
||||
# under the default `hard` gate.
|
||||
# - When your selection split is large (dozens+ items) and / or your
|
||||
# reward is already binary — `hard` is the more conservative choice
|
||||
# and matches the design described in the paper.
|
||||
#
|
||||
# To use: inherit your env config from this file, e.g.
|
||||
# _base_: ../features/soft_gate.yaml
|
||||
# or copy the `evaluation:` block below into your config.
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
_base_: ../_base_/default.yaml
|
||||
|
||||
evaluation:
|
||||
# Three options:
|
||||
# 'hard' — default; exact-match accuracy. Use this to reproduce the paper.
|
||||
# 'soft' — per-item soft / partial-credit score (recommended for the
|
||||
# small-split + continuous-reward scenario described above).
|
||||
# 'mixed' — weighted average: (1 - w) * hard + w * soft, with `w` set by
|
||||
# `gate_mixed_weight` below.
|
||||
gate_metric: soft
|
||||
|
||||
# Only used when gate_metric == 'mixed'. Ignored otherwise.
|
||||
gate_mixed_weight: 0.5
|
||||
@@ -0,0 +1,22 @@
|
||||
_base_: ../_base_/default.yaml
|
||||
|
||||
train:
|
||||
train_size: 0
|
||||
batch_size: 40
|
||||
accumulation: 1
|
||||
|
||||
env:
|
||||
name: livemathematicianbench
|
||||
skill_init: skillopt/envs/livemathematicianbench/skills/initial.md
|
||||
split_mode: split_dir
|
||||
split_dir: data/livemathematicianbench_split
|
||||
data_path: ""
|
||||
split_output_dir: ""
|
||||
max_turns: 1
|
||||
max_completion_tokens: 16384
|
||||
exec_timeout: 300
|
||||
workers: 64
|
||||
limit: 0
|
||||
shuffle_choices: true
|
||||
use_theorem: false
|
||||
use_sketch: false
|
||||
@@ -0,0 +1,34 @@
|
||||
_base_: ../_base_/default.yaml
|
||||
|
||||
model:
|
||||
reasoning_effort: medium
|
||||
|
||||
train:
|
||||
batch_size: 40
|
||||
accumulation: 1
|
||||
|
||||
gradient:
|
||||
minibatch_size: 8
|
||||
merge_batch_size: 8
|
||||
|
||||
optimizer:
|
||||
learning_rate: 4
|
||||
|
||||
env:
|
||||
name: officeqa
|
||||
skill_init: skillopt/envs/officeqa/skills/initial.md
|
||||
split_mode: split_dir
|
||||
split_dir: data/officeqa_split
|
||||
data_dirs:
|
||||
- data/officeqa_docs_official
|
||||
workers: 4
|
||||
max_tool_turns: 24
|
||||
max_completion_tokens: 16384
|
||||
search_mode: offline
|
||||
max_queries_per_turn: 4
|
||||
search_api_url: http://apisix.westus2.cloudapp.azure.com/search_tool/search
|
||||
search_auth_env: OFFICEQA_CUSTOM_SEARCH_AUTH
|
||||
search_provider: duckduckgo
|
||||
search_max_num_results: 4
|
||||
search_timeout_seconds: 20
|
||||
limit: 0
|
||||
@@ -0,0 +1,32 @@
|
||||
_base_: ../_base_/default.yaml
|
||||
|
||||
model:
|
||||
reasoning_effort: medium
|
||||
|
||||
train:
|
||||
train_size: 400
|
||||
batch_size: 40
|
||||
accumulation: 1
|
||||
|
||||
gradient:
|
||||
minibatch_size: 8
|
||||
merge_batch_size: 8
|
||||
|
||||
optimizer:
|
||||
learning_rate: 4
|
||||
|
||||
evaluation:
|
||||
sel_env_num: 0
|
||||
test_env_num: 0
|
||||
|
||||
env:
|
||||
name: searchqa
|
||||
skill_init: skillopt/envs/searchqa/skills/initial.md
|
||||
split_mode: split_dir
|
||||
split_dir: data/searchqa_split
|
||||
data_path: ""
|
||||
split_output_dir: ""
|
||||
max_turns: 1
|
||||
max_completion_tokens: 16384
|
||||
workers: 24
|
||||
limit: 0
|
||||
@@ -0,0 +1,34 @@
|
||||
_base_: ../_base_/default.yaml
|
||||
|
||||
model:
|
||||
reasoning_effort: medium
|
||||
|
||||
train:
|
||||
train_size: 80
|
||||
batch_size: 40
|
||||
accumulation: 1
|
||||
|
||||
gradient:
|
||||
minibatch_size: 8
|
||||
merge_batch_size: 8
|
||||
|
||||
optimizer:
|
||||
learning_rate: 4
|
||||
|
||||
evaluation:
|
||||
sel_env_num: 0
|
||||
test_env_num: 0
|
||||
|
||||
env:
|
||||
name: spreadsheetbench
|
||||
skill_init: skillopt/envs/spreadsheetbench/skills/initial.md
|
||||
split_mode: split_dir
|
||||
split_dir: data/spreadsheetbench_split
|
||||
data_path: ""
|
||||
split_output_dir: ""
|
||||
data_root: data/spreadsheetbench_verified_400
|
||||
mode: multi
|
||||
max_turns: 30
|
||||
max_completion_tokens: 16384
|
||||
exec_timeout: 600
|
||||
workers: 24
|
||||
+223
@@ -0,0 +1,223 @@
|
||||
# Data Manifests
|
||||
|
||||
This directory releases lightweight split manifests for the SkillOpt paper
|
||||
splits. These manifests are not full runnable benchmark payloads. To evaluate a
|
||||
benchmark, first materialize the full examples from the raw data source when
|
||||
needed, then point `--split_dir` at the split directory listed below.
|
||||
|
||||
In this README, "coverage" describes which part of the upstream benchmark the
|
||||
manifest references. It does not mean the released manifest directory contains
|
||||
the full runnable examples.
|
||||
|
||||
## Layout
|
||||
|
||||
Every released manifest directory uses the same file layout:
|
||||
|
||||
```text
|
||||
data/<benchmark>_<manifest_type>/
|
||||
|-- split_manifest.json
|
||||
|-- train/items.json
|
||||
|-- val/items.json
|
||||
`-- test/items.json
|
||||
```
|
||||
|
||||
`split_manifest.json` records source metadata, split counts, and item fields.
|
||||
Each `items.json` contains only stable IDs or source-path hints.
|
||||
|
||||
## Released Splits
|
||||
|
||||
| Manifest directory | Benchmark | Counts | Coverage | Raw data source | `split_dir` |
|
||||
|---|---|---:|---|---|---|
|
||||
| `searchqa_id_split/` | SearchQA | 400 / 200 / 1400 | Official HF dataset IDs | [lucadiliello/searchqa](https://huggingface.co/datasets/lucadiliello/searchqa) | `data/searchqa_split` |
|
||||
| `livemathematicianbench_id_split/` | LiveMathematicianBench | 35 / 18 / 124 | Four official monthly files | [LiveMathematicianBench/LiveMathematicianBench](https://huggingface.co/datasets/LiveMathematicianBench/LiveMathematicianBench) | `data/livemathematicianbench_split` |
|
||||
| `docvqa_id_split/` | DocVQA | 107 / 53 / 374 | 10% subset of validation | [lmms-lab/DocVQA](https://huggingface.co/datasets/lmms-lab/DocVQA) | `data/docvqa/splits` |
|
||||
| `officeqa_id_split/` | OfficeQA | 50 / 24 / 172 | OfficeQA Full | [databricks/officeqa](https://huggingface.co/datasets/databricks/officeqa) | `data/officeqa_split` |
|
||||
| `spreadsheetbench_id_split/` | SpreadsheetBench | 80 / 40 / 280 | SpreadsheetBench Verified 400 | [KAKA22/SpreadsheetBench](https://huggingface.co/datasets/KAKA22/SpreadsheetBench) | `data/spreadsheetbench_split` |
|
||||
| `alfworld_path_split/` | ALFWorld | 39 / 18 / 134 | ALFWorld `json_2.1.1` paths | [alfworld/alfworld](https://github.com/alfworld/alfworld) | `data/alfworld_path_split` |
|
||||
|
||||
Counts are ordered as train / val / test.
|
||||
|
||||
## Direct Use
|
||||
|
||||
Only `alfworld_path_split/` can be used directly as `--split_dir` from this
|
||||
release, because the ALFWorld loader reads `gamefile` and `task_type` from the
|
||||
split items.
|
||||
|
||||
This does not mean the ALFWorld raw data is included. You still need to
|
||||
download ALFWorld separately with `alfworld-download` and set `$ALFWORLD_DATA`
|
||||
to the data root containing `json_2.1.1`.
|
||||
|
||||
The other manifest directories are lookup manifests. They intentionally omit
|
||||
full example fields such as questions, answers, contexts, images, or task
|
||||
instructions. Materialize those benchmarks into the `split_dir` paths listed
|
||||
above before running SkillOpt.
|
||||
|
||||
## Lookup Keys
|
||||
|
||||
The manifests are sufficient to locate the corresponding raw examples after
|
||||
the raw data has been downloaded or otherwise made available:
|
||||
|
||||
| Benchmark | Manifest lookup key |
|
||||
|---|---|
|
||||
| SearchQA | Match `items.json[].id` to the `key` field in `lucadiliello/searchqa`. |
|
||||
| LiveMathematicianBench | Open `source_file`, then match `no`; the manifest `id` is `<month>:<no>`. |
|
||||
| DocVQA | Match `questionId` within the official DocVQA `validation` split; `image_path` records the expected local image path. |
|
||||
| OfficeQA | Match `uid` in `officeqa_full.csv`; `source_files` and `source_docs` identify the supporting document. |
|
||||
| SpreadsheetBench | Match `id`; `spreadsheet_path` identifies the referenced spreadsheet directory. |
|
||||
| ALFWorld | Resolve `gamefile` relative to `$ALFWORLD_DATA`. |
|
||||
|
||||
## Manifest Item Examples
|
||||
|
||||
SearchQA:
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "221c83e6630f4e7983da48fa28da1882"
|
||||
}
|
||||
```
|
||||
|
||||
LiveMathematicianBench:
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "202602:22",
|
||||
"month": "202602",
|
||||
"no": 22,
|
||||
"paper_link": "http://arxiv.org/abs/2602.10700v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
}
|
||||
```
|
||||
|
||||
DocVQA:
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "50877",
|
||||
"questionId": "50877",
|
||||
"docId": "14724",
|
||||
"image_path": "data/docvqa_images/q50877_d14724.png",
|
||||
"source_split": "validation"
|
||||
}
|
||||
```
|
||||
|
||||
OfficeQA:
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "UID0002",
|
||||
"uid": "UID0002",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1944_01.txt"
|
||||
}
|
||||
```
|
||||
|
||||
SpreadsheetBench:
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "32438",
|
||||
"spreadsheet_path": "spreadsheet/32438",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
}
|
||||
```
|
||||
|
||||
ALFWorld:
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "train:0000",
|
||||
"gamefile": "json_2.1.1/train/.../game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
}
|
||||
```
|
||||
|
||||
## Benchmark Notes
|
||||
|
||||
### SearchQA
|
||||
|
||||
`searchqa_id_split/` is an ID-only manifest. Each released `id` exactly matches
|
||||
the `key` field in `lucadiliello/searchqa`.
|
||||
|
||||
Materialized examples must include the fields consumed by the SearchQA
|
||||
environment, including:
|
||||
|
||||
```text
|
||||
question
|
||||
context
|
||||
answers
|
||||
```
|
||||
|
||||
### LiveMathematicianBench
|
||||
|
||||
`livemathematicianbench_id_split/` was generated from these raw files:
|
||||
|
||||
```text
|
||||
data/202511/qa_202511_final.json
|
||||
data/202512/qa_202512_final.json
|
||||
data/202601/qa_202601_final.json
|
||||
data/202602/qa_202602_final.json
|
||||
```
|
||||
|
||||
The manifest stores IDs in the loader format:
|
||||
|
||||
```text
|
||||
<month>:<no>
|
||||
```
|
||||
|
||||
Materialized examples must include:
|
||||
|
||||
```text
|
||||
question
|
||||
choices
|
||||
correct_choice
|
||||
theorem_type
|
||||
theorem
|
||||
sketch
|
||||
paper_link
|
||||
```
|
||||
|
||||
### DocVQA
|
||||
|
||||
`docvqa_id_split/` records `docvqa_validation_10pct`: a 10% subset sampled from
|
||||
the official DocVQA `validation` split.
|
||||
|
||||
```text
|
||||
source_split: validation
|
||||
docvqa_validation_10pct: train=107, val=53, test=374
|
||||
```
|
||||
|
||||
Each manifest item contains question/document IDs plus image location metadata.
|
||||
Materialized examples must provide `question`, `answer` or `ground_truth`, and
|
||||
an `image_path` that resolves locally.
|
||||
|
||||
### OfficeQA
|
||||
|
||||
`officeqa_id_split/` records the split over OfficeQA Full
|
||||
(`officeqa_full.csv`). The official OfficeQA CSVs are gated on Hugging Face, so
|
||||
materialization requires authorized access.
|
||||
|
||||
Each manifest item contains `uid`, `category`, `source_files`, and
|
||||
`source_docs` hints. Materialized examples must include `question` and
|
||||
`ground_truth` or `answer`.
|
||||
|
||||
### SpreadsheetBench
|
||||
|
||||
`spreadsheetbench_id_split/` records the split over SpreadsheetBench Verified
|
||||
400, from `spreadsheetbench_verified_400.tar.gz`.
|
||||
|
||||
Each manifest item contains task identity metadata such as `id`,
|
||||
`spreadsheet_path`, and `instruction_type`. Materialization must also place the
|
||||
referenced spreadsheet directories at:
|
||||
|
||||
```text
|
||||
data/spreadsheetbench_verified_400
|
||||
```
|
||||
|
||||
### ALFWorld
|
||||
|
||||
`alfworld_path_split/` records `gamefile` paths relative to `$ALFWORLD_DATA`.
|
||||
The source payload is `json_2.1.1`, which must be downloaded separately with
|
||||
`alfworld-download`.
|
||||
|
||||
This manifest can be used directly as `--split_dir` after `$ALFWORLD_DATA`
|
||||
points to the local ALFWorld data root containing `json_2.1.1`.
|
||||
@@ -0,0 +1,29 @@
|
||||
{
|
||||
"benchmark": "ALFWorld",
|
||||
"manifest_type": "path_split",
|
||||
"source_repo": "alfworld/alfworld",
|
||||
"source_repo_type": "repository",
|
||||
"source_url": "https://github.com/alfworld/alfworld",
|
||||
"source_file": "json_2.1.1",
|
||||
"source_method": "generated by alfworld-download",
|
||||
"source_split_files": [
|
||||
"split_train.json",
|
||||
"split_val.json",
|
||||
"split_test.json"
|
||||
],
|
||||
"counts": {
|
||||
"train": 39,
|
||||
"val": 18,
|
||||
"test": 134
|
||||
},
|
||||
"item_fields": [
|
||||
"id",
|
||||
"gamefile",
|
||||
"task_type"
|
||||
],
|
||||
"path_root": "$ALFWORLD_DATA",
|
||||
"notes": [
|
||||
"This is a path manifest, not the ALFWorld game payload.",
|
||||
"The gamefile field is relative to ALFWORLD_DATA and must be expanded before direct use as split_dir data."
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,672 @@
|
||||
[
|
||||
{
|
||||
"id": "test:0000",
|
||||
"gamefile": "json_2.1.1/valid_unseen/look_at_obj_in_light-AlarmClock-None-DeskLamp-308/trial_T20190908_222917_366542/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "test:0001",
|
||||
"gamefile": "json_2.1.1/valid_unseen/look_at_obj_in_light-AlarmClock-None-DeskLamp-308/trial_T20190908_222933_607649/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "test:0002",
|
||||
"gamefile": "json_2.1.1/valid_unseen/look_at_obj_in_light-AlarmClock-None-DeskLamp-308/trial_T20190908_222951_616606/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "test:0003",
|
||||
"gamefile": "json_2.1.1/valid_unseen/look_at_obj_in_light-Book-None-DeskLamp-308/trial_T20190908_020029_636862/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "test:0004",
|
||||
"gamefile": "json_2.1.1/valid_unseen/look_at_obj_in_light-Book-None-DeskLamp-308/trial_T20190908_020048_814402/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "test:0005",
|
||||
"gamefile": "json_2.1.1/valid_unseen/look_at_obj_in_light-Book-None-DeskLamp-308/trial_T20190908_144951_587345/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "test:0006",
|
||||
"gamefile": "json_2.1.1/valid_unseen/look_at_obj_in_light-Bowl-None-DeskLamp-308/trial_T20190907_133919_856963/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "test:0007",
|
||||
"gamefile": "json_2.1.1/valid_unseen/look_at_obj_in_light-Bowl-None-DeskLamp-308/trial_T20190907_133935_066606/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "test:0008",
|
||||
"gamefile": "json_2.1.1/valid_unseen/look_at_obj_in_light-Bowl-None-DeskLamp-308/trial_T20190907_133953_562557/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "test:0009",
|
||||
"gamefile": "json_2.1.1/valid_unseen/look_at_obj_in_light-CD-None-DeskLamp-308/trial_T20190908_141942_810052/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "test:0010",
|
||||
"gamefile": "json_2.1.1/valid_unseen/look_at_obj_in_light-CD-None-DeskLamp-308/trial_T20190908_141958_463362/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "test:0011",
|
||||
"gamefile": "json_2.1.1/valid_unseen/look_at_obj_in_light-CD-None-DeskLamp-308/trial_T20190908_142046_281296/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "test:0012",
|
||||
"gamefile": "json_2.1.1/valid_unseen/look_at_obj_in_light-Mug-None-DeskLamp-308/trial_T20190908_161733_213242/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "test:0013",
|
||||
"gamefile": "json_2.1.1/valid_unseen/look_at_obj_in_light-Mug-None-DeskLamp-308/trial_T20190908_201421_021646/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "test:0014",
|
||||
"gamefile": "json_2.1.1/valid_unseen/look_at_obj_in_light-Mug-None-DeskLamp-308/trial_T20190908_201444_037645/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "test:0015",
|
||||
"gamefile": "json_2.1.1/valid_unseen/look_at_obj_in_light-Pencil-None-DeskLamp-308/trial_T20190908_220545_153480/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "test:0016",
|
||||
"gamefile": "json_2.1.1/valid_unseen/look_at_obj_in_light-Pencil-None-DeskLamp-308/trial_T20190908_220604_010430/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "test:0017",
|
||||
"gamefile": "json_2.1.1/valid_unseen/look_at_obj_in_light-Pencil-None-DeskLamp-308/trial_T20190908_220656_510400/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "test:0018",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-Mug-None-Desk-308/trial_T20190908_125200_737896/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0019",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-Mug-None-Desk-308/trial_T20190909_203041_433487/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0020",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-Mug-None-Desk-308/trial_T20190909_210238_431966/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0021",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-Pencil-None-Shelf-308/trial_T20190908_121952_610012/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0022",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-Pencil-None-Shelf-308/trial_T20190908_122024_052056/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0023",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-Pencil-None-Shelf-308/trial_T20190908_122154_042763/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0024",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-PepperShaker-None-Drawer-10/trial_T20190906_184021_215264/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0025",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-PepperShaker-None-Drawer-10/trial_T20190918_154326_823501/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0026",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-PepperShaker-None-Drawer-10/trial_T20190918_154424_844749/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0027",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-SaltShaker-None-Cabinet-10/trial_T20190906_191429_743650/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0028",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-SaltShaker-None-Cabinet-10/trial_T20190906_191445_723170/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0029",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-SaltShaker-None-Cabinet-10/trial_T20190906_191501_563086/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0030",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-SaltShaker-None-Drawer-10/trial_T20190909_021613_077537/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0031",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-SaltShaker-None-Drawer-10/trial_T20190909_021650_880235/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0032",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-SaltShaker-None-Drawer-10/trial_T20190909_021728_339782/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0033",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-SoapBottle-None-Toilet-424/trial_T20190907_004321_405868/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0034",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-SoapBottle-None-Toilet-424/trial_T20190907_004351_281384/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0035",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-SoapBottle-None-Toilet-424/trial_T20190907_004404_604165/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0036",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-Vase-None-Safe-219/trial_T20190908_205204_244321/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0037",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-Vase-None-Safe-219/trial_T20190908_205221_748352/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0038",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-Vase-None-Safe-219/trial_T20190908_205246_776817/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0039",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-Watch-None-Safe-219/trial_T20190907_074524_006355/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0040",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-Watch-None-Safe-219/trial_T20190907_074556_124850/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0041",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_and_place_simple-Watch-None-Safe-219/trial_T20190907_074643_810052/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "test:0042",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Bowl-None-Cabinet-10/trial_T20190909_061130_844814/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0043",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Bowl-None-Cabinet-10/trial_T20190909_061158_110530/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0044",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Bowl-None-Cabinet-10/trial_T20190909_061232_368489/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0045",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Cloth-None-Cabinet-424/trial_T20190908_022321_380927/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0046",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Cloth-None-Cabinet-424/trial_T20190908_022436_073995/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0047",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Cloth-None-CounterTop-424/trial_T20190908_100632_546757/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0048",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Cloth-None-CounterTop-424/trial_T20190908_114340_674467/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0049",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Egg-None-Microwave-10/trial_T20190909_120554_888709/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0050",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Egg-None-Microwave-10/trial_T20190909_120632_691361/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0051",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Egg-None-Microwave-10/trial_T20190909_120712_273910/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0052",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Knife-None-CounterTop-10/trial_T20190909_110347_624008/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0053",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Knife-None-CounterTop-10/trial_T20190909_110445_675754/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0054",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Knife-None-CounterTop-10/trial_T20190909_110531_148235/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0055",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Mug-None-CoffeeMachine-10/trial_T20190907_221208_560499/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0056",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Mug-None-CoffeeMachine-10/trial_T20190907_221300_362511/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0057",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Mug-None-CoffeeMachine-10/trial_T20190907_221355_558505/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0058",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Pan-None-CounterTop-10/trial_T20190908_032434_013084/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0059",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Pan-None-CounterTop-10/trial_T20190908_032518_891433/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0060",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Pan-None-CounterTop-10/trial_T20190908_032543_712058/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0061",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Plate-None-CounterTop-10/trial_T20190908_213356_017769/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0062",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Plate-None-CounterTop-10/trial_T20190908_213420_728917/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0063",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Plate-None-CounterTop-10/trial_T20190908_213533_897289/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0064",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-SoapBar-None-Cabinet-424/trial_T20190908_214926_337906/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0065",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-SoapBar-None-Cabinet-424/trial_T20190908_214946_567644/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0066",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-SoapBar-None-Cabinet-424/trial_T20190908_215019_162873/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0067",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-SoapBar-None-CounterTop-424/trial_T20190907_074045_109439/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0068",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-SoapBar-None-CounterTop-424/trial_T20190907_074106_050405/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0069",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-SoapBar-None-CounterTop-424/trial_T20190907_074124_966890/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0070",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Spatula-None-Drawer-10/trial_T20190907_080730_211959/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0071",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Spatula-None-Drawer-10/trial_T20190907_080800_275989/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0072",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_clean_then_place_in_recep-Spatula-None-Drawer-10/trial_T20190907_080825_222432/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0073",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_cool_then_place_in_recep-Bread-None-CounterTop-10/trial_T20190908_091747_866951/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0074",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_cool_then_place_in_recep-Bread-None-CounterTop-10/trial_T20190908_091811_414150/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0075",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_cool_then_place_in_recep-Bread-None-CounterTop-10/trial_T20190908_091835_825830/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0076",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_cool_then_place_in_recep-Lettuce-None-CounterTop-10/trial_T20190909_123133_763972/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0077",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_cool_then_place_in_recep-Lettuce-None-CounterTop-10/trial_T20190909_174807_646433/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0078",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_cool_then_place_in_recep-Lettuce-None-CounterTop-10/trial_T20190909_174840_771703/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0079",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_cool_then_place_in_recep-Mug-None-Cabinet-10/trial_T20190909_121559_082363/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0080",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_cool_then_place_in_recep-Mug-None-Cabinet-10/trial_T20190909_121635_622676/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0081",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_cool_then_place_in_recep-Mug-None-Cabinet-10/trial_T20190909_121710_650938/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0082",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_cool_then_place_in_recep-Mug-None-CoffeeMachine-10/trial_T20190907_183715_299073/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0083",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_cool_then_place_in_recep-Mug-None-CoffeeMachine-10/trial_T20190907_183807_477267/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0084",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_cool_then_place_in_recep-Mug-None-CoffeeMachine-10/trial_T20190907_183853_958104/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0085",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_cool_then_place_in_recep-Pan-None-CounterTop-10/trial_T20190908_114545_244903/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0086",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_cool_then_place_in_recep-Pan-None-CounterTop-10/trial_T20190908_114622_738670/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0087",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_cool_then_place_in_recep-Pan-None-CounterTop-10/trial_T20190908_114656_768805/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0088",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_cool_then_place_in_recep-Potato-None-Microwave-10/trial_T20190907_033157_424297/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0089",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_cool_then_place_in_recep-Potato-None-Microwave-10/trial_T20190907_033228_194678/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0090",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_cool_then_place_in_recep-Potato-None-Microwave-10/trial_T20190907_033306_962974/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0091",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_cool_then_place_in_recep-Tomato-None-Microwave-10/trial_T20190909_102608_318800/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0092",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_cool_then_place_in_recep-Tomato-None-Microwave-10/trial_T20190909_102644_926781/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0093",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_cool_then_place_in_recep-Tomato-None-Microwave-10/trial_T20190909_102710_795182/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0094",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Apple-None-Fridge-10/trial_T20190906_182259_116320/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0095",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Apple-None-Fridge-10/trial_T20190906_182353_418140/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0096",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Apple-None-Fridge-10/trial_T20190906_182435_622538/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0097",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Apple-None-GarbageCan-10/trial_T20190908_145050_918567/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0098",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Apple-None-GarbageCan-10/trial_T20190908_145143_820541/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0099",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Apple-None-GarbageCan-10/trial_T20190908_145356_918528/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0100",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Cup-None-Cabinet-10/trial_T20190907_083346_800823/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0101",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Cup-None-Cabinet-10/trial_T20190907_083429_887065/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0102",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Cup-None-Cabinet-10/trial_T20190907_083507_594820/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0103",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Egg-None-GarbageCan-10/trial_T20190908_113432_673307/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0104",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Egg-None-GarbageCan-10/trial_T20190908_113523_123938/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0105",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Egg-None-GarbageCan-10/trial_T20190908_113610_425142/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0106",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Mug-None-Cabinet-10/trial_T20190909_021100_341887/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0107",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Mug-None-Cabinet-10/trial_T20190909_021200_669381/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0108",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Mug-None-Cabinet-10/trial_T20190909_021247_306737/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0109",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Mug-None-CoffeeMachine-10/trial_T20190907_171806_406231/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0110",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Mug-None-CoffeeMachine-10/trial_T20190907_171850_960211/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0111",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Mug-None-CoffeeMachine-10/trial_T20190907_171933_349922/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0112",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Potato-None-GarbageCan-10/trial_T20190907_161745_664033/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0113",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Potato-None-GarbageCan-10/trial_T20190907_161853_945788/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0114",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Tomato-None-GarbageCan-10/trial_T20190908_225046_020282/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0115",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Tomato-None-GarbageCan-10/trial_T20190908_225359_617900/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0116",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_heat_then_place_in_recep-Tomato-None-GarbageCan-10/trial_T20190908_225453_272533/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "test:0117",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_two_obj_and_place-CD-None-Safe-308/trial_T20190907_050942_897916/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
},
|
||||
{
|
||||
"id": "test:0118",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_two_obj_and_place-CD-None-Safe-308/trial_T20190907_051013_060265/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
},
|
||||
{
|
||||
"id": "test:0119",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_two_obj_and_place-CD-None-Safe-308/trial_T20190907_051056_585414/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
},
|
||||
{
|
||||
"id": "test:0120",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_two_obj_and_place-KeyChain-None-Safe-219/trial_T20190909_011803_423115/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
},
|
||||
{
|
||||
"id": "test:0121",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_two_obj_and_place-KeyChain-None-Safe-219/trial_T20190909_012027_782483/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
},
|
||||
{
|
||||
"id": "test:0122",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_two_obj_and_place-PepperShaker-None-Drawer-10/trial_T20190908_010306_215435/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
},
|
||||
{
|
||||
"id": "test:0123",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_two_obj_and_place-PepperShaker-None-Drawer-10/trial_T20190912_221016_460197/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
},
|
||||
{
|
||||
"id": "test:0124",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_two_obj_and_place-PepperShaker-None-Drawer-10/trial_T20190912_221141_608117/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
},
|
||||
{
|
||||
"id": "test:0125",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_two_obj_and_place-Pillow-None-Sofa-219/trial_T20190907_163240_345855/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
},
|
||||
{
|
||||
"id": "test:0126",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_two_obj_and_place-Pillow-None-Sofa-219/trial_T20190907_163327_486300/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
},
|
||||
{
|
||||
"id": "test:0127",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_two_obj_and_place-Pillow-None-Sofa-219/trial_T20190907_163408_914117/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
},
|
||||
{
|
||||
"id": "test:0128",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_two_obj_and_place-SoapBar-None-Cabinet-424/trial_T20190909_081720_491733/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
},
|
||||
{
|
||||
"id": "test:0129",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_two_obj_and_place-SoapBar-None-Cabinet-424/trial_T20190909_081746_857594/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
},
|
||||
{
|
||||
"id": "test:0130",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_two_obj_and_place-SoapBar-None-GarbageCan-424/trial_T20190909_064053_839817/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
},
|
||||
{
|
||||
"id": "test:0131",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_two_obj_and_place-SoapBar-None-GarbageCan-424/trial_T20190909_064221_368939/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
},
|
||||
{
|
||||
"id": "test:0132",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_two_obj_and_place-SoapBar-None-GarbageCan-424/trial_T20190909_064309_357168/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
},
|
||||
{
|
||||
"id": "test:0133",
|
||||
"gamefile": "json_2.1.1/valid_unseen/pick_two_obj_and_place-ToiletPaper-None-Cabinet-424/trial_T20190906_202926_527010/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
}
|
||||
]
|
||||
@@ -0,0 +1,197 @@
|
||||
[
|
||||
{
|
||||
"id": "train:0000",
|
||||
"gamefile": "json_2.1.1/train/look_at_obj_in_light-AlarmClock-None-DeskLamp-305/trial_T20190908_082736_108723/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "train:0001",
|
||||
"gamefile": "json_2.1.1/train/look_at_obj_in_light-CD-None-DeskLamp-304/trial_T20190907_185649_782438/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "train:0002",
|
||||
"gamefile": "json_2.1.1/train/look_at_obj_in_light-CD-None-DeskLamp-320/trial_T20190907_224439_174735/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "train:0003",
|
||||
"gamefile": "json_2.1.1/train/look_at_obj_in_light-Pillow-None-DeskLamp-316/trial_T20190908_232421_645610/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "train:0004",
|
||||
"gamefile": "json_2.1.1/train/look_at_obj_in_light-Statue-None-DeskLamp-319/trial_T20190907_035546_167548/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "train:0005",
|
||||
"gamefile": "json_2.1.1/train/pick_and_place_simple-CellPhone-None-Shelf-313/trial_T20190908_123725_452958/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "train:0006",
|
||||
"gamefile": "json_2.1.1/train/pick_and_place_simple-Newspaper-None-Sofa-211/trial_T20190906_175004_203092/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "train:0007",
|
||||
"gamefile": "json_2.1.1/train/pick_and_place_simple-Pencil-None-Desk-302/trial_T20190908_032836_462632/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "train:0008",
|
||||
"gamefile": "json_2.1.1/train/pick_and_place_simple-SoapBar-None-GarbageCan-416/trial_T20190908_020839_714699/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "train:0009",
|
||||
"gamefile": "json_2.1.1/train/pick_and_place_simple-Statue-None-CoffeeTable-222/trial_T20190907_131249_788749/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "train:0010",
|
||||
"gamefile": "json_2.1.1/train/pick_and_place_simple-ToiletPaper-None-ToiletPaperHanger-406/trial_T20190908_122807_136741/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "train:0011",
|
||||
"gamefile": "json_2.1.1/train/pick_and_place_simple-ToiletPaper-None-ToiletPaperHanger-415/trial_T20190908_050443_333939/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "train:0012",
|
||||
"gamefile": "json_2.1.1/train/pick_clean_then_place_in_recep-Apple-None-DiningTable-4/trial_T20190908_104413_450768/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0013",
|
||||
"gamefile": "json_2.1.1/train/pick_clean_then_place_in_recep-DishSponge-None-Shelf-20/trial_T20190907_222429_992578/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0014",
|
||||
"gamefile": "json_2.1.1/train/pick_clean_then_place_in_recep-DishSponge-None-Shelf-401/trial_T20190908_072225_397518/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0015",
|
||||
"gamefile": "json_2.1.1/train/pick_clean_then_place_in_recep-Kettle-None-Cabinet-2/trial_T20190909_043103_418752/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0016",
|
||||
"gamefile": "json_2.1.1/train/pick_clean_then_place_in_recep-Knife-None-Drawer-22/trial_T20190907_224827_746945/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0017",
|
||||
"gamefile": "json_2.1.1/train/pick_clean_then_place_in_recep-Lettuce-None-DiningTable-20/trial_T20190906_191148_519826/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0018",
|
||||
"gamefile": "json_2.1.1/train/pick_clean_then_place_in_recep-Lettuce-None-Fridge-13/trial_T20190908_203022_601787/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0019",
|
||||
"gamefile": "json_2.1.1/train/pick_clean_then_place_in_recep-Plate-None-Fridge-5/trial_T20190909_112954_869911/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0020",
|
||||
"gamefile": "json_2.1.1/train/pick_clean_then_place_in_recep-Spoon-None-DiningTable-18/trial_T20190909_102159_277894/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0021",
|
||||
"gamefile": "json_2.1.1/train/pick_cool_then_place_in_recep-Bread-None-CounterTop-1/trial_T20190908_212439_711334/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0022",
|
||||
"gamefile": "json_2.1.1/train/pick_cool_then_place_in_recep-Bread-None-CounterTop-15/trial_T20190909_085448_256298/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0023",
|
||||
"gamefile": "json_2.1.1/train/pick_cool_then_place_in_recep-Bread-None-CounterTop-16/trial_T20190908_143948_082471/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0024",
|
||||
"gamefile": "json_2.1.1/train/pick_cool_then_place_in_recep-Pan-None-StoveBurner-27/trial_T20190906_212619_469871/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0025",
|
||||
"gamefile": "json_2.1.1/train/pick_cool_then_place_in_recep-Plate-None-DiningTable-17/trial_T20190909_122939_032098/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0026",
|
||||
"gamefile": "json_2.1.1/train/pick_cool_then_place_in_recep-Pot-None-CounterTop-1/trial_T20190909_124252_504581/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0027",
|
||||
"gamefile": "json_2.1.1/train/pick_heat_then_place_in_recep-Apple-None-Fridge-20/trial_T20190908_013911_274341/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0028",
|
||||
"gamefile": "json_2.1.1/train/pick_heat_then_place_in_recep-Egg-None-CounterTop-12/trial_T20190908_215527_416490/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0029",
|
||||
"gamefile": "json_2.1.1/train/pick_heat_then_place_in_recep-Mug-None-CoffeeMachine-1/trial_T20190907_222924_821086/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0030",
|
||||
"gamefile": "json_2.1.1/train/pick_heat_then_place_in_recep-Mug-None-CoffeeMachine-28/trial_T20190908_062730_537428/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0031",
|
||||
"gamefile": "json_2.1.1/train/pick_heat_then_place_in_recep-Plate-None-Cabinet-13/trial_T20190907_062749_759882/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0032",
|
||||
"gamefile": "json_2.1.1/train/pick_heat_then_place_in_recep-Potato-None-Fridge-2/trial_T20190909_030845_198194/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0033",
|
||||
"gamefile": "json_2.1.1/train/pick_heat_then_place_in_recep-Tomato-None-CounterTop-26/trial_T20190907_005525_499114/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "train:0034",
|
||||
"gamefile": "json_2.1.1/train/pick_two_obj_and_place-CD-None-Drawer-319/trial_T20190907_145515_348252/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
},
|
||||
{
|
||||
"id": "train:0035",
|
||||
"gamefile": "json_2.1.1/train/pick_two_obj_and_place-Candle-None-Drawer-427/trial_T20190909_043917_251333/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
},
|
||||
{
|
||||
"id": "train:0036",
|
||||
"gamefile": "json_2.1.1/train/pick_two_obj_and_place-KeyChain-None-ArmChair-222/trial_T20190909_100312_677332/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
},
|
||||
{
|
||||
"id": "train:0037",
|
||||
"gamefile": "json_2.1.1/train/pick_two_obj_and_place-Newspaper-None-Sofa-212/trial_T20190908_112632_208041/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
},
|
||||
{
|
||||
"id": "train:0038",
|
||||
"gamefile": "json_2.1.1/train/pick_two_obj_and_place-SaltShaker-None-SideTable-21/trial_T20190909_041626_844806/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
}
|
||||
]
|
||||
@@ -0,0 +1,92 @@
|
||||
[
|
||||
{
|
||||
"id": "val:0000",
|
||||
"gamefile": "json_2.1.1/valid_seen/look_at_obj_in_light-AlarmClock-None-DeskLamp-323/trial_T20190909_044715_250790/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "val:0001",
|
||||
"gamefile": "json_2.1.1/valid_seen/look_at_obj_in_light-Bowl-None-DeskLamp-301/trial_T20190909_150719_492274/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "val:0002",
|
||||
"gamefile": "json_2.1.1/valid_seen/look_at_obj_in_light-Pillow-None-DeskLamp-323/trial_T20190908_053153_077977/game.tw-pddl",
|
||||
"task_type": "look_at_obj_in_light"
|
||||
},
|
||||
{
|
||||
"id": "val:0003",
|
||||
"gamefile": "json_2.1.1/valid_seen/pick_and_place_simple-Mug-None-SideTable-329/trial_T20190909_032318_169393/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "val:0004",
|
||||
"gamefile": "json_2.1.1/valid_seen/pick_and_place_simple-Mug-None-SideTable-329/trial_T20190909_032340_274147/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "val:0005",
|
||||
"gamefile": "json_2.1.1/valid_seen/pick_and_place_simple-Pencil-None-Desk-310/trial_T20190909_113054_894334/game.tw-pddl",
|
||||
"task_type": "pick_and_place_simple"
|
||||
},
|
||||
{
|
||||
"id": "val:0006",
|
||||
"gamefile": "json_2.1.1/valid_seen/pick_clean_then_place_in_recep-ButterKnife-None-Drawer-30/trial_T20190908_052007_212776/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "val:0007",
|
||||
"gamefile": "json_2.1.1/valid_seen/pick_clean_then_place_in_recep-ButterKnife-None-Drawer-8/trial_T20190909_124425_112757/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "val:0008",
|
||||
"gamefile": "json_2.1.1/valid_seen/pick_clean_then_place_in_recep-SoapBar-None-Cabinet-402/trial_T20190908_055221_984342/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "val:0009",
|
||||
"gamefile": "json_2.1.1/valid_seen/pick_clean_then_place_in_recep-SoapBar-None-Toilet-410/trial_T20190906_201106_979461/game.tw-pddl",
|
||||
"task_type": "pick_clean_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "val:0010",
|
||||
"gamefile": "json_2.1.1/valid_seen/pick_cool_then_place_in_recep-Apple-None-Microwave-19/trial_T20190906_210937_878489/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "val:0011",
|
||||
"gamefile": "json_2.1.1/valid_seen/pick_cool_then_place_in_recep-Plate-None-CounterTop-1/trial_T20190906_205324_559361/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "val:0012",
|
||||
"gamefile": "json_2.1.1/valid_seen/pick_cool_then_place_in_recep-Tomato-None-Microwave-18/trial_T20190909_012524_159092/game.tw-pddl",
|
||||
"task_type": "pick_cool_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "val:0013",
|
||||
"gamefile": "json_2.1.1/valid_seen/pick_heat_then_place_in_recep-Apple-None-DiningTable-26/trial_T20190907_060234_011675/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "val:0014",
|
||||
"gamefile": "json_2.1.1/valid_seen/pick_heat_then_place_in_recep-Tomato-None-Fridge-15/trial_T20190909_020200_054379/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "val:0015",
|
||||
"gamefile": "json_2.1.1/valid_seen/pick_heat_then_place_in_recep-Tomato-None-Fridge-23/trial_T20190909_082320_103350/game.tw-pddl",
|
||||
"task_type": "pick_heat_then_place_in_recep"
|
||||
},
|
||||
{
|
||||
"id": "val:0016",
|
||||
"gamefile": "json_2.1.1/valid_seen/pick_two_obj_and_place-Book-None-Desk-313/trial_T20190908_125930_920681/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
},
|
||||
{
|
||||
"id": "val:0017",
|
||||
"gamefile": "json_2.1.1/valid_seen/pick_two_obj_and_place-CreditCard-None-Safe-323/trial_T20190907_001129_214240/game.tw-pddl",
|
||||
"task_type": "pick_two_obj_and_place"
|
||||
}
|
||||
]
|
||||
@@ -0,0 +1,36 @@
|
||||
{
|
||||
"benchmark": "DocVQA",
|
||||
"manifest_type": "id_split",
|
||||
"source_repo": "lmms-lab/DocVQA",
|
||||
"source_repo_type": "dataset",
|
||||
"source_url": "https://huggingface.co/datasets/lmms-lab/DocVQA",
|
||||
"source_revision": "539088ef8a8ada01ac8e2e6d4e372586748a265e",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"source_split_name": "docvqa_validation_10pct",
|
||||
"split_method": "10% subset sampled from the DocVQA validation split",
|
||||
"counts": {
|
||||
"train": 107,
|
||||
"val": 53,
|
||||
"test": 374
|
||||
},
|
||||
"item_fields": [
|
||||
"id",
|
||||
"questionId",
|
||||
"docId",
|
||||
"image_path",
|
||||
"ucsf_document_id",
|
||||
"ucsf_document_page_no",
|
||||
"topic",
|
||||
"source_dataset",
|
||||
"source_config",
|
||||
"source_split",
|
||||
"sample_seed"
|
||||
],
|
||||
"notes": [
|
||||
"This is a split manifest, not the full DocVQA payload.",
|
||||
"Materialize full CSV rows and image files before evaluation.",
|
||||
"This manifest corresponds to docvqa_validation_10pct.",
|
||||
"All released train/val/test items originate from a 10% subset of the official DocVQA validation split."
|
||||
]
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,691 @@
|
||||
[
|
||||
{
|
||||
"id": "62409",
|
||||
"questionId": "62409",
|
||||
"docId": "8554",
|
||||
"image_path": "data/docvqa_images/q62409_d8554.png",
|
||||
"ucsf_document_id": "pgjw0227",
|
||||
"ucsf_document_page_no": "5",
|
||||
"topic": "table/list",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "50961",
|
||||
"questionId": "50961",
|
||||
"docId": "549",
|
||||
"image_path": "data/docvqa_images/q50961_d549.png",
|
||||
"ucsf_document_id": "qtjf0226",
|
||||
"ucsf_document_page_no": "2",
|
||||
"topic": "free_text",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "46461",
|
||||
"questionId": "46461",
|
||||
"docId": "13361",
|
||||
"image_path": "data/docvqa_images/q46461_d13361.png",
|
||||
"ucsf_document_id": "ysbw0217",
|
||||
"ucsf_document_page_no": "5",
|
||||
"topic": "layout",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "3041",
|
||||
"questionId": "3041",
|
||||
"docId": "1204",
|
||||
"image_path": "data/docvqa_images/q3041_d1204.png",
|
||||
"ucsf_document_id": "xfjv0228",
|
||||
"ucsf_document_page_no": "3",
|
||||
"topic": "form",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "41716",
|
||||
"questionId": "41716",
|
||||
"docId": "11835",
|
||||
"image_path": "data/docvqa_images/q41716_d11835.png",
|
||||
"ucsf_document_id": "qjgn0226",
|
||||
"ucsf_document_page_no": "131",
|
||||
"topic": "form",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "61123",
|
||||
"questionId": "61123",
|
||||
"docId": "7374",
|
||||
"image_path": "data/docvqa_images/q61123_d7374.png",
|
||||
"ucsf_document_id": "mldg0227",
|
||||
"ucsf_document_page_no": "5",
|
||||
"topic": "layout",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "43068",
|
||||
"questionId": "43068",
|
||||
"docId": "12393",
|
||||
"image_path": "data/docvqa_images/q43068_d12393.png",
|
||||
"ucsf_document_id": "rmwn0226",
|
||||
"ucsf_document_page_no": "52",
|
||||
"topic": "figure/diagram",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "51221",
|
||||
"questionId": "51221",
|
||||
"docId": "764",
|
||||
"image_path": "data/docvqa_images/q51221_d764.png",
|
||||
"ucsf_document_id": "kzbn0226",
|
||||
"ucsf_document_page_no": "14",
|
||||
"topic": "layout",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "6397",
|
||||
"questionId": "6397",
|
||||
"docId": "2242",
|
||||
"image_path": "data/docvqa_images/q6397_d2242.png",
|
||||
"ucsf_document_id": "jkcn0000",
|
||||
"ucsf_document_page_no": "2",
|
||||
"topic": "form",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "57428",
|
||||
"questionId": "57428",
|
||||
"docId": "4779",
|
||||
"image_path": "data/docvqa_images/q57428_d4779.png",
|
||||
"ucsf_document_id": "rnbx0223",
|
||||
"ucsf_document_page_no": "208",
|
||||
"topic": "Image/Photo",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "3135",
|
||||
"questionId": "3135",
|
||||
"docId": "1221",
|
||||
"image_path": "data/docvqa_images/q3135_d1221.png",
|
||||
"ucsf_document_id": "ngph0227",
|
||||
"ucsf_document_page_no": "5",
|
||||
"topic": "layout",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "18819",
|
||||
"questionId": "18819",
|
||||
"docId": "5749",
|
||||
"image_path": "data/docvqa_images/q18819_d5749.png",
|
||||
"ucsf_document_id": "jhfd0079",
|
||||
"ucsf_document_page_no": "9",
|
||||
"topic": "form",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "15382",
|
||||
"questionId": "15382",
|
||||
"docId": "4890",
|
||||
"image_path": "data/docvqa_images/q15382_d4890.png",
|
||||
"ucsf_document_id": "kjvw0217",
|
||||
"ucsf_document_page_no": "3",
|
||||
"topic": "table/list",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "5772",
|
||||
"questionId": "5772",
|
||||
"docId": "1940",
|
||||
"image_path": "data/docvqa_images/q5772_d1940.png",
|
||||
"ucsf_document_id": "pzyw0224",
|
||||
"ucsf_document_page_no": "10",
|
||||
"topic": "form",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "49077",
|
||||
"questionId": "49077",
|
||||
"docId": "14179",
|
||||
"image_path": "data/docvqa_images/q49077_d14179.png",
|
||||
"ucsf_document_id": "nrxb0228",
|
||||
"ucsf_document_page_no": "3",
|
||||
"topic": "layout",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "58519",
|
||||
"questionId": "58519",
|
||||
"docId": "5347",
|
||||
"image_path": "data/docvqa_images/q58519_d5347.png",
|
||||
"ucsf_document_id": "sjbw0217",
|
||||
"ucsf_document_page_no": "11",
|
||||
"topic": "table/list",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "50720",
|
||||
"questionId": "50720",
|
||||
"docId": "281",
|
||||
"image_path": "data/docvqa_images/q50720_d281.png",
|
||||
"ucsf_document_id": "nrcj0037",
|
||||
"ucsf_document_page_no": "7",
|
||||
"topic": "table/list",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "56785",
|
||||
"questionId": "56785",
|
||||
"docId": "14289",
|
||||
"image_path": "data/docvqa_images/q56785_d14289.png",
|
||||
"ucsf_document_id": "xkbv0228",
|
||||
"ucsf_document_page_no": "1",
|
||||
"topic": "layout",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "59653",
|
||||
"questionId": "59653",
|
||||
"docId": "6579",
|
||||
"image_path": "data/docvqa_images/q59653_d6579.png",
|
||||
"ucsf_document_id": "mzbx0227",
|
||||
"ucsf_document_page_no": "2",
|
||||
"topic": "layout",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "61791",
|
||||
"questionId": "61791",
|
||||
"docId": "8072",
|
||||
"image_path": "data/docvqa_images/q61791_d8072.png",
|
||||
"ucsf_document_id": "hfmf0227",
|
||||
"ucsf_document_page_no": "1",
|
||||
"topic": "layout",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "37229",
|
||||
"questionId": "37229",
|
||||
"docId": "10742",
|
||||
"image_path": "data/docvqa_images/q37229_d10742.png",
|
||||
"ucsf_document_id": "nkcd0227",
|
||||
"ucsf_document_page_no": "2",
|
||||
"topic": "free_text",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "60407",
|
||||
"questionId": "60407",
|
||||
"docId": "7135",
|
||||
"image_path": "data/docvqa_images/q60407_d7135.png",
|
||||
"ucsf_document_id": "gkpk0226",
|
||||
"ucsf_document_page_no": "1",
|
||||
"topic": "free_text",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "64420",
|
||||
"questionId": "64420",
|
||||
"docId": "10230",
|
||||
"image_path": "data/docvqa_images/q64420_d10230.png",
|
||||
"ucsf_document_id": "jnjm0223",
|
||||
"ucsf_document_page_no": "107",
|
||||
"topic": "table/list",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "47365",
|
||||
"questionId": "47365",
|
||||
"docId": "13813",
|
||||
"image_path": "data/docvqa_images/q47365_d13813.png",
|
||||
"ucsf_document_id": "nxym0227",
|
||||
"ucsf_document_page_no": "28",
|
||||
"topic": "table/list",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "47458",
|
||||
"questionId": "47458",
|
||||
"docId": "13639",
|
||||
"image_path": "data/docvqa_images/q47458_d13639.png",
|
||||
"ucsf_document_id": "skdv0228",
|
||||
"ucsf_document_page_no": "5",
|
||||
"topic": "layout",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "7621",
|
||||
"questionId": "7621",
|
||||
"docId": "2668",
|
||||
"image_path": "data/docvqa_images/q7621_d2668.png",
|
||||
"ucsf_document_id": "flxn0020",
|
||||
"ucsf_document_page_no": "1",
|
||||
"topic": "form",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "53575",
|
||||
"questionId": "53575",
|
||||
"docId": "2766",
|
||||
"image_path": "data/docvqa_images/q53575_d2766.png",
|
||||
"ucsf_document_id": "hsfn0020",
|
||||
"ucsf_document_page_no": "2",
|
||||
"topic": "free_text|table/list",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "60913",
|
||||
"questionId": "60913",
|
||||
"docId": "7349",
|
||||
"image_path": "data/docvqa_images/q60913_d7349.png",
|
||||
"ucsf_document_id": "jzhd0227",
|
||||
"ucsf_document_page_no": "61",
|
||||
"topic": "free_text",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "60454",
|
||||
"questionId": "60454",
|
||||
"docId": "7163",
|
||||
"image_path": "data/docvqa_images/q60454_d7163.png",
|
||||
"ucsf_document_id": "jgyk0226",
|
||||
"ucsf_document_page_no": "1",
|
||||
"topic": "free_text",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "57978",
|
||||
"questionId": "57978",
|
||||
"docId": "4920",
|
||||
"image_path": "data/docvqa_images/q57978_d4920.png",
|
||||
"ucsf_document_id": "lkvw0217",
|
||||
"ucsf_document_page_no": "2",
|
||||
"topic": "table/list",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "64547",
|
||||
"questionId": "64547",
|
||||
"docId": "10361",
|
||||
"image_path": "data/docvqa_images/q64547_d10361.png",
|
||||
"ucsf_document_id": "lpdl0226",
|
||||
"ucsf_document_page_no": "32",
|
||||
"topic": "table/list",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "59481",
|
||||
"questionId": "59481",
|
||||
"docId": "6243",
|
||||
"image_path": "data/docvqa_images/q59481_d6243.png",
|
||||
"ucsf_document_id": "psgv0228",
|
||||
"ucsf_document_page_no": "5",
|
||||
"topic": "free_text",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "61472",
|
||||
"questionId": "61472",
|
||||
"docId": "7757",
|
||||
"image_path": "data/docvqa_images/q61472_d7757.png",
|
||||
"ucsf_document_id": "ymkp0227",
|
||||
"ucsf_document_page_no": "13",
|
||||
"topic": "handwritten|table/list",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "5673",
|
||||
"questionId": "5673",
|
||||
"docId": "1908",
|
||||
"image_path": "data/docvqa_images/q5673_d1908.png",
|
||||
"ucsf_document_id": "lldj0224",
|
||||
"ucsf_document_page_no": "2",
|
||||
"topic": "form",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "49109",
|
||||
"questionId": "49109",
|
||||
"docId": "13644",
|
||||
"image_path": "data/docvqa_images/q49109_d13644.png",
|
||||
"ucsf_document_id": "mzdv0228",
|
||||
"ucsf_document_page_no": "1",
|
||||
"topic": "table/list",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "46123",
|
||||
"questionId": "46123",
|
||||
"docId": "13503",
|
||||
"image_path": "data/docvqa_images/q46123_d13503.png",
|
||||
"ucsf_document_id": "xmww0217",
|
||||
"ucsf_document_page_no": "17",
|
||||
"topic": "table/list",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "48158",
|
||||
"questionId": "48158",
|
||||
"docId": "13976",
|
||||
"image_path": "data/docvqa_images/q48158_d13976.png",
|
||||
"ucsf_document_id": "zqhm0227",
|
||||
"ucsf_document_page_no": "1",
|
||||
"topic": "free_text",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "1955",
|
||||
"questionId": "1955",
|
||||
"docId": "892",
|
||||
"image_path": "data/docvqa_images/q1955_d892.png",
|
||||
"ucsf_document_id": "jsbn0226",
|
||||
"ucsf_document_page_no": "2",
|
||||
"topic": "table/list",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "8127",
|
||||
"questionId": "8127",
|
||||
"docId": "2754",
|
||||
"image_path": "data/docvqa_images/q8127_d2754.png",
|
||||
"ucsf_document_id": "xtvn0020",
|
||||
"ucsf_document_page_no": "2",
|
||||
"topic": "form",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "57431",
|
||||
"questionId": "57431",
|
||||
"docId": "4779",
|
||||
"image_path": "data/docvqa_images/q57431_d4779.png",
|
||||
"ucsf_document_id": "rnbx0223",
|
||||
"ucsf_document_page_no": "208",
|
||||
"topic": "Image/Photo",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "64306",
|
||||
"questionId": "64306",
|
||||
"docId": "10149",
|
||||
"image_path": "data/docvqa_images/q64306_d10149.png",
|
||||
"ucsf_document_id": "lpjm0223",
|
||||
"ucsf_document_page_no": "23",
|
||||
"topic": "table/list",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "64887",
|
||||
"questionId": "64887",
|
||||
"docId": "9754",
|
||||
"image_path": "data/docvqa_images/q64887_d9754.png",
|
||||
"ucsf_document_id": "szpg0227",
|
||||
"ucsf_document_page_no": "9",
|
||||
"topic": "table/list",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "58680",
|
||||
"questionId": "58680",
|
||||
"docId": "5545",
|
||||
"image_path": "data/docvqa_images/q58680_d5545.png",
|
||||
"ucsf_document_id": "hhwh0078",
|
||||
"ucsf_document_page_no": "1",
|
||||
"topic": "table/list",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "5287",
|
||||
"questionId": "5287",
|
||||
"docId": "1785",
|
||||
"image_path": "data/docvqa_images/q5287_d1785.png",
|
||||
"ucsf_document_id": "mtnh0227",
|
||||
"ucsf_document_page_no": "10",
|
||||
"topic": "form",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "55471",
|
||||
"questionId": "55471",
|
||||
"docId": "4340",
|
||||
"image_path": "data/docvqa_images/q55471_d4340.png",
|
||||
"ucsf_document_id": "fsgj0223",
|
||||
"ucsf_document_page_no": "96",
|
||||
"topic": "free_text",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "53095",
|
||||
"questionId": "53095",
|
||||
"docId": "296",
|
||||
"image_path": "data/docvqa_images/q53095_d296.png",
|
||||
"ucsf_document_id": "qhxj0037",
|
||||
"ucsf_document_page_no": "3",
|
||||
"topic": "layout",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "53726",
|
||||
"questionId": "53726",
|
||||
"docId": "2008",
|
||||
"image_path": "data/docvqa_images/q53726_d2008.png",
|
||||
"ucsf_document_id": "hhnf0094",
|
||||
"ucsf_document_page_no": "5",
|
||||
"topic": "layout",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "57321",
|
||||
"questionId": "57321",
|
||||
"docId": "4722",
|
||||
"image_path": "data/docvqa_images/q57321_d4722.png",
|
||||
"ucsf_document_id": "xybx0223",
|
||||
"ucsf_document_page_no": "32",
|
||||
"topic": "table/list",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "26659",
|
||||
"questionId": "26659",
|
||||
"docId": "7470",
|
||||
"image_path": "data/docvqa_images/q26659_d7470.png",
|
||||
"ucsf_document_id": "lhmg0227",
|
||||
"ucsf_document_page_no": "1",
|
||||
"topic": "layout",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "38920",
|
||||
"questionId": "38920",
|
||||
"docId": "11157",
|
||||
"image_path": "data/docvqa_images/q38920_d11157.png",
|
||||
"ucsf_document_id": "klnf0227",
|
||||
"ucsf_document_page_no": "1",
|
||||
"topic": "table/list|layout",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "50837",
|
||||
"questionId": "50837",
|
||||
"docId": "14742",
|
||||
"image_path": "data/docvqa_images/q50837_d14742.png",
|
||||
"ucsf_document_id": "ysmc0228",
|
||||
"ucsf_document_page_no": "4",
|
||||
"topic": "table/list",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "59615",
|
||||
"questionId": "59615",
|
||||
"docId": "6569",
|
||||
"image_path": "data/docvqa_images/q59615_d6569.png",
|
||||
"ucsf_document_id": "hnnp0227",
|
||||
"ucsf_document_page_no": "45",
|
||||
"topic": "handwritten|table/list|layout",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
},
|
||||
{
|
||||
"id": "58687",
|
||||
"questionId": "58687",
|
||||
"docId": "5545",
|
||||
"image_path": "data/docvqa_images/q58687_d5545.png",
|
||||
"ucsf_document_id": "hhwh0078",
|
||||
"ucsf_document_page_no": "1",
|
||||
"topic": "table/list",
|
||||
"source_dataset": "lmms-lab/DocVQA",
|
||||
"source_config": "DocVQA",
|
||||
"source_split": "validation",
|
||||
"sample_seed": "full_validation_5349"
|
||||
}
|
||||
]
|
||||
@@ -0,0 +1,34 @@
|
||||
{
|
||||
"benchmark": "LiveMathematicianBench",
|
||||
"manifest_type": "id_split",
|
||||
"source_repo": "LiveMathematicianBench/LiveMathematicianBench",
|
||||
"source_repo_type": "dataset",
|
||||
"source_url": "https://huggingface.co/datasets/LiveMathematicianBench/LiveMathematicianBench",
|
||||
"source_revision": "b72450f6ce96c26158d64d945a5d31ef7727be41",
|
||||
"source_files": [
|
||||
"data/202511/qa_202511_final.json",
|
||||
"data/202512/qa_202512_final.json",
|
||||
"data/202601/qa_202601_final.json",
|
||||
"data/202602/qa_202602_final.json"
|
||||
],
|
||||
"split_mode": "ratio",
|
||||
"split_ratio": "2:1:7",
|
||||
"split_seed": 42,
|
||||
"counts": {
|
||||
"train": 35,
|
||||
"val": 18,
|
||||
"test": 124
|
||||
},
|
||||
"item_fields": [
|
||||
"id",
|
||||
"month",
|
||||
"no",
|
||||
"paper_link",
|
||||
"source_file"
|
||||
],
|
||||
"id_format": "<month>:<no>",
|
||||
"notes": [
|
||||
"This is an ID split manifest, not the full LiveMathematicianBench payload.",
|
||||
"Materialize full split items from the official LiveMathematicianBench raw qa_*_final.json files before evaluation."
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,870 @@
|
||||
[
|
||||
{
|
||||
"id": "202602:12",
|
||||
"month": "202602",
|
||||
"no": 12,
|
||||
"paper_link": "http://arxiv.org/abs/2602.07171v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:3",
|
||||
"month": "202601",
|
||||
"no": 3,
|
||||
"paper_link": "http://arxiv.org/abs/2601.01447v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:4",
|
||||
"month": "202511",
|
||||
"no": 4,
|
||||
"paper_link": "http://arxiv.org/abs/2511.23123v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:20",
|
||||
"month": "202601",
|
||||
"no": 20,
|
||||
"paper_link": "http://arxiv.org/abs/2601.13212v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:42",
|
||||
"month": "202601",
|
||||
"no": 42,
|
||||
"paper_link": "http://arxiv.org/abs/2601.09348v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:38",
|
||||
"month": "202512",
|
||||
"no": 38,
|
||||
"paper_link": "http://arxiv.org/abs/2512.19831v2",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:4",
|
||||
"month": "202512",
|
||||
"no": 4,
|
||||
"paper_link": "http://arxiv.org/abs/2512.03141v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:4",
|
||||
"month": "202602",
|
||||
"no": 4,
|
||||
"paper_link": "http://arxiv.org/abs/2602.14368v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:15",
|
||||
"month": "202511",
|
||||
"no": 15,
|
||||
"paper_link": "http://arxiv.org/abs/2511.17325v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:32",
|
||||
"month": "202602",
|
||||
"no": 32,
|
||||
"paper_link": "http://arxiv.org/abs/2602.14817v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:51",
|
||||
"month": "202512",
|
||||
"no": 51,
|
||||
"paper_link": "http://arxiv.org/abs/2512.14581v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:26",
|
||||
"month": "202512",
|
||||
"no": 26,
|
||||
"paper_link": "http://arxiv.org/abs/2512.19586v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:13",
|
||||
"month": "202601",
|
||||
"no": 13,
|
||||
"paper_link": "http://arxiv.org/abs/2601.10017v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:1",
|
||||
"month": "202602",
|
||||
"no": 1,
|
||||
"paper_link": "http://arxiv.org/abs/2602.23137v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:18",
|
||||
"month": "202511",
|
||||
"no": 18,
|
||||
"paper_link": "http://arxiv.org/abs/2511.10795v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:5",
|
||||
"month": "202512",
|
||||
"no": 5,
|
||||
"paper_link": "http://arxiv.org/abs/2512.00348v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:19",
|
||||
"month": "202511",
|
||||
"no": 19,
|
||||
"paper_link": "http://arxiv.org/abs/2511.06951v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:40",
|
||||
"month": "202602",
|
||||
"no": 40,
|
||||
"paper_link": "http://arxiv.org/abs/2602.20462v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:29",
|
||||
"month": "202602",
|
||||
"no": 29,
|
||||
"paper_link": "http://arxiv.org/abs/2602.10676v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:35",
|
||||
"month": "202512",
|
||||
"no": 35,
|
||||
"paper_link": "http://arxiv.org/abs/2512.08840v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:48",
|
||||
"month": "202512",
|
||||
"no": 48,
|
||||
"paper_link": "http://arxiv.org/abs/2512.03482v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:52",
|
||||
"month": "202512",
|
||||
"no": 52,
|
||||
"paper_link": "http://arxiv.org/abs/2512.11246v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:44",
|
||||
"month": "202512",
|
||||
"no": 44,
|
||||
"paper_link": "http://arxiv.org/abs/2512.10385v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:28",
|
||||
"month": "202511",
|
||||
"no": 28,
|
||||
"paper_link": "http://arxiv.org/abs/2511.03812v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:43",
|
||||
"month": "202601",
|
||||
"no": 43,
|
||||
"paper_link": "http://arxiv.org/abs/2601.22555v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:9",
|
||||
"month": "202602",
|
||||
"no": 9,
|
||||
"paper_link": "http://arxiv.org/abs/2602.19882v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:23",
|
||||
"month": "202512",
|
||||
"no": 23,
|
||||
"paper_link": "http://arxiv.org/abs/2512.09180v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:21",
|
||||
"month": "202602",
|
||||
"no": 21,
|
||||
"paper_link": "http://arxiv.org/abs/2602.10509v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:5",
|
||||
"month": "202511",
|
||||
"no": 5,
|
||||
"paper_link": "http://arxiv.org/abs/2511.20164v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:35",
|
||||
"month": "202601",
|
||||
"no": 35,
|
||||
"paper_link": "http://arxiv.org/abs/2601.15606v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:50",
|
||||
"month": "202602",
|
||||
"no": 50,
|
||||
"paper_link": "http://arxiv.org/abs/2602.05652v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:13",
|
||||
"month": "202512",
|
||||
"no": 13,
|
||||
"paper_link": "http://arxiv.org/abs/2512.22861v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:49",
|
||||
"month": "202602",
|
||||
"no": 49,
|
||||
"paper_link": "http://arxiv.org/abs/2602.07167v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:18",
|
||||
"month": "202602",
|
||||
"no": 18,
|
||||
"paper_link": "http://arxiv.org/abs/2602.20124v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:15",
|
||||
"month": "202601",
|
||||
"no": 15,
|
||||
"paper_link": "http://arxiv.org/abs/2601.05327v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:21",
|
||||
"month": "202601",
|
||||
"no": 21,
|
||||
"paper_link": "http://arxiv.org/abs/2601.04994v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:32",
|
||||
"month": "202601",
|
||||
"no": 32,
|
||||
"paper_link": "http://arxiv.org/abs/2601.09183v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:34",
|
||||
"month": "202602",
|
||||
"no": 34,
|
||||
"paper_link": "http://arxiv.org/abs/2602.21118v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:20",
|
||||
"month": "202602",
|
||||
"no": 20,
|
||||
"paper_link": "http://arxiv.org/abs/2602.16506v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:5",
|
||||
"month": "202602",
|
||||
"no": 5,
|
||||
"paper_link": "http://arxiv.org/abs/2602.09806v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:40",
|
||||
"month": "202512",
|
||||
"no": 40,
|
||||
"paper_link": "http://arxiv.org/abs/2512.16535v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:22",
|
||||
"month": "202511",
|
||||
"no": 22,
|
||||
"paper_link": "http://arxiv.org/abs/2511.07607v2",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:36",
|
||||
"month": "202601",
|
||||
"no": 36,
|
||||
"paper_link": "http://arxiv.org/abs/2601.12457v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:49",
|
||||
"month": "202512",
|
||||
"no": 49,
|
||||
"paper_link": "http://arxiv.org/abs/2512.21565v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:10",
|
||||
"month": "202511",
|
||||
"no": 10,
|
||||
"paper_link": "http://arxiv.org/abs/2511.06484v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:2",
|
||||
"month": "202601",
|
||||
"no": 2,
|
||||
"paper_link": "http://arxiv.org/abs/2601.07068v4",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:19",
|
||||
"month": "202602",
|
||||
"no": 19,
|
||||
"paper_link": "http://arxiv.org/abs/2602.18179v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:9",
|
||||
"month": "202601",
|
||||
"no": 9,
|
||||
"paper_link": "http://arxiv.org/abs/2601.17765v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:6",
|
||||
"month": "202512",
|
||||
"no": 6,
|
||||
"paper_link": "http://arxiv.org/abs/2512.23079v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:5",
|
||||
"month": "202601",
|
||||
"no": 5,
|
||||
"paper_link": "http://arxiv.org/abs/2601.20344v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:14",
|
||||
"month": "202602",
|
||||
"no": 14,
|
||||
"paper_link": "http://arxiv.org/abs/2602.09177v2",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:17",
|
||||
"month": "202512",
|
||||
"no": 17,
|
||||
"paper_link": "http://arxiv.org/abs/2512.11657v2",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:19",
|
||||
"month": "202512",
|
||||
"no": 19,
|
||||
"paper_link": "http://arxiv.org/abs/2512.16655v2",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:33",
|
||||
"month": "202602",
|
||||
"no": 33,
|
||||
"paper_link": "http://arxiv.org/abs/2602.13734v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:18",
|
||||
"month": "202512",
|
||||
"no": 18,
|
||||
"paper_link": "http://arxiv.org/abs/2512.22960v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:26",
|
||||
"month": "202601",
|
||||
"no": 26,
|
||||
"paper_link": "http://arxiv.org/abs/2601.06814v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:1",
|
||||
"month": "202601",
|
||||
"no": 1,
|
||||
"paper_link": "http://arxiv.org/abs/2601.18276v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:30",
|
||||
"month": "202512",
|
||||
"no": 30,
|
||||
"paper_link": "http://arxiv.org/abs/2512.07260v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:44",
|
||||
"month": "202602",
|
||||
"no": 44,
|
||||
"paper_link": "http://arxiv.org/abs/2602.01138v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:20",
|
||||
"month": "202512",
|
||||
"no": 20,
|
||||
"paper_link": "http://arxiv.org/abs/2512.14575v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:13",
|
||||
"month": "202511",
|
||||
"no": 13,
|
||||
"paper_link": "http://arxiv.org/abs/2511.16910v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:30",
|
||||
"month": "202601",
|
||||
"no": 30,
|
||||
"paper_link": "http://arxiv.org/abs/2601.12140v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:40",
|
||||
"month": "202601",
|
||||
"no": 40,
|
||||
"paper_link": "http://arxiv.org/abs/2601.05146v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:29",
|
||||
"month": "202601",
|
||||
"no": 29,
|
||||
"paper_link": "http://arxiv.org/abs/2601.12846v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:11",
|
||||
"month": "202511",
|
||||
"no": 11,
|
||||
"paper_link": "http://arxiv.org/abs/2511.17548v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:9",
|
||||
"month": "202512",
|
||||
"no": 9,
|
||||
"paper_link": "http://arxiv.org/abs/2512.08817v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:18",
|
||||
"month": "202601",
|
||||
"no": 18,
|
||||
"paper_link": "http://arxiv.org/abs/2601.01797v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:1",
|
||||
"month": "202512",
|
||||
"no": 1,
|
||||
"paper_link": "http://arxiv.org/abs/2512.20055v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:4",
|
||||
"month": "202601",
|
||||
"no": 4,
|
||||
"paper_link": "http://arxiv.org/abs/2601.21223v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:6",
|
||||
"month": "202511",
|
||||
"no": 6,
|
||||
"paper_link": "http://arxiv.org/abs/2511.14959v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:38",
|
||||
"month": "202602",
|
||||
"no": 38,
|
||||
"paper_link": "http://arxiv.org/abs/2602.08398v2",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:10",
|
||||
"month": "202601",
|
||||
"no": 10,
|
||||
"paper_link": "http://arxiv.org/abs/2601.15524v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:11",
|
||||
"month": "202602",
|
||||
"no": 11,
|
||||
"paper_link": "http://arxiv.org/abs/2602.11045v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:45",
|
||||
"month": "202512",
|
||||
"no": 45,
|
||||
"paper_link": "http://arxiv.org/abs/2512.08395v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:12",
|
||||
"month": "202601",
|
||||
"no": 12,
|
||||
"paper_link": "http://arxiv.org/abs/2601.11877v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:47",
|
||||
"month": "202512",
|
||||
"no": 47,
|
||||
"paper_link": "http://arxiv.org/abs/2512.09683v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:21",
|
||||
"month": "202511",
|
||||
"no": 21,
|
||||
"paper_link": "http://arxiv.org/abs/2511.21288v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:16",
|
||||
"month": "202601",
|
||||
"no": 16,
|
||||
"paper_link": "http://arxiv.org/abs/2601.05008v2",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:3",
|
||||
"month": "202512",
|
||||
"no": 3,
|
||||
"paper_link": "http://arxiv.org/abs/2512.13450v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:37",
|
||||
"month": "202601",
|
||||
"no": 37,
|
||||
"paper_link": "http://arxiv.org/abs/2601.09443v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:12",
|
||||
"month": "202511",
|
||||
"no": 12,
|
||||
"paper_link": "http://arxiv.org/abs/2511.04978v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:39",
|
||||
"month": "202512",
|
||||
"no": 39,
|
||||
"paper_link": "http://arxiv.org/abs/2512.19003v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:8",
|
||||
"month": "202601",
|
||||
"no": 8,
|
||||
"paper_link": "http://arxiv.org/abs/2601.19754v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:11",
|
||||
"month": "202601",
|
||||
"no": 11,
|
||||
"paper_link": "http://arxiv.org/abs/2601.13552v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:25",
|
||||
"month": "202511",
|
||||
"no": 25,
|
||||
"paper_link": "http://arxiv.org/abs/2511.10548v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:17",
|
||||
"month": "202601",
|
||||
"no": 17,
|
||||
"paper_link": "http://arxiv.org/abs/2601.02655v2",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:36",
|
||||
"month": "202602",
|
||||
"no": 36,
|
||||
"paper_link": "http://arxiv.org/abs/2602.13001v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:43",
|
||||
"month": "202602",
|
||||
"no": 43,
|
||||
"paper_link": "http://arxiv.org/abs/2602.06897v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:6",
|
||||
"month": "202601",
|
||||
"no": 6,
|
||||
"paper_link": "http://arxiv.org/abs/2601.04747v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:35",
|
||||
"month": "202602",
|
||||
"no": 35,
|
||||
"paper_link": "http://arxiv.org/abs/2602.20938v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:11",
|
||||
"month": "202512",
|
||||
"no": 11,
|
||||
"paper_link": "http://arxiv.org/abs/2512.03294v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:23",
|
||||
"month": "202602",
|
||||
"no": 23,
|
||||
"paper_link": "http://arxiv.org/abs/2602.09201v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:7",
|
||||
"month": "202601",
|
||||
"no": 7,
|
||||
"paper_link": "http://arxiv.org/abs/2601.02859v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:39",
|
||||
"month": "202602",
|
||||
"no": 39,
|
||||
"paper_link": "http://arxiv.org/abs/2602.21659v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:12",
|
||||
"month": "202512",
|
||||
"no": 12,
|
||||
"paper_link": "http://arxiv.org/abs/2512.00690v3",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:2",
|
||||
"month": "202511",
|
||||
"no": 2,
|
||||
"paper_link": "http://arxiv.org/abs/2511.19681v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:43",
|
||||
"month": "202512",
|
||||
"no": 43,
|
||||
"paper_link": "http://arxiv.org/abs/2512.10820v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:24",
|
||||
"month": "202602",
|
||||
"no": 24,
|
||||
"paper_link": "http://arxiv.org/abs/2602.08680v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:34",
|
||||
"month": "202601",
|
||||
"no": 34,
|
||||
"paper_link": "http://arxiv.org/abs/2601.07318v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:28",
|
||||
"month": "202512",
|
||||
"no": 28,
|
||||
"paper_link": "http://arxiv.org/abs/2512.11294v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:27",
|
||||
"month": "202601",
|
||||
"no": 27,
|
||||
"paper_link": "http://arxiv.org/abs/2601.05692v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:42",
|
||||
"month": "202602",
|
||||
"no": 42,
|
||||
"paper_link": "http://arxiv.org/abs/2602.09749v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:22",
|
||||
"month": "202512",
|
||||
"no": 22,
|
||||
"paper_link": "http://arxiv.org/abs/2512.11658v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:17",
|
||||
"month": "202602",
|
||||
"no": 17,
|
||||
"paper_link": "http://arxiv.org/abs/2602.22504v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:48",
|
||||
"month": "202602",
|
||||
"no": 48,
|
||||
"paper_link": "http://arxiv.org/abs/2602.08760v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:28",
|
||||
"month": "202602",
|
||||
"no": 28,
|
||||
"paper_link": "http://arxiv.org/abs/2602.11595v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:3",
|
||||
"month": "202602",
|
||||
"no": 3,
|
||||
"paper_link": "http://arxiv.org/abs/2602.17369v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:31",
|
||||
"month": "202512",
|
||||
"no": 31,
|
||||
"paper_link": "http://arxiv.org/abs/2512.23668v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:27",
|
||||
"month": "202512",
|
||||
"no": 27,
|
||||
"paper_link": "http://arxiv.org/abs/2512.16505v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:24",
|
||||
"month": "202511",
|
||||
"no": 24,
|
||||
"paper_link": "http://arxiv.org/abs/2511.12549v2",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:8",
|
||||
"month": "202511",
|
||||
"no": 8,
|
||||
"paper_link": "http://arxiv.org/abs/2511.12657v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:9",
|
||||
"month": "202511",
|
||||
"no": 9,
|
||||
"paper_link": "http://arxiv.org/abs/2511.09015v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:28",
|
||||
"month": "202601",
|
||||
"no": 28,
|
||||
"paper_link": "http://arxiv.org/abs/2601.14825v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:25",
|
||||
"month": "202602",
|
||||
"no": 25,
|
||||
"paper_link": "http://arxiv.org/abs/2602.16048v3",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:23",
|
||||
"month": "202511",
|
||||
"no": 23,
|
||||
"paper_link": "http://arxiv.org/abs/2511.06595v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:13",
|
||||
"month": "202602",
|
||||
"no": 13,
|
||||
"paper_link": "http://arxiv.org/abs/2602.12261v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:27",
|
||||
"month": "202511",
|
||||
"no": 27,
|
||||
"paper_link": "http://arxiv.org/abs/2511.04407v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:7",
|
||||
"month": "202512",
|
||||
"no": 7,
|
||||
"paper_link": "http://arxiv.org/abs/2512.09490v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:29",
|
||||
"month": "202512",
|
||||
"no": 29,
|
||||
"paper_link": "http://arxiv.org/abs/2512.08562v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:34",
|
||||
"month": "202512",
|
||||
"no": 34,
|
||||
"paper_link": "http://arxiv.org/abs/2512.09598v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:42",
|
||||
"month": "202512",
|
||||
"no": 42,
|
||||
"paper_link": "http://arxiv.org/abs/2512.10845v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:7",
|
||||
"month": "202511",
|
||||
"no": 7,
|
||||
"paper_link": "http://arxiv.org/abs/2511.13976v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:29",
|
||||
"month": "202511",
|
||||
"no": 29,
|
||||
"paper_link": "http://arxiv.org/abs/2511.03722v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:37",
|
||||
"month": "202602",
|
||||
"no": 37,
|
||||
"paper_link": "http://arxiv.org/abs/2602.08644v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
}
|
||||
]
|
||||
@@ -0,0 +1,247 @@
|
||||
[
|
||||
{
|
||||
"id": "202602:22",
|
||||
"month": "202602",
|
||||
"no": 22,
|
||||
"paper_link": "http://arxiv.org/abs/2602.10700v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:8",
|
||||
"month": "202512",
|
||||
"no": 8,
|
||||
"paper_link": "http://arxiv.org/abs/2512.08863v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:16",
|
||||
"month": "202511",
|
||||
"no": 16,
|
||||
"paper_link": "http://arxiv.org/abs/2511.15668v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:44",
|
||||
"month": "202601",
|
||||
"no": 44,
|
||||
"paper_link": "http://arxiv.org/abs/2601.21267v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:14",
|
||||
"month": "202511",
|
||||
"no": 14,
|
||||
"paper_link": "http://arxiv.org/abs/2511.13447v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:30",
|
||||
"month": "202602",
|
||||
"no": 30,
|
||||
"paper_link": "http://arxiv.org/abs/2602.16692v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:2",
|
||||
"month": "202602",
|
||||
"no": 2,
|
||||
"paper_link": "http://arxiv.org/abs/2602.22933v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:41",
|
||||
"month": "202601",
|
||||
"no": 41,
|
||||
"paper_link": "http://arxiv.org/abs/2601.01164v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:23",
|
||||
"month": "202601",
|
||||
"no": 23,
|
||||
"paper_link": "http://arxiv.org/abs/2601.02528v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:20",
|
||||
"month": "202511",
|
||||
"no": 20,
|
||||
"paper_link": "http://arxiv.org/abs/2511.02963v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:22",
|
||||
"month": "202601",
|
||||
"no": 22,
|
||||
"paper_link": "http://arxiv.org/abs/2601.03984v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:14",
|
||||
"month": "202512",
|
||||
"no": 14,
|
||||
"paper_link": "http://arxiv.org/abs/2512.22459v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:26",
|
||||
"month": "202511",
|
||||
"no": 26,
|
||||
"paper_link": "http://arxiv.org/abs/2511.07817v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:3",
|
||||
"month": "202511",
|
||||
"no": 3,
|
||||
"paper_link": "http://arxiv.org/abs/2511.11409v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:33",
|
||||
"month": "202601",
|
||||
"no": 33,
|
||||
"paper_link": "http://arxiv.org/abs/2601.07747v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:7",
|
||||
"month": "202602",
|
||||
"no": 7,
|
||||
"paper_link": "http://arxiv.org/abs/2602.22912v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:27",
|
||||
"month": "202602",
|
||||
"no": 27,
|
||||
"paper_link": "http://arxiv.org/abs/2602.13968v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:31",
|
||||
"month": "202602",
|
||||
"no": 31,
|
||||
"paper_link": "http://arxiv.org/abs/2602.15528v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:41",
|
||||
"month": "202602",
|
||||
"no": 41,
|
||||
"paper_link": "http://arxiv.org/abs/2602.10707v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:25",
|
||||
"month": "202512",
|
||||
"no": 25,
|
||||
"paper_link": "http://arxiv.org/abs/2512.04531v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:10",
|
||||
"month": "202602",
|
||||
"no": 10,
|
||||
"paper_link": "http://arxiv.org/abs/2602.17863v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:16",
|
||||
"month": "202602",
|
||||
"no": 16,
|
||||
"paper_link": "http://arxiv.org/abs/2602.02723v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:16",
|
||||
"month": "202512",
|
||||
"no": 16,
|
||||
"paper_link": "http://arxiv.org/abs/2512.11601v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:2",
|
||||
"month": "202512",
|
||||
"no": 2,
|
||||
"paper_link": "http://arxiv.org/abs/2512.16120v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:24",
|
||||
"month": "202512",
|
||||
"no": 24,
|
||||
"paper_link": "http://arxiv.org/abs/2512.08391v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:32",
|
||||
"month": "202512",
|
||||
"no": 32,
|
||||
"paper_link": "http://arxiv.org/abs/2512.23224v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:47",
|
||||
"month": "202602",
|
||||
"no": 47,
|
||||
"paper_link": "http://arxiv.org/abs/2602.10391v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:46",
|
||||
"month": "202602",
|
||||
"no": 46,
|
||||
"paper_link": "http://arxiv.org/abs/2602.13727v2",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:21",
|
||||
"month": "202512",
|
||||
"no": 21,
|
||||
"paper_link": "http://arxiv.org/abs/2512.12835v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:33",
|
||||
"month": "202512",
|
||||
"no": 33,
|
||||
"paper_link": "http://arxiv.org/abs/2512.19500v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:45",
|
||||
"month": "202602",
|
||||
"no": 45,
|
||||
"paper_link": "http://arxiv.org/abs/2602.23912v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:26",
|
||||
"month": "202602",
|
||||
"no": 26,
|
||||
"paper_link": "http://arxiv.org/abs/2602.14658v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:41",
|
||||
"month": "202512",
|
||||
"no": 41,
|
||||
"paper_link": "http://arxiv.org/abs/2512.15177v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:38",
|
||||
"month": "202601",
|
||||
"no": 38,
|
||||
"paper_link": "http://arxiv.org/abs/2601.07817v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:14",
|
||||
"month": "202601",
|
||||
"no": 14,
|
||||
"paper_link": "http://arxiv.org/abs/2601.08704v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
}
|
||||
]
|
||||
@@ -0,0 +1,128 @@
|
||||
[
|
||||
{
|
||||
"id": "202602:8",
|
||||
"month": "202602",
|
||||
"no": 8,
|
||||
"paper_link": "http://arxiv.org/abs/2602.19529v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:50",
|
||||
"month": "202512",
|
||||
"no": 50,
|
||||
"paper_link": "http://arxiv.org/abs/2512.15277v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:36",
|
||||
"month": "202512",
|
||||
"no": 36,
|
||||
"paper_link": "http://arxiv.org/abs/2512.06696v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:1",
|
||||
"month": "202511",
|
||||
"no": 1,
|
||||
"paper_link": "http://arxiv.org/abs/2511.04651v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:31",
|
||||
"month": "202601",
|
||||
"no": 31,
|
||||
"paper_link": "http://arxiv.org/abs/2601.10298v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202511:17",
|
||||
"month": "202511",
|
||||
"no": 17,
|
||||
"paper_link": "http://arxiv.org/abs/2511.13215v1",
|
||||
"source_file": "data/202511/qa_202511_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:37",
|
||||
"month": "202512",
|
||||
"no": 37,
|
||||
"paper_link": "http://arxiv.org/abs/2512.20498v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:39",
|
||||
"month": "202601",
|
||||
"no": 39,
|
||||
"paper_link": "http://arxiv.org/abs/2601.06601v2",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:25",
|
||||
"month": "202601",
|
||||
"no": 25,
|
||||
"paper_link": "http://arxiv.org/abs/2601.10996v3",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:24",
|
||||
"month": "202601",
|
||||
"no": 24,
|
||||
"paper_link": "http://arxiv.org/abs/2601.12250v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:45",
|
||||
"month": "202601",
|
||||
"no": 45,
|
||||
"paper_link": "http://arxiv.org/abs/2601.12113v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:19",
|
||||
"month": "202601",
|
||||
"no": 19,
|
||||
"paper_link": "http://arxiv.org/abs/2601.00779v1",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:10",
|
||||
"month": "202512",
|
||||
"no": 10,
|
||||
"paper_link": "http://arxiv.org/abs/2512.07073v2",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202601:46",
|
||||
"month": "202601",
|
||||
"no": 46,
|
||||
"paper_link": "http://arxiv.org/abs/2601.07793v2",
|
||||
"source_file": "data/202601/qa_202601_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:15",
|
||||
"month": "202512",
|
||||
"no": 15,
|
||||
"paper_link": "http://arxiv.org/abs/2512.16165v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:15",
|
||||
"month": "202602",
|
||||
"no": 15,
|
||||
"paper_link": "http://arxiv.org/abs/2602.05303v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202602:6",
|
||||
"month": "202602",
|
||||
"no": 6,
|
||||
"paper_link": "http://arxiv.org/abs/2602.01571v1",
|
||||
"source_file": "data/202602/qa_202602_final.json"
|
||||
},
|
||||
{
|
||||
"id": "202512:46",
|
||||
"month": "202512",
|
||||
"no": 46,
|
||||
"paper_link": "http://arxiv.org/abs/2512.05945v1",
|
||||
"source_file": "data/202512/qa_202512_final.json"
|
||||
}
|
||||
]
|
||||
@@ -0,0 +1,27 @@
|
||||
{
|
||||
"benchmark": "OfficeQA",
|
||||
"manifest_type": "id_split",
|
||||
"source_repo": "databricks/officeqa",
|
||||
"source_repo_type": "dataset",
|
||||
"source_url": "https://huggingface.co/datasets/databricks/officeqa",
|
||||
"source_revision": "8ecbf18d3833daf4750a903d14963e4c4c1d4cd8",
|
||||
"source_file": "officeqa_full.csv",
|
||||
"source_split_name": "officeqa_split",
|
||||
"counts": {
|
||||
"train": 50,
|
||||
"val": 24,
|
||||
"test": 172
|
||||
},
|
||||
"item_fields": [
|
||||
"id",
|
||||
"uid",
|
||||
"category",
|
||||
"source_files",
|
||||
"source_docs",
|
||||
"source_split"
|
||||
],
|
||||
"notes": [
|
||||
"This is a split manifest, not the full OfficeQA payload.",
|
||||
"The official OfficeQA CSV is gated on Hugging Face; materialization requires authorized access."
|
||||
]
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,402 @@
|
||||
[
|
||||
{
|
||||
"id": "UID0002",
|
||||
"uid": "UID0002",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1944_01.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/january-1944-6565?page=18",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0007",
|
||||
"uid": "UID0007",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1950_02.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/february-1950-6637?page=15",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0014",
|
||||
"uid": "UID0014",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1942_07.txt\r\ntreasury_bulletin_2001_12.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/july-1942-6547?page=76\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/december-2001-7108?page=17&deep=true",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0017",
|
||||
"uid": "UID0017",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1982_08.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/august-1982-7028?page=13",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0018",
|
||||
"uid": "UID0018",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1985_03.txt\r\ntreasury_bulletin_1986_03.txt\r\ntreasury_bulletin_1987_03.txt\r\ntreasury_bulletin_1988_03.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/march-1985-7040?page=22\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/march-1986-7045?page=26\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/march-1987-7049?page=24\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/march-1988-7052?page=36",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0019",
|
||||
"uid": "UID0019",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_2016_09.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/september-2016-533966?page=54\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/september-2016-533966?page=58",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0026",
|
||||
"uid": "UID0026",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1963_01.txt\r\ntreasury_bulletin_1962_01.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/january-1963-6793?page=88\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/january-1962-6781?page=82&deep=true",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0028",
|
||||
"uid": "UID0028",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1970_06.txt\r\ntreasury_bulletin_1964_12.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/june-1970-6882?page=89&deep=true\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/december-1964-6816?page=25&deep=true",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0030",
|
||||
"uid": "UID0030",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1990_09.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/september-1990-7062?page=19&deep=true",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0031",
|
||||
"uid": "UID0031",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1992_03.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/march-1992-7068?page=158&deep=true",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0033",
|
||||
"uid": "UID0033",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1977_04.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/april-1977-6964?page=9",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0034",
|
||||
"uid": "UID0034",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1992_06.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/june-1992-7069?page=32",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0044",
|
||||
"uid": "UID0044",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1939_04.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/april-1939-6506?page=61",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0046",
|
||||
"uid": "UID0046",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1988_09.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/september-1988-7054?page=37",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0049",
|
||||
"uid": "UID0049",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1942_02.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/february-1942-6542?page=19&deep=true",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0056",
|
||||
"uid": "UID0056",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1991_09.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/september-1991-7066?page=30&deep=true",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0063",
|
||||
"uid": "UID0063",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1990_06.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/june-1990-7061?page=127",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0065",
|
||||
"uid": "UID0065",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1998_06.txt\r\ntreasury_bulletin_1995_12.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/june-1998-7094?page=7\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/december-1995-7084?page=16",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0073",
|
||||
"uid": "UID0073",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1982_03.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/march-1982-7023?page=24",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0079",
|
||||
"uid": "UID0079",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_2011_12.txt\r\ntreasury_bulletin_2013_12.txt\r\ntreasury_bulletin_2015_12.txt\r\ntreasury_bulletin_2017_12.txt\r\ntreasury_bulletin_2019_12.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/december-2011-7149?page=25\r\n \r\n https://fraser.stlouisfed.org/title/treasury-bulletin-407/december-2013-7155?page=24\r\n \r\n https://fraser.stlouisfed.org/title/treasury-bulletin-407/december-2015-519209?page=23\r\n \r\n https://fraser.stlouisfed.org/title/treasury-bulletin-407/december-2017-575188?page=23\r\n \r\n https://fraser.stlouisfed.org/title/treasury-bulletin-407/december-2019-584842?page=22",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0083",
|
||||
"uid": "UID0083",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1981_12.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/december-1981-7020?page=24",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0085",
|
||||
"uid": "UID0085",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_2019_12.txt\r\ntreasury_bulletin_2018_12.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/december-2019-584842?page=23\r\n\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/december-2018-581283?page=22",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0087",
|
||||
"uid": "UID0087",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_2013_12.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/december-2013-7155?page=17",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0092",
|
||||
"uid": "UID0092",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1987_12.txt\r\ntreasury_bulletin_1992_12.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/december-1987-7051?page=69\r\n\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/december-1992-7071?page=84",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0098",
|
||||
"uid": "UID0098",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_2020_12.txt\r\ntreasury_bulletin_2024_12.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/december-2020-598551?page=21\r\n \r\n https://fraser.stlouisfed.org/title/treasury-bulletin-407/december-2024-679984?page=22",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0101",
|
||||
"uid": "UID0101",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_2011_12.txt\r\ntreasury_bulletin_2019_12.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/december-2011-7149?page=25\r\n \r\n https://fraser.stlouisfed.org/title/treasury-bulletin-407/december-2019-584842?page=22",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0110",
|
||||
"uid": "UID0110",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_2020_03.txt\r\ntreasury_bulletin_2016_03.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/march-2020-587316?page=10\r\n \r\n https://fraser.stlouisfed.org/title/treasury-bulletin-407/march-2016-527290?page=9",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0115",
|
||||
"uid": "UID0115",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1980_02.txt\r\ntreasury_bulletin_1981_02.txt\r\ntreasury_bulletin_1982_02.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/february-1980-6998?page=27\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/february-1981-7010?page=38\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/february-1982-7022?page=31",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0122",
|
||||
"uid": "UID0122",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_2001_03.txt\r\ntreasury_bulletin_2002_03.txt\r\ntreasury_bulletin_2003_03.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/march-2001-7105?page=112\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/march-2002-7110?page=115\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/march-2003-7109?page=113",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0123",
|
||||
"uid": "UID0123",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1941_12.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/december-1941-6540?page=91",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0133",
|
||||
"uid": "UID0133",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_2004_09.txt\r\ntreasury_bulletin_2013_12.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/september-2004-7119?page=63\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/december-2013-7155?page=68",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0141",
|
||||
"uid": "UID0141",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1962_04.txt\r\ntreasury_bulletin_1963_04.txt\r\ntreasury_bulletin_1964_04.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/april-1962-6784?page=75\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/april-1963-6796?page=79\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/april-1964-6808?page=82",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0144",
|
||||
"uid": "UID0144",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1980_11.txt\r\ntreasury_bulletin_1981_11.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/november-1980-7007?page=76\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/november-1981-7019?page=67",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0145",
|
||||
"uid": "UID0145",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1943_01.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/january-1943-6553?page=41",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0150",
|
||||
"uid": "UID0150",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1972_04.txt\r\ntreasury_bulletin_1973_04.txt\r\ntreasury_bulletin_1974_04.txt\r\ntreasury_bulletin_1975_04.txt\r\ntreasury_bulletin_1976_04.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/april-1972-6905?page=89\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/april-1973-6916?page=91\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/april-1974-6927?page=88\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/april-1975-6940?page=88\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/april-1976-6952?page=104",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0151",
|
||||
"uid": "UID0151",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1953_03.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/march-1953-6674?page=54",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0162",
|
||||
"uid": "UID0162",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_2011_06.txt\r\ntreasury_bulletin_2012_06.txt\r\ntreasury_bulletin_2013_06.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/june-2011-7129?page=105\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/june-2012-7151?page=105\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/june-2013-7153?page=104",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0163",
|
||||
"uid": "UID0163",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1981_12.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/december-1981-7020?page=28",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0165",
|
||||
"uid": "UID0165",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_2010_03.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/march-2010-7143?page=49",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0166",
|
||||
"uid": "UID0166",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1943_03.txt\r\ntreasury_bulletin_1944_03.txt\r\ntreasury_bulletin_1945_03.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/march-1943-6555?page=68\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/march-1944-6567?page=83\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/march-1945-6577?page=71",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0169",
|
||||
"uid": "UID0169",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1982_03.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/march-1982-7023?page=73",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0189",
|
||||
"uid": "UID0189",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1970_08.txt\r\ntreasury_bulletin_1970_09.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/august-1970-6884?page=70\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/september-1970-6885?page=70",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0195",
|
||||
"uid": "UID0195",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1956_08.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/august-1956-6715?page=59",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0202",
|
||||
"uid": "UID0202",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1939_07.txt\r\ntreasury_bulletin_1939_08.txt\r\ntreasury_bulletin_1939_09.txt\r\ntreasury_bulletin_1939_10.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/july-1939-6509?page=99\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/august-1939-6510?page=107\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/september-1939-6511?page=60\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/october-1939-6520?page=62",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0212",
|
||||
"uid": "UID0212",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1964_01.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/january-1964-6805?page=99",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0222",
|
||||
"uid": "UID0222",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_2001_06.txt\r\ntreasury_bulletin_2006_06.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/june-2001-7106?page=50\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/june-2006-7126?page=50",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0228",
|
||||
"uid": "UID0228",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1956_03.txt\r\ntreasury_bulletin_1956_04.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/march-1956-6710?page=22\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/april-1956-6711?page=22",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0229",
|
||||
"uid": "UID0229",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_2005_03.txt\r\ntreasury_bulletin_2006_03.txt\r\ntreasury_bulletin_2007_03.txt\r\ntreasury_bulletin_2008_03.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/march-2005-7121?page=109\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/march-2006-7125?page=106\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/march-2007-7130?page=109\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/march-2008-7134?page=107",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0238",
|
||||
"uid": "UID0238",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1982_03.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/march-1982-7023?page=44",
|
||||
"source_split": "train"
|
||||
},
|
||||
{
|
||||
"id": "UID0241",
|
||||
"uid": "UID0241",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1963_06.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/june-1963-6798?page=13",
|
||||
"source_split": "train"
|
||||
}
|
||||
]
|
||||
@@ -0,0 +1,194 @@
|
||||
[
|
||||
{
|
||||
"id": "UID0001",
|
||||
"uid": "UID0001",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1941_01.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/january-1941-6529?page=15",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0027",
|
||||
"uid": "UID0027",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1970_06.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/june-1970-6882?page=89&deep=true",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0039",
|
||||
"uid": "UID0039",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_2004_03.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/march-2004-7117?page=20\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/march-2004-7117?page=21&deep=true",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0041",
|
||||
"uid": "UID0041",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1970_10.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/october-1970-6886?page=35",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0052",
|
||||
"uid": "UID0052",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_2000_06.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/june-2000-7102?page=56",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0070",
|
||||
"uid": "UID0070",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1939_01.txt\r\ntreasury_bulletin_1939_02.txt\r\ntreasury_bulletin_1939_03.txt\r\ntreasury_bulletin_1939_04.txt\r\ntreasury_bulletin_1939_05.txt\r\ntreasury_bulletin_1939_06.txt\r\ntreasury_bulletin_1939_07.txt\r\ntreasury_bulletin_1939_08.txt\r\ntreasury_bulletin_1939_09.txt\r\ntreasury_bulletin_1939_10.txt\r\ntreasury_bulletin_1939_11.txt\r\ntreasury_bulletin_1939_12.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/january-1939-6518?page=81\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/february-1939-6505?page=111\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/march-1939-6519?page=117\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/april-1939-6506?page=95\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/may-1939-6507?page=109\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/june-1939-6508?page=117\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/july-1939-6509?page=109\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/august-1939-6510?page=117\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/september-1939-6511?page=66&deep=true \r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/october-1939-6520?page=68\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/november-1939-6512?page=70\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/december-1939-6513?page=72",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0072",
|
||||
"uid": "UID0072",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_2011_12.txt\r\ntreasury_bulletin_2016_12.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/december-2011-7149?page=58\r\n\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/december-2016-535293?page=57",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0086",
|
||||
"uid": "UID0086",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_2022_12.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/december-2022-627778?page=86",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0091",
|
||||
"uid": "UID0091",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1940_12.txt\r\ntreasury_bulletin_1941_12.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/december-1940-6528?page=21\r\n\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/december-1941-6540?page=64",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0109",
|
||||
"uid": "UID0109",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_2015_12.txt\r\ntreasury_bulletin_2020_12.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/december-2015-519209?page=21\r\n \r\n https://fraser.stlouisfed.org/title/treasury-bulletin-407/december-2020-598551?page=24",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0142",
|
||||
"uid": "UID0142",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1944_03.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/march-1944-6567?page=93",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0154",
|
||||
"uid": "UID0154",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1977_03.txt\r\ntreasury_bulletin_1978_03.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/march-1977-6963?page=83\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/march-1978-6975?page=84",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0159",
|
||||
"uid": "UID0159",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_2000_09.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/september-2000-7103?page=109",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0161",
|
||||
"uid": "UID0161",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1980_03.txt\r\ntreasury_bulletin_1985_03.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/march-1980-6999?page=88\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/march-1985-7040?page=48",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0170",
|
||||
"uid": "UID0170",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1960_03.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/march-1960-6759?page=64",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0190",
|
||||
"uid": "UID0190",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1939_10.txt\r\ntreasury_bulletin_1939_11.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/october-1939-6520?page=14\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/november-1939-6512?page=14",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0213",
|
||||
"uid": "UID0213",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1947_04.txt\r\ntreasury_bulletin_1948_04.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/april-1947-6603?page=28\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/april-1948-6615?page=18",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0217",
|
||||
"uid": "UID0217",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1963_10.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/october-1963-6802?page=15",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0220",
|
||||
"uid": "UID0220",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1939_02.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/february-1939-6505?page=25",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0233",
|
||||
"uid": "UID0233",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1948_04.txt\r\ntreasury_bulletin_1958_04.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/april-1948-6615?page=42\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/april-1958-6735?page=54",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0234",
|
||||
"uid": "UID0234",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1958_01.txt\r\ntreasury_bulletin_1958_02.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/january-1958-6732?page=28\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/february-1958-6733?page=32",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0235",
|
||||
"uid": "UID0235",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1948_04.txt\r\ntreasury_bulletin_1948_05.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/april-1948-6615?page=27\r\nhttps://fraser.stlouisfed.org/title/treasury-bulletin-407/may-1948-6616?page=27",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0239",
|
||||
"uid": "UID0239",
|
||||
"category": "easy",
|
||||
"source_files": "treasury_bulletin_1953_01.txt\r\ntreasury_bulletin_1954_01.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/january-1953-6672?page=62\r\nhttp://fraser.stlouisfed.org/title/treasury-bulletin-407/january-1954-6684?page=51",
|
||||
"source_split": "val"
|
||||
},
|
||||
{
|
||||
"id": "UID0240",
|
||||
"uid": "UID0240",
|
||||
"category": "hard",
|
||||
"source_files": "treasury_bulletin_1957_12.txt",
|
||||
"source_docs": "https://fraser.stlouisfed.org/title/treasury-bulletin-407/december-1957-6731?page=26",
|
||||
"source_split": "val"
|
||||
}
|
||||
]
|
||||
@@ -0,0 +1,21 @@
|
||||
{
|
||||
"benchmark": "SearchQA",
|
||||
"manifest_type": "id_split",
|
||||
"source_repo": "lucadiliello/searchqa",
|
||||
"source_repo_type": "dataset",
|
||||
"source_url": "https://huggingface.co/datasets/lucadiliello/searchqa",
|
||||
"source_id_field": "key",
|
||||
"counts": {
|
||||
"train": 400,
|
||||
"val": 200,
|
||||
"test": 1400
|
||||
},
|
||||
"item_fields": [
|
||||
"id"
|
||||
],
|
||||
"notes": [
|
||||
"This is a split manifest, not the full SearchQA payload.",
|
||||
"Materialize full split items from lucadiliello/searchqa before evaluation.",
|
||||
"The IDs in items.json exactly match the key field in lucadiliello/searchqa."
|
||||
]
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,602 @@
|
||||
[
|
||||
{
|
||||
"id": "1758dc50625e46ee814e44a6061f091d"
|
||||
},
|
||||
{
|
||||
"id": "52967b20d9564d5b94756bc0b3d113d6"
|
||||
},
|
||||
{
|
||||
"id": "e77dbfc0d45e476b9ed5b2204b6d9fc8"
|
||||
},
|
||||
{
|
||||
"id": "675f01b5343e46599626104968c0f45a"
|
||||
},
|
||||
{
|
||||
"id": "dccd0f04cd3b41bc8c4182e41a413b3e"
|
||||
},
|
||||
{
|
||||
"id": "efde5820a29949488feaa18ce7e429f9"
|
||||
},
|
||||
{
|
||||
"id": "98773558971e490d8de41088e41a3207"
|
||||
},
|
||||
{
|
||||
"id": "70e3e1e1d41f4795ba45678d5f4b6524"
|
||||
},
|
||||
{
|
||||
"id": "206a9ffb8ed24686a24db1492c9f10dc"
|
||||
},
|
||||
{
|
||||
"id": "e91f1225cdb84ee7a2d773b5177945c1"
|
||||
},
|
||||
{
|
||||
"id": "b68a14a7b43842a0a50a3965d3bbffbb"
|
||||
},
|
||||
{
|
||||
"id": "e2e679989b544e94bf8dbc3466e49003"
|
||||
},
|
||||
{
|
||||
"id": "c51c26c07b384bbe865976b81a86736d"
|
||||
},
|
||||
{
|
||||
"id": "83928d21f3504843aa6294ce9d9c8da1"
|
||||
},
|
||||
{
|
||||
"id": "f822cde36d864c508b426f3e2b5245b2"
|
||||
},
|
||||
{
|
||||
"id": "f03757766e034c539bf2595428c03626"
|
||||
},
|
||||
{
|
||||
"id": "349be67be2e141438287d21cb6a94131"
|
||||
},
|
||||
{
|
||||
"id": "7f3e18272c6d43dd8c6385c4758b685e"
|
||||
},
|
||||
{
|
||||
"id": "66dc8aee931142b1834da63363511491"
|
||||
},
|
||||
{
|
||||
"id": "52ac78645e144b4598958d67920617dc"
|
||||
},
|
||||
{
|
||||
"id": "32a66c67a8d7490b8572387d8f1d762a"
|
||||
},
|
||||
{
|
||||
"id": "b028bea99d2641268532cd2a32d4ad2d"
|
||||
},
|
||||
{
|
||||
"id": "c9f3a971abe64853aca3e0bb1fea6962"
|
||||
},
|
||||
{
|
||||
"id": "5324041c25d04316a800ccba2d79a2c7"
|
||||
},
|
||||
{
|
||||
"id": "9960a3f1ac954282a7ea1d7f7c8a86f9"
|
||||
},
|
||||
{
|
||||
"id": "03bd8fc86a9d465e9370dcdaedb3c422"
|
||||
},
|
||||
{
|
||||
"id": "a1d0d03fcac54483a1e88542b767afd5"
|
||||
},
|
||||
{
|
||||
"id": "c21a41745ab9402da8b1deea67055a8d"
|
||||
},
|
||||
{
|
||||
"id": "d5c48bd600824c8a9328df80a8baf964"
|
||||
},
|
||||
{
|
||||
"id": "746806ac0b1c48cb80d7ab5ad0a63517"
|
||||
},
|
||||
{
|
||||
"id": "e3b96fc1b99247ebada09a84fa5fd205"
|
||||
},
|
||||
{
|
||||
"id": "bbeeec20466a4dab9e4d73d86c964cba"
|
||||
},
|
||||
{
|
||||
"id": "2b85b097d130404daeb26a3349bdc297"
|
||||
},
|
||||
{
|
||||
"id": "fbaae95260474ed58ccbb05ac482089b"
|
||||
},
|
||||
{
|
||||
"id": "28b837ac304845bd961a46cedc9246af"
|
||||
},
|
||||
{
|
||||
"id": "691e3a24e5994c749ff11c6bc43b2d1c"
|
||||
},
|
||||
{
|
||||
"id": "12d5b51c793a43b1aa3b7e1bf5f4bc73"
|
||||
},
|
||||
{
|
||||
"id": "a3a3892296c74fc4b23b2475156901b1"
|
||||
},
|
||||
{
|
||||
"id": "8065f2e8639848f5b2f16a4f2877b7c2"
|
||||
},
|
||||
{
|
||||
"id": "b06e0d2be0a5465ea12ff2bf2b27d387"
|
||||
},
|
||||
{
|
||||
"id": "e50328dc70634863ac39580ef6f8dd5a"
|
||||
},
|
||||
{
|
||||
"id": "17b441db9b444435951d859e6e081f3b"
|
||||
},
|
||||
{
|
||||
"id": "0fc37a4bd8424ecf8eede56e956edb3a"
|
||||
},
|
||||
{
|
||||
"id": "09c769c7262e43eb95a52e874042c02a"
|
||||
},
|
||||
{
|
||||
"id": "2417ce75564b4b8eab1ddd755b10d962"
|
||||
},
|
||||
{
|
||||
"id": "8f4beaa8a7704eba93f753a0e7040165"
|
||||
},
|
||||
{
|
||||
"id": "6d8e99531264435badb67bf6ac209075"
|
||||
},
|
||||
{
|
||||
"id": "940ddc06648f4c9bac6b3bf5df402e5a"
|
||||
},
|
||||
{
|
||||
"id": "fa0fcf36e99644d0b51be95869d87adc"
|
||||
},
|
||||
{
|
||||
"id": "1882f20b777146608474ac0056d54a5e"
|
||||
},
|
||||
{
|
||||
"id": "439e908aa7c7422daf5632af7e0440ec"
|
||||
},
|
||||
{
|
||||
"id": "2d90c36ab1dd4724b1c108b2507316e4"
|
||||
},
|
||||
{
|
||||
"id": "efa5cc2f750c421a807bc4315db0dc1e"
|
||||
},
|
||||
{
|
||||
"id": "3b1c596e1720499ebe790d14f34a67af"
|
||||
},
|
||||
{
|
||||
"id": "c87234a3a1074b5ea5500114ad9aa16d"
|
||||
},
|
||||
{
|
||||
"id": "9d08ada791844519b7a4bb247d50c89c"
|
||||
},
|
||||
{
|
||||
"id": "74c45e35e5a740798fbd528ff08a70e3"
|
||||
},
|
||||
{
|
||||
"id": "650b683d39dd45c99ad7435fc7d33fe3"
|
||||
},
|
||||
{
|
||||
"id": "191ce59c97e143f7b4c33400913d3570"
|
||||
},
|
||||
{
|
||||
"id": "74441a0f7dae42ca9ee19d8be8c9b318"
|
||||
},
|
||||
{
|
||||
"id": "baef3cef84f14ed98655ac73293a56db"
|
||||
},
|
||||
{
|
||||
"id": "099a1ae148184bd5a93c9bc9889306e7"
|
||||
},
|
||||
{
|
||||
"id": "4fb4332a7a0e44e4b5655de1d7c2e5ac"
|
||||
},
|
||||
{
|
||||
"id": "cdf86d1982f64cb49d6595680b1c53b8"
|
||||
},
|
||||
{
|
||||
"id": "fa4cbba4ffed4f4e8ed4c148711cbc78"
|
||||
},
|
||||
{
|
||||
"id": "41c691d3b7e94ffd993647c7e531e94e"
|
||||
},
|
||||
{
|
||||
"id": "df43ea868a6d409ea67eb632fc9f285c"
|
||||
},
|
||||
{
|
||||
"id": "746448dcf441410db997a59f7137a2c8"
|
||||
},
|
||||
{
|
||||
"id": "3d0cad58c3c145f3b98d51c06a65a0c9"
|
||||
},
|
||||
{
|
||||
"id": "9cb4317831024d589622e2a3cf544360"
|
||||
},
|
||||
{
|
||||
"id": "802015dbc9f5406d8df3b468497f6553"
|
||||
},
|
||||
{
|
||||
"id": "5463f1d3fc124cadb2ad016076a699f5"
|
||||
},
|
||||
{
|
||||
"id": "8c77419b4e864778b965753f411bb469"
|
||||
},
|
||||
{
|
||||
"id": "2e112a7c9d97482fa0deb079c6c09172"
|
||||
},
|
||||
{
|
||||
"id": "f0c26f4f2e9f4d1d84b4f13cb201cbeb"
|
||||
},
|
||||
{
|
||||
"id": "231d7d9955d54f60a1f1ff4f0bf1c5b6"
|
||||
},
|
||||
{
|
||||
"id": "9b56a920d67a48e8b33932802d17d374"
|
||||
},
|
||||
{
|
||||
"id": "092a8aab433c4184bde9c8aa08789464"
|
||||
},
|
||||
{
|
||||
"id": "266a4337fa5742c8bbcc5aa79e571e85"
|
||||
},
|
||||
{
|
||||
"id": "d67cecf38f954800a7c651cca70bdc66"
|
||||
},
|
||||
{
|
||||
"id": "e634fa14bd4941ef96e1d21e796a9cd3"
|
||||
},
|
||||
{
|
||||
"id": "dbc29aef324e4c4f9f19082dcb1b5cb1"
|
||||
},
|
||||
{
|
||||
"id": "e688c5081a84488694b111a11a50ef7e"
|
||||
},
|
||||
{
|
||||
"id": "80b114d03dbb4598b5da648bbe174d0f"
|
||||
},
|
||||
{
|
||||
"id": "ce42d55ff5874c6380e018896b16e128"
|
||||
},
|
||||
{
|
||||
"id": "ef2b37212c98453390c9dca445d69de1"
|
||||
},
|
||||
{
|
||||
"id": "a66639dbb5ed4dfb9e65fe359ed8dc03"
|
||||
},
|
||||
{
|
||||
"id": "10eacaf72a4e4ad49cc43d3a710ee24a"
|
||||
},
|
||||
{
|
||||
"id": "b7e3d45f77bd4a4db768505fc4726592"
|
||||
},
|
||||
{
|
||||
"id": "0d1592f640f04ecb9c7add7164f87c5a"
|
||||
},
|
||||
{
|
||||
"id": "296a678ada814b09bdd0a53629f7f3f9"
|
||||
},
|
||||
{
|
||||
"id": "6e6fd6d62ce54031a4e36f88c2d5b63d"
|
||||
},
|
||||
{
|
||||
"id": "8d0dd71f94cc4abba255f3cdd6c2d940"
|
||||
},
|
||||
{
|
||||
"id": "7c26258ff99c4adeb4b9b35001005453"
|
||||
},
|
||||
{
|
||||
"id": "d13c24c99c704e31bbb3aaa585884d56"
|
||||
},
|
||||
{
|
||||
"id": "dbfb81f100dc474f8267e6f9ab0ae51c"
|
||||
},
|
||||
{
|
||||
"id": "3b4a5a2f617b43c8aa37b0890f576749"
|
||||
},
|
||||
{
|
||||
"id": "c49a8bf2d0ad4afd9b0bb0d03c1b9d96"
|
||||
},
|
||||
{
|
||||
"id": "fbfb4851dd494e44b9a21475d7e6dbaa"
|
||||
},
|
||||
{
|
||||
"id": "e0ba941f8e2c4fa6a0dea2b562b0e4be"
|
||||
},
|
||||
{
|
||||
"id": "7c9ec2b1cd8d4a8eaa6e1154e6fb40c4"
|
||||
},
|
||||
{
|
||||
"id": "4ae645d2721c4751b1e71c83870bbe16"
|
||||
},
|
||||
{
|
||||
"id": "c7bd32aee95f485d85d0b6749d542c51"
|
||||
},
|
||||
{
|
||||
"id": "20e140b44e0f4a6484e9cf776b09679b"
|
||||
},
|
||||
{
|
||||
"id": "5deb6bfcdd8c486d9f5f630309968763"
|
||||
},
|
||||
{
|
||||
"id": "18d1fcc3ed394b23ada799aad3beca90"
|
||||
},
|
||||
{
|
||||
"id": "74a655bcaead4857b92a49ec5f93ee23"
|
||||
},
|
||||
{
|
||||
"id": "bfec094ae9da4aaebd91713b39d3008e"
|
||||
},
|
||||
{
|
||||
"id": "80ebe13e36eb413a9acaace84648b36d"
|
||||
},
|
||||
{
|
||||
"id": "982e414b371248788c72732975949957"
|
||||
},
|
||||
{
|
||||
"id": "963ad06f3e314d3b99e0b6b9eaeaf2d8"
|
||||
},
|
||||
{
|
||||
"id": "90b6478bd7ae4145976c474dc42da80f"
|
||||
},
|
||||
{
|
||||
"id": "ea475dcd08d54da6896f56e330fa2a1c"
|
||||
},
|
||||
{
|
||||
"id": "64899cac91e5423fa3617bf576e6df4d"
|
||||
},
|
||||
{
|
||||
"id": "6231d57d040143e3bc56cafaabaa0975"
|
||||
},
|
||||
{
|
||||
"id": "3fdb66ab2b0a49318b053edda6633237"
|
||||
},
|
||||
{
|
||||
"id": "7dff080db1b8438e9ed737df1d7c32e3"
|
||||
},
|
||||
{
|
||||
"id": "878108ec7e9742a090175bd33390e2c5"
|
||||
},
|
||||
{
|
||||
"id": "3b2acfd62bdf4361aa4e51895d12df8b"
|
||||
},
|
||||
{
|
||||
"id": "0335847f4a44451db5a1b775cc019765"
|
||||
},
|
||||
{
|
||||
"id": "b6bfa34339c649398ae3b4540ed95fcf"
|
||||
},
|
||||
{
|
||||
"id": "82fd2c61b9a64b5f8971b28a1cfd8c1b"
|
||||
},
|
||||
{
|
||||
"id": "a9385b35a3a5492283aa461b93fdb361"
|
||||
},
|
||||
{
|
||||
"id": "847f117d4638487d92539d52eb07a999"
|
||||
},
|
||||
{
|
||||
"id": "de103586389245e7858b26b9b9a1ba57"
|
||||
},
|
||||
{
|
||||
"id": "82d95f6b3df64bb48668f7e98ddf5b9c"
|
||||
},
|
||||
{
|
||||
"id": "b1f6907b8f0f41ec82dc2eae7f84666b"
|
||||
},
|
||||
{
|
||||
"id": "1f232b1de1f34793b6557c8688bf4d9c"
|
||||
},
|
||||
{
|
||||
"id": "c16f718bd6554e998a95fe1d07623a13"
|
||||
},
|
||||
{
|
||||
"id": "dc2add9004fd40e69fcbd2eea80c513c"
|
||||
},
|
||||
{
|
||||
"id": "464e204e258340b69eca2b90451f06b8"
|
||||
},
|
||||
{
|
||||
"id": "c12fd75ea3db450bbfcabeb4f62f7e8f"
|
||||
},
|
||||
{
|
||||
"id": "5dbcadd728fa4956a4f739c1168fbf18"
|
||||
},
|
||||
{
|
||||
"id": "1730220302104d54ba2bb781ca219e52"
|
||||
},
|
||||
{
|
||||
"id": "ffcf25fe8cce4198bd483f2faaa17889"
|
||||
},
|
||||
{
|
||||
"id": "98ed5d627d0746e0aecb42eae39c14dc"
|
||||
},
|
||||
{
|
||||
"id": "30718db44c7f44a5b4b0c68842b1053b"
|
||||
},
|
||||
{
|
||||
"id": "4e90a93f4548458bb5d35ac9c767233b"
|
||||
},
|
||||
{
|
||||
"id": "dabccc32402a43a7ab9c82408f3603a1"
|
||||
},
|
||||
{
|
||||
"id": "80a8a1ea275944228e71fd0dfc74fb74"
|
||||
},
|
||||
{
|
||||
"id": "687434d75186430896ea636fec156e6d"
|
||||
},
|
||||
{
|
||||
"id": "28e5233352ba44ea9f7a8d4cf7f6b86a"
|
||||
},
|
||||
{
|
||||
"id": "b11499ee2a1f4a2b8fa8c9f561136430"
|
||||
},
|
||||
{
|
||||
"id": "6efcba5974464dcbaf5d9c5e9322bd9c"
|
||||
},
|
||||
{
|
||||
"id": "a1d460152c3840d98ed3bc45bf80d5c1"
|
||||
},
|
||||
{
|
||||
"id": "0d482f0e297544dcb2ce1e87e4c3b4b4"
|
||||
},
|
||||
{
|
||||
"id": "1a5baeb7a51d453d892ba70a9de2146d"
|
||||
},
|
||||
{
|
||||
"id": "ab84169c66b2415180563269295f53be"
|
||||
},
|
||||
{
|
||||
"id": "08c0e9678f1548aebeada27ab7b954e1"
|
||||
},
|
||||
{
|
||||
"id": "efd2d367aeec4cc4ae72cd3eaa040a0f"
|
||||
},
|
||||
{
|
||||
"id": "ae54626131ed42eaa55755f10ee64a76"
|
||||
},
|
||||
{
|
||||
"id": "1787cbe3e160427ca4e61931daf66e51"
|
||||
},
|
||||
{
|
||||
"id": "24143f01e7dd455e86c34a4a6f4ed3b9"
|
||||
},
|
||||
{
|
||||
"id": "5e92b869dd0b4ecbaf8694041891eb75"
|
||||
},
|
||||
{
|
||||
"id": "8d130f744b214055ad0c4cdce459a4b9"
|
||||
},
|
||||
{
|
||||
"id": "87913f6be2004e938ba60cda550bc675"
|
||||
},
|
||||
{
|
||||
"id": "68c7f892cbd94fc59ebe4ef1e1ad0c89"
|
||||
},
|
||||
{
|
||||
"id": "f7bf36628c7540da907dda84b48aac82"
|
||||
},
|
||||
{
|
||||
"id": "357af030b55046fc856e03fdac36032a"
|
||||
},
|
||||
{
|
||||
"id": "3a90f09909344dfdbb1a73f80c983a87"
|
||||
},
|
||||
{
|
||||
"id": "44fd76ce5f864cc2aef99c75cfd4904e"
|
||||
},
|
||||
{
|
||||
"id": "95e54b962ce1420d82c716ff428d84fe"
|
||||
},
|
||||
{
|
||||
"id": "8e22ba3b6fc34886a86d60dd1a1db025"
|
||||
},
|
||||
{
|
||||
"id": "162917ea12b24d0e874540feecf9b445"
|
||||
},
|
||||
{
|
||||
"id": "c8e87f5da41c47c7a3116bd1ed0cc34d"
|
||||
},
|
||||
{
|
||||
"id": "a3458bd73564491585df515468475a66"
|
||||
},
|
||||
{
|
||||
"id": "e27a4c1fb25c48809feb3cc438b5eb3c"
|
||||
},
|
||||
{
|
||||
"id": "f36338c39e2846959491ea48cbce07f6"
|
||||
},
|
||||
{
|
||||
"id": "eae01f064523415c95e57ff1e8dadc43"
|
||||
},
|
||||
{
|
||||
"id": "8b845d50f998467a898367107377c64b"
|
||||
},
|
||||
{
|
||||
"id": "a50ae9e7cb6a4f1e8dad07a9bbace810"
|
||||
},
|
||||
{
|
||||
"id": "5c2009e4f1384b27b4b4fe9422f048ed"
|
||||
},
|
||||
{
|
||||
"id": "0baaa0fe4fde4bae99bc6d2be230d525"
|
||||
},
|
||||
{
|
||||
"id": "d2563a0ca702457497995a6959ff1c61"
|
||||
},
|
||||
{
|
||||
"id": "b07871d2991d47e899cfa68a91f16b52"
|
||||
},
|
||||
{
|
||||
"id": "b745738e5b3c482aa2ae94ad5688246b"
|
||||
},
|
||||
{
|
||||
"id": "d66f1f0158054f638ff1b8acf3e18a8a"
|
||||
},
|
||||
{
|
||||
"id": "ede0705d606d4435a300e06beab2322f"
|
||||
},
|
||||
{
|
||||
"id": "e9d415f928644b59b5ed19b7011f627d"
|
||||
},
|
||||
{
|
||||
"id": "1a78a8cf476c4a0dbf551558da1d879b"
|
||||
},
|
||||
{
|
||||
"id": "ea7c66d5dae241eda155763541749cd1"
|
||||
},
|
||||
{
|
||||
"id": "78ae271549144680a42139693d1117e9"
|
||||
},
|
||||
{
|
||||
"id": "7eb9a69a23b8469b9f20f1915abff3be"
|
||||
},
|
||||
{
|
||||
"id": "db73ca2786684b9e92deb1fcc796effd"
|
||||
},
|
||||
{
|
||||
"id": "a876fb1accf64f4eb1b20b4a0ee089db"
|
||||
},
|
||||
{
|
||||
"id": "71d6dabd553c471b98debe941341a4c0"
|
||||
},
|
||||
{
|
||||
"id": "1c6add1cbd464f34b6e90485780f1c3e"
|
||||
},
|
||||
{
|
||||
"id": "5e9b0b44825e44d89ba0501c3b819685"
|
||||
},
|
||||
{
|
||||
"id": "4cfc841cb81e4c90822cfb7c0ec9f86e"
|
||||
},
|
||||
{
|
||||
"id": "2933df89dd6146c791eac1a85f0d02c8"
|
||||
},
|
||||
{
|
||||
"id": "5af06dca120a46889964109e94fff0a0"
|
||||
},
|
||||
{
|
||||
"id": "77970683cb6444388e5cfdd743784db6"
|
||||
},
|
||||
{
|
||||
"id": "3e40d5d91615427d976a56932be3c955"
|
||||
},
|
||||
{
|
||||
"id": "70bc65e0c17041729a4eabe86ff36d53"
|
||||
},
|
||||
{
|
||||
"id": "b913e55670474b69a70736d314822155"
|
||||
},
|
||||
{
|
||||
"id": "ada9cbe2dd8e48deb55171645e47a24a"
|
||||
},
|
||||
{
|
||||
"id": "fba813b8837249cbb39a1996dc657ef0"
|
||||
},
|
||||
{
|
||||
"id": "f7ecc32a212e4c3b8471ce26c32f3135"
|
||||
},
|
||||
{
|
||||
"id": "015641bf588b43aebba1aee07a8afde0"
|
||||
},
|
||||
{
|
||||
"id": "bcecb67a14f34fc6896bc5d65c58405a"
|
||||
}
|
||||
]
|
||||
@@ -0,0 +1,24 @@
|
||||
{
|
||||
"benchmark": "SpreadsheetBench",
|
||||
"manifest_type": "id_split",
|
||||
"source_repo": "KAKA22/SpreadsheetBench",
|
||||
"source_repo_type": "dataset",
|
||||
"source_url": "https://huggingface.co/datasets/KAKA22/SpreadsheetBench",
|
||||
"source_revision": "ab0b742b0fc95b946f212d80ac7771b5531272e4",
|
||||
"source_file": "spreadsheetbench_verified_400.tar.gz",
|
||||
"source_split_name": "spreadsheetbench_split",
|
||||
"counts": {
|
||||
"train": 80,
|
||||
"val": 40,
|
||||
"test": 280
|
||||
},
|
||||
"item_fields": [
|
||||
"id",
|
||||
"spreadsheet_path",
|
||||
"instruction_type"
|
||||
],
|
||||
"notes": [
|
||||
"This is a split manifest, not the full SpreadsheetBench payload.",
|
||||
"Materialize full task JSON rows plus spreadsheet files from SpreadsheetBench Verified 400 before evaluation."
|
||||
]
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,402 @@
|
||||
[
|
||||
{
|
||||
"id": "32438",
|
||||
"spreadsheet_path": "spreadsheet/32438",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "398-14",
|
||||
"spreadsheet_path": "spreadsheet/398-14",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "47766",
|
||||
"spreadsheet_path": "spreadsheet/47766",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "48365",
|
||||
"spreadsheet_path": "spreadsheet/48365",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "32255",
|
||||
"spreadsheet_path": "spreadsheet/32255",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "10747",
|
||||
"spreadsheet_path": "spreadsheet/10747",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "50916",
|
||||
"spreadsheet_path": "spreadsheet/50916",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "577-40",
|
||||
"spreadsheet_path": "spreadsheet/577-40",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "35742",
|
||||
"spreadsheet_path": "spreadsheet/35742",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "46121",
|
||||
"spreadsheet_path": "spreadsheet/46121",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "51090",
|
||||
"spreadsheet_path": "spreadsheet/51090",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "51249",
|
||||
"spreadsheet_path": "spreadsheet/51249",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "82-30",
|
||||
"spreadsheet_path": "spreadsheet/82-30",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "56274",
|
||||
"spreadsheet_path": "spreadsheet/56274",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "57445",
|
||||
"spreadsheet_path": "spreadsheet/57445",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "46646",
|
||||
"spreadsheet_path": "spreadsheet/46646",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "105-24",
|
||||
"spreadsheet_path": "spreadsheet/105-24",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "6239",
|
||||
"spreadsheet_path": "spreadsheet/6239",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "414-20",
|
||||
"spreadsheet_path": "spreadsheet/414-20",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "165-23",
|
||||
"spreadsheet_path": "spreadsheet/165-23",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "40892",
|
||||
"spreadsheet_path": "spreadsheet/40892",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "48745",
|
||||
"spreadsheet_path": "spreadsheet/48745",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "32612",
|
||||
"spreadsheet_path": "spreadsheet/32612",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "325-44",
|
||||
"spreadsheet_path": "spreadsheet/325-44",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "262-17",
|
||||
"spreadsheet_path": "spreadsheet/262-17",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "141-20",
|
||||
"spreadsheet_path": "spreadsheet/141-20",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "52216",
|
||||
"spreadsheet_path": "spreadsheet/52216",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "22-47",
|
||||
"spreadsheet_path": "spreadsheet/22-47",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "55421",
|
||||
"spreadsheet_path": "spreadsheet/55421",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "56427",
|
||||
"spreadsheet_path": "spreadsheet/56427",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "36097",
|
||||
"spreadsheet_path": "spreadsheet/36097",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "32902",
|
||||
"spreadsheet_path": "spreadsheet/32902",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "32023",
|
||||
"spreadsheet_path": "spreadsheet/32023",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "1818",
|
||||
"spreadsheet_path": "spreadsheet/1818",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "170-13",
|
||||
"spreadsheet_path": "spreadsheet/170-13",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "66-24",
|
||||
"spreadsheet_path": "spreadsheet/66-24",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "58949",
|
||||
"spreadsheet_path": "spreadsheet/58949",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "42354",
|
||||
"spreadsheet_path": "spreadsheet/42354",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "194-19",
|
||||
"spreadsheet_path": "spreadsheet/194-19",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "31915",
|
||||
"spreadsheet_path": "spreadsheet/31915",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "58499",
|
||||
"spreadsheet_path": "spreadsheet/58499",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "45372",
|
||||
"spreadsheet_path": "spreadsheet/45372",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "11842",
|
||||
"spreadsheet_path": "spreadsheet/11842",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "57558",
|
||||
"spreadsheet_path": "spreadsheet/57558",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "472-15",
|
||||
"spreadsheet_path": "spreadsheet/472-15",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "55060",
|
||||
"spreadsheet_path": "spreadsheet/55060",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "31011",
|
||||
"spreadsheet_path": "spreadsheet/31011",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "408-39",
|
||||
"spreadsheet_path": "spreadsheet/408-39",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "54085",
|
||||
"spreadsheet_path": "spreadsheet/54085",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "39903",
|
||||
"spreadsheet_path": "spreadsheet/39903",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "48983",
|
||||
"spreadsheet_path": "spreadsheet/48983",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "108-24",
|
||||
"spreadsheet_path": "spreadsheet/108-24",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "58484",
|
||||
"spreadsheet_path": "spreadsheet/58484",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "118-50",
|
||||
"spreadsheet_path": "spreadsheet/118-50",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "10452",
|
||||
"spreadsheet_path": "spreadsheet/10452",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "39931",
|
||||
"spreadsheet_path": "spreadsheet/39931",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "3413",
|
||||
"spreadsheet_path": "spreadsheet/3413",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "247-24",
|
||||
"spreadsheet_path": "spreadsheet/247-24",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "56786",
|
||||
"spreadsheet_path": "spreadsheet/56786",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "55965",
|
||||
"spreadsheet_path": "spreadsheet/55965",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "379-36",
|
||||
"spreadsheet_path": "spreadsheet/379-36",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "58109",
|
||||
"spreadsheet_path": "spreadsheet/58109",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "433-47",
|
||||
"spreadsheet_path": "spreadsheet/433-47",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "192-22",
|
||||
"spreadsheet_path": "spreadsheet/192-22",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "49333",
|
||||
"spreadsheet_path": "spreadsheet/49333",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "493-18",
|
||||
"spreadsheet_path": "spreadsheet/493-18",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "54638",
|
||||
"spreadsheet_path": "spreadsheet/54638",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "34033",
|
||||
"spreadsheet_path": "spreadsheet/34033",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "30930",
|
||||
"spreadsheet_path": "spreadsheet/30930",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "585-41",
|
||||
"spreadsheet_path": "spreadsheet/585-41",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "32337",
|
||||
"spreadsheet_path": "spreadsheet/32337",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "55427",
|
||||
"spreadsheet_path": "spreadsheet/55427",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "263-1",
|
||||
"spreadsheet_path": "spreadsheet/263-1",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "254-34",
|
||||
"spreadsheet_path": "spreadsheet/254-34",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "57113",
|
||||
"spreadsheet_path": "spreadsheet/57113",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "57743",
|
||||
"spreadsheet_path": "spreadsheet/57743",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "43589",
|
||||
"spreadsheet_path": "spreadsheet/43589",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "250-20",
|
||||
"spreadsheet_path": "spreadsheet/250-20",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "48080",
|
||||
"spreadsheet_path": "spreadsheet/48080",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "370-43",
|
||||
"spreadsheet_path": "spreadsheet/370-43",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
}
|
||||
]
|
||||
@@ -0,0 +1,202 @@
|
||||
[
|
||||
{
|
||||
"id": "45635",
|
||||
"spreadsheet_path": "spreadsheet/45635",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "560-12",
|
||||
"spreadsheet_path": "spreadsheet/560-12",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "55049",
|
||||
"spreadsheet_path": "spreadsheet/55049",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "9569",
|
||||
"spreadsheet_path": "spreadsheet/9569",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "7902",
|
||||
"spreadsheet_path": "spreadsheet/7902",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "227-40",
|
||||
"spreadsheet_path": "spreadsheet/227-40",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "463-17",
|
||||
"spreadsheet_path": "spreadsheet/463-17",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "54144",
|
||||
"spreadsheet_path": "spreadsheet/54144",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "80-42",
|
||||
"spreadsheet_path": "spreadsheet/80-42",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "2768",
|
||||
"spreadsheet_path": "spreadsheet/2768",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "37456",
|
||||
"spreadsheet_path": "spreadsheet/37456",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "12864",
|
||||
"spreadsheet_path": "spreadsheet/12864",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "55979",
|
||||
"spreadsheet_path": "spreadsheet/55979",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "48620",
|
||||
"spreadsheet_path": "spreadsheet/48620",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "48588",
|
||||
"spreadsheet_path": "spreadsheet/48588",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "395-36",
|
||||
"spreadsheet_path": "spreadsheet/395-36",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "382-10",
|
||||
"spreadsheet_path": "spreadsheet/382-10",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "59595",
|
||||
"spreadsheet_path": "spreadsheet/59595",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "53383",
|
||||
"spreadsheet_path": "spreadsheet/53383",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "48921",
|
||||
"spreadsheet_path": "spreadsheet/48921",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "416-15",
|
||||
"spreadsheet_path": "spreadsheet/416-15",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "47798",
|
||||
"spreadsheet_path": "spreadsheet/47798",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "56563",
|
||||
"spreadsheet_path": "spreadsheet/56563",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "46897",
|
||||
"spreadsheet_path": "spreadsheet/46897",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "9726",
|
||||
"spreadsheet_path": "spreadsheet/9726",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "50768",
|
||||
"spreadsheet_path": "spreadsheet/50768",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "51-12",
|
||||
"spreadsheet_path": "spreadsheet/51-12",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "31628",
|
||||
"spreadsheet_path": "spreadsheet/31628",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "39046",
|
||||
"spreadsheet_path": "spreadsheet/39046",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "8942",
|
||||
"spreadsheet_path": "spreadsheet/8942",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "48527",
|
||||
"spreadsheet_path": "spreadsheet/48527",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "59196",
|
||||
"spreadsheet_path": "spreadsheet/59196",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "6698",
|
||||
"spreadsheet_path": "spreadsheet/6698",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "43436",
|
||||
"spreadsheet_path": "spreadsheet/43436",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "38462",
|
||||
"spreadsheet_path": "spreadsheet/38462",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "402-43",
|
||||
"spreadsheet_path": "spreadsheet/402-43",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "267-18",
|
||||
"spreadsheet_path": "spreadsheet/267-18",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "37378",
|
||||
"spreadsheet_path": "spreadsheet/37378",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "53647",
|
||||
"spreadsheet_path": "spreadsheet/53647",
|
||||
"instruction_type": "Cell-Level Manipulation"
|
||||
},
|
||||
{
|
||||
"id": "142-12",
|
||||
"spreadsheet_path": "spreadsheet/142-12",
|
||||
"instruction_type": "Sheet-Level Manipulation"
|
||||
}
|
||||
]
|
||||
@@ -0,0 +1,69 @@
|
||||
# Contributing to SkillOpt
|
||||
|
||||
Thank you for your interest in contributing to SkillOpt! This guide covers how to get started.
|
||||
|
||||
## Development Setup
|
||||
|
||||
```bash
|
||||
git clone https://github.com/microsoft/SkillOpt.git
|
||||
cd SkillOpt
|
||||
pip install -e ".[dev]"
|
||||
```
|
||||
|
||||
## Ways to Contribute
|
||||
|
||||
### 🐛 Bug Reports
|
||||
|
||||
Open an issue with:
|
||||
- Steps to reproduce
|
||||
- Expected vs actual behavior
|
||||
- Config file used (sanitize API keys)
|
||||
- Python version and OS
|
||||
|
||||
### 🔧 New Benchmark
|
||||
|
||||
See [Add a New Benchmark](guide/new-benchmark.md) for the implementation guide.
|
||||
|
||||
**Checklist:**
|
||||
- [ ] Data loader in `skillopt/envs/<benchmark>/dataloader.py`
|
||||
- [ ] Environment adapter in `skillopt/envs/<benchmark>/adapter.py`
|
||||
- [ ] Config file in `configs/<benchmark>/default.yaml`
|
||||
- [ ] Registration in `scripts/train.py` (`_ENV_REGISTRY`)
|
||||
- [ ] Documentation page in `docs/`
|
||||
|
||||
### 🤖 New Model Backend
|
||||
|
||||
See [Add a New Model Backend](guide/new-backend.md) for the implementation guide.
|
||||
|
||||
**Checklist:**
|
||||
- [ ] Backend in `skillopt/model/<backend>.py`
|
||||
- [ ] Registration in `skillopt/model/__init__.py`
|
||||
- [ ] API key entry in `.env.example`
|
||||
- [ ] Documentation update
|
||||
|
||||
### 📝 Documentation
|
||||
|
||||
Documentation is built with MkDocs Material:
|
||||
|
||||
```bash
|
||||
pip install -e ".[docs]"
|
||||
mkdocs serve # Preview at http://localhost:8000
|
||||
```
|
||||
|
||||
## Code Style
|
||||
|
||||
- Follow existing patterns in the codebase
|
||||
- Use type hints for function signatures
|
||||
- Keep docstrings concise
|
||||
|
||||
## Pull Request Process
|
||||
|
||||
1. Fork the repository
|
||||
2. Create a feature branch: `git checkout -b feature/my-benchmark`
|
||||
3. Make your changes
|
||||
4. Test with an existing benchmark config
|
||||
5. Submit a PR with a clear description
|
||||
|
||||
## License
|
||||
|
||||
By contributing, you agree that your contributions will be licensed under the MIT License.
|
||||
@@ -0,0 +1,109 @@
|
||||
# Configuration Guide
|
||||
|
||||
SkillOpt uses YAML configuration files with a hierarchical override system.
|
||||
|
||||
## Config Structure
|
||||
|
||||
```
|
||||
configs/
|
||||
├── _base_/
|
||||
│ └── default.yaml # Global defaults
|
||||
├── searchqa/
|
||||
│ └── default.yaml # SearchQA overrides
|
||||
├── docvqa/
|
||||
│ └── default.yaml # DocVQA overrides
|
||||
└── alfworld/
|
||||
└── default.yaml # ALFWorld overrides
|
||||
```
|
||||
|
||||
Benchmark configs inherit from `_base_/default.yaml` and override specific values.
|
||||
|
||||
## Key Parameters
|
||||
|
||||
### Model
|
||||
|
||||
```yaml
|
||||
model:
|
||||
backend: azure_openai # azure_openai | openai_chat | claude_code_exec | qwen
|
||||
optimizer: gpt-5.5 # Optimizer model (for reflection)
|
||||
target: gpt-5.5 # Target model (for rollout)
|
||||
```
|
||||
|
||||
### Training
|
||||
|
||||
```yaml
|
||||
train:
|
||||
num_epochs: 4 # Number of training epochs
|
||||
batch_size: 40 # Tasks per step (batch size)
|
||||
accumulation: 1 # Gradient accumulation
|
||||
seed: 42
|
||||
```
|
||||
|
||||
### Gradient (Reflection)
|
||||
|
||||
```yaml
|
||||
gradient:
|
||||
minibatch_size: 8 # Reflect minibatch size
|
||||
analyst_workers: 16 # Parallel reflection workers
|
||||
max_analyst_rounds: 3 # Max rounds of analyst reflection
|
||||
failure_only: false # Only reflect on failures
|
||||
```
|
||||
|
||||
### Optimizer
|
||||
|
||||
```yaml
|
||||
optimizer:
|
||||
learning_rate: 4 # Max edits per step (edit budget)
|
||||
min_learning_rate: 2 # Min edits for decay schedulers
|
||||
lr_scheduler: cosine # constant | linear | cosine | autonomous
|
||||
use_slow_update: true # Momentum-like blending at epoch boundary
|
||||
slow_update_samples: 20 # Samples for slow update evaluation
|
||||
use_meta_skill: true # Cross-epoch strategy memory
|
||||
```
|
||||
|
||||
### Evaluation
|
||||
|
||||
```yaml
|
||||
evaluation:
|
||||
use_gate: true # Validation gating (accept/reject updates)
|
||||
eval_test: true # Run test evaluation after training
|
||||
```
|
||||
|
||||
### Environment (Data)
|
||||
|
||||
```yaml
|
||||
env:
|
||||
name: searchqa # Benchmark name
|
||||
split_mode: ratio # ratio | split_dir
|
||||
split_ratio: "2:1:7" # train:val:test ratio
|
||||
data_path: "" # Path to dataset
|
||||
exec_timeout: 120 # Per-task timeout (seconds)
|
||||
```
|
||||
|
||||
## CLI Overrides
|
||||
|
||||
Override any config value from the command line:
|
||||
|
||||
```bash
|
||||
python scripts/train.py \
|
||||
--config configs/searchqa/default.yaml \
|
||||
optimizer.learning_rate=16 \
|
||||
optimizer.lr_scheduler=linear \
|
||||
gradient.analyst_workers=8
|
||||
```
|
||||
|
||||
## Environment Variables
|
||||
|
||||
Model credentials are loaded from environment variables:
|
||||
|
||||
| Variable | Backend | Description |
|
||||
|---|---|---|
|
||||
| `AZURE_OPENAI_ENDPOINT` | azure_openai | Azure resource endpoint |
|
||||
| `AZURE_OPENAI_API_KEY` | azure_openai | Azure API key |
|
||||
| `OPENAI_API_KEY` | openai | OpenAI API key |
|
||||
| `ANTHROPIC_API_KEY` | claude | Anthropic API key |
|
||||
| `QWEN_API_BASE` | qwen | Local Qwen vLLM endpoint |
|
||||
|
||||
## Full Reference
|
||||
|
||||
See [Configuration Reference](../reference/config.md) for the complete parameter list.
|
||||
@@ -0,0 +1,51 @@
|
||||
# Deep Learning ↔ SkillOpt Analogy
|
||||
|
||||
SkillOpt is designed around a core insight: **optimizing natural-language prompts follows the same structure as training neural networks**. This page maps every DL concept to its SkillOpt counterpart.
|
||||
|
||||
## Complete Mapping
|
||||
|
||||
| Deep Learning | SkillOpt | Description |
|
||||
|---|---|---|
|
||||
| **Model weights** | Skill document (Markdown) | The thing being optimized |
|
||||
| **Forward pass** | Rollout | Target executes tasks using current skill |
|
||||
| **Loss function** | Task evaluator | Scores task execution quality |
|
||||
| **Backpropagation** | Reflect | Optimizer analyzes failures → edit patches |
|
||||
| **Gradients** | Edit patches | Proposed changes to the skill |
|
||||
| **Gradient aggregation** | Patch aggregation | Merge similar edits |
|
||||
| **Gradient clipping** | Edit selection | Cap max edits per step |
|
||||
| **Learning rate** | `learning_rate` | Max number of edits applied per step |
|
||||
| **LR scheduler** | `lr_scheduler` | Decay schedule: cosine, linear, constant |
|
||||
| **SGD step** | Skill update | Apply selected patches to document |
|
||||
| **Validation set** | Selection split | Gate checks improvement before accepting |
|
||||
| **Early stopping** | Gate patience | Reject updates that don't improve |
|
||||
| **Training step** | Step | One rollout → reflect → update cycle |
|
||||
| **Epoch** | Epoch | Full pass with slow update + meta memory |
|
||||
| **Momentum** | Slow update | Longitudinal comparison at epoch boundary |
|
||||
| **Meta-learning** | Meta skill | Cross-epoch optimizer strategy memory |
|
||||
| **Batch size** | `batch_size` | Tasks sampled per rollout |
|
||||
| **Data parallelism** | `analyst_workers` | Parallel reflection workers |
|
||||
| **Training set** | Train split | Items used for rollout |
|
||||
| **Test set** | Test split | Held-out final evaluation |
|
||||
| **Warm-up** | (implicit) | High LR early steps explore broadly |
|
||||
| **Checkpointing** | Skill snapshots | Saved after each accepted step |
|
||||
| **Transfer learning** | Seed skill / cross-benchmark init | Start from pre-trained skill |
|
||||
|
||||
## Why This Analogy Matters
|
||||
|
||||
1. **Familiar mental model**: ML practitioners immediately understand how to tune SkillOpt
|
||||
2. **Principled hyperparameter search**: Grid search over `learning_rate` × `lr_scheduler` works just like in DL
|
||||
3. **Proven mechanisms**: Gating ≈ validation-based selection, patience ≈ early stopping, slow update ≈ momentum — all with strong theoretical motivation
|
||||
|
||||
## Hyperparameter Transfer Rules
|
||||
|
||||
From our experiments, these DL intuitions transfer well:
|
||||
|
||||
!!! success "What transfers"
|
||||
- **Cosine schedule > constant** — same as in DL, cosine annealing helps convergence
|
||||
- **Moderate LR (4-16) > very high/low** — too few edits = slow learning, too many = noisy
|
||||
- **Slow update helps** — longitudinal comparison prevents catastrophic forgetting across epochs
|
||||
- **Meta skill memory improves reflection** — optimizer benefits from cross-epoch strategy notes
|
||||
|
||||
!!! warning "What doesn't transfer"
|
||||
- **Batch size ≠ better** — larger rollout batches have diminishing returns due to API costs
|
||||
- **More epochs ≠ better** — skills converge faster than neural networks (2-4 epochs usually enough)
|
||||
@@ -0,0 +1,110 @@
|
||||
# Your First Experiment
|
||||
|
||||
This guide walks through running a complete SkillOpt training on SearchQA.
|
||||
|
||||
## 1. Choose a Benchmark
|
||||
|
||||
SkillOpt includes ready-to-use configs for several benchmarks:
|
||||
|
||||
| Benchmark | Difficulty | Typical Runtime |
|
||||
|---|---|---|
|
||||
| SearchQA | ⭐ Easy | ~30 min |
|
||||
| DocVQA | ⭐⭐ Medium | ~2 hours |
|
||||
| ALFWorld | ⭐⭐⭐ Hard | ~3 hours |
|
||||
|
||||
We'll use **SearchQA** as it's the fastest to complete.
|
||||
|
||||
## 2. Configure
|
||||
|
||||
Review the config file:
|
||||
|
||||
```bash
|
||||
cat configs/searchqa/default.yaml
|
||||
```
|
||||
|
||||
Key parameters (deep learning analogy in parentheses):
|
||||
|
||||
```yaml
|
||||
train:
|
||||
num_epochs: 4 # (epochs)
|
||||
batch_size: 40 # (batch size)
|
||||
|
||||
optimizer:
|
||||
learning_rate: 4 # (max edits per step)
|
||||
lr_scheduler: cosine # (learning rate schedule)
|
||||
use_slow_update: true # (momentum at epoch boundary)
|
||||
use_meta_skill: true # (cross-epoch optimizer memory)
|
||||
|
||||
gradient:
|
||||
analyst_workers: 16 # (parallel reflection workers)
|
||||
|
||||
evaluation:
|
||||
use_gate: true # (validation gating)
|
||||
```
|
||||
|
||||
## 3. Train
|
||||
|
||||
```bash
|
||||
python scripts/train.py --config configs/searchqa/default.yaml
|
||||
```
|
||||
|
||||
You'll see output like:
|
||||
|
||||
```
|
||||
[Step 1/8] Rollout: 20 items, 4 workers...
|
||||
[Step 1/8] Score: 0.65 → Reflect...
|
||||
[Step 1/8] 6 edit patches generated
|
||||
[Step 1/8] Selected 4 edits (lr=8, cosine → 7.7)
|
||||
[Step 1/8] Gate: val score 0.68 > 0.65 ✓ ACCEPT
|
||||
[Step 2/8] ...
|
||||
```
|
||||
|
||||
## 4. Monitor
|
||||
|
||||
Training outputs are saved to `outputs/<benchmark>/<run_id>/`:
|
||||
|
||||
```
|
||||
outputs/searchqa/2024-01-15_10-30-00/
|
||||
├── steps/
|
||||
│ ├── step_0001/
|
||||
│ │ ├── candidate_skill.md
|
||||
│ │ ├── step_record.json
|
||||
│ │ └── trajectory_digest.json
|
||||
│ └── step_0002/
|
||||
├── slow_update/
|
||||
│ └── epoch_02/
|
||||
├── meta_skill/
|
||||
│ └── epoch_02/
|
||||
├── skills/
|
||||
│ └── step_0001.md
|
||||
├── best_skill.md
|
||||
├── history.json
|
||||
└── config.yaml
|
||||
```
|
||||
|
||||
## 5. Evaluate
|
||||
|
||||
Evaluate the best skill on the test split:
|
||||
|
||||
```bash
|
||||
python scripts/eval_only.py \
|
||||
--config configs/searchqa/default.yaml \
|
||||
--skill outputs/searchqa/<run_id>/skills/best_skill.md
|
||||
```
|
||||
|
||||
## WebUI
|
||||
|
||||
Prefer a graphical interface? Launch the WebUI:
|
||||
|
||||
```bash
|
||||
pip install -e ".[webui]"
|
||||
python -m skillopt_webui.app
|
||||
```
|
||||
|
||||
Then open `http://localhost:7860` in your browser to configure parameters and launch training.
|
||||
|
||||
## Next Steps
|
||||
|
||||
- [Understand the training loop](training-loop.md)
|
||||
- [Configuration reference](../reference/config.md)
|
||||
- [Add a new benchmark](new-benchmark.md)
|
||||
@@ -0,0 +1,89 @@
|
||||
# Installation
|
||||
|
||||
## Requirements
|
||||
|
||||
- Python ≥ 3.10
|
||||
- At least one model API key (Azure OpenAI, OpenAI, Anthropic, or local Qwen)
|
||||
|
||||
## Quick Install
|
||||
|
||||
```bash
|
||||
git clone https://github.com/microsoft/SkillOpt.git
|
||||
cd SkillOpt
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
## Optional Dependencies
|
||||
|
||||
Install extras for specific benchmarks or backends:
|
||||
|
||||
=== "ALFWorld"
|
||||
|
||||
```bash
|
||||
pip install -e ".[alfworld]"
|
||||
```
|
||||
|
||||
=== "Claude Backend"
|
||||
|
||||
```bash
|
||||
pip install -e ".[claude]"
|
||||
```
|
||||
|
||||
=== "Qwen (Local)"
|
||||
|
||||
```bash
|
||||
pip install -e ".[qwen]"
|
||||
```
|
||||
|
||||
=== "WebUI"
|
||||
|
||||
```bash
|
||||
pip install -e ".[webui]"
|
||||
```
|
||||
|
||||
=== "Development"
|
||||
|
||||
```bash
|
||||
pip install -e ".[dev]"
|
||||
```
|
||||
|
||||
=== "All"
|
||||
|
||||
```bash
|
||||
pip install -e ".[alfworld,claude,qwen,webui,dev]"
|
||||
```
|
||||
|
||||
## Environment Variables
|
||||
|
||||
Copy the example `.env` file and fill in your credentials:
|
||||
|
||||
```bash
|
||||
cp .env.example .env
|
||||
```
|
||||
|
||||
Edit `.env` with your API keys:
|
||||
|
||||
```ini
|
||||
# Azure OpenAI (default backend)
|
||||
AZURE_OPENAI_ENDPOINT=https://your-resource.openai.azure.com/
|
||||
AZURE_OPENAI_API_KEY=your-key
|
||||
|
||||
# Or use OpenAI directly
|
||||
OPENAI_API_KEY=sk-...
|
||||
|
||||
# Or Anthropic Claude
|
||||
ANTHROPIC_API_KEY=sk-ant-...
|
||||
```
|
||||
|
||||
!!! tip
|
||||
You only need credentials for the backend you plan to use. Azure OpenAI is the default.
|
||||
|
||||
## Verify Installation
|
||||
|
||||
```bash
|
||||
python -c "import skillopt; print('SkillOpt ready!')"
|
||||
```
|
||||
|
||||
## Next Steps
|
||||
|
||||
→ [Run your first experiment](first-experiment.md)
|
||||
@@ -0,0 +1,143 @@
|
||||
# Local Environment Smoke Tests
|
||||
|
||||
This guide describes a lightweight pattern for testing a custom SkillOpt environment before connecting it to expensive model calls or a full benchmark dataset.
|
||||
|
||||
The goal is to validate the training loop plumbing first:
|
||||
|
||||
- config loading
|
||||
- adapter construction
|
||||
- dataloader splits
|
||||
- rollout output shape
|
||||
- reflection patch shape
|
||||
- merge/rank/update control flow
|
||||
- artifact creation under `out_root`
|
||||
|
||||
Once those are stable, you can switch the same environment to real model calls and larger evaluation splits.
|
||||
|
||||
## 1. Add a tiny fixture split
|
||||
|
||||
Start with a handful of deterministic examples that cover the expected pass/fail cases for your environment. Keep them small enough that a single training step can run locally.
|
||||
|
||||
A minimal fixture item usually needs:
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "example-1",
|
||||
"split": "train",
|
||||
"question": "...",
|
||||
"expected": "..."
|
||||
}
|
||||
```
|
||||
|
||||
Use the split names your adapter maps to SkillOpt phases:
|
||||
|
||||
- `train` for optimization rollouts
|
||||
- `val` or `valid_seen` for selection/gating
|
||||
- `test` or `valid_unseen` for final evaluation
|
||||
|
||||
## 2. Support an offline mock mode
|
||||
|
||||
Add a configuration flag such as `mock: true` to your adapter. In mock mode, `rollout()` should return deterministic responses without calling external model APIs.
|
||||
|
||||
This lets you verify the SkillOpt loop with a fast command such as:
|
||||
|
||||
```bash
|
||||
python scripts/train.py \
|
||||
--config configs/myenv/tiny_mock.yaml
|
||||
```
|
||||
|
||||
Mock mode should still write the same artifacts as a real run, for example:
|
||||
|
||||
- `responses.json`
|
||||
- `rollout_results.json`
|
||||
- `ranked_edits.json`
|
||||
- `candidate_skill.md`
|
||||
- `summary.json`
|
||||
|
||||
## 3. Keep the smoke config tiny
|
||||
|
||||
A CI-friendly smoke config should run a single small step:
|
||||
|
||||
```yaml
|
||||
train:
|
||||
num_epochs: 1
|
||||
train_size: 3
|
||||
batch_size: 3
|
||||
|
||||
gradient:
|
||||
minibatch_size: 1
|
||||
merge_batch_size: 2
|
||||
analyst_workers: 1
|
||||
max_analyst_rounds: 1
|
||||
|
||||
optimizer:
|
||||
learning_rate: 1
|
||||
min_learning_rate: 1
|
||||
lr_scheduler: constant
|
||||
skill_update_mode: patch
|
||||
use_slow_update: false
|
||||
|
||||
evaluation:
|
||||
use_gate: true
|
||||
sel_env_num: 2
|
||||
test_env_num: 2
|
||||
eval_test: false
|
||||
|
||||
env:
|
||||
name: myenv
|
||||
out_root: outputs/myenv_tiny_mock
|
||||
mock: true
|
||||
```
|
||||
|
||||
Prefer a mock config that runs without credentials. That makes it useful for contributors and CI.
|
||||
|
||||
## 4. Validate optimizer JSON before returning it
|
||||
|
||||
If your environment or extension asks an LLM to merge or rank skill edits, validate the returned JSON before passing it back into SkillOpt. This avoids silent fallbacks from empty, malformed, or out-of-range responses.
|
||||
|
||||
Useful checks for edit payloads:
|
||||
|
||||
- response is a JSON object
|
||||
- `edits` is a non-empty list
|
||||
- every edit is an object
|
||||
- every edit has an allowed operation
|
||||
- required fields such as `content` or `target` are present for that operation
|
||||
|
||||
Useful checks for ranking payloads:
|
||||
|
||||
- `selected_indices` exists
|
||||
- indices are integers
|
||||
- indices are unique
|
||||
- indices are within the candidate edit range
|
||||
- selected count does not exceed the edit budget
|
||||
|
||||
On failure, retry with a compact prompt that includes the schema error. If retries fail, raise an explicit error instead of silently accepting malformed output.
|
||||
|
||||
## 5. Run progressively stronger checks
|
||||
|
||||
A good development sequence is:
|
||||
|
||||
```bash
|
||||
python -m py_compile scripts/train.py skillopt/envs/myenv/adapter.py
|
||||
python scripts/train.py --config configs/myenv/tiny_mock.yaml
|
||||
python scripts/train.py --config configs/myenv/tiny.yaml
|
||||
```
|
||||
|
||||
For the real tiny run, verify that:
|
||||
|
||||
- the run completes
|
||||
- `summary.json` is written
|
||||
- `ranked_edits.json` contains the expected ranking metadata
|
||||
- any optimizer bridge log marks the response schema as valid
|
||||
- no generated files are written outside `out_root`
|
||||
|
||||
## 6. Keep custom environments isolated
|
||||
|
||||
When adding a custom environment to the registry, avoid side effects for existing benchmarks:
|
||||
|
||||
- lazy-import optional dependencies
|
||||
- install environment-specific hooks only when `cfg["env"]` matches your environment
|
||||
- keep mock behavior behind an explicit config flag
|
||||
- write generated artifacts only under `out_root`
|
||||
|
||||
This makes it easier to review and test a custom integration without affecting the built-in benchmarks.
|
||||
@@ -0,0 +1,130 @@
|
||||
# Add a New Model Backend
|
||||
|
||||
SkillOpt supports multiple LLM backends. This guide shows how to add your own.
|
||||
|
||||
## Backend Architecture
|
||||
|
||||
```
|
||||
skillopt/model/
|
||||
├── base.py # Abstract base class
|
||||
├── azure_openai.py # Azure OpenAI backend
|
||||
├── openai_model.py # Direct OpenAI backend
|
||||
├── claude.py # Anthropic Claude backend
|
||||
├── qwen.py # Local Qwen (vLLM) backend
|
||||
└── your_backend.py # Your new backend
|
||||
```
|
||||
|
||||
## Step 1: Create the Backend
|
||||
|
||||
Create `skillopt/model/your_backend.py`:
|
||||
|
||||
```python
|
||||
from skillopt.model.base import ModelBackend, ModelResponse
|
||||
|
||||
class YourBackend(ModelBackend):
|
||||
"""Your custom model backend."""
|
||||
|
||||
def __init__(self, cfg: dict):
|
||||
super().__init__(cfg)
|
||||
self.model_name = cfg.get('model_name', 'your-default-model')
|
||||
self.api_key = os.environ.get('YOUR_API_KEY', '')
|
||||
self.client = self._init_client()
|
||||
|
||||
def _init_client(self):
|
||||
"""Initialize API client."""
|
||||
# TODO: Set up your API client
|
||||
pass
|
||||
|
||||
async def generate(
|
||||
self,
|
||||
messages: list[dict],
|
||||
temperature: float = 0.7,
|
||||
max_tokens: int = 4096,
|
||||
**kwargs
|
||||
) -> ModelResponse:
|
||||
"""
|
||||
Generate a completion.
|
||||
|
||||
Args:
|
||||
messages: Chat messages [{"role": "...", "content": "..."}]
|
||||
temperature: Sampling temperature
|
||||
max_tokens: Maximum tokens in response
|
||||
|
||||
Returns:
|
||||
ModelResponse with content, usage, and metadata
|
||||
"""
|
||||
response = await self.client.chat(
|
||||
model=self.model_name,
|
||||
messages=messages,
|
||||
temperature=temperature,
|
||||
max_tokens=max_tokens,
|
||||
)
|
||||
|
||||
return ModelResponse(
|
||||
content=response.text,
|
||||
usage={
|
||||
'prompt_tokens': response.usage.input,
|
||||
'completion_tokens': response.usage.output,
|
||||
},
|
||||
model=self.model_name,
|
||||
)
|
||||
|
||||
async def generate_with_tools(
|
||||
self,
|
||||
messages: list[dict],
|
||||
tools: list[dict],
|
||||
**kwargs
|
||||
) -> ModelResponse:
|
||||
"""Generate with tool/function calling support."""
|
||||
# Optional: implement if your model supports tool use
|
||||
raise NotImplementedError("Tool use not supported")
|
||||
```
|
||||
|
||||
## Step 2: Register the Backend
|
||||
|
||||
Add to `skillopt/model/__init__.py`:
|
||||
|
||||
```python
|
||||
from .your_backend import YourBackend
|
||||
|
||||
BACKEND_REGISTRY = {
|
||||
# ... existing backends ...
|
||||
'your_backend': YourBackend,
|
||||
}
|
||||
```
|
||||
|
||||
## Step 3: Configure
|
||||
|
||||
Use your backend in any config:
|
||||
|
||||
```yaml
|
||||
model:
|
||||
backend: your_backend
|
||||
model_name: your-model-id
|
||||
temperature: 0.7
|
||||
max_tokens: 4096
|
||||
```
|
||||
|
||||
Set credentials via environment variable:
|
||||
|
||||
```bash
|
||||
export YOUR_API_KEY="your-key"
|
||||
```
|
||||
|
||||
## Required Interface
|
||||
|
||||
Your backend must implement these methods:
|
||||
|
||||
| Method | Required | Description |
|
||||
|---|---|---|
|
||||
| `generate()` | ✅ | Basic text generation |
|
||||
| `generate_with_tools()` | Optional | Tool/function calling |
|
||||
| `count_tokens()` | Optional | Token counting for context management |
|
||||
|
||||
## Tips
|
||||
|
||||
!!! tip
|
||||
- Test your backend with `python -c "from skillopt.model.your_backend import YourBackend"` first
|
||||
- Use `async` methods for all API calls — SkillOpt uses asyncio throughout
|
||||
- Implement retry logic with exponential backoff for production use
|
||||
- Add your API key to `.env.example` when submitting a PR
|
||||
@@ -0,0 +1,130 @@
|
||||
# Add a New Benchmark
|
||||
|
||||
Extend SkillOpt with your own benchmark in ~100 lines of code.
|
||||
|
||||
## Overview
|
||||
|
||||
To add a benchmark, you need:
|
||||
|
||||
1. **Data Loader** — Subclass `SplitDataLoader` to load your split data
|
||||
2. **Environment Adapter** — Subclass `EnvAdapter` and implement rollout/reflect hooks
|
||||
3. **Config** — YAML configuration file
|
||||
4. **Registration** — Add your adapter to the train script registry
|
||||
|
||||
## Step 1: Create the Benchmark Package
|
||||
|
||||
```bash
|
||||
mkdir -p skillopt/envs/my_benchmark
|
||||
touch skillopt/envs/my_benchmark/__init__.py
|
||||
```
|
||||
|
||||
## Step 2: Implement the Data Loader
|
||||
|
||||
Create `skillopt/envs/my_benchmark/dataloader.py`:
|
||||
|
||||
```python
|
||||
from skillopt.datasets.base import SplitDataLoader
|
||||
|
||||
|
||||
class MyBenchmarkDataLoader(SplitDataLoader):
|
||||
"""Load benchmark items from raw data and/or split directories."""
|
||||
|
||||
def load_raw_items(self, data_path: str) -> list[dict]:
|
||||
# For ratio mode, parse your source dataset from data_path.
|
||||
# Return list[dict] where each item has at least a unique, deterministic "id".
|
||||
return super().load_raw_items(data_path)
|
||||
|
||||
def load_split_items(self, split_path: str) -> list[dict]:
|
||||
# For split_dir mode, parse one split directory.
|
||||
return super().load_split_items(split_path)
|
||||
```
|
||||
|
||||
## Step 3: Implement the Environment Adapter
|
||||
|
||||
Create `skillopt/envs/my_benchmark/adapter.py`:
|
||||
|
||||
```python
|
||||
from skillopt.envs.base import EnvAdapter
|
||||
from skillopt.envs.my_benchmark.dataloader import MyBenchmarkDataLoader
|
||||
|
||||
class MyBenchmarkAdapter(EnvAdapter):
|
||||
def __init__(self, split_dir: str = "", data_path: str = "", **kwargs):
|
||||
self.dataloader = MyBenchmarkDataLoader(split_dir=split_dir, data_path=data_path, **kwargs)
|
||||
|
||||
def setup(self, cfg: dict) -> None:
|
||||
super().setup(cfg)
|
||||
self.dataloader.setup(cfg)
|
||||
|
||||
def get_dataloader(self):
|
||||
return self.dataloader
|
||||
|
||||
def build_train_env(self, batch_size: int, seed: int, **kwargs):
|
||||
return self.dataloader.build_train_batch(batch_size=batch_size, seed=seed, **kwargs).payload
|
||||
|
||||
def build_eval_env(self, env_num: int, split: str, seed: int, **kwargs):
|
||||
return self.dataloader.build_eval_batch(env_num=env_num, split=split, seed=seed, **kwargs).payload
|
||||
|
||||
def rollout(self, env_manager, skill_content: str, out_dir: str, **kwargs) -> list[dict]:
|
||||
# env_manager is the payload returned by build_train_env/build_eval_env
|
||||
# (commonly list[dict] task items).
|
||||
# Run target model on each item and return list[dict].
|
||||
# Required keys per row: "id", "hard" (0/1), "soft" (0.0-1.0)
|
||||
raise NotImplementedError
|
||||
|
||||
def reflect(self, results: list[dict], skill_content: str, out_dir: str, **kwargs) -> list[dict | None]:
|
||||
# Convert failure/success analysis into RawPatch-like dicts.
|
||||
raise NotImplementedError
|
||||
|
||||
def get_task_types(self) -> list[str]:
|
||||
return ["my_benchmark"]
|
||||
```
|
||||
|
||||
## Step 4: Register the Benchmark
|
||||
|
||||
Add your adapter to `_register_builtins()` in `scripts/train.py`:
|
||||
|
||||
```python
|
||||
from skillopt.envs.my_benchmark.adapter import MyBenchmarkAdapter
|
||||
|
||||
_ENV_REGISTRY["my_benchmark"] = MyBenchmarkAdapter
|
||||
```
|
||||
|
||||
## Step 5: Create Config
|
||||
|
||||
Create `configs/my_benchmark/default.yaml`:
|
||||
|
||||
```yaml
|
||||
_base_: ../_base_/default.yaml
|
||||
|
||||
env:
|
||||
name: my_benchmark
|
||||
data_path: data/my_benchmark
|
||||
split_mode: ratio
|
||||
split_ratio: "2:1:7"
|
||||
|
||||
train:
|
||||
num_epochs: 4
|
||||
batch_size: 40
|
||||
|
||||
optimizer:
|
||||
learning_rate: 4
|
||||
lr_scheduler: cosine
|
||||
use_slow_update: true
|
||||
use_meta_skill: true
|
||||
|
||||
gradient:
|
||||
analyst_workers: 16
|
||||
```
|
||||
|
||||
## Step 6: Run
|
||||
|
||||
```bash
|
||||
python scripts/train.py --config configs/my_benchmark/default.yaml
|
||||
```
|
||||
|
||||
## Tips
|
||||
|
||||
!!! tip
|
||||
- Use a small `batch_size` (10-20) for initial testing
|
||||
- Start from `skillopt/envs/_template/` and adapt from there
|
||||
- Use an existing adapter (for example `skillopt/envs/officeqa/adapter.py`) as a concrete reference
|
||||
@@ -0,0 +1,78 @@
|
||||
# Skill Document
|
||||
|
||||
A **skill document** is a Markdown file that serves as the "prompt weights" of your agent. SkillOpt trains this document through iterative optimization.
|
||||
|
||||
## What is a Skill Document?
|
||||
|
||||
A skill document is a structured set of instructions that tells a language model **how** to approach a specific type of task. It's analogous to learned weights in a neural network — encoding task-specific knowledge in natural language rather than floating-point parameters.
|
||||
|
||||
## Structure
|
||||
|
||||
A typical skill document contains:
|
||||
|
||||
```markdown
|
||||
# Task Strategy
|
||||
|
||||
## General Approach
|
||||
- Break complex problems into sub-steps
|
||||
- Always verify intermediate results
|
||||
|
||||
## Common Patterns
|
||||
- When you see X, try approach Y
|
||||
- Avoid Z because it leads to errors
|
||||
|
||||
## Edge Cases
|
||||
- If the input contains A, handle it specially by...
|
||||
- Watch out for B — it requires C
|
||||
|
||||
## Output Format
|
||||
- Always include reasoning before the answer
|
||||
- Format numbers with proper units
|
||||
```
|
||||
|
||||
## How It Evolves
|
||||
|
||||
During training, the skill document is modified by **edit patches**:
|
||||
|
||||
1. **Additions**: New rules or strategies discovered from failed trajectories
|
||||
2. **Modifications**: Refining existing rules that are partially correct
|
||||
3. **Deletions**: Removing rules that consistently lead to errors
|
||||
|
||||
Each edit is validated through the **gate** mechanism before being permanently accepted.
|
||||
|
||||
## Initial Skill
|
||||
|
||||
You can start training with:
|
||||
|
||||
- **Empty skill**: The system learns everything from scratch
|
||||
- **Seed skill**: Provide initial instructions to bootstrap training
|
||||
- **Pre-trained skill**: Transfer a skill from a related benchmark
|
||||
|
||||
Configure the initial skill in your YAML:
|
||||
|
||||
```yaml
|
||||
train:
|
||||
init_skill: "path/to/initial_skill.md" # or omit for empty
|
||||
```
|
||||
|
||||
## Skill Quality Metrics
|
||||
|
||||
Track your skill's evolution through:
|
||||
|
||||
- **Validation score**: Primary metric on the selection split
|
||||
- **Test score**: Final metric on held-out test data
|
||||
- **Skill length**: Total tokens in the document
|
||||
- **Edit acceptance rate**: Fraction of proposed edits that pass gating
|
||||
|
||||
## Best Practices
|
||||
|
||||
!!! tip "Tips for better skills"
|
||||
1. **Start with a seed skill** (`env.skill_init`) if you have domain knowledge — it converges faster
|
||||
2. **Use cosine LR schedule** — aggressive early exploration + careful late refinement
|
||||
3. **Enable slow update** (`use_slow_update: true`) to prevent forgetting across epochs
|
||||
4. **Enable meta skill** (`use_meta_skill: true`) so the optimizer accumulates strategy memory
|
||||
|
||||
## Next Steps
|
||||
|
||||
- [Deep Learning Analogy](dl-analogy.md)
|
||||
- [Configuration Reference](../reference/config.md)
|
||||
@@ -0,0 +1,92 @@
|
||||
# The Training Loop
|
||||
|
||||
SkillOpt's core insight: **optimizing natural-language skill documents follows the same structure as training neural networks**.
|
||||
|
||||
## Overview
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────┐
|
||||
│ Training Loop │
|
||||
│ │
|
||||
│ for epoch in epochs: │
|
||||
│ for step in steps: │
|
||||
│ 1. Rollout — Target executes tasks │
|
||||
│ 2. Reflect — Optimizer analyzes trajectories │
|
||||
│ 3. Aggregate — Hierarchical merge of patches │
|
||||
│ 4. Select — Rank & clip edits (learning rate) │
|
||||
│ 5. Update — Apply patches to skill doc │
|
||||
│ 6. Gate — Validate & accept/reject │
|
||||
│ │
|
||||
│ Epoch Boundary: │
|
||||
│ • Slow Update (longitudinal comparison & guidance) │
|
||||
│ • Meta Skill (cross-epoch strategy memory) │
|
||||
└─────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
## Stage Details
|
||||
|
||||
### 1. Rollout (Forward Pass)
|
||||
|
||||
The **target** model executes tasks using the current skill document as its prompt. Each task produces a trajectory and a score.
|
||||
|
||||
```python
|
||||
# Analogy: forward pass through the network
|
||||
predictions = model(input, skill_document)
|
||||
scores = evaluate(predictions, ground_truth)
|
||||
```
|
||||
|
||||
### 2. Reflect (Backward Pass)
|
||||
|
||||
The **optimizer** model analyzes failed trajectories and produces **edit patches** — structured suggestions for improving the skill document.
|
||||
|
||||
Two modes:
|
||||
|
||||
- **Shallow**: Analyze each trajectory independently
|
||||
- **Deep**: Cross-reference multiple failures to find systemic issues
|
||||
|
||||
```python
|
||||
# Analogy: computing gradients
|
||||
gradients = loss.backward() # → edit patches
|
||||
```
|
||||
|
||||
### 3. Aggregate
|
||||
|
||||
Semantically similar edit patches are merged to avoid redundant edits.
|
||||
|
||||
### 4. Select (Gradient Clipping)
|
||||
|
||||
Edits are ranked by relevance score. The `learning_rate` parameter caps how many edits are applied per step — just like gradient clipping prevents overshooting.
|
||||
|
||||
```python
|
||||
# Analogy: gradient clipping + optimizer step size
|
||||
selected = top_k(edits, k=learning_rate)
|
||||
```
|
||||
|
||||
The `lr_scheduler` adjusts this over training:
|
||||
|
||||
- **cosine**: Start aggressive, taper smoothly
|
||||
- **linear**: Linear decay
|
||||
- **constant**: Fixed rate
|
||||
|
||||
### 5. Update (Parameter Update)
|
||||
|
||||
Selected edits are applied to the skill document, producing a new version.
|
||||
|
||||
### 6. Gate (Validation)
|
||||
|
||||
The updated skill is evaluated on a **selection split** (analogous to a validation set). The update is only accepted if performance improves.
|
||||
|
||||
## Epoch Boundary Mechanisms
|
||||
|
||||
### Slow Update
|
||||
|
||||
At the end of each epoch (starting from epoch 2), the system performs a **longitudinal comparison**: it rolls out both the previous epoch's skill and the current skill on the same samples, categorizes items as improved/regressed/persistent_fail/stable_success, then generates high-level **guidance** that is injected into the skill document. This prevents catastrophic forgetting of earlier improvements.
|
||||
|
||||
### Meta Skill
|
||||
|
||||
A **meta-skill memory** accumulates high-level strategy notes across the entire training run. At the end of each epoch, the optimizer reflects on what changed between epochs and produces a compact memory that is provided as additional context during future reflection steps.
|
||||
|
||||
## Next Steps
|
||||
|
||||
- [Understand Skill Documents](skill-document.md)
|
||||
- [DL ↔ SkillOpt analogy table](dl-analogy.md)
|
||||
+170
@@ -0,0 +1,170 @@
|
||||
---
|
||||
hide:
|
||||
- navigation
|
||||
---
|
||||
|
||||
<div class="hero" markdown>
|
||||
|
||||
# SkillOpt
|
||||
|
||||
### Train Agent Skills Like Neural Networks
|
||||
|
||||
*Optimize natural-language skill documents through iterative rollout, reflection, and gated validation — with epochs, learning rates, and validation gates — without touching model weights.*
|
||||
|
||||
[Get Started :material-rocket-launch:](guide/installation.md){ .md-button .md-button--primary }
|
||||
[View on GitHub :material-github:](https://github.com/microsoft/SkillOpt){ .md-button }
|
||||
|
||||
</div>
|
||||
|
||||
---
|
||||
|
||||
## How It Works
|
||||
|
||||
<div class="pipeline-container" markdown>
|
||||
<div class="pipeline-wrapper">
|
||||
|
||||
<div class="pipeline-stage" id="stage-rollout">
|
||||
<div class="stage-icon">🎯</div>
|
||||
<div class="stage-label">Rollout</div>
|
||||
<div class="stage-desc">Target executes tasks</div>
|
||||
</div>
|
||||
|
||||
<div class="pipeline-arrow"><div class="flow-line"></div></div>
|
||||
|
||||
<div class="pipeline-stage" id="stage-reflect">
|
||||
<div class="stage-icon">🔍</div>
|
||||
<div class="stage-label">Reflect</div>
|
||||
<div class="stage-desc">Optimizer analyzes trajectories</div>
|
||||
</div>
|
||||
|
||||
<div class="pipeline-arrow"><div class="flow-line"></div></div>
|
||||
|
||||
<div class="pipeline-stage" id="stage-aggregate">
|
||||
<div class="stage-icon">🔗</div>
|
||||
<div class="stage-label">Aggregate</div>
|
||||
<div class="stage-desc">Merge edit patches</div>
|
||||
</div>
|
||||
|
||||
<div class="pipeline-arrow"><div class="flow-line"></div></div>
|
||||
|
||||
<div class="pipeline-stage" id="stage-select">
|
||||
<div class="stage-icon">✂️</div>
|
||||
<div class="stage-label">Select</div>
|
||||
<div class="stage-desc">Rank & clip edits</div>
|
||||
</div>
|
||||
|
||||
<div class="pipeline-arrow"><div class="flow-line"></div></div>
|
||||
|
||||
<div class="pipeline-stage" id="stage-update">
|
||||
<div class="stage-icon">📝</div>
|
||||
<div class="stage-label">Update</div>
|
||||
<div class="stage-desc">Apply to skill doc</div>
|
||||
</div>
|
||||
|
||||
<div class="pipeline-arrow"><div class="flow-line"></div></div>
|
||||
|
||||
<div class="pipeline-stage" id="stage-gate">
|
||||
<div class="stage-icon">🚦</div>
|
||||
<div class="stage-label">Gate</div>
|
||||
<div class="stage-desc">Validate & accept</div>
|
||||
</div>
|
||||
|
||||
</div>
|
||||
|
||||
<div class="pipeline-epoch-bar">
|
||||
<div class="epoch-mechanism">🔄 Slow Update</div>
|
||||
<div class="epoch-mechanism">🧠 Meta Skill</div>
|
||||
<div class="epoch-label">Epoch Boundary</div>
|
||||
</div>
|
||||
|
||||
</div>
|
||||
|
||||
---
|
||||
|
||||
## Deep Learning Analogy
|
||||
|
||||
SkillOpt brings the familiar deep-learning training paradigm to agentic prompt optimization:
|
||||
|
||||
| Deep Learning | SkillOpt |
|
||||
|---|---|
|
||||
| Model weights | Skill document (Markdown) |
|
||||
| Forward pass | Rollout (target executes tasks) |
|
||||
| Loss / gradient | Reflect (optimizer produces edit patches) |
|
||||
| Gradient clipping | Edit selection (`learning_rate` = max edits) |
|
||||
| SGD step | Patch application to skill |
|
||||
| Validation set | Gated evaluation on selection split |
|
||||
| LR schedule | `lr_scheduler`: cosine, linear, constant |
|
||||
| Epochs | Multi-epoch with slow update & meta skill memory |
|
||||
|
||||
---
|
||||
|
||||
## Supported Benchmarks
|
||||
|
||||
| Benchmark | Type | Config |
|
||||
|---|---|---|
|
||||
| **DocVQA** | Document QA | `configs/docvqa/` |
|
||||
| **ALFWorld** | Embodied AI | `configs/alfworld/` |
|
||||
| **OfficeQA** | Enterprise QA | `configs/officeqa/` |
|
||||
| **SearchQA** | Open-domain QA | `configs/searchqa/` |
|
||||
| **LiveMathBench** | Math reasoning | `configs/livemathematicianbench/` |
|
||||
| **SWEBench** | Software Engineering | `configs/swebench/` |
|
||||
| + 5 more | Various | See [docs](guide/first-experiment.md) |
|
||||
|
||||
---
|
||||
|
||||
## Quick Example
|
||||
|
||||
```bash
|
||||
# Install
|
||||
pip install -e .
|
||||
|
||||
# Configure credentials
|
||||
export AZURE_OPENAI_ENDPOINT="https://your-resource.openai.azure.com/"
|
||||
export AZURE_OPENAI_API_KEY="your-key"
|
||||
|
||||
# Train on SearchQA
|
||||
python scripts/train.py --config configs/searchqa/default.yaml
|
||||
|
||||
# Evaluate best skill
|
||||
python scripts/eval_only.py \
|
||||
--config configs/searchqa/default.yaml \
|
||||
--skill outputs/best_skill.md
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
<div class="grid cards" markdown>
|
||||
|
||||
- :material-book-open-variant:{ .lg .middle } **Getting Started**
|
||||
|
||||
---
|
||||
|
||||
Install SkillOpt, configure your API keys, and run your first experiment in 5 minutes.
|
||||
|
||||
[:octicons-arrow-right-24: Installation](guide/installation.md)
|
||||
|
||||
- :material-puzzle:{ .lg .middle } **Add a Benchmark**
|
||||
|
||||
---
|
||||
|
||||
Extend SkillOpt with your own benchmark in ~100 lines of code.
|
||||
|
||||
[:octicons-arrow-right-24: Extension Guide](guide/new-benchmark.md)
|
||||
|
||||
- :material-cog:{ .lg .middle } **Configuration**
|
||||
|
||||
---
|
||||
|
||||
Full reference for all hyperparameters with deep learning analogies.
|
||||
|
||||
[:octicons-arrow-right-24: Config Reference](reference/config.md)
|
||||
|
||||
- :material-monitor-dashboard:{ .lg .middle } **WebUI**
|
||||
|
||||
---
|
||||
|
||||
Configure, launch, and monitor training from your browser.
|
||||
|
||||
[:octicons-arrow-right-24: WebUI Guide](guide/first-experiment.md#webui)
|
||||
|
||||
</div>
|
||||
@@ -0,0 +1,75 @@
|
||||
# API Reference
|
||||
|
||||
## Core Classes
|
||||
|
||||
### `EnvAdapter`
|
||||
|
||||
Abstract base class for benchmark environments (`skillopt/envs/base.py`).
|
||||
|
||||
```python
|
||||
class EnvAdapter(ABC):
|
||||
def setup(self, cfg: dict) -> None
|
||||
def get_dataloader(self) -> BaseDataLoader | None
|
||||
def build_train_env(self, batch_size: int, seed: int, **kwargs)
|
||||
def build_eval_env(self, env_num: int, split: str, seed: int, **kwargs)
|
||||
def rollout(self, env_manager, skill_content: str, out_dir: str, **kwargs) -> list[dict]
|
||||
def reflect(self, results: list[dict], skill_content: str, out_dir: str, **kwargs) -> list[dict | None]
|
||||
def get_task_types(self) -> list[str]
|
||||
```
|
||||
|
||||
The rollout contract expects result rows with at least:
|
||||
|
||||
```python
|
||||
{"id": str, "hard": int, "soft": float}
|
||||
```
|
||||
|
||||
### `BaseDataLoader` / `SplitDataLoader`
|
||||
|
||||
Data loader abstractions (`skillopt/datasets/base.py`).
|
||||
|
||||
```python
|
||||
class BaseDataLoader(ABC):
|
||||
def setup(self, cfg: dict) -> None
|
||||
def build_train_batch(self, batch_size: int, seed: int, **kwargs) -> BatchSpec
|
||||
def build_eval_batch(self, env_num: int, split: str, seed: int, **kwargs) -> BatchSpec
|
||||
|
||||
class SplitDataLoader(BaseDataLoader):
|
||||
def load_raw_items(self, data_path: str) -> list[dict]
|
||||
def load_split_items(self, split_path: str) -> list[dict]
|
||||
def get_split_items(self, split: str) -> list[dict]
|
||||
```
|
||||
|
||||
### `BatchSpec`
|
||||
|
||||
Represents one concrete batch request.
|
||||
|
||||
```python
|
||||
@dataclass(slots=True)
|
||||
class BatchSpec:
|
||||
phase: str
|
||||
split: str
|
||||
seed: int
|
||||
batch_size: int
|
||||
payload: object | None = None
|
||||
metadata: dict[str, Any] = field(default_factory=dict)
|
||||
```
|
||||
|
||||
### `RolloutResult` / `RawPatch`
|
||||
|
||||
Typed helpers for stage I/O in `skillopt/types.py`.
|
||||
|
||||
```python
|
||||
@dataclass
|
||||
class RolloutResult:
|
||||
id: str
|
||||
hard: int
|
||||
soft: float
|
||||
# optional benchmark-specific fields
|
||||
|
||||
@dataclass
|
||||
class RawPatch:
|
||||
patch: Patch
|
||||
source_type: Literal["failure", "success"] = "failure"
|
||||
```
|
||||
|
||||
For detailed source code, see the [`skillopt/`](https://github.com/microsoft/SkillOpt/tree/main/skillopt) directory.
|
||||
@@ -0,0 +1,71 @@
|
||||
# CLI Reference
|
||||
|
||||
## Training
|
||||
|
||||
```bash
|
||||
python scripts/train.py --config <config.yaml> [overrides...]
|
||||
```
|
||||
|
||||
### Arguments
|
||||
|
||||
| Argument | Description |
|
||||
|---|---|
|
||||
| `--config` | Path to YAML config file (required) |
|
||||
| `key=value` | Override any config parameter |
|
||||
|
||||
### Examples
|
||||
|
||||
```bash
|
||||
# Basic training
|
||||
python scripts/train.py --config configs/searchqa/default.yaml
|
||||
|
||||
# With overrides
|
||||
python scripts/train.py \
|
||||
--config configs/searchqa/default.yaml \
|
||||
--cfg-options optimizer.learning_rate=16 optimizer.lr_scheduler=linear
|
||||
|
||||
# With custom initial skill
|
||||
python scripts/train.py \
|
||||
--config configs/searchqa/default.yaml \
|
||||
--cfg-options env.skill_init=skills/my_seed.md
|
||||
```
|
||||
|
||||
## Evaluation
|
||||
|
||||
```bash
|
||||
python scripts/eval_only.py --config <config.yaml> --skill <skill.md>
|
||||
```
|
||||
|
||||
### Arguments
|
||||
|
||||
| Argument | Description |
|
||||
|---|---|
|
||||
| `--config` | Path to YAML config file (required) |
|
||||
| `--skill` | Path to skill document to evaluate (required) |
|
||||
| `--split` | Evaluation split: `test` (default), `valid`, `train` |
|
||||
|
||||
### Examples
|
||||
|
||||
```bash
|
||||
# Evaluate best skill on test set
|
||||
python scripts/eval_only.py \
|
||||
--config configs/searchqa/default.yaml \
|
||||
--skill outputs/searchqa/run_001/skills/best_skill.md
|
||||
|
||||
# Evaluate on validation set
|
||||
python scripts/eval_only.py \
|
||||
--config configs/searchqa/default.yaml \
|
||||
--skill outputs/searchqa/run_001/skills/best_skill.md \
|
||||
--split valid
|
||||
```
|
||||
|
||||
## WebUI
|
||||
|
||||
```bash
|
||||
python -m skillopt_webui.app [--port PORT] [--share]
|
||||
```
|
||||
|
||||
| Argument | Default | Description |
|
||||
|---|---|---|
|
||||
| `--port` | 7860 | Port number |
|
||||
| `--share` | false | Create public Gradio link |
|
||||
@@ -0,0 +1,85 @@
|
||||
# Configuration Reference
|
||||
|
||||
Complete reference for all SkillOpt configuration parameters.
|
||||
|
||||
## Model
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|---|---|---|---|
|
||||
| `model.backend` | str | `azure_openai` | Backend: `azure_openai` / `openai_chat` / `claude_code_exec` / `qwen` |
|
||||
| `model.optimizer` | str | `gpt-5.5` | Optimizer model (for reflection & slow update) |
|
||||
| `model.target` | str | `gpt-5.5` | Target model (for rollout execution) |
|
||||
| `model.reasoning_effort` | str | `medium` | Reasoning effort level |
|
||||
| `model.optimizer_backend` | str | `openai_chat` | Optimizer backend: `openai_chat` / `claude_chat` / `qwen_chat` / `minimax_chat` |
|
||||
| `model.target_backend` | str | `openai_chat` | Target backend: chat backends plus execution harnesses |
|
||||
| `model.qwen_chat_base_url` | str | `http://localhost:8000/v1` | Shared Qwen/vLLM OpenAI-compatible endpoint |
|
||||
| `model.qwen_chat_enable_thinking` | bool | `false` | Shared Qwen thinking flag |
|
||||
| `model.optimizer_qwen_chat_base_url` | str | — | Optimizer-specific Qwen/vLLM endpoint; overrides shared `qwen_chat_base_url` |
|
||||
| `model.target_qwen_chat_base_url` | str | — | Target-specific Qwen/vLLM endpoint; overrides shared `qwen_chat_base_url` |
|
||||
|
||||
## Training (`train`)
|
||||
|
||||
| Parameter | Type | Default | DL Analogy | Description |
|
||||
|---|---|---|---|---|
|
||||
| `train.num_epochs` | int | 4 | Epochs | Number of training epochs |
|
||||
| `train.batch_size` | int | 40 | Batch size | Tasks sampled per step |
|
||||
| `train.accumulation` | int | 1 | Gradient accumulation | Accumulation rounds per step |
|
||||
| `train.seed` | int | 42 | Random seed | Reproducibility seed |
|
||||
|
||||
## Gradient / Reflection (`gradient`)
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|---|---|---|---|
|
||||
| `gradient.minibatch_size` | int | 8 | Reflect minibatch size |
|
||||
| `gradient.merge_batch_size` | int | 8 | Patch merge batch size |
|
||||
| `gradient.analyst_workers` | int | 16 | Parallel reflection workers |
|
||||
| `gradient.max_analyst_rounds` | int | 3 | Max rounds of analyst reflection |
|
||||
| `gradient.failure_only` | bool | `false` | Only reflect on failures |
|
||||
|
||||
## Optimizer (`optimizer`)
|
||||
|
||||
| Parameter | Type | Default | DL Analogy | Description |
|
||||
|---|---|---|---|---|
|
||||
| `optimizer.learning_rate` | int | 4 | Learning rate | Max edit patches per step (edit budget) |
|
||||
| `optimizer.min_learning_rate` | int | 2 | Min LR | Min edits for decay schedulers |
|
||||
| `optimizer.lr_scheduler` | str | `cosine` | LR schedule | `constant` / `linear` / `cosine` / `autonomous` |
|
||||
| `optimizer.skill_update_mode` | str | `patch` | — | `patch` / `rewrite_from_suggestions` / `full_rewrite_minibatch` |
|
||||
| `optimizer.use_slow_update` | bool | `true` | Momentum | Epoch-boundary longitudinal comparison & guidance |
|
||||
| `optimizer.slow_update_samples` | int | 20 | — | Samples for slow update evaluation |
|
||||
| `optimizer.use_meta_skill` | bool | `true` | Meta-learning | Cross-epoch optimizer-side strategy memory |
|
||||
| `optimizer.longitudinal_pair_policy` | str | `mixed` | — | `mixed` / `changed` / `unchanged` |
|
||||
|
||||
## Evaluation (`evaluation`)
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|---|---|---|---|
|
||||
| `evaluation.use_gate` | bool | `true` | Enable validation gating (accept/reject updates) |
|
||||
| `evaluation.eval_test` | bool | `true` | Run test evaluation after training |
|
||||
|
||||
## Environment (`env`)
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|---|---|---|---|
|
||||
| `env.name` | str | — | Benchmark name (e.g., `searchqa`, `docvqa`) |
|
||||
| `env.data_path` | str | — | Path to dataset |
|
||||
| `env.skill_init` | str | — | Path to initial seed skill (optional) |
|
||||
| `env.split_mode` | str | `ratio` | `ratio` or `split_dir` |
|
||||
| `env.split_ratio` | str | `2:1:7` | Train:val:test ratio |
|
||||
| `env.exec_timeout` | int | 120 | Per-task timeout in seconds |
|
||||
| `env.out_root` | str | — | Output directory |
|
||||
|
||||
## Azure OpenAI Credentials
|
||||
|
||||
| Variable | Description |
|
||||
|---|---|
|
||||
| `AZURE_OPENAI_ENDPOINT` / `model.azure_openai_endpoint` | Azure resource endpoint |
|
||||
| `AZURE_OPENAI_API_KEY` / `model.azure_openai_api_key` | Azure API key |
|
||||
| `OPENAI_API_KEY` | OpenAI API key (for `openai_chat` backend) |
|
||||
| `ANTHROPIC_API_KEY` | Anthropic API key (for `claude_code_exec` backend) |
|
||||
| `QWEN_CHAT_BASE_URL` | Shared local vLLM endpoint for `qwen_chat` |
|
||||
| `QWEN_CHAT_MODEL` | Shared served model name for `qwen_chat` |
|
||||
| `QWEN_CHAT_API_KEY` | Optional API key for the shared Qwen endpoint |
|
||||
| `OPTIMIZER_QWEN_CHAT_BASE_URL` | Optimizer-specific local vLLM endpoint |
|
||||
| `OPTIMIZER_QWEN_CHAT_MODEL` | Optimizer-specific served model name |
|
||||
| `TARGET_QWEN_CHAT_BASE_URL` | Target-specific local vLLM endpoint |
|
||||
| `TARGET_QWEN_CHAT_MODEL` | Target-specific served model name |
|
||||
+2739
File diff suppressed because it is too large
Load Diff
+78
@@ -0,0 +1,78 @@
|
||||
site_name: SkillOpt Documentation
|
||||
site_url: https://microsoft.github.io/SkillOpt
|
||||
site_description: "SkillOpt: Agentic Skill Optimization via Reflective Training Loops"
|
||||
repo_url: https://github.com/microsoft/SkillOpt
|
||||
repo_name: microsoft/SkillOpt
|
||||
|
||||
theme:
|
||||
name: material
|
||||
palette:
|
||||
- scheme: default
|
||||
primary: indigo
|
||||
accent: deep purple
|
||||
toggle:
|
||||
icon: material/brightness-7
|
||||
name: Switch to dark mode
|
||||
- scheme: slate
|
||||
primary: indigo
|
||||
accent: deep purple
|
||||
toggle:
|
||||
icon: material/brightness-4
|
||||
name: Switch to light mode
|
||||
features:
|
||||
- navigation.instant
|
||||
- navigation.tracking
|
||||
- navigation.sections
|
||||
- navigation.expand
|
||||
- navigation.top
|
||||
- content.code.copy
|
||||
- content.tabs.link
|
||||
- search.suggest
|
||||
- search.highlight
|
||||
icon:
|
||||
repo: fontawesome/brands/github
|
||||
font:
|
||||
text: Inter
|
||||
code: JetBrains Mono
|
||||
|
||||
|
||||
|
||||
nav:
|
||||
- Home: index.md
|
||||
- Getting Started:
|
||||
- Installation: guide/installation.md
|
||||
- First Experiment: guide/first-experiment.md
|
||||
- Configuration: guide/configuration.md
|
||||
- Core Concepts:
|
||||
- Training Loop: guide/training-loop.md
|
||||
- Skill Document: guide/skill-document.md
|
||||
- Deep Learning Analogy: guide/dl-analogy.md
|
||||
- Extension Guides:
|
||||
- Add a New Benchmark: guide/new-benchmark.md
|
||||
- Local Environment Smoke Tests: guide/local-env-smoke.md
|
||||
- Add a New Model Backend: guide/new-backend.md
|
||||
- Reference:
|
||||
- Configuration Reference: reference/config.md
|
||||
- CLI Reference: reference/cli.md
|
||||
- API Reference: reference/api.md
|
||||
- Contributing: contributing.md
|
||||
|
||||
markdown_extensions:
|
||||
- admonition
|
||||
- pymdownx.details
|
||||
- pymdownx.superfences
|
||||
- pymdownx.tabbed:
|
||||
alternate_style: true
|
||||
- pymdownx.highlight:
|
||||
anchor_linenums: true
|
||||
- pymdownx.inlinehilite
|
||||
- pymdownx.emoji:
|
||||
emoji_index: !!python/name:material.extensions.emoji.twemoji
|
||||
emoji_generator: !!python/name:material.extensions.emoji.to_svg
|
||||
- attr_list
|
||||
- md_in_html
|
||||
- toc:
|
||||
permalink: true
|
||||
|
||||
plugins:
|
||||
- search
|
||||
@@ -0,0 +1,75 @@
|
||||
[build-system]
|
||||
requires = ["setuptools>=68.0", "wheel"]
|
||||
build-backend = "setuptools.build_meta"
|
||||
|
||||
[project]
|
||||
name = "skillopt"
|
||||
version = "0.1.0"
|
||||
description = "SkillOpt: Agentic Skill Optimization via Reflective Training Loops"
|
||||
readme = "README.md"
|
||||
license = {text = "MIT"}
|
||||
requires-python = ">=3.10"
|
||||
authors = [
|
||||
{name = "SkillOpt Team"},
|
||||
]
|
||||
keywords = ["agent", "prompt-optimization", "skill-learning", "LLM", "agentic"]
|
||||
classifiers = [
|
||||
"Development Status :: 3 - Alpha",
|
||||
"Intended Audience :: Science/Research",
|
||||
"License :: OSI Approved :: MIT License",
|
||||
"Programming Language :: Python :: 3",
|
||||
"Programming Language :: Python :: 3.10",
|
||||
"Programming Language :: Python :: 3.11",
|
||||
"Programming Language :: Python :: 3.12",
|
||||
"Topic :: Scientific/Engineering :: Artificial Intelligence",
|
||||
]
|
||||
dependencies = [
|
||||
"openai>=1.30.0",
|
||||
"pyyaml>=6.0",
|
||||
"numpy>=1.24.0",
|
||||
"openpyxl>=3.1.0",
|
||||
"azure-identity>=1.15.0",
|
||||
"azure-core>=1.30.0",
|
||||
"httpx>=0.27.0",
|
||||
]
|
||||
|
||||
[project.optional-dependencies]
|
||||
# Benchmark-specific dependencies
|
||||
alfworld = ["alfworld>=0.4.0", "gymnasium>=0.29.0"]
|
||||
# Claude model backend
|
||||
claude = ["claude-agent-sdk>=0.1.0"]
|
||||
# Qwen local model backend (via vLLM)
|
||||
qwen = ["vllm>=0.4.0"]
|
||||
# Documentation site
|
||||
docs = ["mkdocs-material>=9.5.0", "mkdocstrings[python]>=0.24.0"]
|
||||
# WebUI dashboard
|
||||
webui = ["gradio>=4.0.0"]
|
||||
# Development tools
|
||||
dev = ["ruff>=0.4.0", "pytest>=8.0.0"]
|
||||
# All optional dependencies (except docs/dev/webui)
|
||||
all = [
|
||||
"alfworld>=0.4.0",
|
||||
"gymnasium>=0.29.0",
|
||||
"claude-agent-sdk>=0.1.0",
|
||||
]
|
||||
|
||||
[project.scripts]
|
||||
skillopt-train = "scripts.train:main"
|
||||
skillopt-eval = "scripts.eval_only:main"
|
||||
|
||||
[project.urls]
|
||||
Homepage = "https://github.com/microsoft/SkillOpt"
|
||||
Documentation = "https://microsoft.github.io/SkillOpt"
|
||||
Repository = "https://github.com/microsoft/SkillOpt"
|
||||
Issues = "https://github.com/microsoft/SkillOpt/issues"
|
||||
|
||||
[tool.setuptools.packages.find]
|
||||
include = ["skillopt*", "scripts*"]
|
||||
|
||||
[tool.ruff]
|
||||
line-length = 120
|
||||
target-version = "py310"
|
||||
|
||||
[tool.ruff.lint]
|
||||
select = ["E", "F", "I", "W"]
|
||||
ignore = ["E501"]
|
||||
@@ -0,0 +1,25 @@
|
||||
# ── Core ──────────────────────────────────────────
|
||||
openai>=1.30.0
|
||||
pyyaml>=6.0
|
||||
numpy>=1.24.0
|
||||
openpyxl>=3.1.0
|
||||
azure-identity>=1.15.0
|
||||
azure-core>=1.30.0
|
||||
httpx>=0.27.0
|
||||
|
||||
# ── Optional: ALFWorld benchmark ──────────────────
|
||||
# alfworld>=0.4.0
|
||||
# gymnasium>=0.29.0
|
||||
|
||||
# ── Optional: Claude model backend ────────────────
|
||||
# claude-agent-sdk>=0.1.0
|
||||
|
||||
# ── Optional: Qwen local model (via vLLM) ────────
|
||||
# vllm>=0.4.0
|
||||
|
||||
# ── Optional: WebUI dashboard ────────────────────
|
||||
# gradio>=4.0.0
|
||||
|
||||
# ── Optional: Documentation site ─────────────────
|
||||
# mkdocs-material>=9.5.0
|
||||
# mkdocstrings[python]>=0.24.0
|
||||
@@ -0,0 +1,451 @@
|
||||
#!/usr/bin/env python3
|
||||
"""SkillOpt eval-only: run a single skill on a dataset without training.
|
||||
|
||||
Usage
|
||||
-----
|
||||
python scripts/eval_only.py \
|
||||
--config configs/spreadsheetbench/default.yaml \
|
||||
--skill skillopt/envs/spreadsheetbench/skills/initial.md \
|
||||
--split_dir /path/to/split \
|
||||
--out_root outputs/eval_skill0
|
||||
|
||||
All YAML keys can be overridden from the CLI, same as train.py.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import datetime
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
|
||||
_SCRIPT_DIR = os.path.dirname(os.path.abspath(__file__))
|
||||
_PROJECT_ROOT = os.path.dirname(_SCRIPT_DIR)
|
||||
if _PROJECT_ROOT not in sys.path:
|
||||
sys.path.insert(0, _PROJECT_ROOT)
|
||||
|
||||
from skillopt.model import (
|
||||
configure_azure_openai,
|
||||
configure_claude_code_exec,
|
||||
configure_codex_exec,
|
||||
set_reasoning_effort,
|
||||
set_target_backend,
|
||||
set_target_deployment,
|
||||
set_optimizer_backend,
|
||||
set_optimizer_deployment,
|
||||
)
|
||||
from skillopt.model.common import default_model_for_backend, normalize_backend_name
|
||||
|
||||
_OPENAI_DEFAULT_MODEL_SENTINELS = {"gpt-5.4", "gpt-5.5"}
|
||||
from skillopt.utils import compute_score
|
||||
|
||||
|
||||
# ── Reuse registry from train.py ───────────────────────────────────────────
|
||||
|
||||
_ENV_REGISTRY: dict[str, type] = {}
|
||||
|
||||
|
||||
def _register_builtins() -> None:
|
||||
try:
|
||||
from skillopt.envs.alfworld.adapter import ALFWorldAdapter
|
||||
_ENV_REGISTRY["alfworld"] = ALFWorldAdapter
|
||||
except ImportError:
|
||||
pass
|
||||
try:
|
||||
from skillopt.envs.searchqa.adapter import SearchQAAdapter
|
||||
_ENV_REGISTRY["searchqa"] = SearchQAAdapter
|
||||
except ImportError:
|
||||
pass
|
||||
try:
|
||||
from skillopt.envs.livemathematicianbench.adapter import LiveMathematicianBenchAdapter
|
||||
_ENV_REGISTRY["livemathematicianbench"] = LiveMathematicianBenchAdapter
|
||||
except ImportError:
|
||||
pass
|
||||
try:
|
||||
from skillopt.envs.babyvision.adapter import BabyVisionAdapter
|
||||
_ENV_REGISTRY["babyvision"] = BabyVisionAdapter
|
||||
except ImportError:
|
||||
pass
|
||||
try:
|
||||
from skillopt.envs.spreadsheetbench.adapter import SpreadsheetBenchAdapter
|
||||
_ENV_REGISTRY["spreadsheetbench"] = SpreadsheetBenchAdapter
|
||||
except ImportError:
|
||||
pass
|
||||
try:
|
||||
from skillopt.envs.mmrb.adapter import MMRBAdapter
|
||||
_ENV_REGISTRY["mmrb"] = MMRBAdapter
|
||||
except ImportError:
|
||||
pass
|
||||
try:
|
||||
from skillopt.envs.docvqa.adapter import DocVQAAdapter
|
||||
_ENV_REGISTRY["docvqa"] = DocVQAAdapter
|
||||
except ImportError:
|
||||
pass
|
||||
try:
|
||||
from skillopt.envs.mathverse.adapter import MathVerseAdapter
|
||||
_ENV_REGISTRY["mathverse"] = MathVerseAdapter
|
||||
except ImportError:
|
||||
pass
|
||||
try:
|
||||
from skillopt.envs.officeqa.adapter import OfficeQAAdapter
|
||||
_ENV_REGISTRY["officeqa"] = OfficeQAAdapter
|
||||
except ImportError:
|
||||
pass
|
||||
try:
|
||||
from skillopt.envs.sealqa.adapter import SealQAAdapter
|
||||
_ENV_REGISTRY["sealqa"] = SealQAAdapter
|
||||
except ImportError:
|
||||
pass
|
||||
try:
|
||||
from skillopt.envs.swebench.adapter import SWEBenchAdapter
|
||||
_ENV_REGISTRY["swebench"] = SWEBenchAdapter
|
||||
except ImportError:
|
||||
pass
|
||||
|
||||
|
||||
def get_adapter(cfg: dict):
|
||||
_register_builtins()
|
||||
env_name = cfg.get("env", "alfworld")
|
||||
if env_name not in _ENV_REGISTRY:
|
||||
raise ValueError(
|
||||
f"Unknown environment '{env_name}'. "
|
||||
f"Available: {list(_ENV_REGISTRY.keys())}"
|
||||
)
|
||||
adapter_cls = _ENV_REGISTRY[env_name]
|
||||
|
||||
import inspect
|
||||
sig = inspect.signature(adapter_cls.__init__)
|
||||
accepted = set(sig.parameters.keys()) - {"self"}
|
||||
adapter_kwargs = {k: cfg[k] for k in accepted if k in cfg}
|
||||
return adapter_cls(**adapter_kwargs)
|
||||
|
||||
|
||||
# ── CLI ────────────────────────────────────────────────────────────────────
|
||||
|
||||
_BOOL = lambda x: str(x).lower() in ("true", "1", "yes") # noqa: E731
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
p = argparse.ArgumentParser(description="SkillOpt eval-only")
|
||||
p.add_argument("--config", type=str, required=True)
|
||||
p.add_argument("--skill", type=str, required=True,
|
||||
help="Path to skill .md file to evaluate")
|
||||
p.add_argument("--split", type=str, default="all",
|
||||
help="Which split to eval: train/valid_seen/valid_unseen/all (default: all)")
|
||||
p.add_argument("--cfg-options", nargs="+", default=[],
|
||||
help="Override config: section.key=value")
|
||||
# Legacy flat overrides
|
||||
p.add_argument("--env", type=str)
|
||||
p.add_argument("--backend", type=str,
|
||||
choices=["azure_openai", "codex", "codex_exec", "claude", "claude_chat", "claude_code_exec"])
|
||||
p.add_argument("--optimizer_model", type=str)
|
||||
p.add_argument("--target_model", type=str)
|
||||
p.add_argument("--optimizer_backend", type=str)
|
||||
p.add_argument("--target_backend", type=str)
|
||||
p.add_argument("--reasoning_effort", type=str,
|
||||
choices=["", "low", "medium", "high", "xhigh", "max"])
|
||||
p.add_argument("--azure_endpoint", type=str)
|
||||
p.add_argument("--azure_api_version", type=str)
|
||||
p.add_argument("--azure_api_key", type=str)
|
||||
p.add_argument("--azure_openai_endpoint", type=str)
|
||||
p.add_argument("--azure_openai_api_version", type=str)
|
||||
p.add_argument("--azure_openai_api_key", type=str)
|
||||
p.add_argument("--azure_openai_auth_mode", type=str)
|
||||
p.add_argument("--azure_openai_ad_scope", type=str)
|
||||
p.add_argument("--azure_openai_managed_identity_client_id", type=str)
|
||||
p.add_argument("--optimizer_azure_openai_endpoint", type=str)
|
||||
p.add_argument("--optimizer_azure_openai_api_version", type=str)
|
||||
p.add_argument("--optimizer_azure_openai_api_key", type=str)
|
||||
p.add_argument("--optimizer_azure_openai_auth_mode", type=str)
|
||||
p.add_argument("--optimizer_azure_openai_ad_scope", type=str)
|
||||
p.add_argument("--optimizer_azure_openai_managed_identity_client_id", type=str)
|
||||
p.add_argument("--target_azure_openai_endpoint", type=str)
|
||||
p.add_argument("--target_azure_openai_api_version", type=str)
|
||||
p.add_argument("--target_azure_openai_api_key", type=str)
|
||||
p.add_argument("--target_azure_openai_auth_mode", type=str)
|
||||
p.add_argument("--target_azure_openai_ad_scope", type=str)
|
||||
p.add_argument("--target_azure_openai_managed_identity_client_id", type=str)
|
||||
p.add_argument("--codex_exec_path", type=str)
|
||||
p.add_argument("--codex_exec_sandbox", type=str)
|
||||
p.add_argument("--codex_exec_profile", type=str)
|
||||
p.add_argument("--codex_exec_full_auto", type=_BOOL)
|
||||
p.add_argument("--codex_exec_reasoning_effort", type=str)
|
||||
p.add_argument("--codex_exec_use_sdk", type=str)
|
||||
p.add_argument("--codex_exec_network_access", type=_BOOL)
|
||||
p.add_argument("--codex_exec_web_search", type=_BOOL)
|
||||
p.add_argument("--codex_exec_approval_policy", type=str)
|
||||
p.add_argument("--claude_code_exec_path", type=str)
|
||||
p.add_argument("--claude_code_exec_profile", type=str)
|
||||
p.add_argument("--claude_code_exec_use_sdk", type=str)
|
||||
p.add_argument("--claude_code_exec_effort", type=str)
|
||||
p.add_argument("--claude_code_exec_max_thinking_tokens", type=int)
|
||||
p.add_argument("--out_root", type=str)
|
||||
p.add_argument("--data_path", type=str)
|
||||
p.add_argument("--split_mode", type=str,
|
||||
choices=["ratio", "split_dir"])
|
||||
p.add_argument("--split_ratio", type=str)
|
||||
p.add_argument("--split_seed", type=int)
|
||||
p.add_argument("--split_dir", type=str)
|
||||
p.add_argument("--split_output_dir", type=str)
|
||||
p.add_argument("--data_root", type=str)
|
||||
p.add_argument("--max_turns", type=int)
|
||||
p.add_argument("--workers", type=int)
|
||||
p.add_argument("--max_api_workers", type=int)
|
||||
p.add_argument("--seed", type=int)
|
||||
p.add_argument("--test_env_num", type=int)
|
||||
p.add_argument("--mode", type=str,
|
||||
help="SpreadsheetBench: single/multi/react (default comes from config)")
|
||||
return p.parse_args()
|
||||
|
||||
|
||||
def main() -> None:
|
||||
args = parse_args()
|
||||
|
||||
from skillopt.config import load_config as _load, flatten_config, is_structured
|
||||
|
||||
cfg = _load(args.config, overrides=args.cfg_options)
|
||||
structured = is_structured(cfg)
|
||||
|
||||
# Apply legacy --key value overrides
|
||||
cli = {k: v for k, v in vars(args).items()
|
||||
if v is not None and k not in ("config", "skill", "split", "cfg_options")}
|
||||
if cli:
|
||||
if structured:
|
||||
from skillopt.config import apply_overrides
|
||||
_MAP = {
|
||||
"backend": "model.backend",
|
||||
"optimizer_model": "model.optimizer",
|
||||
"target_model": "model.target",
|
||||
"optimizer_backend": "model.optimizer_backend",
|
||||
"target_backend": "model.target_backend",
|
||||
"reasoning_effort": "model.reasoning_effort",
|
||||
"azure_endpoint": "model.azure_endpoint",
|
||||
"azure_api_version": "model.azure_api_version",
|
||||
"azure_api_key": "model.azure_api_key",
|
||||
"azure_openai_endpoint": "model.azure_openai_endpoint",
|
||||
"azure_openai_api_version": "model.azure_openai_api_version",
|
||||
"azure_openai_api_key": "model.azure_openai_api_key",
|
||||
"azure_openai_auth_mode": "model.azure_openai_auth_mode",
|
||||
"azure_openai_ad_scope": "model.azure_openai_ad_scope",
|
||||
"azure_openai_managed_identity_client_id": "model.azure_openai_managed_identity_client_id",
|
||||
"optimizer_azure_openai_endpoint": "model.optimizer_azure_openai_endpoint",
|
||||
"optimizer_azure_openai_api_version": "model.optimizer_azure_openai_api_version",
|
||||
"optimizer_azure_openai_api_key": "model.optimizer_azure_openai_api_key",
|
||||
"optimizer_azure_openai_auth_mode": "model.optimizer_azure_openai_auth_mode",
|
||||
"optimizer_azure_openai_ad_scope": "model.optimizer_azure_openai_ad_scope",
|
||||
"optimizer_azure_openai_managed_identity_client_id": "model.optimizer_azure_openai_managed_identity_client_id",
|
||||
"target_azure_openai_endpoint": "model.target_azure_openai_endpoint",
|
||||
"target_azure_openai_api_version": "model.target_azure_openai_api_version",
|
||||
"target_azure_openai_api_key": "model.target_azure_openai_api_key",
|
||||
"target_azure_openai_auth_mode": "model.target_azure_openai_auth_mode",
|
||||
"target_azure_openai_ad_scope": "model.target_azure_openai_ad_scope",
|
||||
"target_azure_openai_managed_identity_client_id": "model.target_azure_openai_managed_identity_client_id",
|
||||
"codex_exec_path": "model.codex_exec_path",
|
||||
"codex_exec_sandbox": "model.codex_exec_sandbox",
|
||||
"codex_exec_profile": "model.codex_exec_profile",
|
||||
"codex_exec_full_auto": "model.codex_exec_full_auto",
|
||||
"codex_exec_reasoning_effort": "model.codex_exec_reasoning_effort",
|
||||
"codex_exec_use_sdk": "model.codex_exec_use_sdk",
|
||||
"codex_exec_network_access": "model.codex_exec_network_access",
|
||||
"codex_exec_web_search": "model.codex_exec_web_search",
|
||||
"codex_exec_approval_policy": "model.codex_exec_approval_policy",
|
||||
"claude_code_exec_path": "model.claude_code_exec_path",
|
||||
"claude_code_exec_profile": "model.claude_code_exec_profile",
|
||||
"claude_code_exec_use_sdk": "model.claude_code_exec_use_sdk",
|
||||
"claude_code_exec_effort": "model.claude_code_exec_effort",
|
||||
"claude_code_exec_max_thinking_tokens": "model.claude_code_exec_max_thinking_tokens",
|
||||
"seed": "train.seed",
|
||||
"test_env_num": "evaluation.test_env_num",
|
||||
"env": "env.name",
|
||||
"out_root": "env.out_root",
|
||||
}
|
||||
mapped = []
|
||||
for k, v in cli.items():
|
||||
dotted = _MAP.get(k)
|
||||
if dotted:
|
||||
mapped.append(f"{dotted}={v}")
|
||||
else:
|
||||
mapped.append(f"env.{k}={v}")
|
||||
apply_overrides(cfg, mapped)
|
||||
else:
|
||||
cfg.update(cli)
|
||||
|
||||
cfg = flatten_config(cfg) if structured else cfg
|
||||
|
||||
for new_key, old_key in (
|
||||
("azure_openai_endpoint", "azure_endpoint"),
|
||||
("azure_openai_api_version", "azure_api_version"),
|
||||
("azure_openai_api_key", "azure_api_key"),
|
||||
):
|
||||
if cfg.get(new_key) in (None, "") and cfg.get(old_key) not in (None, ""):
|
||||
cfg[new_key] = cfg[old_key]
|
||||
|
||||
explicit_backend = getattr(args, "backend", None)
|
||||
if explicit_backend is None:
|
||||
for option in args.cfg_options or []:
|
||||
key = str(option).split("=", 1)[0].strip()
|
||||
if key == "model.backend":
|
||||
explicit_backend = str(option).split("=", 1)[1].strip()
|
||||
break
|
||||
|
||||
backend = normalize_backend_name(cfg.get("model_backend") or cfg.get("target_backend") or "azure_openai")
|
||||
|
||||
def _has_model_override(dotted_key: str, legacy_key: str) -> bool:
|
||||
if getattr(args, legacy_key, None) is not None:
|
||||
return True
|
||||
for option in args.cfg_options or []:
|
||||
key = str(option).split("=", 1)[0].strip()
|
||||
if key == dotted_key:
|
||||
return True
|
||||
return False
|
||||
|
||||
if explicit_backend is not None:
|
||||
backend = normalize_backend_name(explicit_backend)
|
||||
cfg["model_backend"] = backend
|
||||
if backend in {"claude", "claude_chat"}:
|
||||
cfg.setdefault("optimizer_backend", "claude_chat")
|
||||
cfg.setdefault("target_backend", "claude_chat")
|
||||
elif backend in {"codex", "codex_exec"}:
|
||||
cfg.setdefault("optimizer_backend", "openai_chat")
|
||||
cfg.setdefault("target_backend", "codex_exec")
|
||||
elif backend == "claude_code_exec":
|
||||
cfg.setdefault("optimizer_backend", "openai_chat")
|
||||
cfg.setdefault("target_backend", "claude_code_exec")
|
||||
else:
|
||||
cfg.setdefault("optimizer_backend", "openai_chat")
|
||||
cfg.setdefault("target_backend", "openai_chat")
|
||||
else:
|
||||
cfg.setdefault("optimizer_backend", "openai_chat")
|
||||
cfg.setdefault("target_backend", "openai_chat")
|
||||
|
||||
if cfg.get("optimizer_backend") == "claude_chat":
|
||||
if (
|
||||
str(cfg.get("optimizer_model", "") or "").strip() in _OPENAI_DEFAULT_MODEL_SENTINELS
|
||||
and not _has_model_override("model.optimizer", "optimizer_model")
|
||||
):
|
||||
cfg["optimizer_model"] = default_model_for_backend("claude_chat")
|
||||
if cfg.get("target_backend") == "claude_chat":
|
||||
if (
|
||||
str(cfg.get("target_model", "") or "").strip() in _OPENAI_DEFAULT_MODEL_SENTINELS
|
||||
and not _has_model_override("model.target", "target_model")
|
||||
):
|
||||
cfg["target_model"] = default_model_for_backend("claude_chat")
|
||||
if cfg.get("target_backend") == "claude_code_exec":
|
||||
if (
|
||||
str(cfg.get("target_model", "") or "").strip() in _OPENAI_DEFAULT_MODEL_SENTINELS
|
||||
and not _has_model_override("model.target", "target_model")
|
||||
):
|
||||
cfg["target_model"] = default_model_for_backend("claude_chat")
|
||||
|
||||
if not cfg.get("out_root"):
|
||||
env = cfg.get("env", "unknown")
|
||||
model = cfg.get("target_model", "unknown").replace("/", "-")
|
||||
ts = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
|
||||
cfg["out_root"] = os.path.join("outputs", f"eval_{env}_{model}_{ts}")
|
||||
|
||||
cfg["out_root"] = os.path.abspath(cfg["out_root"])
|
||||
|
||||
out_root = cfg["out_root"]
|
||||
os.makedirs(out_root, exist_ok=True)
|
||||
|
||||
# Load skill
|
||||
skill_path = os.path.abspath(args.skill)
|
||||
with open(skill_path) as f:
|
||||
skill_content = f.read()
|
||||
print(f" [skill] {skill_path} ({len(skill_content)} chars)")
|
||||
|
||||
# Configure models
|
||||
configure_azure_openai(
|
||||
endpoint=(cfg.get("azure_openai_endpoint") or cfg.get("azure_endpoint") or None),
|
||||
api_version=(cfg.get("azure_openai_api_version") or cfg.get("azure_api_version") or None),
|
||||
api_key=(cfg.get("azure_openai_api_key") or cfg.get("azure_api_key") or None),
|
||||
auth_mode=cfg.get("azure_openai_auth_mode") or None,
|
||||
ad_scope=cfg.get("azure_openai_ad_scope") or None,
|
||||
managed_identity_client_id=cfg.get("azure_openai_managed_identity_client_id") or None,
|
||||
optimizer_endpoint=cfg.get("optimizer_azure_openai_endpoint") or None,
|
||||
optimizer_api_version=cfg.get("optimizer_azure_openai_api_version") or None,
|
||||
optimizer_api_key=cfg.get("optimizer_azure_openai_api_key") or None,
|
||||
optimizer_auth_mode=cfg.get("optimizer_azure_openai_auth_mode") or None,
|
||||
optimizer_ad_scope=cfg.get("optimizer_azure_openai_ad_scope") or None,
|
||||
optimizer_managed_identity_client_id=(
|
||||
cfg.get("optimizer_azure_openai_managed_identity_client_id") or None
|
||||
),
|
||||
target_endpoint=cfg.get("target_azure_openai_endpoint") or None,
|
||||
target_api_version=cfg.get("target_azure_openai_api_version") or None,
|
||||
target_api_key=cfg.get("target_azure_openai_api_key") or None,
|
||||
target_auth_mode=cfg.get("target_azure_openai_auth_mode") or None,
|
||||
target_ad_scope=cfg.get("target_azure_openai_ad_scope") or None,
|
||||
target_managed_identity_client_id=(
|
||||
cfg.get("target_azure_openai_managed_identity_client_id") or None
|
||||
),
|
||||
)
|
||||
set_optimizer_backend(cfg.get("optimizer_backend", "openai_chat"))
|
||||
set_target_backend(cfg.get("target_backend", "openai_chat"))
|
||||
set_optimizer_deployment(cfg.get("optimizer_model", default_model_for_backend(backend)))
|
||||
set_target_deployment(cfg.get("target_model", default_model_for_backend(backend)))
|
||||
configure_codex_exec(
|
||||
path=cfg.get("codex_exec_path", "codex"),
|
||||
sandbox=cfg.get("codex_exec_sandbox", "workspace-write"),
|
||||
profile=cfg.get("codex_exec_profile", ""),
|
||||
full_auto=cfg.get("codex_exec_full_auto", False),
|
||||
reasoning_effort=cfg.get("codex_exec_reasoning_effort", "none"),
|
||||
use_sdk=cfg.get("codex_exec_use_sdk", None),
|
||||
network_access=cfg.get("codex_exec_network_access", False),
|
||||
web_search=cfg.get("codex_exec_web_search", False),
|
||||
approval_policy=cfg.get("codex_exec_approval_policy", "never"),
|
||||
)
|
||||
configure_claude_code_exec(
|
||||
path=cfg.get("claude_code_exec_path", "claude"),
|
||||
profile=cfg.get("claude_code_exec_profile", ""),
|
||||
use_sdk=cfg.get("claude_code_exec_use_sdk", None),
|
||||
effort=cfg.get("claude_code_exec_effort", cfg.get("reasoning_effort", "medium")),
|
||||
max_thinking_tokens=cfg.get("claude_code_exec_max_thinking_tokens", 16384),
|
||||
)
|
||||
set_reasoning_effort(cfg.get("reasoning_effort", "") or None)
|
||||
|
||||
# Build adapter
|
||||
adapter = get_adapter(cfg)
|
||||
adapter.setup(cfg)
|
||||
|
||||
seed = cfg.get("seed", 42)
|
||||
split = args.split or "all"
|
||||
|
||||
if split == "all":
|
||||
items = (
|
||||
adapter.build_eval_env(0, "train", seed)
|
||||
+ adapter.build_eval_env(0, "valid_seen", seed)
|
||||
+ adapter.build_eval_env(0, "valid_unseen", seed)
|
||||
)
|
||||
else:
|
||||
env_num = cfg.get("test_env_num", 0)
|
||||
items = adapter.build_eval_env(env_num, split, seed)
|
||||
|
||||
print(f"\n [eval] split={split} items={len(items)}")
|
||||
print(f" [eval] out_root={out_root}")
|
||||
print(f"{'='*60}")
|
||||
|
||||
# Run rollout
|
||||
results = adapter.rollout(items, skill_content, out_root)
|
||||
|
||||
# Score
|
||||
hard, soft = compute_score(results)
|
||||
print(f"\n{'='*60}")
|
||||
print(f" Results: hard={hard:.4f} soft={soft:.4f} (n={len(results)})")
|
||||
print(f"{'='*60}")
|
||||
|
||||
# Save summary
|
||||
summary = {
|
||||
"skill": skill_path,
|
||||
"split": split,
|
||||
"n_items": len(results),
|
||||
"hard": hard,
|
||||
"soft": soft,
|
||||
}
|
||||
with open(os.path.join(out_root, "eval_summary.json"), "w") as f:
|
||||
json.dump(summary, f, indent=2, ensure_ascii=False)
|
||||
|
||||
print(f" Saved to: {out_root}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
Executable
+60
@@ -0,0 +1,60 @@
|
||||
#!/usr/bin/env bash
|
||||
# ──────────────────────────────────────────────────────────────────────────────
|
||||
# SkillOpt — ALFWorld training launch script
|
||||
#
|
||||
# Prerequisites:
|
||||
# pip install -e ".[alfworld]"
|
||||
# pip install alfworld[full] && alfworld-download
|
||||
#
|
||||
# Usage:
|
||||
# bash scripts/run_alfworld.sh
|
||||
# bash scripts/run_alfworld.sh --num_epochs 2 --edit_budget 6
|
||||
# bash scripts/run_alfworld.sh --split_dir /path/to/alfworld_split
|
||||
# ──────────────────────────────────────────────────────────────────────────────
|
||||
set -euo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
PROJECT_ROOT="$(dirname "${SCRIPT_DIR}")"
|
||||
|
||||
export PYTHONPATH="${PROJECT_ROOT}:${PYTHONPATH:-}"
|
||||
|
||||
# ALFWorld data — uses ~/.cache/alfworld by default
|
||||
export ALFWORLD_DATA="${ALFWORLD_DATA:-${HOME}/.cache/alfworld}"
|
||||
|
||||
if [ ! -d "${ALFWORLD_DATA}/json_2.1.1" ]; then
|
||||
echo "ERROR: ALFWorld data not found at ${ALFWORLD_DATA}/json_2.1.1"
|
||||
echo ""
|
||||
echo "To download ALFWorld data, run:"
|
||||
echo " pip install alfworld[full]"
|
||||
echo " alfworld-download"
|
||||
echo ""
|
||||
echo "Or set ALFWORLD_DATA to the directory containing json_2.1.1/"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
OPTIMIZER_MODEL="${OPTIMIZER_MODEL:-gpt-5.5}"
|
||||
TARGET_MODEL="${TARGET_MODEL:-gpt-5.5}"
|
||||
|
||||
TIMESTAMP=$(date +%Y%m%d_%H%M%S)
|
||||
DEFAULT_OUT_ROOT="${PROJECT_ROOT}/outputs/skillopt_alfworld_${TARGET_MODEL}_${TIMESTAMP}"
|
||||
|
||||
echo "============================================================"
|
||||
echo " SkillOpt — ALFWorld Training"
|
||||
echo "============================================================"
|
||||
echo " Optimizer: ${OPTIMIZER_MODEL}"
|
||||
echo " Target: ${TARGET_MODEL}"
|
||||
echo " ALFWORLD_DATA: ${ALFWORLD_DATA}"
|
||||
echo " Output: ${DEFAULT_OUT_ROOT}"
|
||||
echo "============================================================"
|
||||
|
||||
cd "${PROJECT_ROOT}"
|
||||
|
||||
python scripts/train.py \
|
||||
--config configs/alfworld/default.yaml \
|
||||
--optimizer_model "${OPTIMIZER_MODEL}" \
|
||||
--target_model "${TARGET_MODEL}" \
|
||||
--out_root "${DEFAULT_OUT_ROOT}" \
|
||||
"$@"
|
||||
|
||||
echo ""
|
||||
echo "Done! Results saved to: ${DEFAULT_OUT_ROOT}"
|
||||
Executable
+40
@@ -0,0 +1,40 @@
|
||||
#!/usr/bin/env bash
|
||||
# ──────────────────────────────────────────────────────────────────────────────
|
||||
# SkillOpt — SearchQA training launch script
|
||||
#
|
||||
# Usage:
|
||||
# bash scripts/run_searchqa.sh
|
||||
# bash scripts/run_searchqa.sh --num_epochs 2 --edit_budget 6
|
||||
# bash scripts/run_searchqa.sh --split_dir /path/to/searchqa_split
|
||||
# ──────────────────────────────────────────────────────────────────────────────
|
||||
set -euo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
PROJECT_ROOT="$(dirname "${SCRIPT_DIR}")"
|
||||
|
||||
export PYTHONPATH="${PROJECT_ROOT}:${PYTHONPATH:-}"
|
||||
|
||||
OPTIMIZER_MODEL="${OPTIMIZER_MODEL:-gpt-5.5}"
|
||||
TARGET_MODEL="${TARGET_MODEL:-gpt-5.5}"
|
||||
|
||||
TIMESTAMP=$(date +%Y%m%d_%H%M%S)
|
||||
DEFAULT_OUT_ROOT="${PROJECT_ROOT}/outputs/skillopt_searchqa_${TARGET_MODEL}_${TIMESTAMP}"
|
||||
|
||||
echo "============================================================"
|
||||
echo " SkillOpt — SearchQA Training"
|
||||
echo "============================================================"
|
||||
echo " Optimizer: ${OPTIMIZER_MODEL}"
|
||||
echo " Target: ${TARGET_MODEL}"
|
||||
echo "============================================================"
|
||||
|
||||
cd "${PROJECT_ROOT}"
|
||||
|
||||
python scripts/train.py \
|
||||
--config configs/searchqa/default.yaml \
|
||||
--optimizer_model "${OPTIMIZER_MODEL}" \
|
||||
--target_model "${TARGET_MODEL}" \
|
||||
--out_root "${DEFAULT_OUT_ROOT}" \
|
||||
"$@"
|
||||
|
||||
echo ""
|
||||
echo "Done! Results saved to: ${DEFAULT_OUT_ROOT}"
|
||||
Executable
+39
@@ -0,0 +1,39 @@
|
||||
#!/usr/bin/env bash
|
||||
# ──────────────────────────────────────────────────────────────────────────────
|
||||
# SkillOpt — SpreadsheetBench training launch script
|
||||
#
|
||||
# Usage:
|
||||
# bash scripts/run_spreadsheetbench.sh --split_dir /path/to/split --data_root /path/to/data
|
||||
# bash scripts/run_spreadsheetbench.sh --num_epochs 2 --edit_budget 6
|
||||
# ──────────────────────────────────────────────────────────────────────────────
|
||||
set -euo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
PROJECT_ROOT="$(dirname "${SCRIPT_DIR}")"
|
||||
|
||||
export PYTHONPATH="${PROJECT_ROOT}:${PYTHONPATH:-}"
|
||||
|
||||
OPTIMIZER_MODEL="${OPTIMIZER_MODEL:-gpt-5.5}"
|
||||
TARGET_MODEL="${TARGET_MODEL:-gpt-5.5}"
|
||||
|
||||
TIMESTAMP=$(date +%Y%m%d_%H%M%S)
|
||||
DEFAULT_OUT_ROOT="${PROJECT_ROOT}/outputs/skillopt_spreadsheetbench_${TARGET_MODEL}_${TIMESTAMP}"
|
||||
|
||||
echo "============================================================"
|
||||
echo " SkillOpt — SpreadsheetBench Training"
|
||||
echo "============================================================"
|
||||
echo " Optimizer: ${OPTIMIZER_MODEL}"
|
||||
echo " Target: ${TARGET_MODEL}"
|
||||
echo "============================================================"
|
||||
|
||||
cd "${PROJECT_ROOT}"
|
||||
|
||||
python scripts/train.py \
|
||||
--config configs/spreadsheetbench/default.yaml \
|
||||
--optimizer_model "${OPTIMIZER_MODEL}" \
|
||||
--target_model "${TARGET_MODEL}" \
|
||||
--out_root "${DEFAULT_OUT_ROOT}" \
|
||||
"$@"
|
||||
|
||||
echo ""
|
||||
echo "Done! Results saved to: ${DEFAULT_OUT_ROOT}"
|
||||
@@ -0,0 +1,548 @@
|
||||
#!/usr/bin/env python3
|
||||
"""SkillOpt unified training entry point.
|
||||
|
||||
Usage
|
||||
-----
|
||||
python scripts/train.py --config configs/alfworld/default.yaml
|
||||
|
||||
Any YAML key can be overridden from the command line::
|
||||
|
||||
python scripts/train.py --config configs/alfworld/default.yaml \\
|
||||
--batch_size 40 --num_epochs 2 --seed 123
|
||||
|
||||
Run ``python scripts/train.py --help`` for a full list of options.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import datetime
|
||||
import os
|
||||
import sys
|
||||
|
||||
# Ensure the project root is on sys.path so ``import skillopt`` works
|
||||
# regardless of where the script is invoked from.
|
||||
_SCRIPT_DIR = os.path.dirname(os.path.abspath(__file__))
|
||||
_PROJECT_ROOT = os.path.dirname(_SCRIPT_DIR)
|
||||
if _PROJECT_ROOT not in sys.path:
|
||||
sys.path.insert(0, _PROJECT_ROOT)
|
||||
|
||||
from skillopt.model.common import default_model_for_backend, normalize_backend_name
|
||||
|
||||
_OPENAI_DEFAULT_MODEL_SENTINELS = {"gpt-5.4", "gpt-5.5"}
|
||||
|
||||
|
||||
# ── Environment registry ────────────────────────────────────────────────────
|
||||
|
||||
_ENV_REGISTRY: dict[str, type] = {}
|
||||
|
||||
|
||||
def _register_builtins() -> None:
|
||||
"""Lazy-import built-in adapters so we don't pull heavy deps at CLI parse time."""
|
||||
try:
|
||||
from skillopt.envs.alfworld.adapter import ALFWorldAdapter
|
||||
_ENV_REGISTRY["alfworld"] = ALFWorldAdapter
|
||||
except ImportError:
|
||||
pass # ALFWorld deps not installed — skip
|
||||
try:
|
||||
from skillopt.envs.searchqa.adapter import SearchQAAdapter
|
||||
_ENV_REGISTRY["searchqa"] = SearchQAAdapter
|
||||
except ImportError:
|
||||
pass
|
||||
try:
|
||||
from skillopt.envs.livemathematicianbench.adapter import LiveMathematicianBenchAdapter
|
||||
_ENV_REGISTRY["livemathematicianbench"] = LiveMathematicianBenchAdapter
|
||||
except ImportError:
|
||||
pass
|
||||
try:
|
||||
from skillopt.envs.babyvision.adapter import BabyVisionAdapter
|
||||
_ENV_REGISTRY["babyvision"] = BabyVisionAdapter
|
||||
except ImportError:
|
||||
pass
|
||||
try:
|
||||
from skillopt.envs.spreadsheetbench.adapter import SpreadsheetBenchAdapter
|
||||
_ENV_REGISTRY["spreadsheetbench"] = SpreadsheetBenchAdapter
|
||||
except ImportError:
|
||||
pass
|
||||
try:
|
||||
from skillopt.envs.mmrb.adapter import MMRBAdapter
|
||||
_ENV_REGISTRY["mmrb"] = MMRBAdapter
|
||||
except ImportError:
|
||||
pass
|
||||
try:
|
||||
from skillopt.envs.docvqa.adapter import DocVQAAdapter
|
||||
_ENV_REGISTRY["docvqa"] = DocVQAAdapter
|
||||
except ImportError:
|
||||
pass
|
||||
try:
|
||||
from skillopt.envs.mathverse.adapter import MathVerseAdapter
|
||||
_ENV_REGISTRY["mathverse"] = MathVerseAdapter
|
||||
except ImportError:
|
||||
pass
|
||||
try:
|
||||
from skillopt.envs.officeqa.adapter import OfficeQAAdapter
|
||||
_ENV_REGISTRY["officeqa"] = OfficeQAAdapter
|
||||
except ImportError:
|
||||
pass
|
||||
try:
|
||||
from skillopt.envs.sealqa.adapter import SealQAAdapter
|
||||
_ENV_REGISTRY["sealqa"] = SealQAAdapter
|
||||
except ImportError:
|
||||
pass
|
||||
try:
|
||||
from skillopt.envs.swebench.adapter import SWEBenchAdapter
|
||||
_ENV_REGISTRY["swebench"] = SWEBenchAdapter
|
||||
except ImportError:
|
||||
pass
|
||||
|
||||
|
||||
def get_adapter(cfg: dict):
|
||||
"""Instantiate the environment adapter specified in ``cfg["env"]``."""
|
||||
_register_builtins()
|
||||
env_name = cfg.get("env", "alfworld")
|
||||
if env_name not in _ENV_REGISTRY:
|
||||
raise ValueError(
|
||||
f"Unknown environment '{env_name}'. "
|
||||
f"Available: {list(_ENV_REGISTRY.keys())}"
|
||||
)
|
||||
adapter_cls = _ENV_REGISTRY[env_name]
|
||||
|
||||
# Inspect adapter __init__ signature and only pass accepted kwargs
|
||||
import inspect
|
||||
sig = inspect.signature(adapter_cls.__init__)
|
||||
accepted = set(sig.parameters.keys()) - {"self"}
|
||||
adapter_kwargs: dict = {}
|
||||
for key in accepted:
|
||||
if key in cfg:
|
||||
adapter_kwargs[key] = cfg[key]
|
||||
|
||||
return adapter_cls(**adapter_kwargs)
|
||||
|
||||
|
||||
# ── CLI ──────────────────────────────────────────────────────────────────────
|
||||
|
||||
_BOOL = lambda x: x.lower() in ("true", "1", "yes") # noqa: E731
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
p = argparse.ArgumentParser(
|
||||
description="SkillOpt: Executive Strategy for Self-Evolving Agent Skills",
|
||||
formatter_class=argparse.RawDescriptionHelpFormatter,
|
||||
epilog=__doc__,
|
||||
)
|
||||
p.add_argument("--config", type=str, required=True,
|
||||
help="Path to YAML config file")
|
||||
p.add_argument("--cfg-options", nargs="+", default=[],
|
||||
help="Override config: section.key=value (e.g. train.batch_size=40)")
|
||||
|
||||
# Legacy flat CLI overrides (still work, prefer --cfg-options for new usage)
|
||||
p.add_argument("--env", type=str)
|
||||
p.add_argument("--backend", type=str,
|
||||
choices=["azure_openai", "codex", "codex_exec", "claude", "claude_chat", "claude_code_exec", "qwen", "qwen_chat", "minimax", "minimax_chat"])
|
||||
p.add_argument("--optimizer_model", type=str)
|
||||
p.add_argument("--target_model", type=str)
|
||||
p.add_argument("--optimizer_backend", type=str)
|
||||
p.add_argument("--target_backend", type=str)
|
||||
p.add_argument("--reasoning_effort", type=str,
|
||||
choices=["", "low", "medium", "high", "xhigh", "max"])
|
||||
p.add_argument("--rewrite_reasoning_effort", type=str)
|
||||
p.add_argument("--rewrite_max_completion_tokens", type=int)
|
||||
p.add_argument("--azure_endpoint", type=str)
|
||||
p.add_argument("--azure_api_version", type=str)
|
||||
p.add_argument("--azure_api_key", type=str)
|
||||
p.add_argument("--azure_openai_endpoint", type=str)
|
||||
p.add_argument("--azure_openai_api_version", type=str)
|
||||
p.add_argument("--azure_openai_api_key", type=str)
|
||||
p.add_argument("--azure_openai_auth_mode", type=str)
|
||||
p.add_argument("--azure_openai_ad_scope", type=str)
|
||||
p.add_argument("--azure_openai_managed_identity_client_id", type=str)
|
||||
p.add_argument("--optimizer_azure_openai_endpoint", type=str)
|
||||
p.add_argument("--optimizer_azure_openai_api_version", type=str)
|
||||
p.add_argument("--optimizer_azure_openai_api_key", type=str)
|
||||
p.add_argument("--optimizer_azure_openai_auth_mode", type=str)
|
||||
p.add_argument("--optimizer_azure_openai_ad_scope", type=str)
|
||||
p.add_argument("--optimizer_azure_openai_managed_identity_client_id", type=str)
|
||||
p.add_argument("--target_azure_openai_endpoint", type=str)
|
||||
p.add_argument("--target_azure_openai_api_version", type=str)
|
||||
p.add_argument("--target_azure_openai_api_key", type=str)
|
||||
p.add_argument("--target_azure_openai_auth_mode", type=str)
|
||||
p.add_argument("--target_azure_openai_ad_scope", type=str)
|
||||
p.add_argument("--target_azure_openai_managed_identity_client_id", type=str)
|
||||
p.add_argument("--qwen_chat_base_url", type=str)
|
||||
p.add_argument("--qwen_chat_api_key", type=str)
|
||||
p.add_argument("--qwen_chat_temperature", type=float)
|
||||
p.add_argument("--qwen_chat_timeout_seconds", type=float)
|
||||
p.add_argument("--qwen_chat_max_tokens", type=int)
|
||||
p.add_argument("--qwen_chat_enable_thinking", type=_BOOL)
|
||||
p.add_argument("--optimizer_qwen_chat_base_url", type=str)
|
||||
p.add_argument("--optimizer_qwen_chat_api_key", type=str)
|
||||
p.add_argument("--optimizer_qwen_chat_temperature", type=float)
|
||||
p.add_argument("--optimizer_qwen_chat_timeout_seconds", type=float)
|
||||
p.add_argument("--optimizer_qwen_chat_max_tokens", type=int)
|
||||
p.add_argument("--optimizer_qwen_chat_enable_thinking", type=_BOOL)
|
||||
p.add_argument("--target_qwen_chat_base_url", type=str)
|
||||
p.add_argument("--target_qwen_chat_api_key", type=str)
|
||||
p.add_argument("--target_qwen_chat_temperature", type=float)
|
||||
p.add_argument("--target_qwen_chat_timeout_seconds", type=float)
|
||||
p.add_argument("--target_qwen_chat_max_tokens", type=int)
|
||||
p.add_argument("--target_qwen_chat_enable_thinking", type=_BOOL)
|
||||
p.add_argument("--minimax_base_url", type=str)
|
||||
p.add_argument("--minimax_api_key", type=str)
|
||||
p.add_argument("--minimax_model", type=str)
|
||||
p.add_argument("--minimax_temperature", type=float)
|
||||
p.add_argument("--minimax_max_tokens", type=int)
|
||||
p.add_argument("--minimax_enable_thinking", type=_BOOL)
|
||||
p.add_argument("--codex_exec_path", type=str)
|
||||
p.add_argument("--codex_exec_sandbox", type=str)
|
||||
p.add_argument("--codex_exec_profile", type=str)
|
||||
p.add_argument("--codex_exec_full_auto", type=_BOOL)
|
||||
p.add_argument("--codex_exec_reasoning_effort", type=str)
|
||||
p.add_argument("--codex_exec_use_sdk", type=str)
|
||||
p.add_argument("--codex_exec_network_access", type=_BOOL)
|
||||
p.add_argument("--codex_exec_web_search", type=_BOOL)
|
||||
p.add_argument("--codex_exec_approval_policy", type=str)
|
||||
p.add_argument("--claude_code_exec_path", type=str)
|
||||
p.add_argument("--claude_code_exec_profile", type=str)
|
||||
p.add_argument("--claude_code_exec_use_sdk", type=str)
|
||||
p.add_argument("--claude_code_exec_effort", type=str)
|
||||
p.add_argument("--claude_code_exec_max_thinking_tokens", type=int)
|
||||
p.add_argument("--codex_trace_to_optimizer", type=_BOOL)
|
||||
p.add_argument("--skill_init", type=str)
|
||||
p.add_argument("--num_epochs", type=int)
|
||||
p.add_argument("--train_size", type=int)
|
||||
p.add_argument("--steps_per_epoch", type=int)
|
||||
p.add_argument("--batch_size", type=int)
|
||||
p.add_argument("--accumulation", type=int)
|
||||
p.add_argument("--seed", type=int)
|
||||
p.add_argument("--edit_budget", type=int)
|
||||
p.add_argument("--min_edit_budget", type=int)
|
||||
p.add_argument("--lr_scheduler", type=str,
|
||||
choices=["constant", "linear", "cosine", "autonomous"])
|
||||
p.add_argument("--lr_control_mode", type=str,
|
||||
choices=["fixed", "autonomous", "none"])
|
||||
p.add_argument("--merge_batch_size", type=int)
|
||||
p.add_argument("--max_analyst_rounds", type=int)
|
||||
p.add_argument("--sel_env_num", type=int)
|
||||
p.add_argument("--test_env_num", type=int)
|
||||
p.add_argument("--eval_test", type=_BOOL)
|
||||
p.add_argument("--use_gate", type=_BOOL)
|
||||
p.add_argument("--max_steps", type=int)
|
||||
p.add_argument("--max_api_workers", type=int)
|
||||
p.add_argument("--analyst_workers", type=int)
|
||||
p.add_argument("--failure_only", type=_BOOL)
|
||||
p.add_argument("--minibatch_size", type=int)
|
||||
p.add_argument("--skill_update_mode", type=str,
|
||||
choices=[
|
||||
"patch",
|
||||
"rewrite_from_suggestions",
|
||||
"rewrite",
|
||||
"suggestions",
|
||||
"full_rewrite",
|
||||
"full_rewrite_minibatch",
|
||||
"minibatch_full_rewrite",
|
||||
])
|
||||
p.add_argument("--use_slow_update", type=_BOOL)
|
||||
p.add_argument("--slow_update_samples", type=int)
|
||||
p.add_argument("--longitudinal_pair_policy", type=str,
|
||||
choices=["mixed", "changed", "unchanged"])
|
||||
p.add_argument("--use_meta_skill", type=_BOOL)
|
||||
p.add_argument("--data_path", type=str)
|
||||
p.add_argument("--split_mode", type=str,
|
||||
choices=["ratio", "split_dir"])
|
||||
p.add_argument("--split_ratio", type=str)
|
||||
p.add_argument("--split_seed", type=int)
|
||||
p.add_argument("--split_dir", type=str)
|
||||
p.add_argument("--split_output_dir", type=str)
|
||||
p.add_argument("--data_root", type=str)
|
||||
p.add_argument("--max_turns", type=int)
|
||||
p.add_argument("--workers", type=int)
|
||||
p.add_argument("--limit", type=int)
|
||||
p.add_argument("--shuffle_choices", type=_BOOL)
|
||||
p.add_argument("--use_theorem", type=_BOOL)
|
||||
p.add_argument("--use_sketch", type=_BOOL)
|
||||
p.add_argument("--image_detail", type=str)
|
||||
p.add_argument("--judge_model", type=str)
|
||||
p.add_argument("--judge_max_completion_tokens", type=int)
|
||||
p.add_argument("--judge_retries", type=int)
|
||||
p.add_argument("--out_root", type=str)
|
||||
p.add_argument("--mode", type=str)
|
||||
|
||||
return p.parse_args()
|
||||
|
||||
|
||||
# ── Flat key → structured path mapping (for legacy CLI → structured config) ──
|
||||
|
||||
_LEGACY_TO_STRUCTURED: dict[str, str] = {
|
||||
"backend": "model.backend",
|
||||
"optimizer_model": "model.optimizer",
|
||||
"target_model": "model.target",
|
||||
"optimizer_backend": "model.optimizer_backend",
|
||||
"target_backend": "model.target_backend",
|
||||
"reasoning_effort": "model.reasoning_effort",
|
||||
"rewrite_reasoning_effort": "model.rewrite_reasoning_effort",
|
||||
"rewrite_max_completion_tokens": "model.rewrite_max_completion_tokens",
|
||||
"azure_endpoint": "model.azure_endpoint",
|
||||
"azure_api_version": "model.azure_api_version",
|
||||
"azure_api_key": "model.azure_api_key",
|
||||
"azure_openai_endpoint": "model.azure_openai_endpoint",
|
||||
"azure_openai_api_version": "model.azure_openai_api_version",
|
||||
"azure_openai_api_key": "model.azure_openai_api_key",
|
||||
"azure_openai_auth_mode": "model.azure_openai_auth_mode",
|
||||
"azure_openai_ad_scope": "model.azure_openai_ad_scope",
|
||||
"azure_openai_managed_identity_client_id": "model.azure_openai_managed_identity_client_id",
|
||||
"optimizer_azure_openai_endpoint": "model.optimizer_azure_openai_endpoint",
|
||||
"optimizer_azure_openai_api_version": "model.optimizer_azure_openai_api_version",
|
||||
"optimizer_azure_openai_api_key": "model.optimizer_azure_openai_api_key",
|
||||
"optimizer_azure_openai_auth_mode": "model.optimizer_azure_openai_auth_mode",
|
||||
"optimizer_azure_openai_ad_scope": "model.optimizer_azure_openai_ad_scope",
|
||||
"optimizer_azure_openai_managed_identity_client_id": "model.optimizer_azure_openai_managed_identity_client_id",
|
||||
"target_azure_openai_endpoint": "model.target_azure_openai_endpoint",
|
||||
"target_azure_openai_api_version": "model.target_azure_openai_api_version",
|
||||
"target_azure_openai_api_key": "model.target_azure_openai_api_key",
|
||||
"target_azure_openai_auth_mode": "model.target_azure_openai_auth_mode",
|
||||
"target_azure_openai_ad_scope": "model.target_azure_openai_ad_scope",
|
||||
"target_azure_openai_managed_identity_client_id": "model.target_azure_openai_managed_identity_client_id",
|
||||
"qwen_chat_base_url": "model.qwen_chat_base_url",
|
||||
"qwen_chat_api_key": "model.qwen_chat_api_key",
|
||||
"qwen_chat_temperature": "model.qwen_chat_temperature",
|
||||
"qwen_chat_timeout_seconds": "model.qwen_chat_timeout_seconds",
|
||||
"qwen_chat_max_tokens": "model.qwen_chat_max_tokens",
|
||||
"qwen_chat_enable_thinking": "model.qwen_chat_enable_thinking",
|
||||
"optimizer_qwen_chat_base_url": "model.optimizer_qwen_chat_base_url",
|
||||
"optimizer_qwen_chat_api_key": "model.optimizer_qwen_chat_api_key",
|
||||
"optimizer_qwen_chat_temperature": "model.optimizer_qwen_chat_temperature",
|
||||
"optimizer_qwen_chat_timeout_seconds": "model.optimizer_qwen_chat_timeout_seconds",
|
||||
"optimizer_qwen_chat_max_tokens": "model.optimizer_qwen_chat_max_tokens",
|
||||
"optimizer_qwen_chat_enable_thinking": "model.optimizer_qwen_chat_enable_thinking",
|
||||
"target_qwen_chat_base_url": "model.target_qwen_chat_base_url",
|
||||
"target_qwen_chat_api_key": "model.target_qwen_chat_api_key",
|
||||
"target_qwen_chat_temperature": "model.target_qwen_chat_temperature",
|
||||
"target_qwen_chat_timeout_seconds": "model.target_qwen_chat_timeout_seconds",
|
||||
"target_qwen_chat_max_tokens": "model.target_qwen_chat_max_tokens",
|
||||
"target_qwen_chat_enable_thinking": "model.target_qwen_chat_enable_thinking",
|
||||
"minimax_base_url": "model.minimax_base_url",
|
||||
"minimax_api_key": "model.minimax_api_key",
|
||||
"minimax_model": "model.minimax_model",
|
||||
"minimax_temperature": "model.minimax_temperature",
|
||||
"minimax_max_tokens": "model.minimax_max_tokens",
|
||||
"minimax_enable_thinking": "model.minimax_enable_thinking",
|
||||
"codex_exec_path": "model.codex_exec_path",
|
||||
"codex_exec_sandbox": "model.codex_exec_sandbox",
|
||||
"codex_exec_profile": "model.codex_exec_profile",
|
||||
"codex_exec_full_auto": "model.codex_exec_full_auto",
|
||||
"codex_exec_reasoning_effort": "model.codex_exec_reasoning_effort",
|
||||
"codex_exec_use_sdk": "model.codex_exec_use_sdk",
|
||||
"codex_exec_network_access": "model.codex_exec_network_access",
|
||||
"codex_exec_web_search": "model.codex_exec_web_search",
|
||||
"codex_exec_approval_policy": "model.codex_exec_approval_policy",
|
||||
"claude_code_exec_path": "model.claude_code_exec_path",
|
||||
"claude_code_exec_profile": "model.claude_code_exec_profile",
|
||||
"claude_code_exec_use_sdk": "model.claude_code_exec_use_sdk",
|
||||
"claude_code_exec_effort": "model.claude_code_exec_effort",
|
||||
"claude_code_exec_max_thinking_tokens": "model.claude_code_exec_max_thinking_tokens",
|
||||
"codex_trace_to_optimizer": "model.codex_trace_to_optimizer",
|
||||
"num_epochs": "train.num_epochs",
|
||||
"train_size": "train.train_size",
|
||||
"steps_per_epoch": "train.steps_per_epoch",
|
||||
"batch_size": "train.batch_size",
|
||||
"accumulation": "train.accumulation",
|
||||
"seed": "train.seed",
|
||||
"minibatch_size": "gradient.minibatch_size",
|
||||
"merge_batch_size": "gradient.merge_batch_size",
|
||||
"analyst_workers": "gradient.analyst_workers",
|
||||
"max_analyst_rounds": "gradient.max_analyst_rounds",
|
||||
"failure_only": "gradient.failure_only",
|
||||
"edit_budget": "optimizer.learning_rate",
|
||||
"min_edit_budget": "optimizer.min_learning_rate",
|
||||
"lr_scheduler": "optimizer.lr_scheduler",
|
||||
"lr_control_mode": "optimizer.lr_control_mode",
|
||||
"skill_update_mode": "optimizer.skill_update_mode",
|
||||
"use_slow_update": "optimizer.use_slow_update",
|
||||
"slow_update_samples": "optimizer.slow_update_samples",
|
||||
"longitudinal_pair_policy": "optimizer.longitudinal_pair_policy",
|
||||
"use_meta_skill": "optimizer.use_meta_skill",
|
||||
"use_gate": "evaluation.use_gate",
|
||||
"sel_env_num": "evaluation.sel_env_num",
|
||||
"test_env_num": "evaluation.test_env_num",
|
||||
"eval_test": "evaluation.eval_test",
|
||||
"env": "env.name",
|
||||
"skill_init": "env.skill_init",
|
||||
"out_root": "env.out_root",
|
||||
}
|
||||
|
||||
|
||||
def load_config(args: argparse.Namespace) -> dict:
|
||||
"""Load config with _base_ inheritance, then apply CLI overrides."""
|
||||
from skillopt.config import load_config as _load, flatten_config, is_structured
|
||||
|
||||
cfg = _load(args.config, overrides=args.cfg_options)
|
||||
structured = is_structured(cfg)
|
||||
|
||||
# Apply legacy --key value overrides
|
||||
cli = {k: v for k, v in vars(args).items()
|
||||
if v is not None and k not in ("config", "cfg_options")}
|
||||
if cli:
|
||||
if structured:
|
||||
from skillopt.config import apply_overrides
|
||||
mapped = []
|
||||
for k, v in cli.items():
|
||||
dotted = _LEGACY_TO_STRUCTURED.get(k)
|
||||
if dotted:
|
||||
mapped.append(f"{dotted}={v}")
|
||||
else:
|
||||
mapped.append(f"env.{k}={v}")
|
||||
apply_overrides(cfg, mapped)
|
||||
else:
|
||||
cfg.update(cli)
|
||||
|
||||
# Flatten structured config → flat dict for trainer/adapter
|
||||
flat = flatten_config(cfg) if structured else cfg
|
||||
|
||||
for new_key, old_key in (
|
||||
("azure_openai_endpoint", "azure_endpoint"),
|
||||
("azure_openai_api_version", "azure_api_version"),
|
||||
("azure_openai_api_key", "azure_api_key"),
|
||||
):
|
||||
if flat.get(new_key) in (None, "") and flat.get(old_key) not in (None, ""):
|
||||
flat[new_key] = flat[old_key]
|
||||
|
||||
explicit_backend = getattr(args, "backend", None)
|
||||
if explicit_backend is None:
|
||||
for option in args.cfg_options or []:
|
||||
key = str(option).split("=", 1)[0].strip()
|
||||
if key == "model.backend":
|
||||
explicit_backend = str(option).split("=", 1)[1].strip()
|
||||
break
|
||||
|
||||
backend = normalize_backend_name(flat.get("model_backend") or flat.get("target_backend") or "azure_openai")
|
||||
|
||||
def _has_model_override(dotted_key: str, legacy_key: str) -> bool:
|
||||
if getattr(args, legacy_key, None) is not None:
|
||||
return True
|
||||
for option in args.cfg_options or []:
|
||||
key = str(option).split("=", 1)[0].strip()
|
||||
if key == dotted_key:
|
||||
return True
|
||||
return False
|
||||
|
||||
if explicit_backend is not None:
|
||||
backend = normalize_backend_name(explicit_backend)
|
||||
flat["model_backend"] = backend
|
||||
if backend in {"claude", "claude_chat"}:
|
||||
flat.setdefault("optimizer_backend", "claude_chat")
|
||||
flat.setdefault("target_backend", "claude_chat")
|
||||
elif backend in {"codex", "codex_exec"}:
|
||||
flat.setdefault("optimizer_backend", "openai_chat")
|
||||
flat.setdefault("target_backend", "codex_exec")
|
||||
elif backend == "claude_code_exec":
|
||||
flat.setdefault("optimizer_backend", "openai_chat")
|
||||
flat.setdefault("target_backend", "claude_code_exec")
|
||||
elif backend in {"qwen", "qwen_chat"}:
|
||||
flat.setdefault("optimizer_backend", "openai_chat")
|
||||
flat.setdefault("target_backend", "qwen_chat")
|
||||
elif backend in {"minimax", "minimax_chat"}:
|
||||
flat.setdefault("optimizer_backend", "openai_chat")
|
||||
flat.setdefault("target_backend", "minimax_chat")
|
||||
else:
|
||||
flat.setdefault("optimizer_backend", "openai_chat")
|
||||
flat.setdefault("target_backend", "openai_chat")
|
||||
else:
|
||||
flat.setdefault("optimizer_backend", "openai_chat")
|
||||
flat.setdefault("target_backend", "openai_chat")
|
||||
|
||||
if flat.get("optimizer_backend") == "claude_chat":
|
||||
if (
|
||||
str(flat.get("optimizer_model", "") or "").strip() in _OPENAI_DEFAULT_MODEL_SENTINELS
|
||||
and not _has_model_override("model.optimizer", "optimizer_model")
|
||||
):
|
||||
flat["optimizer_model"] = default_model_for_backend("claude_chat")
|
||||
if flat.get("optimizer_backend") == "qwen_chat":
|
||||
if (
|
||||
str(flat.get("optimizer_model", "") or "").strip() in _OPENAI_DEFAULT_MODEL_SENTINELS
|
||||
and not _has_model_override("model.optimizer", "optimizer_model")
|
||||
):
|
||||
flat["optimizer_model"] = default_model_for_backend("qwen_chat")
|
||||
if flat.get("target_backend") == "claude_chat":
|
||||
if (
|
||||
str(flat.get("target_model", "") or "").strip() in _OPENAI_DEFAULT_MODEL_SENTINELS
|
||||
and not _has_model_override("model.target", "target_model")
|
||||
):
|
||||
flat["target_model"] = default_model_for_backend("claude_chat")
|
||||
if flat.get("target_backend") == "claude_code_exec":
|
||||
if (
|
||||
str(flat.get("target_model", "") or "").strip() in _OPENAI_DEFAULT_MODEL_SENTINELS
|
||||
and not _has_model_override("model.target", "target_model")
|
||||
):
|
||||
flat["target_model"] = default_model_for_backend("claude_chat")
|
||||
if flat.get("target_backend") == "qwen_chat":
|
||||
if (
|
||||
str(flat.get("target_model", "") or "").strip() in _OPENAI_DEFAULT_MODEL_SENTINELS
|
||||
and not _has_model_override("model.target", "target_model")
|
||||
):
|
||||
flat["target_model"] = default_model_for_backend("qwen_chat")
|
||||
if flat.get("target_backend") == "minimax_chat":
|
||||
if (
|
||||
str(flat.get("target_model", "") or "").strip() in _OPENAI_DEFAULT_MODEL_SENTINELS
|
||||
and not _has_model_override("model.target", "target_model")
|
||||
):
|
||||
flat["target_model"] = (
|
||||
flat.get("minimax_model")
|
||||
or default_model_for_backend("minimax_chat")
|
||||
)
|
||||
|
||||
# Auto-generate output root
|
||||
if not flat.get("out_root"):
|
||||
env = flat.get("env", "unknown")
|
||||
model = flat.get("optimizer_model", "unknown").replace("/", "-")
|
||||
ts = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
|
||||
flat["out_root"] = os.path.join("outputs", f"skillopt_{env}_{model}_{ts}")
|
||||
|
||||
flat["out_root"] = os.path.abspath(flat["out_root"])
|
||||
return flat
|
||||
|
||||
|
||||
# ── Main ─────────────────────────────────────────────────────────────────────
|
||||
|
||||
def main() -> None:
|
||||
args = parse_args()
|
||||
cfg = load_config(args)
|
||||
|
||||
print(f"\n{'='*60}")
|
||||
print(f" SkillOpt — Executive Strategy for Self-Evolving Agent Skills")
|
||||
print(f"{'='*60}")
|
||||
print(f" env: {cfg.get('env')}")
|
||||
print(f" optimizer_model: {cfg.get('optimizer_model')}")
|
||||
print(f" target_model: {cfg.get('target_model')}")
|
||||
print(f" optimizer_backend:{cfg.get('optimizer_backend', 'openai_chat')}")
|
||||
print(f" target_backend:{cfg.get('target_backend', 'openai_chat')}")
|
||||
print(f" reasoning: {cfg.get('reasoning_effort') or 'off'}")
|
||||
print(f" rewrite_effort: {cfg.get('rewrite_reasoning_effort') or 'off'}")
|
||||
print(f" epochs: {cfg.get('num_epochs')}")
|
||||
print(f" train_size: {cfg.get('train_size') or 'from dataset'}")
|
||||
print(f" steps/epoch: auto")
|
||||
print(f" batch_size: {cfg.get('batch_size')}")
|
||||
print(f" edit_budget: {cfg.get('edit_budget')}")
|
||||
print(f" lr_scheduler: {cfg.get('lr_scheduler', 'constant')}")
|
||||
print(f" update_mode: {cfg.get('skill_update_mode', 'patch')}")
|
||||
print(f" min_edit_budget:{cfg.get('min_edit_budget', 2)}")
|
||||
print(f" minibatch_size: {cfg.get('minibatch_size')}")
|
||||
print(f" seed: {cfg.get('seed')}")
|
||||
print(f" meta_skill: {cfg.get('use_meta_skill', False)}")
|
||||
print(f" slow_update: {cfg.get('use_slow_update', False)}")
|
||||
print(f" out_root: {cfg.get('out_root')}")
|
||||
print(f"{'='*60}\n")
|
||||
|
||||
# Build adapter
|
||||
adapter = get_adapter(cfg)
|
||||
|
||||
# Build trainer and run
|
||||
from skillopt.engine.trainer import ReflACTTrainer
|
||||
trainer = ReflACTTrainer(cfg, adapter)
|
||||
summary = trainer.train()
|
||||
|
||||
print(f"\n Output saved to: {cfg['out_root']}")
|
||||
if summary.get("test_hard") is not None:
|
||||
print(f" Final test: {summary['test_hard']:.4f}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1 @@
|
||||
<svg id="logomark" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 74.492 100.25"><g id="tiny_-_black" data-name="tiny - black"><path d="M586.72,255.616a3.377,3.377,0,0,1,.448.031,5.917,5.917,0,0,1,3.581,2.79c.454,1.116.314,2.023-1.315,4.141L563.168,293.6l-8.558-10.047,29.348-26.616a4.406,4.406,0,0,1,2.762-1.321m0-1.5a5.766,5.766,0,0,0-3.69,1.643l-.041.032-.038.035L553.6,282.442l-1.077.977.943,1.107,8.558,10.047,1.145,1.344,1.141-1.348,26.267-31.022.022-.027.022-.028c1.574-2.046,2.327-3.622,1.516-5.619a7.309,7.309,0,0,0-4.779-3.714,5.083,5.083,0,0,0-.64-.043Z" transform="translate(-526.086 -245.559)"/><path d="M553.423,284.593l8.977,10.558L597.911,337.9c.873,1.093,1.419,2.186,1.047,3.418a4.092,4.092,0,0,1-2.721,2.837,3.557,3.557,0,0,1-1.045.159,4,4,0,0,1-2.687-1.124L548.01,300.808c-3.5-3.5-2.971-8.151.436-11.558l4.977-4.657m.124-2.17L552.4,283.5l-4.976,4.656c-4.192,4.191-4.372,9.816-.473,13.714l44.521,42.4a5.485,5.485,0,0,0,3.722,1.538,5.1,5.1,0,0,0,1.483-.224,5.59,5.59,0,0,0,3.719-3.838,5.176,5.176,0,0,0-1.31-4.788l-35.53-42.767-8.988-10.571-1.019-1.2Z" transform="translate(-526.086 -245.559)"/><path d="M562.4,295.151l9.556,11.5,5.761-5.356a7.926,7.926,0,0,0,.041-11.743l-43.7-41.923s-1.671-2.029-3.437-2.071a4.49,4.49,0,0,0-4.23,2.718c-.688,1.651-.194,2.809,1.315,4.97l29.306,35.565Z" transform="translate(-526.086 -245.559)"/><path d="M553.7,306.223l-17.116,21.024c-1.255,1.337-2.032,3.683-1.331,5.367a4.587,4.587,0,0,0,4.287,2.841,4.087,4.087,0,0,0,3.082-1.523l20.328-18.9Z" transform="translate(-526.086 -245.559)"/><path d="M592.074,250.547" transform="translate(-526.086 -245.559)" fill="#fff" stroke="#000" stroke-miterlimit="10" stroke-width="0.25"/></g></svg>
|
||||
|
After Width: | Height: | Size: 1.6 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 78 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 11 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 390 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 10 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 407 KiB |
+2739
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,28 @@
|
||||
"""ReflACT: Reflective Agent Tuning.
|
||||
|
||||
A general-purpose framework for iteratively optimizing LLM agent skills
|
||||
through structured reflection and self-improvement.
|
||||
|
||||
Pipeline stages:
|
||||
1. Rollout — execute episodes with current skill
|
||||
2. Reflect — analyze trajectories, generate patches
|
||||
3. Aggregate — hierarchical merge of patches
|
||||
4. Select — rank and select top edits
|
||||
5. Update — apply edits to skill document
|
||||
6. Evaluate — validate candidate skill, accept/reject
|
||||
"""
|
||||
|
||||
__version__ = "0.1.0"
|
||||
|
||||
from skillopt.types import ( # noqa: F401
|
||||
BatchSpec,
|
||||
Edit,
|
||||
EditOp,
|
||||
FailureSummaryEntry,
|
||||
GateAction,
|
||||
GateResult,
|
||||
Patch,
|
||||
RawPatch,
|
||||
RolloutResult,
|
||||
SlowUpdateResult,
|
||||
)
|
||||
@@ -0,0 +1,286 @@
|
||||
"""ReflACT config loading engine — structured YAML with inheritance.
|
||||
|
||||
Supports two config formats:
|
||||
1. **Structured** (new): sections like ``model``, ``train``, ``gradient``,
|
||||
``optimizer``, ``evaluation``, ``env`` — with ``_base_`` inheritance.
|
||||
2. **Flat** (legacy): all keys at top level — fully backward compatible.
|
||||
|
||||
Usage::
|
||||
|
||||
from skillopt.config import load_config, flatten_config
|
||||
|
||||
cfg = load_config("configs/searchqa_default.yaml")
|
||||
flat = flatten_config(cfg) # always returns flat dict for trainer
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import copy
|
||||
import os
|
||||
from typing import Any
|
||||
|
||||
import yaml
|
||||
|
||||
# ── Section names that indicate a structured config ──────────────────────
|
||||
|
||||
_STRUCTURED_SECTIONS = frozenset({
|
||||
"model", "train", "gradient", "optimizer", "evaluation", "env",
|
||||
})
|
||||
|
||||
# ── Structured → flat key mapping ────────────────────────────────────────
|
||||
|
||||
_FLATTEN_MAP: dict[str, str] = {
|
||||
"model.backend": "model_backend",
|
||||
"model.optimizer": "optimizer_model",
|
||||
"model.target": "target_model",
|
||||
"model.optimizer_backend": "optimizer_backend",
|
||||
"model.target_backend": "target_backend",
|
||||
"model.reasoning_effort": "reasoning_effort",
|
||||
"model.rewrite_reasoning_effort": "rewrite_reasoning_effort",
|
||||
"model.rewrite_max_completion_tokens": "rewrite_max_completion_tokens",
|
||||
"model.codex_exec_path": "codex_exec_path",
|
||||
"model.codex_exec_sandbox": "codex_exec_sandbox",
|
||||
"model.codex_exec_profile": "codex_exec_profile",
|
||||
"model.codex_exec_full_auto": "codex_exec_full_auto",
|
||||
"model.codex_exec_reasoning_effort": "codex_exec_reasoning_effort",
|
||||
"model.codex_exec_use_sdk": "codex_exec_use_sdk",
|
||||
"model.codex_exec_network_access": "codex_exec_network_access",
|
||||
"model.codex_exec_web_search": "codex_exec_web_search",
|
||||
"model.codex_exec_approval_policy": "codex_exec_approval_policy",
|
||||
"model.claude_code_exec_path": "claude_code_exec_path",
|
||||
"model.claude_code_exec_profile": "claude_code_exec_profile",
|
||||
"model.claude_code_exec_use_sdk": "claude_code_exec_use_sdk",
|
||||
"model.claude_code_exec_effort": "claude_code_exec_effort",
|
||||
"model.claude_code_exec_max_thinking_tokens": "claude_code_exec_max_thinking_tokens",
|
||||
"model.codex_trace_to_optimizer": "codex_trace_to_optimizer",
|
||||
"model.azure_endpoint": "azure_endpoint",
|
||||
"model.azure_api_version": "azure_api_version",
|
||||
"model.azure_api_key": "azure_api_key",
|
||||
"model.azure_openai_endpoint": "azure_openai_endpoint",
|
||||
"model.azure_openai_api_version": "azure_openai_api_version",
|
||||
"model.azure_openai_api_key": "azure_openai_api_key",
|
||||
"model.azure_openai_auth_mode": "azure_openai_auth_mode",
|
||||
"model.azure_openai_ad_scope": "azure_openai_ad_scope",
|
||||
"model.azure_openai_managed_identity_client_id": "azure_openai_managed_identity_client_id",
|
||||
"model.optimizer_azure_openai_endpoint": "optimizer_azure_openai_endpoint",
|
||||
"model.optimizer_azure_openai_api_version": "optimizer_azure_openai_api_version",
|
||||
"model.optimizer_azure_openai_api_key": "optimizer_azure_openai_api_key",
|
||||
"model.optimizer_azure_openai_auth_mode": "optimizer_azure_openai_auth_mode",
|
||||
"model.optimizer_azure_openai_ad_scope": "optimizer_azure_openai_ad_scope",
|
||||
"model.optimizer_azure_openai_managed_identity_client_id": "optimizer_azure_openai_managed_identity_client_id",
|
||||
"model.target_azure_openai_endpoint": "target_azure_openai_endpoint",
|
||||
"model.target_azure_openai_api_version": "target_azure_openai_api_version",
|
||||
"model.target_azure_openai_api_key": "target_azure_openai_api_key",
|
||||
"model.target_azure_openai_auth_mode": "target_azure_openai_auth_mode",
|
||||
"model.target_azure_openai_ad_scope": "target_azure_openai_ad_scope",
|
||||
"model.target_azure_openai_managed_identity_client_id": "target_azure_openai_managed_identity_client_id",
|
||||
"model.qwen_chat_base_url": "qwen_chat_base_url",
|
||||
"model.qwen_chat_api_key": "qwen_chat_api_key",
|
||||
"model.qwen_chat_temperature": "qwen_chat_temperature",
|
||||
"model.qwen_chat_timeout_seconds": "qwen_chat_timeout_seconds",
|
||||
"model.qwen_chat_max_tokens": "qwen_chat_max_tokens",
|
||||
"model.qwen_chat_enable_thinking": "qwen_chat_enable_thinking",
|
||||
"model.optimizer_qwen_chat_base_url": "optimizer_qwen_chat_base_url",
|
||||
"model.optimizer_qwen_chat_api_key": "optimizer_qwen_chat_api_key",
|
||||
"model.optimizer_qwen_chat_temperature": "optimizer_qwen_chat_temperature",
|
||||
"model.optimizer_qwen_chat_timeout_seconds": "optimizer_qwen_chat_timeout_seconds",
|
||||
"model.optimizer_qwen_chat_max_tokens": "optimizer_qwen_chat_max_tokens",
|
||||
"model.optimizer_qwen_chat_enable_thinking": "optimizer_qwen_chat_enable_thinking",
|
||||
"model.target_qwen_chat_base_url": "target_qwen_chat_base_url",
|
||||
"model.target_qwen_chat_api_key": "target_qwen_chat_api_key",
|
||||
"model.target_qwen_chat_temperature": "target_qwen_chat_temperature",
|
||||
"model.target_qwen_chat_timeout_seconds": "target_qwen_chat_timeout_seconds",
|
||||
"model.target_qwen_chat_max_tokens": "target_qwen_chat_max_tokens",
|
||||
"model.target_qwen_chat_enable_thinking": "target_qwen_chat_enable_thinking",
|
||||
"model.minimax_base_url": "minimax_base_url",
|
||||
"model.minimax_api_key": "minimax_api_key",
|
||||
"model.minimax_model": "minimax_model",
|
||||
"model.minimax_temperature": "minimax_temperature",
|
||||
"model.minimax_max_tokens": "minimax_max_tokens",
|
||||
"model.minimax_enable_thinking": "minimax_enable_thinking",
|
||||
"train.num_epochs": "num_epochs",
|
||||
"train.train_size": "train_size",
|
||||
"train.steps_per_epoch": "steps_per_epoch",
|
||||
"train.batch_size": "batch_size",
|
||||
"train.accumulation": "accumulation",
|
||||
"train.seed": "seed",
|
||||
"gradient.minibatch_size": "minibatch_size",
|
||||
"gradient.merge_batch_size": "merge_batch_size",
|
||||
"gradient.analyst_workers": "analyst_workers",
|
||||
"gradient.failure_only": "failure_only",
|
||||
"gradient.max_analyst_rounds": "max_analyst_rounds",
|
||||
"optimizer.learning_rate": "edit_budget",
|
||||
"optimizer.min_learning_rate": "min_edit_budget",
|
||||
"optimizer.lr_scheduler": "lr_scheduler",
|
||||
"optimizer.lr_control_mode": "lr_control_mode",
|
||||
"optimizer.skill_update_mode": "skill_update_mode",
|
||||
"optimizer.meta_learning_rate": "meta_edit_budget",
|
||||
"optimizer.use_slow_update": "use_slow_update",
|
||||
"optimizer.slow_update_samples": "slow_update_samples",
|
||||
"optimizer.slow_update_gate_with_selection": "slow_update_gate_with_selection",
|
||||
"optimizer.longitudinal_pair_policy": "longitudinal_pair_policy",
|
||||
"optimizer.use_meta_skill": "use_meta_skill",
|
||||
"evaluation.use_gate": "use_gate",
|
||||
"evaluation.gate_metric": "gate_metric",
|
||||
"evaluation.gate_mixed_weight": "gate_mixed_weight",
|
||||
"evaluation.sel_env_num": "sel_env_num",
|
||||
"evaluation.test_env_num": "test_env_num",
|
||||
"evaluation.eval_test": "eval_test",
|
||||
"env.name": "env",
|
||||
"env.skill_init": "skill_init",
|
||||
"env.out_root": "out_root",
|
||||
}
|
||||
|
||||
|
||||
# ── Deep merge ───────────────────────────────────────────────────────────
|
||||
|
||||
def _deep_merge(base: dict, override: dict) -> dict:
|
||||
"""Recursively merge *override* into *base* (returns new dict)."""
|
||||
result = copy.deepcopy(base)
|
||||
for key, val in override.items():
|
||||
if key in result and isinstance(result[key], dict) and isinstance(val, dict):
|
||||
result[key] = _deep_merge(result[key], val)
|
||||
else:
|
||||
result[key] = copy.deepcopy(val)
|
||||
return result
|
||||
|
||||
|
||||
# ── YAML loading with _base_ inheritance ─────────────────────────────────
|
||||
|
||||
def _load_yaml(path: str, _visited: set[str] | None = None) -> dict:
|
||||
"""Load a YAML file, resolving ``_base_`` inheritance recursively."""
|
||||
abs_path = os.path.abspath(path)
|
||||
if _visited is None:
|
||||
_visited = set()
|
||||
if abs_path in _visited:
|
||||
raise ValueError(f"Circular _base_ inheritance: {abs_path}")
|
||||
_visited.add(abs_path)
|
||||
|
||||
with open(abs_path) as f:
|
||||
cfg = yaml.safe_load(f) or {}
|
||||
|
||||
base_ref = cfg.pop("_base_", None)
|
||||
if base_ref:
|
||||
base_path = os.path.join(os.path.dirname(abs_path), base_ref)
|
||||
base_cfg = _load_yaml(base_path, _visited)
|
||||
cfg = _deep_merge(base_cfg, cfg)
|
||||
|
||||
return cfg
|
||||
|
||||
|
||||
# ── Format detection ─────────────────────────────────────────────────────
|
||||
|
||||
def is_structured(cfg: dict) -> bool:
|
||||
"""Return True if *cfg* uses the new structured section format."""
|
||||
return any(
|
||||
key in _STRUCTURED_SECTIONS and isinstance(cfg.get(key), dict)
|
||||
for key in cfg
|
||||
)
|
||||
|
||||
|
||||
# ── Flatten ──────────────────────────────────────────────────────────────
|
||||
|
||||
def flatten_config(cfg: dict) -> dict:
|
||||
"""Convert a structured config to the flat dict expected by the trainer.
|
||||
|
||||
If *cfg* is already flat, returns a shallow copy unchanged.
|
||||
"""
|
||||
if not is_structured(cfg):
|
||||
return dict(cfg)
|
||||
|
||||
flat: dict[str, Any] = {}
|
||||
|
||||
evaluation_section = cfg.get("evaluation", {})
|
||||
if isinstance(evaluation_section, dict) and evaluation_section.get("use_gate") is False:
|
||||
raise ValueError(
|
||||
"Gate validation is mandatory in this branch. Remove "
|
||||
"`evaluation.use_gate: false` from the config."
|
||||
)
|
||||
|
||||
# Apply the explicit mapping
|
||||
for dotted, flat_key in _FLATTEN_MAP.items():
|
||||
section, key = dotted.split(".", 1)
|
||||
section_dict = cfg.get(section, {})
|
||||
if isinstance(section_dict, dict) and key in section_dict:
|
||||
flat[flat_key] = section_dict[key]
|
||||
|
||||
# Pass through env-specific keys not in the explicit mapping
|
||||
env_section = cfg.get("env", {})
|
||||
if isinstance(env_section, dict):
|
||||
mapped_env_keys = {
|
||||
k.split(".", 1)[1]
|
||||
for k in _FLATTEN_MAP
|
||||
if k.startswith("env.")
|
||||
}
|
||||
for key, val in env_section.items():
|
||||
if key not in mapped_env_keys:
|
||||
flat[key] = val
|
||||
|
||||
return flat
|
||||
|
||||
|
||||
# ── Override application ─────────────────────────────────────────────────
|
||||
|
||||
def _cast_value(val_str: str) -> Any:
|
||||
"""Auto-cast a CLI string value to int / float / bool / str."""
|
||||
if val_str.lower() in ("true", "yes"):
|
||||
return True
|
||||
if val_str.lower() in ("false", "no"):
|
||||
return False
|
||||
try:
|
||||
return int(val_str)
|
||||
except ValueError:
|
||||
pass
|
||||
try:
|
||||
return float(val_str)
|
||||
except ValueError:
|
||||
pass
|
||||
return val_str
|
||||
|
||||
|
||||
def apply_overrides(cfg: dict, overrides: list[str]) -> None:
|
||||
"""Apply ``key=value`` overrides to a structured config (in place).
|
||||
|
||||
Supports both ``section.key=value`` (for structured configs) and
|
||||
``key=value`` (for flat configs or flat keys in env section).
|
||||
"""
|
||||
for item in overrides:
|
||||
if "=" not in item:
|
||||
raise ValueError(f"Invalid override (expected key=value): {item!r}")
|
||||
key, val_str = item.split("=", 1)
|
||||
val = _cast_value(val_str)
|
||||
|
||||
if "." in key:
|
||||
section, subkey = key.split(".", 1)
|
||||
if section in cfg and isinstance(cfg[section], dict):
|
||||
cfg[section][subkey] = val
|
||||
else:
|
||||
cfg.setdefault(section, {})[subkey] = val
|
||||
else:
|
||||
# Flat key — apply to top level (for legacy compat)
|
||||
cfg[key] = val
|
||||
|
||||
|
||||
# ── Public API ───────────────────────────────────────────────────────────
|
||||
|
||||
def load_config(
|
||||
path: str,
|
||||
overrides: list[str] | None = None,
|
||||
) -> dict:
|
||||
"""Load a config file with ``_base_`` inheritance and optional overrides.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
path : str
|
||||
Path to the YAML config file.
|
||||
overrides : list[str] | None
|
||||
``key=value`` strings from ``--cfg-options``.
|
||||
|
||||
Returns
|
||||
-------
|
||||
dict
|
||||
The merged config (structured or flat depending on the YAML).
|
||||
"""
|
||||
cfg = _load_yaml(path)
|
||||
if overrides:
|
||||
apply_overrides(cfg, overrides)
|
||||
return cfg
|
||||
@@ -0,0 +1,7 @@
|
||||
"""ReflACT Datasets -- task batch planning and data loading.
|
||||
|
||||
Analogous to the datasets and dataloaders in neural network training:
|
||||
provides batch sampling, epoch planning, and data management for the
|
||||
ReflACT training pipeline.
|
||||
"""
|
||||
from skillopt.datasets.base import BaseDataLoader, BatchSpec, SplitDataLoader # noqa: F401
|
||||
@@ -0,0 +1,512 @@
|
||||
"""Generic task dataloader abstractions for ReflACT.
|
||||
|
||||
ReflACT does not train model parameters directly. Instead, it iterates over
|
||||
task batches, rolls out the current skill, reflects on failures/successes,
|
||||
and updates the skill document. Because of that, the "dataloader" abstraction
|
||||
here is closer to a batch sampler / episode planner than a tensor loader.
|
||||
|
||||
Class hierarchy::
|
||||
|
||||
BaseDataLoader # abstract — simulator-backed envs (e.g. ALFWorld)
|
||||
└── SplitDataLoader # abstract — dataset-backed envs with split_dir
|
||||
|
||||
SplitDataLoader supports two dataset entry modes:
|
||||
|
||||
1. ``split_mode="split_dir"``: consume an existing split directory.
|
||||
2. ``split_mode="ratio"``: build a deterministic split directory from a raw
|
||||
dataset path using an explicit train:val:test ratio.
|
||||
|
||||
In either case, the standardised split layout is:
|
||||
|
||||
split_dir/
|
||||
├── train/ # training items
|
||||
├── val/ # validation / selection items (gate)
|
||||
└── test/ # held-out test items
|
||||
|
||||
Each subdirectory's contents are benchmark-specific. Subclasses only need
|
||||
to implement ``load_split_items(split_path)`` to teach the loader how to
|
||||
read items from one of those directories.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import glob
|
||||
import json
|
||||
import os
|
||||
import random
|
||||
from abc import ABC, abstractmethod
|
||||
from dataclasses import dataclass, field
|
||||
from typing import Any
|
||||
|
||||
|
||||
@dataclass(slots=True)
|
||||
class BatchSpec:
|
||||
"""A concrete batch request consumed by the training loop.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
phase : str
|
||||
``"train"`` or ``"eval"``.
|
||||
split : str
|
||||
Dataset split name, typically ``"train"`` or an eval split.
|
||||
seed : int
|
||||
Random seed used to construct the batch deterministically.
|
||||
batch_size : int
|
||||
Requested number of items / episodes in this batch.
|
||||
payload : object | None
|
||||
Environment-specific batch payload. For dataset-backed environments
|
||||
this is often a list of sampled items; for simulator-backed
|
||||
environments this may be ``None`` and the seed alone can define the
|
||||
batch.
|
||||
metadata : dict[str, Any]
|
||||
Optional structured metadata for logging, resume, or curriculum logic.
|
||||
"""
|
||||
|
||||
phase: str
|
||||
split: str
|
||||
seed: int
|
||||
batch_size: int
|
||||
payload: object | None = None
|
||||
metadata: dict[str, Any] = field(default_factory=dict)
|
||||
|
||||
|
||||
class BaseDataLoader(ABC):
|
||||
"""Abstract base class for task batch planning in ReflACT.
|
||||
|
||||
Subclasses are responsible for defining how a train or eval batch is
|
||||
sampled. The default implementation here provides deterministic epoch seed
|
||||
planning so all loaders share the same reproducibility behavior.
|
||||
"""
|
||||
|
||||
def setup(self, cfg: dict) -> None:
|
||||
"""Optional one-time initialization with the full trainer config."""
|
||||
|
||||
def set_out_root(self, out_root: str) -> None:
|
||||
"""Optional hook for loaders that persist split files or state."""
|
||||
|
||||
def state_dict(self) -> dict[str, Any]:
|
||||
"""Return serializable loader state for resume support."""
|
||||
return {}
|
||||
|
||||
def load_state_dict(self, state: dict[str, Any]) -> None:
|
||||
"""Restore loader state from :meth:`state_dict` output."""
|
||||
|
||||
def get_train_size(self) -> int | None:
|
||||
"""Return the size of the training pool when known."""
|
||||
return None
|
||||
|
||||
@staticmethod
|
||||
def make_base_seeds(steps_per_epoch: int, accumulation: int, seed: int) -> list[int]:
|
||||
"""Return the deterministic seed pool used to define train batches."""
|
||||
batches_per_epoch = steps_per_epoch * accumulation
|
||||
return [seed + i + 1 for i in range(batches_per_epoch)]
|
||||
|
||||
@staticmethod
|
||||
def shuffle_epoch_seeds(base_seeds: list[int], epoch: int, seed: int) -> list[int]:
|
||||
"""Return the per-epoch deterministic shuffle of *base_seeds*."""
|
||||
epoch_rng = random.Random(seed + epoch * 1000)
|
||||
shuffled = list(base_seeds)
|
||||
epoch_rng.shuffle(shuffled)
|
||||
return shuffled
|
||||
|
||||
def plan_train_epoch(
|
||||
self,
|
||||
*,
|
||||
epoch: int,
|
||||
steps_per_epoch: int,
|
||||
accumulation: int,
|
||||
batch_size: int,
|
||||
seed: int,
|
||||
**kwargs,
|
||||
) -> list[BatchSpec]:
|
||||
"""Build the full list of training batches for one epoch."""
|
||||
base_seeds = self.make_base_seeds(
|
||||
steps_per_epoch=steps_per_epoch,
|
||||
accumulation=accumulation,
|
||||
seed=seed,
|
||||
)
|
||||
shuffled_seeds = self.shuffle_epoch_seeds(base_seeds, epoch=epoch, seed=seed)
|
||||
return [
|
||||
self.build_train_batch(batch_size=batch_size, seed=batch_seed, **kwargs)
|
||||
for batch_seed in shuffled_seeds
|
||||
]
|
||||
|
||||
@abstractmethod
|
||||
def build_train_batch(self, batch_size: int, seed: int, **kwargs) -> BatchSpec:
|
||||
"""Construct one training batch specification."""
|
||||
|
||||
@abstractmethod
|
||||
def build_eval_batch(
|
||||
self,
|
||||
env_num: int,
|
||||
split: str,
|
||||
seed: int,
|
||||
**kwargs,
|
||||
) -> BatchSpec:
|
||||
"""Construct one evaluation batch specification."""
|
||||
|
||||
|
||||
# ── Split-based dataloader for dataset-backed environments ──────────────
|
||||
|
||||
# Canonical split names expected under split_dir/
|
||||
SPLIT_NAMES = ("train", "val", "test")
|
||||
|
||||
# Maps legacy / trainer split names → canonical directory names
|
||||
_SPLIT_ALIAS: dict[str, str] = {
|
||||
"train": "train",
|
||||
"valid_seen": "val",
|
||||
"selection": "val",
|
||||
"val": "val",
|
||||
"valid_unseen": "test",
|
||||
"test": "test",
|
||||
}
|
||||
|
||||
|
||||
def _load_json_or_jsonl(path: str) -> list[dict]:
|
||||
"""Load a list of items from a JSON or JSONL file."""
|
||||
with open(path, encoding="utf-8") as f:
|
||||
content = f.read().strip()
|
||||
if not content:
|
||||
return []
|
||||
|
||||
try:
|
||||
data = json.loads(content)
|
||||
except json.JSONDecodeError:
|
||||
data = None
|
||||
|
||||
if isinstance(data, list):
|
||||
return data
|
||||
if isinstance(data, dict):
|
||||
nested = data.get("data")
|
||||
if isinstance(nested, list):
|
||||
return nested
|
||||
return list(data.values())
|
||||
|
||||
items: list[dict] = []
|
||||
for line in content.splitlines():
|
||||
line = line.strip()
|
||||
if line:
|
||||
items.append(json.loads(line))
|
||||
return items
|
||||
|
||||
|
||||
def _parse_split_ratio(text: str) -> tuple[int, int, int]:
|
||||
parts = [part.strip() for part in str(text or "").split(":") if part.strip()]
|
||||
if len(parts) != 3:
|
||||
raise ValueError(
|
||||
f"split_ratio must be in train:val:test form, got {text!r}"
|
||||
)
|
||||
try:
|
||||
train, val, test = (int(part) for part in parts)
|
||||
except ValueError as exc:
|
||||
raise ValueError(
|
||||
f"split_ratio must contain integers, got {text!r}"
|
||||
) from exc
|
||||
if min(train, val, test) <= 0:
|
||||
raise ValueError(f"split_ratio parts must be positive, got {text!r}")
|
||||
return train, val, test
|
||||
|
||||
|
||||
def _compute_split_counts(total: int, ratio: tuple[int, int, int]) -> tuple[int, int, int]:
|
||||
weights = list(ratio)
|
||||
denom = sum(weights)
|
||||
raw = [total * weight / denom for weight in weights]
|
||||
counts = [int(value) for value in raw]
|
||||
remaining = total - sum(counts)
|
||||
order = sorted(
|
||||
range(len(raw)),
|
||||
key=lambda idx: (raw[idx] - counts[idx], weights[idx]),
|
||||
reverse=True,
|
||||
)
|
||||
for idx in order[:remaining]:
|
||||
counts[idx] += 1
|
||||
return counts[0], counts[1], counts[2]
|
||||
|
||||
|
||||
class SplitDataLoader(BaseDataLoader):
|
||||
"""Base class for dataset-backed environments.
|
||||
|
||||
Supported modes:
|
||||
|
||||
- ``split_mode="split_dir"``: load an existing ``train/``, ``val/``,
|
||||
``test/`` directory tree.
|
||||
- ``split_mode="ratio"``: load raw items from ``data_path`` and materialize
|
||||
a deterministic split directory with the requested ratio.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
split_dir: str = "",
|
||||
data_path: str = "",
|
||||
split_mode: str = "ratio",
|
||||
split_ratio: str = "2:1:7",
|
||||
split_seed: int = 42,
|
||||
split_output_dir: str = "",
|
||||
seed: int = 42,
|
||||
limit: int = 0,
|
||||
**kwargs,
|
||||
) -> None:
|
||||
self.split_dir = split_dir
|
||||
self.data_path = data_path
|
||||
self.split_mode = split_mode
|
||||
self.split_ratio = split_ratio
|
||||
self.split_seed = int(split_seed)
|
||||
self.split_output_dir = split_output_dir
|
||||
self.seed = seed
|
||||
self.limit = limit
|
||||
self._splits: dict[str, list[dict]] = {}
|
||||
|
||||
# ── Setup ────────────────────────────────────────────────────────────
|
||||
|
||||
def setup(self, cfg: dict) -> None:
|
||||
if not self.split_mode:
|
||||
self.split_mode = str(cfg.get("split_mode", "ratio") or "ratio")
|
||||
if not self.split_dir:
|
||||
self.split_dir = cfg.get("split_dir", "")
|
||||
if not self.data_path:
|
||||
self.data_path = cfg.get("data_path", "")
|
||||
if not self.split_output_dir:
|
||||
self.split_output_dir = cfg.get("split_output_dir", "")
|
||||
if "split_seed" in cfg and not self.split_seed:
|
||||
self.split_seed = int(cfg.get("split_seed", 0) or 0)
|
||||
if not self.split_seed:
|
||||
self.split_seed = self.seed
|
||||
if not self.split_ratio:
|
||||
self.split_ratio = str(cfg.get("split_ratio", "2:1:7") or "2:1:7")
|
||||
|
||||
mode = str(self.split_mode or "ratio").strip().lower()
|
||||
if mode not in {"ratio", "split_dir"}:
|
||||
raise ValueError(
|
||||
f"{type(self).__name__} split_mode must be 'ratio' or 'split_dir', "
|
||||
f"got {self.split_mode!r}"
|
||||
)
|
||||
self.split_mode = mode
|
||||
|
||||
if self.split_mode == "ratio":
|
||||
self.split_dir = self._materialize_ratio_split(cfg)
|
||||
if not self.split_dir:
|
||||
raise ValueError(
|
||||
f"{type(self).__name__} requires either "
|
||||
"`split_mode=ratio` with `data_path`, or `split_mode=split_dir` "
|
||||
f"with `split_dir` pointing to {'/'.join(SPLIT_NAMES)}/."
|
||||
)
|
||||
self._load_all_splits()
|
||||
|
||||
def _resolve_split_output_dir(self, cfg: dict) -> str:
|
||||
if self.split_output_dir:
|
||||
return os.path.abspath(self.split_output_dir)
|
||||
out_root = os.path.abspath(str(cfg.get("out_root") or os.getcwd()))
|
||||
env_name = str(cfg.get("env") or type(self).__name__.replace("DataLoader", "").lower())
|
||||
ratio_tag = str(self.split_ratio or "2:1:7").replace(":", "-")
|
||||
return os.path.join(out_root, "_generated_splits", f"{env_name}_{ratio_tag}_seed{self.split_seed}")
|
||||
|
||||
def load_raw_items(self, data_path: str) -> list[dict]:
|
||||
"""Load raw items from a dataset path before ratio splitting.
|
||||
|
||||
Subclasses can override when the raw dataset is not a single JSON/JSONL
|
||||
file or when directory layouts require custom normalization.
|
||||
"""
|
||||
if os.path.isdir(data_path):
|
||||
if any(os.path.isdir(os.path.join(data_path, name)) for name in SPLIT_NAMES):
|
||||
raise ValueError(
|
||||
f"{type(self).__name__} got a split directory as data_path. "
|
||||
"Use split_mode=split_dir and pass it as split_dir instead."
|
||||
)
|
||||
candidates = sorted(glob.glob(os.path.join(data_path, "*.json")))
|
||||
candidates += sorted(glob.glob(os.path.join(data_path, "*.jsonl")))
|
||||
if len(candidates) != 1:
|
||||
raise ValueError(
|
||||
f"{type(self).__name__} expected data_path to be one JSON/JSONL file "
|
||||
f"or a directory containing exactly one such file, got: {data_path}"
|
||||
)
|
||||
return _load_json_or_jsonl(candidates[0])
|
||||
return _load_json_or_jsonl(data_path)
|
||||
|
||||
def write_split_items(self, split_path: str, items: list[dict]) -> None:
|
||||
os.makedirs(split_path, exist_ok=True)
|
||||
out_path = os.path.join(split_path, "items.json")
|
||||
with open(out_path, "w", encoding="utf-8") as f:
|
||||
json.dump(items, f, ensure_ascii=False, indent=2)
|
||||
|
||||
def _materialize_ratio_split(self, cfg: dict) -> str:
|
||||
data_path = os.path.abspath(str(self.data_path or "").strip())
|
||||
if not data_path:
|
||||
raise ValueError(
|
||||
f"{type(self).__name__} requires data_path when split_mode=ratio."
|
||||
)
|
||||
|
||||
ratio = _parse_split_ratio(self.split_ratio)
|
||||
items = self.load_raw_items(data_path)
|
||||
if not isinstance(items, list) or not items:
|
||||
raise ValueError(f"No raw items available for ratio split from {data_path}")
|
||||
|
||||
shuffled = list(items)
|
||||
rng = random.Random(self.split_seed)
|
||||
rng.shuffle(shuffled)
|
||||
|
||||
train_n, val_n, test_n = _compute_split_counts(len(shuffled), ratio)
|
||||
train_items = shuffled[:train_n]
|
||||
val_items = shuffled[train_n: train_n + val_n]
|
||||
test_items = shuffled[train_n + val_n: train_n + val_n + test_n]
|
||||
|
||||
split_dir = self._resolve_split_output_dir(cfg)
|
||||
manifest = {
|
||||
"source_data_path": data_path,
|
||||
"split_mode": "ratio",
|
||||
"split_ratio": self.split_ratio,
|
||||
"split_seed": self.split_seed,
|
||||
"counts": {
|
||||
"train": len(train_items),
|
||||
"val": len(val_items),
|
||||
"test": len(test_items),
|
||||
},
|
||||
}
|
||||
os.makedirs(split_dir, exist_ok=True)
|
||||
self.write_split_items(os.path.join(split_dir, "train"), train_items)
|
||||
self.write_split_items(os.path.join(split_dir, "val"), val_items)
|
||||
self.write_split_items(os.path.join(split_dir, "test"), test_items)
|
||||
with open(os.path.join(split_dir, "split_manifest.json"), "w", encoding="utf-8") as f:
|
||||
json.dump(manifest, f, ensure_ascii=False, indent=2)
|
||||
print(
|
||||
f" [{type(self).__name__}] generated ratio split {self.split_ratio} "
|
||||
f"at {split_dir} from {data_path}"
|
||||
)
|
||||
return split_dir
|
||||
|
||||
def _load_all_splits(self) -> None:
|
||||
for name in SPLIT_NAMES:
|
||||
split_path = os.path.join(self.split_dir, name)
|
||||
if not os.path.isdir(split_path):
|
||||
raise ValueError(
|
||||
f"Missing '{name}/' subdirectory in split_dir: {self.split_dir}"
|
||||
)
|
||||
items = self.load_split_items(split_path)
|
||||
if self.limit:
|
||||
items = items[: self.limit]
|
||||
self._splits[name] = items
|
||||
|
||||
counts = " ".join(f"{k}={len(v)}" for k, v in self._splits.items())
|
||||
print(f" [{type(self).__name__}] {counts} (from {self.split_dir})")
|
||||
|
||||
def load_split_items(self, split_path: str) -> list[dict]:
|
||||
"""Load items from one split directory (e.g. ``split_dir/train/``).
|
||||
|
||||
Default: finds the first ``.json`` file in the directory and loads it
|
||||
as a JSON array. Subclasses can override for custom formats.
|
||||
"""
|
||||
json_files = sorted(glob.glob(os.path.join(split_path, "*.json")))
|
||||
if not json_files:
|
||||
raise FileNotFoundError(
|
||||
f"No .json file found in {split_path}"
|
||||
)
|
||||
with open(json_files[0], encoding="utf-8") as f:
|
||||
items = json.load(f)
|
||||
if not isinstance(items, list):
|
||||
raise ValueError(
|
||||
f"Expected JSON array in {json_files[0]}, got {type(items).__name__}"
|
||||
)
|
||||
return items
|
||||
|
||||
# ── Accessors ────────────────────────────────────────────────────────
|
||||
|
||||
@property
|
||||
def train_items(self) -> list[dict]:
|
||||
return self._splits.get("train", [])
|
||||
|
||||
@property
|
||||
def val_items(self) -> list[dict]:
|
||||
return self._splits.get("val", [])
|
||||
|
||||
@property
|
||||
def test_items(self) -> list[dict]:
|
||||
return self._splits.get("test", [])
|
||||
|
||||
def get_split_items(self, split: str) -> list[dict]:
|
||||
"""Resolve a split name (including legacy aliases) to its item list."""
|
||||
canonical = _SPLIT_ALIAS.get(split, split)
|
||||
return list(self._splits.get(canonical, self.val_items))
|
||||
|
||||
def get_train_size(self) -> int:
|
||||
return len(self.train_items)
|
||||
|
||||
def plan_train_epoch(
|
||||
self,
|
||||
*,
|
||||
epoch: int,
|
||||
steps_per_epoch: int,
|
||||
accumulation: int,
|
||||
batch_size: int,
|
||||
seed: int,
|
||||
**kwargs,
|
||||
) -> list[BatchSpec]:
|
||||
"""Build one full epoch that covers the train split in shuffled order.
|
||||
|
||||
For split-backed datasets, an epoch should correspond to one pass over
|
||||
the available training items rather than repeated independent sampling.
|
||||
"""
|
||||
epoch_rng = random.Random(seed + epoch * 1000)
|
||||
items = list(self.train_items)
|
||||
epoch_rng.shuffle(items)
|
||||
|
||||
total_batches = steps_per_epoch * accumulation
|
||||
if total_batches <= 0:
|
||||
return []
|
||||
|
||||
batches: list[BatchSpec] = []
|
||||
cursor = 0
|
||||
for batch_idx in range(total_batches):
|
||||
batch_items = items[cursor: cursor + batch_size]
|
||||
cursor += len(batch_items)
|
||||
|
||||
# Extremely small datasets can leave trailing empty microbatches
|
||||
# when accumulation > 1. Reuse the shuffled prefix in that case so
|
||||
# the trainer still receives the expected batch count.
|
||||
if not batch_items and items:
|
||||
refill_rng = random.Random(seed + epoch * 1000 + batch_idx + 1)
|
||||
batch_items = list(items)
|
||||
refill_rng.shuffle(batch_items)
|
||||
batch_items = batch_items[:batch_size]
|
||||
|
||||
batches.append(
|
||||
BatchSpec(
|
||||
phase="train",
|
||||
split="train",
|
||||
seed=seed + epoch * 1000 + batch_idx + 1,
|
||||
batch_size=len(batch_items),
|
||||
payload=batch_items,
|
||||
)
|
||||
)
|
||||
|
||||
return batches
|
||||
|
||||
# ── Batch construction ───────────────────────────────────────────────
|
||||
|
||||
def build_train_batch(self, batch_size: int, seed: int, **kwargs) -> BatchSpec:
|
||||
rng = random.Random(seed)
|
||||
items = list(self.train_items)
|
||||
rng.shuffle(items)
|
||||
items = items[:batch_size]
|
||||
return BatchSpec(
|
||||
phase="train",
|
||||
split="train",
|
||||
seed=seed,
|
||||
batch_size=len(items),
|
||||
payload=items,
|
||||
)
|
||||
|
||||
def build_eval_batch(
|
||||
self,
|
||||
env_num: int,
|
||||
split: str,
|
||||
seed: int,
|
||||
**kwargs,
|
||||
) -> BatchSpec:
|
||||
items = self.get_split_items(split)
|
||||
if env_num and env_num < len(items):
|
||||
items = items[:env_num]
|
||||
return BatchSpec(
|
||||
phase="eval",
|
||||
split=split,
|
||||
seed=seed,
|
||||
batch_size=len(items),
|
||||
payload=items,
|
||||
)
|
||||
@@ -0,0 +1,9 @@
|
||||
"""ReflACT Engine -- the training runner.
|
||||
|
||||
Analogous to the Runner in mmengine: orchestrates the full training pipeline
|
||||
including rollout, gradient computation, aggregation, optimization, and
|
||||
evaluation.
|
||||
"""
|
||||
from skillopt.engine.trainer import ReflACTTrainer # noqa: F401
|
||||
|
||||
__all__ = ["ReflACTTrainer"]
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1 @@
|
||||
"""ReflACT environment adapters."""
|
||||
@@ -0,0 +1,19 @@
|
||||
# Benchmark Template
|
||||
|
||||
This directory provides scaffold files for adding a new benchmark to SkillOpt.
|
||||
|
||||
## Files
|
||||
|
||||
- `env_template.py` — Environment adapter template
|
||||
- `loader_template.py` — Data loader template
|
||||
- `config_template.yaml` — Config file template
|
||||
|
||||
## Usage
|
||||
|
||||
1. Copy this directory: `cp -r skillopt/envs/_template skillopt/envs/your_benchmark`
|
||||
2. Rename files: remove `_template` suffix
|
||||
3. Implement the `TODO` sections
|
||||
4. Register your adapter in `_ENV_REGISTRY` inside `scripts/train.py`
|
||||
5. Create config at `configs/your_benchmark/default.yaml`
|
||||
|
||||
See the [documentation](../../docs/guide/new-benchmark.md) for the full guide.
|
||||
@@ -0,0 +1,45 @@
|
||||
# ──────────────────────────────────────────────────
|
||||
# SkillOpt Config Template — <Your Benchmark Name>
|
||||
# ──────────────────────────────────────────────────
|
||||
# Copy this file to configs/<your_benchmark>/default.yaml
|
||||
# and customize the values below.
|
||||
|
||||
# Inherit global defaults
|
||||
_base_: ../_base_/default.yaml
|
||||
|
||||
# ── Environment ──────────────────────────────────
|
||||
env:
|
||||
name: your_benchmark # Must match _ENV_REGISTRY key in scripts/train.py
|
||||
data_path: data/your_benchmark # Path to your data
|
||||
split_mode: ratio # "ratio" or "split_dir"
|
||||
split_ratio: "2:1:7" # train:val:test
|
||||
exec_timeout: 120 # Per-task timeout (seconds)
|
||||
|
||||
# ── Training ─────────────────────────────────────
|
||||
train:
|
||||
num_epochs: 4 # Number of epochs
|
||||
batch_size: 40 # Tasks per step (batch size)
|
||||
seed: 42
|
||||
|
||||
# ── Gradient (Reflection) ───────────────────────
|
||||
gradient:
|
||||
analyst_workers: 16 # Parallel reflection workers
|
||||
minibatch_size: 8
|
||||
|
||||
# ── Optimizer ────────────────────────────────────
|
||||
optimizer:
|
||||
learning_rate: 4 # Max edits per step (edit budget)
|
||||
lr_scheduler: cosine # cosine | linear | constant | autonomous
|
||||
use_slow_update: true # Epoch-boundary momentum
|
||||
use_meta_skill: true # Cross-epoch optimizer memory
|
||||
|
||||
# ── Evaluation ───────────────────────────────────
|
||||
evaluation:
|
||||
use_gate: true # Validation gating
|
||||
eval_test: true # Run test eval after training
|
||||
|
||||
# ── Model ────────────────────────────────────────
|
||||
model:
|
||||
backend: azure_openai # azure_openai | openai_chat | claude_code_exec | qwen
|
||||
optimizer: gpt-4o
|
||||
target: gpt-4o
|
||||
@@ -0,0 +1,83 @@
|
||||
"""
|
||||
Benchmark Environment Template
|
||||
===============================
|
||||
Copy this file and implement the TODO sections to add a new benchmark.
|
||||
|
||||
The EnvAdapter is responsible for:
|
||||
1. Building train/eval environment payloads
|
||||
2. Running rollout and returning scored result rows
|
||||
3. Reflecting on results and returning patch candidates
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from skillopt.datasets.base import BatchSpec
|
||||
from skillopt.envs._template.loader_template import TemplateBenchmarkDataLoader
|
||||
from skillopt.envs.base import EnvAdapter
|
||||
|
||||
|
||||
class TemplateBenchmarkAdapter(EnvAdapter):
|
||||
"""
|
||||
Environment adapter for <Your Benchmark Name>.
|
||||
|
||||
Rename this class and implement the abstract methods below.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
split_dir: str = "",
|
||||
data_path: str = "",
|
||||
split_mode: str = "ratio",
|
||||
split_ratio: str = "2:1:7",
|
||||
split_seed: int = 42,
|
||||
split_output_dir: str = "",
|
||||
seed: int = 42,
|
||||
limit: int = 0,
|
||||
**kwargs,
|
||||
) -> None:
|
||||
self.dataloader = TemplateBenchmarkDataLoader(
|
||||
split_dir=split_dir,
|
||||
data_path=data_path,
|
||||
split_mode=split_mode,
|
||||
split_ratio=split_ratio,
|
||||
split_seed=split_seed,
|
||||
split_output_dir=split_output_dir,
|
||||
seed=seed,
|
||||
limit=limit,
|
||||
)
|
||||
# TODO: initialize runtime options, e.g.
|
||||
# self.max_retries = int(kwargs.get("max_retries", 3))
|
||||
# self.timeout_s = int(kwargs.get("timeout_s", 120))
|
||||
|
||||
def setup(self, cfg: dict) -> None:
|
||||
super().setup(cfg)
|
||||
self.dataloader.setup(cfg)
|
||||
|
||||
def get_dataloader(self):
|
||||
return self.dataloader
|
||||
|
||||
def build_env_from_batch(self, batch: BatchSpec, **kwargs):
|
||||
return list(batch.payload or [])
|
||||
|
||||
def build_train_env(self, batch_size: int, seed: int, **kwargs):
|
||||
batch = self.dataloader.build_train_batch(batch_size=batch_size, seed=seed, **kwargs)
|
||||
return self.build_env_from_batch(batch, **kwargs)
|
||||
|
||||
def build_eval_env(self, env_num: int, split: str, seed: int, **kwargs):
|
||||
batch = self.dataloader.build_eval_batch(env_num=env_num, split=split, seed=seed, **kwargs)
|
||||
return self.build_env_from_batch(batch, **kwargs)
|
||||
|
||||
def rollout(self, env_manager, skill_content: str, out_dir: str, **kwargs) -> list[dict]:
|
||||
"""
|
||||
Run one batch and return list[dict] with at least:
|
||||
{"id": str, "hard": int, "soft": float}
|
||||
"""
|
||||
raise NotImplementedError("Implement rollout() for your benchmark")
|
||||
|
||||
def reflect(self, results: list[dict], skill_content: str, out_dir: str, **kwargs) -> list[dict | None]:
|
||||
"""
|
||||
Reflect on rollout results and return patch dicts (or None entries).
|
||||
"""
|
||||
raise NotImplementedError("Implement reflect() for your benchmark")
|
||||
|
||||
def get_task_types(self) -> list[str]:
|
||||
return ["your_benchmark"]
|
||||
@@ -0,0 +1,40 @@
|
||||
"""
|
||||
Benchmark Data Loader Template
|
||||
================================
|
||||
Copy this file and implement the TODO sections to load your benchmark data.
|
||||
|
||||
The SplitDataLoader is responsible for:
|
||||
1. Loading raw data from disk for ratio split mode
|
||||
2. Loading items from train/val/test directories for split_dir mode
|
||||
3. Returning list[dict] items used by the training loop
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from skillopt.datasets.base import SplitDataLoader
|
||||
|
||||
|
||||
class TemplateBenchmarkDataLoader(SplitDataLoader):
|
||||
"""
|
||||
Data loader for <Your Benchmark Name>.
|
||||
|
||||
Rename this class and implement the methods below.
|
||||
"""
|
||||
|
||||
def load_raw_items(self, data_path: str) -> list[dict]:
|
||||
"""
|
||||
Parse raw benchmark data for split_mode="ratio".
|
||||
|
||||
Return a list of normalized item dicts.
|
||||
"""
|
||||
# TODO: parse your raw JSON/JSONL/CSV format and return list[dict]
|
||||
# with deterministic "id" values.
|
||||
return super().load_raw_items(data_path)
|
||||
|
||||
def load_split_items(self, split_path: str) -> list[dict]:
|
||||
"""
|
||||
Parse one split directory for split_mode="split_dir".
|
||||
|
||||
split_path points to train/, val/, or test/.
|
||||
"""
|
||||
# TODO: customize when split directories contain non-standard files.
|
||||
return super().load_split_items(split_path)
|
||||
@@ -0,0 +1,5 @@
|
||||
"""ALFWorld environment adapter for ReflACT."""
|
||||
|
||||
from skillopt.envs.alfworld.adapter import ALFWorldAdapter
|
||||
|
||||
__all__ = ["ALFWorldAdapter"]
|
||||
@@ -0,0 +1,459 @@
|
||||
"""ALFWorld environment adapter for ReflACT.
|
||||
|
||||
Connects the ReflACT training loop to ALFWorld by implementing
|
||||
:class:`~skillopt.envs.base.EnvAdapter`.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
import json
|
||||
import os
|
||||
|
||||
from skillopt.datasets.base import BatchSpec
|
||||
from skillopt.envs.base import EnvAdapter
|
||||
from skillopt.envs.alfworld.dataloader import ALFWorldDataLoader
|
||||
from skillopt.envs.alfworld.rollout import (
|
||||
build_alfworld_env,
|
||||
run_alfworld_batch,
|
||||
TASKS,
|
||||
)
|
||||
from skillopt.gradient.reflect import run_minibatch_reflect
|
||||
from skillopt.utils import compute_score
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ALFWorldBatchRun:
|
||||
"""Lazy ALFWorld batch description.
|
||||
|
||||
The adapter materializes this in rollout chunks so a large evaluation set
|
||||
does not keep every ALFWorld simulator open at once.
|
||||
"""
|
||||
|
||||
env_num: int
|
||||
eval_dataset: str
|
||||
seed: int
|
||||
is_train: bool
|
||||
workers: int
|
||||
specific_gamefiles: list[str] | None = None
|
||||
result_ids: list[str] | None = None
|
||||
items: list[dict] | None = None
|
||||
|
||||
def __iter__(self):
|
||||
return iter(self.items or [])
|
||||
|
||||
def __len__(self) -> int:
|
||||
return int(self.env_num or 0)
|
||||
|
||||
|
||||
class ALFWorldAdapter(EnvAdapter):
|
||||
"""ALFWorld environment adapter.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
max_steps : int
|
||||
Maximum steps per ALFWorld episode (default 50).
|
||||
max_api_workers : int
|
||||
Maximum concurrent API calls during rollout (default 8).
|
||||
analyst_workers : int
|
||||
Parallel workers for analyst stage (default 16).
|
||||
failure_only : bool
|
||||
If True, only run error analyst (skip success analyst).
|
||||
minibatch_size : int
|
||||
Trajectories per analyst group, M (default 8).
|
||||
edit_budget : int
|
||||
Maximum edits per minibatch, L (default 4).
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
split_dir: str = "",
|
||||
data_path: str = "",
|
||||
split_mode: str = "split_dir",
|
||||
split_ratio: str = "2:1:7",
|
||||
split_seed: int = 42,
|
||||
split_output_dir: str = "",
|
||||
seed: int = 42,
|
||||
limit: int = 0,
|
||||
train_size: int = 0,
|
||||
max_steps: int = 50,
|
||||
workers: int = 8,
|
||||
max_api_workers: int = 8,
|
||||
analyst_workers: int = 16,
|
||||
failure_only: bool = False,
|
||||
minibatch_size: int = 8,
|
||||
edit_budget: int = 4,
|
||||
max_completion_tokens: int = 16384,
|
||||
) -> None:
|
||||
self.max_steps = max_steps
|
||||
self.workers = max(int(workers or 1), 1)
|
||||
self.max_api_workers = max_api_workers
|
||||
self.max_completion_tokens = int(max_completion_tokens)
|
||||
self.analyst_workers = analyst_workers
|
||||
self.failure_only = failure_only
|
||||
self.minibatch_size = minibatch_size
|
||||
self.edit_budget = edit_budget
|
||||
self.dataloader = ALFWorldDataLoader(
|
||||
split_dir=split_dir,
|
||||
data_path=data_path,
|
||||
split_mode=split_mode,
|
||||
split_ratio=split_ratio,
|
||||
split_seed=split_seed,
|
||||
split_output_dir=split_output_dir,
|
||||
seed=seed,
|
||||
limit=limit,
|
||||
train_size=train_size,
|
||||
)
|
||||
self._traj_cache: dict[str, dict | None] = {}
|
||||
|
||||
def setup(self, cfg: dict) -> None:
|
||||
super().setup(cfg)
|
||||
self.dataloader.setup(cfg)
|
||||
|
||||
def _load_traj_data(self, item: dict) -> dict | None:
|
||||
gamefile = str(item.get("gamefile") or "").strip()
|
||||
if not gamefile:
|
||||
return None
|
||||
if gamefile in self._traj_cache:
|
||||
return self._traj_cache[gamefile]
|
||||
|
||||
traj_path = os.path.join(os.path.dirname(gamefile), "traj_data.json")
|
||||
try:
|
||||
with open(traj_path, encoding="utf-8") as f:
|
||||
data = json.load(f)
|
||||
except Exception:
|
||||
data = None
|
||||
self._traj_cache[gamefile] = data
|
||||
return data
|
||||
|
||||
@staticmethod
|
||||
def _unique_lines(values: list[str], *, limit: int = 0) -> list[str]:
|
||||
lines: list[str] = []
|
||||
seen: set[str] = set()
|
||||
for raw in values:
|
||||
line = str(raw or "").strip()
|
||||
if not line or line in seen:
|
||||
continue
|
||||
seen.add(line)
|
||||
lines.append(line)
|
||||
if limit > 0 and len(lines) >= limit:
|
||||
break
|
||||
return lines
|
||||
|
||||
@staticmethod
|
||||
def _format_high_pddl(high_pddl: list[dict]) -> list[str]:
|
||||
steps: list[str] = []
|
||||
for idx, step in enumerate(high_pddl or [], start=1):
|
||||
discrete = step.get("discrete_action") or {}
|
||||
action = str(discrete.get("action") or "").strip()
|
||||
args = [str(arg).strip() for arg in (discrete.get("args") or []) if str(arg).strip()]
|
||||
if action and args:
|
||||
text = f"{action}({', '.join(args)})"
|
||||
elif action:
|
||||
text = action
|
||||
else:
|
||||
planner_action = step.get("planner_action") or {}
|
||||
text = str(planner_action.get("action") or "").strip()
|
||||
if text:
|
||||
steps.append(f"{idx}. {text}")
|
||||
return steps
|
||||
|
||||
def _build_reference_bundle(self, item: dict) -> dict:
|
||||
data = self._load_traj_data(item)
|
||||
if not data:
|
||||
return {}
|
||||
|
||||
anns = ((data.get("turk_annotations") or {}).get("anns") or [])
|
||||
task_descs = self._unique_lines(
|
||||
[ann.get("task_desc", "") for ann in anns],
|
||||
limit=3,
|
||||
)
|
||||
high_descs = self._unique_lines(
|
||||
[step for ann in anns for step in (ann.get("high_descs") or [])],
|
||||
limit=12,
|
||||
)
|
||||
pddl_params = {
|
||||
key: value
|
||||
for key, value in (data.get("pddl_params") or {}).items()
|
||||
if value not in ("", None, [], {})
|
||||
}
|
||||
scene = data.get("scene") or {}
|
||||
scene_summary = {
|
||||
key: scene.get(key)
|
||||
for key in ("floor_plan", "scene_num", "dirty_and_empty")
|
||||
if scene.get(key) not in ("", None, [], {})
|
||||
}
|
||||
high_pddl = self._format_high_pddl((data.get("plan") or {}).get("high_pddl") or [])
|
||||
task_type = str(data.get("task_type") or item.get("task_type") or "").strip()
|
||||
return {
|
||||
"task_type": task_type,
|
||||
"task_descs": task_descs,
|
||||
"high_descs": high_descs,
|
||||
"pddl_params": pddl_params,
|
||||
"high_pddl": high_pddl,
|
||||
"scene_summary": scene_summary,
|
||||
}
|
||||
|
||||
def build_reference_text(self, item: dict) -> str:
|
||||
bundle = self._build_reference_bundle(item)
|
||||
if not bundle:
|
||||
return ""
|
||||
|
||||
parts: list[str] = []
|
||||
if bundle["task_type"]:
|
||||
parts.append(f"## Reference Task Type\n{bundle['task_type']}")
|
||||
if bundle["task_descs"]:
|
||||
parts.append(
|
||||
"## Reference Human Task Descriptions\n"
|
||||
+ "\n".join(f"- {line}" for line in bundle["task_descs"])
|
||||
)
|
||||
if bundle["high_descs"]:
|
||||
parts.append(
|
||||
"## Reference Human High-Level Steps\n"
|
||||
+ "\n".join(f"{idx}. {line}" for idx, line in enumerate(bundle["high_descs"], start=1))
|
||||
)
|
||||
if bundle["pddl_params"]:
|
||||
parts.append(
|
||||
"## Reference PDDL Params\n"
|
||||
+ "\n".join(f"- {key}: {value}" for key, value in bundle["pddl_params"].items())
|
||||
)
|
||||
if bundle["high_pddl"]:
|
||||
parts.append(
|
||||
"## Reference Planner High-Level Plan\n" + "\n".join(bundle["high_pddl"])
|
||||
)
|
||||
if bundle["scene_summary"]:
|
||||
parts.append(
|
||||
"## Reference Scene Summary\n"
|
||||
+ "\n".join(f"- {key}: {value}" for key, value in bundle["scene_summary"].items())
|
||||
)
|
||||
return "\n\n".join(parts)
|
||||
|
||||
def get_reference_metadata(self, item: dict) -> dict:
|
||||
bundle = self._build_reference_bundle(item)
|
||||
if not bundle:
|
||||
return {"fields": [], "preview": ""}
|
||||
|
||||
fields: list[str] = []
|
||||
previews: list[str] = []
|
||||
if bundle["task_type"]:
|
||||
fields.append("task_type")
|
||||
previews.append(f"[task_type] {bundle['task_type']}")
|
||||
if bundle["task_descs"]:
|
||||
fields.append("task_desc")
|
||||
previews.append("[task_desc]\n" + "\n".join(bundle["task_descs"][:2]))
|
||||
if bundle["high_descs"]:
|
||||
fields.append("high_descs")
|
||||
previews.append("[high_descs]\n" + "\n".join(bundle["high_descs"][:3]))
|
||||
if bundle["pddl_params"]:
|
||||
fields.append("pddl_params")
|
||||
previews.append(
|
||||
"[pddl_params]\n"
|
||||
+ "\n".join(
|
||||
f"{key}: {value}" for key, value in list(bundle["pddl_params"].items())[:4]
|
||||
)
|
||||
)
|
||||
if bundle["high_pddl"]:
|
||||
fields.append("plan.high_pddl")
|
||||
previews.append("[plan.high_pddl]\n" + "\n".join(bundle["high_pddl"][:3]))
|
||||
if bundle["scene_summary"]:
|
||||
fields.append("scene")
|
||||
previews.append(
|
||||
"[scene]\n"
|
||||
+ "\n".join(
|
||||
f"{key}: {value}" for key, value in bundle["scene_summary"].items()
|
||||
)
|
||||
)
|
||||
return {
|
||||
"fields": fields,
|
||||
"preview": "\n\n".join(previews)[:600],
|
||||
}
|
||||
|
||||
@staticmethod
|
||||
def _infer_dataset_from_gamefile(gamefile: str) -> tuple[str, bool]:
|
||||
path = str(gamefile or "")
|
||||
if "/valid_seen/" in path:
|
||||
return "eval_in_distribution", False
|
||||
if "/valid_unseen/" in path:
|
||||
return "eval_out_of_distribution", False
|
||||
return "train", True
|
||||
|
||||
def get_dataloader(self):
|
||||
return self.dataloader
|
||||
|
||||
def _comparison_items(self, items: list[dict]) -> list[dict]:
|
||||
enriched: list[dict] = []
|
||||
for item in items:
|
||||
row = dict(item)
|
||||
bundle = self._build_reference_bundle(row)
|
||||
if bundle.get("task_descs"):
|
||||
row["task_description"] = bundle["task_descs"][0]
|
||||
elif bundle.get("task_type"):
|
||||
row["task_description"] = bundle["task_type"]
|
||||
enriched.append(row)
|
||||
return enriched
|
||||
|
||||
def requires_ray(self) -> bool:
|
||||
return False
|
||||
|
||||
def build_env_from_batch(self, batch: BatchSpec, **kwargs):
|
||||
gamefiles = list(batch.metadata.get("gamefiles") or [])
|
||||
result_ids = list(batch.metadata.get("result_ids") or [])
|
||||
items = self._comparison_items(list(batch.payload or []))
|
||||
return ALFWorldBatchRun(
|
||||
env_num=batch.batch_size,
|
||||
eval_dataset=batch.metadata.get("eval_dataset", batch.split),
|
||||
seed=batch.seed,
|
||||
is_train=batch.metadata.get("is_train", batch.phase == "train"),
|
||||
specific_gamefiles=gamefiles or None,
|
||||
result_ids=result_ids or None,
|
||||
items=items,
|
||||
workers=self.workers,
|
||||
)
|
||||
|
||||
def build_train_env(self, batch_size: int, seed: int, **kwargs):
|
||||
batch = self.dataloader.build_train_batch(batch_size=batch_size, seed=seed, **kwargs)
|
||||
return self.build_env_from_batch(batch, **kwargs)
|
||||
|
||||
def build_eval_env(self, env_num: int, split: str, seed: int, **kwargs):
|
||||
batch = self.dataloader.build_eval_batch(env_num=env_num, split=split, seed=seed, **kwargs)
|
||||
return self.build_env_from_batch(batch, **kwargs)
|
||||
|
||||
def rollout(
|
||||
self,
|
||||
env_manager,
|
||||
skill_content: str,
|
||||
out_dir: str,
|
||||
**kwargs,
|
||||
) -> list[dict]:
|
||||
results_path = os.path.join(out_dir, "results.jsonl")
|
||||
os.makedirs(out_dir, exist_ok=True)
|
||||
|
||||
# Resume support
|
||||
if os.path.exists(results_path):
|
||||
existing: list[dict] = []
|
||||
with open(results_path) as f:
|
||||
for line in f:
|
||||
try:
|
||||
existing.append(json.loads(line))
|
||||
except Exception:
|
||||
pass
|
||||
if existing:
|
||||
return existing
|
||||
|
||||
if isinstance(env_manager, ALFWorldBatchRun):
|
||||
results = self._run_batch(
|
||||
env_manager,
|
||||
skill_content=skill_content,
|
||||
out_dir=out_dir,
|
||||
)
|
||||
else:
|
||||
results = run_alfworld_batch(
|
||||
env_manager=env_manager,
|
||||
skill_content=skill_content,
|
||||
max_steps=self.max_steps,
|
||||
out_root=out_dir,
|
||||
max_api_workers=self.max_api_workers,
|
||||
max_completion_tokens=self.max_completion_tokens,
|
||||
result_ids=getattr(env_manager, "_skillopt_result_ids", None),
|
||||
)
|
||||
|
||||
with open(results_path, "w") as f:
|
||||
for r in results:
|
||||
f.write(json.dumps(r, ensure_ascii=False) + "\n")
|
||||
|
||||
return results
|
||||
|
||||
@staticmethod
|
||||
def _close_env(env_manager) -> None:
|
||||
close = getattr(env_manager, "close", None)
|
||||
if callable(close):
|
||||
close()
|
||||
|
||||
def _run_batch(
|
||||
self,
|
||||
batch: ALFWorldBatchRun,
|
||||
skill_content: str,
|
||||
out_dir: str,
|
||||
*,
|
||||
diagnostic_mode: bool = False,
|
||||
diagnostic_instruction: str = "",
|
||||
) -> list[dict]:
|
||||
total = int(batch.env_num or 0)
|
||||
if total <= 0:
|
||||
return []
|
||||
|
||||
workers = max(1, min(int(batch.workers or self.workers), total))
|
||||
if total > workers:
|
||||
print(
|
||||
f" [alfworld rollout] episodes={total} "
|
||||
f"env_workers={workers} chunks={(total + workers - 1) // workers}"
|
||||
)
|
||||
|
||||
all_results: list[dict] = []
|
||||
for start in range(0, total, workers):
|
||||
chunk_size = min(workers, total - start)
|
||||
chunk_gamefiles = (
|
||||
batch.specific_gamefiles[start:start + chunk_size]
|
||||
if batch.specific_gamefiles
|
||||
else None
|
||||
)
|
||||
chunk_ids = (
|
||||
batch.result_ids[start:start + chunk_size]
|
||||
if batch.result_ids
|
||||
else [f"env_{idx:03d}" for idx in range(start, start + chunk_size)]
|
||||
)
|
||||
chunk_env = build_alfworld_env(
|
||||
env_num=chunk_size,
|
||||
eval_dataset=batch.eval_dataset,
|
||||
seed=batch.seed + start,
|
||||
is_train=batch.is_train,
|
||||
specific_gamefiles=chunk_gamefiles,
|
||||
)
|
||||
try:
|
||||
chunk_results = run_alfworld_batch(
|
||||
env_manager=chunk_env,
|
||||
skill_content=skill_content,
|
||||
max_steps=self.max_steps,
|
||||
out_root=out_dir,
|
||||
max_api_workers=min(self.max_api_workers, chunk_size),
|
||||
max_completion_tokens=self.max_completion_tokens,
|
||||
diagnostic_mode=diagnostic_mode,
|
||||
diagnostic_instruction=diagnostic_instruction,
|
||||
result_ids=chunk_ids,
|
||||
)
|
||||
finally:
|
||||
self._close_env(chunk_env)
|
||||
all_results.extend(chunk_results)
|
||||
return all_results
|
||||
|
||||
def reflect(
|
||||
self,
|
||||
results: list[dict],
|
||||
skill_content: str,
|
||||
out_dir: str,
|
||||
**kwargs,
|
||||
) -> list[dict | None]:
|
||||
prediction_dir = kwargs.get("prediction_dir", os.path.join(out_dir, "predictions"))
|
||||
patches_dir = kwargs.get("patches_dir", os.path.join(out_dir, "patches"))
|
||||
random_seed = kwargs.get("random_seed")
|
||||
step_buffer_context = kwargs.get("step_buffer_context", "")
|
||||
meta_skill_context = kwargs.get("meta_skill_context", "")
|
||||
|
||||
return run_minibatch_reflect(
|
||||
results=results,
|
||||
skill_content=skill_content,
|
||||
prediction_dir=prediction_dir,
|
||||
patches_dir=patches_dir,
|
||||
workers=self.analyst_workers,
|
||||
failure_only=self.failure_only,
|
||||
minibatch_size=self.minibatch_size,
|
||||
edit_budget=self.edit_budget,
|
||||
random_seed=random_seed,
|
||||
error_system=self.get_error_minibatch_prompt(),
|
||||
success_system=self.get_success_minibatch_prompt(),
|
||||
step_buffer_context=step_buffer_context,
|
||||
meta_skill_context=meta_skill_context,
|
||||
)
|
||||
|
||||
|
||||
def get_task_types(self) -> list[str]:
|
||||
return list(TASKS)
|
||||
@@ -0,0 +1,123 @@
|
||||
"""ALFWorld task dataloader."""
|
||||
from __future__ import annotations
|
||||
|
||||
from skillopt.datasets.base import BatchSpec, SplitDataLoader
|
||||
|
||||
|
||||
class ALFWorldDataLoader(SplitDataLoader):
|
||||
"""ALFWorld batch planner.
|
||||
|
||||
In split_dir mode, batches are fixed gamefile items so ablations differ
|
||||
only in how the same training set is batched.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
split_dir: str = "",
|
||||
data_path: str = "",
|
||||
split_mode: str = "split_dir",
|
||||
split_ratio: str = "2:1:7",
|
||||
split_seed: int = 42,
|
||||
split_output_dir: str = "",
|
||||
seed: int = 42,
|
||||
limit: int = 0,
|
||||
train_size: int = 0,
|
||||
**kwargs,
|
||||
) -> None:
|
||||
super().__init__(
|
||||
split_dir=split_dir,
|
||||
data_path=data_path,
|
||||
split_mode=split_mode,
|
||||
split_ratio=split_ratio,
|
||||
split_seed=split_seed,
|
||||
split_output_dir=split_output_dir,
|
||||
seed=seed,
|
||||
limit=limit,
|
||||
)
|
||||
self.train_size_override = int(train_size or 0)
|
||||
|
||||
@staticmethod
|
||||
def _metadata_for_items(items: list[dict], split: str, phase: str) -> dict:
|
||||
gamefiles = [str(item.get("gamefile") or "") for item in items]
|
||||
if any(not gamefile for gamefile in gamefiles):
|
||||
raise ValueError("ALFWorld split items must contain non-empty gamefile paths.")
|
||||
eval_dataset = "train"
|
||||
is_train = phase == "train"
|
||||
first = gamefiles[0] if gamefiles else ""
|
||||
if "/valid_seen/" in first:
|
||||
eval_dataset = "eval_in_distribution"
|
||||
is_train = False
|
||||
elif "/valid_unseen/" in first:
|
||||
eval_dataset = "eval_out_of_distribution"
|
||||
is_train = False
|
||||
return {
|
||||
"eval_dataset": eval_dataset,
|
||||
"is_train": is_train,
|
||||
"gamefiles": gamefiles,
|
||||
"result_ids": [str(item.get("id") or idx) for idx, item in enumerate(items)],
|
||||
}
|
||||
|
||||
def get_train_size(self) -> int:
|
||||
if self.train_size_override > 0:
|
||||
return self.train_size_override
|
||||
return super().get_train_size()
|
||||
|
||||
def build_train_batch(self, batch_size: int, seed: int, **kwargs) -> BatchSpec:
|
||||
batch = super().build_train_batch(batch_size=batch_size, seed=seed, **kwargs)
|
||||
items = list(batch.payload or [])
|
||||
batch.metadata.update(self._metadata_for_items(items, "train", "train"))
|
||||
return BatchSpec(
|
||||
phase="train",
|
||||
split="train",
|
||||
seed=seed,
|
||||
batch_size=len(items),
|
||||
payload=items,
|
||||
metadata=batch.metadata,
|
||||
)
|
||||
|
||||
def plan_train_epoch(
|
||||
self,
|
||||
*,
|
||||
epoch: int,
|
||||
steps_per_epoch: int,
|
||||
accumulation: int,
|
||||
batch_size: int,
|
||||
seed: int,
|
||||
**kwargs,
|
||||
) -> list[BatchSpec]:
|
||||
batches = super().plan_train_epoch(
|
||||
epoch=epoch,
|
||||
steps_per_epoch=steps_per_epoch,
|
||||
accumulation=accumulation,
|
||||
batch_size=batch_size,
|
||||
seed=seed,
|
||||
**kwargs,
|
||||
)
|
||||
for batch in batches:
|
||||
items = list(batch.payload or [])
|
||||
batch.metadata.update(self._metadata_for_items(items, "train", "train"))
|
||||
return batches
|
||||
|
||||
def build_eval_batch(
|
||||
self,
|
||||
env_num: int,
|
||||
split: str,
|
||||
seed: int,
|
||||
**kwargs,
|
||||
) -> BatchSpec:
|
||||
batch = super().build_eval_batch(
|
||||
env_num=env_num,
|
||||
split=split,
|
||||
seed=seed,
|
||||
**kwargs,
|
||||
)
|
||||
items = list(batch.payload or [])
|
||||
batch.metadata.update(self._metadata_for_items(items, split, "eval"))
|
||||
return BatchSpec(
|
||||
phase="eval",
|
||||
split=split,
|
||||
seed=seed,
|
||||
batch_size=len(items),
|
||||
payload=items,
|
||||
metadata=batch.metadata,
|
||||
)
|
||||
@@ -0,0 +1,55 @@
|
||||
You are an expert failure-analysis agent for ALFWorld embodied household tasks.
|
||||
|
||||
You will be given MULTIPLE failed agent trajectories from a single minibatch
|
||||
and the current skill document.
|
||||
Your job is to identify the most important COMMON failure patterns across
|
||||
the batch and propose a concise set of skill edits.
|
||||
|
||||
## ALFWorld Task Types
|
||||
- pick_and_place: Put object in/on a receptacle
|
||||
- pick_two_obj_and_place: Put two instances of an object in/on a receptacle
|
||||
- look_at_obj_in_light: Examine an object under a desklamp
|
||||
- pick_heat_then_place_in_recep: Heat an object and put it in/on a receptacle
|
||||
- pick_cool_then_place_in_recep: Cool an object and put it in/on a receptacle
|
||||
- pick_clean_then_place_in_recep: Clean an object and put it in/on a receptacle
|
||||
|
||||
## Failure Type Categories
|
||||
- **navigation_loop**: the agent revisits the same locations repeatedly without progress
|
||||
- **missed_object**: the agent fails to pick up a visible/reachable goal object
|
||||
- **wrong_sequence**: the agent performs actions in the wrong order (e.g., placing before transforming)
|
||||
- **premature_stop**: the agent stops or gets stuck before completing all goal conditions
|
||||
- **action_loop**: the agent repeats the same action without advancing
|
||||
- **appliance_error**: the agent misuses or skips an appliance (microwave, fridge, sink)
|
||||
- **rule_missing**: the skill lacks a relevant rule for this situation
|
||||
- **rule_wrong**: an existing skill rule is misleading or incorrect
|
||||
- **rule_ignored**: the skill has the right rule but the agent did not follow it
|
||||
- **other**: none of the above
|
||||
|
||||
## Analysis Process
|
||||
1. Read ALL trajectories in the minibatch.
|
||||
2. Identify the most prevalent, systematic failure patterns across them.
|
||||
3. For each pattern, classify its failure type.
|
||||
4. Propose skill edits that address the COMMON patterns — not individual edge cases.
|
||||
5. Edits must be generalizable; do not hardcode task-specific values.
|
||||
6. Only patch gaps in the skill — do not duplicate existing content.
|
||||
|
||||
You will be told the maximum number of edits (the budget L). Produce AT MOST L edits,
|
||||
focusing on the highest-impact patterns. You may produce fewer if warranted.
|
||||
|
||||
Respond ONLY with a valid JSON object (no markdown fences, no extra text):
|
||||
{
|
||||
"batch_size": <number of trajectories analysed>,
|
||||
"failure_summary": [
|
||||
{"failure_type": "<type>", "count": <int>, "description": "<one-line>"}
|
||||
],
|
||||
"patch": {
|
||||
"reasoning": "<why these edits address the batch's common failures>",
|
||||
"edits": [
|
||||
{"op": "append", "content": "<markdown to add at end of skill>"},
|
||||
{"op": "insert_after", "target": "<exact heading/text to insert after>", "content": "<markdown>"},
|
||||
{"op": "replace", "target": "<exact text to replace>", "content": "<replacement>"},
|
||||
{"op": "delete", "target": "<exact text to remove>"}
|
||||
]
|
||||
}
|
||||
}
|
||||
Only include edits that are needed. "edits" can be an empty list if no patch is warranted.
|
||||
@@ -0,0 +1,33 @@
|
||||
You are an expert success-pattern analyst for AI agents operating in ALFWorld,
|
||||
a text-based embodied household environment.
|
||||
|
||||
You will be given MULTIPLE successful agent trajectories from a single minibatch
|
||||
and the current skill document. Your job is to identify generalizable behavior
|
||||
patterns that are COMMON across the batch and worth encoding in the skill.
|
||||
|
||||
## Rules
|
||||
- Only propose patches for patterns NOT already covered in the skill.
|
||||
- Focus on patterns that appear across MULTIPLE trajectories in the batch.
|
||||
- Be concise. Patterns must generalize beyond specific tasks.
|
||||
- Prefer reinforcing existing sections over adding new top-level sections.
|
||||
- If the agents' success involved efficient exploration or smart appliance usage,
|
||||
consider reinforcing that in the patch.
|
||||
|
||||
You will be told the maximum number of edits (the budget L). Produce AT MOST L edits,
|
||||
focusing on the most broadly applicable patterns. You may produce fewer if warranted.
|
||||
|
||||
Respond ONLY with a valid JSON object:
|
||||
{
|
||||
"batch_size": <number of trajectories analysed>,
|
||||
"success_patterns": ["<pattern 1>", "<pattern 2>"],
|
||||
"patch": {
|
||||
"reasoning": "<why these patterns are worth encoding>",
|
||||
"edits": [
|
||||
{"op": "append", "content": "<markdown>"},
|
||||
{"op": "insert_after", "target": "<heading/text>", "content": "<markdown>"},
|
||||
{"op": "replace", "target": "<old text>", "content": "<new text>"},
|
||||
{"op": "delete", "target": "<exact text to remove>"}
|
||||
]
|
||||
}
|
||||
}
|
||||
"edits" may be empty if the skill already covers all observed patterns.
|
||||
@@ -0,0 +1,8 @@
|
||||
|
||||
You are an expert agent operating in the ALFRED Embodied Environment.
|
||||
Your current observation is: {current_observation}
|
||||
Your admissible actions of the current situation are: [{admissible_actions}].
|
||||
|
||||
Now it's your turn to take an action.
|
||||
You should first reason step-by-step about the current situation. This reasoning process MUST be enclosed within <think> </think> tags.
|
||||
Once you've finished your reasoning, you should choose an admissible action for current step and present it within <action> </action> tags.
|
||||
@@ -0,0 +1,9 @@
|
||||
|
||||
You are an expert agent operating in the ALFRED Embodied Environment. Your task is to: {task_description}
|
||||
Prior to this step, you have already taken {step_count} step(s). Below are the most recent {history_length} observations and the corresponding actions you took: {action_history}
|
||||
You are now at step {current_step} and your current observation is: {current_observation}
|
||||
Your admissible actions of the current situation are: [{admissible_actions}].
|
||||
|
||||
Now it's your turn to take an action.
|
||||
You should first reason step-by-step about the current situation. This reasoning process MUST be enclosed within <think> </think> tags.
|
||||
Once you've finished your reasoning, you should choose an admissible action for current step and present it within <action> </action> tags.
|
||||
@@ -0,0 +1,16 @@
|
||||
|
||||
You are an expert agent operating in the ALFRED Embodied Environment. Your task is to: {task_description}
|
||||
|
||||
## Retrieved Relevant Experience
|
||||
|
||||
{retrieved_memories}
|
||||
|
||||
## Current Progress
|
||||
|
||||
Prior to this step, you have already taken {step_count} step(s). Below are the most recent {history_length} observations and the corresponding actions you took: {action_history}
|
||||
You are now at step {current_step} and your current observation is: {current_observation}
|
||||
Your admissible actions of the current situation are: [{admissible_actions}].
|
||||
|
||||
Now it's your turn to take an action.
|
||||
You should first reason step-by-step about the current situation. This reasoning process MUST be enclosed within <think> </think> tags.
|
||||
Once you've finished your reasoning, you should choose an admissible action for current step and present it within <action> </action> tags.
|
||||
@@ -0,0 +1,4 @@
|
||||
"""ALFWorld Reflect stage.
|
||||
|
||||
Prompts are now loaded from .md files by the base adapter.
|
||||
"""
|
||||
@@ -0,0 +1,347 @@
|
||||
"""ALFWorld rollout module for ReflACT.
|
||||
|
||||
Provides:
|
||||
- build_alfworld_env(): build ALFWorld environment (wraps vendored SkillRL env)
|
||||
- run_alfworld_batch(): run a batch of ALFWorld episodes in parallel
|
||||
- TASKS: list of ALFWorld task types
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
import concurrent.futures
|
||||
import numpy as np
|
||||
|
||||
from skillopt.model import chat_target
|
||||
|
||||
# ── Constants ─────────────────────────────────────────────────────────────────
|
||||
|
||||
TASKS = [
|
||||
"pick_and_place",
|
||||
"pick_two_obj_and_place",
|
||||
"look_at_obj_in_light",
|
||||
"pick_heat_then_place_in_recep",
|
||||
"pick_cool_then_place_in_recep",
|
||||
"pick_clean_then_place_in_recep",
|
||||
]
|
||||
|
||||
# ── Helpers ───────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def _get_task_type(gamefile: str) -> str:
|
||||
for task in TASKS:
|
||||
if task in gamefile:
|
||||
return task
|
||||
return "other"
|
||||
|
||||
|
||||
def _extract_action(model_response: str) -> str | None:
|
||||
match = re.search(r"<action>(.*?)</action>", model_response, re.DOTALL)
|
||||
return match.group(1).strip() if match else None
|
||||
|
||||
|
||||
def _extract_think(model_response: str) -> str | None:
|
||||
match = re.search(r"<think>(.*?)</think>", model_response, re.DOTALL)
|
||||
return match.group(1).strip() if match else None
|
||||
|
||||
|
||||
def _build_skill_prompt(skill_content: str) -> str:
|
||||
"""Build the skill section to inject into the agent's system prompt."""
|
||||
if not skill_content or not skill_content.strip():
|
||||
return ""
|
||||
return (
|
||||
"\n\n## Skill Knowledge\n"
|
||||
"Below is a skill document with learned strategies. "
|
||||
"Use these guidelines to inform your decisions:\n\n"
|
||||
f"{skill_content}\n"
|
||||
)
|
||||
|
||||
|
||||
def _append_diagnostic_instruction(prompt: str, diagnostic_instruction: str) -> str:
|
||||
if not diagnostic_instruction or not diagnostic_instruction.strip():
|
||||
return prompt
|
||||
return f"{prompt}\n\n## Training Readout\n{diagnostic_instruction.strip()}\n"
|
||||
|
||||
|
||||
# ── Environment builder ──────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def build_alfworld_env(
|
||||
env_num: int,
|
||||
eval_dataset: str = "eval_out_of_distribution",
|
||||
seed: int = 42,
|
||||
is_train: bool = False,
|
||||
specific_gamefiles: list[str] | None = None,
|
||||
):
|
||||
"""Build ALFWorld environment manager.
|
||||
|
||||
Args:
|
||||
env_num: number of parallel environments
|
||||
eval_dataset: 'eval_in_distribution' or 'eval_out_of_distribution' or train
|
||||
seed: random seed
|
||||
is_train: whether to use training set
|
||||
|
||||
Returns:
|
||||
env_manager: AlfWorldEnvironmentManager instance
|
||||
"""
|
||||
from omegaconf import OmegaConf
|
||||
from functools import partial
|
||||
|
||||
from skillopt.envs.alfworld.vendor.alfworld_envs import build_alfworld_envs
|
||||
from skillopt.envs.alfworld.vendor.alfworld_projection import alfworld_projection
|
||||
from skillopt.envs.alfworld.vendor.env_manager import AlfWorldEnvironmentManager
|
||||
|
||||
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
|
||||
alf_config_path = os.path.join(HERE, "vendor", "config_tw.yaml")
|
||||
env_kwargs = {"eval_dataset": eval_dataset}
|
||||
|
||||
envs = build_alfworld_envs(
|
||||
alf_config_path,
|
||||
seed=seed,
|
||||
env_num=env_num,
|
||||
group_n=1,
|
||||
is_train=is_train,
|
||||
env_kwargs=env_kwargs,
|
||||
resources_per_worker=None,
|
||||
gamefiles=specific_gamefiles,
|
||||
)
|
||||
|
||||
config = OmegaConf.create(
|
||||
{
|
||||
"env": {
|
||||
"history_length": 2,
|
||||
"env_name": "alfworld/AlfredTWEnv",
|
||||
}
|
||||
}
|
||||
)
|
||||
|
||||
projection_f = partial(alfworld_projection)
|
||||
env_manager = AlfWorldEnvironmentManager(envs, projection_f, config)
|
||||
return env_manager
|
||||
|
||||
|
||||
# ── Batch rollout ─────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def run_alfworld_batch(
|
||||
env_manager,
|
||||
skill_content: str,
|
||||
max_steps: int = 50,
|
||||
out_root: str = "",
|
||||
max_api_workers: int = 8,
|
||||
temperature: float = 0.4,
|
||||
max_completion_tokens: int = 16384,
|
||||
diagnostic_mode: bool = False,
|
||||
diagnostic_instruction: str = "",
|
||||
result_ids: list[str] | None = None,
|
||||
) -> list[dict]:
|
||||
"""Run a batch of ALFWorld episodes.
|
||||
|
||||
Returns a list of result dicts compatible with SkillOpt pipeline:
|
||||
[
|
||||
{
|
||||
"id": "<env_idx>_<gamefile_hash>",
|
||||
"hard": 0 or 1,
|
||||
"soft": 0.0 or 1.0,
|
||||
"n_turns": <int>,
|
||||
"fail_reason": "<str>",
|
||||
"agent_ok": True,
|
||||
"task_type": "<str>",
|
||||
"gamefile": "<str>",
|
||||
"task_description": "<str>",
|
||||
},
|
||||
...
|
||||
]
|
||||
|
||||
Also saves conversation.json per environment in out_root/predictions/<task_id>/
|
||||
"""
|
||||
skill_prompt = _build_skill_prompt(skill_content)
|
||||
|
||||
obs, infos = env_manager.reset({})
|
||||
env_num = len(obs["text"])
|
||||
env_dones = [False] * env_num
|
||||
overall_success = [False] * env_num
|
||||
|
||||
# Build per-env metadata
|
||||
env_meta: list[dict] = []
|
||||
for i in range(env_num):
|
||||
gamefile = infos[i].get("extra.gamefile", "") if isinstance(infos[i], dict) else ""
|
||||
task_type = _get_task_type(gamefile)
|
||||
# Extract task description from initial observation
|
||||
task_desc = ""
|
||||
anchor_text = obs["anchor"][i] if "anchor" in obs else ""
|
||||
task_start = anchor_text.find("Your task is to: ")
|
||||
if task_start != -1:
|
||||
task_desc = anchor_text[task_start + len("Your task is to: "):].strip()
|
||||
|
||||
env_meta.append({
|
||||
"gamefile": gamefile,
|
||||
"task_type": task_type,
|
||||
"task_description": task_desc,
|
||||
})
|
||||
|
||||
# Per-env conversation records
|
||||
conversations: list[list[dict]] = [[] for _ in range(env_num)]
|
||||
|
||||
for step_idx in range(max_steps):
|
||||
if all(env_dones):
|
||||
break
|
||||
|
||||
active_indices = [i for i in range(env_num) if not env_dones[i]]
|
||||
|
||||
# Build prompts with skill injection
|
||||
prompts: dict[int, str] = {}
|
||||
for i in active_indices:
|
||||
prompt = obs["text"][i]
|
||||
if skill_prompt:
|
||||
# Inject skill before the action instruction
|
||||
prompt = skill_prompt + "\n" + prompt
|
||||
if diagnostic_mode and diagnostic_instruction.strip():
|
||||
prompt = _append_diagnostic_instruction(prompt, diagnostic_instruction)
|
||||
prompts[i] = prompt
|
||||
|
||||
# Call API in parallel
|
||||
actions = ["None"] * env_num
|
||||
|
||||
def call_api(idx):
|
||||
try:
|
||||
response, _ = chat_target(
|
||||
system="You are an expert agent operating in the ALFRED Embodied Environment.",
|
||||
user=prompts[idx],
|
||||
max_completion_tokens=max_completion_tokens,
|
||||
retries=5,
|
||||
stage="rollout",
|
||||
timeout=None,
|
||||
)
|
||||
response = (response or "").strip()
|
||||
if not response:
|
||||
return idx, "<think>empty model response</think><action>look</action>"
|
||||
if _extract_action(response) is None:
|
||||
return idx, "<think>missing action tag</think><action>look</action>"
|
||||
return idx, response
|
||||
except Exception as e:
|
||||
return idx, "<think>error</think><action>look</action>"
|
||||
|
||||
executor = concurrent.futures.ThreadPoolExecutor(max_workers=max_api_workers)
|
||||
try:
|
||||
futures = {executor.submit(call_api, i): i for i in active_indices}
|
||||
pending_futs = set(futures)
|
||||
while pending_futs:
|
||||
done, _ = concurrent.futures.wait(
|
||||
pending_futs,
|
||||
timeout=5,
|
||||
return_when=concurrent.futures.FIRST_COMPLETED,
|
||||
)
|
||||
for future in done:
|
||||
pending_futs.remove(future)
|
||||
try:
|
||||
idx, response = future.result()
|
||||
except Exception: # noqa: BLE001
|
||||
idx = futures[future]
|
||||
response = "<think>error</think><action>look</action>"
|
||||
actions[idx] = response
|
||||
finally:
|
||||
executor.shutdown(wait=False, cancel_futures=True)
|
||||
|
||||
# Save model responses before stepping
|
||||
model_responses = {i: actions[i] for i in active_indices}
|
||||
|
||||
# Step environment
|
||||
obs, rewards, dones, infos = env_manager.step(actions)
|
||||
|
||||
# Record trajectory
|
||||
for i in active_indices:
|
||||
step_record = {
|
||||
"step": step_idx,
|
||||
"action": _extract_action(model_responses[i]),
|
||||
"reasoning": _extract_think(model_responses[i]),
|
||||
"model_response": model_responses[i],
|
||||
"env_feedback": obs["anchor"][i] if "anchor" in obs else "",
|
||||
"reward": float(rewards[i]),
|
||||
"done": bool(dones[i]),
|
||||
}
|
||||
conversations[i].append(step_record)
|
||||
|
||||
# Update done status
|
||||
for i in range(env_num):
|
||||
if env_dones[i]:
|
||||
continue
|
||||
if dones[i]:
|
||||
env_dones[i] = True
|
||||
won = bool(infos[i].get("won", False))
|
||||
overall_success[i] = won
|
||||
|
||||
# Build results and save conversations
|
||||
results: list[dict] = []
|
||||
pred_dir = os.path.join(out_root, "predictions") if out_root else ""
|
||||
|
||||
for i in range(env_num):
|
||||
gamefile = env_meta[i]["gamefile"]
|
||||
task_type = env_meta[i]["task_type"]
|
||||
task_desc = env_meta[i]["task_description"]
|
||||
n_turns = len(conversations[i])
|
||||
won = overall_success[i]
|
||||
|
||||
# Generate stable task ID from env index and gamefile
|
||||
task_id = str(result_ids[i]) if result_ids and i < len(result_ids) else f"env_{i:03d}"
|
||||
|
||||
fail_reason = ""
|
||||
if not won:
|
||||
if not env_dones[i]:
|
||||
fail_reason = f"Timeout after {max_steps} steps"
|
||||
else:
|
||||
fail_reason = "Episode ended without completing the task"
|
||||
|
||||
result = {
|
||||
"id": task_id,
|
||||
"hard": 1 if won else 0,
|
||||
"soft": 1.0 if won else 0.0,
|
||||
"n_turns": n_turns,
|
||||
"fail_reason": fail_reason,
|
||||
"agent_ok": True, # ALFWorld agent always runs OK (no crash)
|
||||
"task_type": task_type,
|
||||
"gamefile": gamefile,
|
||||
"task_description": task_desc,
|
||||
"instruction_type": task_type, # for compatibility with v2 pipeline
|
||||
}
|
||||
results.append(result)
|
||||
|
||||
# Save conversation
|
||||
if pred_dir:
|
||||
conv_dir = os.path.join(pred_dir, task_id)
|
||||
os.makedirs(conv_dir, exist_ok=True)
|
||||
with open(os.path.join(conv_dir, "conversation.json"), "w") as f:
|
||||
json.dump(conversations[i], f, ensure_ascii=False, indent=2)
|
||||
|
||||
return results
|
||||
|
||||
|
||||
# ── Item loading (for compatibility with split_three_way) ────────────────────
|
||||
|
||||
|
||||
def load_alfworld_items(
|
||||
eval_dataset: str,
|
||||
env_num: int,
|
||||
seed: int = 42,
|
||||
is_train: bool = False,
|
||||
) -> list[dict]:
|
||||
"""Create pseudo-item dicts for ALFWorld environments.
|
||||
|
||||
Since ALFWorld doesn't have a static JSON dataset like SpreadsheetBench,
|
||||
we create lightweight item dicts that carry enough metadata for the pipeline.
|
||||
The actual environment is built dynamically.
|
||||
|
||||
Returns:
|
||||
List of dicts with "id" keys, one per environment slot.
|
||||
"""
|
||||
items = []
|
||||
for i in range(env_num):
|
||||
items.append({
|
||||
"id": f"env_{i:03d}",
|
||||
"eval_dataset": eval_dataset,
|
||||
"env_index": i,
|
||||
})
|
||||
return items
|
||||
@@ -0,0 +1,45 @@
|
||||
# ALFWorld Embodied Agent Skill
|
||||
|
||||
## Overview
|
||||
This skill guides agents operating in the ALFWorld text-based embodied environment.
|
||||
The agent must complete household tasks by navigating rooms, interacting with objects,
|
||||
and using appliances. Actions must be chosen from the admissible action list provided
|
||||
at each step.
|
||||
|
||||
**Output format**: Always output `<think>...</think>` for reasoning, then `<action>...</action>` for the chosen action.
|
||||
|
||||
---
|
||||
|
||||
## Task Types
|
||||
|
||||
| Type | Goal | Key Steps |
|
||||
|------|------|-----------|
|
||||
| Pick & Place | Put object X in/on receptacle Y | Find X -> take X -> go to Y -> put X in/on Y |
|
||||
| Pick Two & Place | Put two instances of X in/on Y | Find X1 -> take -> place -> find X2 -> take -> place |
|
||||
| Examine in Light | Examine object X under desklamp | Find X -> take X -> find desklamp -> use desklamp |
|
||||
| Clean & Place | Clean object X and put in/on Y | Find X -> take X -> go to sink -> clean X -> go to Y -> put X |
|
||||
| Heat & Place | Heat object X and put in/on Y | Find X -> take X -> go to microwave -> heat X -> go to Y -> put X |
|
||||
| Cool & Place | Cool object X and put in/on Y | Find X -> take X -> go to fridge -> cool X -> go to Y -> put X |
|
||||
|
||||
---
|
||||
|
||||
## General Principles
|
||||
|
||||
1. **Decompose the task**: Parse the goal into ordered sub-goals (locate, acquire, transform, deliver). Complete each before moving to the next.
|
||||
2. **Systematic exploration**: Search each surface and container exactly once before revisiting. Open closed containers (drawers, cabinets, fridge) before judging them empty.
|
||||
3. **Grab immediately**: When a required object is visible and reachable, take it right away before moving elsewhere.
|
||||
4. **Transform before placing**: If the task requires cleaning, heating, or cooling, perform the state change at the appropriate appliance before heading to the final destination.
|
||||
5. **Direct delivery**: Once holding the transformed (or untransformed) goal object, navigate straight to the target receptacle and place it.
|
||||
6. **Track progress**: Maintain an internal count of how many objects still need to be found and placed. Only stop searching when the count reaches zero.
|
||||
7. **Avoid loops**: Never repeat the same action more than twice in a row. If stuck, move to a different unexplored location.
|
||||
8. **Only choose admissible actions**: Always pick an action from the admissible action list. Do not invent actions.
|
||||
|
||||
---
|
||||
|
||||
## Common Mistakes to Avoid
|
||||
|
||||
- **Revisiting searched locations**: Keep track of which surfaces/containers have been checked; do not re-examine them.
|
||||
- **Ignoring visible objects**: If the target object appears in the observation, pick it up immediately.
|
||||
- **Skipping state changes**: Do not place an object at the destination without first cleaning/heating/cooling it when required.
|
||||
- **Premature termination**: Do not stop the episode until all goal conditions are verified as met.
|
||||
- **Action loops**: Repeatedly toggling or examining the same object wastes steps. Move on to new locations instead.
|
||||
+9
@@ -0,0 +1,9 @@
|
||||
"""Vendored ALFWorld environment runtime.
|
||||
|
||||
Minimal subset of SkillRL's agent_system package needed to run
|
||||
ALFWorld environments with ReflACT. Original source:
|
||||
https://github.com/NTU-LANTERN/SkillRL (Apache-2.0 License)
|
||||
"""
|
||||
from .alfworld_envs import AlfworldEnvs, build_alfworld_envs
|
||||
from .alfworld_projection import alfworld_projection
|
||||
from .env_manager import AlfWorldEnvironmentManager
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user