Files
colibri/c/tests/test_decode_batch.c
FABIOTESS d7b855f43f serve stage 2: per-request grammars end-to-end — response_format -> SUBMIT -> grammar-forced drafts in the multiplexed server
The OpenAI gateway previously 400'd every response_format; the mux engine path
ran with speculation disabled entirely. Now:

- openai_server.py: response_format {"type":"json_object"} (generic ws-tolerant
  JSON grammar), {"type":"json_schema"} (schema forwarded as-is, compiled
  engine-side by schema_gbnf.h), and a raw-GBNF {"type":"gbnf"} extension.
  Draft-source semantics throughout: a schema the engine cannot compile costs
  the speedup, never the request and never the output.
- SUBMIT protocol: optional 7th field gbytes; grammar text appended to the
  payload after the prompt. 6-field headers unchanged (back-compatible).
- Engine: per-slot GrDraft (grammar_setup_text/grammar_teardown split out of
  the env-driven setup); walkers fed on every emitted token. Grammar-forced
  drafting in run_serve_mux for greedy requests: a drafting slot leaves the
  shared batch for one forward and runs the proven single-sequence verify path
  (kv_bind + step_all) — the same primitives prefill already uses per
  submission — then rejoins; rejected drafts' KV entries are overwritten by the
  next forward exactly like the existing prefix-truncation path. Sampling
  requests never draft (verification under sampling needs rejection resampling;
  out of scope).

Tests: 7-field SUBMIT parse cases; response_format->grammar plumbing incl.
fail cases; test doubles updated; generic JSON grammar parse+walk validated
against grammar.h. make test-c green; python suite green except the known
environmental memory_available failure (#150 fixes it).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 17:27:33 +01:00

63 lines
2.2 KiB
C

#include <assert.h>
#include <stdio.h>
#include "../decode_batch.h"
static void test_rows_use_their_own_sequence_storage(void)
{
float sequence_a[4 * 3] = {0};
float sequence_b[4 * 3] = {0};
float *a2 = coli_kv_row(sequence_a, 2, 3);
float *b1 = coli_kv_row(sequence_b, 1, 3);
a2[0] = 20.0f;
b1[2] = 12.0f;
assert(a2 == &sequence_a[6]);
assert(b1 == &sequence_b[3]);
assert(sequence_a[6] == 20.0f);
assert(sequence_b[5] == 12.0f);
assert(sequence_a[5] == 0.0f);
assert(sequence_b[6] == 0.0f);
}
static void test_const_reader_selects_the_same_row(void)
{
float storage[5 * 7] = {0};
const float *row = coli_kv_row(storage, 4, 7);
assert(row == &storage[28]);
}
static void test_submit_header(void)
{
ColiSubmit sub;
assert(coli_submit_parse("SUBMIT 42 3 17 64 0.7 0.95", &sub));
assert(sub.id == 42 && sub.slot == 3 && sub.bytes == 17);
assert(sub.max_tokens == 64 && sub.temperature > .69f && sub.top_p > .94f);
assert(!coli_submit_parse("SUBMIT 1 -1 2 3 0.7 1", &sub));
assert(!coli_submit_parse("SUBMIT 1 0 2 0 0.7 1", &sub));
assert(!coli_submit_parse("SUBMIT 1 0 2 3 4 1", &sub));
assert(!coli_submit_parse("SUBMIT 0 0 2 3 1 1", &sub));
assert(!coli_submit_parse("SUBMIT 1 0 2 3 nan 1", &sub));
assert(!coli_submit_parse("SUBMIT 1 0 2 3 1 inf", &sub));
assert(coli_submit_parse("SUBMIT 1 0 16777216 3 1 1", &sub));
assert(!coli_submit_parse("SUBMIT 1 0 16777217 3 1 1", &sub));
assert(!coli_submit_parse("SUBMIT 1 0 2 3 1 1 trailing", &sub));
/* optional 7th field: per-request grammar length (0 when absent) */
assert(coli_submit_parse("SUBMIT 42 3 17 64 0.7 0.95", &sub) && sub.gbytes == 0);
assert(coli_submit_parse("SUBMIT 42 3 17 64 0.7 0.95 512", &sub) && sub.gbytes == 512);
assert(coli_submit_parse("SUBMIT 42 3 17 64 0.7 0.95 1048576", &sub));
assert(!coli_submit_parse("SUBMIT 42 3 17 64 0.7 0.95 1048577", &sub));
assert(!coli_submit_parse("SUBMIT 42 3 17 64 0.7 0.95 512 extra", &sub));
}
int main(void)
{
test_rows_use_their_own_sequence_storage();
test_const_reader_selects_the_same_row();
test_submit_header();
puts("decode batch helper tests: ok");
return 0;
}