Skip to content

Latest commit

 

History

History
110 lines (91 loc) · 5.04 KB

File metadata and controls

110 lines (91 loc) · 5.04 KB

v0.71.6 manual QA log — synth data + build runner + IRT (#75 + step-6 smokes)

Environment: Windows 11 + RTX 3050 4 GB (Ampere). Ollama daemon local; model pulled for the live smokes: qwen2.5:0.5b (≈398 MB, runs on CPU/GPU). Date: 2026-06-02.

This log records the manual end-to-end smokes for the v0.71.6 patch (closes #231, #232, #167, #213, #75). Automated coverage lives in tests/test_v0716.py; this file is the live-provider / real-CLI evidence that pytest green ≠ "works on a real soup.yaml".


#75 — synth data QA (data from-traces / data augment / data generate)

soup data from-traces (no provider — offline) — PASS

Constructed an OpenAI-trace JSONL with a thumbs-up and a thumbs-down output for the same prompt:

soup data from-traces --logs traces.jsonl --format openai --signal thumbs_up -o pairs.jsonl

Wrote 1 preference pair(s); the pair is a correct {prompt, chosen, rejected, source} row. PII reminder panel printed. ✅

soup data generate (live Ollama) — PASS

soup data generate --provider ollama --model qwen2.5:0.5b -p "Write a short factual Q&A about Python lists" -n 1 -o gen.jsonl

Generated 2 examples, alpaca format, sane content (list-vs-tuple Q&A). ✅

soup data augment (live Ollama) — BUG FOUND + FIXED

First run raised ImportError: cannot import name 'OllamaProvider' from soup_cli.data.providers.ollama. Root cause: _load_augment_provider (commands/data.py) constructed OllamaProvider / AnthropicProvider / VllmProvider classes that never existed — the provider modules ship functions (generate_ollama, validate_ollama_url, …). So data augment was broken for all three providers since v0.25.0, AND it hardcoded llama3.1:8b with no --model flag.

Fix (v0.71.6 #75): _load_augment_provider now returns an _AugmentProvider adapter that delegates to the v0.53.7 hardened make_judge_provider_fn (SSRF-validated URLs, env-only Anthropic key) and unwraps {"text": …} to the .generate(prompt) -> str contract the augment strategies expect. Added --model / --base-url options. After the fix:

soup data augment -i in.jsonl -o out.jsonl --strategy rephrase --provider ollama --model qwen2.5:0.5b --count 1

Augmentation complete: 1 → 2 (rephrase via ollama). ✅ Pipeline runs end-to-end. (Observation, not a bug: qwen2.5:0.5b is small enough that its rephrase output sometimes echoes the input verbatim — a model-quality artefact, not a wiring issue. A larger model produces visibly varied paraphrases.)

Regression coverage: tests/test_v0716.py::TestAugmentProviderFix.


#231 — live soup build runner — PASS

3-stage DAG (raw incremental seed → clean incremental drop_empty → counted table token_count):

First run materialised raw (3 rows), clean (2 rows after drop_empty), counted (n_tokens added). Re-run after changing one source row showed the incremental optimisation working:

raw     incremental  rows_in=3 rows_out=3 transformed=1  +0 ~1 -0 =2
clean   incremental  rows_in=3 rows_out=3 transformed=1  +0 ~1 -0 =2
counted table        rows_in=3 rows_out=3 transformed=3  —

transformed=1 proves only the changed row was re-run; the 2 unchanged rows carried over from the SQLite state store. ✅


#232 — live Magpie generator — PASS

soup data gen-magpie --base qwen2.5:0.5b --provider ollama --target 2 --output magpie.jsonl --no-quality-filter --max-tokens 96

Rows kept: 2, Filtered: 0, Duplicates: 0, Attempts: 2. The output rows are real {messages:[user, assistant]} pairs harvested via the chat-template prefix (ChatML for the Qwen family) — the model generated diverse user instructions from the bare <|im_start|>user\n prefix and corresponding assistant responses through raw /api/generate completion. ✅ Genuine prefix-harvest Magpie against a live model.


#213 — 2PL / 3PL IRT — PASS

Response matrix: 6 respondents × {const (always correct), split (clean ability threshold), easy}.

soup eval irt-subset resp.jsonl --size small --model 2pl   → selects ["split"]
soup eval irt-subset resp.jsonl --size tiny  --model 3pl   → runs, guessing bounded

The 2PL fit correctly selects the most-informative item (split, the balanced threshold) over the zero-information const item. Missing respondent_id on a 2PL/3PL run fails with a friendly IRT fit failed: rows[0] missing respondent_id (exit 2). 1PL stays on the {item_id, correct} single-respondent path (back-compat). ✅


#167 — tokenizer-aware memorization probe — PASS (helper)

split_prefix / score_memorization now accept a tokenizer (HF id/path or a duck-typed object) and split / overlap on sub-word token ids instead of whitespace words. Exercised in tests/test_v0716.py::TestSplitPrefixTokenizer

  • TestScoreMemorizationTokenizer with a char-level fake tokenizer (exact round-trip). The live soup diagnose --tokenizer wiring lands with the live probe runner (#165, v0.71.7) — this patch ships the tokenizer-aware variant of the probe helper that #165 will call.