Data formats, the Axolotl/LF-parity pipeline, data tools, synthetic generation/forge, quality scorecards, trace tooling, remote datasets, mixing, recipe DAGs, and the v0.69 data-engineering surfaces.
Contents:
- Data Engineering Pro
- Production Trace Ecosystem (
soup ingest) - Prompt Mining (
soup prune-prompt) - Active-Learning Sampler (
soup data active-sample) - Synthetic Data Generation
- Data Augmentation
- Trace-to-Preference
- Config Migration
- Data Formats
- Data Pipeline Pro
- Data Tools
- Demo Datasets (
soup data demo) - Trace-to-Preference: LLM-Judge Filter
- Synthetic Data Forge
- Data Quality Scorecard
- Remote Datasets (S3 / GCS / Azure / OCI)
- Semantic dedup (
soup data dedup --semantic) - Topic map (
soup data topics) - Canaries (
soup data canary insertcheck) - Data Recipe DAG
- Data Mixing Optimizer (BETA)
- AOT Tokenization with
soup data preprocess - Data Recipe DAG Runner (
soup data recipe --execute)
soup data dedup removes near-duplicates with MinHash by default — fast, no
torch, but lexical: it compares shared token shingles, so two rows that say
the same thing in different words look unrelated to it.
--semantic compares embedding cosine instead:
soup data dedup train.jsonl --semantic -o clean.jsonl
soup data dedup train.jsonl --semantic --threshold 0.85 --field text -o clean.jsonl
soup data dedup train.jsonl --semantic --embed-model sentence-transformers/all-mpnet-base-v2Requires the [train] extra (it reuses transformers; there is no new
dependency) and downloads a small embedding model on first use. Plain MinHash
dedup stays on the light core.
What it buys you. Measured against MinHash on the same rows (all-MiniLM-L6-v2):
| pair | cosine | MinHash | --semantic |
|---|---|---|---|
| exact duplicate | 1.000 | caught | caught |
| "sorts a list of integers" / "sorts an array of integers" | 0.908 | missed | caught |
| "which sorts a list of ints" (reworded) | 0.880 | missed | caught |
| "Add two numbers" / "Multiply two numbers" | 0.759 | kept | kept (correct) |
So --semantic catches rewordings MinHash's shingling scores as distinct.
Heavier paraphrases are not reliably separable. Measured, paraphrase cosines (0.49–0.76) overlap with genuinely-distinct rows (0.54–0.76):
- "reverse a string" / "invert the order of characters" — a true paraphrase — scores 0.491
- "Add two numbers" / "Multiply two numbers" — two rows you must keep — scores 0.759
A real paraphrase can score lower than two rows that must both survive, so no
threshold cleanly separates them. Lowering --threshold to chase paraphrase
recall deletes real training rows — silent data loss, which is worse than keeping
a duplicate. The 0.8 default is deliberately conservative. Raise or lower it only
against your own data, and check what got dropped.
--threshold means Jaccard for MinHash and cosine for --semantic. They are
different scales; a value tuned for one is not meaningful for the other.
See what you are actually training on:
soup data topics train.jsonl # 'auto' picks the cluster count
soup data topics train.jsonl --clusters 8 -o topics.jsonEmbeds every row, clusters with k-means, and labels each cluster with c-TF-IDF terms — terms frequent in that cluster and rare elsewhere, so filler words like "the" never become a label. Prints a coverage table plus a warning for any topic under 2% of the data:
Topic map — 4200 rows, 6 clusters
┏━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━┳━━━━━━━━━━┓
┃ Topic ┃ Rows ┃ Coverage ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━╇━━━━━━━━━━┩
│ function / python / code │ 3444 │ 82.0% │
│ theorem / proof / math │ 252 │ 6.0% │
│ refuse / harmful / safe │ 42 │ 1.0% │
└──────────────────────────┴──────┴──────────┘
topic 'refuse / harmful / safe' is thin: 1.0% of rows (42/4200)
Labels are emergent term clusters, not a classification against a fixed
taxonomy: "82% code" means 82% of rows landed in a cluster whose top terms look
like code. Requires [train].
Prove whether a model memorized your data — for leak detection and provenance.
# 1. insert K unique secrets (keep the manifest OUT of your repo)
soup data canary insert train.jsonl -o canaried.jsonl --count 16 --manifest secrets.json
# 2. train on canaried.jsonl as usual, then:
soup data canary check --manifest secrets.json --base ./my-model --adapter ./loracheck measures the model's loss on each inserted secret and ranks it against
never-inserted controls drawn from the same secret space and sharing the same
carrier prompt — so a low loss means the secret is unusually likely, not the
prompt. Exit 2 on MAJOR, so CI can gate on a leak.
This is loss-vs-controls (Carlini et al., The Secret Sharer), not "ask the model and see if it says the secret": a model can memorize a canary and still not emit it under greedy decoding, so "nothing came back" would be false reassurance.
The verdict asks whether more canaries look memorized than chance explains (binomial tail, α=0.05) rather than whether any single one dipped low — with 16 canaries, an "any single one" rule fires on a clean model about 15% of the time, which would make a CI gate useless.
Measured on SmolLM2-135M:
| model | loss | percentiles | verdict |
|---|---|---|---|
| trained on the canaries | 1.7–2.5 | all 0.0% | MAJOR (exit 2) |
| never saw them | 4.1–6.2 | 1.6%–93% | OK (exit 0) |
The manifest is the sensitive artifact, not the dataset: anyone holding it can
reproduce the secrets. It is written 0600 on POSIX and must not be committed
alongside the data it protects. check --output embeds the same secrets.
Exposure is a sampled-control approximation, not full-space rank enumeration — "no exposure" is not proof of no memorization.
The v0.69.0 release ships 5 surfaces that turn dataset prep from "throw a JSONL at the trainer" into a first-class engineering workflow.
# dbt-for-SFT — DAG of dataset transforms with incremental materialization
cat > build.yaml << 'EOF'
models:
- {name: raw, kind: incremental, source: data/raw.jsonl, transform: identity}
- {name: filtered, kind: incremental, refs: [raw], transform: filter_low_quality}
- {name: tokenized, kind: incremental, refs: [filtered], transform: tokenize}
EOF
soup build build.yaml --dry-run # validate topology + plan
soup build build.yaml --output-dir built/ # live materialise (v0.71.6)
# Expectations suite — Great Expectations for chat data
cat > suite.yaml << 'EOF'
expectations:
- {name: expect_no_pii}
- {name: expect_token_length_between, args: {min_tokens: 16, max_tokens: 4096}}
- {name: expect_no_refusal_pattern}
EOF
soup expect data.jsonl suite.yaml # exit 3 on suite failure
# Magpie synthetic data — chat-template-prefix harvest (live, v0.71.6)
soup data gen-magpie --base meta-llama/Llama-3.1-8B-Instruct \
--provider ollama --target 1000 --output magpie.jsonl --quality-filter
# Persona-Hub diversity — prompt × persona × style matrix sampling
soup data persona-mix --prompts prompts.jsonl --n 500 --output mixed.jsonl
# Brain-rot detector (arXiv 2510.13928) — refuses to train on excessive slop
soup data brain-rot data.jsonl --strict --max-major-fraction 0.10
# Best-of-N rejection sampling — local sampling stays the default
soup data best-of-n --base HuggingFaceTB/SmolLM2-135M-Instruct \
--prompts prompts.jsonl --n 8 --judge ollama://llama3.1 \
-o best_of_n.jsonl --emit-pairs pairs.jsonl
# Or draw the N candidates from a running Ollama / vLLM raw-completion endpoint
soup data best-of-n --provider ollama --model qwen2.5:7b \
--base-url http://localhost:11434 --prompts prompts.jsonl --n 8 \
--judge ollama://llama3.1 -o best_of_n.jsonl
# Every non-blank prompt row is validated; accepted SFT rows record source_line
# in their _best_of_n provenance so input/output completeness can be checked.
# If sampling or judging stops, continue from the last fsynced prompt group.
soup data best-of-n --provider ollama --model qwen2.5:7b \
--prompts prompts.jsonl --n 8 --judge ollama://llama3.1 \
-o best_of_n.jsonl --resume
# Two-phase / air-gapped workflow: sample first, without constructing a judge.
soup data best-of-n --base Qwen/Qwen3.8-27B --revision <commit> \
--prompts prompts.jsonl --n 8 --export-candidates candidates.jsonl
# An offline human, deterministic program, CI job, or Codex writes one judgment
# per candidate group, copying prompt_id and group_digest from candidates.jsonl:
# {"prompt_id":"...","group_digest":"...","winner_idx":2,
# "scores":[0.1,0.4,0.9,...],"verifier":{"name":"Codex","version":"offline-v1"}}
soup data best-of-n --candidate-artifact candidates.jsonl \
--judgments judgments.jsonl -o best_of_n.jsonl --emit-pairs pairs.jsonl
# Evol-Instruct (WizardLM depth/breadth, v0.71.31) — grow instruction diversity
soup data evolve --input seeds.jsonl --provider ollama --model llama3.1 \
--strategy depth --rounds 2 -o evolved.jsonlEvery command applies the project-wide TOCTOU policy (os.lstat + S_ISLNK symlink rejection before any open) and cwd containment via the shared paths.enforce_under_cwd_and_no_symlink helper. All five are LIVE: soup build materialises with five built-in transforms (identity / drop_empty / lowercase / strip / dedup_exact) and SQLite-tracked incremental re-transform (v0.71.6); soup data gen-magpie and provider-backed best-of-n harvest via raw completion against --provider ollama|vllm (SSRF-validated; anthropic rejected because the Messages API has no raw-completion endpoint). Provider-backed best-of-n records the sampler provider and model in each row's _best_of_n provenance; omit --provider to retain the local Transformers --base path.
best-of-n fsyncs each completed prompt group to a private recovery journal
(<output>.checkpoint.jsonl by default). --resume reuses only a sequential
prefix whose prompt and run-configuration digest matches exactly, so completed
prompts are not sampled or judged twice. Before final publication, Soup snapshots
any prior SFT, DPO, and manifest targets. Each file remains an atomic replacement,
the manifest is written last, and any failed replacement restores the complete old
generation or removes the newly created set. The manifest binds the exact SHA-256
hashes and row counts as one generation. Keep or archive the checkpoint after
success if reproducible rematerialization is useful; it contains dataset content
and should be protected like the generated dataset.
The candidate artifact preserves every ordered candidate with prompt, candidate,
group, and whole-artifact SHA-256 bindings plus a public sampler specification.
The offline phase validates complete one-to-one coverage before writing anything:
missing or duplicate prompt ids, changed group digests, invalid winner indexes,
score-count drift, non-finite scores, and winner/score disagreement all fail
closed. It does not construct a sampler or judge. SFT and DPO rows retain the
candidate-artifact and judgment-file digests, group id, public sampler settings,
and bounded verifier identity. Endpoint URLs and local model paths are never
copied into those artifacts. Reusing the same two input files produces identical
training-row bytes. The offline command writes <output>.manifest.json last (or
the explicit --manifest path) and binds the exact SFT/DPO hashes, row counts,
candidate artifact, judgment file, and whether DPO output was requested. Treat
the manifest as the commit marker: missing or mismatched manifests identify an
interrupted or replaced generation.
The candidate artifact and verified judgment file are the durable recovery
boundary for offline materialization. This phase performs no sampling, so it has
no progress checkpoint: --resume and --checkpoint are rejected. After any
late write failure, keep those two inputs and rerun the exact offline command.
Soup removes the previous manifest before replacing outputs and publishes the
new manifest last. A prior, manifest-authenticated DPO beside the manifest is
removed when the replacement run requests SFT only.
Consumers must verify the final manifest and open exactly the SFT/DPO files it lists. They must never discover training inputs by globbing neighboring JSONL files: an unlisted sidecar, including an older DPO stored elsewhere, is not part of the committed generation.
Candidate export durably checkpoints each completed prompt group at
<artifact>.checkpoint.jsonl. If sampling stops, rerun the same command with
--resume; Soup authenticates the checkpoint against the prompts and sampler
before continuing at the first incomplete group. Candidate and judgment inputs
are validated through a temporary disk index, and final SFT/DPO files are staged
incrementally, so memory does not grow with the complete artifact size.
For a local model directory, the checkpoint binds the exact regular-file names,
sizes, and contents through a privacy-safe fingerprint; replacing weights at the
same path therefore invalidates resume before the model is loaded. Prompt source
lines and provider endpoints are bound as well without exposing private paths or
URLs. Streamed SFT/DPO replacements are committed as one rollback-protected set,
and an SFT-only replacement retires a prior manifest-bound DPO in that same
transaction.
Use a dotted-path string (module.path:function_name) as the transform
value to import a custom transform at build time:
models:
- {name: clean, kind: table, source: data/raw.jsonl, transform: my_pkg.transforms:clean_row}
- {name: enriched, kind: table, refs: [clean], transform: my_pkg.transforms:enrich}The target function must accept exactly two positional arguments (row, config)
and return a dict or None. Soup resolves the dotted path lazily (the module
is imported only when the build actually runs) and caches the result so repeated
references to the same path do not re-import.
Trusted-input posture. The dotted-path syntax causes Soup to import an
arbitrary Python module and call a function from it. Treat transform values
as trusted input: do not feed untrusted or operator-controlled YAML into soup build on shared CI hosts. An attacker who controls the manifest can execute
arbitrary code during the build phase. If you must accept user-supplied manifests,
validate them against a allowlist of permitted transform paths before passing
them to the resolver.
Closing the data flywheel without leaving your existing observability stack. soup ingest parses JSONL exports from every major SaaS dashboard and emits a normalised trace stream that soup data from-traces (v0.26) consumes.
# Six supported sources — adapters for the major SaaS vendors + raw OTel
soup ingest --source langfuse --logs ./langfuse-export.jsonl --output traces.jsonl
soup ingest --source langsmith --logs ./langsmith-runs.jsonl
soup ingest --source helicone --logs ./helicone-requests.jsonl
soup ingest --source openpipe --logs ./openpipe-export.jsonl
soup ingest --source otel --logs ./otel-spans.jsonl
soup ingest --source openai-stored --logs ./oai-stored-completions.jsonlThe CLI never makes the network call — operators export from their SaaS dashboard or vendor API, then point soup ingest at the local file. Auth env vars (LANGFUSE_KEY / LANGSMITH_API_KEY / HELICONE_API_KEY / OPENPIPE_API_KEY / OPENAI_API_KEY / OTEL_EXPORTER_OTLP_HEADERS) are advisory only — Soup surfaces which one is unset so operators wire creds before the SaaS-side export. A PII reminder fires on every ingest run (matches v0.26.0 Trace-to-Preference policy).
Production LLM apps often pin a multi-paragraph system prompt to every request. Fine-tuning with that prefix wastes tokens (the model learns to copy what's already in context). soup prune-prompt finds the longest character prefix shared by ≥ 95% of rows and strips it, so the FT model internalises the behaviour instead.
soup prune-prompt --input traces.jsonl --output pruned.jsonl --min-frequency 0.95Binary-search over up-to-32 candidate templates finds the longest qualifying prefix (a longer threshold-meeting prefix may exist beyond the universal one — Soup does not early-exit on the 100% match). Two-pass file read with a 100 000-row DoS cap.
Tokenizer-aware mode (v0.71.5). Pass --tokenizer <id-or-path> (a HuggingFace repo id, a local path, or anything AutoTokenizer.from_pretrained accepts) to detect the shared prefix in token space and decode only the remaining ids:
soup prune-prompt --input traces.jsonl --output pruned.jsonl --tokenizer Qwen/Qwen2.5-0.5BChar-level stripping can cut a BPE multi-byte sequence in half when the shared prefix ends mid-token; token-aware pruning finds the longest shared token-id prefix and decodes the remainder, so the boundary always lands on a real token. Per-row encoding is capped at 50 000 tokens. Omit --tokenizer to keep the original character-level behaviour.
Surface the most uncertain prod traces for human review. Two modes via the input data shape:
- Single RM:
rm_score: 0.5→ uncertainty 1.0 (peak);rm_score: 0.0or1.0→ uncertainty 0.0. - Dual RM:
rm_scores: [s1, s2]→ uncertainty =|s1 - s2|(pairwise disagreement).
soup data active-sample --input traces.jsonl --output for-review.jsonl --budget 100The output JSONL is a drop-in prompt set for soup eval human (v0.19). Budget is bounded [1, 100 000].
Webhooks (v0.71.5). soup ingest, soup prune-prompt, soup ab, and soup data active-sample all accept --slack-url / --discord-url and POST a one-line summary on completion through the same SSRF-hardened validator as soup drift-alarm (scheme allowlist, loopback-only HTTP, RFC1918 / link-local / reserved / multicast rejected; the post never raises, so a flaky webhook can't fail the command). soup ab only fires when the sequential test actually decides (reject_h0 / accept_h0), not while it's still continue-ing.
Generate training data using LLMs:
# Generate using OpenAI API
soup data generate --prompt "Create math word problems" --count 100 --format alpaca
# Use a different model
soup data generate --prompt "Medical Q&A pairs" --model gpt-4o --count 500
# Deduplicate against existing data
soup data generate --prompt "..." --count 200 --dedup-with existing.jsonl
# Use seed examples to guide style
soup data generate --prompt "..." --seed examples.jsonl --count 100
# Use a local OpenAI-compatible server (soup serve, Ollama, etc.)
soup data generate --prompt "..." --provider server --api-base http://localhost:11434/v1# Generate via local Ollama instance
soup data generate --prompt "..." --provider ollama --model llama3.1
soup data generate --prompt "..." --ollama-model llama3.1 # shorthand
# Generate via Anthropic Claude API (set ANTHROPIC_API_KEY env var)
soup data generate --prompt "..." --provider anthropic --model claude-3-haiku-20240307
# Generate via local vLLM server
soup data generate --prompt "..." --provider vllm --model meta-llama/Llama-3.1-8B-Instruct# Code instruction pairs (Python, JS, Go, Rust, Java)
soup data generate --prompt "..." --template code --language Python --task-type function
# Multi-turn conversations
soup data generate --prompt "..." --template conversation --turns 6 --topic "science"
# QA from context document
soup data generate --prompt "..." --template qa --context document.txt
# Preference data (DPO/KTO/ORPO)
soup data generate --prompt "..." --template preference --pref-task dpo
# Chain-of-thought reasoning (GRPO)
soup data generate --prompt "..." --template reasoning --domain math# Auto-validate after generation (remove malformed entries)
soup data generate --prompt "..." --validate
# Auto-filter by quality (coherence scoring)
soup data generate --prompt "..." --filter
# Auto-dedup (MinHash, requires: pip install "soup-cli[data]")
soup data generate --prompt "..." --dedup
# Full quality pipeline: validate + filter + dedup
soup data generate --prompt "..." --quality-pipelineAugment an existing dataset using an LLM — rephrase for diversity, translate for multilingual coverage, or apply a style transform.
# Rephrase each example N times for more diversity
soup data augment ./data/train.jsonl --strategy rephrase --count 3 \
--output ./data/train_augmented.jsonl
# Translate into multiple languages
soup data augment ./data/train.jsonl --strategy translate --lang es,fr,de \
--output ./data/train_multilingual.jsonl
# Style transfer (formal / casual / technical / etc.)
soup data augment ./data/train.jsonl --strategy style --styles formal,casual \
--output ./data/train_styled.jsonl
# Local provider (Ollama / vLLM) — loopback-only, pick the model + base URL
soup data augment ./data/train.jsonl --strategy rephrase --count 2 \
--provider ollama --model qwen2.5:0.5b --output ./data/train_local.jsonlWorks with any provider supported by soup data generate (OpenAI, Ollama, vLLM, local server). --model and --base-url select a specific local model/endpoint; the Ollama/vLLM paths are loopback-only (SSRF-hardened). --count is capped at 10; --lang and --styles each capped at 10 entries × 32 chars.
Harvest DPO / KTO-ready preference pairs from your production inference logs — no manual labeling.
# LangChain logs + thumbs-up signal
soup data from-traces --logs ./logs/langchain.jsonl \
--format langchain --signal thumbs_up --output prefs.jsonl
# OpenAI API logs + regeneration signal (second response wins)
soup data from-traces --logs ./logs/openai.jsonl \
--format openai --signal regeneration --output prefs.jsonl
# Soup-serve logs + user-edit signal (edited response wins over original)
soup data from-traces --logs ./logs/soup-serve.jsonl \
--format soup_serve --signal user_edit --output prefs.jsonl
# Preview generated pairs before training
soup data review prefs.jsonl --sample 10Supported log formats: langchain, openai, soup_serve
Supported signals: thumbs_up (rating-based), regeneration (latest wins), user_edit (edited wins)
Trace files are capped at 100,000 lines to prevent OOM on production logs. A PII warning panel appears on every run — redact sensitive fields before harvesting.
Switch from other tools with one command:
# Import from LLaMA-Factory
soup migrate --from llamafactory llama3_lora_sft.yaml
# Import from Axolotl
soup migrate --from axolotl axolotl_config.yml
# Import from Unsloth notebook
soup migrate --from unsloth finetune.ipynb
# Preview without writing
soup migrate --from llamafactory config.yaml --dry-runAutomatically maps model, LoRA, training params, quantization, and task type. Warns about unsupported features.
Soup supports these formats (auto-detected). Files can be JSONL, JSON, CSV, Parquet, or TXT.
Alpaca:
{"instruction": "Explain gravity", "input": "", "output": "Gravity is..."}ShareGPT:
{"conversations": [{"from": "human", "value": "Hi"}, {"from": "gpt", "value": "Hello!"}]}ChatML:
{"messages": [{"role": "user", "content": "Hi"}, {"role": "assistant", "content": "Hello!"}]}DPO / ORPO / SimPO / IPO (preference pairs):
{"prompt": "Explain gravity", "chosen": "Gravity is a force...", "rejected": "I don't know"}KTO (unpaired preferences):
{"prompt": "Explain gravity", "completion": "Gravity is a force...", "label": true}LLaVA (vision):
{"image": "photo.jpg", "conversations": [{"from": "human", "value": "<image>\nDescribe this."}, {"from": "gpt", "value": "A cat."}]}ShareGPT4V (vision):
{"image": "chart.png", "conversations": [{"from": "human", "value": "<image>\nExplain this chart."}, {"from": "gpt", "value": "Revenue growth."}]}Plaintext (pre-training):
{"text": "Raw text document for continued pre-training..."}Or use .txt files directly (one document per line).
Embedding (sentence embedding pairs/triplets):
{"anchor": "What is Python?", "positive": "Python is a programming language."}
{"anchor": "What is Python?", "positive": "A programming language.", "negative": "A type of snake."}Audio (speech + conversation):
{"audio": "recording.wav", "messages": [{"role": "user", "content": "Transcribe."}, {"role": "assistant", "content": "Hello world."}]}ASR (Whisper transcription — data.format: asr, v0.71.32):
{"audio": "clip.wav", "text": "hello world"}Audio paths resolve under data.audio_dir (containment-checked). Used by
task: asr and soup infer --task asr. See Training → ASR.
PRM (process reward, stepwise-supervised):
{"prompt": "Solve 2+2", "completions": ["First, add", "Result is 4"], "labels": [true, true]}Pre-tokenized (skip tokenize stage):
{"input_ids": [1, 2, 3, ...], "labels": [-100, 2, 3, ...], "attention_mask": [1, 1, 1, ...]}Use with data.format: pre_tokenized and data.tokenized_path: ./.soup-tokenized/<key> after running soup data preprocess.
Input/Output (template-free, segment-level loss control):
{"segments": [{"text": "Q: hi", "label": false}, {"text": "A: hello", "label": true}]}Video:
{"video": "clip.mp4", "messages": [{"role": "user", "content": "Describe this clip."}]}Multimodal (typed content parts — text / image / audio / video in one message):
{"messages": [{"role": "user", "content": [{"type": "text", "text": "What's in this?"}, {"type": "image", "url": "x.png"}]}]}Soup speaks the same dataset surface as Axolotl + LlamaFactory + Unsloth — remote URIs, streaming, sharding, multi-dataset interleaving, vocab expansion, and document ingestion all live in one schema.
Remote datasets (schema gate live; fsspec backend wiring lands in v0.42.1):
data:
train: s3://my-bucket/datasets/train.jsonl # also gs:// gcs:// az:// abfs:// abfss:// oci://
streaming: true
buffer_size: 8192
shards: 4HuggingFace Hub names (a single data.train like org/dataset): data.streaming: true
is forwarded to datasets.load_dataset(..., streaming=True) and buffer_size shuffles
that stream, then Soup materialises up to 1M rows — the same shape as remote (#689).
An all-hub list with streaming: true is still refused (#459). buffer_size shuffles the train split only; a capped validation split takes the first N rows unshuffled.
Multi-dataset interleave (v0.42.0 schema, wired into training-time loading in #443; extended to streaming and HF-hub dataset names in #459):
data:
train:
- dolma.jsonl
- wikipedia.jsonl
interleave: { strategy: probs, probs: [0.7, 0.3] } # also: concat / under / over
eval_on_each_dataset: truedata.train as a list requires data.interleave (and vice versa). training.packing /
training.multipack must be off. With the probs strategy, len(data.train) must equal
len(probs).
Every list entry is classified once — local file path, remote URI, or HF-hub
dataset name (no recognised file suffix — .jsonl / .json / .csv / .parquet /
.txt only count as local; a hub name with a version number like teknium/OpenHermes-2.5
still classifies as hub) — and the classes may not mix within one list. Any entry
containing "://" classifies as remote regardless of scheme — a scheme outside the
allowlist (e.g. https://, http://, ftp://) is refused by name at load time rather
than silently falling through to local/hub classification:
- All entries local files and/or remote URIs: with
data.streaming: false(default), entries must be local files only (the original #443 path, eager-loaded and combined in-process). Withdata.streaming: true, entries may also be remote URIs, and combining delegates to HFdatasets.interleave_datasets/concatenate_datasetsinstead — a remote URI entry always requiresdata.streaming: true(there is no non-streaming multi-remote-file loader). Each remote entry is canonicalised through the same SSRF-hardenedvalidate_remote_uriallowlist used everywhere else in Soup (bucket regex, no userinfo / query / fragment) before it reaches the streaming loader. The streaming path supports the same file types as local loading —.jsonl/.json/.csv/.parquet/.txt— chosen per entry by suffix, so flipping onlydata.streaming: truekeeps reading the same file format instead of misparsing it as JSON; an unrecognised suffix refuses by name rather than reaching the HF loader. - All entries HF-hub dataset names (e.g.
teknium/OpenHermes-2.5): each name's owntrainsplit is loaded and combined the same way as the local-file path. Always eager —data.streaming: trueis not yet supported for an all-hub-name list (streaming several differently-shaped hub datasets through their own split negotiation is unimplemented and refuses at parse time, by name). A hub entry's ownvalidationsplit is used for the combined val set only when every entry provides one (combined the same way); if only some entries provide one it is ignored (warned) anddata.val_splitis derived instead: per source beforeover/probspad it, same as the local-file path below, or from the combined train rows forconcat/under. A partial hub split is not a decided mixture. - A mix of hub names with local/remote entries in the same list always refuses — there is no decided answer for how a hub split and a local file's row count should reconcile.
The strategy names mean the same thing on the streaming path as on the local path, though not byte-identically (a streaming source's size generally can't be known ahead of time):
| strategy | local (eager) | streaming (delegated) |
|---|---|---|
concat |
every source's rows, in order | concatenate_datasets(streams) |
under |
truncate every source to the smallest source's size | interleave_datasets(streams, stopping_strategy="first_exhausted") |
over |
upsample every source to the largest source's size (cycled) | interleave_datasets(streams, stopping_strategy="all_exhausted") |
probs |
exact apportionment to the requested ratio | interleave_datasets(streams, probabilities=probs, stopping_strategy="first_exhausted") — converges to the same ratio, sampled rather than exact |
On the local (eager) and all-hub-name paths, data.val_split is applied per source before
over/probs pad it with copies of its own rows, so a padded row can never land on both
sides of the split; concat/under never duplicate rows and still split the combined
result as before. This does not reach the streaming path below, which still splits after
combining and can still duplicate a row across train and val under over/probs.
Splitting before padding also means the requested val_split fraction is no longer exact
under over/probs: it is taken from each source's own (smaller, unpadded) row count, so
the held-out share of the final, padded total comes out lower than requested. For example,
two sources of 1000 and 100 rows with over and val_split: 0.1 yield 110 val rows out of
1910 total (5.8%), not the 200/2000 (10%) a single-source split would give. concat/under
are unaffected (they never pad). This is the trade-off for closing the duplicate-row leak,
not a separate bug: holding out an exact 10% of the padded total would mean some val rows
are copies of val rows already counted, or of train rows.
Train and val also end up with different source mixtures once over/probs pads: val is
carved from each source's original, unpadded rows, while train sees the padded, rebalanced
mix. Anyone who oversampled specifically to correct a source imbalance gets a validation set
that still reflects the original, un-rebalanced skew, not the mixture train now trains on.
Vocab expansion + advanced masking:
data:
add_new_tokens: ["<reasoning>", "</reasoning>"]
new_special_tokens: ["<|tool_call|>"]
resize_vocab: true
mask_history: true
split_thinking: true # Qwen3-style <think> reasoning-block masking
image_min_pixels: 256
image_max_pixels: 4096
image_resize_algorithm: bicubic
video_fps: 24
video_maxlen: 32
video_dir: ./videosAOT preprocessing:
# Tokenize once, reuse the cache across runs.
soup data preprocess soup.yaml --output ./.soup-tokenized
# Then in soup.yaml:
# data:
# format: pre_tokenized
# tokenized_path: ./.soup-tokenized/<16-char-cache-key>Document ingestion (PDF / DOCX / MD / TXT → JSONL):
soup data ingest report.pdf --output report.jsonl
soup data ingest README.md
soup data ingest notes.docxCustom prompt strategies (schema only — runtime invocation in v0.42.1):
data:
prompt_strategy: my_pkg.transforms:rephrase# Inspect a dataset
soup data inspect ./data/train.jsonl
# Validate format (auto-detects if --format not specified)
soup data validate ./data/train.jsonl
soup data validate ./data/train.jsonl --format alpaca
# Convert between formats
soup data convert ./data/train.jsonl --to sharegpt --output converted.jsonl
# Merge multiple datasets
soup data merge data1.jsonl data2.jsonl --output merged.jsonl --shuffle
# Remove near-duplicates (requires: pip install "soup-cli[data]")
soup data dedup ./data/train.jsonl --threshold 0.8
# Extended statistics (length distribution, token counts, languages)
soup data stats ./data/train.jsonl
# Filter by quality (perplexity + coherence scoring)
soup data filter ./data/train.jsonl --coherence 0.3
soup data filter ./data/train.jsonl --perplexity 500 --coherence 0.3
soup data filter ./data/train.jsonl --score-only # add scores without filteringTiny JSONL fixtures bundled with Soup so you can warm up soup train without
hunting for data:
# List available bundles
soup data demo
# Copy one into the current directory
soup data demo alpaca_demo --output ./alpaca.jsonlBundles: alpaca_demo, sharegpt_demo, dpo_demo, grpo_demo. Output path
must stay under cwd; existing files are not overwritten.
soup data from-traces --judge filters harvested preference pairs through an LLM judge:
soup data from-traces \
--logs ./prod-traces.jsonl --format langchain --signal thumbs_up \
--output ./prefs.jsonl \
--judge --judge-provider ollama --judge-model llama3 \
--min-confidence 0.7The judge scores chosen and rejected independently against its rubric (default helpfulness/accuracy/safety on a 1-5 scale). Pairs whose normalised (chosen - rejected) confidence falls below --min-confidence are dropped. Per-pair backend exceptions are counted (not crashed) and reported. Provider allowlist {openai, server, ollama} validated at the CLI boundary; SSRF protection on --judge-api-base carries over from soup eval judge.
Multi-stage synthetic data pipeline with full provenance — every synthetic row links back to the source document, the judge call, and the filter score:
# Pipeline: chunk docs → judge → active-prune → JSONL + provenance manifest
soup data forge \
--docs ./my_docs/ \
--task sft \
--target-rows 1000 \
--uncertainty-threshold 0.4 \
--output forge_dataset.jsonl \
--provenance forge_provenance.jsonThree tasks supported: sft (Q&A pairs), preference (chosen/rejected), tool (tool-call hypotheses). Active learning prunes rows whose judge reply is too close to the source chunk (low Jaccard distance), keeping only uncertain / informative samples. The provenance manifest is a separate JSON file mapping every row id to {source_doc, judge_id, chunk_id, filter_score} so you have a complete audit trail for compliance.
Document discovery is one level deep over .txt / .md / .json / .jsonl; dotfiles + symlinked directories are skipped. All paths are cwd-contained, all writes are atomic via staged-tempfile + os.replace, and write targets are rejected if they're symlinks. Judge providers are live: --judge-provider ollama (localhost-only), --judge-provider anthropic (env-only API key), --judge-provider vllm (scheme-validated). Per-call judge exceptions logged at DEBUG.
Alternative teacher hubs (v0.71.5). --hub modelscope|modelers pre-fetches the --teacher from that hub when the teacher is a routable repo id (owner/name); --hub hf (default) is a no-op and leaves the teacher as a provenance label. If --hub is non-HF but --teacher is not a repo id (e.g. the default local-judge), Soup prints a loud yellow warning rather than silently dropping the flag.
Composite, lightweight data-quality triage — no GPU, no 200 MB Presidio model:
# Single-shot composite scorecard
soup data score --input training.jsonl
# Standalone subcommands — JSONL-in, enriched JSONL-out
soup data pii --input training.jsonl --output pii_flagged.jsonl
soup data toxicity --input training.jsonl --output tox_flagged.jsonl --threshold 0.1
soup data langdetect --input training.jsonl --output tagged.jsonl
soup data educational --input training.jsonl --output scored.jsonl
soup data decontaminate --input training.jsonl --benchmarks mmlu,gsm8k,humaneval --output clean.jsonlThe scorecard reports PII flagged, toxic flagged, language distribution, mean educational value, and decontamination removed. PII detection uses a narrow ReDoS-hardened regex set (email / phone / SSN / credit-card) with a 50 KB pre-cap on every input. Language detection is a stopword heuristic across six languages. Toxicity is a keyword baseline; the Llama-Guard-3-1B variant + FineWeb-Edu classifier ship behind [data-pro] extras. Decontamination uses n-gram containment against benchmark corpora: use --benchmarks mmlu,gsm8k for built-in allowlist, or --benchmark-file custom_benchmark.jsonl for your own corpus.
Point data.train at any object in the v0.42.0 fsspec allowlist and soup train will stream it through fsspec.open after running the URI through the same SSRF-hardened validator used everywhere else in Soup (bucket regex, no userinfo / query / fragment):
data:
train: s3://my-bucket/datasets/train.jsonl
format: alpaca
streaming: true # opt-in HF datasets streaming with shuffle
buffer_size: 10000 # shuffle buffer (requires streaming=true)Recognised schemes: s3://, gs://, gcs://, az://, abfs://, abfss://, oci://. The matching backend SDK (s3fs / gcsfs / adlfs / ocifs) is lazy-imported — install only what you need or grab the convenience extra:
pip install soup-cli[remote] # fsspec + s3fs + gcsfs + adlfsMaterialised rows are capped at 1M to defend against pathological remote objects; use a local split for larger jobs.
soup data recipe my_recipe.yamlnodes:
- name: seed1
kind: seed
config: {path: prompts.jsonl}
- name: llm1
kind: llm_text
- name: judge1
kind: judge
- name: samp1
kind: sampler
edges:
- [seed1, llm1]
- [llm1, judge1]
- [judge1, samp1]Closed node-kind allowlist (seed / llm_text / code / judge / validator / sampler); Kahn's topological sort via collections.deque (deterministic, O(N+E)); cycle / self-loop / duplicate-edge / dangling-edge / unknown-kind rejection. _MAX_NODES=256, _MAX_EDGES=1024, _MAX_FILE_BYTES=1MiB. The recipe file must stay under cwd and must not be a symlink (os.lstat + S_ISLNK TOCTOU defence). Live offline runner against a local model lands in v0.45.1.
Search for the dataset mixture weights that minimise eval loss on a short proxy run.
soup data mix --optimize --budget 1h \
--datasets dolma.jsonl,wikipedia.jsonl,arxiv.jsonl \
--num-probes 8 --output mix_recipe.yamlWrites a YAML recipe you can splice into your soup.yaml: data.train renders as the full ranked dataset list (index-aligned with data.interleave.probs), and data.interleave carries the searched mixture weights — as of #443, data.interleave is fully wired into training-time dataset loading, so soup train consumes the real N-dataset mixture this search found rather than collapsing to one path. --budget accepts 60s / 5m / 1h / 24h. Per-candidate proxy failures are isolated (DEBUG-logged, sentinel high loss recorded) so a single OOM combo does not abort the whole search; partial=True is surfaced in the report when the budget cap trips mid-loop.
Re-apply a previously written recipe:
soup data mix --apply mix_recipe.yamlLive wiring of the proxy training loop into a short soup train run is the v0.48.1 deliverable; v0.48.0 ships a synthetic offline proxy (quadratic penalty around the uniform simplex) so the budget tracker, optimiser surface, and recipe writer can be exercised without GPUs. scikit-optimize is opt-in via OptimizerProtocol; the default fallback is a deterministic Dirichlet sampler.
Pre-tokenize your dataset once and cache Arrow shards keyed by
(dataset, tokenizer, max_length, format):
soup data preprocess soup.yaml --output ./tokenized_cacheSFT and Pretrain trainers short-circuit at schema validation when
format: pre_tokenized + tokenized_path: ./tokenized_cache is set, eliminating
the per-epoch tokenization tax. Cache keys ensure resume safety; partial runs pick
up from the last completed shard.
Execute a Data Recipe DAG end-to-end:
soup data recipe path/to/recipe.yaml --execute --output ./outSix node kinds now run live: seed (JSONL load), llm_text (LLM generation via any provider), code (execution via RLVR sandbox), judge (binary scoring), validator (regex or JSON schema), sampler (deterministic selection). Checkpoint written per node; resume rehydrates from per-node sidecars. Failed rows logged with redacted reasons (paths stripped, capped at 256 chars).
Chat-template compatibility report — catches the top silent fine-tuning failures before a single training step:
soup data doctor ./data/train.jsonl --model meta-llama/Llama-3.1-8B-Instruct
# Render N sample rows with per-token trained/masked colouring, through the REAL
# collator path (answer-only / per-message-train-field / RAFT span-mask)
soup data doctor ./data/train.jsonl --model meta-llama/Llama-3.1-8B-Instruct --show-mask 5Eight checks, same OK/MINOR/MAJOR taxonomy as soup diagnose (exit 0 on OK/MINOR,
exit 2 on MAJOR): chat_template (tokenizer has one), template_render (renders
cleanly on a sample), generation_markers ({% generation %} support),
eos_in_labels — the #1 "model never stops generating" bug: every trained
assistant turn must actually contain an EOS/EOT token, checked across the whole
trained span, not just the last turn — bos_duplication (template + tokenizer both
prepending BOS), system_role (Mistral-style templates that reject a leading system
turn), unknown_roles, and truncation_risk (p95 rendered length vs
data.max_length). --train-on-responses-only / --train-on-messages-with-train-field
select the same masking strategy soup train would use, so the report and
--show-mask preview can never disagree about what's actually trained.
Catches the top silent degradations in DPO/ORPO/SimPO/IPO/BCO/KTO preference data:
soup data lint ./data/prefs.jsonl
soup data lint ./data/prefs.jsonl --model meta-llama/Llama-3.1-8B-Instruct # exact token-length bias, not word countFive checks: length_bias — the #1 silent DPO degradation: chosen
systematically longer than rejected, reported as a Cohen's d effect size —
label_imbalance (KTO desirable:undesirable ratio), near_duplicates
(MinHash/LSH, reuses the soup data dedup kernel; requires
pip install "soup-cli[data]", degrades to an advisory skip otherwise),
identical_pairs (chosen == rejected — zero preference signal), and
prompt_leak (the prompt echoed verbatim inside the completion, a common
synthetic-data pipeline bug). Same OK/MINOR/MAJOR taxonomy and exit codes as
soup data doctor.