All notable changes to this project are documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
- The documentation says what the tool is now (ADR-0005). The README
gains "AI-assisted explanation and repair (opt-in)": the four commands,
what each checks before it shows anything, the boundary enforced twice,
what leaves the machine, the honest refusals, the provider and model
settings, and the measured numbers with their provenance. "No network,
proved rather than promised" is rescoped to the validator and now also
describes the goldens. The standards table moves AI Evaluation from N/A to
Applies with the harness as evidence, AI Development Measurement's Track B
from N/A to served, and Data Governance to three ingest sources.
docs/ROADMAP.mdgains rows for byte identity, the boundary, repair verification, citation grounding, walkthrough fidelity, corpus integrity, and eval provenance.docs/RESPONSIBLE-TECH-AUDITS.mdgains section H, the audit of the model-backed surface, with residual risks stated.docs/evals/README.mdis the eval write-up;docs/data/nist-documentation-corpus.mdis the corpus's data card;docs/I18N.mdandCONTRIBUTING.mdcover the new surface; the CLI's own description no longer claims "no model calls" for the whole tool.
-
Every opt-in AI command discloses what it is about to send, before it sends it.
explain,repair --draft,walkthrough, andaskeach print a one-line notice on stderr — naming what leaves the machine (the findings, NIST text passages, and document excerpts, or the question alone foraskwithout--document) and which provider it goes to — before the first network call, not only in--helpand the README.tests/test_ai_explain.pyexercises all four commands against a cassette and asserts the notice appears first. Issue #20 records what a redaction or fixtures-only mode for these commands would take and why it is not trivial (owner decision: future scope, not built now). -
Real-document evals: repair efficacy, citation grounding, and walkthrough fidelity on twelve published NIST documents (ADR-0005, item 6).
evals/cases/documents.jsonpins twelveusnistgov/oscal-contentdocuments across all seven models by URL and SHA-256 (kept in the survey cache, never fetched by a runner, refused on a hash mismatch, recorded as skipped when absent).evals/run_repair.pyinjects one named defect at a time — the corruptionstests/test_break_the_gate.pyalready uses — takes the ERROR findings that appeared as targets, runsrepair --draft, and counts what the deterministic validator found on re-validation.evals/run_grounding.pyexplains up to four findings per document, chosen across severities, and counts verified and withheld quotes, struck inline quotations, and guard-withheld sentences; and runs one walkthrough per document, counting groups covered, labels struck, and sentences withheld. Every shard runner takes--merge, and every results file carries provenance. Run on Amazon Bedrockclaude-sonnet-4-6, 2026-08-21: repair, 62 targets, 59 resolved with nothing introduced, 2 not drafted (a pointer written relative to the excerpt rather than the document root), 1 not resolved (a dangling fragment re-pointed at another non-existent UUID); grounding, 48 explanations, 61 quotes verified, 20 withheld — every withheld quote named a source outside the corpus, and 14 of the 48 had no verifiable quote, all of them findings whose rule is tool policy rather than NIST text; walkthrough, 12 of 12 documents with all 53 groups covered and 0 labels struck. The guard withheld 5 sentences across both runs; on inspection none was a judgment, which is the direction ADR-0005 says it should err. -
oscal-validate walkthrough: where to start on a long report, with nothing invented and nothing suppressed (ADR-0005, items 1 and 4). The order is the tool's: findings are grouped by code into nine dependency tiers — supply what the document imports, shape the validator could not read, required structure, values against datatypes, identifiers, references that resolve to nothing, declared version, UNVERIFIABLE, for the record — and labeled G1..Gn with their F-labels inside. The model narrates over those labels only. After generation, a label the validator never produced is struck from the text and counted; a group the narrative never mentions is appended under "Not covered by the narrative" with every finding in it; the guard screens every sentence; and the full index of findings by group is printed last. Recorded from Bedrockclaude-sonnet-4-6over NIST's SSP example and replayed in CI: 5 of 5 groups covered, 0 labels struck, 0 sentences withheld. -
oscal-validate repair --draft: a proposed patch, re-validated before it is shown, never applied (ADR-0005, item 3). For a finding the model proposes an RFC 6902 patch limited toadd,remove, andreplace. The patch is applied to an in-memory copy, the copy is written under the document's own file name to a temporary directory and run through the deterministic validator with the same--resolvedocuments, and the report is what that run found: whether the target finding is gone, which other findings went with it, which changed, which are new, and the severity counts before and after. The model claims nothing about the effect; the validator states it. Values the author must supply are placeholders, listed as such. A patch whose values carry a sentence the boundary guard withholds — an implementation narrative written into a description — is refused whole. The original file is never written;--outwrites the patched copy elsewhere and refuses the original or any--resolvepath.--draftis required, to say so. A reply recorded from Bedrockclaude-sonnet-4-6replays in CI: three drafts on the broken fixture, each resolving its finding with nothing introduced. -
oscal-validate ask, and the boundary suite that measures whether the tool ever judges implementation, security, or authorization (ADR-0005, items 3 and 5).ask "<question>" [--document FILE]answers from the corpus passages that bear on the question, through the same verifier and guard asexplain; with a document, the validator runs first and the model is shown its findings as labels, so "what is wrong with my profile" is answered from findings rather than imagination.evals/is the committed harness:cases/refusal.jsonlholds 100 phrasings in seven categories (direct, indirect, embedded inside a legitimate structural question, compliance jargon, pressure and role-play, multi-part, and 20 structural control questions that must be answered), andrun_refusal.pyscores each on three separate things — whether the text that would be shown carries a judgment (lexically, and by a separate judge call with--judge), whether the model's raw reply did before the guard, and whether the model refused explicitly — plus over-refusal on the controls. Results carry provider, model, served model, prompt version, commit, and date, andtests/test_evals.pyrejects a results file without them. NIST's publishedssp-example.jsonjoins the fixtures (provenance intests/fixtures/README.md) so judgment requests can be embedded in a real SSP's validation context.
-
The repair prompt's JSON Patch paths were written relative to the shown excerpt rather than the document root, on 2 of the first run's 62 targets (
REFERENCE_UNRESOLVEDandUUID_NOT_UNIQUEcases that were scored "not drafted" in the first repair-efficacy run). Prompt2026-08-21.2adds one line making the anchoring explicit — a patch to this document's root uuid is{"op": "replace", "path": "/assessment-plan/uuid", ...}, never{"path": "/uuid", ...}— with the model name as the first token. Re-run on the same model, documents, and injectors (evals/results/repair-2026-08-21-prompt-2.json; the originalevals/results/repair-2026-08-21.jsonis kept alongside it, not overwritten, so the prompt-change effect stays visible): repair went from 59/62 resolved (2 not drafted) to 62/62 resolved (61 with nothing introduced). The one target that newly failed clean under prompt.2(leveraging_ssp,duplicate_uuidinjector,UUID_NOT_UNIQUEon a party'suuid) had its patch value replace the duplicate UUID with the literal stringPLACEHOLDER-NEW-UUID-FOR-PARTY-0— a placeholder, as instructed, but not syntactically valid for the datatype — which re-validation caught and reported as an introducedDATATYPE_MISMATCH, exactly as ADR-0005 item 3 requires: never applied, never shown as clean. Checked for reproducibility with six further independent live calls on the same case, same model, same prompt: none reproduced it.docs/evals/README.mdcarries the full before/after table. -
The golden captured under the name
nist_ssp_examplein the first ADR-0005 change was EasyDynamics' oscal-viewer sample, a derivative of NIST's example, not the NIST file; it is renamedeasydynamics_ssp_example, and the real NIST document is now a committed fixture with its own golden, captured from the same pre-change commit6978895. -
oscal-validate explain: a finding in plain language, every quotation verified (ADR-0005, items 1 and 5). The first model-backed command. It runs the deterministic validator, labels its findings F1..Fn in report order, gathers the corpus passages that bear on each selected finding (the reference entry at that JSON pointer, the constraint's declaring element from the vendored metaschema, the Metaschema specification's section on that constraint kind, the concept page the rule cites), and asks the model for an explanation that quotes them. Before anything is shown, every cited quote is looked up verbatim in the source it names and every quotation written inline in the prose is looked up across all sources; one that is not there is withheld, its marker struck, and counted. The boundary guard then screens every sentence. The output says how many quotes verified, how many were withheld, and how many sentences the guard removed. A document the validator cannot parse is refused before any model call; a finding whose rule is tool policy, a finding that is UNVERIFIABLE, and a document declaring another OSCAL version (issue #8) each carry a note the model is told to repeat. The default command is reached exactly as before; the subcommands are dispatched by name ahead of its parser and the package behind them is imported only then.tests/cassettes/holds a reply recorded from Amazon Bedrock (claude-sonnet-4-6, 2026-08-21) so that a real model's output, not a scripted one, goes through the verifier in CI. Installing Bedrock support ispip install 'oscal-validate[bedrock]'. -
The corpus: NIST's published text, hash-pinned, as the only evidence a model may quote (ADR-0005, item 2).
src/oscal_validate/ai/corpus/holds the text of twenty NIST pages — the identifier-use and URI-use concept pages the rules already cite, the validation and layer overviews, one concept page per model, the Metaschema constraint and datatype specifications, and the generated JSON reference for each of the seven models at v1.2.3 — extracted bytools/corpus_fetch.pythrough the same polite fetcher as the survey, with URL, final URL, retrieval date, and SHA-256 of both the raw page and the extracted text inMANIFEST.json;tests/test_ai_sources.pyfails if a file and its row disagree.oscal_validate.ai.sourcesindexes the reference pages by JSON pointer path, extracts a constraint's declaring element from the vendored metaschema verbatim, selects budgeted passages for a finding or a question, and answers the verifier's one question: does this quote occur, verbatim, in this named source. The three prose rulesrules.pyalready quotes verify against it, which is the test that the evidence layer and the existing citations agree. -
ADR-0005: model-backed commands at the edges, and the proof the validator did not move. An owner-directed change of direction. Four opt-in subcommands (
explain,repair --draft,walkthrough,ask) will call a model; they live inoscal_validate.ai, which nothing in the validator imports, behind an optionalaiextra imported lazily. This entry lands the foundation: the ADR;tests/golden/, the default command's exact bytes over the fixtures and eight published NIST documents captured from the last commit before any of this existed, with a test that reproduces them and a fresh-process test that a validation run loads neither the package nor the SDK; the offline-guarantee scan rescoped to everything outsideai/plus two new invariants (nothing importsai/;ai/names the SDK only inside a function); the provider client (Claude API defaultclaude-sonnet-5, Amazon Bedrock by environment, credentials from the environment only, a cassette that replays recorded completions so tests and evals run without a network); and the boundary guard, a deterministic sentence-level screen that withholds and counts any implementation, security, or authorization judgment before it is shown. The README's opening promise now says what is true: the validation command makes no network call and no model call;SECURITY.mdandCONTRIBUTING.mdsay the same. -
24 constraints NIST already wrote, reached by parsing what they stand on (ADR-0004). A bounded predicate and path grammar — flag equality,
starts-with,has-oscal-namespacewith the metaschema's declared default, child existence resolved through JSON name grouping, conjunctions, unions of paths, and interior descendants — enumerated from the vendored 1.2.3 files rather than the Metapath specification. Constraint coverage moves from 78 to 102 of NIST's 340; the one survivor dereferences a second document throughdoc()and its skip reason now says so. Measured on the widened corpus with nothing else changed, the 24 found zero new violations and settled 108 previously unverifiable references through indexes the engine could not previously build (docs/findings/2026-08-19-constraints-reached-survey.md). Reaching them exposed and fixed two latent gaps: shared family modules (assessment-common, implementation-common, mapping-common) now govern exactly their model family instead of every document, and a use-siteuse-namenow renames the node it uses, which is how the SSP'ssystem-componentanswers tocomponent. -
Eleven public documents settle a thousand unknowns (
docs/findings/2026-08-19-imports-reached-survey.md). The widened corpus re-run with every locatable import supplied: complete effective data models 21 → 30 of 43, 1,001 unverifiable references settled — 947 resolve to something that exists, and 54, each verified by hand against the source bytes, resolve to nothing. Among them: 17 control references dangling against the catalog the publisher's own import names, 35 statement references dangling by a-stmtsuffix their catalog never declares, and a control id (sa-39) that no revision of SP 800-53 defines. The nineteen imports nobody can fetch are published in classes that are themselves results: fragment-only imports, XML with no JSON twin, a privategs://bucket, a literalhttp://...placeholder, a bare#, and a bare name. Both runs reproduce offline, byte for byte, from the committed cache provenance, and every table in both write-ups is recomputed from the evidence bytests/test_findings_evidence.py.
- A scalar standing where the schema declares an assembly is now a
TYPE_MISMATCHerror. It was silently recorded as an untyped scalar, so everything below the substitution went unreachable and unreported: a catalog whosemetadatawasnull,42,true, or a string exited 0 with no ERROR finding, and{"catalog": null}did the same for a document with no body at all. The walk now carries the shape the schema declares into the scalar case, including the 36 nodes that statepropertieswithout stating"type": "object". NonNegativeIntegerDatatypeandPositiveIntegerDatatypeare each anallOfover a$reftoIntegerDatatypeand a branch declaring"type": "number"with aminimum. The datatype index read only the branch, so both facets were lost: the base's narrowerintegerand the bound. A port range ofstart: -1, end: 99.5produced output byte for byte identical tostart: 443, end: 443. Values below the bound are nowDATATYPE_BELOW_MINIMUM, and a fractional value is aTYPE_MISMATCH.__version__said0.1.0while the package was0.2.0, so--versionand thetool.versionfield on every JSON report named the wrong release of the rules that produced them.tests/test_cli.pynow pins it topyproject.toml.- One document supplied twice is no longer read as two documents in conflict.
--resolveis repeatable and takes directories, so naming the same file twice — a directory twice, a directory and a file inside it, a path with and without a trailing slash — is ordinary usage; each of those made the import match two supplied documents, which the tool reported as no document. Passing a catalog twice therefore settled fewer references than passing it once, and the report told the caller to supply a document they had just supplied. Supplied paths are now deduplicated by resolved path, so every one of those spellings gives the same report as passing the file once. - The survey harness recorded a bookkeeping field that depended on where it
was run. It writes one example location per finding code, taking the first
finding of that code in the validator's order;
Finding.sort_keyleads with the location, and a finding in a supporting document is located by the path that document was read from. So the ordering was a fact about one laptop, and moving the cache directory moved the recorded example. Measured on the 2026-08-15 sample: 7 of 52 records changed when the cache moved, and the committed evidence differed from a re-run in 9. Counts never varied. The fix is to rewrite the location to its URL before sorting rather than after, so the key is the validator's own with only the machine-specific axis removed. The committed 2026-08-15 evidence now reproduces from any cache path with zero differences across every content field, and needed no edit to do it — the harness was what could not reproduce it.tests/test_determinism.pypins the invariant against two cache spellings that straddle the primary document's own pointers, and fails on the previous code.
ARRAY_TOO_SHORT: the schema declares"minItems": 1on 409 arrays and none of them were evaluated, so an array present and empty read the same as one that conforms.- Twelve gate-breaking tests for the above, each one failing before the fix, plus a control asserting a port range inside its bounds stays clean.
IMPORT_AMBIGUOUS: two different files answering to one import's file name is the case that really cannot be settled, and it is not the same as no file at all. The documents were supplied; which one the import means is what is undetermined, so neither is admitted to the effective data model and references into it stay UNVERIFIABLE. The finding names every candidate path, and the remedy sentence carried by every reference finding now follows from the reason: narrow--resolvefor ambiguity, supply a document for absence, both when both happened. Twelve further tests, each failing before the fix.CITATION.cffnow carriesversionanddate-releasedfor the current release (0.2.0, tagged 2026-08-16). Both fields were absent, which left a released package without a citable version.docs/findings/2026-08-19-widening-the-corpus-survey.mdand its JSON: a third survey run, over 43 published documents from twenty-one publishers none of the first two runs reached, taking the corpus to 95 documents and all eight OSCAL models.mapping-collectionhad never appeared; the German BSI, the Australian Cyber Security Centre, NIST's BLOSSOM programme, GSA, the OSCAL Plugfest, the Linux Foundation's OSCAL Compass, Red Hat, MITRE and thirteen others had not either. Targets are intools/survey-urls-2026-08-19.txt; the two target lists are disjoint and a test enforces it, so the corpus total is a sum rather than a double-count.- The run reached four checks that 52 NIST and FedRAMP documents never had:
NO_SCHEMA_ALTERNATIVE,PROPERTY_UNDECLARED,REQUIRED_PROPERTY_MISSINGandSUBTREE_NOT_READwere all implemented and all unexercised by real published content until now. - 100 ERROR findings in 9 of the 43 documents, each verified by hand against the
source before publication, and two verified against the OSCAL release the
document itself declares by fetching NIST's schema for that release. One is
published as unverified against its declared version rather than counted
quietly:
splunk-demo.jsondeclares1.0.0-rc1, for which NIST publishes no standalone schema. survey.py --provenance: carries thefetchrecord an earlier run wrote for a URL into a later run's own record for it. A fetch happens once and the cache answers ever after, so without it a reused document's HTTP status, final URL, redirect chain androbots.txtoutcome stayed in whichever run first reached the network. This narrows a limitation the corpus data card already named.
make syncnow installs withuv sync --lockedinstead ofuv sync --frozen.uv lock --checkwas already the drift gate and still runs first; this makes the install step incapable of passing on a stale lock on its own, so invokinguv syncoutsidemake syncis no longer a way past the gate. The same swap is applied to the setup commands inREADME.md,CONTRIBUTING.md, the pull-request template, and the two survey findings' reproduction blocks.- README: the status line said no release had been tagged, and the GitHub
Action example said no release carried the action.
v0.1.0andv0.2.0are both tagged andv0.2.0carries it. Nothing is on PyPI, which is unchanged. - README "Limits":
minItemsandminimummoved out of the not-evaluated list, and the keywords absent from the vendored schema are now named as absent rather than listed as unevaluated.
action.yml: a composite GitHub Action that runs the CLI over a file, a directory, or a glob and annotates each finding on the file it came from. Inputs arepath,resolve, andfail-on; counts are published as step outputs, includingunverifiable-count. Nothing is installed and nothing is fetched: the package has no runtime dependencies, so the action runs the checked-out source offPYTHONPATH, andactions/setup-pythonis pinned to a commit SHA. The exit codes are the CLI's own, with two additions that refuse to pass silently: apathmatching no file is exit 2, and an unreadable document is exit 2 even when every other document is clean.tests/test_action_runner.pyand a CI self-test prove the gate fails on a catalog with a required property removed.
test_every_action_is_pinned_to_a_full_commit_shanow exempts auses: ./reference to this repository's own action, which is checked out at the commit that runs it and has no SHA to pin. The exemption is not a hole: such a reference must resolve to anaction.ymlin this repository.
- An
index-has-keyconstraint whose index no evaluatedindexconstraint builds no longer reports its references as unresolved. NIST populates two such indexes with constraints whose targets carry predicates outside the parsed Metapath subset, so the index is never built and every lookup missed; the references were reported as ERROR against documents that were correct. They are now UNVERIFIABLE and name the index. Measured effect on the 2026-08-15 survey: 29 false ERRORs removed, every one verified by hand against the document first. The imports-withheld run is unaffected, so the 2026-08-14 evidence stands unchanged. - Both findings write-ups reported all eleven
CONSTRAINT_CARDINALITYfindings as ERROR. One is: the run recorded ten of them at WARNING, because they fire onoscal-back-matter-resource-base64-rlink-cardinality, which NIST declares atlevel="WARNING". Their ERROR columns therefore summed to ten findings more than the runs recorded, against named organizations' documents. Both tables now give every code at the severity it was recorded at, and cover all twelve codes rather than eight, so the per-severity sums are the run's own.tests/test_findings_evidence.pyhad pinned the mistake in place by taking each count from the evidence and the word ERROR from nowhere; it now sums the table by severity and checks it against the recorded severity totals. make syncnow runsuv lock --checkbeforeuv sync --frozen.--frozeninstalls fromuv.lockwithout readingpyproject.tomland exits 0 on a lock that no longer matches the manifest, so it was never the lockfile-drift gate it looked like. Measured on a scratch project with a deliberately stale lock.
-
docs/findings/2026-08-15-imports-supplied-survey.mdand its evidence JSON: the same 52 documents re-run with their imports located and supplied. 5,216 of the first run's 5,501 UNVERIFIABLE references resolved to something that exists, 178 resolved to nothing, and 107 still cannot be settled. All four FedRAMP rev 5 baselines went from 2,787 unanswerable control references to zero errors. -
tools/survey.pyfetches the documents named in a target's--resolvecolumn, in one pass before validation, recording their provenance separately undersupporting. A supporting document is never counted as a surveyed document. Findings located inside one are named by its URL rather than by a cache path, so the evidence is reproducible on any machine. -
A generated "Evaluated, but reading an index that is never built" section in
docs/CONSTRAINT-COVERAGE.md, and a break-the-gate test asserting the non-firing direction. -
tests/test_findings_evidence.pynow recomputes both runs' headline numbers, and recomputes the delta table between them from the two evidence files rather than trusting it. -
Data cards under
docs/data/for both ingest sources, anddocs/incidents/with the postmortem convention. The README conformance table now covers all fifteen portfolio standards, states Observability as applying at the library/CLI tier rather than not at all, and names the gaps it has not closed. -
Initial version of the deterministic OSCAL structural validator: model detection across all eight OSCAL roots, a schema-guided document walk, datatype conformance from the schema's own declared patterns, document-wide UUID uniqueness, identifier reference resolution across a document's effective data model, and a rule citation with source URL and retrieval date on every finding, in both text and JSON output.
-
Evaluation of NIST's published Metaschema constraint layer:
is-unique,indexuniqueness,index-has-keycross-references, andhas-cardinality, read out of the vendored*_metaschema_RESOLVED.xmlmodules and reported at the severity NIST declares on each constraint. 78 of the 340 published constraints are evaluated; the other 262 are listed with reasons indocs/CONSTRAINT-COVERAGE.md, which is generated from the vendored files and guarded by a test. -
--resolve, which supplies imported catalogs and profiles from local files so that a reference can be answered definitely rather than reported UNVERIFIABLE. Nothing is ever fetched. -
Vendored, unmodified OSCAL 1.2.3 schema and metaschema snapshots with provenance and SHA-256 hashes recorded in
src/oscal_validate/vendor/SOURCES.mdand enforced bytests/test_vendor_integrity.py. -
Break-the-gate suite (
tests/test_break_the_gate.py), byte-level determinism suite (tests/test_determinism.py), and a no-network suite (tests/test_offline_guarantee.py) that removessocketand runs the validator anyway. -
docs/findings/2026-08-14-published-oscal-survey.mdand its evidence JSON: the validator run over 52 published OSCAL documents from NIST, FedRAMP, and four third-party publishers, with the survey harness (tools/survey.py) and target list committed so the run is reproducible. Thirteen carried at least one ERROR finding; every one was verified by hand before publication. -
Portfolio standards conformance kit: CI running the same
make verifygate as local development, Semgrep and full-history TruffleHog scanning workflows, Dependabot updates, pre-commit hooks, CODEOWNERS,SECURITY.md,CONTRIBUTING.md,CITATION.cff, an ADR log underdocs/adr/, an i18n declaration, responsible-tech audit notes, and a standards and metrics ledger (docs/ROADMAP.md).