All notable changes to this project are recorded here, in Keep a Changelog form. The project uses SemVer once it starts tagging releases; nothing has been released yet, so everything to date lives under Unreleased. The commit history is the fine-grained record; this file is the human-readable one.
-
deploy/: the prepared, unapplied deployment shape. A SAM template (JSON, so the test suite reads it with the standard library) for one Lambda behind a Function URL with CORS locked to the Pages origin, reserved concurrency of 2, IAM limited to invoking the one configured model, an invocations alarm and a monthly budget; a build script that assembles the package and never talks to AWS; a README that says it is not applied and lists the decisions it does not make.tests/test_deploy.pyholds the template to the code it would run. -
The opt-in question form on institution pages, off by default.
disclosed site --ask-endpoint URLadds, to each institution page, a labelled form and one inline script with nosrcwhose only network call is inside the submit handler; without the flag the build is byte-for-byte what it was and carries no script. Rendering usestextContentonly; a failed or rate-limited request leaves the page unchanged.tests/test_ask_widget.pyproves all of it from the built bytes. -
ADR 0006: runtime AI at the edges. An owner-directed change of direction, recorded before the code: an optional question-answering layer (
disclosed.ask) in which the model structures a question and narrates the project's own classified records, never sees a reported value, cites a record for every claim, passes a verifier before display, refuses performance judgement, and never collapses the five classifications.AGENTS.mdstates the working rules. -
corpus/: the federal definitions the AI layer may quote. The College Scorecard glossary and data dictionary and the IPEDS HD2023/IC2023 dictionaries, kept as fetched (hash and retrieval date inmanifest.json), reduced to 3,545 passages bydisclosed corpus, replayed byte-for-byte in the test suite.disclosed.ask.definitionsmaps every graded field to the passage that defines its exact variable, and separately to related glossary entries with a note when they define a different measure. Quotes verify verbatim or are withheld. -
The evidence store and the question-structuring step of
disclosed.ask.evidence.buildreduces the committed inputs, in about a second and with no key, to every classification the project has made: 7,095 institutions, 153,486 records across two Scorecard snapshots and three IPEDS years, with the applicability condition behind every IPEDSnot_applicable, the 18 field-level drift measurements from the snapshot series, and 15 cross-source sector contradictions computed over the full census rather than the 600. Areportedvalue is never carried; only animplausibleone is.provideris the SDK seam (first-party or Bedrock, credentials from the environment only, scripted fake for tests);structureturns a question into a typed lookup whose field vocabulary is the schema's enum;lookupresolves the institution exactly, gathers the pack per intent, and refuses with fixed text: performance or ranking, outside disclosure, not in the frame, ambiguous, unclassified measure, unclear. -
Grounded narration, the verifier, and the service.
narrateasks the model for claims that each cite a record id and for verbatim quotes;verifywithholds and counts every claim it cannot prove against the pack: uncited or foreign citations, a classification word none of the cited records is in, an absence rendered as a non-state ("has no", "unavailable"), a number the model was never given, a judgement or recommendation; quotes verify verbatim against the corpus or are withheld.serviceruns the path with a per-client hourly limit and a hard daily cap before the first model call, labels every answer AI-generated and unofficial, keeps no request body, and carries the provenance of every quote; a Lambda Function URL handler and a stdlib development server share it.disclosed askanddisclosed serveon the CLI. -
Five-way classification of every published value (
REPORTED,IMPLAUSIBLE,SUPPRESSED,NOT_APPLICABLE,MISSING), with written rationales for every credible range. -
College Scorecard adapter (600-institution committed capture) and IPEDS adapter (full directory plus institutional characteristics), with the sector disagreement between the two federal sources reported and deliberately left unresolved.
-
Disclosure drift measured as a change in rate against the applicable population, not a change in raw counts, with three committed IPEDS collection years as real history.
-
National corpus (
data/national.json) for the fields IPEDS covers, with ascopeblock in every payload;disclosed nationalrefuses to build from a run that did not cover the population. -
Static site generator with per-institution, per-state, methodology, and national pages; Lighthouse accessibility gate at 100 with every non-document resource budgeted at zero.
-
Citable CSV export with a Table Schema generated in the same pass, plus
CITATION.cff. -
Daily scheduled snapshot workflow accruing per-field disclosure counts in git.
-
Provenance for every page the Scorecard adapter fetches.
disclosed fetchwalks the API and writes a capture envelope: the records, plus for each page the request URL with the key redacted, the fetch time, HTTP status, byte count, SHA-256, attempts, and the rate-limit headers the API returned.Retry-Afteris honoured, consecutive fetches are paused, and a--cache-dirlets a rerun touch no network.grade --sourcereplays the envelope and labels it national only when its own counts prove the walk was exhaustive; the daily job now grades from such a capture with no key in the environment, keeps the raw capture as a ninety-day artifact, and commits the provenance summary beside each snapshot. A dispatch-onlycensusworkflow commits a full capture to the branch it was run on, refusing if the replay is not national or the key is in the file. -
A real Scorecard census (#17), beside the 600-institution sample, never in place of it. Every Scorecard figure this project published came from 600 institutions in 13 states, 51% of them Californian, because the API returns institutions grouped by state and nobody had paged it to exhaustion.
censuswas dispatched for real on 2026-08-21 and committeddata/census/scorecard.json: 6,273 institutions, provenance-proven exhaustive, no key in the file.disclosed census-reportreduces it and the committed sample todata/scorecard-census.json-- per-field national coverage plus both frames' composition (institutions by state and by sector) side by side, so "51% Californian" is answered with a table rather than asserted away. The re-derived headline: 4,363 of 6,273, or 69.6%, publish no admission rate at all, five points higher than the sample's 64.5% -- the sample, if anything, understated non-disclosure. The sample also turns out skewed by sector and not just by state: 47.8% public against the census's 32.6%, and private for-profits at a fifth of the sample against over a third of the census. The site gains a/census/page (disclosed.site.scorecard_census_page) with a pointer from the home page; the README's "What is a sample and what is national" table gains a third row and the sector comparison.tests/test_census_replay.pygates the reduction byte-for-byte against the committed capture, the same disciplinetests/test_replay.pyholdsdata/national.jsonto. -
Portfolio standards conformance set: security workflow (gitleaks, semgrep, pip-audit), Dependabot config, pre-commit hooks, committed
uv.lockand.python-version, ADR log,SECURITY.md,CONTRIBUTING.md, roadmap metrics ledger, and responsible-tech audit record.
- A Scorecard walk that cannot confirm exhaustion now fails instead of reporting national figures.
- Drift no longer reports a shrinking directory as a reporting collapse.
- Site claims are computed from the report payload rather than from constants.
- The SAST gate could not fail.
semgrep --severity=ERRORran 141 rules and found nothing; the same scan without the floor runs 321 and finds three, allWARNING, and all of them theurllibcalls the job was added to watch. The floor is gone and the three are waived at their lines with the reason written beside them.semgrep scan ... src testswas also skipping every one of the fifteen test modules, silently, under semgrep's bundled ignore list; a repository.semgrepignoretakes the scan from 15 targets to 30. - The site-origin guard was outside every gate it protects.
.github/scripts/sat outside the ruff targets, outside strict mypy'sfiles, and outside the coverage source, so the one executable deciding whether the published site may name the origin it names was the only Python file here that nothing read. It is now linted, typed, and covered, with tests that break each of its three checks in turn. - The zero-subresource budget was audited over a fixture, not the site. The claim was
"every one of the 616 generated pages in
make verify";make verifybuilt six, from a report with no implausible finding in it, so the markup rendered around a finding was never parsed. The committed report and national artifact are now rendered whole and every page of the result is checked, and the page count in the prose is checked against the build. - An unrecognized classification word counted against institutions. A report written by a
newer version puts a word the reader does not know into
fields; both aggregators counted it as applicable-but-not-reported, so the published national rate and every drift rate for that field fell for a reason that had nothing to do with any publisher. It is now counted exactly where an absent field is counted: nowhere. - The daily snapshot still reached nobody. With the diff fixed (#18), five more runs
(2026-08-16 to 2026-08-20) wrote a snapshot, committed it, and had the push rejected by
master's protections for lacking the checks they require; a sixteenth recorded a
verifystatus on its own commit (ADR 0002) and was refused for the four checks it had not run. The Actions app cannot be a bypass actor on a user-owned repository and a token-opened pull request never acquires a check, so the job now pushes its commit to a staging ref, dispatchesverify.ymlandsecurity.ymlon that ref, waits for both to pass on that exact SHA, fast-forwards master, and dispatches the site rebuild that a token push would otherwise never start. No step records a check. The post-condition checksorigin/master, not the runner's clone. Sixteen graded days remain unrecoverable and are recorded as such (ADR 0003,docs/RESPONSIBLE-TECH-AUDITS.md). - ADR 0003 also reached nobody, on its first real run. Merged as #26 and dispatched for real
(run 32473991532), the push was refused twice, seconds after
gh api .../check-runsshowed all five required checks green on the exact SHA being pushed. A GitHub Actions check run's check suite is scoped to the branch that triggered it; dispatchingverify.ymlandsecurity.ymlonsnapshot/stagingearns checks that satisfy nothing bound tomaster, no matter how identical the SHA. Commit statuses carry no such scoping, which is why ADR 0002's rejected single self-recorded status had satisfied the ruleset's one check when five real, correctly-SHA'd check runs did not. The job keeps ADR 0003's real dispatch-and-watch verification and, only after each dispatched job has been watched to completion, transcribes that job's own already-earned result to a commit status quoting the run it came from (ADR 0004). Eighteen graded days are now unrecoverable, not sixteen.