This is for a state DOT, a national RTAP, a regional program, or a country program that wants this tool for its own agencies, under its own name and domain — not for adding one more agency to the maintainer's public instance at gtfsscorecard.org (for that, see docs/add-your-agency.md).
Because the system is static artifacts plus a stateless pipeline (no
database, no backend server), standing up your own instance is a fork, a
config edit, and a deploy — not a rewrite. Budget half a day for the steps
below; the excellence bar this quickstart is written against (EXP-15,
docs/ideation/03-expansions.md) is a correctly-cited, branded instance in a
day.
You get: your own domain, your own agency registry, your own branding (site name, organization name, contact address) in every page's metadata and feed, and a daily-refreshed static site you host and own.
You don't get (yet): a different scoring rubric. Every fork scores
against the same shared, versioned methodology (docs/rubric.md,
RUBRIC_VERSION in pipeline/src/scorecard_pipeline/__init__.py) that cites
the California Transit Data Guidelines v4.0 and the MobilityData validator's
own rule taxonomy. If your program has different quality guidance you want
scored against, that is real, larger follow-up work (a pluggable region
rubric — see docs/decisions/0030-forkable-instance-config.md and RR:E11 in
docs/RESEARCH-ROADMAP.md), not something you can configure today. Two
consequences worth knowing going in: (1) findings will cite Caltrans/Cal-ITP
guidance by name until that work lands — accurate as the origin of the
threshold, but not your program's own guidance; (2) because the rubric is
shared, your instance's grades stay comparable to every other instance's,
which is deliberate — the alternative is silent methodology drift.
Fork ChelseaKR/gtfs-scorecard on GitHub, then clone your fork.
Replace the contents of registry/intake.yaml with
your own agencies. For a clean fork, also remove the existing country
subdirectories and reduce registry/index.yaml to
list only registry/intake.yaml; the loader deliberately rejects unlisted
shards. As your registry grows, add your own country/subdivision shards back to
the explicit manifest. Each entry is documented in
docs/add-your-agency.md; scorecard sync --country <cc> --state <state> can propose entries for a whole state or country from the
Mobility Database V2 catalog instead of hand-entering each one:
cd pipeline
uv run scorecard sync \
--country US \
--state Washington \
--out /tmp/washington-proposals.yaml \
--source-metadata-out /tmp/washington-source-metadata.jsonThe command omits authentication-gated Schedule feeds and identities already
tracked by catalog id or normalized URL. The metadata sidecar fingerprints the
exact CSV bytes, each normalized external identity's current public registry
assignment, the proposal-tool source and jurisdiction data, the versioned
public receipt schema, and rendered proposal bytes. It also records source
counts, filters, and one mechanical disposition for every recognized Schedule
source row, including excluded, filtered, already tracked, duplicate,
conflicted, and proposed records. The complete sidecar is schema-validated
before either output is written. Raw feed URLs, authentication details, and
contacts are omitted from the ledger. Use --catalog <URL-or-path> when you
need a pinned or local catalog snapshot.
sync only writes proposals. It has no --apply option and does not edit the
registry. Review source status, location, reuse terms, attribution, and feed
identity before copying any proposal into an intake shard. Catalog metadata
supports that review; it does not grant permission to reuse or republish a
feed. When --source all is used, the JSON sidecar covers the Mobility
Database source rows and per-source counts only and explicitly excludes
Transitland Atlas. Its output hash still binds the combined proposal file; pin
Transitland separately if you need to reconstruct that combined run. The
candidate ledger is likewise Mobility Database-only before cross-source
deduplication, and says so explicitly.
This sidecar is catalog-processing provenance, not a remediation closure receipt. It records no participant action and makes no causal claim. Existing 1.1 sidecars remain valid against the frozen unversioned schema. New sidecars name the immutable 1.2 schema URL; regenerate one when you need the stricter assignment and decision guarantees.
cp instance.example.yaml instance.yamlEdit instance.yaml at the repo root: base_url (your domain), site_name,
org_name, contact_email, and tagline. Every field is independently
optional — set only what you want to change and leave the rest to inherit the
maintainer's defaults, though for a real fork you will want at least
base_url, site_name, and org_name set. This one file drives:
- Canonical URLs,
robots.txt,sitemap.xml, and JSON-LD across every rendered page. - The Atom change feed's author, generator tag, and entry-id tag authority.
- The data attribution string embedded in every published JSON artifact.
- The MCP server's default data source (
docs/mcp.md) and the CLI's offline-preview link rewriting.
Commit instance.yaml. It is meant to be checked into your fork (like the
registry/ directory), not kept as a local override.
What this does not yet rebrand. Page <title> tags, on-page prose (the
board one-pager's footer line, the press and procurement pages), and the
hand-authored marketing pages in web/*.html (the home page, try.html,
submit.html, subscribe.html) still say "GTFS Scorecard" literally in
several dozen places — that sweep is real follow-up work, not done in this
pass (docs/decisions/0030-forkable-instance-config.md). Search for "GTFS Scorecard" across pipeline/src/scorecard_pipeline/render_site.py,
pages_tools.py, and web/*.html if you want to rebrand those too; each is
a literal string, safe to edit directly, just not yet threaded through
instance.yaml.
Follow docs/deploy.md. The short version: GitHub Actions plus
GitHub Pages runs the whole pipeline for free, which is how the maintainer's
own instance runs day to day. The optional AWS pieces (artifacts CDN, the
feed-health email digest, self-serve submission) are independent stacks you
apply only if you want them; deploy.md already calls out the Actions
variables a fork sets differently from the upstream default (for example
ARTIFACTS_CDN, so a fork with its own CDN doesn't inherit the maintainer's).
Point your domain's DNS at GitHub Pages (or your CDN, if you applied
infra/artifacts) and you're live.
- Run the pipeline once locally (
uv run scorecard run --allfrompipeline/, see the mainREADME.md) and check a rendered agency page's view-source: the JSON-LDpublisher/creatorblock should show yourorg_name, and the canonical<link>,og:url, and sitemap entries should show yourbase_url— not the maintainer's. - Check
web/changes/feed.xmlafter a render for yoursite_namein the<author>and<generator>tags. - Check
pipeline/tests/test_instance.pypasses against yourinstance.yaml(uv run pytest tests/test_instance.pyfrompipeline/) — it validates the config loader, not your specific values, but is a quick sanity check that the file parses. - Open a scored agency page and confirm the findings correctly disclose which guideline they cite (Caltrans/Cal-ITP today, per the scope note above).
Because the rubric is shared and versioned (RUBRIC_VERSION), pulling
upstream changes into your fork periodically keeps your instance's grades
comparable to the reference instance's and to any other fork's. A rubric
version bump is the signal a trend on your instance reflects a methodology
change, not a feed change — the same signal docs/rubric.md's "Governed
upgrades" section describes for the maintainer's own instance.
This is young: you are likely the first or an early external adopter. Open an
issue on the upstream repo (ChelseaKR/gtfs-scorecard) if something in this
guide is wrong or missing — that feedback is exactly what turns "Later" work
like the pluggable region rubric into something worth prioritizing.