From scattered files to a permanent record.
This archive was not downloaded from one place — no such place exists. It was assembled document by document during years of reporting, then verified, extracted, and organized into what you are browsing now. Here is the whole pipeline.
Collect — from everywhere the record hides
Documents come from dozens of scattered sources: agency websites (before reports vanish from them), the Legislature's bill-tracking system, the Federal Audit Clearinghouse, GAO and federal Inspectors General, GovInfo, IRS public filings, federal and territorial court systems, records requests — and the Wayback Machine, for documents the government already deleted. Every document was gathered in the course of VI Update's investigative reporting.
sources: ~12 sweep families · some documents exist nowhere else online anymoreFingerprint — every file gets an identity
Each file is hashed with SHA-256 — a cryptographic fingerprint that changes if even
one byte of the document changes. The fingerprint becomes the document's permanent ID
(the LF-… you see in every address), deduplicates copies collected from
different sources, and lets anyone prove our copy is unaltered.
Review — a person decides what publishes
Nothing reaches this site automatically. Every document was classified — public government works we host, third-party material we only link to, and material we exclude — and every class was reviewed and signed off by a person. Before anything ships, documents are screened for privacy concerns and for anything that could compromise a source. Journalistic work product stays out; the public record goes in.
host / link / exclude classes · privacy & source-protection sweep · human sign-off at every gateExtract — every readable word, indexed
Digital documents give up their text directly; paper scans go through OCR. Today 99.9% of the archive is word-for-word readable — 87,513 of 88,942 documents, 471 million counted words. The rest are scans still in the OCR queue. That is what powers search, the extracted-text reader on every document page, and the copy button. Machine-extracted text can contain recognition errors, which is why every page carries a warning to verify wording against the archival copy before quoting.
digital extraction + OCR · 470,952,455 words counted · the extraction is a tool — the PDF is the recordIdentify — real titles from authoritative registers
Filenames lie; registers do not. Each document is matched against authoritative records: acts against the Legislature's own bill-tracking data (statutory titles, which Legislature, date passed), audits against the audit census (issuer, report number, date), nonprofit filings against IRS records (organization names). Thousands of documents carry their true official titles because of these joins.
acts register · audit census · IRS records · thousands of documents re-titled from source registersOrganize — shelves, tags, and doors
Every document lands on a shelf (a collection plus a sub-shelf recording where it actually came from — audits by issuer, acts by Legislature, filings by organization, rescues by the site they were recovered from), gets tagged with the institutions it concerns, its era, its identifiers (act numbers, report numbers, EINs), and its topics. That is what makes everything on this site clickable — one tap on any value finds everything related. The full shelf list lives in the Card Catalog.
15 collections · 393 sub-shelves · controlled topic vocabulary · identifier cross-referencesPublish — permanently, in waves
Documents go live in reviewed waves, each document at a permanent address that will not change, with its first-page preview, its metadata, its extracted text, a link to the original source where one still exists — and the download of the complete archival copy. When an agency deletes a report, it stays here.
permanent /doc/LF-… addresses · original-source links preserved · integrity line on every pageMeasure — the archive checks itself, and the world
Every night, automatically: local files are re-hashed against their recorded fingerprints; origin links are re-fetched and compared byte-for-byte, so a government document quietly revised at its source is caught rather than missed; registered sources are watched for dying certificates and lapsing domains; and sequential record series are counted against what the issuing institution itself says should exist — so this archive can state what it holds out of, not just what it holds. Documents with no surviving copy in any public web archive are identified and given backup priority. The results live on the stats page, re-measured at every publish.
nightly fixity + origin recheck · frames ("N of M") · web-archive survival measurement · a third offsite copy verified byte-for-byteGrow — including what you send in
Submissions are reviewed by a person before anything happens with them. Nothing about a submission is published — not a name, not the fact that anyone sent anything. The archive also keeps growing from ongoing reporting: new audits, new acts, new filings, captured as they appear (and before they disappear).
human review · anonymity respected · refresh cycles capture new documentsThe engine room
This site is the public face of a much larger apparatus built for VI Update's reporting. The documents here are cross-checked against purpose-built research systems that track the territory's government from every angle:
The USVI Data Warehouse
Millions of rows of territorial financial records — every government checkbook payment, payroll, budgets, vendors, contracts — reconciled against federal award data down to the dollar, and cross-referenced person-by-person and entity-by-entity. When a document here mentions a payment, the warehouse usually knows the check. A public snapshot is queryable as the Shadow Ledger — the "Data" door in the navigation.
The Audit Census
The first complete inventory of every findable audit of the V.I. government ever published — thousands of documents across a dozen federal and territorial issuers, reaching back nearly a century — plus a findings database that traces individual audit findings year over year, including ones repeated for a decade or more. Its core corpus is publicly queryable in the Shadow Ledger (the audit census · the findings, row by row), and audit pages here link straight to their parsed rows — each direction resolves the other by the document's permanent ID.
The Acts Register
A structured register of V.I. legislation — every act since 1998 with its number, session, type and whether it became law, joined to bill numbers where the Legislature's own tracking recorded them, plus parsed sections and full text. No public repository like it existed, so one was built. The reconciliation that traces each act through to the Code section it changed is still being built.
The Hearings Archive
Thousands of hours of public meetings, hearings, and briefings — Legislature sessions, boards, agencies — captured and made full-text searchable, so testimony can be found by the words that were said.
The Court Index
Tens of thousands of USVI federal and territorial dockets and opinions, indexed and cross-referenced against the financial data — parties, vendors, officials.
The Federal Accountability Dashboard
Federal money flowing to the territory, tracked at access.viupdate.com — grants, disaster obligations, and the agencies responsible for them.
Each system feeds this archive and checks it. They are separate by design — this site shares the documents; the systems behind it power the reporting.
Verify a document yourself
Download any document and compute its SHA-256 — on Windows:
certutil -hashfile document.pdf SHA256; on Mac/Linux:
shasum -a 256 document.pdf. The result should match the fingerprint
printed on the document's page here, character for character. If it does, your copy
is bit-for-bit identical to the archival copy. If it ever does not —
tell us immediately.
What we host, link, and leave out
- We host: U.S. government works (public domain), acts and edicts of the V.I. government (uncopyrightable by law), court filings and opinions, government-published financial reports, public IRS filings — and third-party work that carries an open licence or whose rights we have otherwise determined.
- We link: copyrighted third-party material whose rights are not settled — news articles, paywalled research, commercial reports. Cited, never hosted.
- We leave out: private individuals' records where hosting would serve no public purpose, anything that could expose a source, and our own unpublished work product.
Believe something here is in error — or should not be here? Write us. Corrections get made in the open — except where confirming or denying a change could put someone at risk; those may be made quietly, and neither confirmed nor denied.
Using what you find here
Everything published here is yours to take — copy it, republish it, mirror it, and please do. A public record with one copy is one fire from gone. Nothing on this site asks you to seek our permission, and most of it was never ours to grant or refuse in the first place.
Two layers sit in front of you, and only one of them is ours. The documents are what they were when we found them; we own none of them and assert nothing about them. What this project adds is description, tagging and arrangement — and extracted text and machine transcripts, which are mechanical rather than authored and so have nothing in them to own. To the extent any copyright subsists in that layer, we waive it worldwide under CC0 1.0. No attribution is required; it is appreciated, and a link back helps the next person reach the source.
That is CC0 and deliberately not the Public Domain Mark. The Mark asserts that a work is public domain — a claim about the underlying record that is not ours to make. CC0 gives away only what might be ours, and says nothing about the document.
Why the records themselves are free is answered per document, not once. Acts and edicts of government are uncopyrightable whoever wrote them (Georgia v. Public.Resource.Org, 590 U.S. 255 (2020)), and the Virgin Islands Code reserves copyright in exactly one thing — the Madras fabric design (1 V.I.C. § 101a). Federal works are public domain under 17 U.S.C. § 105, which covers the federal government only and not the territories. Records of the Legislature and its committees are public under 3 V.I.C. § 881, which reaches "any department, board, council or committee of any branch of government" and expressly permits the news media to publish them — while the open-meetings chapter does not reach them at all, because 1 V.I.C. § 253(b) writes the Legislature and its committees out of it by name. Boards and commissions, among them GERS and the Public Service Commission, are covered by 1 V.I.C. § 254 and 3 V.I.C. § 884 in addition.
Where a document's own basis has been recorded, it is printed on that document's page.