From scattered files to a permanent record.
This archive was not downloaded from one place — no such place exists. It was assembled document by document during years of reporting, then verified, extracted, and organized into what you are browsing now. Here is the whole pipeline.
Collect — from everywhere the record hides
Documents come from dozens of scattered sources: agency websites (before reports vanish from them), the Legislature's bill-tracking system, the Federal Audit Clearinghouse, GAO and federal Inspectors General, GovInfo, IRS public filings, federal and territorial court systems, records requests — and the Wayback Machine, for documents the government already deleted. Every document was gathered in the course of VI Update's investigative reporting.
sources: ~12 sweep families · some documents exist nowhere else online anymoreFingerprint — every file gets an identity
Each file is hashed with SHA-256 — a cryptographic fingerprint that changes if even
one byte of the document changes. The fingerprint becomes the document's permanent ID
(the LF-… you see in every address), deduplicates copies collected from
different sources, and lets anyone prove our copy is unaltered.
Review — a person decides what publishes
Nothing reaches this site automatically. Every document was classified — public government works we host, third-party material we only link to, and material we exclude — and every class was reviewed and signed off by a person. Before anything ships, documents are screened for privacy concerns and for anything that could compromise a source. Journalistic work product stays out; the public record goes in.
host / link / exclude classes · privacy & source-protection sweep · human sign-off at every gateExtract — every readable word, indexed
Digital documents give up their text directly; paper scans go through OCR. Today the archive is essentially 100% word-for-word readable — tens of millions of words. That is what powers search, the extracted-text reader on every document page, and the copy button. Machine-extracted text can contain recognition errors, which is why every page carries a warning to verify wording against the archival copy before quoting.
digital extraction + OCR · ~42,000,000 words · the extraction is a tool — the PDF is the recordIdentify — real titles from authoritative registers
Filenames lie; registers do not. Each document is matched against authoritative records: acts against the Legislature's own bill-tracking data (statutory titles, which Legislature, date passed), audits against the audit census (issuer, report number, date), nonprofit filings against IRS records (organization names). Thousands of documents carry their true official titles because of these joins.
acts register · audit census · IRS records · 4,000+ documents re-titled from source registersOrganize — shelves, tags, and doors
Every document lands on a shelf (eight collections, each with sub-shelves — audits by issuer, acts by Legislature, filings by organization), gets tagged with the institutions it concerns, its era, its identifiers (act numbers, report numbers, EINs), and its topics. That is what makes everything on this site clickable — one tap on any value finds everything related.
8 collections · 245 sub-shelves · controlled topic vocabulary · identifier cross-referencesPublish — permanently, in waves
Documents go live in reviewed waves, each document at a permanent address that will not change, with its first-page preview, its metadata, its extracted text, a link to the original source where one still exists — and the download of the complete archival copy. When an agency deletes a report, it stays here.
permanent /doc?id=LF-… addresses · original-source links preserved · integrity line on every pageGrow — including what you send in
Submissions are reviewed by a person before anything happens with them. Nothing about a submission is published — not a name, not the fact that anyone sent anything. The archive also keeps growing from ongoing reporting: new audits, new acts, new filings, captured as they appear (and before they disappear).
human review · anonymity respected · refresh cycles capture new documentsThe engine room
This site is the public face of a much larger apparatus built for VI Update's reporting. The documents here are cross-checked against purpose-built research systems that track the territory's government from every angle:
The USVI Data Warehouse
Millions of rows of territorial financial records — every government checkbook payment, payroll, budgets, vendors, contracts — reconciled against federal award data down to the dollar, and cross-referenced person-by-person and entity-by-entity. When a document here mentions a payment, the warehouse usually knows the check. A public snapshot is queryable as the Shadow Ledger — the "Data" door in the navigation.
The Audit Census
The first complete inventory of every findable audit of the V.I. government ever published — thousands of documents across a dozen federal and territorial issuers, reaching back nearly a century — plus a findings database that traces individual audit findings year over year, including ones repeated for a decade or more. Its core corpus is publicly queryable in the Shadow Ledger (the audit census · the findings, row by row), and audit pages here link straight to their parsed rows — each direction resolves the other by the document's permanent ID.
The Acts Register
A structured register of V.I. legislation — thousands of acts with their bills, veto and override history, and parsed sections, full-text searchable. No public repository like it existed, so one was built.
The Hearings Archive
Thousands of hours of public meetings, hearings, and briefings — Legislature sessions, boards, agencies — captured and made full-text searchable, so testimony can be found by the words that were said.
The Court Index
Tens of thousands of USVI federal and territorial dockets and opinions, indexed and cross-referenced against the financial data — parties, vendors, officials.
The Federal Accountability Dashboard
Federal money flowing to the territory, tracked at access.viupdate.com — grants, disaster obligations, and the agencies responsible for them.
Each system feeds this archive and checks it. They are separate by design — this site shares the documents; the systems behind it power the reporting.
Verify a document yourself
Download any document and compute its SHA-256 — on Windows:
certutil -hashfile document.pdf SHA256; on Mac/Linux:
shasum -a 256 document.pdf. The result should match the fingerprint
printed on the document's page here, character for character. If it does, your copy
is bit-for-bit identical to the archival copy. If it ever does not —
tell us immediately.
What we host, link, and leave out
- We host: U.S. government works (public domain), acts and edicts of the V.I. government (uncopyrightable by law), court filings and opinions, government-published financial reports, public IRS filings.
- We link: copyrighted third-party material — news articles, academic papers, commercial reports. Cited, never hosted.
- We leave out: private individuals' records where hosting would serve no public purpose, anything that could expose a source, and our own unpublished work product.
Believe something here is in error — or should not be here? Write us. Corrections get made in the open — except where confirming or denying a change could put someone at risk; those may be made quietly, and neither confirmed nor denied.