Pandemic Darlings The pandemic economy, in original documents
Home About Pandemic Darlings

The Archive

About Pandemic Darlings

An archive of pandemic-era relief, built from the original documents: what was spent, who touched it, and what the record says happened.

What this is

Pandemic Darlings collects the public record of the pandemic relief programs and publishes it in one place, free to read. Court filings, agency reports, inspector-general audits, congressional material, company filings, contracts, arbitration records, and the official loan- and award-level data files sit at the bottom of the site. Everything else here, the program guides and the derived tables, is built on top of them and links back to them. When a page and a document disagree, the document wins and the page gets fixed.

The site is not affiliated with any government agency. Nothing here is legal advice.

What is in the archive

As of September 25, 2026: 11,966 court-filing pages, 4,312 source documents, 19 statute references, and 23 derived data tables built from 8 federal source releases. The section indexes carry the current counts; pages are still being added, and the numbers on this page are the ones measured when it was built.

Court filings come with the extracted text of each document and, where the record allows it, the original file. The court-filings section also holds the agency and congressional record: SBA program forms, notices, and FAQs; SBA Office of Inspector General reports; Congressional Research Service reports; Select Subcommittee staff reports and letters; Federal Reserve facility disclosures; and bill texts. The source-documents section holds a second body of government and program records, most of them catalogued by issuer, type, and date while their text is still being converted. The data section holds the official loan-level data, 11,468,210 PPP loans in the SBA's file and 3,680,124 COVID EIDL awards from USASpending, plus the RRF and SVOG award files and a corpus of 135,405 Department of Justice press releases, with derived tables you can download as CSV.

Where the documents come from

Courts. Federal dockets and filings come from PACER and from the free RECAP archive at CourtListener; state-court records come from county clerk portals. Acquisition is free-first: a paid PACER purchase is the last resort, made one document at a time when no free copy exists, and every acquisition is logged with its source.

Agencies and Congress. Records are taken from the issuing agency's own site, from Congress.gov, and from FOIA releases.

The contemporaneous record. Company releases, industry letters, and program guidance from 2020 through 2022 are preserved with their dates, so that later claims can be checked against what was said at the time.

How files are processed

  1. Preserve. Every document is archived as filed. The original file is the record; nothing on this site replaces it. Originals are served where the record allows it (the link on a page reads "View original PDF" and states the file's size in bytes); the rest are published as extracted text.
  2. Extract. Text comes from the document's own text layer where one exists. Scanned documents are run through Tesseract OCR. Pages that fail quality checks (broken encodings, garbled glyphs) are re-processed page by page rather than shipped dirty. Where no usable text exists yet, the page says "Full text unavailable".
  3. Publish. Each document becomes a readable page with its extracted text and its record facts: court or issuer, the filing or document date, the source, and the case it belongs to. Titles are normalized from raw filing names by script, with model assistance and a human-maintained override table checked against the document. Duplicate or low-quality conversions are kept but marked not-for-indexing rather than deleted.
  4. Index. The search index is rebuilt over the whole corpus on every publish.

The methodology page covers the rest: the authority tiers, what the labels on a page mean, how to tell a verbatim title from a composed one, and how to cite a page.

Where models are used, and where they are not

This archive is built with substantial AI assistance. Anthropic's Claude models (the Claude 4 generation for most of the build, and later Claude models from mid-2026) assist document titling and cross-referencing and help run the verification passes. Tesseract handles OCR. The documents themselves are never AI-generated; numbers come from the underlying data files. The AI use page lists each step and says where a model is involved.

Corrections

When the record proves a page wrong, the page is fixed. What to send when you find an error is on the corrections page.

Data and reproducibility

Derived datasets ship as CSV downloads with their provenance: the official source files they were built from, the cleaning steps applied, and the known quirks in the underlying data. Every file in the data release is listed with its SHA-256 checksum, byte size, and row count. Tables reconcile to the underlying datasets, not to secondhand summaries.

Reuse and citation

The underlying documents are public records. Tables and text from this site may be quoted with a link back to the page they came from. Each data page carries a suggested citation; the form for citing a document page is on the methodology page.

Phases

Phase 1, the archive, is publishing documents, dockets, and datasets as they are ready. Phase 2, original analysis built on it, follows.

Back to top