Congressional disclosures · data dictionary

What is in this dataset, what is missing, and how to check us.

The coverage tables below were measured at 2026-09-10 17:33 UTC, and the queries are printed beside them. Identity evidence is dated separately. Current counts are served live at /v1/coverage. The original sources are the House Clerk and Senate eFD systems; every row we serve links to the document it was parsed from, so you can check any single row against the government’s own copy without asking us anything.

These coverage tables measure warehouse stages, not current served disclosure identities. The separately dated identity evidence distinguishes captured API rows from physical warehouse rebuilds.

Identity, versions and duplicate rows

disclosure_id / identity_basis
disclosure_id identifies a filing-derived transaction, not a physical warehouse row or a unique PDF line. Treat the returned identifier as opaque. The basis government_filing_facts uses source system, filer, document, filed instrument, transaction date and type, amount range, and owner. The instrument key uses the ticker when present, otherwise normalized asset text. Amount identity uses both bounds when present, otherwise normalized range text. Parser line, span, version and run identifiers are excluded. Identical business fields alone cannot prove two receipt lines are one transaction.
version_id / amendments
version_id identifies an immutable observed version. Parser-only changes preserve disclosure identity; a restatement can issue a new version under that identity. An amendment changing identity inputs can produce a different disclosure identifier. Use the returned history and change events, not a ticker/date match, to reconcile updates. On an incremental read, a supersede event removes only the matching stored version; a withdrawn identifier is not a current transaction.
content_hash
SHA-256 of the contract's canonical JSON array, basis veridion-disclosure-content-v1. It includes both identifiers, filing facts and receipt coordinates, including source line number. The same version must retain its hash. Because version_id is included, hashes across different versions are not a test of unchanged filing facts. Midpoint, delay, operational timestamps and parser/methodology versions are excluded. A matching hash verifies the calculation, not the truth or completeness of the government filing. The OpenAPI contract lists the exact fields and history/change response shapes.

Do not collapse business-field matches

Within a pinned snapshot, count returned disclosure identifiers. A key made from document, filer, ticker, date and transaction type can discard different amounts, owners or instruments. The response amount fields are amount.range_low, amount.range_high, amount.currency and amount.verbatim_text, not amount.amount_range. Adding those fields still does not establish a replacement for the API identity or receipt review.

Captured at , generation 962: 781 rows, 781 distinct disclosure identifiers. The withdrawn key produced 708 groups by silently ignoring amounts. Its 33 collision groups included 18 differing only in asset labels, 7 only in owner, and 2 in both; those categories did not compare amounts. 26 groups differed in amount bounds.

Using the actual amount bounds and currency still produced 14 collision groups: 10 differed in asset labels and 4 in owner. Neither grouping justifies dropping rows. This capture is not a completeness or duplicate-transaction audit. Capture measurement, keys and artifact hashes.

Duplicate-writer observation

Measured at : the run scheduled for rebuilt 1,080 physical House PDF warehouse rows. Of these, 979 matched active lined rows, 62 reopened closed versions and 39 matched older lineless copies. These are not counts of extra served API transactions. At that measurement, repair verification was pending and API impact was unestablished. Pre-existing business-field matches require separate receipt review. Dated export measurement and limits.

Section 1

Three row counts, and why they differ

The most common way to be wrong about this dataset is to quote one number for all three of these. They are not interchangeable, and the gap between them is the honest part.

StageRowsDocumentsFilers

First-party warehouse rows

Collected by us from the House Clerk and Senate eFD systems. Not superseded.

53,3186,458349

Passes serving admission

Loses 201 rows to the admission rule below, including 42 whose filing date precedes their transaction date. The loss has been exactly 201 on both days we measured it.

53,1176,442348

Passes the qualified projection

Adds a required amount range and owner determination. Owner type is the binding constraint, and it costs 104 filers and 2,499 documents.

29,5653,943244

Serving admission

A row reaches the serving path only with a non-empty member slug, member name and document ID; both dates present; a receipt URL matching the House Clerk or Senate eFD pattern exactly; and a filing date on or after its transaction date. That last rule is why 42 chronologically impossible rows exist in the warehouse and are not served. The predicate below is the rule as written, not a paraphrase of it.

›The admission predicate, verbatim
where source in ('pdf', 'official-senate-efd')
  and superseded_at is null
  and nullif(trim(member_slug), '') is not null
  and nullif(trim(member), '')      is not null
  and nullif(trim(doc_id), '')      is not null
  and transaction_date is not null
  and filing_date      is not null
  and filing_date >= transaction_date
  and pdf_url ~ '^https://(disclosures-clerk[.]house[.]gov/.+[.]pdf([?].*)?|efdsearch[.]senate[.]gov/.+)$'
›Queries for the three counts
-- Stage 1: first-party warehouse
select count(*), count(distinct doc_id), count(distinct member_slug)
from warehouse
where source in ('pdf', 'official-senate-efd') and superseded_at is null;

-- Stage 2: serving admission
select count(*), count(distinct doc_id), count(distinct member_slug)
from warehouse
where source in ('pdf', 'official-senate-efd')
  and superseded_at is null
  and nullif(trim(member_slug), '') is not null
  and nullif(trim(member), '')      is not null
  and nullif(trim(doc_id), '')      is not null
  and transaction_date is not null
  and filing_date      is not null
  and filing_date >= transaction_date
  and pdf_url ~ '^https://(disclosures-clerk[.]house[.]gov/.+[.]pdf([?].*)?|efdsearch[.]senate[.]gov/.+)$';

-- Stage 3: qualified projection
select count(*), count(distinct doc_id), count(distinct member_slug)
from warehouse
where source in ('pdf', 'official-senate-efd')
  and superseded_at is null
  and nullif(trim(member_slug), '') is not null
  and nullif(trim(member), '')      is not null
  and nullif(trim(doc_id), '')      is not null
  and transaction_date is not null
  and filing_date      is not null
  and filing_date >= transaction_date
  and pdf_url ~ '^https://(disclosures-clerk[.]house[.]gov/.+[.]pdf([?].*)?|efdsearch[.]senate[.]gov/.+)$'
  and amount_low is not null
  and owner_type is not null;
›Query for the 42 impossible rows
select count(*)
from warehouse
where source in ('pdf', 'official-senate-efd') and superseded_at is null
  and filing_date < transaction_date;

Queries on this page are printed against the alias warehouse. The physical relation name is never printed on a public page: a build gate refuses bare warehouse-relation references outside reviewed files, and it is the same gate that keeps any row we did not collect ourselves off this page. The predicate is what you need to check the logic; the name is not.

Section 2

Field fill rates, measured

Measured across the 53,117 rows that pass serving admission. A blank field is a fact about the filing or our parser, not a rounding error, so we publish the rate rather than the impression.

FieldPresentNote
chamber100.0%House or Senate.
transaction_type100.0%Purchase, sale, or exchange. Senate rows also carry Sale (Full) or Sale (Partial). House rows do not, and that is our defect, not the form's: the House PTR marks partial sales as "S (partial)" and our parser drops the marker. Measured 2026-09-14: 22,804 House sale rows, none preserving it. Read a House sale as "sale, extent not preserved" until the restatement lands.
amount_low / amount_high100.0%Disclosure brackets, never exact amounts. The filing discloses a range.
asset99.6%The filer's own text. The form forbids ticker-only entries.
party90.8%
ticker75.3%Seven distinct statuses, not one null. Many disclosed assets have no ticker to find.
source_line_number65.4%Present only on rows read by the current parser.
parse_run_id65.4%Same cohort as source_line_number.
owner_type55.7%The largest completeness gap in the dataset. Optional on the form, so some is absent at source rather than unparsed — we have not yet separated the two causes.
›Query for the fill rates
-- One row; each column is the share of admitted rows with the field present.
select
  round(100.0 * count(*) filter (where nullif(trim(coalesce(chamber, '')), '') is not null) / count(*), 1) as chamber,
  round(100.0 * count(*) filter (where amount_low is not null and amount_high is not null) / count(*), 1) as amount,
  round(100.0 * count(*) filter (where nullif(trim(coalesce(asset, '')), '') is not null)   / count(*), 1) as asset,
  round(100.0 * count(*) filter (where nullif(trim(coalesce(party, '')), '') is not null)   / count(*), 1) as party,
  round(100.0 * count(*) filter (where nullif(trim(coalesce(ticker, '')), '') is not null)  / count(*), 1) as ticker,
  round(100.0 * count(*) filter (where source_line_number is not null)                      / count(*), 1) as source_line_number,
  round(100.0 * count(*) filter (where owner_type is not null)                              / count(*), 1) as owner_type
from warehouse
where source in ('pdf', 'official-senate-efd')
  and superseded_at is null
  and nullif(trim(member_slug), '') is not null
  and nullif(trim(member), '')      is not null
  and nullif(trim(doc_id), '')      is not null
  and transaction_date is not null
  and filing_date      is not null
  and filing_date >= transaction_date
  and pdf_url ~ '^https://(disclosures-clerk[.]house[.]gov/.+[.]pdf([?].*)?|efdsearch[.]senate[.]gov/.+)$';

Why ticker is only 75.3%

Seven distinct statuses rather than one null. Municipal bonds, private partnerships and notes have no ticker to find, so source_asset_without_listed_ticker is a description, not a failure. The counts below are on the same predicate as every other table on this page and sum exactly to 53,117.

  • listed_ticker22,978
  • source_asset_without_listed_ticker11,750
  • (null — rows predating the status field)11,338
  • listed_symbol5,655
  • not_disclosed1,357
  • receipt_asset_without_ticker26
  • invalid_ticker_format13
›Query for ticker statuses
select coalesce(ticker_status, '(null)') as status, count(*)
from warehouse
where source in ('pdf', 'official-senate-efd')
  and superseded_at is null
  and nullif(trim(member_slug), '') is not null
  and nullif(trim(member), '')      is not null
  and nullif(trim(doc_id), '')      is not null
  and transaction_date is not null
  and filing_date      is not null
  and filing_date >= transaction_date
  and pdf_url ~ '^https://(disclosures-clerk[.]house[.]gov/.+[.]pdf([?].*)?|efdsearch[.]senate[.]gov/.+)$'
group by 1 order by 2 desc;

Section 3

Machine-readability by year of transaction

The share of admitted rows where each field is present, by the year the transaction occurred. “All four” means ticker, owner, amount range and asset name all present on the same row. Ticker identification rose from 61.9% to 91.0% across 2014–2025; owner type did not follow it, and all four together have never cleared 63% in any year.

A blank is one of two things and this table cannot tell them apart: the field was blank on the filed document, or our parser did not read it. This measures the machine-readability of the record — a joint property of the filings and of any parser reading them. It says nothing about filer conduct.

YearRowsTickerOwnerAll four
2012partial40.0%100.0%0.0%
2013partial16861.9%37.5%22.0%
20142,12561.9%41.8%31.2%
20153,14669.3%59.5%45.2%
20163,37365.2%59.9%42.7%
20173,37664.5%53.9%39.5%
20183,81961.7%58.6%42.1%
20195,10961.5%47.6%30.6%
20205,35567.2%55.9%37.9%
20214,53275.3%58.1%42.8%
20223,53381.7%61.1%50.8%
20234,45690.8%53.8%48.2%
20242,91788.9%68.4%62.1%
20258,01491.0%54.0%48.3%
2026partial3,19083.7%54.1%42.1%

2012 and 2013 are too sparse to read as rates. 2026 is the year in progress. Row counts are what we collected and admitted, not filing volume.

›Query for the by-year table
select
  extract(year from transaction_date)::int as year,
  count(*) as rows,
  round(100.0 * count(*) filter (where nullif(trim(coalesce(ticker, '')), '') is not null) / count(*), 1) as ticker,
  round(100.0 * count(*) filter (where owner_type is not null) / count(*), 1) as owner,
  round(100.0 * count(*) filter (
      where nullif(trim(coalesce(ticker, '')), '') is not null
        and owner_type is not null
        and amount_low is not null
        and nullif(trim(coalesce(asset, '')), '') is not null) / count(*), 1) as all_four
from warehouse
where source in ('pdf', 'official-senate-efd')
  and superseded_at is null
  and nullif(trim(member_slug), '') is not null
  and nullif(trim(member), '')      is not null
  and nullif(trim(doc_id), '')      is not null
  and transaction_date is not null
  and filing_date      is not null
  and filing_date >= transaction_date
  and pdf_url ~ '^https://(disclosures-clerk[.]house[.]gov/.+[.]pdf([?].*)?|efdsearch[.]senate[.]gov/.+)$'
group by 1 order by 1;

The per-year numerators behind this table, the plotting script, and why the document and filer totals do not sum from it, are published at /data-api/dictionary/field-completeness.

Section 4

Three dates matter, and we only have two

We carry the transaction_date and the filing_date, and the interval between them as delay_days, uncorrected.

We do not carry the notification date, and it changes what the interval means.

A periodic transaction report is due by the earlier of 45 days from the transaction or 30 days from the filer being notified of it. The official form carries a “Date Notified of Transaction” column beside the transaction date. We do not parse it.

So an interval at or under 45 days does not establish that a filing was timely — the 30-day clock may have expired first. The limitation is ours, not the public record’s: the data needed to compute real compliance is published, and we have not captured it. Any compliance rate built on this dataset, ours or anyone else’s, is measuring the 45-day clock alone and should say so. The API field within_statutory_window means exactly delay_days <= 45 and nothing more.

Section 5

What is excluded, and what is not

Excluded from publication entirely

Rows we did not collect ourselves from the government system of record. We hold commercially licensed rows; we do not serve, export or license any of them, and none are counted anywhere on this page.

Excluded by serving admission

The 201 rows described in Section 1. Superseded rows carry a reason and a parse-run reference and are omitted from every count here.

Not excluded — and this is the part we would rather you heard from us

There is no transcription-defect exclusion. The longest transaction-to-filing interval in the first-party warehouse is 3,698 days, and a row that old is more likely a single-digit year misread than a decade-late filing. No detector for this is deployed and nothing is filtered on it today. An earlier internal claim that 160 such rows were excluded was withdrawn: it came from a query that was not recorded and could not be reproduced. Treat long intervals as unverified.

Section 6

Known gaps

  1. House partial-sale markers are not preserved. The House PTR types a partial sale as "S (partial)"; our parser normalises it to Sale. Found 2026-09-14 on Clerk PDF 20035408, which carries 25 such rows. Every one of the 22,804 House sale rows served that day reads the same whether or not the filer marked it partial. Senate rows are unaffected. The fix adds a transaction_extent field and restates the House corpus; until then, do not infer full versus partial from a House sale row.
  2. The notification date is not captured. It is a column on the official form; we do not parse it, so the 30-day clock is unevaluated in our data.
  3. Owner type is present on 55.7% of admitted rows, and we have not separated absent-at-source from unparsed.
  4. No transcription-defect detector is deployed. Long intervals are unverified. The longest in the warehouse is 3,698 days.
  5. 42 rows in the warehouse have a filing date before their transaction date. They are refused at the serving boundary but not yet corrected at source.
  6. Line-level provenance covers 65.4% of admitted rows. Older rows are receipted but carry no line number.
  7. All deadline reasoning on this page is from House guidance. The corpus includes 7,025 Senate rows and the Senate's procedural rules are not separately verified here.
  8. Coverage is what we matched to a document, not everything ever filed. We do not know what we are missing, and we do not claim completeness.
  9. Row counts per year reflect what we collected and admitted, not filing volume. A year with fewer rows may be a year we collected less of, and this page cannot tell you which.

What this is not

Not investment advice, not a compliance determination, and not an accusation. A long filing interval has innocent explanations: broker reporting delays, amended filings, newly seated members reporting historical holdings, and spouse or dependent-child transactions the filer learned of late. We publish the interval between two dates printed on a government document, and infer nothing about motive or performance.

Snapshot history

  • 2026-09-10 · 2026-09-10 17:33 UTC · this page. Admitted 53,117. Added the by-year table and printed every query.
  • 2026-09-09 · admitted 52,705. Ticker statuses were measured on a wider predicate and did not sum; corrected above.
Weekly Veridion brief

Rating changes, public disclosure activity, methodology notes, and product updates. One email per week. No advertising list resale.