# CalMatters Digital Democracy: API & Docs (working preview)

The campaign money behind state politics, ready to question and cite. Every figure starts in a state's own disclosure filings and keeps a link back to the report it came from. This is a working preview: good enough to find leads, and still being checked before anything is published from it. Treat a figure as a lead and confirm it against its filing before you publish it.

This page is the Markdown version of https://flows.takuma.ai/access, for agents and scripts.

## Two ways in

- **Explore:** ask a question in plain English at https://flows.takuma.ai/ (the playground), e.g. "who funds Regina Hinojosa?" or "where did Uber's money go?". Nothing to set up.
- **Build:** query the same data in SQL, download a whole state as one SQLite file, or hand it to an agent. Needs a token (below).

## Coverage
| State | Contributions | Dollars | Dates | Published (UTC) |
|---|---|---|---|---|
| Hawaii (`hi`) | 16,831 | $11,162,960 | 2024-11-06 to 2026-08-08 | 2026-09-30 00:27:45 |
| Texas (`tx`) | 3,586,758 | $793,806,451 | 2024-11-06 to 2026-09-25 | 2026-10-01 19:48:39 |
Every contribution links to the filed report it came from.
### Known limits

- Each state covers the current election season, through the latest data the state has published.
- More states arrive one at a time, each checked before it is published here.
- Links between committees and candidates, employers, and industry codes are still filling in and vary by state. Industry codes are beta.
- Money a conduit only passed along is left out, so nothing counts twice. Small gifts reported only as lump sums appear as flows with no giver (`kind = 'unitemized'`), and as a total on each committee.
- Matching of names to people and organizations is checked by blind review of random samples and is still being refined.

## How the data is made

State filings → (1) pull in fresh data, every row linked to its report → (2) keep what is current: amended reports replace what they amend → (3) resolve who is who: decision models judge which records are the same person or organization, and keep the reason → (4) categorize: language models read what filings leave out, like employers and industries (beta) → the commons. Read it three ways: the playground, the SQL API, or a state download.

## Getting in

1. **Get a token from us.** One per organization, covering the states agreed with you and nothing else. Keep it out of repositories; if it leaks, tell us and we will issue a new one. Some tokens read the three views described below (`flows`, `entities`, `ties`); others also read the tables underneath. Every token can download a state; which file you get depends on which kind yours is. We will tell you which yours is.
2. **Send it on every request** as a header. Base address: `https://flows.takuma.ai/api/v1`

   ```http
   Authorization: Bearer <your token>
   ```
3. **Check what it covers.**

   ```bash
   curl -H "Authorization: Bearer $DD_TOKEN" https://flows.takuma.ai/api/v1/coverage
   ```

## Ask in SQL

Send one `SELECT` (or `WITH … SELECT`) as JSON. The data is read-only. Each answer returns at most 10,000 rows and must finish within 30 seconds. Add `"format": "csv"` for a CSV file instead of JSON. Ask in the three views: `flows` (one row per gift), `entities` (everyone who gave or received) and `ties` (who belongs to whom). They are described under "The views and tables".

```bash
curl -H "Authorization: Bearer $DD_TOKEN" -H "Content-Type: application/json" \
  -d '{"sql": "select giver, sum(amount) as total from flows where giver_id is not null group by giver_id, giver order by total desc limit 20"}' \
  https://flows.takuma.ai/api/v1/sql
```

```python
import os, requests, pandas as pd
API, H = "https://flows.takuma.ai/api/v1", {"Authorization": f"Bearer {os.environ['DD_TOKEN']}"}
def q(sql):
    r = requests.post(f"{API}/sql", json={"sql": sql}, headers=H).json()
    if "error" in r: raise RuntimeError(r["error"])
    return pd.DataFrame(r["rows"], columns=r["columns"])
q("select receiver, sum(amount) as total from flows where state = 'tx' group by receiver_id, receiver order by total desc limit 20")
```

### Which view answers which question

| You want | Use |
|---|---|
| Find a person, company or committee by name | `/search`, or `entities` by `name` |
| Who funded a candidate | `flows` for the committees that `ties` links to the candidate |
| Where an entity's money went | `flows`, filtered by `giver_id` |
| Top givers or receivers, in a state | `flows`, grouped by giver or receiver; to rank givers, add `WHERE kind <> 'unitemized'` |
| Giving rolled up by employer or industry | `flows`, grouped by `employer` or `industry` |
| Who someone works for, or which candidate a committee belongs to | `ties` |
| Small gifts reported only as lump sums | `flows` where `kind = 'unitemized'`, one row per filed lump with its filing; `entities.unitemized` for a committee's total |
| The gifts themselves, to cite | `flows`, with `filing` and `report_id` |

The database is SQLite, so dates are text (`'2025-06-30'`). `flows` already applies the counting rule, so a plain `sum(amount)` matches what committees report, small gifts included. To rank or count givers, add `WHERE kind <> 'unitemized'`. Filter by id rather than `lower(name) like '%term%'` where you can; `/search` gives you the ids.

### Good to know

- State codes are lowercase in every column (`'tx'`), and dates are `'YYYY-MM-DD'`. Write `'TX'` or `'01/31/2026'` and they are read as `'tx'` and `'2026-01-31'`; the answer's `adjusted` list says so. An empty answer carries a `hint` when the query shows why (an id in the wrong column, a name matched exactly).
- Every answer carries `as_of`: for each state, the latest gift in the data, when it was published, and its `data_version` and `pipeline_version`. It also carries `version`, the version of this API. States amend their filings all the time, so the same query can give a different answer later: compare the versions to see whether the data, the pipeline or both changed. A list that ends in `order by ... limit n` and stops inside a tie says so in `tie`.
- An answer holds up to 10,000 rows. When there are more it says `"truncated": true` and gives `next_offset`: send the same query with that as `"offset"` for the next rows. CSV answers carry the same in the `X-Truncated` and `X-Next-Offset` headers.
- Each `giver_id` is meant to be one person or organization, and every gift is counted once. The matching is not perfect yet: the same giver can still show up as two entities, under two spellings or ids (a union and its PAC filed under slightly different names, say). We are refining it continuously. Until then, ask `/search/smart` for the giver: it finds every entity that is really them and returns their ids together, so filtering `flows` by all of them gives their whole giving.
- Employers in Texas are resolved for the employer names behind $500 or more of person gifts, about 98% of employer-reported dollars; smaller employer names each stand as their own record for now.
- Covers state-level campaign finance disclosures; federal races are not included.
- Not in this data: spending, votes, bills, and money moved between committees.

## Find an entity

Two ways to find an entity's id, with one answer shape. `/search` lists every entity whose name matches your words, most gifts first: instant, and yours to choose from. `/search/smart` takes a question in plain English and keeps only the entities it means, using common sense (both spellings of one committee, not its employees, not a namesake in another state), with a reason for each. A model answers it, so it takes a few seconds, and it only ever chooses among what `/search` would list. Both look only inside the states your token covers.

```bash
curl -H "Authorization: Bearer $DD_TOKEN" "https://flows.takuma.ai/api/v1/search?q=realtors"

curl -H "Authorization: Bearer $DD_TOKEN" -H "Content-Type: application/json" \
  -d '{"question": "where did the realtors give"}' https://flows.takuma.ai/api/v1/search/smart
```

```json
{
  "status": "found",
  "results": [
    {"id": "jexd-xbcg:5d4418bf53a84a27", "as": "giver", "name": "Hawaii Realtors Political Action Committee",
     "type": "organization", "state": "hi", "city": "Honolulu", "gifts": 82,
     "side": "giver", "why": "Hawaii Realtors PAC, listed under its full name"},
    {"id": "jexd-xbcg:dcd5c799cdba9ba2", "as": "giver", "name": "Hawaii Realtors PAC",
     "type": "organization", "state": "hi", "city": "Honolulu", "gifts": 82,
     "side": "giver", "why": "Hawaii Realtors PAC, listed under its short name"}
  ],
  "left_out": [
    {"id": "jexd-xbcg:580bc658ad6f4d8b:emp", "as": "employer", "name": "Hawaii Assocation of Realtors",
     "type": "employer", "state": "hi", "city": "Honolulu", "gifts": 3,
     "why": "people who work at the realtors association, not the group's own gifts"}
  ],
  "searches": [{"q": "realtors", "side": "giver", "state": "hi", "found": 3}],
  "note": "..."
}
```

- `/search` takes `q` (the words), and optionally `side` (`giver` or `receiver`), `state` and `limit` (up to 100). It returns `results` only.
- Every result has `id`, `as` (the `flows` column its id goes in, below), `name`, `type`, `state`, `city` and `gifts`.
- `/search/smart` adds `side` and `why` to each result. `side` is `giver`, `receiver`, `other_giver` or `other_receiver`; the second ones appear when a question compares two. `left_out` lists near matches it set aside, each with a reason. `status` is `found`, `names_no_one` (the question asks about everyone, like top donors in a state: query `flows` without an id), `not_found` or `not_applicable` (the question is not about campaign money).

Put each id in the column that matches its `as`:

| `as` | In a query |
|---|---|
| `giver` | `WHERE giver_id IN ('…')` |
| `employer` | `WHERE employer_id IN ('…')` |
| `receiver` | `WHERE receiver_id IN ('…')` |
| `candidate` | `WHERE candidate_id IN ('…')` |
| `industry` | `WHERE industry IN ('…')` |

```sql
-- where did the realtors give: the ids /search/smart kept
select receiver, count(*) as gifts, sum(amount) as total
from flows
where giver_id in ('jexd-xbcg:5d4418bf53a84a27', 'jexd-xbcg:dcd5c799cdba9ba2')
group by receiver_id
order by total desc limit 25;
```

## Download a state

The whole state as one SQLite file. It opens in any SQLite tool, DuckDB, pandas or Datasette, and is rebuilt each time the state is published. A token that reads everything gets the raw tables described below, and the `published` table inside says when. A token that reads only the views gets a file with three tables, `entities`, `flows` and `ties`, with the same columns as the views.

```bash
curl -H "Authorization: Bearer $DD_TOKEN" -o tx.sqlite https://flows.takuma.ai/api/v1/download/tx
sqlite3 tx.sqlite "select count(*), sum(amount) from contributions"   # a token that reads everything
sqlite3 tx.sqlite "select count(*), sum(amount) from flows"            # a token that reads only the views
```

## The views and tables

Three views are the recommended way in, and they are defined on every SQL connection. Entities give money to entities (`flows`), and entities belong to one another (`ties`). Every flow links to the filing it came from.

### `flows`
One row per counted gift, or per unitemized line.
- `date`, `amount` (dollars), `state` (two-letter, lowercase), `kind` (`gift`, `in kind` or `unitemized`)
- `giver_id`, `giver`: who gave, resolved, so every spelling of one person or organization shares one id; null on an unitemized row
- `giver_type` (`person`, `org` or `committee`, or empty where the filing doesn't say; null on an unitemized row), `giver_city`, `giver_state`: where the giver lives, as filed (the state code lowercase)
- `employer_id`, `employer`: the resolved employer a person gave under, when the filing names a real one
- `industry`, `sector`: industry codes, beta; null when not yet classified
- `receiver_id`, `receiver` (as filed), `receiver_type` (`candidate`, `pac`, `party` or `other`, where known)
- `candidate_id`, `candidate`: the candidate who controls that committee; null when there is none or it is not linked yet
- `supported_candidate_id`, `supported_candidate` (where published, Texas today): the candidate a committee backs without belonging to them, like Greg Abbott's supporting committee. Money "to a candidate" can mean both.
- `why`: one plain sentence saying how the giver was identified: matched to earlier filings, a new person, or a model's judgment. On an unitemized row it says the committee filed the gifts as a lump with no giver named
- `filing` (a link to the filed report), `report_id`, `source_note` (why a gift has no link)

`flows` already applies the counting rule: direct and in-kind gifts are in, and refunds and transfers are left out. Each lump of small gifts a committee filed is one row with `kind = 'unitemized'`, dated and linked to its report, with no giver: `giver_id`, `giver` and `giver_type` are null. A plain `sum(amount)` includes them and matches what committees report. To rank or count givers, add `WHERE kind <> 'unitemized'`. Each committee's total of these is also in `entities.unitemized`.

### `entities`
Everyone who gave or received.
- `id`, `name`
- `type`: `person`, `organization`, `committee`, `employer` or `candidate` (a candidate's id is the one `flows.candidate_id` and `ties` use)
- `state`, `city`, `industry` (beta)
- `given`, `received`: everything it gave and received in itemized gifts, in dollars
- `unitemized`: a committee's total of small gifts reported only as a lump sum, with no donor names. The lumps themselves are rows of `flows` with `kind = 'unitemized'`

### `ties`
Who belongs to whom, where it is published. One link per row, from `entity_id` to `other_id`.
- `tie = 'candidate'`: the committee's candidate is `other_id`
- `tie = 'employer'`: the person's employer is `other_id`
- `tie = 'supports'`: the committee backs the candidate `other_id`

### The tables underneath

Tokens that read everything can also query the tables the views are built from. They are faster for rankings over a big state, and they keep the raw column names. A token that reads only the views gets an error naming the views if it asks for one of these. The `entities` view keeps the table's own column names too, so older queries still run.

Names (`entities`, `entities_search`) give ids; ids pick rows in `contributions`; `contributions` is rolled up into the totals tables, which are rebuilt at every publish so they always agree.

## Brief for an agent

Paste this into a conversation first, then ask in plain English.

````markdown
# Digital Democracy campaign contributions

Query the data with SQL (SQLite dialect). Covers state-level campaign finance disclosures; federal races are not included.

- **Endpoint:** `POST https://flows.takuma.ai/api/v1/sql`
- **Header:** `Authorization: Bearer $DD_TOKEN`
- **Body:** `{"sql": "..."}`, one SELECT per request
- **Limits:** at most 10,000 rows, 30 seconds. Answers come back as `{columns, rows}`.
- **Download:** `GET https://flows.takuma.ai/api/v1/download/:state` returns the state as one SQLite file. A token that reads everything gets the tables below; a views-only token gets `entities`, `flows` and `ties`, with the views' columns.
- **Find ids first:** `GET https://flows.takuma.ai/api/v1/search?q=words` lists every entity whose name matches, most gifts first. `POST https://flows.takuma.ai/api/v1/search/smart` with `{"question": "..."}` keeps only the ones a plain-English question means, each with a reason; it takes a few seconds.

## Views (ask in these)

Three views are defined on every connection. Some tokens read only these; others also read the tables underneath.

### `flows`
One row per counted gift, giver to receiver: direct and in-kind gifts, plus one row for each lump of small gifts a committee filed (`kind = 'unitemized'`, dated and linked to its report, with no giver). Refunds and transfers are left out. Plain `sum(amount)` includes the lumps and matches what committees report; to rank or count givers, add `WHERE kind <> 'unitemized'`.
- `date` ('YYYY-MM-DD'), `amount` (dollars), `state` (two-letter, lowercase), `kind` ('gift' | 'in kind' | 'unitemized')
- `giver_id`, `giver`, `giver_type` ('person' | 'org' | 'committee', or empty where the filing doesn't say), `giver_city`, `giver_state`; `giver_id`, `giver` and `giver_type` are null on an unitemized row
- `employer_id`, `employer` (Texas: resolved for employer names behind $500 or more of person gifts, about 98% of employer-reported dollars; smaller names stand as their own record for now)
- `industry`, `sector` (beta codes; null when not classified)
- `receiver_id`, `receiver` (as filed), `receiver_type` ('candidate' | 'pac' | 'party' | 'other', where known)
- `candidate_id`, `candidate`: the candidate who controls the receiving committee; null when none or not linked yet
- `supported_candidate_id`, `supported_candidate` (where published, Texas today): the candidate a committee backs without belonging to them, like Greg Abbott's supporting committee. Money "to a candidate" can mean both.
- `why`: a plain sentence saying how the giver was identified (matched to earlier filings, a new person, or a model's judgment)
- `filing` (link to the filed report), `report_id`, `source_note` (why a row has no link)

### `entities`
Everyone who gave or received.
- `id`, `name`, `type` ('person' | 'organization' | 'committee' | 'employer' | 'candidate'), `state`, `city`, `industry`
- `given`, `received` (itemized dollars), `unitemized` (a committee's total of small gifts reported only as a lump sum; the lumps themselves are rows of `flows` with `kind = 'unitemized'`)

### `ties`
Who belongs to whom, where published.
- `entity_id`, `tie`, `other_id`
- `tie = 'candidate'`: this committee's candidate is `other_id`. `'employer'`: this person's employer is `other_id`. `'supports'`: this committee backs candidate `other_id`.

## From a question to ids

Both return `results`: each has `id`, `as`, `name`, `type`, `state`, `city`, `gifts`. `/search/smart` adds `side` (giver, receiver, other_giver, other_receiver) and `why` to each, a `left_out` list of near matches with reasons, and `status` (`found`, `names_no_one`, `not_found`, `not_applicable`). Put each id in the column that matches `as`:
- `as = 'giver'`: `WHERE giver_id IN ('...')`
- `as = 'employer'`: `WHERE employer_id IN ('...')`
- `as = 'receiver'`: `WHERE receiver_id IN ('...')`
- `as = 'candidate'`: `WHERE candidate_id IN ('...')`
- `as = 'industry'`: `WHERE industry IN ('...')`

## The tables underneath

Only for tokens that read everything. The `entities` view keeps this table's column names too (`kind`, `jurisdiction`, `search_name`), so queries written against the table still run.

### `contributions`
One row per counted contribution in the current season (conduit pass-throughs left out), plus one per unitemized lump, which has no donor. The table under `flows`.
- `flow_id`, `jurisdiction` (two-letter state, lowercase), `date` ('YYYY-MM-DD'), `amount` (dollars)
- `donor_id`, `donor_name`, `donor_kind` ('person' | 'org'), `donor_city`, `donor_state`
- `donor_industry`, `donor_sector` (beta codes; null when not classified)
- `donor_employer_id`, `donor_employer`, `donor_employer_industry`
- `recipient_id`, `recipient_name` (as filed)
- `candidate_id`, `candidate_name`: the candidate who controls the committee; null when none or not linked yet
- `recipient_tie_type`: 'candidate' | 'pac' | 'party' | 'other', where known
- `why`, `employer_why`: a plain sentence each saying how the donor, and the employer, were identified
- `report_id`, `source_url` (the filed report), `source_note` (why a row has no link)

### `entities` (the table's own columns)
Every donor, employer and recipient.
- `id`, `name`, `search_name` (lowercase), `kind` ('person' | 'org' | 'employer' | 'committee'), `jurisdiction`, `city`, `industry`

### `entities_search`
Full-text index over entity names. Query with MATCH, not LIKE.
- `id`, `name`, `search_name`, `kind`, `jurisdiction`

### `donor_totals` · `recipient_totals`
One row per (jurisdiction, donor) or (jurisdiction, recipient), pre-summed.
- `jurisdiction`, `donor_id` / `recipient_id`, `donor_name` / `recipient_name`, `total`, `gift_count`, `first_date`, `last_date`

### `donor_recipient_totals`
One row per (jurisdiction, donor, recipient): "who gave to X", "everything donor Y gave".
- `jurisdiction`, `donor_id`, `donor_name`, `recipient_id`, `recipient_name`, `total`, `gift_count`, `first_date`, `last_date`

### `employer_totals`
One row per (jurisdiction, employer), pre-summed.
- `jurisdiction`, `employer_id`, `employer_name`, `employer_industry`, `total`, `donor_count`, `gift_count`, `first_date`, `last_date`

### `candidate_totals`
One row per (jurisdiction, candidate), every committee they control added up. Committees with no candidate appear under their own name; `kind` tells which.
- `jurisdiction`, `key`, `kind`, `candidate_id`, `recipient_id`, `name`, `total`, `gift_count`, `donor_count`, `first_date`, `last_date`

### `unitemized_totals`
One row per (jurisdiction, committee): small gifts it reported only as a lump sum.
- `jurisdiction`, `recipient_id`, `recipient_name`, `total`, `report_count`

### `published`
One row per state: `jurisdiction`, `published_at`, `contributions`, `entities`, `data_version`, `pipeline_version`

## Rules for good answers

- Ask in `flows`, `entities` and `ties`. Get ids from `/search` or `/search/smart`, then filter `flows` by id.
- Money to a candidate: use `candidate_id`, or the committees `ties` links to the candidate.
- Plain `sum(amount)` over `flows` includes small lump-sum gifts (`kind = 'unitemized'`) and matches what committees report. To rank or count givers, add `WHERE kind <> 'unitemized'`: those rows have no giver. Each committee's lump total is also in `entities.unitemized`.
- Each `giver_id` is meant to be one person or organization, and every gift is counted once. The matching is imperfect and being refined: one giver can still show up as two entities (two ids). Get a giver's ids from `/search/smart` and filter by all of them, and when an answer ranks or counts givers, say that one giver may appear twice.
- Grouping a big state live can be slow. With a full token, rank from the `*_totals` tables, and search names with `entities_search MATCH 'term'`, never `LIKE '%term%'`.
- Say which state and date range an answer covers, and how many rows are behind a total.
- Cite from `flows`: every row has `filing`, the state's filed report.
- Industry codes are beta. Candidate ties and employers are still filling in; expect nulls, and say how much of a total they cover.
- State codes are lowercase in every column ('tx'); dates are 'YYYY-MM-DD'. Other spellings are read that way and listed in the answer's `adjusted`. An empty answer may carry a `hint`: read it before concluding there is nothing. Say the `as_of` dates and versions with any total, and report a `tie` rather than an arbitrary cut.
- One query per request, at most 10,000 rows per answer. When `truncated` is true, send the same query again with `"offset"` set to the answer's `next_offset`.
- Not in this data: spending, votes, bills, money moved between committees. Say so rather than guess.
````

## Example questions

```sql
-- find a committee by name, then who funds it through its candidate tie
select id, name from entities where type = 'committee' and lower(name) like '%green%';

select giver, giver_type, count(*) as gifts, sum(amount) as total
from flows
where kind <> 'unitemized'  -- leave out lumps of small gifts, which have no giver
  and receiver_id in (select entity_id from ties where tie = 'candidate' and other_id = 'cand-hi-...')  -- the candidate's id, from ties
group by giver_id, giver
order by total desc limit 25;

-- where one giver's money went, with the filings (look the giver up first, then use the id)
select id, name, type from entities where lower(name) like '%carpenters%';
select date, receiver, amount, filing
from flows
where giver_id = '...'  -- the id from the lookup above
order by date;

-- giving rolled up to employers
select employer, count(distinct giver_id) as givers, sum(amount) as total
from flows
where state = 'hi' and employer is not null and kind <> 'unitemized'
group by employer_id, employer
order by total desc limit 50;

-- top givers in a state
select giver, sum(amount) as total, count(*) as gifts
from flows
where state = 'hi' and kind <> 'unitemized'
group by giver_id, giver
order by total desc limit 25;

-- small gifts reported only as a lump sum, by committee (the lumps are in flows where kind = 'unitemized')
select name, unitemized
from entities
where unitemized > 0
order by unitemized desc limit 25;

-- in-state versus out-of-state money to a committee
select case when upper(giver_state) = upper(state) then 'in state' else 'out of state' end as origin,
       sum(amount) as total
from flows
where receiver_id = 'r-hi-...'  -- the committee's id, from /search or entities
  and kind <> 'unitemized'  -- lumps have no giver_state
group by 1;
```

## Tracing a number to its filing

A total → the `flows` rows behind it (by id) → each row's `report_id` and `filing` → the state's filed report (a PDF or report page, as the state publishes it). When a row has no link, `source_note` says why. Cite the rows, not the total.

## Tell us what you find

If a number looks wrong, something is missing, or an answer surprises you, email dd@takuma.ai. A screenshot of what you saw and a line on what you expected is all we need. Two donors that should be one, or one that should be two, are especially worth sending: each becomes a test case the matching is checked against from then on.

## Release notes

### 1.0.0 · 1 October 2026

The first stable release of the pipeline, proven on Texas. It covers every contribution in a state since 6 November 2024, each traced to the filing it came from. Donors are resolved into the people and organizations behind them and tied to their employers, their industries and the candidates they fund.

**In this release**

- **Texas, the full season** since 6 November 2024, updated from the Texas Ethics Commission's exports.
- **One record per donor.** The same person or organization across many filings is a single donor. A company and its own PAC count as one, marked `own_pac` so they can be split (40 of 40 correct in a blind check).
- **Employers and industries.** Donors are linked to their employers, covering about 98% of employer-reported dollars. Through their employers, gifts carrying 60% of individuals' dollars now have an industry.
- **Candidate ties**, including committees that exist to support a single candidate.
- **A reason on every gift**, in a plain sentence: how it was counted and how its employer was matched.
- **Three ways in:** the playground, plain-English and keyword search, and direct queries on three views.
- **Versions on every answer:** the API version, plus each state's data and pipeline versions. States keep amending their filings, so when an answer changes, the versions say why.

**How well it holds up**

- **Organization matching is measured.** A blind check of 100 organizations across the table found none wrongly combined. A few minor misses remain, so trace any figure to its filing and use your judgment before publishing.
- **Industry codes are beta**, about 87% right in a blind check, with most misses in a neighboring category.

**Known limits**

- Employers on gifts under $500 aren't matched yet.
- Some company branch offices still count as separate employers.
- Committees appear separately as giver and as recipient, and transfers from a candidate's committee to the candidate aren't labeled yet.
- Hawaii's data is still a pre-release, with known limits.
- Federal races aren't covered.

**Next**

- **More states** on the same code.
- **Across states:** we've started sketching the conflict resolution needed to join donors across states and up to the federal level.

## Data changelog

Every data version of every state, newest first. Every answer names its versions in `as_of`, so a different answer to the same query can be traced to the change here.

| Version | State | Date | Pipeline | What changed in the data |
|---|---|---|---|---|
| `texas-v1` | Texas | 1 Oct 2026 | v1.0.0+2c61615 | Compared with the 29 September preview: fewer organizations, as duplicates merged and each company combined with its own PAC (marked own_pac). Gifts from people carry employers and industries. Candidates gain money from their supporting committees, so totals for candidates like Dan Patrick rise. |
---
Research and development by Takuma Productions (https://www.takuma.ai) for CalMatters Digital Democracy.
