What the Texas Medicaid open-data release actually contains
In February 2026, HHS released years of provider-level Medicaid claims data as open data. The Texas file alone runs to millions of rows. It is a genuinely useful public dataset, as long as you are clear about what it does and does not contain.
What the release is
The data comes from T-MSIS, the Transformed Medicaid Statistical Information System, which is the national system states report their Medicaid claims into. The open release aggregates it to the provider level and covers 2018 through 2024. We worked with the Texas extract, a single wide CSV of aggregated, provider-level claims.
What one row represents
Each row is a combination of a billing provider, a servicing provider, a procedure code, and a month, along with how many beneficiaries and claims it covered and how much was paid. It is an aggregate, not an individual claim. A handful of the columns:
| Column | What it holds |
|---|---|
| billing_npi | Billing provider National Provider Identifier |
| servicing_npi | Servicing provider NPI |
| proc_code | HCPCS procedure code |
| yrmonth | Year-month of the date of service (e.g. 202301) |
| num_benes | Number of Medicaid beneficiaries |
| num_claims | Number of claims |
| paid_amt | Total paid amount |
| billing_org_name | Billing provider organization name |
| billing_city / billing_zip | Billing provider location |
What it can tell you
- How much a given procedure was billed and paid across the state, by month.
- Provider-level benchmarking: volume and paid amounts for a specific NPI against peers.
- Geographic patterns in spend and utilization, down to the city and ZIP.
- How many beneficiaries a service reached, which separates high-volume codes from rare ones.
What it cannot tell you
This is the part that matters most, and it is easy to get wrong. The dataset records what was paid. It does not record adjudication outcomes: there are no denial flags, no denial reason codes, and no per-claim status. It is aggregated, so there is no patient-level detail and no way to reconstruct an individual claim.
A claim we deliberately do not make
Why we loaded it
Inside the EHR, this data backs provider benchmarking and statewide analytics: what a procedure typically pays, how a provider's volume compares, where utilization concentrates. It is a strong reference layer for context. It is not, and should not be presented as, a source of denial or appeal outcomes.
Yeam deploys AI medical employees into clinics.
Reception, documentation, coding, and billing, handled by agents so staff can focus on patients.