Yeamyeam
← All posts
Data

What the Texas Medicaid open-data release actually contains

Zarak ShahAug 27, 20265 min read

In February 2026, HHS released years of provider-level Medicaid claims data as open data. The Texas file alone runs to millions of rows. It is a genuinely useful public dataset, as long as you are clear about what it does and does not contain.

What the release is

The data comes from T-MSIS, the Transformed Medicaid Statistical Information System, which is the national system states report their Medicaid claims into. The open release aggregates it to the provider level and covers 2018 through 2024. We worked with the Texas extract, a single wide CSV of aggregated, provider-level claims.

What one row represents

Each row is a combination of a billing provider, a servicing provider, a procedure code, and a month, along with how many beneficiaries and claims it covered and how much was paid. It is an aggregate, not an individual claim. A handful of the columns:

ColumnWhat it holds
billing_npiBilling provider National Provider Identifier
servicing_npiServicing provider NPI
proc_codeHCPCS procedure code
yrmonthYear-month of the date of service (e.g. 202301)
num_benesNumber of Medicaid beneficiaries
num_claimsNumber of claims
paid_amtTotal paid amount
billing_org_nameBilling provider organization name
billing_city / billing_zipBilling provider location
A selection of the 34 columns in the Texas T-MSIS extract.

What it can tell you

  • How much a given procedure was billed and paid across the state, by month.
  • Provider-level benchmarking: volume and paid amounts for a specific NPI against peers.
  • Geographic patterns in spend and utilization, down to the city and ZIP.
  • How many beneficiaries a service reached, which separates high-volume codes from rare ones.

What it cannot tell you

This is the part that matters most, and it is easy to get wrong. The dataset records what was paid. It does not record adjudication outcomes: there are no denial flags, no denial reason codes, and no per-claim status. It is aggregated, so there is no patient-level detail and no way to reconstruct an individual claim.

A claim we deliberately do not make

Because there are no denial outcomes in this data, it cannot be used to measure a denial rate or to claim a reduction in one. Denial and appeal figures belong to different sources (the CMS Transparency in Coverage PUF and KFF's analysis of it), which we cover in a separate post. Mixing the two would be a mistake.

Why we loaded it

Inside the EHR, this data backs provider benchmarking and statewide analytics: what a procedure typically pays, how a provider's volume compares, where utilization concentrates. It is a strong reference layer for context. It is not, and should not be presented as, a source of denial or appeal outcomes.

Yeam deploys AI medical employees into clinics.

Reception, documentation, coding, and billing, handled by agents so staff can focus on patients.