Health Services Research with Claims Data
Claims Data Sources
Where Claims Data Originates
Every time you visit a doctor, fill a prescription, or stay in a hospital, a claim is generated. This is a request for payment sent from a healthcare provider to an insurer. While their main purpose is billing, these claims create a massive trail of data. Researchers use this administrative data to understand healthcare trends, costs, and outcomes on a large scale. The data isn't created for research, which comes with limitations, but its breadth makes it incredibly powerful.
Government-Funded Programs
The largest and most comprehensive claims datasets in the United States come from government insurance programs. These datasets cover millions of people and are a cornerstone of health services research.
Medicare is the federal health insurance program primarily for people aged 65 or older, younger people with certain disabilities, and people with End-Stage Renal Disease.
Because it's a single, national program, Medicare data is relatively standardized. The data is managed by the Centers for Medicare & Medicaid Services (CMS) and is available to researchers through a request process. The core of the data is organized into different files based on the type of service.
| File Type | Description |
|---|---|
| Part A | Covers inpatient hospital stays, care in a skilled nursing facility, hospice care, and some home health care. |
| Part B | Covers doctors' services, outpatient care, medical supplies, and preventive services. |
| Part D | Helps cover the cost of prescription drugs. |
| Master Beneficiary Summary File (MBSF) | Contains demographic information about each beneficiary, such as date of birth, sex, and enrollment details. This file acts as an anchor to link the other files together. |
Next is Medicaid, a joint federal and state program that helps with medical costs for millions of Americans with limited income and resources.
Unlike Medicare, Medicaid is administered by individual states according to federal requirements. This means the data isn't uniform across the country.
Historically, researchers used the Medicaid Analytic eXtract (MAX) data. This has been replaced by the Transformed Medicaid Statistical Information System (T-MSIS), which aims to provide more timely and complete data. Medicaid data is crucial for studying low-income populations, children, and pregnant women, groups that are less represented in other datasets. However, researchers must account for state-by-state variations in eligibility, benefits, and data quality. A person's eligibility can also change from month to month, creating gaps in their data.
Private and Pooled Databases
While government programs cover a large portion of the population, most non-elderly Americans are insured through private, commercial health plans, often provided by an employer. This data is also a valuable resource for research.
Commercial claims
noun
Data from private insurance companies, such as UnitedHealthcare, Aetna, or Blue Cross Blue Shield.
This data is proprietary, meaning it's owned by the insurance companies or data aggregators like Optum and Merative. Researchers typically have to purchase access, which can be expensive. Commercial claims databases are essential for studying the health and healthcare utilization of the employed population and their dependents. However, they aren't nationally representative. The covered population tends to be younger and healthier than the Medicare population, and the data is limited to the customers of a specific insurer or a specific network of employers.
To get a more complete picture, some states have created All-Payer Claims Databases (APCDs). These are large-scale databases that systematically collect medical and pharmacy claims from a variety of public and private payers.
The goal of an APCD is to increase transparency by allowing for comparisons of healthcare costs, quality, and utilization across different payers and providers. While incredibly valuable for state-level analyses, not all states have an APCD, and the standards for data collection can differ between those that do.
Finally, another key resource is the Healthcare Cost and Utilization Project (HCUP). Sponsored by the Agency for Healthcare Research and Quality (AHRQ), HCUP is a family of healthcare databases and related software tools.
HCUP is not a single database but rather a collection of datasets. It uses billing data from hospital encounters, regardless of payer, to create a uniform, national-level information source. The data is collected from states and organizations that participate in the HCUP partnership. This makes it particularly useful for research on hospital care. Key databases within HCUP include the National Inpatient Sample (NIS), which is the largest publicly available all-payer inpatient care database in the United States.
What is the primary, original purpose of a healthcare claim?
A researcher wants to study healthcare trends across the entire United States for a population over 65. Which dataset is most suitable due to its national standardization?
Understanding the source of claims data is the first step in using it effectively. Each dataset has unique strengths and limitations based on the population it covers and how the data is collected.
