I build data systems that hold up under regulatory scrutiny.
12+ years in healthcare data, most recently Medicaid rate-setting, supplemental payment calculations and administration, and federal compliance demonstrations for state agencies. That depth is in the data, not the policy. Underneath it: 8 years building production pipelines, data lakes, and cloud migrations, fluent in Python, SQL, SAS, and Azure.
what i do
pipelines & lakes
Production data pipelines and lake architecture in Python, SQL, and PySpark, designed for correctness and lineage, not just throughput.
cloud & on-prem
Equally at home on-premises and in the cloud, including the migrations and integrations that connect the two without breaking downstream consumers.
regulated data modeling
Payment, claims, and risk data modeled at production scale, where a schema mistake has compliance consequences, not just a bad report.
automation & tooling
Internal tools and automation that remove manual steps from a pipeline, built the same way whether it's for an enterprise client or my own infrastructure.
selected work
all projects →Annual Inpatient Hospital Rate Setting & Policy Simulation
2025-2026A claim-level rate-setting pipeline and four-scenario policy simulation using two years of inpatient claims to update hospital payment rates, quantify provider-level tradeoffs, and support selection of a model affecting roughly $148 million in modeled payments.
Synthetic Data Generator
2026A four-stage pipeline that generates privacy-safe synthetic datasets, preserving cross-column statistical correlation, for testing in a locked-down environment where production data never leaves the room.
Price Transparency Data Lake
2023–presentA data lake and automated pipelines ingesting machine-readable pricing files from 80+ healthcare payers under the federal Transparency in Coverage Act, externally-published data at volume, with automated quality monitoring replacing manual review.
$ whoami