Work

Work that connects healthcare questions, data, and decisions.

These projects span real-world evidence, pharmacovigilance, healthcare operations, biomedical visualization, and public-health analytics. Each one reflects a different part of how I approach complex healthcare questions.

Real-World Evidence · PharmacovigilanceCompleted

GLP-1 Pharmacovigilance

Sex-stratified pharmacovigilance signal detection for GLP-1 receptor agonists using 18.5M+ harmonized FDA Adverse Event Reporting System records. Applies disproportionality analysis methods to identify hypothesis-generating signals across demographic and temporal dimensions. Results are exploratory and do not support causal inference.

18.5M+FAERS records harmonized

Methods

Disproportionality analysisReporting Odds RatioProportional Reporting RatioSex-stratified analysis

Technologies

PythonPostgreSQLRFDA FAERS
Analytical Product · Real-World EvidenceCompleted

Real-World Evidence Studio

An analytical product concept that guides researchers from a structured clinical question through cohort definition, SQL generation, diagnostics, and interpretable evidence outputs. Designed to make real-world evidence generation more transparent and reproducible. Currently in active development.

Methods

Cohort designPropensity score methodsCausal inferenceOMOP Common Data Model

Technologies

PythonSQLROMOP CDM
Consulting · Life Sciences StrategyCompleted

Healthcare AI Strategy and Business Case

Business case development and strategic analysis for a conversational AI platform in early pregnancy care, conducted through the Penn Graduate Consulting Club. Work included economic modeling across care delivery pathways, payer and customer segment analysis, and scenario and sensitivity frameworks. Engagement details are confidential.

Methods

Economic modelingCare pathway analysisPayer and customer segment analysisSensitivity analysis

Technologies

ExcelFinancial modeling
Biomedical Data Science · Pipeline EngineeringCompleted

Genetic Risk Map

A reproducible Python pipeline for computing gene-level association scores and relative score tiers from MAGMA or GWAS gene-level summary statistics. Built as a portfolio rebuild of BMIN 5100 coursework, with modular source modules, an input validation layer, configurable thresholds, and a 31-test pytest suite. Research and educational use only. Outputs are relative rankings within the submitted gene set, not validated clinical risk estimates.

50Example genes in sample dataset
31Unit tests

Methods

Gene-level association scoringZ-score normalizationPercentile-based stratificationInput validation

Technologies

PythonpandasNumPySciPy
Public Health · Machine LearningCompleted

HPV Vaccination Analytics

This BMIN 5030: Data Science for Biomedical Informatics course project examined whether selected access and socioeconomic variables were associated with reported HPV vaccine receipt in NHANES 2021-2023 data. Implements logistic regression and XGBoost with SMOTE class balancing in R. Includes a detailed methodological reflection on the original workflow's limitations and what a more rigorous rebuild would require.

2Outcome classes
3Predictors used
2Models evaluated

Methods

Logistic regressionXGBoost classificationSMOTEBinary outcome modeling

Technologies

RtidyversenhanesAcaret

In progress

Selected case studies are being expanded with detailed methods, findings, and interactive elements. If you would like to discuss a specific project before then, reach out directly.

Let's connect.

Whether you have a collaboration in mind, a healthcare analytics question, or are curious about the work, I would be glad to hear from you.