T2D Risk Screener
A calibrated diabetes pre-screening tool on Korean NHIS health-exam data (~1M patients) — flags patients who need a glucose test, with a per-prediction SHAP explanation behind every score.
Early-career data scientist who builds rigorous, deployable ML — from gradient-boosted risk models to transformer NLP to computer vision. BS in Computational & Data Sciences (summa cum laude), based in Seoul.
I build and deploy end-to-end machine-learning systems — gradient-boosted risk models, transformer-based NLP, computer vision, and reusable tooling — with an emphasis on the parts that decide whether a model is actually useful: honest evaluation, calibrated probabilities, explainability, and a real deployment rather than a notebook. I'm a Computational & Data Sciences graduate of George Mason University (summa cum laude), based in Seoul. A native English speaker with intermediate Korean — and a former U.S. military Korean linguist — I've spent several years working professionally in Korea. I'm open to roles across the data and software spectrum: data science and analytics, machine learning and MLOps, data, backend, and full-stack engineering, DevOps, and technical product management — in any industry, with a particular interest in healthcare and pharma analytics.
A calibrated diabetes pre-screening tool on Korean NHIS health-exam data (~1M patients) — flags patients who need a glucose test, with a per-prediction SHAP explanation behind every score.
A rigorous ablation showing biomedical-domain pretraining gives measurable but modest gains on adverse-drug-event screening — with paired-bootstrap CIs that keep training noise and test-set noise apart.
A fine-tuned image classifier that predicts a photo's country from street-level imagery — and beats humans by roughly 6× on the same 1,000-image test set (McNemar χ²=444.7, p<0.001).
A published Python package (PyPI) extending NetworkX for network visualization — interactive Plotly/Sankey graphs and GeoPandas trade-flow maps over FAOSTAT data.