Director of Data Science · Causal Inference · Applied AI
Evidenceoverassumption.
I build systems that answer what actually works — causal inference for education, and the data and AI infrastructure that makes it hold up in production.
Location
Based in California · Working globally
Status
Open to work
hanjixiong.com
Work
Selectedwork.
Five systems, one throughline: turn messy real-world data into decisions that hold up under scrutiny.
- 01
Causal Inference for Education
Research/Research leadEstimating what actually moves learning outcomes — quasi-experimental designs, sensitivity analysis, and honest uncertainty bounds for enrollment and intervention decisions.
2024 —
PythonDoWhyEconML - 02
Program Evaluation Platform
Data Platform/ArchitectA production pipeline that turns raw enrollment and assessment data into decision-ready causal estimates, with automated balance checks and reproducible specifications.
2023 — 2024
AirflowdbtSnowflake - 03
Applied AI Systems
AI Engineering/BuilderLLM-backed assistants wired into real operational workflows: evaluation harnesses, retrieval, guardrails, and cost discipline rather than demo-ware.
2025 —
PythonTypeScriptRAG - 04
Market Intelligence Scanner
Automation/BuilderA parallelized research engine that screens dozens of tickers nightly, scores momentum, valuation and sentiment, and delivers a ranked brief before the open.
2026
Pythonpandasyfinance - 05
Knowledge Digest Pipeline
Automation/BuilderNightly ingestion across research feeds and technical blogs, distilled into a morning brief that stays readable instead of becoming another unread backlog.
2026
PythonRSSLLM - 06
Portfolio Risk Engine
Quant/BuilderRisk-aware allocation across max-Sharpe, minimum-variance and hierarchical risk parity, with drawdown, VaR and correlation diagnostics run on every rebalance.
2026
PyPortfolioOptskfoliopandas-ta
About
On thework itself.
Focus
- Causal inference & program evaluation
- Education & enrollment analytics
- Applied AI in production workflows
- Data platform architecture
Based in
California, USA
I lead data science work where the cost of being wrong is measured in students, not clicks.
Most of my career has been spent on a single stubborn question: what actually improves learning outcomes, and how do you know? That question forces rigor — because the easy answers are almost always confounded.
Along the way I built the infrastructure to support it: pipelines that stay reproducible, evaluation harnesses that catch regressions before students do, and increasingly, AI systems that earn their place in the workflow instead of decorating it.
I care about evidence over narrative, and about building things that someone else can maintain after I hand them over.
10+
Years in data science
3
Causal programs shipped
∞
Bad hypotheses retired
Writing
Notes onthe craft.
Long-form writing on causal inference, evaluation methods, and putting AI into production workflows. No hot takes — only things worth rereading.
- Aug 2026
Why causal beats correlation in education
A correlation between attendance and outcomes tells you almost nothing about what to fund. Here is the design toolkit I reach for instead.
Causal Inference11 min - Jun 2026
Evaluating programs when you cannot randomize
Difference-in-differences, synthetic control, and propensity weighting — when each one lies to you, and how to catch it early.
Methods14 min - Mar 2026
Building AI that survives production
Most AI pilots die at the handoff. Evaluation harnesses, cost ceilings and refusal behavior are what keep them alive.
AI Engineering9 min - Dec 2025
What I look for when I hire data scientists
Not frameworks. Judgment about which question is worth answering, and the discipline to say when the data cannot answer it.
Leadership7 min