M.S. in Data Science, University of Washington
I love solving problems by harnessing the power of data and great storytelling! Interests include data engineering, applied AI and LLM automations, and experimentation.
- Sequel2SQL — LLM + RAG framework for SQL error diagnosis and correction using a database agent; improved baseline performance by 6% on the BIRD-CRITIC benchmark. UW capstone sponsored by Microsoft. Python, PydanticAI, ChromaDB, PostgreSQL — Repo
- E-Commerce Analytics Pipeline — Serverless AWS pipeline (S3, Lambda, Glue, Athena) using medallion architecture with data quality validation at each layer, serving analytics over 500K+ events. Spark, Parquet — Report
- google-places-analysis — Analyzed business review response behavior on Google Places data using matched treatment/control groups and paired t-tests across 760K+ records; found responder businesses received 9 additional reviews on average. Python, pandas, scipy, statsmodels — Report
- aws-cp-creds — Shell utility that automates copying AWS credentials. — Repo
| Category | Tools |
|---|---|
| Languages | Python, SQL, Shell Scripting |
| AI / LLMs | PydanticAI, MCP, RAG, AWS Bedrock |
| Statistics & Experimentation | scipy, statsmodels, A/B testing, hypothesis testing |
| Data Engineering | dbt, Airflow, Spark/PySpark, ETL/ELT |
| ML & Data | scikit-learn, pandas, polars |
| Databases | PostgreSQL, Snowflake |
| Cloud | AWS (S3, Lambda, Glue, Athena, Redshift) |
| Visualization & Apps | Tableau, Streamlit, FastAPI |
| Tools | Git |



