Computer Science at Duke University with a minor in Machine Learning and Artificial Intelligence. I build AI systems end to end, then design the evaluation needed to show where they work, where they fail, and whether their outputs deserve trust.
I am targeting Summer 2027 AI & machine learning and SDE internships.
My strongest work sits at the intersection of applied ML, LLM systems, evaluation, and product engineering. I enjoy taking a system from experiment to deployed interface, adding measurable quality gates, and documenting the tradeoffs honestly.
Current focus areas include LLM-as-judge methodology, execution-based evaluation, human preference studies, open-weight inference with vLLM and SLURM, and preference alignment with QLoRA and DPO.
| Project | What it proves | Stack and evidence |
|---|---|---|
| AI Model Advisor (live system) | Production model governance across artifact security, adversarial safety, institutional evaluation, and public benchmarks | Flask, Postgres, Docker, LLM-as-judge, execution scoring, vLLM, SLURM, 1,272 Python tests |
| LyricMood (live demo) | Fine-tuned text classification with interpretable predictions and semantic retrieval over 76,000 songs | DistilBERT, ONNX, FastAPI, SHAP, Qdrant, 106 tests |
| Deadline Tracker (live app) | A production web app built around a real daily workflow, with natural-language entry and university integrations | Next.js, TypeScript, Supabase, Vercel, PWA, 781 tests |
| Trading Agent Analyzer (live demo) | A usable web product around a multi-agent research engine, with durable run history and cost controls | LangGraph, Streamlit, Postgres, Docker, DeepSeek |
At Duke Code+, I own major parts of the efficacy and report-card experience for AI Model Advisor. My work includes rubric and execution-based LLM evaluation, a frozen evaluation contract, a six-rater judge validation study, DCC-hosted open-weight inference, cost-versus-performance analysis, and the frontend used to launch and compare runs.
I also contribute to sickle-cell-disease data research at the Duke Global Health Institute.
| Area | Tools |
|---|---|
| ML and alignment | PyTorch, Hugging Face Transformers, TRL, PEFT, QLoRA, DPO, ONNX |
| LLM systems | LiteLLM, LangGraph, LLM-as-judge, vLLM, SLURM, structured evaluation |
| Backend and data | Python, FastAPI, Flask, Postgres, Supabase, Qdrant |
| Product and operations | TypeScript, Next.js, React, Docker, GitHub Actions, Vercel |

