I make voice & speech AI measurable β building the evals, benchmarks, and observability that show how good our voice agents really are, and where they break. Product + data + agents.
What I do
- ποΈ Voice agents β real-time speech-to-speech, Voice Live, STT / TTS quality
- π Eval & benchmarking β accuracy, latency & quality harnesses; competitive bake-offs
- π Observability β usage & quality dashboards over large-scale telemetry
- π€ Agentic tooling β MCP servers, Copilot CLI skills, natural-language data agents
π§° Stack: Python Β· KQL / Kusto Β· Azure AI Foundry Β· MCP Β· GitHub Copilot CLI
βοΈ Recent writing Β· Azure AI Foundry blog
![]() Post-Stream Refinement: GA |
![]() Multilingual Real-Time Transcription |
![]() Voice Live Evaluation Harness |
![]() Post-Stream Refinement: Preview |
π reading-dna β two AI models compete to recommend books from your Goodreads history π






