I study how AI systems fail—and how evaluations can capture those failures more reliably.
My work focuses on AI evaluation, benchmark design, agent experiments, and reproducible failure analysis. I previously worked on Seed model evaluation at ByteDance and now continue this work as an independent researcher and builder.
- LLM and agent evaluation
- Benchmark design and evaluation health
- Failure analysis and reproducible experiments
- Evaluation tooling and research workflows
- More Effort, No Miracle — a multi-model study of reasoning-effort saturation on DeepSWE (code and derived data)
I publish research notes in English and Chinese.
Open to collaboration on AI evaluation, benchmark research, and agent reliability.


