Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 

Repository files navigation

Xiaobo Wang | 王晓博

PhD Student @ USTC · Research Intern @ BIGAI

Reward Modeling · LLM Alignment · Continual Learning · Agent Memory

Homepage · Google Scholar · Hugging Face · Email

Unique visitor-days

About

I work on building language models that keep improving from experience through reliable reward signals, robust alignment, continual adaptation, and memory.

I am currently a PhD student at the University of Science and Technology of China and a research intern at the Beijing Institute for General Artificial Intelligence (BIGAI). I am open to research collaboration.

Selected Research

  • SAVE — On-policy feedback for reward model self-supervised improvement. [Paper] [Project]
  • UAPO — Uncertainty-aware preference optimization. EMNLP 2025. [Paper] [Data]
  • PoliCon — Evaluating LLMs on diverse political consensus objectives. ICLR 2026. [Paper] [Code]
  • ICE — Learning knowledge from self-induced contextual distributions. ICLR 2025. [Paper] [Code]
  • RAM — An ever-improving memory system that learns from communication. [Paper] [Code]

Current Projects

About

Xiaobo Wang's GitHub profile

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors