On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators
-
Updated
Aug 31, 2026 - Python
On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators
Tools for OpenDataArena: Fair, Open, and Transparent Arena for Data
Singer is a complete training and evaluation pipeline for classical Chinese regulated verse (格律诗).
600ms-budget full-duplex voice agent — cascade ASR·RAG·LLM·TTS (458ms) + end-to-end Omni (250ms) | Voice · ASR · TTS · LLM · Infra · Posttraining
A general purpose rig for fine-tuning LLMs
A Data-Centric RL and On-Policy Distillation Framework for LLM Post-Training
Modern, improved reproduction of the prominent GPT 2 (124M)
TARDIS is a causal video-generation framework for temporal consistency that transports predictable content from previous frames via confidence-weighted motion transport, projects the residual into a transport-orbit tangent space (handled by a deterministic corrector) versus a normal quotient space (handled by diffusion)
To associate your repository with the posttraining topic, visit your repo's landing page and select "manage topics."