Algorithm Researcher · UniDT
Algorithm research for large language model training.
I am an Algorithm Researcher at UniDT. My work covers pretraining, supervised fine-tuning, reinforcement learning, and evaluation, with additional research in robustness, uncertainty, privacy, and applied statistics.
SELECTED WORK
Publications & Projects
BLOGS
Latest Blogs
How continued pretraining and supervised fine-tuning teach different skills, with fill-in-the-middle, multi-token prediction, and segment-weighted losses.
Read PDFHow student-generated rollouts and GOLD enable cross-model teaching, with practical observations on early training instability and recovery.
Read PDFHow group rewards drive policy updates, with clipping, DAPO, GSPO, and practical reward design.
Read PDFAn intuitive introduction to feature-wise conditioning, neural-network integration, and training.
Read PDFEXPERIENCE
Current and previous roles
EDUCATION