arxiv:2609.32444
Kaixiang Zhao
kzhao5
AI & ML interests
None yet
Recent Activity
authored a paper 10 days ago
Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It upvoted a paper 10 days ago
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM ReasoningOrganizations
None yet