cxiaoh
Matuxiaoh
AI & ML interests
Base model
Recent Activity
upvoted a paper 3 days ago
ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL upvoted a paper 3 days ago
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement upvoted a paper about 1 month ago
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-ImprovementOrganizations
None yet