Meera Joshi
mjoshi-01
ยท
AI & ML interests
self-play
Recent Activity
upvoted a paper 1 day ago
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL upvoted a paper 1 day ago
RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling upvoted a paper 1 day ago
Nereus: Adaptive Parallelism for LLM Post-TrainingOrganizations
None yet