Rubric-based RL that normalizes judged quality over only the responses satisfying the hard constraints.
🛵 Should they come looking for me, I inten
Aman Behera
beingamanforever
·
AI & ML interests
Long Horizon Agentic RL, OPD, Generative Engine Optimization, Performance Optimisation
Recent Activity
liked a dataset 1 day ago
krutrim-ai-labs/ocr_rotation_bench liked a model 1 day ago
qualcomm/MobileNet-v3-Small liked a dataset 4 days ago
jbarrow/CommonFormsOrganizations
None yet