Rubric-based RL that normalizes judged quality over only the responses satisfying the hard constraints.
🛵 Should they come looking for me, I inten
Aman Behera
beingamanforever
·
AI & ML interests
Long Horizon Agentic RL, OPD, Generative Engine Optimization, Performance Optimisation
Recent Activity
liked a dataset 1 day ago
jbarrow/CommonForms liked a model 3 days ago
Contrastive-LM/CLM-v0.1-8B liked a model 22 days ago
infly/Infinity-Parser2-ProOrganizations
None yet