Léo Martin
leo-mart
·
AI & ML interests
Reinforcement learning, reward modeling, RLHF, policy optimization, offline RL
Recent Activity
liked a dataset about 11 hours ago
Gene829/gene-reinforcement-learning-instruct upvoted a paper about 11 hours ago
SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL upvoted a paper about 11 hours ago
Tactile-JEPA: Topology-Aware Self-Supervised Representation Learning for Distributed Tactile SensorsOrganizations
None yet