Where Competence Breaks: Tiered Evaluation of Financial Task Performance
Tiered, gated evaluation of finance agent tasks
None defined yet.
World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments
Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows
Tiered, gated evaluation of finance agent tasks
A PDF-grounding benchmark for healthcare document work
Explore trust scores for AI benchmark results
Scores a personal assistant by what it did on the device
Evaluate AI models on journal entry audit tasks
Co-evolutionary adversarial training demo (DA vs CA)
Explore and compare RL task trajectories
RL env & benchmark for enterprise BA agents
Cached replays of 140 agent-to-agent negotiation rollouts
RL environment for sales & revenue-ops agents
RL environment & benchmark for clinical EHR agents
Run and evaluate simulated iPhone assistant tasks
Generate a personalized ad and receive a quality score
Interactive demo for the MedMosaic medical-audio benchmark