view article Article Measuring benchmark optimization in speech recognition +5 tlebryk02, bezzam, aliceebaird, dayllon, jpc, jens-hume-ai, tzirakis • 5 days ago • 53
view article Article We changed one line and the benchmark score moved 0.21 AUROC FINAL-Bench • 4 days ago • 10
view article Article Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers +1 tomaarsen, NohTow, raphaelsty • 8 days ago • 94
SWE-bench Collection SWE-bench (Lite, Verified, Multimodal, Multilingual) all in one place! • 5 items • Updated Dec 14, 2025 • 16
view article Article State of Open Models: Summer 2026 Observations +1 AdinaY, multimodalart, irenesolaiman • 12 days ago • 155
view article Article LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge LiquidAI • 14 days ago • 49
Muse Glimmer Collection Muse Glimmer 30B: multimodal agentic model for local deployment. BF16 weights, GGUF k-quants, ExecuTorch builds, DFlash drafter. • 4 items • Updated 16 days ago • 103
view article Article Meta is back with Muse Glimmer: local, agentic, multimodal, and open source +2 pcuenq, merve, burtenshaw, ariG23498 • 16 days ago • 108
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks? Paper • 2605.11086 • Published May 11 • 3
view article Article Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident +2 hlarcher, XciD, raphael-gl, chris-rannou • about 1 month ago • 480
view article Article LFM2.5-Encoders for Fast Long-Context Inference on CPU LiquidAI • 29 days ago • 68