How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data Paper • 2604.13977 • Published Apr 15 • 1
Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections Paper • 2603.12180 • Published Mar 12 • 65
Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL Paper • 2602.03773 • Published Feb 3 • 14
Imprecise Label Learning: A Unified Framework for Learning with Various Imprecise Label Configurations Paper • 2305.12715 • Published May 22, 2023
SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations Paper • 2512.14080 • Published Dec 16, 2025 • 10
FFB: A Fair Fairness Benchmark for In-Processing Group Fairness Methods Paper • 2306.09468 • Published Jun 15, 2023 • 1
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks Paper • 2410.18210 • Published Oct 23, 2024
Large Reasoning Models Learn Better Alignment from Flawed Thinking Paper • 2510.00938 • Published Oct 1, 2025 • 60
Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations Paper • 2411.10414 • Published Nov 15, 2024 • 1
Large Reasoning Models Learn Better Alignment from Flawed Thinking Paper • 2510.00938 • Published Oct 1, 2025 • 60
view post Post 9593 We're kick-starting the process of Transformers v5, with @ArthurZ and @cyrilvallez !v5 should be significant: we're using it as a milestone for performance optimizations, saner defaults, and a much cleaner code base worthy of 2025.Fun fact: v4.0.0-rc-1 came out on Nov 19, 2020, nearly five years ago! See translation 6 replies · 🚀 19 19 👍 10 10 🔥 6 6 + Reply
view post Post 2784 New post is live!This time we cover some major updates to transformers.🤗 See translation 2 replies · 🤗 1 1 + Reply