DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 3 days ago • 126 • 3
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 3 days ago • 126
Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training Paper • 2609.14306 • Published 7 days ago • 10
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090 Paper • 2608.27370 • Published 24 days ago • 40 • 7
Improving the matrix multiplication exponent with modern optimization and AlphaEvolve Paper • 2608.16884 • Published Aug 17 • 19
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning Paper • 2608.09888 • Published Aug 10 • 784
Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration Paper • 2605.17423 • Published May 17 • 31
Seedance 2.0: Advancing Video Generation for World Complexity Paper • 2604.14148 • Published Apr 15 • 171
Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning Paper • 2604.12374 • Published Apr 14 • 39
Large Language Models Align with the Human Brain during Creative Thinking Paper • 2604.03480 • Published Apr 3 • 7
Beyond the Assistant Turn: User Turn Generation as a Probe of Interaction Awareness in Language Models Paper • 2604.02315 • Published Apr 3 • 5