DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 5 days ago • 159
view article Article Trained 210M text-to-image model from scratch on one GPU: what actually mattered ivanmikhnenkov • 13 days ago • 10
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation Paper • 2609.11638 • Published 12 days ago • 695
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published 15 days ago • 370
Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling Paper • 2608.30821 • Published 22 days ago • 67
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 21 days ago • 119
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Paper • 2609.02749 • Published 20 days ago • 401
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published 19 days ago • 243
Compile by Training: Turning Natural-Language Specifications into Local Neural Functions Paper • 2609.04199 • Published 19 days ago • 326
view article Article NeoMME: an efficient Multimodal-native and Multilingual Encoder Hcompany • 19 days ago • 110
view article Article Training a coding model to paint watercolours with TRL and OpenEnv sergiopaniego • 19 days ago • 71
view article Article Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps +1 iamleonie, burtenshaw, sergiopaniego • 19 days ago • 130
view article Article Open Yap 1K: 1,000 hours of full-duplex natural conversation, free for commercial use TheAgenticDataCompany • 18 days ago • 20
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Paper • 2608.30320 • Published 22 days ago • 63
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 22 days ago • 97
DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution Paper • 2608.31106 • Published 22 days ago • 96