Tail-Influence Sampling for CVaR Policy Evaluation Paper • 2609.38096 • Published 11 days ago • 28
Tail-Influence Sampling for CVaR Policy Evaluation Paper • 2609.38096 • Published 11 days ago • 28
LeanPolish: Verified Supervision for Lean Proof Compression Paper • 2609.38384 • Published 11 days ago
Tail-Influence Sampling for CVaR Policy Evaluation Paper • 2609.38096 • Published 11 days ago • 28
Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning Paper • 2504.11354 • Published Apr 15, 2025 • 7
Quaternion recurrent neural network with real-time recurrent learning and maximum correntropy criterion Paper • 2402.14227 • Published Feb 22, 2024
Does This Action Still Explain the Task? Reverse Scoring for Diffusion Language Model Agents Paper • 2609.38536 • Published 11 days ago • 11
Composable Decoding on the Probability Simplex: Theory and Implementation Paper • 2609.34992 • Published 12 days ago • 11
The Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning Paper • 2610.00332 • Published 11 days ago • 10
iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs Paper • 2609.24646 • Published 19 days ago • 9