A Safe Action Is Not Enough: Feasible-Future Decoding for Vision-Language-Action Policies Paper • 2610.05166 • Published 2 days ago • 4
A Safe Action Is Not Enough: Feasible-Future Decoding for Vision-Language-Action Policies Paper • 2610.05166 • Published 2 days ago • 4
Composable Decoding on the Probability Simplex: Theory and Implementation Paper • 2609.34992 • Published 10 days ago • 11
Does This Action Still Explain the Task? Reverse Scoring for Diffusion Language Model Agents Paper • 2609.38536 • Published 9 days ago • 11
The Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning Paper • 2610.00332 • Published 9 days ago • 10
The Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning Paper • 2610.00332 • Published 9 days ago • 10
Composable Decoding on the Probability Simplex: Theory and Implementation Paper • 2609.34992 • Published 10 days ago • 11
Does This Action Still Explain the Task? Reverse Scoring for Diffusion Language Model Agents Paper • 2609.38536 • Published 9 days ago • 11
iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs Paper • 2609.24646 • Published 17 days ago • 9
The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus Paper • 2603.20105 • Published Mar 20 • 37
Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening Paper • 2601.21590 • Published Jan 29 • 14
Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers Paper • 2602.18292 • Published Feb 20 • 14
iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs Paper • 2609.24646 • Published 17 days ago • 9
Beam Search as Test-Time Self-Distillation via Counterfactual Contexts Paper • 2609.37041 • Published 9 days ago • 8
Beam Search as Test-Time Self-Distillation via Counterfactual Contexts Paper • 2609.37041 • Published 9 days ago • 8