Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning Paper • 2609.03430 • Published 2 days ago • 158
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Paper • 2609.04098 • Published 2 days ago • 72
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 2 days ago • 73
Post-Training Language Models for Gold-Medal Performance in Coding Competitions Paper • 2609.02849 • Published 3 days ago • 9
Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model Paper • 2607.22083 • Published Jul 27 • 10
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Paper • 2608.30320 • Published 5 days ago • 51
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 5 days ago • 139
Self-Improving Pretraining: using post-trained models to pretrain better models Paper • 2601.21343 • Published Jan 29 • 21
A Programming Paradigm for Spatiotemporal Composability Paper • 2608.25512 • Published 10 days ago • 15
D^3-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation Paper • 2608.24987 • Published 11 days ago • 26
Agent-G^2: Gaussian Guidance for Agentic Reinforcement Learning Paper • 2608.23318 • Published 12 days ago • 32
Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher Paper • 2608.26872 • Published 9 days ago • 82
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 12 days ago • 205
Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation Paper • 2608.15062 • Published 10 days ago • 11