-
Diffusion Augmented Agents: A Framework for Efficient Exploration and Transfer Learning
Paper • 2407.20798 • Published • 24 -
Offline Reinforcement Learning for LLM Multi-Step Reasoning
Paper • 2412.16145 • Published • 38 -
REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models
Paper • 2501.03262 • Published • 102 -
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Paper • 2502.18449 • Published • 75
Collections
Discover the best community collections!
Collections including paper arxiv:2607.12395
-
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
Paper • 2607.02980 • Published • 84 -
Gemma 4 Technical Report
Paper • 2607.02770 • Published • 83 -
SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe
Paper • 2607.03451 • Published • 35 -
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Paper • 2607.05804 • Published • 20
-
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
Paper • 2602.10693 • Published • 39 -
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
Paper • 2603.09229 • Published • 85 -
DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use
Paper • 2603.11076 • Published • 5 -
LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning
Paper • 2603.21065 • Published • 80
-
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
Paper • 2607.12395 • Published • 91 -
Meshy T2: Fast Native Mesh Generation with Flow Matching
Paper • 2607.28675 • Published • 61 -
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
Paper • 2608.01964 • Published • 187
-
In-Context World Modeling for Robotic Control
Paper • 2606.26025 • Published • 40 -
Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs
Paper • 2606.27378 • Published • 60 -
Infinite Worlds with Versatile Interactions
Paper • 2607.07534 • Published • 46 -
MentalThink: Shaping Thoughts in Mental SVG World
Paper • 2607.03530 • Published • 13
-
LTX-2: Efficient Joint Audio-Visual Foundation Model
Paper • 2601.03233 • Published • 197 -
MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head
Paper • 2601.07832 • Published • 53 -
Motion Attribution for Video Generation
Paper • 2601.08828 • Published • 72 -
Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
Paper • 2601.19895 • Published • 27
-
Diffusion Augmented Agents: A Framework for Efficient Exploration and Transfer Learning
Paper • 2407.20798 • Published • 24 -
Offline Reinforcement Learning for LLM Multi-Step Reasoning
Paper • 2412.16145 • Published • 38 -
REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models
Paper • 2501.03262 • Published • 102 -
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Paper • 2502.18449 • Published • 75
-
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
Paper • 2607.12395 • Published • 91 -
Meshy T2: Fast Native Mesh Generation with Flow Matching
Paper • 2607.28675 • Published • 61 -
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
Paper • 2608.01964 • Published • 187
-
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
Paper • 2607.02980 • Published • 84 -
Gemma 4 Technical Report
Paper • 2607.02770 • Published • 83 -
SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe
Paper • 2607.03451 • Published • 35 -
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Paper • 2607.05804 • Published • 20
-
In-Context World Modeling for Robotic Control
Paper • 2606.26025 • Published • 40 -
Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs
Paper • 2606.27378 • Published • 60 -
Infinite Worlds with Versatile Interactions
Paper • 2607.07534 • Published • 46 -
MentalThink: Shaping Thoughts in Mental SVG World
Paper • 2607.03530 • Published • 13
-
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
Paper • 2602.10693 • Published • 39 -
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
Paper • 2603.09229 • Published • 85 -
DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use
Paper • 2603.11076 • Published • 5 -
LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning
Paper • 2603.21065 • Published • 80
-
LTX-2: Efficient Joint Audio-Visual Foundation Model
Paper • 2601.03233 • Published • 197 -
MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head
Paper • 2601.07832 • Published • 53 -
Motion Attribution for Video Generation
Paper • 2601.08828 • Published • 72 -
Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
Paper • 2601.19895 • Published • 27