-
LTX-2: Efficient Joint Audio-Visual Foundation Model
Paper • 2601.03233 • Published • 196 -
MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head
Paper • 2601.07832 • Published • 53 -
Motion Attribution for Video Generation
Paper • 2601.08828 • Published • 72 -
Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
Paper • 2601.19895 • Published • 27
Collections
Discover the best community collections!
Collections including paper arxiv:2606.13392
-
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
Paper • 2607.02980 • Published • 84 -
Gemma 4 Technical Report
Paper • 2607.02770 • Published • 81 -
SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe
Paper • 2607.03451 • Published • 35 -
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Paper • 2607.05804 • Published • 20
-
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence
Paper • 2606.14777 • Published • 217 -
VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models
Paper • 2606.16140 • Published • 127 -
FastContext: Training Efficient Repository Explorer for Coding Agents
Paper • 2606.14066 • Published • 95 -
MiniMax Sparse Attention
Paper • 2606.13392 • Published • 168
-
LFM2 Technical Report
Paper • 2511.23404 • Published • 73 -
MiniMax Sparse Attention
Paper • 2606.13392 • Published • 168 -
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
Paper • 2603.12201 • Published • 69 -
FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention
Paper • 2606.09079 • Published • 68
-
Why Fine-Tuning Encourages Hallucinations and How to Fix It
Paper • 2604.15574 • Published • 25 -
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Paper • 2604.24763 • Published • 70 -
Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora
Paper • 2604.24819 • Published • 90 -
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
Paper • 2604.26752 • Published • 114
-
MiniMax-01: Scaling Foundation Models with Lightning Attention
Paper • 2501.08313 • Published • 304 -
Lizard: An Efficient Linearization Framework for Large Language Models
Paper • 2507.09025 • Published • 19 -
On the Expressiveness of Softmax Attention: A Recurrent Neural Network Perspective
Paper • 2507.23632 • Published • 6 -
Causal Attention with Lookahead Keys
Paper • 2509.07301 • Published • 22
-
Code as Agent Harness
Paper • 2605.18747 • Published • 224 -
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 198 -
From Context to Skills: Can Language Models Learn from Context Skillfully?
Paper • 2604.27660 • Published • 73 -
PhysBrain 1.0 Technical Report
Paper • 2605.15298 • Published • 61
-
MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild
Paper • 2603.17187 • Published • 141 -
Attention Residuals
Paper • 2603.15031 • Published • 195 -
MOSS-TTS Technical Report
Paper • 2603.18090 • Published • 16 -
MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens
Paper • 2603.23516 • Published • 52
-
LTX-2: Efficient Joint Audio-Visual Foundation Model
Paper • 2601.03233 • Published • 196 -
MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head
Paper • 2601.07832 • Published • 53 -
Motion Attribution for Video Generation
Paper • 2601.08828 • Published • 72 -
Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
Paper • 2601.19895 • Published • 27
-
MiniMax-01: Scaling Foundation Models with Lightning Attention
Paper • 2501.08313 • Published • 304 -
Lizard: An Efficient Linearization Framework for Large Language Models
Paper • 2507.09025 • Published • 19 -
On the Expressiveness of Softmax Attention: A Recurrent Neural Network Perspective
Paper • 2507.23632 • Published • 6 -
Causal Attention with Lookahead Keys
Paper • 2509.07301 • Published • 22
-
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
Paper • 2607.02980 • Published • 84 -
Gemma 4 Technical Report
Paper • 2607.02770 • Published • 81 -
SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe
Paper • 2607.03451 • Published • 35 -
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Paper • 2607.05804 • Published • 20
-
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence
Paper • 2606.14777 • Published • 217 -
VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models
Paper • 2606.16140 • Published • 127 -
FastContext: Training Efficient Repository Explorer for Coding Agents
Paper • 2606.14066 • Published • 95 -
MiniMax Sparse Attention
Paper • 2606.13392 • Published • 168
-
LFM2 Technical Report
Paper • 2511.23404 • Published • 73 -
MiniMax Sparse Attention
Paper • 2606.13392 • Published • 168 -
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
Paper • 2603.12201 • Published • 69 -
FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention
Paper • 2606.09079 • Published • 68
-
Code as Agent Harness
Paper • 2605.18747 • Published • 224 -
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 198 -
From Context to Skills: Can Language Models Learn from Context Skillfully?
Paper • 2604.27660 • Published • 73 -
PhysBrain 1.0 Technical Report
Paper • 2605.15298 • Published • 61
-
Why Fine-Tuning Encourages Hallucinations and How to Fix It
Paper • 2604.15574 • Published • 25 -
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Paper • 2604.24763 • Published • 70 -
Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora
Paper • 2604.24819 • Published • 90 -
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
Paper • 2604.26752 • Published • 114
-
MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild
Paper • 2603.17187 • Published • 141 -
Attention Residuals
Paper • 2603.15031 • Published • 195 -
MOSS-TTS Technical Report
Paper • 2603.18090 • Published • 16 -
MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens
Paper • 2603.23516 • Published • 52