This repository curates papers and blogs on long-context language modeling, covering surveys; efficient attention; KV-cache optimization; recurrent transformers and state-space models; position encoding & length extrapolation; long-context training; long-term memory; retrieval-augmented generation; in-context learning; context and model compression; long reasoning (long CoT); long video & image; long-horizon agents; long-text generation; inference acceleration; benchmarks & evaluation; and technical reports.
π₯ Must-read papers for LLM-based Long Context Modeling.
π₯β‘π₯ Thanks for all the great contributors on GitHub!
ππ€π I have the privilege of joining [LCLM-Horizon] and collaborating with them on providing a very complete and comprehensive scholarly survey (A Comprehensive Survey on Long Context Language Modeling) and repository (A-Comprehensive-Survey-For-Long-Context-Language-Modeling) dedicated to Long Context Language Modeling. I look forward to collaborating with them to advance research and deepen understanding in this area!
Taxonomy at a glance
flowchart LR
LCLM["Long-Context Modeling"]
LCLM --> A["Attention & KV Cache"]
LCLM --> T["Training & Alignment"]
LCLM --> M["Memory & RAG"]
LCLM --> C["Compression"]
LCLM --> R["Reasoning & Generation"]
LCLM --> V["Multimodal / Video"]
LCLM --> E["Evaluation & Acceleration"]
A --> A1["Sparse / Linear / IO-aware Attention"]
A --> A2["Eviction / Quantization / Offloading"]
T --> T1["Continual Pretraining / Long-SFT"]
T --> T2["Adaptation & RL for Long Context"]
M --> M1["Long-Term Memory"]
M --> M2["RAG / Hybrid Long-Context"]
C --> C1["Context Compression"]
C --> C2["Model Compression"]
R --> R1["Long CoT"]
R --> R2["Long-Form Text Generation"]
If you find our repository and survey useful for your research, please consider citing the following paper:
@article{liu2025comprehensive,
title={A Comprehensive Survey on Long Context Language Modeling},
author={Liu, Jiaheng and Zhu, Dawei and Bai, Zhiqi and He, Yancheng and Liao, Huanxuan and Que, Haoran and Wang, Zekun and Zhang, Chenchen and Zhang, Ge and Zhang, Jiebin and others},
journal={arXiv preprint arXiv:2503.17407},
year={2025}
}- π’ News
- π Papers
- 1. Survey Papers
- 2. Efficient Attention
- 3. KV-Cache Optimization
- 4. Recurrent Transformers
- 5. State Space Models & Hybrids
- 6. Position Encoding & Length Extrapolation
- 7. Long-Context Training
- 8. Long-Term Memory
- 9. Retrieval-Augmented Generation
- 10. In-Context Learning (Many-shot / Long-ICL)
- 11. Context Compression
- 12. Model Compression for Long Context
- 13. Long Reasoning (Long CoT)
- 14. Long Video & Image
- 15. Long-Horizon Agents
- 16. Long-form Text Generation
- 17. Inference Acceleration & Serving
- 18. Benchmarks & Evaluation
- 19. Technical Reports (Long-Context Models)
- 20. Blogs & Tutorials
- Acknowledgements
-
[2026.07.24]
-
[2026.07.23]
- Paper: Parameter-free Adaptive Sparse Attention via Compression-Based Content Selection
- Paper: Learning What Matters: Supervising Sparse Attention Routing with Causal Evidence Sets
- Paper: AttriMem: Attribution-Guided Process Feedback for Agent Memory Learning
- Paper: Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers
- Paper: Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering
- Paper: SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
- Paper: Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems
- Paper: MemTools: A Unified Research Framework for Interoperable Agent Memory
- Paper: Delivery, Not Storage: Cue-Anchored Working Memory as a Harness Property for Coding Agents
-
[2026.07.22]
- Paper: ArbiGraph: Arbitrarily Scalable Verifiable Task Graphs for Evaluating Context Management
- Paper: SLPO: Scaling Latent Reasoning via a Surrogate Policy
- Paper: Self Gradient Forcing: Native Long Video Extrapolation
- Paper: JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety
- Paper: PRO-LONG: Programmatic Memory Enables Long-Horizon Reasoning
- Paper: ArbiGraph: Arbitrarily Scalable Verifiable Task Graphs for Evaluating Context Management
-
[2026.07.21]
-
[2026.07.20]
- Paper: C$^2$KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference
- Paper: Is Progressive Disclosure All You Need for Long-Context Agents?
- Paper: AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report
- Paper: Surprise Forcing: What to Remember, When to Skip in Long Video Generation
- Paper: ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning
- Paper: FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications
- Paper: How Agent Skills Fail under Long Contexts: A White-Box Study in Code Auditing
-
[2026.07.19]
-
[2026.07.18]
-
[2026.07.17]
- Paper: SlotMem: Character-Addressable Internal Memory for Narrative Long Video Generation
- Paper: FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation
- Paper: Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos
- Paper: Recursive Harness Self-Improvement
- Paper: ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning
- Paper: DSWorld: A Data Science World Model for Efficient Autonomous Agents
- Paper: SlotMem: Character-Addressable Internal Memory for Narrative Long Video Generation
-
[2026.07.16]
-
[2026.07.15]
-
[2026.07.14]
- Paper: ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams
- Paper: Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
- Paper: MemoHarness: Agent Harnesses That Learn from Experience
- Paper: Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI Agents
- Paper: MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations
- Paper: VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression
- Paper: FOLIO: Focused Semantic Memory for Streaming Video Understanding
-
[2026.07.13]
- Paper: Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos
- Paper: LightMem-Ego: Your AI Memory for Everyday Life
- Paper: ToFu: A White-Box, Token-Efficient Agent Harness for Researchers
- Paper: StructAgent: Harness Long-horizon Digital Agents with Unified Causal Structure
- Paper: SLVMBench: Skill Learning from Video Memory
Month Papers
-
[2026.07.12]
-
[2026.07.11]
-
[2026.07.10]
-
[2026.07.09]
- Paper: OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators
- Paper: Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
- Paper: What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents
- Paper: Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
-
[2026.07.08]
- Paper: Infinite Worlds with Versatile Interactions
- Paper: The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
- Paper: Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE
- Paper: Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing
- Paper: Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity
- Paper: AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning
-
[2026.07.07]
- Paper: AlayaWorld: Long-Horizon and Playable Video World Generation
- Paper: Imagined Rollouts are Kinematic, Not Dynamic: A Diagnosis of Long-Horizon World-Model Failure
- Paper: TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
- Paper: DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression
- Paper: FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference
- Paper: AlayaWorld: Long-Horizon and Playable Video World Generation
-
[2026.07.06]
- Paper: Multiplayer Interactive World Models with Representation Autoencoders
- Paper: Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval
- Paper: CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents
- Paper: KVpop -- Key-Value Cache Compression with Predictive Online Pruning
- Paper: Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory
-
[2026.07.04]
-
[2026.07.03]
-
[2026.07.02]
-
[2026.07.01]
-
[2026.06.30]
- Paper: ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL
- Paper: RaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM Inference
- Paper: SeKV: Resolution-Adaptive KV Cache with Hierarchical Semantic Memory for Long-Context LLM Inference
- Paper: CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts
- Paper: MemLearner: Learning to Query Context memory for Video World Models
-
[2026.06.29]
- Paper: SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions
- Paper: Diagnosing and Mitigating Context Rot in Long-horizon Search
- Paper: Predict, Reuse, and Repair: Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding
- Paper: Morphing into Hybrid Attention Models
- Paper: LLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via a Proprioceptive Dashboard
-
[2026.06.28]
-
[2026.06.26]
-
[2026.06.25]
- Paper: DMV-Bench: Diagnosing Long-Horizon Multimodal Agents' Visual Memory with Incidental Cue Injection
- Paper: Information-Aware KV Cache Compression for Long Reasoning
- Paper: ProtoKV: Streaming Video Understanding under Delayed Query with Summary-State Memory
- Paper: Erase-then-Delta Attention: Decoupling Erase and Write Addresses in Delta-Rule Linear Attention
- Paper: DMV-Bench: Diagnosing Long-Horizon Multimodal Agents' Visual Memory with Incidental Cue Injection
-
[2026.06.24]
-
[2026.06.23]
-
[2026.06.22]
-
[2026.06.20]
-
[2026.06.19]
-
[2026.06.18]
- Paper: Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning
- Paper: ADaPT: Token-Level Decoupling for Efficient Large Reasoning Models
- Paper: CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs
- Paper: HydraHead: From Head-Level Functional Heterogeneity to Specialized Attention Hybridization
- Paper: Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning
-
[2026.06.17]
-
[2026.06.16]
-
[2026.06.15]
- Paper: TokenPilot: Cache-Efficient Context Management for LLM Agents
- Paper: Taylor-Calibrate: Principled Initialization for Hybrid Linear Attention Distillation
- Paper: Long-Context Modeling via GSS-Transformer Hybrid Architecture with Learnable Mixing
- Paper: Tangram: Unlocking Non-Uniform KV Cache Compression for Efficient Multi-turn LLM Serving
-
[2026.06.14]
-
[2026.06.12]
-
[2026.06.11]
- Paper: EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments
- Paper: MiniMax Sparse Attention
- Paper: Learning What to Remember: A Cognitively Grounded Multi-Factor Value Model for Agentic Memory
- Paper: Can I Buy Your KV Cache?
- Paper: Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learning
-
[2026.06.10]
-
[2026.06.09]
-
[2026.06.08]
- Paper: IS-CoT: Breaking the Long-form Generation Collapse via Interleaved Structural Thinking
- Paper: Memory Beyond Recall: A Dual-Process Cognitive Memory System for Self-Evolving LLM Agents
- Paper: H2HMem: A Multimodal Memory Benchmark for Agents in Human-Human Interactions
- Paper: FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention
- Paper: End-to-End Context Compression at Scale
-
[2026.06.07]
- Paper: Sparrow: Sparse Rollout for Stable and Efficient Long-context RL of Large Language Models
- Paper: Look Less, Reason More: Block-wise Attention Skipping for Efficient Multimodal LLMs
- Paper: From Player to Master: Enhancing Test-Time Learning of LLM Agents via Reinforcement Learning over Memory (ICML 2026)
-
[2026.06.06]
-
[2026.06.05]
-
[2026.06.04]
-
[2026.06.03]
- Paper: Cartridges at Scale: Training Modular KV Caches over Large Document Collections
- Paper: LazyAttention: Efficient Retrieval-Augmented Generation with Deferred Positional Encoding (ICML 2026)
- Paper: SparDA: Sparse Decoupled Attention for Efficient Long-Context LLM Inference
- Paper: Depth-Attention: Cross-Layer Value Mixing for Language Models
- Paper: Video2LoRA: Parametric Video Internalization for Vision-Language Models
- Paper: Plan, Watch, Recover: A Benchmark and Architectures for Proactive Procedural Assistance
- Paper: Rethinking Continual Experience Internalization for Self-Evolving LLM Agents
- Paper: Learning While Acting: A Skill-Enhanced Test-Time Co-Evolution Framework for Online Lifelong Learning Agents
-
[2026.06.02]
- Paper: HybridThinker: Efficient Chain-of-Thought Reasoning via Compressed Memory and Transient Thought Steps
- Paper: MemTrain: Self-Supervised Context Memory Training
- Paper: KVarN: Variance-Normalized KV-Cache Quantization Mitigates Error Accumulation in Reasoning Tasks
- Paper: Value-Aware Stochastic KV Cache Eviction for Reasoning Models
-
[2026.06.01]
- Paper: RESTORE: Improving Visual Token Reduction via Rectifying Distortions for Efficient Multimodal LLM Inference (ICML 2026)
- Paper: PaSBench-Video: A Streaming Video Benchmark for Proactive Safety Warning
- Paper: LayerRoute: Input-Conditioned Adaptive Layer Skipping via LoRA Fine-Tuning for Agentic Language Models
- Paper: Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning
- Paper: Do Transformers Need Three Projections? Systematic Study of QKV Variants
Paper entries live under papers/ so this README stays under GitHub's homepage size limit.
For an interactive chapter reader (search + in-page paper cards), open the
project homepage.
Attention, recurrence & systems
Training, position & memory
Compression, reasoning & multimodal
Evaluation & reports
Please contact me if I miss your names in the list, I will add you back ASAP!