Paperglide public library — simplified academic papers Browse every academic paper Paperglide has simplified. Each page offers a TL;DR, semantic sections, and a link to the original source. Browse by topic cs.AI paperscs.LG paperscs.CL paperscs.CV paperscs.NE papersstat.ML papers All simplified papers Looped World Models (LoopWM)MiniMax Sparse AttentionGeometric Action Model for Robot Policy LearningDOPD: Dual On-Policy DistillationFastContext: Training Efficient Repository Explorer for Coding AgentsAttention Is All You NeedScaling Laws for Language ModelsDeepSeek-V3 Technical ReportDeepSeek-V4: Towards Highly Efficient Million-Token Context IntelligenceDenoising Diffusion Probabilistic ModelsABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPUOmniOpt: Taxonomy, Geometry, and Benchmarking of Modern OptimizersTimeLens2: Generalist Video Temporal Grounding with Multimodal LLMsUI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent LearningDataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data PipelinesRAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLMProgram-as-Weights: A Programming Paradigm for Fuzzy FunctionsRynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation ModelAREX: Towards a Recursively Self-Improving Agent for Deep ResearchAutoIndex: Learning Representation Programs for RetrievalVideoChat3: Fully Open Video MLLM for Efficient and Generalist Video UnderstandingMulti-Turn On-Policy Distillation with Prefix ReplayVidu S1: A Real-Time Interactive Video Generation ModelPredictive Divergence Masks for LLM RLWeak-to-Strong Generalization via Direct On-Policy DistillationThe Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement LearningLongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU BudgetABot-N1: Toward a General Visual Language Navigation Foundation ModelDeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable EnvironmentEvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary WorldAccurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural ReasoningSEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement LearningABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal MemoryOmniOpt: Taxonomy, Geometry, and Benchmarking of Modern OptimizersRead It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image GenerationSearch Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual GenerationText Template Tokens Are Implicit Semantic Registers in Diffusion TransformersSWE-Pruner Pro: The Coder LLM Already Knows What to PruneGenerative World Renderer at the Speed of PlayUI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent LearningProgram-as-Weights: A Programming Paradigm for Fuzzy FunctionsSynthDocBench: Controlled Benchmark for Long-Context Visual Document UnderstandingXiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World TrajectoriesMage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and EditingLoop the Loopies!Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied LearningSLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPODLong-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based GradingResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and BlogSearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent CollaborationPixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel SpacexHC: Expanded Hyper-ConnectionsKnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and SkillHOMIE: Human-object Centric Video Personalization via Multimodal Intelligent EnchancementAlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical ReportReferTrack: Referring Then Tracking for Embodied Visual TrackingDual Latent Memory in Vision-Language-Action Models for Robotic ManipulationDataPrep-Bench: Benchmarking LLMs as Training Data PreparatorsVideo Generation Models are General-Purpose Vision LearnersScalable Visual Pretraining for Language IntelligenceOvisOCR2 Technical ReportCura 1T: Specialized Model for Agentic HealthcareHierarchical Sparse Attention Done Right: Toward Infinite Context ModelingABot-Earth 0.5: Generative 3D Earth Model4D Human-Scene Reconstruction from Low-Overlap CapturesEvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive EnvironmentsAgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM AgentsEmbodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous RobotsResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference OutcomesVisual Contrastive Self-DistillationScaling Mixture-of-Experts Video Pretraining for Embodied IntelligenceOrca: The World is in Your MindApple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical IntelligenceAgentCompass: A Unified Evaluation Infrastructure for Agent CapabilitiesBadWAM: When World-Action Models Dream Right but Act WrongKwai Keye-VL-2.0 Technical ReportLightMem-Ego: Your AI Memory for Everyday LifeGigaChat Audio: Time-aware Large Audio Language ModelGemma 4 Technical ReportBridging Interleaved Multi-Modal Reasoning as a Unified Decision ProcessSubliminal Clocks: Latent Time Modelling in Diffusion Language ModelsShow, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM TextMANCE: Manifold Aware Concept ErasureVision as Unified Multimodal GenerationVision Pretraining for Dense Spatial PerceptionGigaWorld-1: A Roadmap to Build World Models for Robot Policy EvaluationThe Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement LearningPolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image GuardrailsKeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video GenerationStale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement LearningGigaAM Multilingual: Foundation Model for Underrepresented LanguagesSkill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving SkillsOrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion TransformersMultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video GenerationSelf Gradient Forcing: Native Long Video ExtrapolationAgents' Last ExamOn-Policy Delta DistillationQwen-Music Technical ReportAgentic Abstention: Do Agents Know When to Stop Instead of Act?Wan-Streamer v0.2: Higher Resolution, Same LatencyRecGPT-V3 Technical ReportReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video StreamsREBASE: Reference-Background Subspace Elimination for Training-Free In-Context SegmentationRESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal ResourcesFlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality MimicryMolt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement LearningAgenticDataBench: A Comprehensive Benchmark for Data AgentsMulti-Turn Agentic Scientific Literature Search via Workflow InductionEVA-Client: A Unified Data Collection, Inference, and Deployment Framework for Embodied Policies on Real RobotsGigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearchMetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry PriorsFrom Human-Centric to Agentic Code Review: The Impact of Different Generations of Generative AI Technology on Review QualityGroup Entropy-Controlled Policy OptimizationSeerGuard: A Safety Framework for Mobile GUI Agents via World Model PredictionBeyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and RankingRecursive Harness Self-ImprovementVLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action HorizonUniClawBench: A Universal Benchmark for Proactive Agents on Real-World TasksFrom Pixels to States: Rethinking Interactive World Models as Game EnginesUniVR: Thinking in Visual Space for Unified Visual ReasoningConcurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump ProcessesSpectral Rewiring for Exploration, Purification, and Model MergingEvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic EnvironmentsSWE-Explore: Benchmarking How Coding Agents Explore RepositoriesMulti-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged SamplingInfinite Worlds with Versatile InteractionsIdeas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea GenerationBlind-Spots-Bench: Evaluating Blind Spots in Multimodal ModelsRxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and ImaginationMoebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level PerformanceInternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional GeneralizationAgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM AgentsQwen-AgentWorld: Language World Models for General AgentsMemSyco-Bench: Benchmarking Sycophancy in Agent MemoryDSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive GenerationLongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion ModelsACID: Action Consistency via Inverse Dynamics for Planning with World ModelsAdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and VerificationFunction-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation ModelsMuScriptor: An Open Model for Multi-Instrument Music TranscriptionUnderstanding Reasoning from Pretraining to Post-TrainingAn Exam for Active ObserversAre We Ready For An Agent-Native Memory System?Dockerless: Environment-Free Program Verifier for Coding AgentsELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE ServingMultimodal Continuous Reasoning via Asymmetric Mutual Variational LearningSkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of VibeLight-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term MemoryKVpop -- Key-Value Cache Compression with Predictive Online PruningNVIDIA-labs OO Agents: Native Python Object-Oriented AgentsAudio Interaction ModelMiniMax Sparse AttentionToward Generalist Autonomous Research via Hypothesis-Tree RefinementSeed2.0 Model Card: Towards Intelligence Frontier for Real-World ComplexityWorldDirector: Building Controllable World Simulators with Persistent Dynamic MemoryDo All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language RetrievalMultiplayer Interactive World Models with Representation AutoencodersLook Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA ModelsKnow Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue ResolutionSciForma: Structure-Faithful Generation of Scientific DiagramsEnvironment-free Synthetic Data Generation for API-Calling AgentsTencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task ConstructionColor Pass-Through via Camera-Display CouplingScaling Native Multimodal Pre-Training From ScratchAgentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture ProblemsPlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool EcosystemsSpatialClaw: Rethinking Action Interface for Agentic Spatial ReasoningParallelized Autoregressive Decoding for Omni-Modal Dense Video CaptioningLLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RLUnified Audio Intelligence Without Regressing on Text IntelligenceKronQ: LLM Quantization via Kronecker-Factored HessianTrust Region Policy DistillationSelf-Improvements in Modern Agentic Systems: A SurveyVideo = World + Event StreamDemystifying On-Policy Distillation: Roles, Pathologies, and RegulationsDiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D GenerationRedesign Mixture-of-Experts Routers with Manifold Power IterationDomain Arithmetic: One-Shot VLA Adaptation under Environmental ShiftsVideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query RefinementAutoMem: Automated Learning of Memory as a Cognitive SkilldOPSD: On-Policy Self-Distillation for Diffusion Language ModelsPerceptual Flow Matching for Few-Step Generative ModelingEdgeBench: Unveiling Scaling Laws of Learning from Real-World EnvironmentsMeasuring the Gap Between Human and LLM Research IdeasRoboTTT: Context Scaling for Robot PoliciesDiscrete Diffusion Models: A Unified Framework from Tokenization to GenerationSelf-Supervised Learning of Structured Dynamics from VideosLLMs Get Lost in Evolving User IntentSANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video GenerationS1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and GenerationCosmos 3: Omnimodal World Models for Physical AIOCC-RAG: Optimal Cognitive Core for Faithful Question AnsweringYour UnEmbedding Matrix is Secretly a Feature Lens for Text EmbeddingsRole-Agent: Bootstrapping LLM Agents via Dual-Role EvolutionInterleaveThinker: Reinforcing Agentic Interleaved GenerationBlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative DecodingRobust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?LiveEdit: Towards Real-Time Diffusion-Based Streaming Video EditingFORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search AgentsOpenRath: Session-Centered Runtime State for Agent Systems