Paperglide public library — simplified academic papers Browse every academic paper Paperglide has simplified. Each page offers a TL;DR, semantic sections, and a link to the original source. Browse by topic cs.AI paperscs.LG paperscs.CL paperscs.CV paperscs.NE papersstat.ML papers All simplified papers Looped World Models (LoopWM)Geometric Action Model for Robot Policy LearningDOPD: Dual On-Policy DistillationFastContext: Training Efficient Repository Explorer for Coding AgentsScaling Laws for Language ModelsDeepSeek-V4: Towards Highly Efficient Million-Token Context IntelligenceStudentSim: Training LLM-based Student SimulatorsNeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing HarnessQwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous DrivingCompile by Training: Turning Natural-Language Specifications into Local Neural FunctionsHarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?Terminal-Universe: Turning Agent Trajectories into Scalable Terminal EnvironmentsLLaDA-Image: Building Strong Image Generators with Fully Open Training RecipesAuK Technical Report: An Open-Source Foundational Model for Speech Generation and EditingRandom Attention: Rethinking KV Cache Eviction for Efficient ReasoningShow-Harness: Just a VLM Agent Can Play RobotsUnlocking Lossless Speedups in LLMs via Discrete DiffusionOmni Interaction Agent Technical ReportScaling Automatic Research Agents via World ModelsLatentPress: Context Compression Beyond Text and VisionBilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM SystemsRoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation LearningAgentGrad: Intervention-guided Prompt Optimization for Multi Agent SystemsEliciting Weak-to-Strong Generalization with On-Policy Reverse DistillationSMELT: Scaling Laws for Compute-Matched MoE Looped TransformersFlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning ExperienceProgrammable World ModelMacaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRARethinking On-Policy Distillation of Large Language Models II: One Training ExampleWhy Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLMIt Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement LearningBDH-CQ: In-Context Learning with Recurrent Latent ReasoningPuffin-World: Scaling a Unified Multimodal Model with Native 3D World StatesEnvHarness: Awakening Static Worlds for Agent LearningOpenART: Scaling Agent Red Teaming via Open-Ended Environment EvolutionSpark-to-Paper: End-to-End Research Paper Generation as a Composable SkillUI-Venus-2 Technical ReportLanguage Models Can Control Their Own AttentionOpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model PretrainingAspire: Can Models Self-Evolve from Vague Goals?Iris: Climbing to the Search FrontierDriveZero: End-to-End Driving Beyond Human DemonstrationsStateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness ScalingMiles v0.1: Production-Level Post-TrainingH3-World: Turning Language Understanding into World ControlMarigold V2: Revisiting Diffusion Transformers for Monocular Depth EstimationComBodied Agents: a New Paradigm of Human-Centric Agentic AIApodex 1.1: Scaling Agentic Intelligence for Complex WorkMask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking RolloutScal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D ReconstructionZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-trainingEditable Visual DesignMotion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken DialogueSelf-Supervised Visual On-Policy DistillationDemystifying Agent Skills: Why They Work-Until They Don'tKnowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-TrainingSwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot TasksENEAS: Embedding-guided Neural Ensemble for Adaptive SegmentationLongHorizon-Harness: Advancing Long-Horizon Agents for Real-World TasksThe Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video GenerationFrom Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request MixCo-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human DesignSWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code RefactoringAgentic Game Development as a Verifiable Trajectory Data Engine for Scaling World ModelsOn the Design Fundamentals of Pixel Text Representation LearningOn-Policy Self-Distillation without Any SupervisionWHALE: A Simple Recipe for Joint Harness-Weight OptimizationBeyond Retrieval: Progressive Latent Memory Evolution for Streaming Video UnderstandingWearableQA: A Benchmark for Health Reasoning over Real-World Wearable DataHarnessEval-W: Agentifying the Evaluation of Visual WorldsSemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code GenerationDr. Claw: An AI Scientist Workspace for Vibe ResearchFACET: Preserving Source Intent and Executable State in Terminal Task SynthesisBeyond Pixels: From Video Priors to 4D WorldsDoes On-Policy Distillation Really Distill? From Noisy Teacher to Self-ImprovementT1: Terminal Agent Reinforcement Learning for Long-Horizon TasksLast Translation BenchmarkAI4AI at Test-Time: Strong-to-Weak Capability Transfer via HarnessesAlaya-EVOKE: From Linear-Scaling Supervision to Endless WorldThe Attention Triangle in Audio-Video ModelsDAPD: Dual-Anchored Policy DistillationCORE: Improving Compositional Reasoning in MLLM Embedding via Reranker DistillationAgentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU RequirementsDreamX-Creator: Democratizing Native Audio-Video Generation at 2K ResolutionGigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System ArchitectureEvaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and SearchingNeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and InferenceWorldReward: Reward Modeling for Camera-Conditioned World ModelsDRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent TrainingLLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM RoutersAnnotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMsDART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling AgentsUncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to SystemZipTok3D: High-Fidelity 3D Tokenization with Compact Token PrefixesPACE: Towards Surfacing Hidden Conflicts in User RequestsCo-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RLWorldSculpt: Generating Compositional Worlds from Grounded VideosLucida: Parse, Generate, and Place for Composable Real-to-Sim Scene ModelingSafin-1: Safety from Within through Memory-Native State EvolutionSWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering AgentsJoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive DiffusionMechanist: AI as a Scientific Instrument for Discovering the Mechanisms of IntelligenceDreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic ManipulationVBVR-Pro: A Scalable and Verifiable Suite for Native Visual ReasoningFlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlowPAWBench: How Far Are We from Probabilistically Aligned World Modeling?AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language ModelingHunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and EditingSkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill LibrariesAgentOPSD: Recursive Self-Distillation for Agentic Reinforcement LearningEnvironment Evolution for Terminal AgentsUrbanGround: From Local Perception to Spatial Agency in a Real-Scale CityOuroboros: A Self-Developing Frontier Coding Agent with Reviewed Core EvolutionDiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error MemoryEnoki: Efficient Multi-Level Hallucination DetectionLearn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy OptimizationEchoWM: Open and Enterable Omnimodal World ModelsTTPO: Test-Time Policy OptimizationSelf-OPD: On-Policy Distillation for Flow Matching Models without TeacherA Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLossPrincipia: Relational Physics Tests for Video ModelsSAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive ExecutionSelect, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMsAsk Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive OptimizationDon't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and InferenceCausal Foundation ModelsScores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research AgentsDarwinX: Evolving Agent Harnesses Through Natural SelectionGenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative ModelingKimi K3: Open Frontier IntelligenceAskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis4DAnyone: Create Anyone in 4D from a Casual Monocular VideoABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit AssignmentSWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?WeMM-Embedding: WeChat Multi-Modal Embedding Technical ReportWhat Makes Good Agentic Data? An ACE Lens on Data Generation for LLM AgentsPercolation Dynamics in Optimization : Variance Cascades and Discrete Scale InvarianceBeneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMsQwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI AgentsTLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live StreamingASI-Bench: At the Dawn of Artificial SuperintelligenceStealing Reasoning Traces from Proprietary LLM APIsLarge Discovery Models: Empirically-grounded Model-Based Open-Ended SearchProgressive Agent Skill Generation via Reinforcement LearningInterpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive RetrievalVibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution TracesLet Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc CalibrationUniMate: One Unified Model to Animate Diverse SkeletonsMaxKernel: Agentic Kernel Generation for TPUsOn-Policy Self-Distillation in Diffusion ModelsWorldClaw: Agentic 3D Open-World Generation at ScaleMetis: Memory Foundation ModelToolArtist: Tool-Using Unified Multimodal Models for Agentic Image GenerationHarness-of-Harness: Multi-Day Autonomous Software Development with Continual ImprovementEM^2Mem: Event-Centric Multimodal Memory for Large Language ModelsVeriPhy: Agentic Physical Reasoning for World Model Evaluation and RefinementSyncWorld: Visual Calibration Enables World Models as Zero-Shot SimulatorsTowards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and RecipesJIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness EvolutionUnlocking the Potential of Image Editing via Concept Scaling and Dense SupervisionVideo-DeepResearch: Towards the Next-Generation Multimodal Deepresearch AgentBeyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and DevelopmentEmbodied-Navigator: Point, Think, Memorize, and Align for Efficient NavigationNormalized Low-Rank AdaptationEmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA AgentsUEmbed: Unified Sparse and Dense Multimodal EmbeddingsIntern-S2-Preview: Scientific Agentic Foundation ModelTraining Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical ReportSPADE: Self-Play in Adaptive Synthetic Executable EnvironmentsInfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View SynthesisMOSS-VL Technical ReportA Common Measure of Communication for Speech Brain-Computer InterfacesRISE: Recursive Improvement via Self-Extrapolating Policy DistillationVerify Before You Distill: Prompt-Level Teacher Gating for On-Policy DistillationHow Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer ReviewKnowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming RecommendationSFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMsAgent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher MemoryUI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context DemonstrationsWithEveryone: Unified Planning and Identity Grounding for Group Image GenerationGraph Engineering in the Era of LLM Agents: From Individual Intelligence to System IntelligenceGameWAM: A World Action Model for Video GamesOn the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training StabilityPCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement LearningThe Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring MisleadsBeyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent LearningMotif 3: Technical ReportMobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World TasksSecOPD: Mitigating Adaptive Prompt Injections by On-Policy DistillationGST-Bench: Can VLMs Develop Global Spatial Awareness from Video?PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon ObjectivesClawGym II: Exploring Black-Box RL on Agent HarnessLet's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-ExpertsPaperGym: Rubric-Centered Evolution for Research-Plan GenerationABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPUEnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement LearningAutoDesign: Meta-Harness Optimization for Long-Horizon Agentic DesignInfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter