ICLR 2026 Accepted Papers
The full list of 5,356 papers accepted at ICLR 2026 (International Conference on Learning Representations). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
Poster: 5,118Oral: 223ICLR 2026 ConditionalPoster: 14ICLR 2026 ConditionalOral: 1
- UniCon: Unified Framework for Efficient Contrastive Alignment via KernelsPoster
- UniEdit-Flow: Unleashing Inversion and Editing in the Era of Flow ModelsPoster
- UniF$^2$ace: A $\underline{Uni}$fied $\underline{F}$ine-grained $\underline{Face}$ Understanding and Generation ModelPoster
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and GenerationPoster
- UniHM: Unified Dexterous Hand Manipulation with Vision Language ModelPoster
- UniHand: A Unified Model for Diverse Controlled 4D Hand Motion ModelingPoster
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and EditingPoster
- UniOD: A Universal Model for Outlier Detection across Diverse DomainsPoster
- UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMsPoster
- UniRestorer: Universal Image Restoration via Adaptively Estimating Image Degradation at Proper GranularityPoster
- UniSS: Unified Expressive Speech-to-Speech Translation with Your VoicePoster
- UniSplat: Unified Spatio-Temporal Fusion via 3D Latent Scaffolds for Dynamic Driving Scene ReconstructionPoster
- UniTrack: Differentiable Graph Representation Learning for Multi-Object TrackingPoster
- UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic EncodingPoster
- Unified 3D Scene Understanding Through Physical World ModelingPoster
- Unified Analyses for Hierarchical Federated Learning: Topology Selection under Data HeterogeneityPoster
- Unified Biomolecular Trajectory Generation via Pretrained Variational BridgePoster
- Unified Brain Surface and Volume RegistrationPoster
- Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Diffusion Diffusion ProcessPoster
- Unified In-Context Video EditingPoster
- Unified Multi-Modal Interactive and Reactive 3D Motion Generation via Rectified FlowPoster
- Unified Privacy Guarantees for Decentralized Learning via Matrix FactorizationPoster
- Unified Vision-Language-Action ModelPoster
- Unified Vision–Language Modeling via Concept Space AlignmentPoster
- Unified and Efficient Multi-view Clustering from Probabilistic PerspectivePoster
- Uniform Discrete Diffusion with Metric Path for Video GenerationPoster
- Unifying Diffusion and Autoregression for Generalizable Vision-Language-Action ModelPoster
- Unifying Formal Explanations: A Complexity-Theoretic PerspectivePoster
- Unifying Stable Optimization and Reference Regularization in RLHFPoster
- Universal Beta SplattingPoster
- Universal Inverse Distillation for Matching Models with Real-Data Supervision (No GANs)Oral
- Universal Model Routing for Efficient LLM InferencePoster
- Universal Multi-Domain Translation via Diffusion RoutersPoster
- Universal Properties of Activation Sparsity in Modern Large Language ModelsPoster
- Universal Value-Function UncertaintiesPoster
- Unlearning Evaluation through Subset Statistical IndependencePoster
- Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model OutputsPoster
- Unlearning during Training: Domain-Specific Gradient Ascent for Domain GeneralizationPoster
- Unleashing Guidance Without Classifiers for Human-Object Interaction AnimationPoster
- Unleashing LLMs in Bayesian Optimization: Preference-Guided Framework for Scientific DiscoveryPoster
- Unleashing Perception-Time Scaling to Multimodal Reasoning ModelsPoster
- Unleashing Scientific Reasoning for Bio-experimental Protocol Generation via Structured Component-based Reward MechanismPoster
- Unlocking Full Efficiency of Token Filtering in Large Language Model TrainingPoster
- Unlocking Long-Horizon Agentic Search with Large-Scale End-to-End RLPoster
- Unlocking the Essence of Beauty: Advanced Aesthetic Reasoning with Relative-Absolute Policy OptimizationPoster
- Unlocking the Potential of Weighting Methods in Federated Learning Through Communication CompressionPoster
- Unlocking the Power of Co-Occurrence in CLIP: A DualPrompt-Driven Method for Training-Free Zero-Shot Multi-Label ClassificationPoster
- Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to DeliberationPoster
- Unlocking the Value of Text: Event-Driven Reasoning and Multi-Level Alignment for Time Series ForecastingPoster
- Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language ModelsPoster
- Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio ClassificationPoster
- Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE FrameworkPoster
- Unraveling the Complexity of Memory in RL Agents: an Approach for Classification and EvaluationPoster
- Unsupervised Learning of Efficient Exploration: Pre-training Adaptive Policies via Self-Imposed GoalsPoster
- Unsupervised Representation Learning - an Invariant Risk Minimization PerspectivePoster
- Unsupervised Representation Learning for 3D Mesh Parameterization with Semantic and Visibility ObjectivesPoster
- Untraceable DeepFakes via Traceable Fingerprint EliminationPoster
- Unveiling Downstream Performance Scaling of LLMs: A Clustering-Based PerspectivePoster
- Unveiling Perceptual Artifacts: A Fine-Grained Benchmark for Interpretable AI-Generated Image DetectionPoster
- Unveiling Super Experts in Mixture-of-Experts Large Language ModelsPoster
- Unveiling the Cognitive Compass: Theory-of-Mind–Guided Multimodal Emotion ReasoningPoster
- Unveiling the Mechanism of Continuous Representation Full-Waveform Inversion: A Wave Based Neural Tangent Kernel FrameworkPoster
- Unveiling the Potential of Diffusion Large Language Model in Controllable GenerationPoster
- Urban Socio-Semantic Segmentation with Vision-Language ReasoningICLR 2026 ConditionalPoster
- UrbanFeel:A Comprehensive Benchmark for Temporal and Perceptual Understanding of City Scenes through Human PerspectivePoster
- UrbanGS: Efficient and Scalable Architecture for Geometrically Accurate Large-Scene ReconstructionPoster
- UrbanGraph: Physics-Informed Spatio-Temporal Dynamic Heterogeneous Graphs for Urban Microclimate PredictionPoster
- UrbanVerse: Scaling Urban Simulation by Watching City-Tour VideosPoster
- Use the Online Network If You Can: Towards Fast and Stable Reinforcement LearningPoster
- Using Reinforcement Learning to Train Large Language Models to Explain Human DecisionsPoster
- Using cognitive models to reveal value trade-offs in language modelsPoster
- Using maximal information auxiliary variables to improve synthetic data generation based on TabPFN foundation modelsPoster
- V2P-Bench: Evaluating Video-Language Understanding with Visual Prompts for Better Human-Model InteractionPoster
- VADv2: End-to-End Autonomous Driving via Probabilistic PlanningPoster
- VARestorer: One-Step VAR Distillation for Real-World Image Super-ResolutionPoster
- VCWorld: A Biological World Model for Virtual Cell SimulationPoster
- VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language ModelsPoster
- VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic RoutingPoster
- VERIFY: A Novel Multi-Domain Dataset Grounding LTL in Contextual Natural Language via Provable Intermediate LogicPoster
- VERINA: Benchmarking Verifiable Code GenerationPoster
- VFScale: Intrinsic Reasoning through Verifier-Free Test-time Scalable Diffusion ModelPoster
- VGR: Visual Grounded ReasoningPoster
- VINCIE: Unlocking In-context Image Editing from VideoPoster
- VIRTUE: Visual-Interactive Text-Image Universal EmbedderPoster
- VITA: Vision-to-Action Flow Matching PolicyPoster
- VITA: Zero-Shot Value Functions via Test-Time Adaptation of Vision–Language ModelsPoster
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-languagePoster
- VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic ManipulationPoster
- VLM-Guided Adaptive Negative Prompting for Creative GenerationPoster
- VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?Poster
- VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action ModelsPoster
- VLMgineer: Vision-Language Models as Robotic ToolsmithsPoster
- VLSU: Mapping the Limits of Joint Multimodal Understanding for AI SafetyPoster
- VMDiff: Visual Mixing Diffusion for Limitless Cross-Object SynthesisPoster
- VMoBA: Mixture-of-Block Attention for Video Diffusion ModelsPoster
- VOGUE: Unified Understanding, Generation, and Editing for VideosPoster
- VPI-Bench: Visual Prompt Injection Attacks for Computer-Use AgentsPoster
- VQ-Transplant: Efficient VQ-Module Integration for Pre-trained Visual TokenizersPoster
- VSF: Simple, Efficient, and Effective Negative Guidance in Few-Step Image Generation Models By Value Sign FlipPoster
- VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool UsePoster
- VUDG: A Dataset for Video Understanding Domain GeneralizationPoster
- Value FlowsPoster
- Value Gradient Flow: Behavior-Regularized RL without RegularizationPoster
- Value Matching: Scalable and Gradient-Free Reward-Guided Flow AdaptationPoster
- Variance-Dependent Regret Lower Bounds for Contextual BanditsPoster
- Variation in Verification: Understanding Verification Dynamics in Large Language ModelsPoster
- Variation-aware Flexible 3D Gaussian EditingPoster
- Variational Autoencoding Discrete Diffusion with Enhanced Dimensional Correlations ModelingPoster
- Variational Deep Learning via Implicit RegularizationPoster
- Variational Inference for Cyclic LearningPoster
- Variational Reasoning for Language ModelsPoster
- VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek PotteryPoster
- VenusX: Unlocking Fine-Grained Functional Understanding of ProteinsPoster
- VeriCoT: Neuro-symbolic Chain-of-Thought Validation via Logical Consistency ChecksPoster
- VeriEquivBench: An Equivalence Score for Ground-Truth-Free Evaluation of Formally Verifiable CodePoster
- VeriRole: Verifiable Role-Awareness through Hint-Guided Reinforcement LearningPoster
- VeriTrail: Closed-Domain Hallucination Detection with TraceabilityPoster
- Verification and Co-Alignment via Heterogeneous Consistency for Preference-Aligned LLM AnnotationsPoster
- Verification of the Implicit World Model in a Generative Model via Adversarial SequencesPoster
- Verifier-free Test-Time Sampling for Vision Language Action ModelsPoster
- VerifyBench: Benchmarking Reference-based Reward Systems for Large Language ModelsPoster
- Verifying Chain-of-Thought Reasoning via Its Computational GraphOral
- Veritas: Generalizable Deepfake Detection via Pattern-Aware ReasoningOral
- ViMo: A Generative Visual GUI World Model for App AgentsPoster
- ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language ModelsPoster
- ViPO: Visual Preference Optimization at ScalePoster
- ViPRA: Video Prediction for Robot ActionsPoster
- ViTSP: A Vision Language Models Guided Framework for Large-Scale Traveling Salesman ProblemsPoster
- VibeVoice: Expressive Podcast Generation with Next-Token DiffusionOral
- Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction–Reasoning SynergyOral
- Vid2World: Crafting Video Diffusion Models to Interactive World ModelsPoster
- VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy TasksPoster
- VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RLPoster
- Video Scene Segmentation with Genre and Duration SignalsPoster
- Video Unlearning via Low-Rank Refusal VectorPoster
- Video-As-Prompt: Unified Semantic Control for Video GenerationPoster
- Video-GPT via Next Clip DiffusionPoster
- Video-KTR: Reinforcing Video Reasoning via Key Token AttributionPoster
- Video-LevelGauge: Investigating Contextual Positional Bias in Video Language Models.Poster
- Video-STAR: Reinforcing Open-Vocabulary Action Recognition with ToolsPoster
- VideoAgentTrek: Computer-Use Pretraining from Unlabeled VideosPoster
- VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial ReasoningPoster
- VideoChat-Flash: Hierarchical Compression for Long-Context Video ModelingPoster
- VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video UnderstandingPoster
- VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in VideoPoster
- VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video ReasoningPoster
- VideoNSA: Native Sparse Attention Scales Video UnderstandingPoster
- VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video GenerationPoster
- VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?Poster
- VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video ReasoningPoster
- Virne: A Comprehensive Benchmark for RL-based Network Resource Allocation in NFVPoster
- Virtual Community: An Open World for Humans, Robots, and SocietyPoster
- VisCoder2: Building Multi-Language Visualization Coding AgentsPoster
- VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding ModelsPoster
- VisJudge-Bench: Aesthetics and Quality Assessment of VisualizationsPoster
- VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMsPoster
- Vision Language Models are BiasedPoster
- Vision-Language-Action Instruction Tuning: From Understanding to ManipulationPoster
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsPoster
- Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-PlayPoster
- VisionLaw: Inferring Interpretable Intrinsic Dynamics from Visual Observations via Bilevel OptimizationPoster
- VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement LearningPoster
- VisionTrim: Unified Vision Token Compression for Training-Free MLLM AccelerationPoster
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsPoster
- VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual ReasoningPoster
- Visual Autoregressive Modeling for Instruction-Guided Image EditingPoster
- Visual Jigsaw Post-Training Improves MLLMsPoster
- Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual FlowPoster
- Visual Planning: Let's Think Only with ImagesOral
- Visual Prompt-Agnostic EvolutionPoster
- Visual Self-Refine: A Pixel-Guided Paradigm for Accurate Chart ParsingPoster
- Visual symbolic mechanisms: Emergent symbol processing in Vision Language ModelsOral
- VisualPRM400K: An Effective Dataset for Training Multimodal Process Reward ModelsPoster
- VisualPrompter: Semantic-Aware Prompt Optimization with Visual Feedback for Text-to-Image SynthesisPoster
- VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world ApplicationsPoster
- Vivid-VR: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video RestorationPoster
- Vlaser: Vision-Language-Action Model with Synergistic Embodied ReasoningPoster
- VoG: Enhancing LLM Reasoning through Stepwise Verification on Knowledge GraphsPoster
- VoMP: Predicting Volumetric Mechanical Property FieldsPoster
- VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic AugmentationPoster
- VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language ModelsPoster
- Vulcan: Crafting Compact Class-Specific Vision Transformers For Edge IntelligencePoster
- W-EDIT: A Wavelet-Based Frequency-Aware Framework for Text-Driven Image EditingPoster
- WAFT: Warping-Alone Field Transforms for Optical FlowOral
- WALT: Web Agents that Learn ToolsPoster
- WARC-Bench: Web Archive based Benchmark for GUI Subtask ExecutionsPoster
- WARP: Weight Teleportation for Attack-Resilient Unlearning ProtocolsPoster
- WATS: Wavelet-Aware Temperature Scaling for Reliable Graph Neural NetworksPoster
- WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLMOral
- WFR-FM: Simulation-Free Dynamic Unbalanced Optimal TransportPoster
- WILD-Diffusion: A WDRO Inspired Training Method for Diffusion Models under Limited DataPoster
- WIMFRIS: WIndow Mamba Fusion and Parameter Efficient Tuning for Referring Image SegmentationPoster
- WIMLE: Uncertainty‑Aware World Models with IMLE for Sample‑Efficient Continuous ControlPoster
- WINA: Weight Informed Neuron Activation for Accelerating Large Language Model InferencePoster
- WMPO: World Model-based Policy Optimization for Vision-Language-Action ModelsPoster
- WOW-Seg: A Word-free Open World Segmentation ModelPoster
- WRING Out The Bias: A Rotation-Based Alternative To Projection DebiasingPoster
- WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-trainingOral
- WSVD: Weighted Low-Rank Approximation for Fast and Efficient Execution of Low-Precision Vision-Language ModelsPoster
- Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMsPoster
- Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM FinetuningICLR 2026 ConditionalOral
- WaterDrum: Watermark-based Data-centric Unlearning MetricPoster
- Watermark-based Attribution of AI-Generated ContentPoster
- Watermarking Diffusion Language ModelsPoster
- WavePolyp: Video Polyp Segmentation via Hierarchical Wavelet-Based Feature Aggregation and Inter-Frame Divergence PerceptionPoster
- WavefrontDiffusion: Dynamic Decoding Schedule for Improved ReasoningPoster
- Wavelet Predictive Representations for Non-Stationary Reinforcement LearningPoster
- We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical ReasoningPoster
- WeTok: Powerful Discrete Tokenization for High-Fidelity Visual ReconstructionPoster
- Weak Correlations as the Underlying Principle for Linearization of Gradient-Based Learning SystemsPoster
- Weak-to-Strong DiffusionPoster
- Weak-to-Strong Generalization with Failure TrajectoriesPoster
- WearVox: An Egocentric Multichannel Voice Assistant Benchmark for WearablesPoster
- Web-CogReasoner: Towards Knowledge-Induced Cognitive Reasoning for Web AgentsPoster
- WebArbiter: A Generative Reasoning Process Reward Model for Web AgentsPoster
- WebDS: An End-to-End Benchmark for Web-based Data SciencePoster
- WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development QualityOral
- WebFactory: Automated Compression of Foundational Language Intelligence into Grounded Web AgentsPoster
- WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement LearningPoster
- WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement LearningPoster
- WebSeer: Training Deeper Search Agents through Reinforcement Learning with Self-ReflectionPoster
- WebShaper: Agentically Data Synthesizing via Information-Seeking FormalizationPoster
- WebWatcher: Breaking New Frontiers of Vision-Language Deep Research AgentPoster
- WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep ResearchPoster
- Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining LevelsPoster
- Weight Decay may matter more than µP for Learning Rate Transfer in PracticePoster
- Weight Space Representation Learning on Diverse NeRF ArchitecturesPoster
- Weight-Space Linear Recurrent Neural NetworksPoster
- Welfarist Formulations for Diverse Similarity SearchPoster
- What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token MergingPoster
- What Do Large Language Models Know About Opinions?Poster
- What Exactly Does Guidance Do in Masked Discrete Diffusion ModelsPoster
- What Generative Search Engines Like and How to Optimize Web Content CooperativelyPoster
- What Happens Next? Anticipating Future Motion by Generating Point TrajectoriesPoster
- What Layers When: Learning to Skip Compute in LLMs with Residual GatesPoster
- What Matters for Batch Online Reinforcement Learning in Robotics?Poster
- What Scales in Cross-Entropy Scaling Law?Poster
- What happens when generative AI models train recursively on each others' outputs?Poster
- What matters for Representation Alignment: Global Information or Spatial Structure?Poster
- What's In My Human Feedback? Learning Interpretable Descriptions of Preference DataOral
- What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generationPoster
- Whatever Remains Must Be True: Filtering Drives Reasoning in LLMs, Shaping DiversityPoster
- When Agents “Misremember” Collectively: Exploring the Mandela Effect in LLM-based Multi-Agent SystemsPoster
- When Bias Helps Learning: Bridging Initial Prejudice and TrainabilityPoster
- When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization DatasetsPoster
- When Does Divide and Conquer Work for Long Context LLM? A Noise Decomposition FrameworkPoster
- When Flatness Does (Not) Guarantee Adversarial RobustnessPoster
- When Foundation Models are One-Liners: Limitations and Future Directions for Time Series Anomaly DetectionPoster
- When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM TrainingPoster
- When Is Diversity Rewarded in Cooperative Multi-Agent Learning?Poster
- When LLMs get significantly worse: A statistical approach to detect model degradationsPoster
- When Language Models Lose Their Mind: The Consequences of Brain MisalignmentPoster
- When Large Multimodal Models Confront Evolving Knowledge: Challenges and ExplorationsPoster
- When MLLMs Meets Compression Distortion: A Coding Paradigm Tailored to MLLMsPoster
- When Machine Learning Gets Personal: Evaluating Prediction and ExplanationPoster
- When More is Less: Understanding Chain-of-Thought Length in LLMsPoster
- When Priors Backfire: On the Vulnerability of Unlearnable Examples to PretrainingPoster
- When Reasoning Meets Compression: Understanding the Effects of LLMs Compression on Large Reasoning ModelsPoster
- When Scores Learn Geometry: Rate Separations under the Manifold HypothesisPoster
- When Shift Happens - Confounding Is to BlamePoster
- When Silence Is Golden: Can LLMs Learn to Abstain in Temporal QA and Beyond?Poster
- When Style Breaks Safety: Defending LLMs Against Superficial Style AlignmentPoster
- When Thinking Backfires: Mechanistic Insights into Reason-induced MisalignmentPoster
- When Weak LLMs Speak with Confidence, Preference Alignment Gets StrongerPoster
- When a Robot is More Capable than a Human: Learning from Constrained DemonstratorsPoster
- When and Where to Reset Matters for Long-Term Test-Time AdaptationPoster
- When to Ensemble: Identifying Token-Level Points for Stable and Fast LLM EnsemblingPoster
- When to Retrain after Drift: A Data-Only Test of Post-Drift Data Size SufficiencyPoster
- When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented GenerationPoster
- When would Vision-Proprioception Policies Fail in Robotic Manipulation?Poster
- Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient TracingPoster
- Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning DistillationPoster
- Who Matters Matters: Agent-Specific Conservative Offline MARLPoster
- WholeBodyVLA: Towards Unified Latent VLA for Whole-body Loco-manipulation ControlPoster
- Why Adversarially Train Diffusion Models?Poster
- Why Ask One When You Can Ask $k$? Learning-to-Defer to the Top-$k$ ExpertsPoster
- Why Attention Patterns Exist: A Unifying Temporal Perspective AnalysisPoster
- Why DPO is a Misspecified Estimator and How to Fix ItOral
- Why Do Unlearnable Examples Work: A Novel Perspective of Mutual InformationPoster
- Why High-rank Neural Networks Generalize?: An Algebraic Framework with RKHSsPoster
- Why Keep Your Doubts to Yourself? Trading Visual Uncertainties in Multi-Agent Bandit SystemsPoster
- Why Less is More (Sometimes): A Theory of Data CurationPoster
- Why Low-Precision Transformer Training Fails: An Analysis on Flash AttentionOral
- Why Prototypes Collapse: Diagnosing and Preventing Partial Collapse in Prototypical Self-Supervised LearningPoster
- Why Reinforcement Fine-Tuning Enables MLLMs Preserve Prior Knowledge Better: A Data PerspectivePoster
- Why We Need New Benchmarks for Local Intrinsic Dimension EstimationPoster
- Why is Your Language Model a Poor Implicit Reward Model?Poster
- Wide-In, Narrow-Out: Revokable Decoding for Efficient and Effective DLLMsPoster
- WideSearch: Benchmarking Agentic Broad Info-SeekingPoster
- Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling CurriculumPoster
- Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmasPoster
- WinT3R: Window-Based Streaming Reconstruction with Camera Token PoolPoster
- Winter Soldier: Backdooring Language Models at Pre-Training with Indirect Data PoisoningPoster
- WithAnyone: Toward Controllable and ID Consistent Image GenerationICLR 2026 ConditionalPoster
- WoW!: World Models in a Closed-Loop WorldOral
- World2Minecraft: Occupancy-Driven simulated scenes ConstructionPoster
- WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed BenchmarkPoster
- WorldGym: World Model as An Environment for Policy EvaluationPoster
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMsPoster
- WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous DrivingPoster
- WorldTree: Towards 4D Dynamic Worlds from Monocular Video using Tree-ChainsPoster
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action ModelPoster
- XIL: Cross-Expanding Incremental LearningPoster
- XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language ModelsPoster
- XQC: Well-conditioned Optimization Accelerates Deep Reinforcement LearningPoster
- YoNoSplat: You Only Need One Model for Feedforward 3D Gaussian SplattingPoster
- You Point, I Learn: Online Adaptation of Interactive Segmentation Models for Handling Distribution Shifts in Medical ImagingPoster
- Your Agent May Misevolve: Emergent Risks in Self-evolving LLM AgentsPoster
- Your Language Model Secretly Contains Personality SubnetworksPoster
- Your Models Have Thought Enough: Training Large Reasoning Models to Stop OverthinkingPoster
- Your VAR Model is Secretly an Efficient and Explainable Generative ClassifierPoster
- Youtu-GraphRAG: Vertically Unified Agents for Graph Retrieval-Augmented Complex ReasoningPoster
- YuE: Scaling Open Foundation Models for Long-Form Music GenerationPoster
- ZIP-RC: Zero-overhead Inference-time Prediction of Reward and Cost for Adaptive and Interpretable GenerationPoster
- Zebra-CoT: A Dataset for Interleaved Vision-Language ReasoningPoster
- Zephyrus: An Agentic Framework for Weather SciencePoster
- Zero-Sacrifice Persistent-Robustness Adversarial Defense for Pre-Trained EncodersPoster
- Zero-Shot Adaptation of Behavioral Foundation Models to Unseen DynamicsPoster
- Zero-shot Forecasting by Simulation AlonePoster
- Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction RecognitionPoster
- Zero-shot Human Pose Estimation using Diffusion-based Inverse solversPoster
- ZeroGR: A Generalizable and Scalable Framework for Zero-Shot Generative RetrievalPoster
- ZeroSiam: An Efficient Siamese for Test-Time Entropy Optimization without CollapsePoster
- ZeroTuning: Unlocking the Initial Token's Power to Enhance Large Language Models Without TrainingPoster
- Zeros can be Informative: Masked Binary U-Net for Image Segmentation on Tensor CoresPoster
- ``Noisier'’ Noise Contrastive Estimation is (Almost) Maximum LikelihoodPoster
- cadrille: Multi-modal CAD Reconstruction with Reinforcement LearningOral
- d$^2$Cache: Accelerating Diffusion-Based LLMs via Dual Adaptive CachingPoster
- dParallel: Learnable Parallel Decoding for dLLMsPoster
- e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMsPoster
- f-INE: A Hypothesis Testing Framework for Estimating Influence under Training RandomnessPoster
- floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RLPoster
- gLSTM: Mitigating Over-Squashing by Increasing Storage CapacityPoster
- gen2seg: Generative Models Enable Generalizable Instance SegmentationPoster
- h-MINT: Modeling Pocket-Ligand Binding with Hierarchical Molecular Interaction NetworkPoster
- iFusion: Integrating Dynamic Interest Streams via Diffusion Model for Click-Through Rate PredictionPoster
- iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal ModelsPoster
- jqBench: a benchmark for reading and editing JSON from natural language and/or examplesPoster
- lmgame-Bench: How Good are LLMs at Playing Games?Poster
- mCLM: A Modular Chemical Language Model that Generates Functional and Makeable MoleculesOral
- mR3: Multilingual Rubric-Agnostic Reward Reasoning ModelsPoster
- pFedMMA: Personalized Federated Fine-Tuning with Multi-Modal Adapter for Vision-Language ModelsPoster
- pi-Flow: Policy-Based Few-Step Generation via Imitation DistillationPoster
- pySpatial: Generating 3D Visual Programs for Zero-Shot Spatial ReasoningPoster
- reAR: Rethinking Visual Autoregressive Models via Token-wise Consistency RegularizationPoster
- scDFM: Distributional Flow Matching Model for Robust Single-Cell Perturbation PredictionPoster
- sleep2vec: Unified Cross-Modal Alignment for Heterogeneous Nocturnal BiosignalsPoster
- ssToken: Self-modulated and Semantic-aware Token Selection for LLM Fine-tuningPoster
- station2radar: query‑conditioned gaussian splatting for precipitation fieldPoster
- t-SNE Exaggerates Clusters, ProvablyPoster
- vAttention: Verified Sparse Attention via SamplingPoster
- vCache: Verified Semantic Prompt CachingPoster
- villa-X: Enhancing Latent Action Modeling in Vision-Language-Action ModelsPoster
- wd1: Weighted Policy Optimization for Reasoning in Diffusion Language ModelsPoster
- xLSTM Scaling Laws: Competitive Performance with Linear Time-ComplexityPoster
- xRFM: Accurate, scalable, and interpretable feature learning models for tabular dataPoster
ICLR accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.