← Search

Sungjin Ahn

41 accepted papers

2026

Extendable Planning via Multiscale Diffusion

AAAI 2026technical

Long-horizon planning is crucial in complex environments, but diffusion-based planners like Diffuser are limited by the trajectory lengths observed during training. This creates a dilemma: long trajectories are needed for effective planning, yet they degrade model performance. In this paper, we intr

Cited by 0SourcePDFScholar
2026

Latent Veracity Inference for Identifying Errors in Stepwise Reasoning

ICLR 2026poster

Chain-of-Thought (CoT) reasoning has advanced the capabilities and transparency of language models (LMs); however, reasoning chains can contain inaccurate statements that reduce performance and trustworthiness. To address this, we propose to augment each reasoning step in a CoT with a latent veracit…

Cited by 0SourceScholar
2026

Loopholing Discrete Diffusion: Deterministic Bypass of the Sampling Wall

ICLR 2026poster

Discrete diffusion models offer a promising alternative to autoregressive generation through parallel decoding, but they suffer from a sampling wall: once categorical sampling occurs, rich distributional information collapses into one-hot vectors and cannot be propagated across steps. We introduce L…

Cited by 0SourcecodeScholar
2026

Understanding LoRA as Knowledge Memory: An Empirical Analysis

ICML 2026poster

Continuous knowledge updating for pre-trained large language models (LLMs) is increasingly necessary yet remains challenging. Although inference-time methods like In-Context Learning (ICL) and Retrieval-Augmented Generation (RAG) are popular, they face constraints in context budgets, costs, and retr…

Cited by 0SourceScholar
2025

Adaptive Inference-Time Scaling via Cyclic Diffusion Search

NeurIPS 2025poster

Diffusion models have demonstrated strong generative capabilities across domains ranging from image synthesis to complex reasoning tasks. However, most inference-time scaling methods rely on fixed denoising schedules, limiting their ability to allocate computation based on instance difficulty or tas…

Cited by 0SourceScholar
2025

Dreamweaver: Learning Compositional World Models from Pixels

ICLR 2025poster

Humans have an innate ability to decompose their perceptions of the world into objects and their attributes, such as colors, shapes, and movement patterns. This cognitive process enables us to imagine novel futures by recombining familiar concepts. However, replicating this ability in artificial int…

2025

Fast Monte Carlo Tree Diffusion: 100× Speedup via Parallel and Sparse Planning

NeurIPS 2025spotlight

Diffusion models have recently emerged as a powerful approach for trajectory planning. However, their inherently non-sequential nature limits their effectiveness in long-horizon reasoning tasks at test time. The recently proposed Monte Carlo Tree Diffusion (MCTD) offers a promising solution by combi…

Cited by 0SourceScholar
2025

Monte Carlo Tree Diffusion for System 2 Planning

ICML 2025spotlight

Diffusion models have recently emerged as a powerful tool for planning. However, unlike Monte Carlo Tree Search (MCTS)—whose performance naturally improves with inference-time computation scaling—standard diffusion‐based planners offer only limited avenues for the scalability. In this paper, we intr…

Cited by 3SourcePDFScholar
2025

MrSteve: Instruction-Following Agents in Minecraft with What-Where-When Memory

ICLR 2025poster

Significant advances have been made in developing general-purpose embodied AI in environments like Minecraft through the adoption of LLM-augmented hierarchical approaches. While these approaches, which combine high-level planners with low-level controllers, show promise, low-level controllers freque…

Cited by 2SourcePDFScholar
2024

Dr. Strategy: Model-Based Generalist Agents with Strategic Dreaming

ICML 2024poster

Model-based reinforcement learning (MBRL) has been a primary approach to ameliorating the sample efficiency issue as well as to make a generalist agent. However, there has not been much effort toward enhancing the strategy of dreaming itself. Therefore, it is a question *whether and how an agent can…

Cited by 6SourcePDFScholar
2024

Learning to Compose: Improving Object Centric Learning by Injecting Compositionality

ICLR 2024poster

Learning compositional representation is a key aspect of object-centric learning as it enables flexible systematic generalization and supports complex visual reasoning. However, most of the existing approaches rely on auto-encoding objective, while the compositionality is implicitly imposed by the a…

2024

Parallelized Spatiotemporal Slot Binding for Videos

ICML 2024poster

While modern best practices advocate for scalable architectures that support long-range interactions, object-centric models are yet to fully embrace these architectures. In particular, existing object-centric models for handling sequential inputs, due to their reliance on RNN-based implementation, s…

Cited by 0SourcePDFScholar
2024

PlanDQ: Hierarchical Plan Orchestration via D-Conductor and Q-Performer

ICML 2024poster

Despite the recent advancements in offline RL, no unified algorithm could achieve superior performance across a broad range of tasks. Offline *value function learning*, in particular, struggles with sparse-reward, long-horizon tasks due to the difficulty of solving credit assignment and extrapolatio…

2023

An Investigation into Pre-Training Object-Centric Representations for Reinforcement Learning

ICML 2023poster

Unsupervised object-centric representation (OCR) learning has recently drawn attention as a new paradigm of visual representation. This is because of its *potential* of being an effective pre-training technique for various downstream tasks in terms of sample efficiency, systematic generalization, an…

Cited by 42SourcePDFScholar
2023

Imagine the Unseen World: A Benchmark for Systematic Generalization in Visual World Models

NeurIPS 2023poster

Systematic compositionality, or the ability to adapt to novel situations by creating a mental model of the world using reusable pieces of knowledge, remains a significant challenge in machine learning. While there has been considerable progress in the language domain, efforts towards systematic visu…

Cited by 3SourcePDFScholar
2022

DreamerPro: Reconstruction-Free Model-Based Reinforcement Learning with Prototypical Representations

ICML 2022spotlight

Reconstruction-based Model-Based Reinforcement Learning (MBRL) agents, such as Dreamer, often fail to discard task-irrelevant visual distractions that are prevalent in natural scenes. In this paper, we propose a reconstruction-free MBRL agent, called DreamerPro, that can enhance robustness to distra…

2022

Simple Unsupervised Object-Centric Learning for Complex and Naturalistic Videos

NeurIPS 2022accept

Unsupervised object-centric learning aims to represent the modular, compositional, and causal structure of a scene as a set of object representations and thereby promises to resolve many critical limitations of traditional single-vector representations such as poor systematic generalization. Althoug…

Cited by 130SourcePDFScholar
2021

Structured World Belief for Reinforcement Learning in POMDP

ICML 2021spotlight

Object-centric world models provide structured representation of the scene and can be an important backbone in reinforcement learning and planning. However, existing approaches suffer in partially-observable environments due to the lack of belief states. In this paper, we propose Structured World Be…

Cited by 37SourcePDFScholar
2020

Improving Generative Imagination in Object-Centric World Models

ICML 2020poster

The remarkable recent advances in object-centric generative world models raise a few questions. First, while many of the recent achievements are indispensable for making a general and versatile world model, it is quite unclear how these ingredients can be integrated into a unified framework. Second,…

Cited by 88SourcePDFScholar
2020

SCALOR: Generative World Models with Scalable Object Representations

ICLR 2020poster

Scalability in terms of object density in a scene is a primary challenge in unsupervised sequential object-oriented representation learning. Most of the previous models have been shown to work only on scenes with a few objects. In this paper, we propose SCALOR, a probabilistic generative world model…

Cited by 140SourceScholar
2020

SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition

ICLR 2020poster

The ability to decompose complex multi-object scenes into meaningful abstractions like objects is fundamental to achieve higher-level cognition. Previous approaches for unsupervised object-oriented scene representation learning are either based on spatial-attention or scene-mixture approaches and li…

Cited by 269SourceScholar
2019

Neural Multisensory Scene Inference

NeurIPS 2019poster

For embodied agents to infer representations of the underlying 3D physical world they inhabit, they should efficiently combine multisensory cues from numerous trials, e.g., by looking at and touching objects. Despite its importance, multisensory 3D scene representation learning has received less att…

2018

Bayesian Model-Agnostic Meta-Learning

NeurIPS 2018spotlight

Due to the inherent model uncertainty, learning to infer Bayesian posterior from a few-shot dataset is an important step towards robust meta-learning. In this paper, we propose a novel Bayesian model-agnostic meta-learning method. The proposed method combines efficient gradient-based meta-learning w…

Cited by 536SourcePDFScholar