← Search

Byoung-Tak Zhang

57 accepted papers

2026

Climb With SHERPA: Heuristic-Guided Reinforcement Learning via Segmented Experience Relay

RA-L 2026

In sparse-reward, long-horizon domains, reinforcement learning (RL) often suffers from slow convergence and instability, complicating robotic manipulation. Previous heuristic-guided approaches have relied on step-level actions and imitation loss, but struggle to maintain temporal coherence or solve

Cited by 0SourceScholar
2026

Kinematics-Driven Gaussian Shape Deformation for Blurry Monocular Dynamic Scenes

ICML 2026poster

Reconstructing dynamic 3D scenes from blurry monocular videos is challenging because motion-induced blur entangles object motion and geometry, hindering geometric consistency. We present Kinematics-GS, a kinematics-aware framework that models blur as motion-aligned deformation and introduces a kinem…

Cited by 0SourceScholar
2026

Learning Coordinate-based Convolutional Kernels for Continuous SE(3) Equivariant and Efficient Point Cloud Analysis

CVPR 2026

A symmetry on rigid motion is one of the salient factors in efficient learning of 3D point cloud problems. Group convolution has been a representative method to extract equivariant features, but its realizations have struggled to retain both rigorous symmetry and scalability simultaneously. We advoc

Cited by 0SourceScholar
2026

Neural Collapse-Informed Initialization with Perturbation Injection in Classification-based Metric Learning

AAAI 2026technical

Recent studies have revealed Neural Collapse (NC) in deep classifiers, where last-layer weights and features align into an equiangular tight frame (ETF), concentrating class information along specific embedding directions. However, conventional fine-tuning typically disregards this structure, initi

Cited by 0SourcePDFScholar
2026

PeriUn: Enhancing Unlearning by Selectively Forgetting Peripheral Samples

AAAI 2026technical

Once trained, neural networks memorize information in diffusely encoded parameters, making it difficult to forget in support of the right to be forgotten. Unlearning aims to remove the influence of data, with performance measured against a retrained model that excludes the data. However, understandi

Cited by 0SourcePDFScholar
2026

Towards Spatially Consistent Image Generation: On Incorporating Intrinsic Scene Properties into Diffusion Models

AAAI 2026technical

Image generation models trained on large datasets can synthesize high-quality images but often produce spatially inconsistent and distorted images due to limited information about the underlying structures and spatial layouts. In this work, we leverage intrinsic scene properties (e.g., depth, segmen

Cited by 0SourcePDFScholar
2026

Voronoi-Based Second-Order Descriptor with Whitened Metric in LiDAR Place Recognition

ICRA 2026poster

The pooling layer plays a vital role in aggregating local descriptors into the metrizable global descriptor in the LiDAR Place Recognition (LPR). In particular, the second-order pooling is capable of capturing higher-order interactions among local descriptors. However, its existing methods in the LP…

2025

CDIS : Cross-Dimensional Class-Agnostic 3D Instance Segmentation via 2D Mask Tracking and 3D-2D Projection Merging

IROS 2025

Class-agnostic 3D instance segmentation is critical for robotic systems operating in unknown environments, enabling perception of previously unseen objects for reliable manipulation and navigation. Existing approaches typically project per-frame 2D instance masks into 3D and merge them, which often

Cited by 0SourceScholar
2025

CLIP-RT: Learning Language-Conditioned Robotic Policies from Natural Language Supervision

RSS 2025poster

Teaching robots desired skills in real-world environments remains challenging, especially for non-experts. Current robot learning methods often require expert demonstrations or complex programming, limiting their accessibility to non-experts. We posit that natural language offers an intuitive and ac…

Cited by 1PDFScholar
2025

Confidence-guided Refinement Reasoning for Zero-shot Question Answering

EMNLP 2025

We propose Confidence-guided Refinement Reasoning (C2R), a novel training-free framework applicable to question-answering (QA) tasks across text, image, and video domains. C2R strategically constructs and refines sub-questions and their answers (sub-QAs), deriving a better confidence score for the t

Cited by 0SourcePDFScholar
2025

DA-Fusion: Deformable Attention-Based RGB-D Fusion Transformer for Unseen Object Instance Segmentation

ICRA 2025

In logistics automation, precise segmentation of unseen objects is crucial for efficient robotic manipulation in cluttered environments. Tasks such as bin-picking and shelfpicking require robust perception to handle occlusions, varying object shapes, and complex spatial arrangements. Traditional RGB

Cited by 2SourceScholar
2025

How Classifier Features Transfer to Downstream: An Asymptotic Analysis in a Two-Layer Model

NeurIPS 2025poster

Neural networks learn effective feature representations, which can be transferred to new tasks without additional training. While larger datasets are known to improve feature transfer, the theoretical conditions for the success of such transfer remain unclear. This work investigates feature transfer…

Cited by 0SourceScholar
2025

OCK: Unsupervised Dynamic Video Prediction with Object-Centric Kinematics

ICCV 2025poster

Human perception involves decomposing complex multi-object scenes into time-static object appearance (i.e., size, shape, color) and time-varying object motion (i.e., position, velocity, acceleration). For machines to achieve human-like intelligence in real-world interactions, understanding these phy…

Cited by 0SourcePDFScholar
2025

On the Consistency of Video Large Language Models in Temporal Comprehension

CVPR 2025poster

Video large language models (Video-LLMs) can temporally ground language queries and retrieve video moments. Yet, such temporal comprehension capabilities are neither well-studied nor understood. So we conduct a study on prediction consistency -- a key indicator for robustness and trustworthiness of…

2025

Socratic Planner: Self-QA-Based Zero-Shot Planning for Embodied Instruction Following

ICRA 2025

Embodied Instruction Following (EIF) is the task of executing natural language instructions by navigating and interacting with objects in interactive environments. A key challenge in EIF is compositional task planning, typically addressed through supervised learning or few-shot in-context learning w

Cited by 8SourceScholar
2024

DUEL: Duplicate Elimination on Active Memory for Self-Supervised Class-Imbalanced Learning

AAAI 2024technical

Recent machine learning algorithms have been developed using well-curated datasets, which often require substantial cost and resources. On the other hand, the direct use of raw data often leads to overfitting towards frequently occurring class information. To address class imbalances cost-efficientl…

Cited by 1SourcePDFScholar
2024

Efficient Monte Carlo Tree Search via On-the-Fly State-Conditioned Action Abstraction

UAI 2024poster

Monte Carlo Tree Search (MCTS) has showcased its efficacy across a broad spectrum of decision-making problems. However, its performance often degrades under vast combinatorial action space, especially where an action is composed of multiple sub-actions. In this work, we propose an action abstraction…

2024

Fine-Grained Causal Dynamics Learning with Quantization for Improving Robustness in Reinforcement Learning

ICML 2024poster

Causal dynamics learning has recently emerged as a promising approach to enhancing robustness in reinforcement learning (RL). Typically, the goal is to build a dynamics model that makes predictions based on the causal relationships among the entities. Despite the fact that causal connections often m…

2024

Multi-Object RANSAC: Efficient Plane Clustering Method in a Clutter

ICRA 2024poster

In this paper, we propose a novel method for plane clustering specialized in cluttered scenes using an RGB-D camera and validate its effectiveness through robot grasping experiments. Unlike existing methods, which focus on large- scale indoor structures, our approach—Multi-Object RANSAC emphasizes c…

Cited by 1SourceScholar
2024

PGA: Personalizing Grasping Agents with Single Human-Robot Interaction

IROS 2024poster

Language-Conditioned Robotic Grasping (LCRG) aims to develop robots that comprehend and grasp objects based on natural language instructions. While the ability to understand personal objects like my wallet facilitates more natural interaction with human users, current LCRG systems only allow generic…

Cited by 2SourcecodeScholar
2024

PROGrasp: Pragmatic Human-Robot Communication for Object Grasping

ICRA 2024poster

Interactive Object Grasping (IOG) is the task of identifying and grasping the desired object via human-robot natural language interaction. Current IOG systems assume that a human user initially specifies the target object’s category (e.g., bottle). Inspired by pragmatics, where humans often convey t…

Cited by 7SourcecodeScholar
2024

Seg2Grasp: A Robust Modular Suction Grasping in Bin Picking

IROS 2024poster

Current bin picking methods that rely heavily on end-to-end learning often falter when confronted with unfamiliar or complex objects in unstructured environments. To overcome these limitations, we introduce Seg2Grasp, a modular pipeline designed for robust suction grasping in dynamic and cluttered b…

Cited by 0SourceScholar
2024

Unveiling the Significance of Toddler-Inspired Reward Transition in Goal-Oriented Reinforcement Learning

AAAI 2024technical

Toddlers evolve from free exploration with sparse feedback to exploiting prior experiences for goal-directed learning with denser rewards. Drawing inspiration from this Toddler-Inspired Reward Transition, we set out to explore the implications of varying reward transitions when incorporated into Rei…

Cited by 3SourcePDFScholar
2023

EXOT: Exit-aware Object Tracker for Safe Robotic Manipulation of Moving Object

ICRA 2023poster

Current robotic hand manipulation narrowly operates with objects in predictable positions in limited environments. Thus, when the location of the target object deviates severely from the expected location, a robot sometimes responds in an unexpected way, especially when it operates with a human. For…

Cited by 0SourcecodeScholar
2023

GVCCI: Lifelong Learning of Visual Grounding for Language-Guided Robotic Manipulation

IROS 2023poster

Language-Guided Robotic Manipulation (LGRM) is a challenging task as it requires a robot to understand human instructions to manipulate everyday objects. Recent approaches in LGRM rely on pre-trained Visual Grounding (VG) models to detect objects without adapting to manipulation environments. This r…

Cited by 7SourcecodeScholar
2023

Learning Geometry-Aware Representations by Sketching

CVPR 2023poster

Understanding geometric concepts, such as distance and shape, is essential for understanding the real world and also for many vision tasks. To incorporate such information into a visual representation of a scene, we propose learning to represent the scene by sketching, inspired by human behavior. Ou…

Cited by 7SourcePDFScholar
2023

Neural Collage Transfer: Artistic Reconstruction via Material Manipulation

ICCV 2023poster

Collage is a creative art form that uses diverse material scraps as a base unit to compose a single image. Although pixel-wise generation techniques can reproduce a target image in collage style, it is not a suitable method due to the solid stroke-by-stroke nature of the collage form. While some p…

Cited by 3PDFcodeScholar
2023

The Dialog Must Go On: Improving Visual Dialog via Generative Self-Training

CVPR 2023poster

Visual dialog (VisDial) is a task of answering a sequence of questions grounded in an image, using the dialog history as context. Prior work has trained the dialog agents solely on VisDial data via supervised learning or leveraged pre-training on related vision-and-language datasets. This paper pres…

2022

From Scratch to Sketch: Deep Decoupled Hierarchical Reinforcement Learning for Robotic Sketching Agent

ICRA 2022poster

We present an automated learning framework for a robotic sketching agent that is capable of learning stroke-based rendering and motor control simultaneously. We formulate the robotic sketching problem as a deep decoupled hierarchical reinforcement learning; two policies for stroke-based rendering an…

Cited by 13SourceScholar
2022

Hypergraph Transformer: Weakly-Supervised Multi-hop Reasoning for Knowledge-based Visual Question Answering

ACL 2022long

Knowledge-based visual question answering (QA) aims to answer a question which requires visually-grounded external knowledge beyond image content itself. Answering complex questions that require multi-hop reasoning under weak supervision is considered as a challenging problem since i) no supervision…

2022

Modal-specific Pseudo Query Generation for Video Corpus Moment Retrieval

EMNLP 2022main

Video corpus moment retrieval (VCMR) is the task to retrieve the most relevant video moment from a large video corpus using a natural language query.For narrative videos, e.g., drama or movies, the holistic understanding of temporal dynamics and multimodal reasoning are crucial.Previous works have s…

2022

PlaceNet: Neural Spatial Representation Learning with Multimodal Attention

IJCAI 2022poster

Spatial representation capable of learning a myriad of environmental features is a significant challenge for natural spatial understanding of mobile AI agents. Deep generative models have the potential of discovering rich representations of observed 3D scenes. However, previous approaches have bee…

2022

Robust Imitation via Mirror Descent Inverse Reinforcement Learning

NeurIPS 2022accept

Recently, adversarial imitation learning has shown a scalable reward acquisition method for inverse reinforcement learning (IRL) problems. However, estimated reward signals often become uncertain and fail to train a reliable statistical model since the existing methods tend to solve hard optimizatio…

Cited by 5SourcePDFScholar
2022

SelecMix: Debiased Learning by Contradicting-pair Sampling

NeurIPS 2022accept

Neural networks trained with ERM (empirical risk minimization) sometimes learn unintended decision rules, in particular when their training data is biased, i.e., when training labels are strongly correlated with undesirable features. To prevent a network from learning such features, recent methods a…

2021

Attend What You Need: Motion-Appearance Synergistic Networks for Video Question Answering

ACL 2021long

Video Question Answering is a task which requires an AI agent to answer questions grounded in video. This task entails three key challenges: (1) understand the intention of various questions, (2) capturing various elements of the input video (e.g., object, action, causality), and (3) cross-modal gro…

2021

Devil’s Advocate: Novel Boosting Ensemble Method from Psychological Findings for Text Classification

EMNLP 2021finding

We present a new form of ensemble method–Devil’s Advocate, which uses a deliberately dissenting model to force other submodels within the ensemble to better collaborate. Our method consists of two different training settings: one follows the conventional training process (Norm), and the other is tra…

2021

DramaQA: Character-Centered Video Story Understanding with Hierarchical QA

AAAI 2021technical

Despite recent progress on computer vision and natural language processing, developing a machine that can understand video story is still hard to achieve due to the intrinsic difficulty of video story. Moreover, researches on how to evaluate the degree of video understanding based on human cognitive…

2021

Goal-Aware Cross-Entropy for Multi-Target Reinforcement Learning

NeurIPS 2021poster

Learning in a multi-target environment without prior knowledge about the targets requires a large amount of samples and makes generalization difficult. To solve this problem, it is important to be able to discriminate targets through semantic understanding. In this paper, we propose goal-aware cross…

2021

Message Passing Adaptive Resonance Theory for Online Active Semi-supervised Learning

ICML 2021spotlight

Active learning is widely used to reduce labeling effort and training time by repeatedly querying only the most beneficial samples from unlabeled data. In real-world problems where data cannot be stored indefinitely due to limited storage or privacy issues, the query selection and the model update s…

Cited by 17SourcePDFScholar
2021

Multimodal Anomaly Detection based on Deep Auto-Encoder for Object Slip Perception of Mobile Manipulation Robots

ICRA 2021poster

Object slip perception is essential for mobile manipulation robots to perform manipulation tasks reliably in the dynamic real-world. Traditional approaches to robot arms’ slip perception use tactile or vision sensors. However, mobile robots still have to deal with noise in their sensor signals cause…

Cited by 18SourceScholar
2021

Reasoning Visual Dialog with Sparse Graph Learning and Knowledge Transfer

EMNLP 2021finding

Visual dialog is a task of answering a sequence of questions grounded in an image using the previous dialog history as context. In this paper, we study how to address two fundamental challenges for this task: (1) reasoning over underlying semantic structures among dialog rounds and (2) identifying s…

2020

Hypergraph Attention Networks for Multimodal Learning

CVPR 2020poster

One of the fundamental problems that arise in multimodal learning tasks is the disparity of information levels between different modalities. To resolve this problem, we propose Hypergraph Attention Networks (HANs), which define a common semantic space among the modalities with symbolic graphs and ex…

Cited by 120PDFcodeScholar
2020

Label Propagation Adaptive Resonance Theory for Semi-Supervised Continuous Learning

ICASSP 2020accepted

Semi-supervised learning and continuous learning are fundamental paradigms for human-level intelligence. To deal with real-world problems where labels are rarely given and the opportunity to access the same data is limited, it is necessary to apply these two paradigms in a joined fashion. In this pa…

Cited by 0SourceScholar
2018

Answerer in Questioner's Mind: Information Theoretic Approach to Goal-Oriented Visual Dialog

NeurIPS 2018spotlight

Goal-oriented dialog has been given attention due to its numerous applications in artificial intelligence. Goal-oriented dialogue tasks occur when a questioner asks an action-oriented question and an answerer responds with the intent of letting the questioner know a correct action to take. To ask t…

Cited by 43SourcePDFScholar
2018

Multimodal Dual Attention Memory for Video Story Question Answering

ECCV 2018poster

We propose a video story question-answering (QA) architecture, Multimodal Dual Attention Memory (MDAM). The key idea is to use a dual attention mechanism with late fusion. MDAM uses self-attention to learn the latent concepts in scene frames and captions. Given a question, MDAM uses the second atten…

Cited by 97SourcePDFScholar
2018

Robust Human Following by Deep Bayesian Trajectory Prediction for Home Service Robots

ICRA 2018poster

The capability of following a person is crucial in service-oriented robots for human assistance and cooperation. Though a vast variety of following systems exist, they lack robustness against dynamic changes of the environment and relocating to continue following a lost target. Here we present a rob…

Cited by 50SourceScholar
2017

Hadamard Product for Low-rank Bilinear Pooling

ICLR 2017poster

Bilinear models provide rich representations compared with linear models. They have been applied in various visual tasks, such as object recognition, segmentation, and visual question-answering, to get state-of-the-art performances taking advantage of the expanded representations. However, bilinear…

Cited by 921SourcecodeScholar
2017

Overcoming Catastrophic Forgetting by Incremental Moment Matching

NeurIPS 2017spotlight

Catastrophic forgetting is a problem of neural networks that loses the information of the first task after training the second task. Here, we propose a method, i.e. incremental moment matching (IMM), to resolve this problem. IMM incrementally matches the moment of the posterior distribution of the n…

2016

Multimodal Residual Learning for Visual QA

NeurIPS 2016poster

Deep neural networks continue to advance the state-of-the-art of image recognition tasks with various methods. However, applications of these methods to multimodality remain limited. We present Multimodal Residual Networks (MRN) for the multimodal residual learning of visual question-answering, whic…