← Search

Seungyeon Kim

25 accepted papers

2026

Motion Manifold Flow Primitives for Task-Conditioned Trajectory Generation under Complex Task-Motion Dependencies

ICRA 2026poster

Effective movement primitives should be capable of encoding and generating a rich repertoire of trajectories -- typically collected from human demonstrations -- conditioned on task-defining parameters such as vision or language inputs. While recent methods based on the motion manifold hypothesis, wh…

2025

Diverse Policy Learning via Random Obstacle Deployment for Zero-Shot Adaptation

RA-L 2025

In this letter, we propose a novel reinforcement learning framework that enables zero-shot policy adaptation in environments with unseen, dynamically changing obstacles. Adopting the idea that learning a policy capable of generating diverse actions is key to achieving such adaptability, our primary

Cited by 1SourceScholar
2025

Faster Cascades via Speculative Decoding

ICLR 2025oral

Cascades and speculative decoding are two common approaches to improving language models' inference efficiency. Both approaches interleave two models, but via fundamentally distinct mechanisms: deferral rule that invokes the larger model only for “hard” inputs, while speculative decoding uses spec…

Cited by 4SourcePDFScholar
2025

Motion Manifold Flow Primitives for Task-Conditioned Trajectory Generation Under Complex Task-Motion Dependencies

RA-L 2025

Effective movement primitives should be capable of encoding and generating a rich repertoire of trajectories conditioned on task-defining parameters such as vision or language inputs. While recent methods based on the motion manifold hypothesis, which assumes that a set of trajectories lies on a low

Cited by 3SourceScholar
2025

OPPA: Online Planner's Parameter Adaptation for Enhanced Mobile Robot Navigation

ICRA 2025

Autonomous navigation in mobile robots has made significant advancements; however, traditional methods often struggle to adapt in real-time to dynamic or unstructured environments. This paper presents the Online Planner's Parameter Adaptation (OPPA) framework, which enhances both adaptability and sa

Cited by 1SourceScholar
2025

Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA

ICLR 2025poster

Large language models (LLMs) are expensive to deploy. Parameter sharing offers a possible path towards reducing their size and cost, but its effectiveness in modern LLMs remains fairly limited. In this work, we revisit "layer tying" as form of parameter sharing in Transformers, and introduce novel m…

Cited by 5SourcePDFScholar
2025

ScrewSplat: An End-to-End Method for Articulated Object Recognition

CoRL 2025oral

Articulated object recognition -- the task of identifying both the geometry and kinematic joints of objects with movable parts -- is essential for enabling robots to interact with everyday objects such as doors and laptops. However, existing approaches often rely on strong assumptions, such as a kno…

Cited by 0SourceScholar
2024

Analysis of Plan-based Retrieval for Grounded Text Generation

EMNLP 2024main

In text generation, hallucinations refer to the generation of seemingly coherent text that contradicts established knowledge. One compelling hypothesis is that hallucinations occur when a language model is given a generation task outside its parametric knowledge (due to rarity, recency, domain, etc.…

Cited by 1SourcePDFScholar
2024

T$^2$SQNet: A Recognition Model for Manipulating Partially Observed Transparent Tableware Objects

CoRL 2024poster

Recognizing and manipulating transparent tableware from partial view RGB image observations is made challenging by the difficulty in obtaining reliable depth measurements of transparent objects. In this paper we present the Transparent Tableware SuperQuadric Network (T$^2$SQNet), a neural network m…

Cited by 0SourceScholar
2024

USTAD: Unified Single-model Training Achieving Diverse Scores for Information Retrieval

ICML 2024poster

Modern information retrieval (IR) systems consists of multiple stages like retrieval and ranking, with Transformer-based models achieving state-of-the-art performance at each stage. In this paper, we challenge the tradition of using separate models for different stages and ask if a single Transforme…

Cited by 0SourcePDFScholar
2023

ARC Joint: Anthropomorphic Rolling Contact Joint With Kinematically Variable Torsional Stiffness

RA-L 2023

As compliant joints not only compensate for the lack of actuated degrees of freedom of an under-actuated system and improve grasp stability but also prevent system failure from unexpected contacts, various types of compliant joints have been applied to end-effectors. Although joint compliance increa

Cited by 6SourceScholar
2023

Efficient Training of Language Models using Few-Shot Learning

ICML 2023poster

Large deep learning models have achieved state-of-the-art performance across various natural language processing (NLP) tasks and demonstrated remarkable few-shot learning performance. However, training them is often challenging and resource-intensive. In this paper, we study an efficient approach to…

Cited by 15SourcePDFScholar
2023

Leveraging 3D Reconstruction for Mechanical Search on Cluttered Shelves

CoRL 2023poster

Finding and grasping a target object on a cluttered shelf, especially when the target is occluded by other unknown objects and initially invisible, remains a significant challenge in robotic manipulation. While there have been advances in finding the target object by rearranging surrounding objects…

Cited by 4SourcecodeScholar
2023

Supervision Complexity and its Role in Knowledge Distillation

ICLR 2023poster

Despite the popularity and efficacy of knowledge distillation, there is limited understanding of why it helps. In order to study the generalization behavior of a distilled student, we propose a new theoretical framework that leverages supervision complexity: a measure of alignment between teacher-pr…

Cited by 14SourcePDFScholar
2023

Teacher Guided Training: An Efficient Framework for Knowledge Transfer

ICLR 2023poster

The remarkable performance gains realized by large pretrained models, e.g., GPT-3, hinge on the massive amounts of data they are exposed to during training. Analogously, distilling such large models to compact models for efficient deployment also necessitates a large amount of (labeled or unlabeled)…

Cited by 2SourcePDFScholar
2022

A Statistical Manifold Framework for Point Cloud Data

ICML 2022spotlight

Many problems in machine learning involve data sets in which each data point is a point cloud in $\mathbb{R}^D$. A growing number of applications require a means of measuring not only distances between point clouds, but also angles, volumes, derivatives, and other more advanced concepts. To formulat…

2022

Contact State Estimation for Peg-in-Hole Assembly Using Gaussian Mixture Model

RA-L 2022

Recently, the robotic assembly has been expanded into an unstructured environment. This environment includes uncertainties that may cause unexpected situations such as a failure of the assembly. Such problems can be prevented or monitored by a robust contact state (CS) estimation method. In that sen

Cited by 32SourceScholar
2022

In defense of dual-encoders for neural ranking

ICML 2022spotlight

Transformer-based models such as BERT have proven successful in information retrieval problem, which seek to identify relevant documents for a given query. There are two broad flavours of such models: cross-attention (CA) models, which learn a joint embedding for the query and document, and dual-enc…

Cited by 33SourcePDFScholar
2022

SE(2)-Equivariant Pushing Dynamics Models for Tabletop Object Manipulations

CoRL 2022oral

For tabletop object manipulation tasks, learning an accurate pushing dynamics model, which predicts the objects' motions when a robot pushes an object, is very important. In this work, we claim that an ideal pushing dynamics model should have the SE(2)-equivariance property, i.e., if tabletop object…

Cited by 13SourcecodeScholar
2021

A statistical perspective on distillation

ICML 2021spotlight

Knowledge distillation is a technique for improving a “student” model by replacing its one-hot training labels with a label distribution obtained from a “teacher” model. Despite its broad success, several basic questions — e.g., Why does distillation help? Why do more accurate teachers not necessari…

Cited by 107SourcePDFScholar
2021

Evaluations and Methods for Explanation through Robustness Analysis

ICLR 2021poster

Feature based explanations, that provide importance of each feature towards the model prediction, is arguably one of the most intuitive ways to explain a model. In this paper, we establish a novel set of evaluation criteria for such feature based explanations by robustness analysis. In contrast to e…

Cited by 69SourcePDFScholar
2021

RankDistil: Knowledge Distillation for Ranking

AISTATS 2021poster

Knowledge distillation is an approach to improve the performance of a student model by using the knowledge of a complex teacher. Despite its success in several deep learning applications, the study of distillation is mostly confined to classification settings. In particular, the use of distillation…

Cited by 38SourcePDFScholar
2020

Why are Adaptive Methods Good for Attention Models?

NeurIPS 2020poster

While stochastic gradient descent (SGD) is still the de facto algorithm in deep learning, adaptive methods like Clipped SGD/Adam have been observed to outperform SGD across important tasks, such as attention models. The settings under which SGD performs poorly in comparison to adaptive methods are n…