← Search

Sujin Jang

11 accepted papers

2026

DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation

ICRA 2026poster

In dynamic environments such as warehouses, hospitals, and homes, robots must seamlessly transition between gross motion and precise manipulations to complete complex tasks. However, current Vision-Language-Action (VLA) frameworks, largely adapted from pre-trained Vision-Language Models (VLMs), ofte…

2025

3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View Transformation

CVPR 2025poster

The resolution of voxel queries significantly influences the quality of view transformation in camera-based 3D occupancy prediction. However, computational constraints and the practical necessity for real-time deployment require smaller query resolutions, which inevitably leads to an information los…

Cited by 2SourcePDFScholar
2025

Active Test-time Vision-Language Navigation

NeurIPS 2025poster

Vision-Language Navigation (VLN) policies trained on offline datasets often exhibit degraded task performance when deployed in unfamiliar navigation environments at test time, where agents are typically evaluated without access to external interaction or feedback. Entropy minimization has emerged as…

Cited by 0SourceScholar
2025

Test-Time Adaptation for Online Vision-Language Navigation with Feedback-based Reinforcement Learning

ICML 2025poster

Navigating in an unfamiliar environment during deployment poses a critical challenge for a vision-language navigation (VLN) agent. Yet, test-time adaptation (TTA) remains relatively underexplored in robotic navigation, leading us to the fundamental question: what are the key properties of TTA for on…

Cited by 0SourcePDFScholar
2024

CMDA: Cross-Modal and Domain Adversarial Adaptation for LiDAR-Based 3D Object Detection

AAAI 2024technical

Recent LiDAR-based 3D Object Detection (3DOD) methods show promising results, but they often do not generalize well to target domains outside the source (or training) data distribution. To reduce such domain gaps and thus to make 3DOD models more generalizable, we introduce a novel unsupervised doma…

Cited by 4SourcePDFScholar
2024

Unified Domain Generalization and Adaptation for Multi-View 3D Object Detection

NeurIPS 2024poster

Recent advances in 3D object detection leveraging multi-view cameras have demonstrated their practical and economical value in various challenging vision tasks. However, typical supervised learning approaches face challenges in achieving satisfactory adaptation toward unseen and unlabeled target dat…

Cited by 1SourcePDFScholar
2024

Unveiling the Hidden: Online Vectorized HD Map Construction with Clip-Level Token Interaction and Propagation

NeurIPS 2024poster

Predicting and constructing road geometric information (e.g., lane lines, road markers) is a crucial task for safe autonomous driving, while such static map elements can be repeatedly occluded by various dynamic objects on the road. Recent studies have shown significantly improved vectorized high-de…

2023

STXD: Structural and Temporal Cross-Modal Distillation for Multi-View 3D Object Detection

NeurIPS 2023poster

3D object detection (3DOD) from multi-view images is an economically appealing alternative to expensive LiDAR-based detectors, but also an extremely challenging task due to the absence of precise spatial cues. Recent studies have leveraged the teacher-student paradigm for cross-modal distillation, w…

Cited by 8SourcePDFScholar
2022

DaDA: Distortion-aware Domain Adaptation for Unsupervised Semantic Segmentation

NeurIPS 2022accept

Distributional shifts in photometry and texture have been extensively studied for unsupervised domain adaptation, but their counterparts in optical distortion have been largely neglected. In this work, we tackle the task of unsupervised domain adaptation for semantic image segmentation where unknown…

2022

Towards a Snake-Like Flexible Robot With Variable Stiffness Using an SMA Spring-Based Friction Change Mechanism

RA-L 2022

A flexible surgical robot that can adjust its stiffness guarantees safe operation by satisfying both the high flexibility required when the robot approaches a surgical target and the more stiffness necessary for the end effector to perform surgical procedures. Therefore, this paper proposes a flexib

Cited by 17SourceScholar
2015

A Collaborative Filtering Approach to Real-Time Hand Pose Estimation

ICCV 2015poster

Collaborative filtering aims to predict unknown user ratings in a recommender system by collectively assessing known user preferences. In this paper, we first draw analogies between collaborative filtering and the pose estimation problem. Specifically, we recast the hand pose estimation problem as t…

Cited by 71PDFScholar