← Search

Daehyun Ji

12 accepted papers

2026

DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation

ICRA 2026poster

In dynamic environments such as warehouses, hospitals, and homes, robots must seamlessly transition between gross motion and precise manipulations to complete complex tasks. However, current Vision-Language-Action (VLA) frameworks, largely adapted from pre-trained Vision-Language Models (VLMs), ofte…

2025

3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View Transformation

CVPR 2025poster

The resolution of voxel queries significantly influences the quality of view transformation in camera-based 3D occupancy prediction. However, computational constraints and the practical necessity for real-time deployment require smaller query resolutions, which inevitably leads to an information los…

Cited by 2SourcePDFScholar
2025

CDP: Constrained Diffusion Policies with Mirror Diffusion Model for Safety-Assured Imitation Learning

IROS 2025

This paper presents a novel imitation learning framework, called constrained diffusion policy (CDP). The primary objective of CDP is to ensure that learned policies strictly adhere to safety constraints while imitating expert demonstrations. To achieve this, we define a polytopic constraint that rep

Cited by 0SourceScholar
2025

Diffusion Transformer meets Multi-level Wavelet Spectrum for Single Image Super-Resolution

ICCV 2025poster

Discrete Wavelet Transform (DWT) has been widely explored to enhance the performance of image super-resolution (SR). Despite some DWT-based methods improving SR by capturing fine-grained frequency signals, most existing approaches neglect the interrelations among multi-scale frequency sub-bands, res…

Cited by 0SourcePDFScholar
2025

Test-Time Adaptation for Online Vision-Language Navigation with Feedback-based Reinforcement Learning

ICML 2025poster

Navigating in an unfamiliar environment during deployment poses a critical challenge for a vision-language navigation (VLN) agent. Yet, test-time adaptation (TTA) remains relatively underexplored in robotic navigation, leading us to the fundamental question: what are the key properties of TTA for on…

Cited by 0SourcePDFScholar
2024

CMDA: Cross-Modal and Domain Adversarial Adaptation for LiDAR-Based 3D Object Detection

AAAI 2024technical

Recent LiDAR-based 3D Object Detection (3DOD) methods show promising results, but they often do not generalize well to target domains outside the source (or training) data distribution. To reduce such domain gaps and thus to make 3DOD models more generalizable, we introduce a novel unsupervised doma…

Cited by 4SourcePDFScholar
2024

Unified Domain Generalization and Adaptation for Multi-View 3D Object Detection

NeurIPS 2024poster

Recent advances in 3D object detection leveraging multi-view cameras have demonstrated their practical and economical value in various challenging vision tasks. However, typical supervised learning approaches face challenges in achieving satisfactory adaptation toward unseen and unlabeled target dat…

Cited by 1SourcePDFScholar
2024

Unveiling the Hidden: Online Vectorized HD Map Construction with Clip-Level Token Interaction and Propagation

NeurIPS 2024poster

Predicting and constructing road geometric information (e.g., lane lines, road markers) is a crucial task for safe autonomous driving, while such static map elements can be repeatedly occluded by various dynamic objects on the road. Recent studies have shown significantly improved vectorized high-de…

2023

D-3DLD: Depth-Aware Voxel Space Mapping for Monocular 3D Lane Detection with Uncertainty

ICASSP 2023accepted

The estimation of 3D lanes from monocular RGB images is a fundamentally ill-posed problem. Previous studies have assumed that all lanes are on a flat ground plane. However, we argue that the algorithms based on this assumption have difficulty in detecting various lanes in actual driving environments…

Cited by 0SourceScholar
2023

STXD: Structural and Temporal Cross-Modal Distillation for Multi-View 3D Object Detection

NeurIPS 2023poster

3D object detection (3DOD) from multi-view images is an economically appealing alternative to expensive LiDAR-based detectors, but also an extremely challenging task due to the absence of precise spatial cues. Recent studies have leveraged the teacher-student paradigm for cross-modal distillation, w…

Cited by 8SourcePDFScholar
2022

Semi-Supervised 360° Depth Estimation from Multiple Fisheye Cameras with Pixel-Level Selective Loss

ICASSP 2022accepted

In this paper, we study a practical omnidirectional depth estimation with neural networks that enables effective learning on real world data obtained using wide-baseline multiple fish-eye cameras. Most previous approaches only used synthetic data providing dense and accurate depth ground truth (GT).…

Cited by 0SourceScholar
2020

Segmenting 2K-Videos at 36.5 FPS with 24.3 GFLOPs: Accurate and Lightweight Realtime Semantic Segmentation Network

ICRA 2020poster

We propose a fast and lightweight end-to-end convolutional network architecture for real-time segmentation of high resolution videos, NfS-SegNet, that can segement 2K-videos at 36.5 FPS with 24.3 GFLOPS. This speed and computation-efficiency is due to following reasons: 1) The encoder network, NfS-N…

Cited by 11SourceScholar