← Search

Dongwook Lee

12 accepted papers

2026

DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation

ICRA 2026poster

In dynamic environments such as warehouses, hospitals, and homes, robots must seamlessly transition between gross motion and precise manipulations to complete complex tasks. However, current Vision-Language-Action (VLA) frameworks, largely adapted from pre-trained Vision-Language Models (VLMs), ofte…

2025

3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View Transformation

CVPR 2025poster

The resolution of voxel queries significantly influences the quality of view transformation in camera-based 3D occupancy prediction. However, computational constraints and the practical necessity for real-time deployment require smaller query resolutions, which inevitably leads to an information los…

Cited by 2SourcePDFScholar
2025

Diffusion Transformer meets Multi-level Wavelet Spectrum for Single Image Super-Resolution

ICCV 2025poster

Discrete Wavelet Transform (DWT) has been widely explored to enhance the performance of image super-resolution (SR). Despite some DWT-based methods improving SR by capturing fine-grained frequency signals, most existing approaches neglect the interrelations among multi-scale frequency sub-bands, res…

Cited by 0SourcePDFScholar
2025

Test-Time Adaptation for Online Vision-Language Navigation with Feedback-based Reinforcement Learning

ICML 2025poster

Navigating in an unfamiliar environment during deployment poses a critical challenge for a vision-language navigation (VLN) agent. Yet, test-time adaptation (TTA) remains relatively underexplored in robotic navigation, leading us to the fundamental question: what are the key properties of TTA for on…

Cited by 0SourcePDFScholar
2024

CMDA: Cross-Modal and Domain Adversarial Adaptation for LiDAR-Based 3D Object Detection

AAAI 2024technical

Recent LiDAR-based 3D Object Detection (3DOD) methods show promising results, but they often do not generalize well to target domains outside the source (or training) data distribution. To reduce such domain gaps and thus to make 3DOD models more generalizable, we introduce a novel unsupervised doma…

Cited by 4SourcePDFScholar
2024

Compositional Conservatism: A Transductive Approach in Offline Reinforcement Learning

ICLR 2024poster

Offline reinforcement learning (RL) is a compelling framework for learning optimal policies from past experiences without additional interaction with the environment. Nevertheless, offline RL inevitably faces the problem of distributional shifts, where the states and actions encountered during polic…

2024

Efficient Learning on Successive Test Time Augmentation

ICASSP 2024accepted

Test time augmentation (TTA) has been a promising tool for improving the robustness against out-of-distribution data at inference time. Recent TTA methods try to learn predictive transformations which are supposed to provide the best performance gain on each test sample. However, existing methods ar…

Cited by 0SourceScholar
2024

Unified Domain Generalization and Adaptation for Multi-View 3D Object Detection

NeurIPS 2024poster

Recent advances in 3D object detection leveraging multi-view cameras have demonstrated their practical and economical value in various challenging vision tasks. However, typical supervised learning approaches face challenges in achieving satisfactory adaptation toward unseen and unlabeled target dat…

Cited by 1SourcePDFScholar
2023

STXD: Structural and Temporal Cross-Modal Distillation for Multi-View 3D Object Detection

NeurIPS 2023poster

3D object detection (3DOD) from multi-view images is an economically appealing alternative to expensive LiDAR-based detectors, but also an extremely challenging task due to the absence of precise spatial cues. Recent studies have leveraged the teacher-student paradigm for cross-modal distillation, w…

Cited by 8SourcePDFScholar
2022

Semi-Supervised 360° Depth Estimation from Multiple Fisheye Cameras with Pixel-Level Selective Loss

ICASSP 2022accepted

In this paper, we study a practical omnidirectional depth estimation with neural networks that enables effective learning on real world data obtained using wide-baseline multiple fish-eye cameras. Most previous approaches only used synthetic data providing dense and accurate depth ground truth (GT).…

Cited by 0SourceScholar
2021

RaScaNet: Learning Tiny Models by Raster-Scanning Images

CVPR 2021poster

Deploying deep convolutional neural networks on ultra-low power systems is challenging due to the extremely limited resources. Especially, the memory becomes a bottleneck as the systems put a hard limit on the size of on-chip memory. Because peak memory explosion in the lower layers is critical even…

Cited by 16PDFcodeScholar