← Search

Dongseok Shim

13 accepted papers

2026

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models

CVPR 2026

Scaling multimodal alignment between video and audio is challenging, particularly due to limited data and the mismatch between text descriptions and frame-level video information. In this work, we tackle the scaling challenge in multimodal-to-audio generation, examining whether models trained on sho

Cited by 0SourcecodeScholar
2026

PTC-Depth: Pose-Refined Monocular Depth Estimation with Temporal Consistency

CVPR 2026

Monocular depth estimation (MDE) has been widely adopted in the perception systems of autonomous vehicles and mobile robots. However, existing approaches often struggle to maintain temporal consistency in depth estimation across consecutive frames. This inconsistency not only causes jitter but can a

Cited by 0SourceScholar
2026

Single-View 3D-Aware Representations for Reinforcement Learning by Cross-View Neural Radiance Fields

ICRA 2026poster

Reinforcement learning (RL) has enabled robots to develop complex skills, but its success in image-based tasks often depends on effective representation learning. Prior works have primarily focused on 2D representations, often overlooking the inherent 3D geometric structure of the world, or have att…

Cited by 0SourceScholar
2025

EUGens: Efficient, Unified and General Dense Layers

NeurIPS 2025poster

Efficient neural networks are essential for scaling machine learning models to real-time applications and resource-constrained environments. Fully-connected feedforward layers (FFLs) introduce computation and parameter count bottlenecks within neural network architectures. To address this challenge…

Cited by 0SourceScholar
2025

Single-View 3D-Aware Representations for Reinforcement Learning by Cross-View Neural Radiance Fields

RA-L 2025

Reinforcement learning (RL) has enabled robots to develop complex skills, but its success in image-based tasks often depends on effective representation learning. Prior works have primarily focused on 2D representations, often overlooking the inherent 3D geometric structure of the world, or have att

Cited by 0SourcecodeScholar
2024

Mono-Camera-Only Target Chasing for a Drone in a Dense Environment by Cross-Modal Learning

RA-L 2024

Chasing a dynamic target in a dense environment is one of the challenging applications of autonomous drones. The task requires multi-modal data, such as RGB and depth, to accomplish safe and robust maneuver. However, using different types of modalities can be difficult due to the limited capacity of

Cited by 4SourceScholar
2023

DiffuPose: Monocular 3D Human Pose Estimation via Denoising Diffusion Probabilistic Model

IROS 2023poster

Thanks to the development of 2D keypoint detectors, monocular 3D human pose estimation (HPE) via 2D-to-3D uplifting approaches have achieved remarkable improvements. Still, monocular 3D HPE is a challenging problem due to the inherent depth ambiguities and occlusions. To handle this problem, many pr…

Cited by 39SourcecodeScholar
2023

SwinDepth: Unsupervised Depth Estimation using Monocular Sequences via Swin Transformer and Densely Cascaded Network

ICRA 2023poster

Monocular depth estimation plays a critical role in various computer vision and robotics applications such as localization, mapping, and 3D object detection. Recently, learning-based algorithms achieve huge success in depth estimation by training models with a large amount of data in a supervised ma…

Cited by 35SourcecodeScholar
2022

S2P: State-conditioned Image Synthesis for Data Augmentation in Offline Reinforcement Learning

NeurIPS 2022accept

Offline reinforcement learning (Offline RL) suffers from the innate distributional shift as it cannot interact with the physical environment during training. To alleviate such limitation, state-based offline RL leverages a learned dynamics model from the logged experience and augments the predicted…

2021

Learning a Geometric Representation for Data-Efficient Depth Estimation via Gradient Field and Contrastive Loss

ICRA 2021poster

Estimating a depth map from a single RGB image has been investigated widely for localization, mapping, and 3- dimensional object detection. Recent studies on a single-view depth estimation are mostly based on deep Convolutional neural Networks (ConvNets) which require a large amount of training data…

Cited by 1SourcecodeScholar