← Search

Sihaeng Lee

7 accepted papers

2025

MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations

CVPR 2025highlight

In this work, we tackle action-scene hallucination in Video Large Language Models (Video-LLMs), where models incorrectly predict actions based on the scene context or scenes based on observed actions. We observe that existing Video-LLMs often suffer from action-scene hallucination due to two main fa…

Cited by 2SourcePDFScholar
2025

ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data Streams

CVPR 2025poster

The rapid growth of video-text data presents challenges in storage and computation during training. Online learning, which processes streaming data in real-time, offers a promising solution to these issues while also allowing swift adaptations in scenarios demanding real-time responsiveness. One str…

2022

Fully Convolutional Transformer with Local-Global Attention

IROS 2022poster

In an attempt to imitate the success of transformers in the field of natural language processing into computer vision tasks, vision transformers (ViTs) have recently gained attention. Performance breakthroughs have been achieved in coarse-grained tasks like classification. However, dense prediction…

Cited by 1SourceScholar
2022

L-Verse: Bidirectional Generation Between Image and Text

CVPR 2022oral

Far beyond learning long-range interactions of natural language, transformers are becoming the de-facto standard for many vision tasks with their power and scalability. Especially with cross-modal tasks between image and text, vector quantized variational autoencoders (VQ-VAEs) are widely used to ma…

Cited by 33PDFcodeScholar
2022

Multi-Scaled and Densely Connected Locally Convolutional Layers for Depth Completion

IROS 2022poster

The depth completion task aims to predict a dense depth map from a sparse LiDAR point cloud and an RGB image. This task is critical because an accurate depth map can be used as prior information to solve many computer vision tasks, such as downstream tasks in autonomous vehicles and robot vision. Pr…

Cited by 3SourceScholar
2021

Patch-Wise Attention Network for Monocular Depth Estimation

AAAI 2021technical

In computer vision, monocular depth estimation is the problem of obtaining a high-quality depth map from a two-dimensional image. This map provides information on three-dimensional scene geometry, which is necessary for various applications in academia and industry, such as robotics and autonomous d…

Cited by 76SourcePDFScholar
2015

Joint Fine-Tuning in Deep Neural Networks for Facial Expression Recognition

ICCV 2015poster

Temporal information has useful features for recognizing facial expressions. However, to manually design useful features requires a lot of effort. In this paper, to reduce this effort, a deep learning technique, which is regarded as a tool to automatically extract useful features from raw data, is a…

Cited by 967PDFScholar