← Search

Jinhyung Kim

8 accepted papers

2025

MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations

CVPR 2025highlight

In this work, we tackle action-scene hallucination in Video Large Language Models (Video-LLMs), where models incorrectly predict actions based on the scene context or scenes based on observed actions. We observe that existing Video-LLMs often suffer from action-scene hallucination due to two main fa…

Cited by 2SourcePDFScholar
2024

Brain-Inspired Hyperdimensional Computing in the Wild: Lightweight Symbolic Learning for Sensorimotor Controls of Wheeled Robots

ICRA 2024poster

Efficiency and performance are significant challenges in applying Machine Learning (ML) to robotics, especially in energy-constrained real-world scenarios. In this context, Hyperdimensional Computing offers an energy-efficient alternative but has been underexplored in robotics. We introduce ReactHD,…

Cited by 3SourceScholar
2024

Expediting Contrastive Language-Image Pretraining via Self-Distilled Encoders

AAAI 2024technical

Recent advances in vision language pretraining (VLP) have been largely attributed to the large-scale data collected from the web. However, uncurated dataset contains weakly correlated image-text pairs, causing data inefficiency. To address the issue, knowledge distillation have been explored at the…

2023

Exploring Temporally Dynamic Data Augmentation for Video Recognition

ICLR 2023top-25%

Data augmentation has recently emerged as an essential component of modern training recipes for visual recognition tasks. However, data augmentation for video recognition has been rarely explored despite its effectiveness. Few existing augmentation recipes for video recognition naively extend the im…

Cited by 13SourcePDFScholar
2023

Frequency Selective Augmentation for Video Representation Learning

AAAI 2023technical

Recent self-supervised video representation learning methods focus on maximizing the similarity between multiple augmented views from the same video and largely rely on the quality of generated views. However, most existing methods lack a mechanism to prevent representation learning from bias toward…

Cited by 5SourcePDFScholar
2023

Misalign, Contrast then Distill: Rethinking Misalignments in Language-Image Pre-training

ICCV 2023poster

Contrastive Language-Image Pretraining has emerged as a prominent approach for training vision and text encoders with uncurated image-text pairs from the web. To enhance data-efficiency, recent efforts have introduced additional supervision terms that involve random-augmented views of the image. How…

Cited by 8PDFcodeScholar
2020

READ: Reciprocal Attention Discriminator for Image-to-Video Re-Identification

ECCV 2020poster

Person re-identification (re-ID) is the problem of visually identifying a person given a database of identities. In this work, we focus on image-to-video re-ID which compares a single query image to videos in the gallery. The main challenge is the asymmetry association of an image and a video, and o…

2020

Regularization on Spatio-Temporally Smoothed Feature for Action Recognition

CVPR 2020poster

Deep neural networks for video action recognition frequently require 3D convolutional filters and often encounter overfitting due to a larger number of parameters. In this paper, we propose Random Mean Scaling (RMS), a simple and effective regularization method, to relieve the overfitting problem in…

Cited by 34PDFScholar