← Search

Jinkyu Kim

32 accepted papers

2026

AxisGuide: Grounding Robot Action Coordinate System in RGB Observations for Robust Visuomotor Manipulation

RSS 2026poster

Visuomotor manipulation policies trained via large-scale behavior cloning have achieved strong semantic scene understanding, yet often fail to reliably execute correct low-level actions under distribution shifts. For example, even in a simple pick-up task with identical scene layouts, camera viewpoi…

Cited by 0SourceScholar
2026

Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding

CVPR 2026

Large Vision-Language Models (LVLMs) have shown strong performance across various multimodal tasks by leveraging the reasoning capabilities of Large Language Models (LLMs). However, processing visually complex and information-rich images, such as infographics or document layouts, requires these mode

Cited by 0SourcecodeScholar
2026

Multi-Modal Locomotion Mode Recognition in the Real World for Robotic Hip Complex Exoskeletons

ICRA 2026poster

Lower limb exoskeletons assist users by supporting joint movements. Since joint motion patterns vary depending on how the user moves, accurately recognizing the type of movement (locomotion mode) is crucial for controlling the exoskeleton and ensuring user safety. Inspired by how humans use multiple…

Cited by 0SourceScholar
2026

The Truth Stays in the Family: Enhancing Contextual Truthfulness via Inherited Heads in Model Lineages

ICML 2026poster

Recent advances in large language models (LLMs) have led to the emergence of specialized multimodal LLMs (MLLMs), forming distinct model families that share a common foundation language models. Despite this evolutionary trend, it remains unexplored whether a fundamental behavioral link exists betwee…

Cited by 0SourceScholar
2025

3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View Transformation

CVPR 2025poster

The resolution of voxel queries significantly influences the quality of view transformation in camera-based 3D occupancy prediction. However, computational constraints and the practical necessity for real-time deployment require smaller query resolutions, which inevitably leads to an information los…

Cited by 2SourcePDFScholar
2025

DiffExp: Efficient Exploration in Reward Fine-tuning for Text-to-Image Diffusion Models

AAAI 2025technical

Fine-tuning text-to-image diffusion models to maximize rewards has proven effective for enhancing model performance. However, reward fine-tuning methods often suffer from slow convergence due to online sample generation. Therefore, obtaining diverse samples with strong reward signals is crucial for…

Cited by 0SourcePDFScholar
2025

Multi-Modal Locomotion Mode Recognition in the Real World for Robotic Hip Complex Exoskeletons

RA-L 2025

Lower limb exoskeletons assist users by supporting joint movements. Since joint motion patterns vary depending on how the user moves, accurately recognizing the type of movement (locomotion mode) is crucial for controlling the exoskeleton and ensuring user safety. Inspired by how humans use multiple

Cited by 0SourceScholar
2025

Watermarking for Factuality: Guiding Vision-Language Models Toward Truth via Tri-layer Contrastive Decoding

EMNLP 2025

Large Vision-Language Models (LVLMs) have recently shown promising results on various multimodal tasks, even achieving human-comparable performance in certain cases. Nevertheless, LVLMs remain prone to hallucinations–they often rely heavily on a single modality or memorize training data without prop

Cited by 0SourcePDFScholar
2024

CMDA: Cross-Modal and Domain Adversarial Adaptation for LiDAR-Based 3D Object Detection

AAAI 2024technical

Recent LiDAR-based 3D Object Detection (3DOD) methods show promising results, but they often do not generalize well to target domains outside the source (or training) data distribution. To reduce such domain gaps and thus to make 3DOD models more generalizable, we introduce a novel unsupervised doma…

Cited by 4SourcePDFScholar
2024

Communication-Efficient Federated Learning with Accelerated Client Gradient

CVPR 2024poster

Federated learning often suffers from slow and unstable convergence due to the heterogeneous characteristics of participating client datasets. Such a tendency is aggravated when the client participation ratio is low since the information collected from the clients has large variations. To address th…

2024

Finetuning Pre-trained Model with Limited Data for LiDAR-based 3D Object Detection by Bridging Domain Gaps

IROS 2024poster

LiDAR-based 3D object detectors have been largely utilized in various applications, including autonomous vehicles or mobile robots. However, LiDAR-based detectors often fail to adapt well to target domains with different sensor configurations (e.g., types of sensors, spatial resolution, or FOVs) and…

Cited by 0SourceScholar
2024

Higher-order Relational Reasoning for Pedestrian Trajectory Prediction

CVPR 2024poster

Social relations have substantial impacts on the potential trajectories of each individual. Modeling these dynamics has been a central solution for more precise and accurate trajectory forecasting. However previous works ignore the importance of `social depth' meaning the influences flowing from dif…

Cited by 13SourcePDFScholar
2024

Just Add $100 More: Augmenting Pseudo-LiDAR Point Cloud for Resolving Class-imbalance Problem

NeurIPS 2024poster

Typical LiDAR-based 3D object detection models are trained with real-world data collection, which is often imbalanced over classes. To deal with it, augmentation techniques are commonly used, such as copying ground truth LiDAR points and pasting them into scenes. However, existing methods struggle w…

2024

Learning Temporal Cues by Predicting Objects Move for Multi-camera 3D Object Detection

ICRA 2024poster

In autonomous driving and robotics, there is a growing interest in utilizing short-term historical data to enhance multi-camera 3D object detection, leveraging the continuous and correlated nature of input video streams. Recent work has focused on spatially aligning BEV-based features over timesteps…

Cited by 0SourceScholar
2024

MEVG : Multi-event Video Generation with Text-to-Video Models

ECCV 2024poster

"We introduce a novel diffusion-based video generation method, generating a video showing multiple events given multiple individual sentences from the user. Our method does not require a large-scale video dataset since our method uses a pre-trained diffusion-based text-to-video generative model with…

2024

Unified Domain Generalization and Adaptation for Multi-View 3D Object Detection

NeurIPS 2024poster

Recent advances in 3D object detection leveraging multi-view cameras have demonstrated their practical and economical value in various challenging vision tasks. However, typical supervised learning approaches face challenges in achieving satisfactory adaptation toward unseen and unlabeled target dat…

Cited by 1SourcePDFScholar
2023

The Power of Sound (TPoS): Audio Reactive Video Generation with Stable Diffusion

ICCV 2023poster

In recent years, video generation has become a prominent generative tool and has drawn significant attention. However, there is little consideration in audio-to-video generation, though audio contains unique qualities like temporal semantics and magnitude. Hence, we propose The Power of Sound (TPoS)…

Cited by 38PDFcodeScholar
2022

Bridging the Domain Gap towards Generalization in Automatic Colorization

ECCV 2022poster

"We propose a novel automatic colorization technique that learns domain-invariance across multiple source domains and is able to leverage such invariance to colorize grayscale images in unseen target domains. This would be particularly useful for colorizing sketches, line arts, or line drawings, whi…

2022

Grounding Visual Representations with Texts for Domain Generalization

ECCV 2022poster

"Reducing the representational discrepancy between source and target domains is a key component to maximize the model generalization. In this work, we advocate for leveraging natural language supervision for the domain generalization task. We introduce two modules to ground visual representations wi…

2022

Occupancy Flow Fields for Motion Forecasting in Autonomous Driving

RA-L 2022

We propose Occupancy Flow Fields, a new representation for motion forecasting of multiple agents, an important task in autonomous driving.Our representation is a spatio-temporal grid with each grid cell containing both the probability of the cell being occupied by any agent, and a two-dimensional fl

Cited by 99SourceScholar
2022

Sound-Guided Semantic Image Manipulation

CVPR 2022poster

The recent success of the generative model shows that leveraging the multi-modal embedding space can manipulate an image using text information. However, manipulating an image with other sources rather than text, such as sound, is not easy due to the dynamic characteristics of the sources. Especiall…

Cited by 60PDFcodeScholar
2022

Sound-Guided Semantic Video Generation

ECCV 2022poster

"The recent success in StyleGAN demonstrates that pre-trained StyleGAN latent space is useful for realistic video generation. However, the generated motion in the video is usually not semantically meaningful due to the difficulty of determining the direction and magnitude in the StyleGAN latent spac…

Cited by 40SourcePDFScholar
2022

StopNet: Scalable Trajectory and Occupancy Prediction for Urban Autonomous Driving

ICRA 2022poster

We introduce a motion forecasting (behavior prediction) method that meets the latency requirements for autonomous driving in dense urban environments without sacrificing accuracy. A whole-scene sparse input representation allows StopNet to scale to predicting trajectories for hundreds of road agents…

Cited by 27SourceScholar
2021

SelfReg: Self-Supervised Contrastive Regularization for Domain Generalization

ICCV 2021poster

In general, an experimental environment for deep learning assumes that the training and the test dataset are sampled from the same distribution. However, in real-world situations, a difference in the distribution between two datasets, i.e. domain shift, may occur, which becomes a major factor impedi…

Cited by 360PDFcodeScholar
2020

Advisable Learning for Self-Driving Vehicles by Internalizing Observation-to-Action Rules

CVPR 2020poster

Humans learn to drive through both practice and theory, e.g. by studying the rules, while most self-driving systems are limited to the former. Being able to incorporate human knowledge of typical causal driving behaviour should benefit autonomous systems. We propose a new approach that learns vehicl…

Cited by 65PDFcodeScholar
2019

Grounding Human-To-Vehicle Advice for Self-Driving Vehicles

CVPR 2019poster

Recent success suggests that deep neural control networks are likely to be a key component of self-driving vehicles. These networks are trained on large datasets to imitate human actions, but they lack semantic understanding of image contents. This makes them brittle and potentially unsafe in situat…

Cited by 131PDFScholar
2018

Textual Explanations for Self-Driving Vehicles

ECCV 2018poster

Deep neural perception and control networks have become key components of self-driving vehicles. User acceptance is likely to benefit from easy-to-interpret textual explanations which allow end-users to understand what triggered a particular behavior. Explanations may be triggered by the neural cont…