← Search

Donghyun Kim

64 accepted papers

2026

Block-based Learned Image Compression without Blocking Artifacts

CVPR 2026

Learned image compression (LIC) outperforms traditional codecs but suffers from excessive peak memory usage when handling high-resolution images. Consequently, block-based LIC has been studied to reduce peak memory and peak computational cost, but it often introduces blocking artifacts that degrade

Cited by 0SourceScholar
2026

Collaborative Planning with Concurrent Synchronization for Operationally Constrained UAV-UGV Teams

ICRA 2026poster

Collaborative planning under operational constraints is an essential capability for heterogeneous robot teams tackling complex large-scale real-world tasks. Unmanned Aerial Vehicles (UAVs) offer rapid environmental coverage, but flight time is often limited by energy constraints, whereas Unmanned Gr…

2026

GuideTWSI: A Diverse Tactile Walking Surface Indicator Dataset from Synthetic and Real-World Images for Blind and Low-Vision Navigation

ICRA 2026poster

Tactile Walking Surface Indicators (TWSIs) are safety-critical landmarks that blind and low-vision (BLV) pedestrians use to locate crossings and hazard zones. From our observation sessions with BLV guide dog handlers, trainers, and an O&M specialist, we confirmed the critical importance of reliable …

2026

Occlusion-Robust Relative Pose Estimation for Multi-Robot Systems Via Geometric-Aware Diffusion Matching

ICRA 2026poster

Relative pose estimation is crucial for coordinated multi-robot navigation. However, robots in close proximity often face intra-team occlusions, where teammates partially block each other's field of view, while dynamic environments further introduce environmental occlusions. Classical relative pose …

Cited by 0Scholar
2026

Physics-Informed Audio-Geometry-Grid Representation Learning for Universal Sound Source Localization

ICLR 2026poster

Sound source localization (SSL) is a fundamental task in spatial audio understanding, yet most deep neural network-based methods are constrained by fixed array geometries and predefined directional grids, limiting generalizability and scalability. To address these issues, we propose _audio-geometry-…

Cited by 0SourcecodeScholar
2026

Universal Compressed Image Restoration via Codec-Aware Conditioning with Reinforcement Learning

AAAI 2026technical

We address the task of universal compressed image restoration, which involves recovering high-quality images degraded by a wide range of codecs and compression levels. While prior methods have made significant progress, they typically target specific degradation types and struggle to generalize acro

Cited by 0SourcePDFScholar
2025

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning

ICCV 2025poster

An image captioning model flexibly switching its language pattern, e.g., descriptiveness and length, should be useful since it can be applied to diverse applications. However, despite the dramatic improvement in generative vision-language models, fine-grained control over the properties of generated…

2025

Difference Inversion: Interpolate and Isolate the Difference with Token Consistency for Image Analogy Generation

CVPR 2025poster

How can we generate an image B' that satisfies A:A'::B:B', given the input images A,A' and B? Recent works have tackled this challenge through approaches like visual in-context learning or visual instruction. However, these methods are typically limited to specific models (InstructPix2Pix. Inpaintin…

Cited by 0SourcePDFScholar
2025

Large Language Models as Realistic Microservice Trace Generators

EMNLP 2025

Workload traces are essential to understand complex computer systems’ behavior and manage processing and memory resources. Since real-world traces are hard to obtain, synthetic trace generation is a promising alternative. This paper proposes a first-of-a-kind approach that relies on training a large

2025

MA-CIR: A Multimodal Arithmetic Benchmark for Composed Image Retrieval

ICCV 2025poster

Composed Image Retrieval (CIR) seeks to retrieve a target image by using a reference image and conditioning text specifying desired modifications. While recent approaches have shown steady performance improvements on existing CIR benchmarks, we argue that it remains unclear whether these gains genui…

2025

PLATYPUS: Progressive Local Surface Estimator for Arbitrary-Scale Point Cloud Upsampling

AAAI 2025technical

3D point clouds are increasingly vital for applications like autonomous driving and robotics, yet the raw data captured by sensors often suffer from noise and sparsity, creating challenges for downstream tasks. Consequently, point cloud upsampling becomes essential for improving density and uniformi…

Cited by 1SourcePDFScholar
2025

RA-SGG: Retrieval-Augmented Scene Graph Generation Framework via Multi-Prototype Learning

AAAI 2025technical

Scene Graph Generation (SGG) research has suffered from two fundamental challenges: the long-tailed predicate distribution and semantic ambiguity between predicates. These challenges lead to a bias towards head predicates in SGG models, favoring dominant general predicates while overlooking fine-gra…

2025

StitchLLM: Serving LLMs, One Block at a Time

ACL 2025long

The rapid evolution of large language models (LLMs) has revolutionized natural language processing (NLP) tasks such as text generation, translation, and comprehension. However, the increasing computational demands and inference costs of these models present significant challenges. This study investi…

Cited by 0SourcePDFScholar
2025

Weakly Supervised Video Scene Graph Generation via Natural Language Supervision

ICLR 2025poster

Existing Video Scene Graph Generation (VidSGG) studies are trained in a fully supervised manner, which requires all frames in a video to be annotated, thereby incurring high annotation cost compared to Image Scene Graph Generation (ImgSGG). Although the annotation cost of VidSGG can be alleviated by…

2025

ZIUM: Zero-Shot Intent-Aware Adversarial Attack on Unlearned Models

ICCV 2025poster

Machine unlearning (MU) removes specific data points or concepts from deep learning models to enhance privacy and prevent sensitive content generation. Adversarial prompts can exploit unlearned models to generate content containing removed concepts, posing a significant security risk. However, exist…

Cited by 0SourcePDFScholar
2024

Adaptive Self-training Framework for Fine-grained Scene Graph Generation

ICLR 2024poster

Scene graph generation (SGG) models have suffered from inherent problems regarding the benchmark datasets such as the long-tailed predicate distribution and missing annotation problems. In this work, we aim to alleviate the long-tailed problem of SGG by utilizing unannotated triplets. To this end, w…

2024

Efficient and Versatile Robust Fine-Tuning of Zero-shot Models

ECCV 2024poster

"Large-scale image-text pre-trained models enable zero-shot classification and provide consistent accuracy across various data distributions. Nonetheless, optimizing these models in downstream tasks typically requires fine-tuning, which reduces generalization to out-of-distribution (OOD) data and de…

Cited by 3SourcePDFScholar
2024

GripFlexer: Development of hybrid gripper with a novel shape that can perform in narrow spaces

IROS 2024poster

In recent years, the role of robots across industries has become increasingly diverse, and they are now required to perform complex missions beyond simple repetitive tasks. However, robots used in confined spaces that humans cannot reach or in disaster field missions have challenges in performing va…

Cited by 2SourceScholar
2024

LLM4SGG: Large Language Models for Weakly Supervised Scene Graph Generation

CVPR 2024poster

Weakly-Supervised Scene Graph Generation (WSSGG) research has recently emerged as an alternative to the fully-supervised approach that heavily relies on costly annotations. In this regard studies on WSSGG have utilized image captions to obtain unlocalized triplets while primarily focusing on groundi…

2024

MATE: Meet At The Embedding - Connecting Images with Long Texts

EMNLP 2024finding

While advancements in Vision Language Models (VLMs) have significantly improved the alignment of visual and textual data, these models primarily focus on aligning images with short descriptive captions. This focus limits their ability to handle complex text interactions, particularly with longer tex…

Cited by 6SourcePDFScholar
2024

StaccaToe: A Single-Leg Robot that Mimics the Human Leg and Toe

IROS 2024poster

We introduce StaccaToe, a human-scale, electric motor-powered single-leg robot designed to rival the agility of human locomotion through two distinctive attributes: an actuated toe and a co-actuation configuration inspired by the human leg. Leveraging the foundational design of HyperLeg’s lower leg…

Cited by 1SourceScholar
2024

Unified Domain Generalization and Adaptation for Multi-View 3D Object Detection

NeurIPS 2024poster

Recent advances in 3D object detection leveraging multi-view cameras have demonstrated their practical and economical value in various challenging vision tasks. However, typical supervised learning approaches face challenges in achieving satisfactory adaptation toward unseen and unlabeled target dat…

Cited by 1SourcePDFScholar
2024

Visual Delta Generator with Large Multi-modal Models for Semi-supervised Composed Image Retrieval

CVPR 2024poster

Composed Image Retrieval (CIR) is a task that retrieves images similar to a query based on a provided textual modification. Current techniques rely on supervised learning for CIR models using labeled triplets of the <reference image text target image>. These specific triplets are not as commonly ava…

Cited by 14SourcePDFScholar
2024

What How and When Should Object Detectors Update in Continually Changing Test Domains?

CVPR 2024poster

It is a well-known fact that the performance of deep learning models deteriorates when they encounter a distribution shift at test time. Test-time adaptation (TTA) algorithms have been proposed to adapt the model online while inferring test data. However existing research predominantly focuses on cl…

2023

Anthropomorphic robot hand using the principle of sweat and fingerprints of human hands

ICRA 2023poster

In our daily life, when a small amount of sweat or water forms on a person's hand, we can empirically feel that the friction force of the hand increases, and the objects are gripped well. However, if sweat or water forms heavily, we can feel the friction decrease when holding an object. In this stud…

Cited by 4SourceScholar
2023

CDAC: Cross-domain Attention Consistency in Transformer for Domain Adaptive Semantic Segmentation

ICCV 2023poster

While transformers have greatly boosted performance in semantic segmentation, domain adaptive transformers are not yet well explored. We identify that the domain gap can cause discrepancies in self-attention. Due to this gap, the transformer attends to spurious regions or pixels, which deteriorates…

Cited by 22PDFcodeScholar
2023

CODA-Prompt: COntinual Decomposed Attention-Based Prompting for Rehearsal-Free Continual Learning

CVPR 2023poster

Computer vision models suffer from a phenomenon known as catastrophic forgetting when learning novel concepts from continuously shifting training data. Typical solutions for this continual learning problem require extensive rehearsal of previously seen data, which increases memory costs and may viol…

2023

ConStruct-VL: Data-Free Continual Structured VL Concepts Learning

CVPR 2023poster

Recently, large-scale pre-trained Vision-and-Language (VL) foundation models have demonstrated remarkable capabilities in many zero-shot downstream tasks, achieving competitive results for recognizing objects defined by as little as short text prompts. However, it has also been shown that VL models…

2023

Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL Models

NeurIPS 2023spotlight

Vision and Language (VL) models offer an effective method for aligning representation spaces of images and text allowing for numerous applications such as cross-modal retrieval, visual and multi-hop question answering, captioning, and many more. However, the aligned image-text spaces learned by all…

Cited by 50SourcePDFScholar
2023

Event Camera-Based Visual Odometry for Dynamic Motion Tracking of a Legged Robot Using Adaptive Time Surface

IROS 2023poster

Our paper proposes a direct sparse visual odometry method that combines event and RGBD data to estimate the pose of agile-legged robots during dynamic locomotion and acrobatic behaviors. Event cameras offer high temporal resolution and dynamic range, which can eliminate the issue of blurred RGB imag…

Cited by 6SourceScholar
2023

Going Beyond Nouns With Vision & Language Models Using Synthetic Data

ICCV 2023poster

Large-scale pre-trained Vision & Language (VL) models have shown remarkable performance in many applications, enabling replacing a fixed set of supported classes with zero-shot open vocabulary reasoning over (almost arbitrary) natural language prompts. However, recent works have uncovered a fundamen…

Cited by 51PDFcodeScholar
2023

Learning Human Action Recognition Representations Without Real Humans

NeurIPS 2023poster

Pre-training on massive video datasets has become essential to achieve high action recognition performance on smaller downstream datasets. However, most large-scale video datasets contain images of people and hence are accompanied with issues related to privacy, ethics, and data protection, often pr…

2023

SCOB: Universal Text Understanding via Character-wise Supervised Contrastive Learning with Online Text Rendering for Bridging Domain Gap

ICCV 2023poster

Inspired by the great success of language model (LM)-based pre-training, recent studies in visual document understanding have explored LM-based pre-training methods for modeling text within document images. Among them, pre-training that reads all text from an image has shown promise, but often exhib…

Cited by 2PDFcodeScholar
2023

System Configuration and Navigation of a Guide Dog Robot: Toward Animal Guide Dog-Level Guiding Work

ICRA 2023poster

A robot guide dog has compelling advantages over animal guide dogs for its cost-effectiveness, the potential for mass production, and low maintenance burden. However, despite the long history of guide dog robot research, previous studies were conducted with little or no consideration of how the guid…

Cited by 33SourceScholar
2023

Towards Efficient Image Compression Without Autoregressive Models

NeurIPS 2023poster

Recently, learned image compression (LIC) has garnered increasing interest with its rapidly improving performance surpassing conventional codecs. A key ingredient of LIC is a hyperprior-based entropy model, where the underlying joint probability of the latent image features is modeled as a product o…

Cited by 13SourcePDFScholar
2022

A Broad Study of Pre-training for Domain Generalization and Adaptation

ECCV 2022poster

"Deep models must learn robust and transferable representations in order to perform well on new domains. While domain transfer methods (\eg, domain adaptation, domain generalization) have been proposed to learn transferable representations across domains, they are typically applied to ResNet backbon…

2022

A Unified Framework for Domain Adaptive Pose Estimation

ECCV 2022poster

"While pose estimation is an important computer vision task, it requires expensive annotation and suffers from domain shift. In this paper, we investigate the problem of domain adaptive 2D pose estimation that transfers knowledge learned on a synthetic source domain to a target domain without superv…

2022

Active Learning on Pre-trained Language Model with Task-Independent Triplet Loss

AAAI 2022technical

Active learning attempts to maximize a task model’s performance gain by obtaining a set of informative samples from an unlabeled data pool. Previous active learning methods usually rely on specific network architectures or task-dependent sample acquisition algorithms. Moreover, when selecting a batc…

Cited by 18SourcePDFScholar
2022

BROS: A Pre-trained Language Model Focusing on Text and Layout for Better Key Information Extraction from Documents

AAAI 2022technical

Key information extraction (KIE) from document images requires understanding the contextual and spatial semantics of texts in two-dimensional (2D) space. Many recent studies try to solve the task by developing pre-trained language models focusing on combining visual features from document images wit…

2022

Emp-RFT: Empathetic Response Generation via Recognizing Feature Transitions between Utterances

NAACL 2022long

Each utterance in multi-turn empathetic dialogues has features such as emotion, keywords, and utterance-level meaning. Feature transitions between utterances occur naturally. However, existing approaches fail to perceive the transitions because they extract features for the context at the coarse-gra…

2022

Ring-pull Type Soft Wearable Robotic Glove for Hand Strength Assistance

RA-L 2022

This letter proposes and verifies a new method, the ring-pull mechanism, to overcome the disadvantages of existing wearable robotic gloves. By attaching a ring to the metacarpopha-langeal joint of the finger, the ring-pull mechanism supplements the grasping force of the user, while reducing the weig

Cited by 5SourceScholar
2021

CDS: Cross-Domain Self-Supervised Pre-Training

ICCV 2021poster

We present a two-stage pre-training approach that improves the generalization ability of standard single-domain pre-training. While standard pre-training on a single large dataset (such as ImageNet) can provide a good initial representation for transfer learning tasks, this approach may result in bi…

Cited by 57PDFScholar
2021

Knowledge Transfer across Imaging Modalities Via Simultaneous Learning of Adaptive Autoencoders for High-Fidelity Mobile Robot Vision

IROS 2021poster

Enabling mobile robots for solving challenging and diverse shape, texture, and motion related tasks with high fidelity vision requires the integration of novel multimodal imaging sensors and advanced fusion techniques. However, it is associated with high cost, power, hardware modification, and compu…

Cited by 6SourceScholar
2021

Learning Cross-Modal Contrastive Features for Video Domain Adaptation

ICCV 2021poster

Learning transferable and domain adaptive feature representations from videos is important for video-relevant tasks such as action recognition. Existing video domain adaptation methods mainly rely on adversarial feature alignment, which has been derived from the RGB image space. However, video data…

Cited by 94PDFScholar
2021

OpenMatch: Open-Set Semi-supervised Learning with Open-set Consistency Regularization

NeurIPS 2021poster

Semi-supervised learning (SSL) is an effective means to leverage unlabeled data to improve a model’s performance. Typical SSL methods like FixMatch assume that labeled and unlabeled data share the same label space. However, in practice, unlabeled data can contain categories unseen in the labeled set…

2021

Tune It the Right Way: Unsupervised Validation of Domain Adaptation via Soft Neighborhood Density

ICCV 2021poster

Unsupervised domain adaptation (UDA) methods can dramatically improve generalization on unlabeled target domains. However, optimal hyper-parameter selection is critical to achieving high accuracy and avoiding negative transfer. Supervised hyper-parameter validation is not possible without labeled ta…

Cited by 75PDFcodeScholar
2020

Bi-Modal Hemispherical Sensors for Dynamic Locomotion and Manipulation

IROS 2020poster

The ability to measure multi-axis contact forces and contact surface normals in real time is critical to allow robots to improve their dexterous manipulation and locomotion abilities. This paper presents a new fingertip sensor for 3-axis contact force and contact location detection, as well as impro…

Cited by 15SourceScholar
2020

Learning to Scale Multilingual Representations for Vision-Language Tasks

ECCV 2020poster

Current multilingual vision-language models either require a large number of additional parameters for each supported language, or suffer performance degradation as languages are added. In this paper, we propose a Scalable Multilingual Aligned Language Representation (SMALR) that supports many langu…

Cited by 37SourcePDFScholar
2020

Robust Autonomous Navigation of a Small-Scale Quadruped Robot in Real-World Environments

IROS 2020poster

Animal-level agility and robustness in robots cannot be accomplished by solely relying on blind locomotion controllers. A significant portion of a robot’s ability to traverse terrain comes from reacting to the external world through visual sensing. However, embedding the sensors and compute that pro…

Cited by 45SourceScholar
2020

Universal Domain Adaptation through Self Supervision

NeurIPS 2020poster

Unsupervised domain adaptation methods traditionally assume that all source categories are present in the target domain. In practice, little may be known about the category overlap between the two domains. While some methods address target settings with either partial or open-set categories, they as…

2019

Bi-Modal Hemispherical Sensor: A Unifying Solution for Three Axis Force and Contact Angle Measurement

IROS 2019poster

In robotic tasks that require physical interactions such as manipulation and legged locomotion, it is important to simultaneously measure contact forces and contact angles. This paper presents a unified solution for simultaneously measuring three axis contact forces and contact angles for legged loc…

Cited by 15SourceScholar
2019

Semi-Supervised Domain Adaptation via Minimax Entropy

ICCV 2019poster

Contemporary domain adaptation methods are very effective at aligning feature distributions of source and target domains without any target supervision. However, we show that these techniques perform poorly when even a few labeled examples are available in the target domain. To address this semi-sup…

Cited by 848PDFScholar
2018

Excitation Backprop for RNNs

CVPR 2018poster

Deep models are state-of-the-art or many vision tasks including video action recognition and video captioning. Models are trained to caption or classify activity in videos, but little is known about the evidence used to make such decisions. Grounding decisions made by deep networks has been studied…

2018

Fast Kinodynamic Bipedal Locomotion Planning with Moving Obstacles

IROS 2018poster

In this paper, we present a sampling-based kino-dynamic planning framework for a bipedal robot in complex environments. Unlike other footstep planning algorithms which typically plan footstep locations and the biped dynamics in separate steps, we handle both simultaneously. Three primary advantages…

Cited by 5SourceScholar