← Search

Junsik Kim

17 accepted papers

2026

ELiC: Efficient LiDAR Geometry Compression via Cross-Bit-depth Feature Propagation and Bag-of-Encoders

CVPR 2026

Hierarchical LiDAR geometry compression encodes voxel occupancies from low to high bit-depths, yet prior methods treat each depth independently and re-estimate local context from coordinates at every level, limiting compression efficiency. We present ELiC, a real-time framework that combines cross-b

Cited by 0SourcecodeScholar
2026

Gravity Compensation Strategy of Space Robotic Arm for On-Ground Testing Using Gyroscope-Type Coupling Interface and Cable-Driven Parallel Robot

RA-L 2026

Space robotic arms are designed to operate in microgravity environments, so their own weight makes them difficult to move on Earth's 1g surface. A cable-driven parallel robot (CDPR) can handle partially the arm's self-weight, enabling operation on the ground. The CDPR connects to a specific point on

Cited by 0SourceScholar
2026

Virtual Multiplex Staining for Histological Images Using a Marker-Wise Conditioned Diffusion Model

AAAI 2026technical

Multiplex imaging is revolutionizing pathology by enabling the simultaneous visualization of multiple biomarkers within tissue samples, providing molecular-level insights that traditional hematoxylin and eosin (H&E) staining cannot provide. However, the complexity and cost of multiplex data acquisit

Cited by 0SourcePDFScholar
2025

RSCF: Relation-Semantics Consistent Filter for Entity Embedding of Knowledge Graph

ACL 2025long

In knowledge graph embedding, leveraging relation specific entity transformation has markedly enhanced performance. However, the consistency of embedding differences before and after transformation remains unaddressed, risking the loss of valuable inductive bias inherent in the embeddings. This inco…

Cited by 0SourcePDFScholar
2024

Joint-Task Regularization for Partially Labeled Multi-Task Learning

CVPR 2024poster

Multi-task learning has become increasingly popular in the machine learning field but its practicality is hindered by the need for large labeled datasets. Most multi-task learning methods depend on fully labeled datasets wherein each input example is accompanied by ground-truth labels for all target…

2024

R^2-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding

ECCV 2024poster

"Video temporal grounding (VTG) is a fine-grained video understanding problem that aims to ground relevant clips in untrimmed videos given natural language queries. Most existing VTG models are built upon frame-wise final-layer CLIP features, aided by additional temporal backbones (, SlowFast) with…

2023

Sound Source Localization is All about Cross-Modal Alignment

ICCV 2023poster

Humans can easily perceive the direction of sound sources in a visual scene, termed sound source localization. Recent studies on learning-based sound source localization have mainly explored the problem from a localization perspective. However, prior arts and existing benchmarks do not account for…

Cited by 18PDFScholar
2022

Camera-Tracklet-Aware Contrastive Learning for Unsupervised Vehicle Re-Identification

ICRA 2022poster

Recently, vehicle re-identification methods based on deep learning constitute remarkable achievement. However, this achievement requires large-scale and well-annotated datasets. In constructing the dataset, assigning globally available identities (Ids) to vehicles captured from a great number of cam…

Cited by 12SourcecodeScholar
2022

Learning Sound Localization Better from Semantically Similar Samples

ICASSP 2022accepted

The objective of this work is to localize the sound sources in visual scenes. Existing audio-visual works employ contrastive learning by assigning corresponding audio-visual pairs from the same source as positives while randomly mismatched pairs as negatives. However, these negative pairs may contai…

Cited by 0SourceScholar
2022

ML-BPM: Multi-Teacher Learning with Bidirectional Photometric Mixing for Open Compound Domain Adaptation in Semantic Segmentation

ECCV 2022poster

"Open compound domain adaptation (OCDA) considers the target domain as the compound of multiple unknown homogeneous subdomains. The goal of OCDA is to minimize the domain gap between the source domain and the compound target domain, which brings the benefit of the model generalization to the unseen…

Cited by 14SourcePDFScholar
2021

Motion-blurred Video Interpolation and Extrapolation

AAAI 2021technical

Abrupt motion of camera or objects in a scene result in a blurry video, and therefore recovering high quality video requires two types of enhancements: visual enhancement and temporal upsampling. A broad range of research attempted to recover clean frames from blurred image sequences or temporally u…

Cited by 20SourcePDFScholar
2021

Optical Flow Estimation from a Single Motion-blurred Image

AAAI 2021technical

In most of computer vision applications, motion blur is regarded as an undesirable artifact. However, it has been shown that motion blur in an image may have practical interests in fundamental computer vision problems. In this work, we propose a novel framework to estimate optical flow from a single…

Cited by 19SourcePDFScholar
2019

Variational Prototyping-Encoder: One-Shot Learning With Prototypical Images

CVPR 2019poster

In daily life, graphic symbols, such as traffic signs and brand logos, are ubiquitously utilized around us due to its intuitive expression beyond language boundary. We tackle an open-set graphic symbol recognition problem by one-shot classification with prototypical images as a single training examp…

Cited by 92PDFcodeScholar
2018

Learning to Localize Sound Source in Visual Scenes

CVPR 2018poster

Visual events are usually accompanied by sounds in our daily lives. We pose the question: Can the machine learn the correspondence between visual scene and the sound, and localize the sound source only by observing sound and visual scene pairs like human? In this paper, we propose a novel unsupervis…

Cited by 397SourcePDFScholar
2017

Pixel-Level Matching for Video Object Segmentation Using Convolutional Neural Networks

ICCV 2017poster

We propose a novel video object segmentation algorithm based on pixel-level matching using Convolutional Neural Networks (CNN). Our network aims to distinguish the target area from the background on the basis of the pixel-level similarity between two object units. The proposed network represents a t…

Cited by 219PDFScholar
2017

VPGNet: Vanishing Point Guided Network for Lane and Road Marking Detection and Recognition

ICCV 2017poster

In this paper, we propose a unified end-to-end trainable multi-task network that jointly handles lane and road marking detection and recognition that is guided by a vanishing point under adverse weather conditions. We tackle rainy and low illumination conditions, which have not been extensively stud…

Cited by 556PDFcodeScholar