← Search

Kuk-Jin Yoon

91 accepted papers

2026

Bootstrapping Video Semantic Segmentation Model via Distillation-assisted Test-Time Adaptation

CVPR 2026

Fully supervised Video Semantic Segmentation (VSS) relies heavily on densely annotated video data, limiting practical applicability. Alternatively, applying pre-trained Image Semantic Segmentation (ISS) models frame-by-frame avoids annotation costs but ignores crucial temporal coherence. Recent foun

Cited by 0SourcecodeScholar
2026

DSERT-RoLL: Robust Multi-Modal Perception for Diverse Driving Conditions with Stereo Event-RGB-Thermal Cameras, 4D Radar, and Dual-LiDAR

CVPR 2026

In this paper, we present DSERT-RoLL, a driving dataset that incorporates stereo event, RGB, and thermal cameras together with 4D radar and dual LiDAR, collected across diverse weather and illumination conditions. The dataset provides precise 2D and 3D bounding boxes with track IDs and ego vehicle o

Cited by 0SourcecodeScholar
2026

Event6D: Event-based Novel Object 6D Pose Tracking

CVPR 2026

Event cameras provide microsecond latency, making them suitable for 6D object pose tracking in fast, dynamic scenes where conventional RGB and depth pipelines suffer from motion blur and large pixel displacements. We introduce EventTrack6D, an event-depth tracking framework that generalizes to novel

Cited by 0SourcecodeScholar
2026

ExPose: Reinforcing Video Generation Models for Extreme Pose Estimation

CVPR 2026

Pose estimation remains challenging under sparse views, especially when visual overlap across images is extremely limited. Recent advances in video generation models offer a promising solution by enabling keyframe interpolation, which can enrich contextual cues and improve pose estimation performanc

Cited by 0SourcecodeScholar
2026

Improving Black-Box Generative Attacks via Generator Semantic Consistency

ICLR 2026poster

Transfer attacks optimize on a surrogate and deploy to a black-box target. While iterative optimization attacks in this paradigm are limited by their per-input cost limits efficiency and scalability due to multistep gradient updates for each input, generative attacks alleviate these by producing adv…

Cited by 0SourcecodeScholar
2026

Multimodal Distribution Matching for Vision-Language Dataset Distillation

CVPR 2026

Dataset distillation compresses large training sets into compact synthetic datasets while preserving downstream performance. As modern systems increasingly operate on paired vision-language inputs, multimodal distillation must preserve representation quality and cross-modal alignment under tight com

Cited by 0SourcecodeScholar
2026

Preference-Aligned LoRA Merging: Preserving Subspace Coverage and Addressing Directional Anisotropy

CVPR 2026

Merging multiple Low-Rank Adaptation (LoRA) modules is promising for constructing general-purpose systems, yet challenging because LoRA update directions span different subspaces and contribute unevenly. When merged naively, such mismatches can weaken the directions most critical to certain task los

Cited by 0SourcecodeScholar
2026

ReSplat: Degradation-agnostic Feed-forward Gaussian Splatting via Self-guided Residual Diffusion

ICLR 2026poster

Recent advances in novel view synthesis (NVS) have predominantly focused on ideal, clear input settings, limiting their applicability in real-world environments with common degradations such as blur, low-light, haze, rain, and snow. While some approaches address NVS under specific degradation types,…

Cited by 0SourceScholar
2026

Test-Time Training for LiDAR Semantic Segmentation under Corruption via Geometric Inlier Discrimination

CVPR 2026

LiDAR semantic segmentation must remain robust under various sensor and environmental corruptions to be reliable in safety-critical applications. Existing test-time adaptation methods, including approaches based on pseudo-labels and normalization statistics, have shown promising results but can stil

Cited by 0SourcecodeScholar
2025

Any6D: Model-free 6D Pose Estimation of Novel Objects

CVPR 2025poster

We introduce Any6D, a model-free framework for 6D object pose estimation that requires only a single RGB-D anchor image to estimate both the 6D pose and size of unknown objects in novel scenes. Unlike existing methods that rely on textured 3D models or multiple viewpoints, Any6D leverages a joint ob…

Cited by 0SourcePDFScholar
2025

DC-TTA: Divide-and-Conquer Framework for Test-Time Adaptation of Interactive Segmentation

ICCV 2025poster

Interactive segmentation (IS) allows users to iteratively refine object boundaries with minimal cues, such as positive and negative clicks. While the Segment Anything Model (SAM) has garnered attention in the IS community for its promptable segmentation capabilities, it often struggles in specialize…

2025

Doppler-Aware LiDAR-RADAR Fusion for Weather-Robust 3D Detection

ICCV 2025poster

Robust 3D object detection across diverse weather con- ditions is crucial for safe autonomous driving, and RADAR is increasingly leveraged for its resilience in adverse weather. Recent advancements have explored 4D RADAR and LiDAR-RADAR fusion to enhance 3D perception capabilities, specifically targ…

2025

Ev-3DOD: Pushing the Temporal Boundaries of 3D Object Detection with Event Cameras

CVPR 2025highlight

Detecting 3D objects in point clouds plays a crucial role in autonomous driving systems. Recently, advanced multi-modal methods incorporating camera information have achieved notable performance. For a safe and effective autonomous driving system, algorithms that excel not only in accuracy but also…

2025

Event-guided Unified Framework for Low-light Video Enhancement, Frame Interpolation, and Deblurring

ICCV 2025poster

In low-light environments, longer exposure times are commonly used to enhance image visibility; however, this inevitably leads to motion blur. Even with a long exposure time, videos captured in low-light environments still suffer from issues such as low visibility, low contrast, and color distortion…

Cited by 0SourcePDFScholar
2025

From Sharp to Blur: Unsupervised Domain Adaptation for 2D Human Pose Estimation Under Extreme Motion Blur Using Event Cameras

ICCV 2025poster

Human pose estimation is critical for applications such as rehabilitation, sports analytics, and AR/VR systems. However, rapid motion and low-light conditions often introduce motion blur, significantly degrading pose estimation due to the domain gap between sharp and blurred images. Most datasets as…

2025

Generative Active Learning for Long-tail Trajectory Prediction via Controllable Diffusion Model

ICCV 2025poster

While data-driven trajectory prediction has enhanced the reliability of autonomous driving systems, it still struggles with rarely observed long-tail scenarios. Prior works addressed this by modifying model architectures, such as using hypernetworks. In contrast, we propose refining the training pro…

Cited by 0SourcePDFScholar
2025

Interaction-Merged Motion Planning: Effectively Leveraging Diverse Motion Datasets for Robust Planning

ICCV 2025poster

Motion planning is a crucial component of autonomous robot driving. While various trajectory datasets exist, effectively utilizing them for a target domain remains challenging due to differences in agent interactions and environmental characteristics. Conventional approaches, such as domain adaptati…

Cited by 0SourcePDFScholar
2025

Learning Large Motion Estimation from Intermediate Representations with a High-Resolution Optical Flow Dataset Featuring Long-Range Dynamic Motion

ICCV 2025poster

With advancements in sensor and display technologies, high-resolution imagery is becoming increasingly prevalent in diverse applications. As a result, optical flow estimation needs to adapt to larger image resolutions, where even moderate movements lead to substantial pixel displacements, making lon…

2025

Multi-modal Knowledge Distillation-based Human Trajectory Forecasting

CVPR 2025poster

Pedestrian trajectory forecasting is crucial in various applications such as autonomous driving and mobile robot navigation. In such applications, camera-based perception enables the extraction of additional modalities (human pose, text) to enhance prediction accuracy. Indeed, we find that textual d…

2025

Non-differentiable Reward Optimization for Diffusion-based Autonomous Motion Planning

IROS 2025

Safe and effective motion planning is crucial for autonomous robots. Diffusion models excel at capturing complex agent interactions, a fundamental aspect of decision-making in dynamic environments. Recent studies have successfully applied diffusion models to motion planning, demonstrating their comp

Cited by 2SourceScholar
2025

Resolving Token-Space Gradient Conflicts: Token Space Manipulation for Transformer-Based Multi-Task Learning

ICCV 2025poster

Multi-Task Learning (MTL) enables multiple tasks to be learned within a shared network, but differences in objectives across tasks can cause negative transfer, where the learning of one task degrades another task's performance. While pre-trained transformers significantly improve MTL performance, th…

Cited by 0SourcePDFScholar
2025

Robust Adverse Weather Removal via Spectral-based Spatial Grouping

ICCV 2025poster

Adverse weather conditions cause diverse and complex degradation patterns, driving the development of All-in-One (AiO) models. However, recent AiO solutions still struggle to capture diverse degradations, since global filtering methods like direct operations on the frequency domain fail to handle hi…

2025

Synchronizing Task Behavior: Aligning Multiple Tasks during Test-Time Training

ICCV 2025poster

Generalizing neural networks to unseen target domains is a significant challenge in real-world deployments. Test-time training (TTT) addresses this by using an auxiliary self-supervised task to reduce the domain gap caused by distribution shifts between the source and target. However, we find that w…

Cited by 0SourcePDFScholar
2025

Unleashing the Temporal Potential of Stereo Event Cameras for Continuous-Time 3D Object Detection

ICCV 2025poster

3D object detection is essential for autonomous systems, enabling precise localization and dimension estimation. While LiDAR and RGB cameras are widely used, their fixed frame rates create perception gaps in high-speed scenarios. Event cameras, with their asynchronous nature and high temporal resolu…

2025

VR-Drive: Viewpoint-Robust End-to-End Driving with Feed-Forward 3D Gaussian Splatting

NeurIPS 2025poster

End-to-end autonomous driving (E2E-AD) has emerged as a promising paradigm that unifies perception, prediction, and planning into a holistic, data-driven framework. However, achieving robustness to varying camera viewpoints, a common real-world challenge due to diverse vehicle configurations, remain…

Cited by 0SourceScholar
2025

WarpHE4D: Dense 4D Head Map toward Full Head Reconstruction

ICCV 2025poster

We address the 3D head reconstruction problem and the facial correspondence search problem in a unified framework, named as WarpHE4D. The underlying idea is to establish correspondences between the facial image and the fixed UV texture map by exploiting powerful self-supervised visual representation…

2024

A Benchmark Dataset for Event-Guided Human Pose Estimation and Tracking in Extreme Conditions

NeurIPS 2024poster

Multi-person pose estimation and tracking have been actively researched by the computer vision community due to their practical applicability. However, existing human pose estimation and tracking datasets have only been successful in typical scenarios, such as those without motion blur or with well-…

2024

Class Tokens Infusion for Weakly Supervised Semantic Segmentation

CVPR 2024poster

Weakly Supervised Semantic Segmentation (WSSS) relies on Class Activation Maps (CAMs) to extract spatial information from image-level labels. With the success of Vision Transformer (ViT) the migration of ViT is actively conducted in WSSS. This work proposes a novel WSSS framework with Class Token In…

2024

FACL-Attack: Frequency-Aware Contrastive Learning for Transferable Adversarial Attacks

AAAI 2024technical

Deep neural networks are known to be vulnerable to security risks due to the inherent transferable nature of adversarial examples. Despite the success of recent generative model-based attacks demonstrating strong transferability, it still remains a challenge to design an efficient attack strategy in…

Cited by 6SourcePDFScholar
2024

From SAM to CAMs: Exploring Segment Anything Model for Weakly Supervised Semantic Segmentation

CVPR 2024poster

Weakly Supervised Semantic Segmentation (WSSS) aims to learn the concept of segmentation using image-level class labels. Recent WSSS works have shown promising results by using the Segment Anything Model (SAM) a foundation model for segmentation during the inference phase. However we observe that th…

2024

Improving Transferability for Cross-Domain Trajectory Prediction via Neural Stochastic Differential Equation

AAAI 2024technical

Multi-agent trajectory prediction is crucial for various practical applications, spurring the construction of many large-scale trajectory datasets, including vehicles and pedestrians. However, discrepancies exist among datasets due to external factors and data acquisition strategies. External facto…

2024

Multi-agent Long-term 3D Human Pose Forecasting via Interaction-aware Trajectory Conditioning

CVPR 2024highlight

Human pose forecasting garners attention for its diverse applications. However challenges in modeling the multi-modal nature of human motion and intricate interactions among agents persist particularly with longer timescales and more agents. In this paper we propose an interaction-aware trajectory-c…

2024

T4P: Test-Time Training of Trajectory Prediction via Masked Autoencoder and Actor-specific Token Memory

CVPR 2024poster

Trajectory prediction is a challenging problem that requires considering interactions among multiple actors and the surrounding environment. While data-driven approaches have been used to address this complex problem they suffer from unreliable predictions under distribution shifts during test time.…

2024

TALoS: Enhancing Semantic Scene Completion via Test-time Adaptation on the Line of Sight

NeurIPS 2024poster

Semantic Scene Completion (SSC) aims to perform geometric completion and semantic segmentation simultaneously. Despite the promising results achieved by existing studies, the inherently ill-posed nature of the task presents significant challenges in diverse driving scenarios. This paper introduces T…

2024

TTA-EVF: Test-Time Adaptation for Event-based Video Frame Interpolation via Reliable Pixel and Sample Estimation

CVPR 2024poster

Video Frame Interpolation (VFI) which aims at generating high-frame-rate videos from low-frame-rate inputs is a highly challenging task. The emergence of bio-inspired sensors known as event cameras which boast microsecond-level temporal resolution has ushered in a transformative era for VFI. Nonethe…

2024

Towards Robust 3D Object Detection with LiDAR and 4D Radar Fusion in Various Weather Conditions

CVPR 2024poster

Detecting objects in 3D under various (normal and adverse) weather conditions is essential for safe autonomous driving systems. Recent approaches have focused on employing weather-insensitive 4D radar sensors and leveraging them with other modalities such as LiDAR. However they fuse multi-modal info…

2024

Weakly Supervised Point Cloud Semantic Segmentation via Artificial Oracle

CVPR 2024poster

Manual annotation of every point in a point cloud is a costly and labor-intensive process. While weakly supervised point cloud semantic segmentation (WSPCSS) with sparse annotation shows promise the limited information from initial sparse labels can place an upper bound on performance. As a new rese…

2023

Cross-Guided Optimization of Radiance Fields With Multi-View Image Super-Resolution for High-Resolution Novel View Synthesis

CVPR 2023poster

Novel View Synthesis (NVS) aims at synthesizing an image from an arbitrary viewpoint using multi-view images and camera poses. Among the methods for NVS, Neural Radiance Fields (NeRF) is capable of NVS for an arbitrary resolution as it learns a continuous volumetric representation. However, radiance…

Cited by 14SourcePDFScholar
2023

Event-Based Video Frame Interpolation With Cross-Modal Asymmetric Bidirectional Motion Fields

CVPR 2023highlight

Video Frame Interpolation (VFI) aims to generate intermediate video frames between consecutive input frames. Since the event cameras are bio-inspired sensors that only encode brightness changes with a micro-second temporal resolution, several works utilized the event camera to enhance the performanc…

2023

Label-Free Event-based Object Recognition via Joint Learning with Image Reconstruction from Events

ICCV 2023oral

Recognizing objects from sparse and noisy events becomes extremely difficult when paired images and category labels do not exist. In this paper, we study label-free event-based object recognition where category labels and paired images are not available. To this end, we propose a joint formulation o…

Cited by 21PDFcodeScholar
2023

Learning Point Cloud Completion without Complete Point Clouds: A Pose-Aware Approach

ICCV 2023poster

Point cloud completion is to restore complete 3D scenes and objects from incomplete observations or limited sensor data. Existing fully-supervised methods rely on paired datasets of incomplete and complete point clouds, which are labor-intensive to obtain. Unpaired methods have been proposed, but st…

Cited by 7PDFScholar
2023

Leveraging Future Relationship Reasoning for Vehicle Trajectory Prediction

ICLR 2023poster

Understanding the interaction between multiple agents is crucial for realistic vehicle trajectory prediction. Existing methods have attempted to infer the interaction from the observed past trajectories of agents using pooling, attention, or graph-based methods, which rely on a deterministic approa…

Cited by 78SourcePDFScholar
2023

MATE: Masked Autoencoders are Online 3D Test-Time Learners

ICCV 2023poster

Our MATE is the first Test-Time-Training (TTT) method designed for 3D data, which makes deep networks trained for point cloud classification robust to distribution shifts occurring in test data. Like existing TTT methods from the 2D image domain, MATE also leverages test data for adaptation. Its tes…

Cited by 20PDFcodeScholar
2023

Pixel-Wise Warping for Deep Image Stitching

AAAI 2023technical

Existing image stitching approaches based on global or local homography estimation are not free from the parallax problem and suffer from undesired artifacts. In this paper, instead of relying on the homography-based warp, we propose a novel deep image stitching framework exploiting the pixel-wise w…

Cited by 11SourcePDFScholar
2023

Single Domain Generalization for LiDAR Semantic Segmentation

CVPR 2023poster

With the success of the 3D deep learning models, various perception technologies for autonomous driving have been developed in the LiDAR domain. While these models perform well in the trained source domain, they struggle in unseen domains with a domain gap. In this paper, we propose a single domain…

2023

TTA-COPE: Test-Time Adaptation for Category-Level Object Pose Estimation

CVPR 2023poster

Test-time adaptation methods have been gaining attention recently as a practical solution for addressing source-to-target domain gaps by gradually updating the model without requiring labels on the target data. In this paper, we propose a method of test-time adaptation for category-level object pose…

Cited by 39SourcePDFScholar
2023

Weakly Supervised Semantic Segmentation via Adversarial Learning of Classifier and Reconstructor

CVPR 2023poster

In Weakly Supervised Semantic Segmentation (WSSS), Class Activation Maps (CAMs) usually 1) do not cover the whole object and 2) be activated on irrelevant regions. To address the issues, we propose a novel WSSS framework via adversarial learning of a classifier and an image reconstructor. When an im…

2022

Adversarial Erasing Framework via Triplet with Gated Pyramid Pooling Layer for Weakly Supervised Semantic Segmentation

ECCV 2022poster

"Weakly supervised semantic segmentation (WSSS) has employed Class Activation Maps (CAMs) to localize the objects. However, the CAMs typically do not fit along the object boundaries and highlight only the most-discriminative regions. To resolve the problems, we propose a Gated Pyramid Pooling (GPP)…

2022

BIPS: Bi-modal Indoor Panorama Synthesis via Residual Depth-Aided Adversarial Learning

ECCV 2022poster

"Providing omnidirectional depth along with RGB information is important for numerous applications. However, as omnidirectional RGB-D data is not always available, synthesizing RGB-D panorama data from limited information of a scene can be useful. Therefore, some prior works tried to synthesize RGB…

2022

Facial Depth and Normal Estimation Using Single Dual-Pixel Camera

ECCV 2022poster

"Recently, Dual-Pixel (DP) sensors have been adopted in many imaging devices. However, despite their various advantages, DP sensors are used just for faster auto-focus and aesthetic image captures, and research on their usage for 3D facial understanding has been limited due to the lack of datasets a…

2022

MM-TTA: Multi-Modal Test-Time Adaptation for 3D Semantic Segmentation

CVPR 2022poster

Test-time adaptation approaches have recently emerged as a practical solution for handling domain shift without access to the source domain data. In this paper, we propose and explore a new multi-modal extension of test-time adaptation for 3D semantic segmentation. We find that, directly applying ex…

Cited by 87PDFScholar
2022

Multi-Source Domain Alignment for Domain Invariant Segmentation in Unknown Targets

IROS 2022poster

Semantic segmentation provides scene understanding capability by performing pixel-wise classification of objects within an image. However, the sensitivity of such algorithms towards domain changes requires fine-tuning using an annotated dataset for each novel domain, which is expensive to construct…

Cited by 1SourceScholar
2022

SphereSR: 360deg Image Super-Resolution With Arbitrary Projection via Continuous Spherical Image Representation

CVPR 2022oral

The 360deg imaging has recently gained much attention; however, its angular resolution is relatively lower than that of a narrow field-of-view (FOV) perspective image as it is captured using a fisheye lens with the same sensor size. Therefore, it is beneficial to super-resolve a 360deg image. Severa…

Cited by 58PDFScholar
2022

Stereo Depth From Events Cameras: Concentrate and Focus on the Future

CVPR 2022poster

Neuromorphic cameras or event cameras mimic human vision by reporting changes in the intensity in a scene, instead of reporting the whole scene at once in a form of an image frame as performed by conventional cameras. Events are streamed data that are often dense when either the scene changes or the…

Cited by 51PDFcodeScholar
2022

UDA-COPE: Unsupervised Domain Adaptation for Category-Level Object Pose Estimation

CVPR 2022poster

Learning to estimate object pose often requires ground-truth (GT) labels, such as CAD model and absolute-scale object pose, which is expensive and laborious to obtain in the real world. To tackle this problem, we propose an unsupervised domain adaptation (UDA) for category-level object pose estimati…

Cited by 43PDFScholar
2021

Adversarially-trained Hierarchical Feature Extractor for Vehicle Re-identification

ICRA 2021poster

Vehicle Re-identification (Re-ID) aims to retrieve all instances of query vehicle images present in an image pool. However viewpoint, illumination, and occlusion variations along with subtle differences between two unique images pose a significant challenge towards achieving an effective system. In…

Cited by 5SourcecodeScholar
2021

Dual Transfer Learning for Event-Based End-Task Prediction via Pluggable Event to Image Translation

ICCV 2021poster

Event cameras are novel sensors that perceive the per-pixel intensity changes and output asynchronous event streams with high dynamic range and less motion blur. It has been shown that events alone can be used for end-task learning, e.g., semantic segmentation, based on encoder-decoder-like networks…

Cited by 43PDFcodeScholar
2021

EvDistill: Asynchronous Events To End-Task Learning via Bidirectional Reconstruction-Guided Cross-Modal Knowledge Distillation

CVPR 2021poster

Event cameras sense per-pixel intensity changes and produce asynchronous event streams with high dynamic range and less motion blur, showing advantages over the conventional cameras. A hurdle of training event-based models is the lack of large qualitative labeled data. Prior works learning end-tasks…

Cited by 85PDFcodeScholar
2021

Improvement of Optical Flow Estimation by Using the Hampel Filter for Low-End Embedded Systems

RA-L 2021

Owing to the recent advances in the field of deep-learning-based approaches, state-of-the-art performance has been achieved for optical flow estimation. However, non-deep-learning-based improvement in the optical flow estimation performance is still required because many platforms, such as small UGV

Cited by 7SourceScholar
2021

Learning Icosahedral Spherical Probability Map Based on Bingham Mixture Model for Vanishing Point Estimation

ICCV 2021poster

Existing vanishing point (VP) estimation methods rely on pre-extracted image lines and/or prior knowledge of the number of VPs. However, in practice, this information may be insufficient or unavailable. To solve this problem, we propose a network that treats a perspective image as input and predicts…

Cited by 9PDFScholar
2021

Scanline Resolution-Invariant Depth Completion Using a Single Image and Sparse LiDAR Point Cloud

RA-L 2021

Most existing deep learning-based depth completion methods are only suitable for high (e.g. 64-scanline) resolution LiDAR measurements, and they usually fail to predict a reliable dense depth map with low resolution (4, 8, or 16-scanline) LiDAR. However, it is of great interest to reduce the number

Cited by 13SourceScholar
2021

Unlocking the Potential of Ordinary Classifier: Class-Specific Adversarial Erasing Framework for Weakly Supervised Semantic Segmentation

ICCV 2021poster

Weakly supervised semantic segmentation (WSSS) using image-level classification labels usually utilizes the Class Activation Maps (CAMs) to localize objects of interest in images. While pointing out that CAMs only highlight the most discriminative regions of the classes of interest, adversarial eras…

Cited by 161PDFcodeScholar
2020

Deceiving Image-to-Image Translation Networks for Autonomous Driving With Adversarial Perturbations

RA-L 2020

Deep neural networks (DNNs) have achieved impressive performance on handling computer vision problems. However, it has been found that DNNs are vulnerable to adversarial examples. For such reason, adversarial perturbations have been recently studied in several respects. However, most previous works

Cited by 29SourceScholar
2020

EventSR: From Asynchronous Events to Image Reconstruction, Restoration, and Super-Resolution via End-to-End Adversarial Learning

CVPR 2020poster

Event cameras sense intensity changes and have many advantages over conventional cameras. To take advantage of event cameras, some methods have been proposed to reconstruct intensity images from event streams. However, the outputs are still in low resolution (LR), noisy, and unrealistic. The low-qua…

Cited by 123PDFcodeScholar
2020

Loop-Net: Joint Unsupervised Disparity and Optical Flow Estimation of Stereo Videos With Spatiotemporal Loop Consistency

RA-L 2020

Most of existing deep learning-based depth and optical flow estimation methods require the supervision of a lot of ground truth data, and hardly generalize to video frames, resulting in temporal inconsistency. In this letter, we propose a joint framework that estimates disparity and optical flow of

Cited by 9SourceScholar
2019

Event-Based High Dynamic Range Image and Very High Frame Rate Video Generation Using Conditional Generative Adversarial Networks

CVPR 2019poster

Event cameras have a lot of advantages over traditional cameras, such as low latency, high temporal resolution, and high dynamic range. However, since the outputs of event cameras are the sequences of asynchronous events over time rather than actual intensity images, existing algorithms could not be…

Cited by 241PDFScholar
2019

SpherePHD: Applying CNNs on a Spherical PolyHeDron Representation of 360deg Images

CVPR 2019poster

Omni-directional cameras have many advantages overconventional cameras in that they have a much wider field-of-view (FOV). Accordingly, several approaches have beenproposed recently to apply convolutional neural networks(CNNs) to omni-directional images for various visual tasks.However, most of…

Cited by 130PDFScholar
2017

Automatic Content-Aware Projection for 360deg Videos

ICCV 2017poster

To watch 360 videos on normal 2D displays, we need to project the selected part of the 360 image onto the 2D display plane. In this paper, we propose a fully-automated framework for generating content-aware 2D normal-view perspective videos from 360 videos. Especially, we focus on the projection ste…

Cited by 1PDFScholar
2017

Joint Layout Estimation and Global Multi-View Registration for Indoor Reconstruction

ICCV 2017poster

In this paper, we propose an approach to jointly solve scene layout estimation and global registration problems for accurate indoor 3D reconstruction. Given a sequence of range data, we build a set of scene fragments using KinectFusion and register them through pose graph optimization. Afterwards, w…

Cited by 28PDFScholar
2016

Online Multi-Object Tracking via Structural Constraint Event Aggregation

CVPR 2016poster

Multi-object tracking (MOT) becomes more challenging when objects of interest have similar appearances. In that case, the motion cues are particularly useful for discriminating multiple objects. However, for online 2D MOT in scenes acquired from moving cameras, observable motion cues are complicated…

Cited by 226PDFScholar