← Search

Seungryong Kim

92 accepted papers

2026

3D Scene Prompting for Scene-Consistent Camera-Controllable Video Generation

ICLR 2026poster

We present 3DScenePrompt, a framework for camera-controllable video generation that maintains scene consistency when extending arbitrary-length input videos along user-specified trajectories. Unlike existing video generative methods limited to conditioning on a single image or just a few frames, we…

Cited by 0SourcecodeScholar
2026

4DP-QA: Scalable QA for 4D Perception in Vision Language Models

CVPR 2026

Despite recent advances, Vision Language Models (VLMs) still struggle to grasp the dynamics of the world. We note that the ability to reason about a 4D scene, challenging in itself, is further complicated by two factors. First, VLMs observe motion indirectly via its projection onto 2D images. Second

Cited by 0SourceScholar
2026

A Noise is Worth Diffusion Guidance

ICLR 2026poster

Diffusion models have demonstrated remarkable image generation capabilities, but their performance heavily relies on classifier-free guidance (CFG). While CFG significantly enhances image quality, evaluating both conditional and unconditional models at every denoising step leads to substantial compu…

Cited by 0SourcecodeScholar
2026

Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation

ICLR 2026poster

We introduce a diffusion-based framework that generates aligned novel view images and geometries via a warping‐and‐inpainting methodology. Unlike prior methods that require dense posed images or pose-embedded generative models limited to in‐domain views, our method leverages off‐the‐shelf geometry p…

Cited by 0SourcecodeScholar
2026

AnthroTAP: Learning Point Tracking with Real-World Motion

CVPR 2026

Point tracking models often struggle to generalize to real-world videos because large-scale training data is predominantly synthetic--the only source currently feasible to produce at scale. Collecting real-world annotations, however, is prohibitively expensive, as it requires tracking hundreds of po

Cited by 0SourcecodeScholar
2026

Attribute-Preserving Pseudo-Labeling for Diffusion-Based Face Swapping

CVPR 2026

Face swapping aims to transfer the identity of a source face onto a target face while preserving target-specific attributes such as pose, expression, lighting, skin tone, and makeup. However, since real ground truth for face swapping is unavailable, achieving both accurate identity transfer and high

Cited by 0SourceScholar
2026

CHIMERA: Controllable High-quality Image-Mask Extraction for Reliable Diffusion-based Anomaly Synthesis

AAAI 2026technical

We present CHIMERA, a novel framework for generating realistic, generalizable, and prompt-driven industrial anomalies from natural language instructions. Our method addresses two key challenges in text-guided anomaly synthesis: (1) the scarcity of scalable, high-quality paired anomaly data and (2) t

Cited by 0SourcePDFScholar
2026

Correspondence-Attention Alignment for Multi-View Diffusion Models

CVPR 2026

Multi-view diffusion models have recently emerged as a powerful paradigm for novel view synthesis, yet the underlying mechanism that enables their view consistency remains unclear. In this work, we first verify that the attention maps of these models acquire geometric correspondence throughout train

Cited by 0SourcecodeScholar
2026

Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression

ICML 2026poster

Recent advances in autoregressive video diffusion have enabled real-time frame streaming, yet existing solutions still suffer from temporal repetition, drift, and motion deceleration. We find that naïvely applying StreamingLLM-style attention sinks to video diffusion leads to fidelity degradation an…

Cited by 0SourceScholar
2026

Emergent Outlier View Rejection in Visual Geometry Grounded Transformers

CVPR 2026

Reliable 3D reconstruction from in-the-wild image collections is often hindered by noisy images--irrelevant inputs with little or no view overlap with others. While traditional Structure-from-Motion pipelines handle such cases through geometric verification and outlier rejection, feed-forward 3D rec

Cited by 0SourcecodeScholar
2026

Exploring Conditions for Diffusion Models in Robotic Control

CVPR 2026

While pre-trained visual representations have significantly advanced imitation learning, they are often task-agnostic as they remain frozen during policy learning. In this work, we explore leveraging pre-trained text-to-image diffusion models to obtain task-adaptive visual representations for roboti

Cited by 0SourceScholar
2026

Learning Compact 3D Representations from Feed-Forward Novel View Synthesis

CVPR 2026

Reconstructing and understanding 3D scenes from unposed sparse views in a feed-forward manner remains as a challenging task in 3D computer vision. Recent approaches use per-pixel 3D Gaussian Splatting for reconstruction, followed by a 2D-to-3D feature lifting stage for scene understanding. However,

Cited by 0SourcecodeScholar
2026

Lookahead Unmasking Elicits Reliable Decoding in Diffusion Language Models

ICML 2026poster

Masked Diffusion Models (MDMs) as language models generate by iteratively unmasking tokens, yet their performance crucially depends on the inference-time order of unmasking. Conventional methods such as confidence-based sampling are short-sighted, focusing on local optimization which neglects test-t…

Cited by 0SourceScholar
2026

MATRIX: Mask Track Alignment for Interaction-aware Video Generation

ICLR 2026poster

Video DiTs have advanced video generation, yet they still struggle to model multi-instance or subject-object interactions. This raises a key question: How do these models internally represent interactions? To answer this, we curate MATRIX-11K, a video dataset with interaction-aware captions and mul…

Cited by 0SourcecodeScholar
2026

MV-TAP: Tracking Any Point in Multi-View Videos

CVPR 2026

Multi-view camera systems enable rich observations of complex real-world scenes, and understanding dynamic objects in multi-view settings has become central to various applications. Point tracking serves as a key mechanism for capturing dynamic motion. However, conventional single-view approaches of

Cited by 0SourcecodeScholar
2026

TAG: Tangential Amplifying Guidance for Hallucination-Resistant Sampling

ICML 2026poster

Recent diffusion models achieve the state-of-the-art performance in image generation, but often suffer from semantic inconsistencies or *hallucinations*. While various inference-time guidance methods can enhance generation, they often operate *indirectly* by relying on external signals or architectu…

Cited by 0SourceScholar
2026

Text-Aware Image Restoration with Diffusion Models

ICLR 2026poster

While diffusion models have achieved remarkable success in natural image restoration, they often fail to faithfully recover textual regions, frequently producing plausible yet incorrect text-like patterns, a phenomenon we term text-image hallucination. To address this limitation, we propose Text-Awa…

Cited by 0SourcecodeScholar
2026

Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry

AAAI 2026technical

We introduce a novel framework for video camera trajectory editing, enabling the re-synthesis of monocular videos along user-defined camera paths. This task is challenging due to its ill-posed nature and the limited multi-view video data for training. Traditional reconstruction methods struggle with

Cited by 0SourcePDFScholar
2026

VideoMaMa: Mask-Guided Video Matting via Generative Prior

CVPR 2026

Generalizing video matting models to real-world videos remains a significant challenge due to the scarcity of labeled data. To address this, we present Video Mask-to-Matte Model VideoMaMa that converts coarse segmentation masks into pixel accurate alpha mattes, by leveraging pretrained video diffusi

Cited by 0SourcecodeScholar
2026

WaTeRFlow: Watermark Temporal Robustness via Flow Consistency

CVPR 2026

Image watermarking supports authenticity and provenance, yet many schemes are still easy to bypass with various distortions and powerful generative edits. Deep learning-based watermarking has improved robustness to diffusion-based image editing, but a gap remains when a watermarked image is converte

Cited by 0SourceScholar
2025

AM-Adapter: Appearance Matching Adapter for Exemplar-based Semantic Image Synthesis in-the-Wild

ICCV 2025poster

Exemplar-based semantic image synthesis generates images aligned with semantic content while preserving the appearance of an exemplar. Conventional structure-guidance models like ControlNet, are limited as they rely solely on text prompts to control appearance and cannot utilize exemplar images as i…

Cited by 0SourcePDFScholar
2025

Active Test-time Vision-Language Navigation

NeurIPS 2025poster

Vision-Language Navigation (VLN) policies trained on offline datasets often exhibit degraded task performance when deployed in unfamiliar navigation environments at test time, where agents are typically evaluated without access to external interaction or feedback. Entropy minimization has emerged as…

Cited by 0SourceScholar
2025

ControlFace: Harnessing Facial Parametric Control for Face Rigging

CVPR 2025poster

Manipulation of facial images to meet specific controls such as pose, expression, and lighting, also referred to as face rigging is a complex task in computer vision. Existing methods are limited by their reliance on image datasets, which necessitates individual-specific fine-tuning and limits their…

Cited by 0SourcePDFScholar
2025

Cross-View Completion Models are Zero-shot Correspondence Estimators

CVPR 2025highlight

In this work, we analyze new aspects of cross-view completion, mainly through the analogy of cross-view completion and traditional self-supervised correspondence learning algorithms. Based on our analysis, we reveal that the cross-attention map of Croco-v2, best reflects this correspondence informat…

Cited by 2SourcePDFScholar
2025

Emergent Temporal Correspondences from Video Diffusion Transformers

NeurIPS 2025poster

Recent advancements in video diffusion models based on Diffusion Transformers (DiTs) have achieved remarkable success in generating temporally coherent videos. Yet, a fundamental question persists: how do these models internally establish and represent temporal correspondences across frames? We int…

Cited by 0SourceScholar
2025

Enhancing 3D Reconstruction for Dynamic Scenes

NeurIPS 2025poster

In this work, we address the task of 3D reconstruction in dynamic scenes, where object motions frequently degrade the quality of previous 3D pointmap regression methods, such as DUSt3R, that are originally designed for static 3D scene reconstruction. Although these methods provide an elegant and pow…

Cited by 0SourceScholar
2025

Exploring Temporally-Aware Features for Point Tracking

CVPR 2025poster

Point tracking in videos is a fundamental task with applications in robotics, video editing, and more. While many vision tasks benefit from pre-trained feature backbones to improve generalizability, point tracking has primarily relied on simpler backbones trained from scratch on synthetic data, whic…

2025

First Attentions Last: Better Exploiting First Attentions for Efficient Parallel Training

NeurIPS 2025poster

As training billion-scale transformers becomes increasingly common, employing multiple distributed GPUs along with parallel training methods has become a standard practice. However, existing transformer designs suffer from significant communication overhead, especially in Tensor Parallelism (TP), wh…

Cited by 0SourceScholar
2025

Identity-preserving Distillation Sampling by Fixed-Point Iterator

CVPR 2025poster

Score distillation sampling (SDS) demonstrates a powerful capability for text-conditioned 2D image and 3D object generation by distilling the knowledge from learned score functions. However, SDS often suffers from blurriness caused by noisy gradients. When SDS meets the image editing, such degradati…

Cited by 0SourcePDFScholar
2025

MoDiTalker: Motion-Disentangled Diffusion Model for High-Fidelity Talking Head Generation

AAAI 2025technical

Conventional GAN-based models for talking head generation often suffer from limited quality and unstable training. Recent approaches based on diffusion models have attempted to address these limitations and improve fidelity. However, they still face challenges, such as intensive sampling times and d…

2025

Multi-Granularity Video Object Segmentation

AAAI 2025technical

Current benchmarks for video segmentation are limited to annotating only salient objects (i.e., foreground instances). Despite their impressive architectural designs, previous works trained on these benchmarks have struggled to adapt to realworld scenarios. Thus, developing a new video segmentation…

Cited by 0SourcePDFScholar
2025

PF3plat: Pose-Free Feed-Forward 3D Gaussian Splatting for Novel View Synthesis

ICML 2025poster

We consider the problem of novel view synthesis from unposed images in a single feed-forward. Our framework capitalizes on fast speed, scalability, and high-quality 3D reconstruction and view synthesis capabilities of 3DGS, where we further extend it to offer a practical solution that relaxes common…

Cited by 0SourcePDFScholar
2025

S4M: Boosting Semi-Supervised Instance Segmentation with SAM

ICCV 2025poster

Semi-supervised instance segmentation poses challenges due to limited labeled data, causing difficulties in accurately localizing distinct object instances. Current teacher-student frameworks still suffer from performance constraints due to unreliable pseudo-label quality stemming from limited label…

Cited by 0SourcePDFScholar
2025

Seg4Diff: Unveiling Open-Vocabulary Semantic Segmentation in Text-to-Image Diffusion Transformers

NeurIPS 2025poster

Text-to-image diffusion models excel at translating language prompts into photorealistic images by implicitly grounding textual concepts through their cross-modal attention mechanisms. Recent multi-modal diffusion transformers extend this by introducing joint self-attention over concatenated image a…

Cited by 0SourceScholar
2025

Visual Persona: Foundation Model for Full-Body Human Customization

CVPR 2025poster

We introduce Visual Persona, a foundation model for text-to-image full-body human customization that, given a single in-the-wild human image, generates diverse images of the individual guided by text descriptions. Unlike prior methods that focus solely on preserving facial identity, our approach cap…

Cited by 0SourcePDFScholar
2025

Where and How to Perturb: On the Design of Perturbation Guidance in Diffusion and Flow Models

NeurIPS 2025poster

Recent guidance methods in diffusion models steer reverse sampling by perturbing the model to construct an implicit weak model and guide generation away from it. Among these approaches, attention perturbation has demonstrated strong empirical performance in unconditional scenarios where classifier-f…

Cited by 0SourceScholar
2024

CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation

CVPR 2024highlight

Open-vocabulary semantic segmentation presents the challenge of labeling each pixel within an image based on a wide range of text descriptions. In this work we introduce a novel cost-based approach to adapt vision-language foundation models notably CLIP for the intricate task of semantic segmentatio…

2024

Context Enhanced Transformer for Single Image Object Detection in Video Data

AAAI 2024technical

With the increasing importance of video data in real-world applications, there is a rising need for efficient object detection methods that utilize temporal information. While existing video object detection (VOD) techniques employ various strategies to address this challenge, they typically depend…

2024

Diffusion Model for Dense Matching

ICLR 2024oral

The objective for establishing dense correspondence between paired images con- sists of two terms: a data term and a prior term. While conventional techniques focused on defining hand-designed prior terms, which are difficult to formulate, re- cent approaches have focused on learning the data term w…

2024

DreamMatcher: Appearance Matching Self-Attention for Semantically-Consistent Text-to-Image Personalization

CVPR 2024poster

The objective of text-to-image (T2I) personalization is to customize a diffusion model to a user-provided reference concept generating diverse images of the concept aligned with the target prompts. Conventional methods representing the reference concepts using unique text embeddings often fail to ac…

2024

Few-Shot Neural Radiance Fields under Unconstrained Illumination

AAAI 2024technical

In this paper, we introduce a new challenge for synthesizing novel view images in practical environments with limited input multi-view images and varying lighting conditions. Neural radiance fields (NeRF), one of the pioneering works for this task, demand an extensive set of multi-view images taken…

Cited by 2SourcePDFScholar
2024

FlowTrack: Revisiting Optical Flow for Long-Range Dense Tracking

CVPR 2024poster

In the domain of video tracking existing methods often grapple with a trade-off between spatial density and temporal range. Current approaches in dense optical flow estimators excel in providing spatially dense tracking but are limited to short temporal spans. Conversely recent advancements in long-…

Cited by 9SourcePDFScholar
2024

GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping

NeurIPS 2024poster

Generating novel views from a single image remains a challenging task due to the complexity of 3D scenes and the limited diversity in the existing multi-view datasets to train a model on. Recent research combining large-scale text-to-image (T2I) models with monocular depth estimation (MDE) has shown…

2024

Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D Generation

ICLR 2024poster

Text-to-3D generation has shown rapid progress in recent days with the advent of score distillation sampling (SDS), a methodology of using pretrained text-to-2D diffusion models to optimize a neural radiance field (NeRF) in a zero-shot setting. However, the lack of 3D awareness in the 2D diffusion m…

2024

MaskingDepth: Masked Consistency Regularization for Semi-Supervised Monocular Depth Estimation

IROS 2024poster

We propose MaskingDepth, a semi-supervised learning framework for monocular depth estimation. MaskingDepth is designed to enforce consistency between the depths obtained from strongly-augmented images and the pseudo-depths derived from weakly-augmented images, which enables mitigating the reliance o…

Cited by 0SourcecodeScholar
2024

Retrieval-Augmented Score Distillation for Text-to-3D Generation

ICML 2024poster

Text-to-3D generation has achieved significant success by incorporating powerful 2D diffusion models, but insufficient 3D prior knowledge also leads to the inconsistency of 3D geometry. Recently, since large-scale multi-view datasets have been released, fine-tuning the diffusion model on the multi-v…

2024

Towards Open-Vocabulary Semantic Segmentation Without Semantic Labels

NeurIPS 2024poster

Large-scale vision-language models like CLIP have demonstrated impressive open-vocabulary capabilities for image-level tasks, excelling in recognizing what objects are present. However, they struggle with pixel-level recognition tasks like semantic segmentation, which require understanding where the…

Cited by 2SourcePDFScholar
2024

Unifying Correspondence Pose and NeRF for Generalized Pose-Free Novel View Synthesis

CVPR 2024highlight

This work delves into the task of pose-free novel view synthesis from stereo pairs a challenging and pioneering task in 3D vision. Our innovative framework unlike any before seamlessly integrates 2D correspondence matching camera pose estimation and NeRF rendering fostering a synergistic enhancement…

Cited by 6SourcePDFScholar
2024

Unifying Feature and Cost Aggregation with Transformers for Semantic and Visual Correspondence

ICLR 2024poster

This paper introduces a Transformer-based integrative feature and cost aggregation network designed for dense matching tasks. In the context of dense matching, many works benefit from one of two forms of aggregation: feature aggregation, which pertains to the alignment of similar features, or cost a…

Cited by 7SourcePDFScholar
2023

Debiasing Scores and Prompts of 2D Diffusion for View-consistent Text-to-3D Generation

NeurIPS 2023poster

Existing score-distilling text-to-3D generation techniques, despite their considerable promise, often encounter the view inconsistency problem. One of the most notable issues is the Janus problem, where the most canonical view of an object (\textit{e.g}., face or head) appears in other views. In thi…

2023

DäRF: Boosting Radiance Fields from Sparse Input Views with Monocular Depth Adaptation

NeurIPS 2023poster

Neural radiance field (NeRF) shows powerful performance in novel view synthesis and 3D geometry reconstruction, but it suffers from critical performance degradation when the number of known viewpoints is drastically reduced. Existing works attempt to overcome this problem by employing external prior…

Cited by 18SourcePDFScholar
2023

GeCoNeRF: Few-shot Neural Radiance Fields via Geometric Consistency

ICML 2023poster

We present a novel framework to regularize Neural Radiance Field (NeRF) in a few-shot setting with a geometry-aware consistency regularization. The proposed approach leverages a rendered depth map at unobserved viewpoint to warp sparse input images to the unobserved viewpoint and impose them as pseu…

2023

Improving Sample Quality of Diffusion Models Using Self-Attention Guidance

ICCV 2023poster

Denoising diffusion models (DDMs) have attracted attention for their exceptional generation quality and diversity. This success is largely attributed to the use of class- or text-conditional diffusion guidance methods, such as classifier and classifier-free guidance. In this paper, we present a more…

Cited by 96PDFScholar
2023

LANIT: Language-Driven Image-to-Image Translation for Unlabeled Data

CVPR 2023poster

Existing techniques for image-to-image translation commonly have suffered from two critical problems: heavy reliance on per-sample domain annotation and/or inability to handle multiple attributes per image. Recent truly-unsupervised methods adopt clustering approaches to easily provide per-sample on…

2023

MIDMs: Matching Interleaved Diffusion Models for Exemplar-Based Image Translation

AAAI 2023technical

We present a novel method for exemplar-based image translation, called matching interleaved diffusion models (MIDMs). Most existing methods for this task were formulated as GAN-based matching-then-generation framework. However, in this framework, matching errors induced by the difficulty of semantic…

2023

PartMix: Regularization Strategy To Learn Part Discovery for Visible-Infrared Person Re-Identification

CVPR 2023poster

Modern data augmentation using a mixture-based technique can regularize the models from overfitting to the training data in various computer vision applications, but a proper data augmentation technique tailored for the part-based Visible-Infrared person Re-IDentification (VI-ReID) models remains un…

Cited by 93SourcePDFScholar
2022

Call for Customized Conversation: Customized Conversation Grounding Persona and Knowledge

AAAI 2022technical

Humans usually have conversations by making use of prior knowledge about a topic and background information of the people whom they are talking to. However, existing conversational agents and datasets do not consider such comprehensive information, and thus they have a limitation in generating the u…

2022

ConMatch: Semi-Supervised Learning with Confidence-Guided Consistency Regularization

ECCV 2022poster

"We present a novel semi-supervised learning framework that intelligently leverages the consistency regularization between the model’s predictions from two strongly-augmented views of an image, weighted by a confidence of pseudo-label, dubbed ConMatch. While the latest semi-supervised learning metho…

2022

Cost Aggregation with 4D Convolutional Swin Transformer for Few-Shot Segmentation

ECCV 2022poster

"We present a novel cost aggregation network, called Volumetric Aggregation with Transformers (VAT), for few-shot segmentation. The use of transformers can benefit correlation map aggregation through self-attention over a global receptive field. However, the tokenization of a correlation map for tra…

2022

Deep Translation Prior: Test-Time Training for Photorealistic Style Transfer

AAAI 2022technical

Recent techniques to solve photorealistic style transfer within deep convolutional neural networks (CNNs) generally require intensive training from large-scale datasets, thus having limited applicability and poor generalization ability to unseen images or styles. To overcome this, we propose a novel…

2022

InstaFormer: Instance-Aware Image-to-Image Translation With Transformer

CVPR 2022poster

We present a novel Transformer-based network architecture for instance-aware image-to-image translation, dubbed InstaFormer, to effectively integrate global- and instance-level information. By considering extracted content features from an image as tokens, our networks discover global consensus of c…

Cited by 65PDFcodeScholar
2022

Joint Learning of Feature Extraction and Cost Aggregation for Semantic Correspondence

ICASSP 2022accepted

Establishing dense correspondences across semantically similar images is one of the challenging tasks due to the significant intra-class variations and background clutters. To solve these problems, numerous methods have been proposed, focused on learning feature extractor or cost aggregation indepen…

Cited by 0SourceScholar
2022

Meta-confidence estimation for stereo matching

ICRA 2022poster

We propose a novel framework to estimate the confidence of a disparity map taking into account, for the first time, the uncertainty affecting the confidence estimation process itself. Conversely to other tasks such as disparity estimation, the uncertainty of confidence directly hints that the confid…

Cited by 2SourceScholar
2022

Neural Matching Fields: Implicit Representation of Matching Fields for Visual Correspondence

NeurIPS 2022accept

Existing pipelines of semantic correspondence commonly include extracting high-level semantic features for the invariance against intra-class variations and background clutters. This architecture, however, inevitably results in a low-resolution matching field that additionally requires an ad-hoc int…

2022

Semi-Supervised Learning of Semantic Correspondence With Pseudo-Labels

CVPR 2022poster

Establishing dense correspondences across semantically similar images remains a challenging task due to the significant intra-class variations and background clutters. Traditionally, a supervised loss was used for training the matching networks, which requires tremendous manually-labeled data, while…

Cited by 21PDFScholar
2022

Semi-Supervised Learning with Mutual Distillation for Monocular Depth Estimation

ICRA 2022poster

We propose a semi-supervised learning framework for monocular depth estimation. Compared to existing semi-supervised learning methods, which inherit limitations of both sparse supervised and unsupervised loss functions, we achieve the complementary advantages of both loss functions, by building two…

Cited by 17SourceScholar
2022

You Truly Understand What I Need : Intellectual and Friendly Dialog Agents grounding Persona and Knowledge

EMNLP 2022finding

To build a conversational agent that interacts fluently with humans, previous studies blend knowledge or personal profile into the pre-trained language model. However, the model that considers knowledge and persona at the same time is still limited, leading to hallucination and a passive way of usin…

2021

Adaptive Confidence Thresholding for Monocular Depth Estimation

ICCV 2021poster

Self-supervised monocular depth estimation has become an appealing solution to the lack of ground truth labels, but its reconstruction loss often produces over-smoothed results across object boundaries and is incapable of handling occlusion explicitly. In this paper, we propose a new approach to lev…

Cited by 35PDFcodeScholar
2021

CATs: Cost Aggregation Transformers for Visual Correspondence

NeurIPS 2021poster

We propose a novel cost aggregation network, called Cost Aggregation Transformers (CATs), to find dense correspondences between semantically similar images with additional challenges posed by large intra-class appearance and geometric variations. Cost aggregation is a highly important process in mat…

2021

Cross-Domain Grouping and Alignment for Domain Adaptive Semantic Segmentation

AAAI 2021technical

Existing techniques to adapt semantic segmentation networks across source and target domains within deep convolutional neural networks (CNNs) deal with all the samples from the two domains in a global or category-aware manner. They do not consider an inter-class variation within the target domain it…

2021

Learning Canonical 3D Object Representation for Fine-Grained Recognition

ICCV 2021poster

We propose a novel framework for fine-grained object recognition that learns to recover object variation in 3D space from a single image, trained on an image collection without using any ground-truth 3D annotation. We accomplish this by representing an object as a composition of 3D shape and its app…

Cited by 16PDFScholar
2021

Mining Better Samples for Contrastive Learning of Temporal Correspondence

CVPR 2021poster

We present a novel framework for contrastive learning of pixel-level representation using only unlabeled video. Without the need of ground-truth annotation, our method is capable of collecting well-defined positive correspondences by measuring their confidences and well-defined negative ones by appr…

Cited by 34PDFScholar
2021

RobustNet: Improving Domain Generalization in Urban-Scene Segmentation via Instance Selective Whitening

CVPR 2021poster

Enhancing the generalization capability of deep neural networks to unseen domains is crucial for safety-critical applications in the real world such as autonomous driving. To address this issue, this paper proposes a novel instance selective whitening loss to improve the robustness of the segmentati…

Cited by 343PDFcodeScholar
2020

Cylindrical Convolutional Networks for Joint Object Detection and Viewpoint Estimation

CVPR 2020poster

Existing techniques to encode spatial invariance within deep convolutional neural networks only model 2D transformation fields. This does not account for the fact that objects in a 2D space are a projection of 3D ones, and thus they have limited ability to severe object viewpoint changes. To overcom…

Cited by 19PDFScholar
2020

DUNIT: Detection-Based Unsupervised Image-to-Image Translation

CVPR 2020poster

Image-to-image translation has made great strides in recent years, with current techniques being able to handle unpaired training images and to account for the multi-modality of the translation problem. Despite this, most methods treat the image as a whole, which makes the results they produce for c…

Cited by 94PDFcodeScholar
2019

Joint Learning of Semantic Alignment and Object Landmark Detection

ICCV 2019poster

Convolutional neural networks (CNNs) based approaches for semantic alignment and object landmark detection have improved their performance significantly. Current efforts for the two tasks focus on addressing the lack of massive training data through weakly- or unsupervised learning frameworks. In th…

Cited by 21PDFScholar
2019

LAF-Net: Locally Adaptive Fusion Networks for Stereo Confidence Estimation

CVPR 2019oral

We present a novel method that estimates confidence map of an initial disparity by making full use of tri-modal input, including matching cost, disparity, and color image through deep networks. The proposed network, termed as Locally Adaptive Fusion Networks (LAF-Net), learns locally-varying attent…

Cited by 65PDFScholar
2018

PARN: Pyramidal Affine Regression Networks for Dense Semantic Correspondence

ECCV 2018poster

This paper presents a deep architecture for dense semantic correspondence, called pyramidal affine regression networks (PARN), that estimates locally-varying affine transformation fields across images. To deal with intra-class appearance and shape variations that commonly exist among different insta…

Cited by 69SourcePDFScholar
2018

Recurrent Transformer Networks for Semantic Correspondence

NeurIPS 2018spotlight

We present recurrent transformer networks (RTNs) for obtaining dense correspondences between semantically similar images. Our networks accomplish this through an iterative process of estimating spatial transformations between the input images and using these transformations to generate aligned convo…

Cited by 115SourcePDFScholar
2018

Spatiotemporal Attention Based Deep Neural Networks for Emotion Recognition

ICASSP 2018accepted

We propose a spatiotemporal attention based deep neural networks for dimensional emotion recognition in facial videos. To learn the spatiotemporal attention that selectively focuses on emotional sailient parts within facial videos, we formulate the spatiotemporal encoder-decoder network using Convol…

Cited by 0SourceScholar
2017

FCSS: Fully Convolutional Self-Similarity for Dense Semantic Correspondence

CVPR 2017poster

We present a descriptor, called fully convolutional self-similarity (FCSS), for dense semantic correspondence. To robustly match points among different instances within the same object class, we formulate FCSS using local self-similarity (LSS) within a fully convolutional network. In contrast to exi…

Cited by 176PDFScholar
2015

DASC: Dense Adaptive Self-Correlation Descriptor for Multi-Modal and Multi-Spectral Correspondence

CVPR 2015poster

Establishing dense visual correspondence between multiple images is a fundamental task in many applications of computer vision and computational photography. Classical approaches, which aim to estimate dense stereo and optical flow fields for images adjacent in viewpoint or in time, have been dramat…

Cited by 119SourcePDFScholar