← Search

Jia-Bin Huang

85 accepted papers

2026

Coupled Diffusion Sampling for Training-Free Multi-View Image Editing

CVPR 2026

Given a collection of multi-view images, we perform consistent multi-view editing with a training-free framework using pre-trained 2D editing models and a generative multi-view model. While 2D editing models can independently edit each image in a set of multi-view images of a 3D scene, they do not m

Cited by 0SourceScholar
2026

Generalizable Sparse-View 3D Reconstruction from Unconstrained Images

CVPR 2026

Reconstructing 3D scenes from sparse, unposed images remains challenging under real-world conditions with varying illumination and transient occlusions. Existing methods rely on scene-specific optimization with appearance embeddings or dynamic masks, requiring extensive per-scene training and failin

Cited by 0SourcecodeScholar
2026

Generative Video Motion Editing with 3D Point Tracks

CVPR 2026

Camera and object motions are central to a video's narrative. However, precisely editing these captured motions remains a significant challenge, especially under complex object movements. Current motion-controlled image-to-video (I2V) approaches often lack full-scene context for consistent video edi

Cited by 0SourceScholar
2026

SIMPACT: Simulation-Enabled Action Planning using Vision-Language Models

CVPR 2026

Vision-Language Models (VLMs) exhibit remarkable common-sense and semantic reasoning capabilities. However, they lack a grounded understanding of physical dynamics. This limitation arises from training VLMs on static internet-scale visual-language data that contain no causal interactions or action-c

Cited by 0SourceScholar
2026

TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videos

CVPR 2026

Learning new robot tasks on new platforms and in new scenes from only a handful of demonstrations remains challenging. While videos of other embodiments---humans and different robots---are abundant, differences in embodiment, camera, and environment hinder their direct use. We address the small-data

Cited by 0SourcecodeScholar
2026

UniVerse: A Unified Modulation Framework for Segmentation-Free, Disentangled Multi-Concept Personalization

CVPR 2026

Personalized visual understanding has advanced significantly, yet existing approaches struggle to localize and extract specific concepts when input images contain multiple objects. Many prior methods rely heavily on segmentation-based supervision or exhibit poor compositional generalization, limitin

Cited by 0SourcecodeScholar
2025

Bridging Diffusion Models and 3D Representations: A 3D Consistent Super-Resolution Framework

ICCV 2025poster

We propose 3D Super Resolution (3DSR), a novel 3D Gaussian-splatting-based super-resolution framework that leverages off-the-shelf diffusion-based 2D super-resolution models. 3DSR encourages 3D consistency across views via the use of an explicit 3D Gaussian-splatting-based scene representation. This…

Cited by 0SourcePDFScholar
2025

Generative Multiview Relighting for 3D Reconstruction under Extreme Illumination Variation

CVPR 2025highlight

Reconstructing the geometry and appearance of objects from photographs taken in different environments is difficult as the illumination and therefore the object appearance vary across captured images. This is particularly challenging for more specular objects whose appearance strongly depends on the…

Cited by 1SourcePDFScholar
2025

Generative Omnimatte: Learning to Decompose Video into Layers

CVPR 2025highlight

Given a video and a set of input object masks, an omnimatte method aims to decompose the video into semantically meaningful layers containing individual objects along with their associated effects, such as shadows and reflections.Existing omnimatte methods assume a static background or accurate pose…

Cited by 3SourcePDFScholar
2025

IRIS: Inverse Rendering of Indoor Scenes from Low Dynamic Range Images

CVPR 2025poster

Inverse rendering seeks to recover 3D geometry, surface material, and lighting from captured images, enabling advanced applications such as novel-view synthesis, relighting, and virtual object insertion. However, most existing techniques rely on high dynamic range (HDR) images as input, limiting acc…

Cited by 4SourcePDFScholar
2025

Imagine, Verify, Execute: Memory-guided Agentic Exploration with Vision-Language Models

CoRL 2025poster

Exploration is key for general-purpose robotic learning, particularly in open-ended environments where explicit guidance or task-specific feedback is limited. Vision-language models (VLMs), which can reason about object semantics, spatial relations, and potential outcomes, offer a promising foundati…

Cited by 0SourceScholar
2025

LIRM: Large Inverse Rendering Model for Progressive Reconstruction of Shape, Materials and View-dependent Radiance Fields

CVPR 2025poster

We present Large Inverse Rendering Model (LIRM), a transformer architecture that jointly reconstructs high-quality shape, materials, and radiance fields with view-dependent effects in less than a second. Our model builds upon the recent Large Reconstruction Models (LRMs) that achieve state-of-the-ar…

Cited by 0SourcePDFScholar
2025

MaDCoW: Marginal Distortion Correction for Wide-Angle Photography with Arbitrary Objects

CVPR 2025poster

We introduce MaDCoW, a method for correcting marginal distortion of arbitrary objects in wide-angle photography. People often use wide-angle photography to convey natural scenes--smartphones typically default to wide-angle photography--but depicting very wide-field-of-view scenes produces distorted…

Cited by 0SourcePDFScholar
2025

Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions

CVPR 2025poster

We explore how body shapes influence human motion synthesis, an aspect often overlooked in existing text-to-motion generation methods due to the ease of learning a homogenized, canonical body shape. However, this homogenization can distort the natural correlations between different body shapes and t…

Cited by 1SourcePDFScholar
2025

Textured Gaussians for Enhanced 3D Scene Appearance Modeling

CVPR 2025poster

3D Gaussian Splatting (3DGS) has recently emerged as a state-of-the-art 3D reconstruction and rendering technique due to its high-quality results and fast training and rendering time. However, pixels covered by the same Gaussian are always shaded in the same color up to a Gaussian falloff scaling fa…

Cited by 3SourcePDFScholar
2025

VideoGigaGAN: Towards Detail-rich Video Super-Resolution

CVPR 2025poster

Video super-resolution (VSR) models achieve temporal consistency but often produce blurrier results than their image-based counterparts due to limited generative capacity. This prompts the question: can we adapt a generative image upsampler for VSR while preserving temporal consistency? We introduce…

Cited by 17SourcePDFScholar
2024

Fast View Synthesis of Casual Videos with Soup-of-Planes

ECCV 2024poster

"Novel view synthesis from an in-the-wild video is difficult due to challenges like scene dynamics and lack of parallax. While existing methods have shown promising results with implicit neural radiance fields, they are slow to train and render. This paper revisits explicit video representations to…

2024

Flash-Splat: 3D Reflection Removal with Flash Cues and Gaussian Splats

ECCV 2024poster

"We introduce a simple yet effective approach for separating transmitted and reflected light. Our key insight is that the powerful novel view synthesis capabilities provided by modern inverse rendering methods (e.g., 3D Gaussian splatting) allow one to perform flash/no-flash reflection separation us…

Cited by 8SourcePDFScholar
2024

FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis

CVPR 2024highlight

Diffusion models have transformed the image-to-image (I2I) synthesis and are now permeating into videos. However the advancement of video-to-video (V2V) synthesis has been hampered by the challenge of maintaining temporal consistency across video frames. This paper proposes a consistent V2V synthesi…

Cited by 41SourcePDFScholar
2024

In-N-Out: Faithful 3D GAN Inversion with Volumetric Decomposition for Face Editing

CVPR 2024poster

3D-aware GANs offer new capabilities for view synthesis while preserving the editing functionalities of their 2D counterparts. GAN inversion is a crucial step that seeks the latent code to reconstruct input images or videos subsequently enabling diverse editing tasks through manipulation of this lat…

Cited by 3SourcePDFScholar
2024

LTM: Lightweight Textured Mesh Extraction and Refinement of Large Unbounded Scenes for Efficient Storage and Real-time Rendering

CVPR 2024poster

Advancements in neural signed distance fields (SDFs) have enabled modeling 3D surface geometry from a set of 2D images of real-world scenes. Baking neural SDFs can extract explicit mesh with appearance baked into texture maps as neural features. The baked meshes still have a large memory footprint a…

Cited by 7SourcePDFScholar
2024

On the Content Bias in Frechet Video Distance

CVPR 2024poster

Frechet Video Distance (FVD) a prominent metric for evaluating video generation models is known to conflict with human perception occasionally. In this paper we aim to explore the extent of FVD's bias toward frame quality over temporal realism and identify its sources. We first quantify the FVD's se…

Cited by 35SourcePDFScholar
2024

Rethinking Score Distillation as a Bridge Between Image Distributions

NeurIPS 2024poster

Score distillation sampling (SDS) has proven to be an important tool, enabling the use of large-scale diffusion priors for tasks operating in data-poor domains. Unfortunately, SDS has a number of characteristic artifacts that limit its utility in general-purpose applications. In this paper, we make…

Cited by 12SourcePDFScholar
2024

Taming Latent Diffusion Model for Neural Radiance Field Inpainting

ECCV 2024poster

"Neural Radiance Field (NeRF) is a representation for 3D reconstruction from multi-view images. Despite some recent work showing preliminary success in editing a reconstructed NeRF with diffusion prior, they remain struggling to synthesize reasonable geometry in completely uncovered regions. One maj…

Cited by 10SourcePDFScholar
2024

TextureDreamer: Image-Guided Texture Synthesis Through Geometry-Aware Diffusion

CVPR 2024poster

We present TextureDreamer a novel image-guided texture synthesis method to transfer relightable textures from a small number of input images (3 to 5) to target 3D shapes across arbitrary categories. Texture creation is a pivotal challenge in vision and graphics. Industrial companies hire experienced…

2023

3D Motion Magnification: Visualizing Subtle Motions from Time-Varying Radiance Fields

ICCV 2023poster

Motion magnification helps us visualize subtle, imperceptible motion. However, prior methods only work for 2D videos captured with a fixed camera. We present a 3D motion magnification method that can magnify subtle motions from scenes captured by a moving camera, while supporting novel view renderin…

Cited by 7PDFScholar
2023

A Safety-Performance Metric Enabling Computational Awareness in Autonomous Robots

RA-L 2023

This letter takes a first step towards the analysis of safety <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">and</i> performance critical computational tasks for autonomous robots. Our contribution is a safety-performance (SP) metric that ensures sa

Cited by 5SourceScholar
2023

ClimateNeRF: Extreme Weather Synthesis in Neural Radiance Field

ICCV 2023poster

Physical simulations produce excellent predictions of weather effects. Neural radiance fields produce SOTA scene models. We describe a novel NeRF-editing procedure that can fuse physical simulations with NeRF models of scenes, producing realistic movies of physical phenomena in those scenes. Our app…

Cited by 32PDFScholar
2023

Consistent View Synthesis With Pose-Guided Diffusion Models

CVPR 2023poster

Novel view synthesis from a single image has been a cornerstone problem for many Virtual Reality applications that provide immersive experiences. However, most existing techniques can only synthesize novel views within a limited range of camera motion or fail to generate consistent and high-quality…

2023

DC2: Dual-Camera Defocus Control by Learning To Refocus

CVPR 2023poster

Smartphone cameras today are increasingly approaching the versatility and quality of professional cameras through a combination of hardware and software advancements. However, fixed aperture remains a key limitation, preventing users from controlling the depth of field (DoF) of captured images. At t…

Cited by 13SourcePDFScholar
2023

HyperReel: High-Fidelity 6-DoF Video With Ray-Conditioned Sampling

CVPR 2023highlight

Volumetric scene representations enable photorealistic view synthesis for static scenes and form the basis of several existing 6-DoF video techniques. However, the volume rendering procedures that drive these representations necessitate careful trade-offs in terms of quality, rendering speed, and me…

2023

Neural-PBIR Reconstruction of Shape, Material, and Illumination

ICCV 2023poster

Reconstructing the shape and spatially varying surface appearances of a physical-world object as well as its surrounding illumination based on 2D images (e.g., photographs) of the object has been a long-standing problem in computer vision and graphics. In this paper, we introduce an accurate and hig…

Cited by 30PDFcodeScholar
2023

OmnimatteRF: Robust Omnimatte with 3D Background Modeling

ICCV 2023poster

Video matting has broad applications, from adding interesting effects to casually captured movies to assisting video production professionals. Matting with associated effects such as shadows and reflections has also attracted increasing research activity, and methods like Omnimatte have been propos…

Cited by 7PDFcodeScholar
2023

Preserve Your Own Correlation: A Noise Prior for Video Diffusion Models

ICCV 2023poster

Despite tremendous progress in generating high-quality images using diffusion models, synthesizing a sequence of animated frames that are both photorealistic and temporally coherent is still in its infancy. While off-the-shelf billion-scale datasets for image generation are available, collecting sim…

Cited by 262PDFScholar
2023

Progressively Optimized Local Radiance Fields for Robust View Synthesis

CVPR 2023poster

We present an algorithm for reconstructing the radiance field of a large-scale scene from a single casually captured video. The task poses two core challenges. First, most existing radiance field reconstruction approaches rely on accurate pre-estimated camera poses from Structure-from-Motion algorit…

Cited by 106SourcePDFScholar
2023

Robust Dynamic Radiance Fields

CVPR 2023poster

Dynamic radiance field reconstruction methods aim to model the time-varying structure and appearance of a dynamic scene. Existing methods, however, assume that accurate camera poses can be reliably estimated by Structure from Motion (SfM) algorithms. These methods, thus, are unreliable as SfM algori…

2023

Shape-Aware Text-Driven Layered Video Editing

CVPR 2023poster

Temporal consistency is essential for video editing applications. Existing work on layered representation of videos allows propagating edits consistently to each frame. These methods, however, can only edit object appearance rather than object shape changes due to the limitation of using a fixed UV…

2022

Boosting View Synthesis With Residual Transfer

CVPR 2022poster

Volumetric view synthesis methods with neural representations, such as NeRF and NeX, have recently demonstrated high-quality novel view synthesis. Optimizing these representations is slow, however, and even fully trained models cannot reproduce all fine details in the input views. We present a simpl…

Cited by 5PDFcodeScholar
2022

Learning Instance-Specific Adaptation for Cross-Domain Segmentation

ECCV 2022poster

"We propose a test-time adaptation method for cross-domain image segmentation. Our method is simple: Given a new unseen instance at the test time, we adapt a pre-trained model by conducting instance-specific BatchNorm (statistics) calibration. Our approach has two core components. First, we replace…

Cited by 16SourcePDFScholar
2022

Learning Neural Light Fields With Ray-Space Embedding

CVPR 2022poster

Neural radiance fields (NeRFs) produce state-of-the-art view synthesis results, but are slow to render, requiring hundreds of network evaluations per pixel to approximate a volume rendering integral. Baking NeRFs into explicit data structures enables efficient rendering, but results in large memory…

Cited by 117PDFScholar
2022

Long Video Generation with Time-Agnostic VQGAN and Time-Sensitive Transformer

ECCV 2022poster

"Videos are created to express emotion, exchange information, and share experiences. Video synthesis has intrigued researchers for a long time. Despite the rapid progress driven by advances in visual synthesis, most existing studies focus on improving the frames’ quality and the transitions between…

2022

Neural Global Shutter: Learn To Restore Video From a Rolling Shutter Camera With Global Reset Feature

CVPR 2022poster

Most computer vision systems assume distortion-free images as inputs. The widely used rolling-shutter (RS) image sensors, however, suffer from geometric distortion when the camera and object undergo motion during capture. Extensive researches have been conducted on correcting RS distortions. However…

Cited by 15PDFcodeScholar
2021

DropLoss for Long-Tail Instance Segmentation

AAAI 2021technical

Long-tailed class distributions are prevalent among the practical applications of object detection and instance segmentation. Prior work in long-tail instance segmentation addresses the imbalance of losses between rare and frequent categories by reducing the penalty for a model incorrectly predictin…

2021

Hybrid Neural Fusion for Full-Frame Video Stabilization

ICCV 2021poster

Existing video stabilization methods often generate visible distortion or require aggressive cropping of frame boundaries, resulting in smaller field of views. In this work, we present a frame synthesis algorithm to achieve full-frame video stabilization. We first estimate dense warp fields from nei…

Cited by 59PDFcodeScholar
2021

PseudoSeg: Designing Pseudo Labels for Semantic Segmentation

ICLR 2021poster

Recent advances in semi-supervised learning (SSL) demonstrate that a combination of consistency regularization and pseudo-labeling can effectively improve image classification accuracy in the low-data regime. Compared to classification, semantic segmentation tasks require much more intensive labelin…

2020

3D Photography Using Context-Aware Layered Depth Inpainting

CVPR 2020poster

We propose a method for converting a single RGB-D input image into a 3D photo, i.e., a multi-layer representation for novel view synthesis that contains hallucinated color and depth structures in regions occluded in the original view. We use a Layered Depth Image with explicit pixel connectivity as…

Cited by 349PDFcodeScholar
2020

Cross-Domain Few-Shot Classification via Learned Feature-Wise Transformation

ICLR 2020spotlight

Few-shot classification aims to recognize novel categories with only few labeled images in each class. Existing metric-based few-shot classification algorithms predict categories by comparing the feature embeddings of query images with those from a few labeled images (support examples) using a learn…

Cited by 524SourcecodeScholar
2020

DRG: Dual Relation Graph for Human-Object Interaction Detection

ECCV 2020poster

We tackle the challenging problem of human-object interaction (HOI) detection. Existing methods either recognize the interaction of each human-object pair in isolation or perform joint inference based on complex appearance-based features. In this paper, we leverage an abstract spatial-semantic repre…

2020

FeatMatch: Feature-Based Augmentation for Semi-Supervised Learning

ECCV 2020poster

Recent state-of-the-art semi-supervised learning (SSL) methods use a combination of image-based transformations and consistency regularization as core components. Such methods, however, are limited to simple transformations such as traditional data augmentation or convex combinations of two images.…

Cited by 162SourcePDFScholar
2020

Learning Monocular Visual Odometry via Self-Supervised Long-Term Modeling

ECCV 2020poster

Monocular visual odometry (VO) suffers severely from error accumulation during frame-to-frame pose estimation. In this paper, we present a self-supervised learning method for VO with special consideration for consistency over longer sequences. To this end, we model the long-term dependency in pose p…

2020

NAS-DIP: Learning Deep Image Prior with Neural Architecture Search

ECCV 2020poster

Recent work has shown that the structure of deep convolutional neural networks can be used as a structured image prior for solving various inverse image restoration tasks. Instead of using hand-designed architectures, we propose to search for neural architectures that capture stronger image priors.…

2020

Single-Image HDR Reconstruction by Learning to Reverse the Camera Pipeline

CVPR 2020poster

Recovering a high dynamic range (HDR) image from a single low dynamic range (LDR) input image is challenging due to missing details in under-/over-exposed regions caused by quantization and saturation of camera sensors. In contrast to existing learning-based methods, our core idea is to incorporate…

Cited by 307PDFcodeScholar
2019

A Closer Look at Few-shot Classification

ICLR 2019poster

Few-shot classification aims to learn a classifier to recognize unseen classes during training with limited labeled examples. While significant progress has been made, the growing complexity of network designs, meta-learning algorithms, and differences in implementation details make a fair comparison d…

2019

CrDoCo: Pixel-Level Domain Transfer With Cross-Domain Consistency

CVPR 2019poster

Unsupervised domain adaptation algorithms aim to transfer the knowledge learned from one domain to another (e.g., synthetic to real images). The adapted representations often do not capture pixel-level domain shifts that are crucial for dense prediction tasks (e.g., semantic segmentation). In this p…

Cited by 378PDFScholar
2019

SAIL-VOS: Semantic Amodal Instance Level Video Object Segmentation - A Synthetic Dataset and Baselines

CVPR 2019poster

We introduce SAIL-VOS (Semantic Amodal Instance Level Video Object Segmentation), a new dataset aiming to stimulate semantic amodal segmentation research. Humans can effortlessly recognize partially occluded objects and reliably estimate their spatial extent beyond the visible. However, few modern c…

Cited by 113PDFScholar
2019

Why Can't I Dance in the Mall? Learning to Mitigate Scene Bias in Action Recognition

NeurIPS 2019poster

Human activities often occur in specific scene contexts, e.g., playing basketball on a basketball court. Training a model using existing video datasets thus inevitably captures and leverages such bias (instead of using the actual discriminative cues). The learned representation may not generalize we…

2018

DF-Net: Unsupervised Joint Learning of Depth and Flow using Cross-Task Consistency

ECCV 2018poster

We present an unsupervised learning framework for simultaneously training single-view depth prediction and optical flow estimation models using unlabeled video sequences. Existing unsupervised methods often exploit brightness constancy and spatial smoothness priors to train depth or flow models. In…

2018

DeepMVS: Learning Multi-View Stereopsis

CVPR 2018poster

We present DeepMVS, a deep convolutional neural network (ConvNet) for multi-view stereo reconstruction. Taking an arbitrary number of posed images as input, we first produce a set of plane-sweep volumes and use the proposed DeepMVS network to predict high-quality disparity maps. The key contribution…

2018

Diverse Image-to-Image Translation via Disentangled Representations

ECCV 2018poster

Image-to-image translation aims to learn the mapping between two visual domains. There are two main challenges for many applications: 1) the lack of aligned training pairs and 2) multiple possible outputs from a single input image. In this work, we present an approach based on disentangled represent…

2018

Learning Blind Video Temporal Consistency

ECCV 2018poster

Applying image processing algorithms independently to each frame of a video often leads to undesired inconsistent results over time. Developing temporally consistent video-based extensions, however, requires domain knowledge for individual tasks and is unable to generalize to other applications. In…

2018

Unsupervised Video Object Segmentation using Motion Saliency-Guided Spatio-Temporal Propagation

ECCV 2018poster

Unsupervised video segmentation plays an important role in a wide variety of applications from object identification to compression. However, to date, fast motion, motion blur and occlusions pose significant challenges. To address these challenges for unsupervised video segmentation, we develop a no…

Cited by 125SourcePDFScholar
2017

Deep Laplacian Pyramid Networks for Fast and Accurate Super-Resolution

CVPR 2017poster

Convolutional neural networks have recently demonstrated high-quality reconstruction for single-image super-resolution. In this paper, we propose the Laplacian Pyramid Super-Resolution Network (LapSRN) to progressively reconstruct the sub-band residuals of high-resolution images. At each pyramid lev…

Cited by 3356PDFScholar
2017

Semi-Supervised Learning for Optical Flow with Generative Adversarial Networks

NeurIPS 2017poster

Convolutional neural networks (CNNs) have recently been applied to the optical flow estimation problem. As training the CNNs requires sufficiently large ground truth training data, existing approaches resort to synthetic, unrealistic datasets. On the other hand, unsupervised methods are capable of l…

Cited by 134SourcePDFScholar
2017

Unsupervised Representation Learning by Sorting Sequences

ICCV 2017poster

We present an unsupervised representation learning approach using videos without semantic labels. We leverage the temporal coherence as a supervisory signal by formulating representation learning as a sequence sorting task. We take temporally shuffled frames (i.e. in non-chronological order) as inpu…

Cited by 570PDFcodeScholar
2016

A Comparative Study for Single Image Blind Deblurring

CVPR 2016spotlight

Numerous single image blind deblurring algorithms have been proposed to restore latent sharp images under camera motion. However, these algorithms are mainly evaluated using either synthetic datasets or few selected real blurred images. It is thus unclear how these algorithms would perform on images…

Cited by 524PDFScholar
2016

Weakly Supervised Object Localization With Progressive Domain Adaptation

CVPR 2016poster

We address the problem of weakly supervised object localization where only image-level annotations are available for training. Many existing approaches tackle this problem through object proposal mining. However, a substantial amount of noise in object proposals causes ambiguities for learning discr…

Cited by 257PDFScholar