← Search

Boxin Shi

132 accepted papers

2026

240FPS Stereo Vision from Monocular Mixed Spikes

CVPR 2026

Stereo vision is fundamental for enabling machines to perceive and interact with the world. While monocular stereo methods offer hardware compactness, they struggle with generalization due to reliance on data-driven priors. Binocular and multi-view systems improve accuracy but incur higher hardware

Cited by 0SourcecodeScholar
2026

AE2VID: Event-based Video Reconstruction via Aperture Modulation

CVPR 2026

Event-based video reconstruction seeks to recover high-speed, high-dynamic-range videos from event streams. While existing approaches rely exclusively on motion-triggered events, these events are inherently sparse and primarily capture dynamic regions. Therefore, they often suffer from error accumul

Cited by 0SourcecodeScholar
2026

Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner

CVPR 2026

Recent advancements in video generation highlight that realistic audio-visual synchronization is crucial for engaging content creation. However, existing video editing methods largely overlook audio-visual synchronization and lack the fine-grained spatial and temporal controllability required for pr

Cited by 0SourcecodeScholar
2026

HFR and HDR Video from Multi-Attenuated Spikes Using a Rapidly Rotating SpokeND Filter

CVPR 2026

Capturing scenes with both high dynamic range (HDR) and high-speed motion remains challenging for conventional cameras. Existing alternating-exposure approaches exacerbate temporal resolution loss, making them unsuitable for high-speed scenes. Consequently, current solutions typically compromise eit

Cited by 0SourceScholar
2026

Light of Normals: Unified Feature Representation for Universal Photometric Stereo

ICLR 2026poster

Universal photometric stereo (PS) is defined by two factors: it must (i) operate under arbitrary, unknown lighting conditions and (ii) avoid reliance on specific illumination models. Despite progress (e.g., SDM UniPS), two challenges remain. First, current encoders cannot guarantee that illumination…

Cited by 0SourcecodeScholar
2026

Lighting-grounded Video Generation with Renderer-based Agent Reasoning

CVPR 2026

Diffusion models have achieved remarkable progress in video generation, but their controllability remains a major limitation. Key scene factors such as layout, lighting, and camera trajectory are often entangled or only weakly modeled, restricting their applicability in domains like filmmaking and v

Cited by 0SourceScholar
2026

STAGE: Storyboard-Anchored Generation for Cinematic Multi-shot Narrative

CVPR 2026

While recent advancements in generative models have achieved remarkable visual fidelity in video synthesis, creating coherent multi-shot narratives remains a significant challenge. To address this, keyframe-based approaches have emerged as a promising alternative to computationally intensive end-to-

Cited by 0SourceScholar
2026

TRIDENT: A Trimodal Cascade Generative Framework for Drug and RNA-Conditioned Cellular Morphology Synthesis

CVPR 2026

Accurately modeling the relationship between perturbations, transcriptional responses, and phenotypic changes is essential for building an AI Virtual Cell (AIVC). However, existing methods typically constrained to modeling direct associations, such as *Perturbation -> RNA* or *Perturbation -> Morpho

Cited by 0SourceScholar
2026

Texvent: Asynchronous Event Data Simulation via Text Prompt

CVPR 2026

Current event simulation methods focus on employing videos to synthesize new event data, suffering from costly video capture and limited scalability across viewpoints, motions, and lighting. To this end, we propose a Text-to-event simulation framework (Texvent) that can directly generate asynchronou

Cited by 0SourcecodeScholar
2025

Active Hyperspectral Imaging Using an Event Camera

CVPR 2025highlight

Hyperspectral imaging plays a critical role in numerous scientific and industrial fields. Conventional hyperspectral imaging systems often struggle with the trade-off between capture speed, spectral resolution, and bandwidth, particularly in dynamic environments. In this work, we present a novel eve…

Cited by 0SourcePDFScholar
2025

AdaptiveAE: An Adaptive Exposure Strategy for HDR Capturing in Dynamic Scenes

ICCV 2025poster

Mainstream high dynamic range imaging techniques typically rely on fusing multiple images captured with different exposure setups (shutter speed and ISO). A good balance between shutter speed and ISO is crucial for achieving high-quality HDR, as high ISO values introduce significant noise, while lon…

Cited by 0SourcePDFScholar
2025

Asynchronous Event Error-Minimizing Noise for Safeguarding Event Dataset

ICCV 2025poster

With more event datasets being released online, safeguarding the event dataset against unauthorized usage has become a serious concern for data owners. Unlearnable Examples are proposed to prevent the unauthorized exploitation of image datasets. However, it's unclear how to create unlearnable asynch…

2025

Audio-Sync Video Generation with Multi-Stream Temporal Control

NeurIPS 2025poster

Audio is inherently temporal and closely synchronized with the visual world, making it a naturally aligned and expressive control signal for controllable video generation (e.g., movies). Beyond control, directly translating audio into video is essential for understanding and visualizing rich audio n…

Cited by 0SourceScholar
2025

BokehDiff: Neural Lens Blur with One-Step Diffusion

ICCV 2025poster

We introduce Bokehdiff, a novel lens blur rendering method that achieves physically accurate and visually appealing outcomes, with the help of generative diffusion prior. Previous methods are bounded by the accuracy of depth estimation, generating artifacts in depth discontinuities. Our method emplo…

2025

Dense Metric Depth Estimation via Event-based Differential Focus Volume Prompting

NeurIPS 2025poster

Dense metric depth estimation has witnessed great developments in recent years. While single-image-based methods have demonstrated commendable performance in certain circumstances, they may encounter challenges regarding scale ambiguities and visual illusions in real world. Traditional depth-from-fo…

Cited by 0SourcecodeScholar
2025

Event-guided HDR Reconstruction with Diffusion Priors

ICCV 2025poster

Events provide High Dynamic Range (HDR) intensity change which can guide Low Dynamic Range (LDR) image for HDR reconstruction. However, events only provide temporal intensity differences and it is still ill-posed in over-/under-exposed areas due to missing initial reference brightness and color info…

2025

EventPSR: Surface Normal and Reflectance Estimation from Photometric Stereo Using an Event Camera

CVPR 2025highlight

Simultaneously acquisition of the surface normal and reflectance parameters is a crucial but challenging technique in the field of computer vision and graphics. It requires capturing multiple high dynamic range (HDR) images in existing methods using frame-based cameras. In this paper, we propose Eve…

Cited by 0SourcePDFScholar
2025

EventUPS: Uncalibrated Photometric Stereo Using an Event Camera

ICCV 2025poster

We present EventUPS, the first uncalibrated photometric stereo (UPS) method using an event camera--a neuromorphic sensor that asynchronously detects brightness changes with microsecond resolution. Traditional frame-based UPS methods are hindered by high bandwidth demands and limited use in dynamic s…

Cited by 0SourcePDFScholar
2025

PIDSR: Complementary Polarized Image Demosaicing and Super-Resolution

CVPR 2025poster

Polarization cameras can capture multiple polarized images with different polarizer angles in a single shot, bringing convenience to polarization-based downstream tasks. However, their direct outputs are color-polarization filter array (CPFA) raw images, requiring demosaicing to reconstruct full-res…

2025

PanoWan: Lifting Diffusion Video Generation Models to 360$^\circ$ with Latitude/Longitude-aware Mechanisms

NeurIPS 2025poster

Panoramic video generation enables immersive 360$^\circ$ content creation, valuable in applications that demand scene-consistent world exploration. However, existing panoramic video generation models struggle to leverage pre-trained generative priors from conventional text-to-video models for high-q…

Cited by 0SourceScholar
2025

PhyS-EdiT: Physics-aware Semantic Image Editing with Text Description

CVPR 2025poster

Achieving joint control over material properties, lighting, and high-level semantics in images is essential for applications in digital media, advertising, and interactive design. Existing methods often isolate these properties, lacking a cohesive approach to manipulating materials, lighting, and se…

Cited by 0SourcePDFScholar
2025

PlaNet: Learning to Mitigate Atmospheric Turbulence in Planetary Images

AAAI 2025technical

Obtaining planetary images with good visual quality is not an easy task since they are usually degenerated by atmospheric turbulence during the imaging procedure. Existing atmospheric turbulence mitigation methods designed for conventional images cannot be applied to planetary images, since the obje…

Cited by 0SourcePDFScholar
2025

PolGS: Polarimetric Gaussian Splatting for Fast Reflective Surface Reconstruction

ICCV 2025poster

Efficient shape reconstruction for surfaces with complex reflectance properties is crucial for real-time virtual reality. While 3D Gaussian Splatting (3DGS)-based methods offer fast novel view rendering by leveraging their explicit surface representation, their reconstruction quality lags behind tha…

Cited by 0SourcePDFScholar
2025

PolarAnything: Diffusion-based Polarimetric Image Synthesis

ICCV 2025poster

Polarization images facilitate image enhancement and 3D reconstruction tasks, but the limited accessibility of polarization cameras hinders their broader application. This gap drives the need for synthesizing photorealistic polarization images. The existing polarization simulator Mitsuba relies on a…

Cited by 0SourcePDFScholar
2025

Polarimetric Neural Field via Unified Complex-Valued Wave Representation

ICCV 2025poster

Polarization has found applications in various computer vision tasks by providing additional physical cues. However, due to the limitations of current imaging systems, polarimetric parameters are typically stored in discrete form, which is non-differentiable and limits their applicability in polariz…

Cited by 0SourcePDFScholar
2025

SpecTRe-GS: Modeling Highly Specular Surfaces with Reflected Nearby Objects by Tracing Rays in 3D Gaussian Splatting

CVPR 2025highlight

3D Gaussian Splatting (3DGS), a recently emerged multi-view 3D reconstruction technique, has shown significant advantages in real-time rendering and explicit editing. However, 3DGS encounters challenges in the accurate modeling of both high-frequency view-dependent appearances and global illuminatio…

Cited by 0SourcePDFScholar
2025

SpikeDiff: Zero-shot High-Quality Video Reconstruction from Chromatic Spike Camera and Sub-millisecond Spike Streams

ICCV 2025poster

High-speed video reconstruction from neuromorphic spike cameras offers a promising alternative to traditional frame-based imaging, providing superior temporal resolution and dynamic range with reduced power consumption. Nevertheless, reconstructing high-quality colored videos from spikes captured in…

Cited by 0SourcePDFScholar
2025

Unified Reconstruction of Static and Dynamic Scenes from Events

CVPR 2025highlight

This paper addresses the challenge that current event-based video reconstruction methods cannot produce static background information. Recent research has uncovered the potential of event cameras in capturing static scenes. Nonetheless, image quality deteriorates due to noise interference and detail…

2025

V2V: Scaling Event-Based Vision through Efficient Video-to-Voxel Simulation

NeurIPS 2025poster

Event-based cameras offer unique advantages such as high temporal resolution, high dynamic range, and low power consumption. However, the massive storage requirements and I/O burdens of existing synthetic data generation pipelines and the scarcity of real data prevent event-based training datasets f…

Cited by 0SourcecodeScholar
2025

VIRES: Video Instance Repainting via Sketch and Text Guided Generation

CVPR 2025poster

We introduce VIRES, a video instance repainting method with sketch and text guidance, enabling video instance repainting, replacement, generation, and removal. Existing approaches struggle with temporal consistency and accurate alignment with the provided sketch sequence. VIRES leverages the generat…

Cited by 0SourcePDFScholar
2025

Zero-Shot Low-Light Image Enhancement via Latent Diffusion Models

AAAI 2025technical

Low-light image enhancement (LLIE) aims to improve visibility and signal-to-noise ratio in images captured under poor lighting conditions. While deep learning has shown promise in this domain, current approaches require extensive paired training data, limiting their practical utility. We present a n…

2024

Colorizing Monochromatic Radiance Fields

AAAI 2024technical

Though Neural Radiance Fields (NeRF) can produce colorful 3D representations of the world by using a set of 2D images, such ability becomes non-existent when only monochromatic images are provided. Since color is necessary in representing the world, reproducing color from monochromatic radiance fiel…

2024

Complementing Event Streams and RGB Frames for Hand Mesh Reconstruction

CVPR 2024poster

Reliable hand mesh reconstruction (HMR) from commonly-used color and depth sensors is challenging especially under scenarios with varied illuminations and fast motions. Event camera is a highly promising alternative for its high dynamic range and dense temporal resolution properties but it lacks key…

Cited by 8SourcePDFScholar
2024

DiLiGenRT: A Photometric Stereo Dataset with Quantified Roughness and Translucency

CVPR 2024poster

Photometric stereo faces challenges from non-Lambertian reflectance in real-world scenarios. Systematically measuring the reliability of photometric stereo methods in handling such complex reflectance necessitates a real-world dataset with quantitatively controlled reflectances. This paper introduce…

2024

EvDiG: Event-guided Direct and Global Components Separation

CVPR 2024poster

Separating the direct and global components of a scene aids in shape recovery and basic material understanding. Conventional methods capture multiple frames under high frequency illumination patterns or shadows requiring the scene to keep stationary during the image acquisition process. Single-frame…

Cited by 0SourcePDFScholar
2024

EventPS: Real-Time Photometric Stereo Using an Event Camera

CVPR 2024poster

Photometric stereo is a well-established technique to estimate the surface normal of an object. However the requirement of capturing multiple high dynamic range images under different illumination conditions limits the speed and real-time applications. This paper introduces EventPS a novel approach…

Cited by 12SourcePDFScholar
2024

Imaging Interiors: An Implicit Solution to Electromagnetic Inverse Scattering Problems

ECCV 2024poster

"Electromagnetic Inverse Scattering Problems (EISP) have gained wide applications in computational imaging. By solving EISP, the internal relative permittivity of the scatterer can be non-invasively determined based on the scattered electromagnetic fields. Despite previous efforts to address EISP, a…

2024

L-DiffER: Single Image Reflection Removal with Language-based Diffusion Model

ECCV 2024poster

"In this paper, we introduce L-DiffER, a language-based diffusion model designed for the ill-posed single image reflection removal task. Although having shown impressive performance for image generation, existing language-based diffusion models struggle with precise control and faithfulness in image…

Cited by 4SourcePDFScholar
2024

Latency Correction for Event-guided Deblurring and Frame Interpolation

CVPR 2024poster

Event cameras with their high temporal resolution dynamic range and low power consumption are particularly good at time-sensitive applications like deblurring and frame interpolation. However their performance is hindered by latency variability especially under low-light conditions and with fast-mov…

Cited by 9SourcePDFScholar
2024

NB-GTR: Narrow-Band Guided Turbulence Removal

CVPR 2024poster

The removal of atmospheric turbulence is crucial for long-distance imaging. Leveraging the stochastic nature of atmospheric turbulence numerous algorithms have been developed that employ multi-frame input to mitigate the turbulence. However when limited to a single frame existing algorithms face sub…

Cited by 2SourcePDFScholar
2024

NeRSP: Neural 3D Reconstruction for Reflective Objects with Sparse Polarized Images

CVPR 2024poster

We present NeRSP a Neural 3D reconstruction technique for Reflective surfaces with Sparse Polarized images. Reflective surface reconstruction is extremely challenging as specular reflections are view-dependent and thus violate the multiview consistency for multiview stereo. On the other hand sparse…

Cited by 7SourcePDFScholar
2024

Pano-NeRF: Synthesizing High Dynamic Range Novel Views with Geometry from Sparse Low Dynamic Range Panoramic Images

AAAI 2024technical

Panoramic imaging research on geometry recovery and High Dynamic Range (HDR) reconstruction becomes a trend with the development of Extended Reality (XR). Neural Radiance Fields (NeRF) provide a promising scene representation for both tasks without requiring extensive prior data. How- ever, in the c…

2024

Quality-Improved and Property-Preserved Polarimetric Imaging via Complementarily Fusing

NeurIPS 2024poster

Polarimetric imaging is a challenging problem in the field of polarization-based vision, since setting a short exposure time reduces the signal-to-noise ratio, making the degree of polarization (DoP) and the angle of polarization (AoP) severely degenerated, while if setting a relatively long exposur…

Cited by 0SourcePDFScholar
2024

Real-data-driven 2000 FPS Color Video from Mosaicked Chromatic Spikes

ECCV 2024poster

"The spike camera continuously records scene radiance with high-speed, high dynamic range, and low data redundancy properties, as a promising replacement for frame-based high-speed cameras. Previous methods for reconstructing color videos from monochromatic spikes are constrained in capturing full-t…

Cited by 1SourcePDFScholar
2024

Real-time 3D-aware Portrait Video Relighting

CVPR 2024highlight

Synthesizing realistic videos of talking faces under custom lighting conditions and viewing angles benefits various downstream applications like video conferencing. However most existing relighting methods are either time-consuming or unable to adjust the viewpoints. In this paper we present the fir…

2024

SfPUEL: Shape from Polarization under Unknown Environment Light

NeurIPS 2024poster

Shape from polarization (SfP) benefits from advancements like polarization cameras for single-shot normal estimation, but its performance heavily relies on light conditions. This paper proposes SfPUEL, an end-to-end SfP method to jointly estimate surface normal and material under unknown environment…

2024

Spatio-Temporal Interactive Learning for Efficient Image Reconstruction of Spiking Cameras

NeurIPS 2024poster

The spiking camera is an emerging neuromorphic vision sensor that records high-speed motion scenes by asynchronously firing continuous binary spike streams. Prevailing image reconstruction methods, generating intermediate frames from these spike streams, often rely on complex step-by-step network ar…

Cited by 1SourcePDFScholar
2024

Spin-UP: Spin Light for Natural Light Uncalibrated Photometric Stereo

CVPR 2024poster

Natural Light Uncalibrated Photometric Stereo (NaUPS) relieves the strict environment and light assumptions in classical Uncalibrated Photometric Stereo (UPS) methods. However due to the intrinsic ill-posedness and high-dimensional ambiguities addressing NaUPS is still an open question. Existing wor…

2024

Towards HDR and HFR Video from Rolling-Mixed-Bit Spikings

CVPR 2024poster

The spiking cameras offer the benefits of high dynamic range (HDR) high temporal resolution and low data redundancy. However reconstructing HDR videos in high-speed conditions using single-bit spikings presents challenges due to the limited bit depth. Increasing the bit depth of the spikings is adva…

Cited by 3SourcePDFScholar
2024

VMINer: Versatile Multi-view Inverse Rendering with Near- and Far-field Light Sources

CVPR 2024highlight

This paper introduces a versatile multi-view inverse rendering framework with near- and far-field light sources. Tackling the fundamental challenge of inherent ambiguity in inverse rendering our framework adopts a lightweight yet inclusive lighting model for different near- and far-field lights thus…

Cited by 0SourcePDFScholar
2024

Zero-Shot Event-Intensity Asymmetric Stereo via Visual Prompting from Image Domain

NeurIPS 2024poster

Event-intensity asymmetric stereo systems have emerged as a promising approach for robust 3D perception in dynamic and challenging environments by integrating event cameras with frame-based sensors in different views. However, existing methods often suffer from overfitting and poor generalization du…

Cited by 2SourcePDFScholar
2023

Affective Image Filter: Reflecting Emotions from Text to Images

ICCV 2023poster

Understanding the emotions in text and presenting them visually is a very challenging problem that requires a deep understanding of natural language and high-quality image synthesis simultaneously. In this work, we propose Affective Image Filter (AIF), a novel model that is able to understand the vi…

Cited by 14PDFScholar
2023

Coherent Event Guided Low-Light Video Enhancement

ICCV 2023poster

With frame-based cameras, capturing fast-moving scenes without suffering from blur often comes at the cost of low SNR and low contrast. Worse still, the photometric constancy that enhancement techniques heavily relied on is fragile for frames with short exposure. Event cameras can record brightness…

Cited by 32PDFcodeScholar
2023

Complementary Intrinsics From Neural Radiance Fields and CNNs for Outdoor Scene Relighting

CVPR 2023poster

Relighting an outdoor scene is challenging due to the diverse illuminations and salient cast shadows. Intrinsic image decomposition on outdoor photo collections could partly solve this problem by weakly supervised labels with albedo and normal consistency from multi-view stereo. With neural radiance…

Cited by 9SourcePDFScholar
2023

DANI-Net: Uncalibrated Photometric Stereo by Differentiable Shadow Handling, Anisotropic Reflectance Modeling, and Neural Inverse Rendering

CVPR 2023poster

Uncalibrated photometric stereo (UPS) is challenging due to the inherent ambiguity brought by the unknown light. Although the ambiguity is alleviated on non-Lambertian objects, the problem is still difficult to solve for more general objects with complex shapes introducing irregular shadows and gene…

2023

DiLiGenT-Pi: Photometric Stereo for Planar Surfaces with Rich Details - Benchmark Dataset and Beyond

ICCV 2023poster

Photometric stereo aims to recover detailed surface shapes from images captured under varying illuminations. However, existing real-world datasets primarily focus on evaluating photometric stereo for general non-Lambertian reflectances and feature bulgy shapes that have a certain height. As shape de…

Cited by 12PDFcodeScholar
2023

High-Fidelity Event-Radiance Recovery via Transient Event Frequency

CVPR 2023poster

High-fidelity radiance recovery plays a crucial role in scene information reconstruction and understanding. Conventional cameras suffer from limited sensitivity in dynamic range, bit depth, and spectral response, etc. In this paper, we propose to use event cameras with bio-inspired silicon sensors,…

2023

L-CAD: Language-based Colorization with Any-level Descriptions using Diffusion Priors

NeurIPS 2023spotlight

Language-based colorization produces plausible and visually pleasing colors under the guidance of user-friendly natural language descriptions. Previous methods implicitly assume that users provide comprehensive color descriptions for most of the objects in the image, which leads to suboptimal perfor…

2023

L-CoIns: Language-Based Colorization With Instance Awareness

CVPR 2023poster

Language-based colorization produces plausible colors consistent with the language description provided by the user. Recent studies introduce additional annotation to prevent color-object coupling and mismatch issues, but they still have difficulty in distinguishing instances corresponding to the sa…

Cited by 28SourcePDFScholar
2023

Learning Event Guided High Dynamic Range Video Reconstruction

CVPR 2023poster

Limited by the trade-off between frame rate and exposure time when capturing moving scenes with conventional cameras, frame based HDR video reconstruction suffers from scene-dependent exposure ratio balancing and ghosting artifacts. Event cameras provide an alternative visual representation with a m…

2023

LuminAIRe: Illumination-Aware Conditional Image Repainting for Lighting-Realistic Generation

NeurIPS 2023poster

We present the ilLumination-Aware conditional Image Repainting (LuminAIRe) task to address the unrealistic lighting effects in recent conditional image repainting (CIR) methods. The environment lighting and 3D geometry conditions are explicitly estimated from given background images and parsing mask…

Cited by 5SourcePDFScholar
2023

Non-Lambertian Multispectral Photometric Stereo via Spectral Reflectance Decomposition

IJCAI 2023poster

Multispectral photometric stereo (MPS) aims at recovering the surface normal of a scene from a single-shot multispectral image captured under multispectral illuminations. Existing MPS methods adopt the Lambertian reflectance model to make the problem tractable, but it greatly limits their applicatio…

Cited by 9SourcePDFScholar
2023

ReLeaPS : Reinforcement Learning-based Illumination Planning for Generalized Photometric Stereo

ICCV 2023poster

Illumination planning in photometric stereo aims to find a balance between tween surface normal estimation accuracy and image capturing efficiency by selecting optimal light configurations. It depends on factors such as the unknown shape and general reflectance of the target object, global illuminat…

Cited by 2PDFScholar
2023

Slow and Weak Attractor Computation Embedded in Fast and Strong E-I Balanced Neural Dynamics

NeurIPS 2023spotlight

Attractor networks require neuronal connections to be highly structured in order to maintain attractor states that represent information, while excitation and inhibition balanced networks (E-INNs) require neuronal connections to be random and sparse to generate irregular neuronal firings. Despite be…

Cited by 3SourcePDFScholar
2022

Data Association between Event Streams and Intensity Frames under Diverse Baselines

ECCV 2022poster

"This paper proposes a learning-based framework to associate event streams and intensity frames under diverse camera baselines, to simultaneously benefit to camera pose estimation under large baseline and depth estimation under small baseline. Based on the observation that event streams are globally…

Cited by 11SourcePDFScholar
2022

DiLiGenT102: A Photometric Stereo Benchmark Dataset With Controlled Shape and Material Variation

CVPR 2022poster

Evaluating photometric stereo using real-world dataset is important yet difficult. Existing datasets are insufficient due to their limited scale and random distributions in shape and material. This paper presents a new real-world photometric stereo dataset with "ground truth" normal maps, which is 1…

Cited by 32PDFcodeScholar
2022

Estimating Spatially-Varying Lighting in Urban Scenes with Disentangled Representation

ECCV 2022poster

"We present an end-to-end network for spatially-varying outdoor lighting estimation in urban scenes given a single limited field-of-view LDR image and any assigned 2D pixel position. We use three disentangled latent spaces learned by our network to represent sky light, sun light, and lighting-indepe…

Cited by 15SourcePDFScholar
2022

L-CoDe:Language-Based Colorization Using Color-Object Decoupled Conditions

AAAI 2022technical

Colorizing a grayscale image is inherently an ill-posed problem with multi-modal uncertainty. Language-based colorization offers a natural way of interaction to reduce such uncertainty via a user-provided caption. However, the color-object coupling and mismatch issues make the mapping from word to c…

Cited by 41SourcePDFScholar
2022

L-CoDer: Language-Based Colorization with Color-Object Decoupling Transformer

ECCV 2022poster

"Language-based colorization requires the colorized image to be consistent with the the user-provided language caption. A most recent work proposes to decouple the language into color and object conditions in solving the problem. Though decent progress has been made, its performance is limited by th…

2022

NEST: Neural Event Stack for Event-Based Image Enhancement

ECCV 2022poster

"Event cameras demonstrate unique characteristics such as high temporal resolution, low latency, and high dynamic range to improve performance for various image enhancement tasks. However, event streams cannot be applied to neural networks directly due to their sparse nature. To integrate events int…

2022

Real-Time Intermediate Flow Estimation for Video Frame Interpolation

ECCV 2022poster

"Real-time video frame interpolation (VFI) is very useful in video processing, media players, and display devices. We propose RIFE, a Real-time Intermediate Flow Estimation algorithm for VFI. To realize a high-quality flow-based VFI method, RIFE uses a neural network named IFNet that can estimate th…

2021

EvIntSR-Net: Event Guided Multiple Latent Frames Reconstruction and Super-Resolution

ICCV 2021poster

An event camera detects the scene radiance changes and sends a sequence of asynchronous event streams with high dynamic range, high temporal resolution, and low latency. However, the spatial resolution of event cameras is limited as a trade-off for these outstanding properties. To reconstruct high-r…

Cited by 54PDFScholar
2021

EventZoom: Learning To Denoise and Super Resolve Neuromorphic Events

CVPR 2021poster

We address the problem of jointly denoising and super resolving neuromorphic events, a novel visual signal that represents thresholded temporal gradients in a space-time window. The challenge for event signal processing is that they are asynchronously generated, and do not carry absolute intensity b…

Cited by 82PDFScholar
2021

High-Speed Image Reconstruction Through Short-Term Plasticity for Spiking Cameras

CVPR 2021poster

Fovea, located in the centre of the retina, is specialized for high-acuity vision. Mimicking the sampling mechanism of the fovea, a retina-inspired camera, named spiking camera, is developed to record the external information with a sampling rate of 40,000 Hz, and outputs asynchronous binary spike s…

Cited by 73PDFScholar
2021

Multispectral Photometric Stereo for Spatially-Varying Spectral Reflectances: A Well Posed Problem?

CVPR 2021poster

Multispectral photometric stereo (MPS) aims at recovering the surface normal of a scene from a single-shot multispectral image, which is known as an ill-posed problem. To make the problem well-posed, existing MPS methods rely on restrictive assumptions, such as shape prior, surfaces having a monochr…

Cited by 14PDFcodeScholar
2021

Normal Integration via Inverse Plane Fitting With Minimum Point-to-Plane Distance

CVPR 2021poster

This paper presents a surface normal integration method that solves an inverse problem of local plane fitting. Surface reconstruction from normal maps is essential in photometric shape reconstruction. To this end, we formulate normal integration in the camera coordinates and jointly solve for 3D poi…

Cited by 23PDFcodeScholar
2021

Single Image Reflection Removal With Absorption Effect

CVPR 2021poster

In this paper, we consider the absorption effect for the problem of single image reflection removal. We show that the absorption effect can be numerically approximated by the average of refractive amplitude coefficient map. We then reformulate the image formation model and propose a two-step solutio…

Cited by 54PDFcodeScholar
2020

AdderNet: Do We Really Need Multiplications in Deep Learning?

CVPR 2020oral

Compared with cheap addition operation, multiplication operation is of much higher computation complexity. The widely-used convolutions in deep neural networks are exactly cross-correlation to measure the similarity between input feature and convolution filters, which involves massive multiplication…

Cited by 286PDFcodeScholar
2020

CARS: Continuous Evolution for Efficient Neural Architecture Search

CVPR 2020poster

Searching techniques in most of existing neural architecture search (NAS) algorithms are mainly dominated by differentiable methods for the efficiency reason. In contrast, we develop an efficient continuous evolutionary approach for searching neural networks. Architectures in the population that sha…

Cited by 310PDFcodeScholar
2020

Conditional Image Repainting via Semantic Bridge and Piecewise Value Function

ECCV 2020poster

We study conditional image repainting where a model is trained to generate visual content conditioned on user inputs, and composite the generated content seamlessly onto a user provided image while preserving the semantics of users' inputs. The content generation community have been pursuing to lowe…

Cited by 6SourcePDFScholar
2020

DIST: Rendering Deep Implicit Signed Distance Function With Differentiable Sphere Tracing

CVPR 2020poster

We propose a differentiable sphere tracing algorithm to bridge the gap between inverse graphics methods and the recently proposed deep learning based implicit signed distance function. Due to the nature of the implicit function, the rendering process requires tremendous function queries, which is pa…

Cited by 350PDFcodeScholar
2020

Deep Shape from Polarization

ECCV 2020poster

This paper makes a first attempt to bring the Shape from Polarization (SfP) problem to the realm of deep learning. The previous state-of-the-art methods for SfP have been purely physics-based. We see value in these principled models, and blend these physical models as priors into a neural network ar…

2020

Frequency Domain Compact 3D Convolutional Neural Networks

CVPR 2020poster

This paper studies the compression and acceleration of 3-dimensional convolutional neural networks (3D CNNs). To reduce the memory cost and computational complexity of deep neural networks, a number of algorithms have been explored by discovering redundant parameters in pre-trained networks. However…

Cited by 32PDFScholar
2020

Group Contextual Encoding for 3D Point Clouds

NeurIPS 2020poster

Global context is crucial for 3D point cloud scene understanding tasks. In this work, we extended the contextual encoding layer that was originally designed for 2D tasks to 3D Point Cloud scenarios. The encoding layer learns a set of code words in the feature space of the 3D point cloud to characte…

2020

Joint Filtering of Intensity Images and Neuromorphic Events for High-Resolution Noise-Robust Imaging

CVPR 2020poster

We present a novel computational imaging system with high resolution and low noise. Our system consists of a traditional video camera which captures high-resolution intensity images, and an event camera which encodes high-speed motion as a stream of asynchronous binary events. To process the hybrid…

Cited by 110PDFScholar
2020

MISC: Multi-Condition Injection and Spatially-Adaptive Compositing for Conditional Person Image Synthesis

CVPR 2020poster

In this paper, we explore synthesizing person images with multiple conditions for various backgrounds. To this end, we propose a framework named "MISC" for conditional image generation and image compositing. For conditional image generation, we improve the existing condition injection mechanisms by…

Cited by 37PDFScholar
2020

Stereoscopic Flash and No-Flash Photography for Shape and Albedo Recovery

CVPR 2020poster

We present a minimal imaging setup that harnesses both geometric and photometric approaches for shape and albedo recovery. We adopt a stereo camera and a flashlight to capture a stereo image pair and a flash/no-flash pair. From the stereo image pair, we recover a rough shape that captures low-freque…

Cited by 12PDFScholar
2020

UnModNet: Learning to Unwrap a Modulo Image for High Dynamic Range Imaging

NeurIPS 2020poster

A conventional camera often suffers from over- or under-exposure when recording a real-world scene with a very high dynamic range (HDR). In contrast, a modulo camera with a Markov random field (MRF) based unwrapping algorithm can theoretically accomplish unbounded dynamic range but shows degenerate…

Cited by 18SourcePDFScholar
2020

What Does Plate Glass Reveal About Camera Calibration?

CVPR 2020poster

This paper aims to calibrate the orientation of glass and the field of view of the camera from a single reflection-contaminated image. We show how a reflective amplitude coefficient map can be used as a calibration cue. Different from existing methods, the proposed solution is free from image conten…

Cited by 19PDFScholar
2020

What is Learned in Deep Uncalibrated Photometric Stereo?

ECCV 2020poster

This paper targets at discovering what a deep uncalibrated photometric stereo network learns to resolve the problem’s inherent ambiguity, and designing an effective network architecture based on the new insight to improve the performance. The recently proposed deep uncalibrated photometric stereo me…

Cited by 63SourcePDFScholar
2019

Data-Free Learning of Student Networks

ICCV 2019poster

Learning portable neural networks is very essential for computer vision for the purpose that pre-trained heavy deep models can be well applied on edge devices such as mobile phones and micro sensors. Most existing deep neural network compression and speed-up methods are very effective for training c…

Cited by 442PDFcodeScholar
2019

LegoNet: Efficient Convolutional Neural Networks with Lego Filters

ICML 2019oral

This paper aims to build efficient convolutional neural networks using a set of Lego filters. Many successful building blocks, e.g., inception and residual modules, have been designed to refresh state-of-the-art records of CNNs on visual recognition tasks. Beyond these high-level modules, we suggest…

2019

Reflection Separation using a Pair of Unpolarized and Polarized Images

NeurIPS 2019spotlight

When we take photos through glass windows or doors, the transmitted background scene is often blended with undesirable reflection. Separating two layers apart to enhance the image quality is of vital importance for both human and machine perception. In this paper, we propose to exploit physical cons…

2019

SPLINE-Net: Sparse Photometric Stereo Through Lighting Interpolation and Normal Estimation Networks

ICCV 2019poster

This paper solves the Sparse Photometric stereo through Lighting Interpolation and Normal Estimation using a generative Network (SPLINE-Net). SPLINE-Net contains a lighting interpolation network to generate dense lighting observations given a sparse set of lights as inputs followed by a normal estim…

Cited by 90PDFScholar
2019

Self-Calibrating Deep Photometric Stereo Networks

CVPR 2019oral

This paper proposes an uncalibrated photometric stereo method for non-Lambertian scenes based on deep learning. Unlike previous approaches that heavily rely on assumptions of specific reflectances and light source distributions, our method is able to determine both shape and light directions of a sc…

Cited by 182PDFcodeScholar
2018

CRRN: Multi-Scale Guided Concurrent Reflection Removal Network

CVPR 2018poster

Removing the undesired reflections from images taken through the glass is of broad application to various computer vision tasks. Non-learning based methods utilize different handcrafted priors such as the separable sparse gradients caused by different levels of blurs, which often fail due to their l…

2018

Uncalibrated Photometric Stereo Under Natural Illumination

CVPR 2018poster

This paper presents a photometric stereo method that works with unknown natural illuminations without any calibration object. To solve this challenging problem, we propose the use of an equivalent directional lighting model for small surface patches consisting of slowly varying normals, and solve ea…

Cited by 43SourcePDFScholar
2017

A Microfacet-Based Reflectance Model for Photometric Stereo With Highly Specular Surfaces

ICCV 2017poster

A precise, stable and invertible model for surface reflectance is the key to the success of photometric stereo with real world materials. Recent developments in the field have enabled shape recovery techniques for surfaces of various types, but an effective solution to directly estimating the surfac…

Cited by 28PDFScholar
2016

A Benchmark Dataset and Evaluation for Non-Lambertian and Uncalibrated Photometric Stereo

CVPR 2016poster

Recent progress on photometric stereo extends the technique to deal with general materials and unknown illumination conditions. However, due to the lack of suitable benchmark data with ground truth shapes (normals), quantitative comparison and evaluation is difficult to achieve. In this paper, we fi…

Cited by 358PDFScholar