← Search

Wangmeng Zuo

164 accepted papers

2026

ArtHOI: Taming Foundation Models for Monocular 4D Reconstruction of Hand-Articulated-Object Interactions

CVPR 2026

Existing hand-object interactions (HOI) methods are largely limited to rigid objects, while 4D reconstruction methods of articulated objects generally require pre-scanning the object or even multi-view videos. It remains an unexplored but significant challenge to reconstruct 4D human-articulated-obj

Cited by 0SourcecodeScholar
2026

CGL: Advancing Continual GUI Learning via Reinforcement Fine-Tuning

CVPR 2026

Graphical User Interface (GUI) Agents, benefiting from recent advances in multimodal large language models (MLLM), have achieved significant development. However, due to the frequent updates of GUI applications, adapting to new tasks without forgetting old tasks in GUI continual learning remains an

Cited by 0SourceScholar
2026

CREval: An Automated Interpretable Evaluation for Creative Image Manipulation under Complex Instructions

CVPR 2026

Instruction-based multimodal image manipulation has recently made rapid progress. However, existing evaluation methods lack a systematic and human-aligned framework for assessing model performance on complex and creative editing tasks. To address this gap, we propose CREval, a fully automated questi

Cited by 0SourcecodeScholar
2026

Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining

CVPR 2026

Large-scale video-language pretraining enables strong generalization across multimodal tasks but often incurs prohibitive computational costs. Although recent advances in masked visual modeling help mitigate this issue, they still suffer from two fundamental limitations: severe visual information lo

Cited by 0SourcecodeScholar
2026

Color3D: Controllable and Consistent 3D Colorization with Personalized Colorizer

ICLR 2026poster

In this work, we present Color3D, a highly adaptable framework for colorizing both static and dynamic 3D scenes from monochromatic inputs, delivering visually diverse and chromatically vibrant reconstructions with flexible user-guided control. In contrast to existing methods that focus solely on sta…

Cited by 0SourcecodeScholar
2026

DeAltHDR: Learning HDR Video Reconstruction from Degraded Alternating Exposure Sequences

ICLR 2026poster

High dynamic range (HDR) video can be reconstructed from low dynamic range (LDR) sequences with alternating exposures. However, most existing methods overlook the degradations (e.g., noise and blur) in LDR frames, focusing only on the brightness and position differences between them. To address this…

Cited by 0SourceScholar
2026

Deblur4DGS: 4D Gaussian Splatting from Blurry Monocular Video

AAAI 2026technical

Recent 4D reconstruction methods have yielded impressive results but rely on sharp videos as supervision. However, motion blur often occurs in videos due to camera shake and object movement, while existing methods render blurry results when using such videos for reconstructing 4D models. Although a

Cited by 0SourcePDFScholar
2026

Don't Let Your Robot Be Harmful: Responsible Robotic Manipulation Via Safety-As-Policy

ICRA 2026poster

Unthinking execution of human instructions in robotic manipulation can lead to severe safety risks, such as poisonings, fires, and even explosions. In this paper, we present responsible robotic manipulation, which requires robots to consider potential hazards in the real-world environment while comp…

2026

HP-Edit: A Human-Preference Post-Training Framework for Image Editing

CVPR 2026

Common image editing tasks typically adopt powerful generative diffusion models as the leading paradigm for real-world content editing. Meanwhile, although reinforcement learning (RL) methods such as Diffusion-DPO and Flow-GRPO have further improved generation quality, efficiently applying Reinforce

Cited by 0SourceScholar
2026

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models

CVPR 2026

Aligning text-to-video diffusion models with human preferences is crucial for generating high-quality videos. Existing Direct Preference Otimization (DPO) methods rely on multi-sample ranking and task-specific critic models, which is inefficient and often yields ambiguous global supervision. To addr

Cited by 0SourcecodeScholar
2026

MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling

ICLR 2026poster

Existing video generation models predominantly emphasize appearance fidelity while exhibiting limited ability to synthesize complex human motions, such as whole-body movements, long-range dynamics, and fine-grained human–environment interactions. This often leads to unrealistic or physically implaus…

Cited by 0SourceScholar
2026

Perceptual Quality Assessment of 3D Gaussian Splatting: A Subjective Dataset and Prediction Metric

AAAI 2026technical

With the rapid advancement of 3D visualization, 3D Gaussian Splatting (3DGS) has emerged as a leading technique for real-time, high-fidelity rendering. While prior research has emphasized algorithmic performance and visual fidelity, the perceptual quality of 3DGS-rendered content, especially under v

Cited by 0SourcePDFScholar
2026

PreferThinker: Reasoning-based Personalized Image Preference Assessment

ICLR 2026poster

Personalized image preference assessment aims to evaluate an individual user's image preferences by relying only on a small set of reference images as prior information. Existing methods mainly focus on general preference assessment, training models with large-scale data to tackle well-defined task…

Cited by 0SourceScholar
2026

Premier: Personalized Preference Modulation with Learnable User Embedding in Text-to-Image Generation

CVPR 2026

Text-to-image generation has advanced rapidly, yet it still struggles to capture the nuanced user preferences. Existing approaches typically rely on multimodal large language models to infer user preferences, but the derived prompts or latent codes rarely reflect them faithfully, leading to suboptim

Cited by 0SourceScholar
2026

RefSTAR: Blind Face Image Restoration with Reference Selection, Transfer, and Reconstruction

AAAI 2026technical

Introducing high-quality references can largely alleviate the uncertainty in blind face image restoration tasks, yet the equivocal utilization of reference priors makes it still a struggle to well preserve the human identity. We attribute the identity inconsistency to two deficiencies of existing re

Cited by 0SourcePDFScholar
2026

UniRestorer: Universal Image Restoration via Adaptively Estimating Image Degradation at Proper Granularity

ICLR 2026poster

Recently, considerable progress has been made in all-in-one image restoration. Generally, existing methods can be degradation-agnostic or degradation-aware. However, the former are limited in leveraging degradation estimation-based priors, and the latter suffer from the inevitable error in degradati…

Cited by 0SourcecodeScholar
2025

ACE: Anti-Editing Concept Erasure in Text-to-Image Models

CVPR 2025poster

Recent advance in text-to-image diffusion models have significantly facilitated the generation of high-quality images, but also raising concerns about the illegal creation of harmful content, such as copyrighted images. Existing concept erasure methods achieve superior results in preventing the prod…

2025

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

NeurIPS 2025poster

In this paper, we introduce BMMR, a large-scale bilingual, multimodal, multi-disciplinary reasoning dataset for the community to develop and evaluate large multimodal models (LMMs). BMMR comprises 100k university-level questions drawn from 300 UNESCO-defined subjects, spanning diverse formats—multip…

Cited by 0SourceScholar
2025

CASP: Consistency-aware Audio-induced Saliency Prediction Model for Omnidirectional Video

CVPR 2025poster

Omnidirectional videos (ODVs) present distinct challenges for accurate audio-visual saliency prediction due to their immersive nature, which combines spatial audio with panoramic visuals to enhance the user experience. While auditory cues are crucial for guiding visual attention across the panoramic…

Cited by 0SourcePDFScholar
2025

DreamPhysics: Learning Physics-Based 3D Dynamics with Video Diffusion Priors

AAAI 2025technical

Dynamic 3D interaction has been attracting a lot of attention recently. However, creating such 4D content remains challenging. One solution is to animate 3D scenes with physics-based simulation, which requires manually assigning precise physical properties to the object or the simulated results woul…

2025

Exposure Bracketing Is All You Need For A High-Quality Image

ICLR 2025poster

It is highly desired but challenging to acquire high-quality photos with clear content in low-light environments. Although multi-image processing methods (using burst, dual-exposure, or multi-exposure images) have made significant progress in addressing this issue, they typically focus on specific r…

2025

Federated Residual Low-Rank Adaptation of Large Language Models

ICLR 2025poster

Low-Rank Adaptation (LoRA) presents an effective solution for federated fine-tuning of Large Language Models (LLMs), as it substantially reduces communication overhead. However, a straightforward combination of FedAvg and LoRA results in suboptimal performance, especially under data heterogeneity. W…

Cited by 0SourcePDFScholar
2025

FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors

ICCV 2025poster

Interactive image editing allows users to modify images through visual interaction operations such as drawing, clicking, and dragging. Existing methods construct such supervision signals from videos, as they capture how objects change with various physical interactions. However, these models are usu…

2025

Generative Inbetweening through Frame-wise Conditions-Driven Video Generation

CVPR 2025poster

Generative inbetweening aims to generate intermediate frame sequences by utilizing two key frames as input. Although remarkable progress has been made in video generation models, generative inbetweening still faces challenges in maintaining temporal stability due to the ambiguous interpolation path…

2025

Integrating Visual Interpretation and Linguistic Reasoning for Geometric Problem Solving

ICCV 2025poster

Current large vision-language models (LVLMs) typically employ a connector module to link visual features with text embeddings of large language models (LLMs) and use end-to-end training to achieve multi-modal understanding in a unified process. Well alignment needs high-quality pre-training data and…

2025

LabUtopia: High-Fidelity Simulation and Hierarchical Benchmark for Scientific Embodied Agents

NeurIPS 2025poster

Scientific embodied agents play a crucial role in modern laboratories by automating complex experimental workflows. Compared to typical household environments, laboratory settings impose significantly higher demands on perception of physical-chemical transformations and long-horizon planning, making…

Cited by 0SourcecodeScholar
2025

MC^2: Multi-concept Guidance for Customized Multi-concept Generation

CVPR 2025poster

Customized text-to-image generation, which synthesizes images based on user-specified concepts, has made significant progress in handling individual concepts. However, when extended to multiple concepts, existing methods often struggle with properly integrating different models and avoiding the unin…

2025

MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM

NeurIPS 2025poster

Multimodal hallucination in multimodal large language models (MLLMs) restricts the correctness of MLLMs. However, multimodal hallucinations are multi-sourced and arise from diverse causes. Existing benchmarks fail to adequately distinguish between perception-induced hallucinations and reasoning-indu…

Cited by 0SourceScholar
2025

MV-VTON: Multi-View Virtual Try-On with Diffusion Models

AAAI 2025technical

The goal of image-based virtual try-on is to generate an image of the target person naturally wearing the given clothing. However, existing methods solely focus on the frontal try-on using the frontal clothing. When the views of the clothing and person are significantly inconsistent, particularly wh…

2025

On the Importance of Language-driven Representation Learning for Heterogeneous Federated Learning

ICLR 2025poster

Non-Independent and Identically Distributed (Non-IID) training data significantly challenge federated learning (FL), impairing the performance of the global model in distributed frameworks. Inspired by the superior performance and generalizability of language-driven representation learning in centra…

Cited by 0SourcePDFScholar
2025

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning

NeurIPS 2025poster

Recent advances in large language models have significantly improved textual reasoning through the effective use of Chain-of-Thought (CoT) and reinforcement learning. However, extending these successes to vision-language tasks remains challenging due to inherent limitations in text-only CoT, such as…

Cited by 0SourceScholar
2025

QR-LoRA: Efficient and Disentangled Fine-tuning via QR Decomposition for Customized Generation

ICCV 2025poster

Existing text-to-image models often rely on parame- ter fine-tuning techniques such as Low-Rank Adaptation (LoRA) to customize visual attributes. However, when com- bining multiple LoRA models for content-style fusion tasks, unstructured modifications of weight matrices often lead to undesired featu…

Cited by 0SourcePDFScholar
2025

ReMP-AD: Retrieval-enhanced Multi-modal Prompt Fusion for Few-Shot Industrial Visual Anomaly Detection

ICCV 2025poster

Industrial visual inspection is crucial for detecting defects in manufactured products, but it traditionally relies on human operators, leading to inefficiencies. Industrial Visual Anomaly Detection (IVAD) has emerged as a promising solution, with methods such as zero-shot, few-shot, and reconstruct…

2025

Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers

ICCV 2025poster

Multimodal Diffusion Transformers (MM-DiTs) have achieved remarkable progress in text-driven visual generation. However, even state-of-the-art MM-DiT models like FLUX struggle with achieving precise alignment between text prompts and generated content. We identify two key issues in the attention mec…

2025

Rethinking Transformer-Based Blind-Spot Network for Self-Supervised Image Denoising

AAAI 2025technical

Blind-spot networks (BSN) have been prevalent neural architectures in self-supervised image denoising (SSID). However, most existing BSNs are conducted with convolution layers. Although transformers have shown the potential to overcome the limitations of convolutions in many image restoration tasks,…

2025

RoomEditor: High-Fidelity Furniture Synthesis with Parameter-Sharing U-Net

NeurIPS 2025poster

Virtual furniture synthesis, a critical task in image composition, aims to seamlessly integrate reference objects into indoor scenes while preserving geometric coherence and visual realism. Despite its significant potential in home design applications, this field remains underexplored due to two maj…

Cited by 0SourcecodeScholar
2025

S2Gaussian: Sparse-View Super-Resolution 3D Gaussian Splatting

CVPR 2025poster

In this paper, we aim ambitiously for a realistic yet challenging problem, namely, how to reconstruct high-quality 3D scenes from sparse low-resolution views that simultaneously suffer from deficient perspectives and contents. Whereas existing methods only deal with either sparse views or low-resolu…

2025

Triad: Empowering LMM-based Anomaly Detection with Expert-guided Region-of-Interest Tokenizer and Manufacturing Process

ICCV 2025poster

Although recent methods have tried to introduce large multimodal models (LMMs) into industrial anomaly detection (IAD), their generalization in the IAD field is far inferior to that for general purposes. We summarize the main reasons for this gap into two aspects. On one hand, general-purpose LMMs l…

2025

VQA4CIR: Boosting Composed Image Retrieval with Visual Question Answering

AAAI 2025technical

Albeit progress has been made in Composed Image Retrieval (CIR), we empirically find that a certain percentage of failure retrieval results are not consistent with their relative captions. To address this issue, this work provides a Visual Question Answering (VQA) perspective to boost the performanc…

2025

VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models

AAAI 2025technical

Text-to-image diffusion models (T2I) have demonstrated unprecedented capabilities in creating realistic and aesthetic images. On the contrary, text-to-video diffusion models (T2V) still lag far behind in frame quality and text alignment, owing to insufficient quality and quantity of training videos.…

2025

Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning

ICLR 2025poster

As large-scale models evolve, language instructions are increasingly utilized in multi-modal tasks. Due to human language habits, these instructions often contain ambiguities in real-world scenarios, necessitating the integration of visual context or common sense for accurate interpretation. However…

2024

Arbitrary-Scale Video Super-Resolution with Structural and Textural Priors

ECCV 2024poster

"Arbitrary-scale video super-resolution (AVSR) aims to enhance the resolution of video frames, potentially at various scaling factors, which presents several challenges regarding spatial detail reproduction, temporal consistency, and computational complexity. In this paper, we first describe a stron…

2024

Combining Generative and Geometry Priors for Wide-Angle Portrait Correction

ECCV 2024poster

"Wide-angle lens distortion in portrait photography presents a significant challenge for capturing photo-realistic and aesthetically pleasing images. Such distortions are especially noticeable in facial regions. In this work, we propose encapsulating the generative face prior as a guided natural man…

2024

ControlVideo: Training-free Controllable Text-to-video Generation

ICLR 2024poster

Text-driven diffusion models have unlocked unprecedented abilities in image generation, whereas their video counterpart lags behind due to the excessive training cost. To avert the training burden, we propose a training-free ControlVideo to produce high-quality videos based on the provided text prom…

2024

Decoupled Textual Embeddings for Customized Image Generation

AAAI 2024technical

Customized text-to-image generation, which aims to learn user-specified concepts with a few images, has drawn significant attention recently. However, existing methods usually suffer from overfitting issues and entangle the subject-unrelated information (e.g., background and pose) with the learned c…

2024

DreamControl: Control-Based Text-to-3D Generation with 3D Self-Prior

CVPR 2024poster

3D generation has raised great attention in recent years. With the success of text-to-image diffusion models the 2D-lifting technique becomes a promising route to controllable 3D generation. However these methods tend to present inconsistent geometry which is also known as the Janus problem. We obse…

2024

Evaluation of Text-to-Video Generation Models: A Dynamics Perspective

NeurIPS 2024poster

Comprehensive and constructive evaluation protocols play an important role when developing sophisticated text-to-video (T2V) generation models. Existing evaluation protocols primarily focus on temporal consistency and content continuity, yet largely ignore dynamics of video content. Such dynamics is…

2024

GLAD: Towards Better Reconstruction with Global and Local Adaptive Diffusion Models for Unsupervised Anomaly Detection

ECCV 2024poster

"Diffusion models have shown superior performance on unsupervised anomaly detection tasks. Since trained with normal data only, diffusion models tend to reconstruct normal counterparts of test images with certain noises added. However, these methods treat all potential anomalies equally, which may c…

2024

Improved Generation of Adversarial Examples Against Safety-aligned LLMs

NeurIPS 2024poster

Adversarial prompts (or say, adversarial examples) generated using gradient-based methods exhibit outstanding performance in performing automatic jailbreak attacks against safety-aligned LLMs. Nevertheless, due to the discrete nature of texts, the input gradient of LLMs struggles to precisely reflec…

2024

Improving Image Restoration through Removing Degradations in Textual Representations

CVPR 2024poster

In this paper we introduce a new perspective for improving image restoration by removing degradation in the textual representations of a given degraded image. Intuitively restoration is much easier on text modality than image one. For example it can be easily conducted by removing degradation-relate…

2024

Learning Real-World Image De-weathering with Imperfect Supervision

AAAI 2024technical

Real-world image de-weathering aims at removing various undesirable weather-related artifacts. Owing to the impossibility of capturing image pairs concurrently, existing real-world de-weathering datasets often exhibit inconsistent illumination, position, and textures between the ground-truth images…

2024

Multi-modal Crowd Counting via a Broker Modality

ECCV 2024poster

"Multi-modal crowd counting involves estimating crowd density from both visual and thermal/depth images. This task is challenging due to the significant gap between these distinct modalities. In this paper, we propose a novel approach by introducing an auxiliary broker modality and on this basis fra…

2024

PLACE: Adaptive Layout-Semantic Fusion for Semantic Image Synthesis

CVPR 2024highlight

Recent advancements in large-scale pre-trained text-to-image models have led to remarkable progress in semantic image synthesis. Nevertheless synthesizing high-quality images with consistent semantics and layout remains a challenge. In this paper we propose the adaPtive LAyout-semantiC fusion modulE…

2024

SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models

ACL 2024findings

In the rapidly evolving landscape of Large Language Models (LLMs), ensuring robust safety measures is paramount. To meet this crucial need, we propose SALAD-Bench, a safety benchmark specifically designed for evaluating LLMs, attack, and defense methods. Distinguished by its breadth, SALAD-Bench tra…

2024

Self-Supervised High Dynamic Range Imaging with Multi-Exposure Images in Dynamic Scenes

ICLR 2024poster

Merging multi-exposure images is a common approach for obtaining high dynamic range (HDR) images, with the primary challenge being the avoidance of ghosting artifacts in dynamic scenes. Recent methods have proposed using deep neural networks for deghosting. However, the methods typically rely on suf…

2024

Self-Supervised Video Desmoking for Laparoscopic Surgery

ECCV 2024oral

"Due to the difficulty of collecting real paired data, most existing desmoking methods train the models by synthesizing smoke, generalizing poorly to real surgical scenarios. Although a few works have explored single-image real-world desmoking in unpaired learning manners, they still encounter chall…

2024

Sentence-level Prompts Benefit Composed Image Retrieval

ICLR 2024spotlight

Composed image retrieval (CIR) is the task of retrieving specific images by using a query that involves both a reference image and a relative caption. Most existing CIR models adopt the late-fusion strategy to combine visual and language features. Besides, several approaches have also been suggested…

2024

ShoeModel: Learning to Wear on the User-specified Shoes via Diffusion Model

ECCV 2024poster

"With the development of the large-scale diffusion model, Artificial Intelligence Generated Content (AIGC) techniques are popular recently. However, how to truly make it serve our daily lives remains an open question. To this end, in this paper, we focus on employing AIGC techniques in one filed of…

Cited by 2SourcePDFScholar
2024

SmartControl: Enhancing ControlNet for Handling Rough Visual Conditions

ECCV 2024poster

"Recent text-to-image generation methods such as ControlNet have achieved remarkable success in controlling image layouts, where the generated images by the default model are constrained to strictly follow the visual conditions (e.g., depth maps). However, in practice, the conditions usually provide…

2024

TextField3D: Towards Enhancing Open-Vocabulary 3D Generation with Noisy Text Fields

ICLR 2024poster

Recent works learn 3D representation explicitly under text-3D guidance. However, limited text-3D data restricts the vocabulary scale and text control of generations. Generators may easily fall into a stereotype concept for certain text prompts, thus losing open-vocabulary generation ability. To tack…

Cited by 12SourcePDFScholar
2024

UniM2AE: Multi-modal Masked Autoencoders with Unified 3D Representation for 3D Perception in Autonomous Driving

ECCV 2024poster

"Masked Autoencoders (MAE) play a pivotal role in learning potent representations, delivering outstanding results across various 3D perception tasks essential for autonomous driving. In real-world driving scenarios, it’s commonplace to deploy multiple sensors for comprehensive environment perception…

2024

VQ-FONT: Few-Shot Font Generation with Structure-Aware Enhancement and Quantization

AAAI 2024technical

Few-shot font generation is challenging, as it needs to capture the fine-grained stroke styles from a limited set of reference glyphs, and then transfer to other characters, which are expected to have similar styles. However, due to the diversity and complexity of Chinese font styles, the synthesize…

2023

Beyond Image Borders: Learning Feature Extrapolation for Unbounded Image Composition

ICCV 2023poster

For improving image composition and aesthetic quality, most existing methods modulate the captured images by striking out redundant content near the image borders. However, such image cropping methods are limited in the range of image views. Some methods have been suggested to extrapolate the images…

Cited by 2PDFcodeScholar
2023

Black-Box Tuning of Vision-Language Models with Effective Gradient Approximation

EMNLP 2023long findings

Parameter-efficient fine-tuning (PEFT) methods have provided an effective way for adapting large vision-language models to specific tasks or scenarios. Typically, they learn a very small scale of parameters for pre-trained models in a white-box formulation, which assumes model architectures to be kn…

Cited by 0SourcecodeScholar
2023

CLIP2Point: Transfer CLIP to Point Cloud Classification with Image-Depth Pre-Training

ICCV 2023poster

Pre-training across 3D vision and language remains under development because of limited training data. Recent works attempt to transfer vision-language (V-L) pre-training methods to 3D vision. However, the domain gap between 3D and images is unsolved, so that V-L pre-trained models are restricted in…

Cited by 167PDFcodeScholar
2023

Diverse Data Augmentation with Diffusions for Effective Test-time Prompt Tuning

ICCV 2023poster

Benefiting from prompt tuning, recent years have witnessed the promising performance of pre-trained vision-language models, e.g., CLIP, on versatile downstream tasks. In this paper, we focus on a particular setting of learning adaptive prompts on the fly for each test sample from an unseen new domai…

Cited by 93PDFcodeScholar
2023

ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation

ICCV 2023oral

In addition to the unprecedented ability in imaginary creation, large text-to-image models are expected to take customized concepts in image generation. Existing works generally learn such concepts in an optimization-based manner, yet bringing excessive computation or memory burden. In this paper, w…

Cited by 361PDFcodeScholar
2023

ImaginaryNet: Learning Object Detectors without Real Images and Annotations

ICLR 2023poster

Without the demand of training in reality, humans are able of detecting a new category of object simply based on the language description on its visual characteristics. Empowering deep learning with this ability undoubtedly enables the neural network to handle complex vision tasks, e.g., object dete…

2023

Improving Adversarial Transferability via Intermediate-level Perturbation Decay

NeurIPS 2023poster

Intermediate-level attacks that attempt to perturb feature representations following an adversarial direction drastically have shown favorable performance in crafting transferable adversarial examples. Existing methods in this category are normally formulated with two separate stages, where a direct…

2023

Inferring and Leveraging Parts From Object Shape for Improving Semantic Image Synthesis

CVPR 2023poster

Despite the progress in semantic image synthesis, it remains a challenging problem to generate photo-realistic parts from input semantic map. Integrating part segmentation map can undoubtedly benefit image synthesis, but is bothersome and inconvenient to be provided by users. To improve part synthes…

2023

Joint Video Multi-Frame Interpolation and Deblurring Under Unknown Exposure Time

CVPR 2023poster

Natural videos captured by consumer cameras often suffer from low framerate and motion blur due to the combination of dynamic scene complexity, lens and sensor imperfection, and less than ideal exposure setting. As a result, computational methods that jointly perform video frame interpolation and de…

2023

Learning Federated Visual Prompt in Null Space for MRI Reconstruction

CVPR 2023poster

Federated Magnetic Resonance Imaging (MRI) reconstruction enables multiple hospitals to collaborate distributedly without aggregating local data, thereby protecting patient privacy. However, the data heterogeneity caused by different MRI protocols, insufficient local training data, and limited commu…

2023

Learning Generative Structure Prior for Blind Text Image Super-Resolution

CVPR 2023poster

Blind text image super-resolution (SR) is challenging as one needs to cope with diverse font styles and unknown degradation. To address the problem, existing methods perform character recognition in parallel to regularize the SR task, either through a loss constraint or intermediate feature conditio…

2023

Learning Single Image Defocus Deblurring with Misaligned Training Pairs

AAAI 2023technical

By adopting popular pixel-wise loss, existing methods for defocus deblurring heavily rely on well aligned training image pairs. Although training pairs of ground-truth and blurry images are carefully collected, e.g., DPDD dataset, misalignment is inevitable between training pairs, making existing me…

2023

Making Substitute Models More Bayesian Can Enhance Transferability of Adversarial Examples

ICLR 2023poster

The transferability of adversarial examples across deep neural networks (DNNs) is the crux of many black-box attacks. Many prior efforts have been devoted to improving the transferability via increasing the diversity in inputs of some substitute models. In this paper, by contrast, we opt for the div…

2023

MetaF2N: Blind Image Super-Resolution by Learning Efficient Model Adaptation from Faces

ICCV 2023poster

Due to their highly structured characteristics, faces are easier to recover than natural scenes for blind image super-resolution. Therefore, we can extract the degradation representation of an image from the low-quality and recovered face pairs. Using the degradation representation, realistic low-qu…

Cited by 7PDFcodeScholar
2023

Physics-Guided ISO-Dependent Sensor Noise Modeling for Extreme Low-Light Photography

CVPR 2023poster

Although deep neural networks have achieved astonishing performance in many vision tasks, existing learning-based methods are far inferior to the physical model-based solutions in extreme low-light sensor noise modeling. To tap the potential of learning-based sensor noise modeling, we investigate th…

2023

Self-supervised Learning to Bring Dual Reversed Rolling Shutter Images Alive

ICCV 2023poster

Modern consumer cameras usually employ the rolling shutter (RS) mechanism, where images are captured by scanning scenes row-by-row, yielding RS distortions for dynamic scenes. To correct RS distortions, existing methods adopt a fully supervised learning manner, where high framerate global shutter (G…

Cited by 5PDFScholar
2023

Spatially Adaptive Self-Supervised Learning for Real-World Image Denoising

CVPR 2023poster

Significant progress has been made in self-supervised image denoising (SSID) in the recent few years. However, most methods focus on dealing with spatially independent noise, and they have little practicality on real-world sRGB images with spatially correlated noise. Although pixel-shuffle downsampl…

2023

Symmetry-Aware Transformer-Based Mirror Detection

AAAI 2023technical

Mirror detection aims to identify the mirror regions in the given input image. Existing works mainly focus on integrating the semantic features and structural features to mine specific relations between mirror and non-mirror regions, or introducing mirror properties like depth or chirality to help a…

2023

Texts as Images in Prompt Tuning for Multi-Label Image Recognition

CVPR 2023poster

Prompt tuning has been employed as an efficient way to adapt large vision-language pre-trained models (e.g. CLIP) to various downstream tasks in data-limited or label-limited settings. Nonetheless, visual data (e.g., images) is by default prerequisite for learning prompts in existing methods. In thi…

2023

Towards Evaluating Transfer-based Attacks Systematically, Practically, and Fairly

NeurIPS 2023poster

The adversarial vulnerability of deep neural networks (DNNs) has drawn great attention due to the security risk of applying these models in real-world applications. Based on transferability of adversarial examples, an increasing number of transfer-based methods have been developed to fool black-box…

Cited by 3SourcePDFScholar
2023

Towards Instance-adaptive Inference for Federated Learning

ICCV 2023poster

Federated learning (FL) is a distributed learning paradigm that enables multiple clients to learn a powerful global model by aggregating local training. However, the performance of the global model is often hampered by non-i.i.d. distribution among the clients, requiring extensive efforts to mitigat…

Cited by 22PDFcodeScholar
2022

Adversarial Contrastive Learning via Asymmetric InfoNCE

ECCV 2022poster

"Contrastive learning (CL) has recently been applied to adversarial learning tasks. Such practice considers adversarial perturbations as additional positive samples of an instance, and by maximizing their agreements with each other, yields better adversarial robustness. However, this mechanism can b…

2022

From Face to Natural Image: Learning Real Degradation for Blind Image Super-Resolution

ECCV 2022poster

"How to design proper training pairs is critical for super-resolving real-world low-quality (LQ) images, which suffers from the difficulties in either acquiring paired ground-truth high-quality (HQ) images or synthesizing photo-realistic degraded LQ observations. Recent works mainly focus on modelin…

2022

Incorporating Semi-Supervised and Positive-Unlabeled Learning for Boosting Full Reference Image Quality Assessment

CVPR 2022poster

Full-reference (FR) image quality assessment (IQA) evaluates the visual quality of a distorted image by measuring its perceptual difference with pristine-quality reference, and has been widely used in low level vision tasks. Pairwise labeled data with mean opinion score (MOS) are required in trainin…

Cited by 33PDFcodeScholar
2022

Localization Distillation for Dense Object Detection

CVPR 2022poster

Knowledge distillation (KD) has witnessed its powerful capability in learning compact models in object detection. Previous KD methods for object detection mostly focus on imitating deep features within the imitation regions instead of logit mimicking on classification due to the inefficiency in dist…

Cited by 239PDFcodeScholar
2022

Retrieval-Based Spatially Adaptive Normalization for Semantic Image Synthesis

CVPR 2022poster

Semantic image synthesis is a challenging task with many practical applications. Albeit remarkable progress has been made in semantic image synthesis with spatially-adaptive normalization and existing methods normalize the feature activations under the coarse-level guidance (e.g., semantic class). H…

Cited by 33PDFcodeScholar
2022

Self-Supervised Image Restoration with Blurry and Noisy Pairs

NeurIPS 2022accept

When taking photos under an environment with insufficient light, the exposure time and the sensor gain usually require to be carefully chosen to obtain images with satisfying visual quality. For example, the images with high ISO usually have inescapable noise, while the long-exposure ones may be blu…

2022

Self-Supervised Learning for Real-World Super-Resolution from Dual Zoomed Observations

ECCV 2022poster

"In this paper, we consider two challenging issues in reference-based super-resolution (RefSR), (i) how to choose a proper reference image, and (ii) how to learn real-world RefSR in a self-supervised manner. Particularly, we present a novel self-supervised learning approach for real-world image SR f…

2022

Self-supervised Learning and Adaptation for Single Image Dehazing

IJCAI 2022poster

Existing deep image dehazing methods usually depend on supervised learning with a large number of hazy-clean image pairs which are expensive or difficult to collect. Moreover, dehazing performance of the learned model may deteriorate significantly when the training hazy-clean image pairs are insuffi…

2022

Semantic-Shape Adaptive Feature Modulation for Semantic Image Synthesis

CVPR 2022poster

Recent years have witnessed substantial progress in semantic image synthesis, it is still challenging in synthesizing photo-realistic images with rich details. Most previous methods focus on exploiting the given semantic map, which just captures an object-level layout for an image. Obviously, a fine…

Cited by 35PDFcodeScholar
2022

Towards Diverse and Faithful One-shot Adaption of Generative Adversarial Networks

NeurIPS 2022accept

One-shot generative domain adaption aims to transfer a pre-trained generator on one domain to a new domain using one reference image only. However, it remains very challenging for the adapted generator (i) to generate diverse images inherited from the pre-trained generator while (ii) faithfully acqu…

2022

Unidirectional Video Denoising by Mimicking Backward Recurrent Modules with Look-Ahead Forward Ones

ECCV 2022poster

"While significant progress has been made in deep video denoising, it remains very challenging for exploiting historical and future frames. Bidirectional recurrent networks (BiRNN) have exhibited appealing performance in several video restoration tasks. However, BiRNN is intrinsically offline becaus…

2022

W2N: Switching from Weak Supervision to Noisy Supervision for Object Detection

ECCV 2022poster

"Weakly-supervised object detection (WSOD) aims to train an object detector only requiring the image-level annotations. Recently, some works have managed to select the accurate boxes generated from a well-trained WSOD network to supervise a semi-supervised detection framework for better performance.…

2021

Boosting Weakly Supervised Object Detection via Learning Bounding Box Adjusters

ICCV 2021poster

Weakly-supervised object detection (WSOD) has emerged as an inspiring recent topic to avoid expensive instance-level object annotations. However, the bounding boxes of most existing WSOD methods are mainly determined by precomputed proposals, thereby being limited in precise object localization. In…

Cited by 60PDFcodeScholar
2021

Bringing Events Into Video Deblurring With Non-Consecutively Blurry Frames

ICCV 2021poster

Recently, video deblurring has attracted considerable research attention, and several works suggest that events at high time rate can benefit deblurring. In this paper, we develop a principled framework D2Nets for video deblurring to exploit non-consecutively blurry frames, and propose a flexible ev…

Cited by 73PDFcodeScholar
2021

Learning RAW-to-sRGB Mappings With Inaccurately Aligned Supervision

ICCV 2021poster

Learning RAW-to-sRGB mapping has drawn increasing attention in recent years, wherein an input raw image is trained to imitate the target sRGB image captured by another camera. However, the severe color inconsistency makes it very challenging to generate well-aligned training pairs of input raw and t…

Cited by 52PDFcodeScholar
2021

Learning Scalable lY=-Constrained Near-Lossless Image Compression via Joint Lossy Image and Residual Compression

CVPR 2021poster

We propose a novel joint lossy image and residual compression framework for learning l_infinity-constrained near-lossless image compression. Specifically, we obtain a lossy reconstruction of the raw image through lossy image compression and uniformly quantize the corresponding residual to satisfy a…

Cited by 33PDFScholar
2021

Learning Semantic Person Image Generation by Region-Adaptive Normalization

CVPR 2021poster

Human pose transfer has received great attention due to its wide applications, yet is still a challenging task that is not well solved. Recent works have achieved great success to transfer the person image from the source to the target pose. However, most of them cannot well capture the semantic app…

Cited by 81PDFcodeScholar
2021

Orthogonal Jacobian Regularization for Unsupervised Disentanglement in Image Generation

ICCV 2021poster

Unsupervised disentanglement learning is a crucial issue for understanding and exploiting deep generative models. Recently, SeFa tries to find latent disentangled directions by performing SVD on the first projection of a pre-trained GAN. However, it is only applied to the first layer and works in a…

Cited by 72PDFcodeScholar
2021

Variational Attention: Propagating Domain-Specific Knowledge for Multi-Domain Learning in Crowd Counting

ICCV 2021poster

In crowd counting, due to the problem of laborious labelling, it is perceived intractability of collecting a new large-scale dataset which has plentiful images with large diversity in density, scene, etc. Thus, for learning a general model, training with data from multiple different datasets might b…

Cited by 57PDFcodeScholar
2021

VirFace: Enhancing Face Recognition via Unlabeled Shallow Data

CVPR 2021poster

Recently, exploiting the effect of the unlabeled data for face recognition attracts increasing attention. However, there are still few works considering the situation that the unlabeled data is shallow which widely exists in real-world scenarios. The existing semi-supervised face recognition methods…

Cited by 22PDFScholar
2020

Blind Face Restoration via Deep Multi-scale Component Dictionaries

ECCV 2020poster

Recent reference-based face restoration methods have received considerable attention due to their great capability in recovering high-frequency details on real low-quality images. However, most of these methods require a high-quality reference image of the same identity, making them only applicable…

2020

Component Divide-and-Conquer for Real-World Image Super-Resolution

ECCV 2020poster

In this paper, we present a large-scale Diverse Real-world image Super-Resolution dataset, i.e., DRealSR, as well as a divide-and-conquer Super-Resolution (SR) network, exploring the utility of guiding SR model with low-level image components. DRealSR establishes a new SR benchmark with diverse real…

2020

Cross-Scale Internal Graph Neural Network for Image Super-Resolution

NeurIPS 2020poster

Non-local self-similarity in natural images has been well studied as an effective prior in image restoration. However, for single image super-resolution (SISR), most existing deep non-local methods (e.g., non-local neural networks) only exploit similar patches within the same scale of the low-resolu…

2020

ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks

CVPR 2020poster

Recently, channel attention mechanism has demonstrated to offer great potential in improving the performance of deep convolutional neural networks (CNNs). However, most existing methods dedicate to developing more sophisticated attention modules for achieving better performance, which inevitably inc…

Cited by 7872PDFcodeScholar
2020

Enhanced Blind Face Restoration With Multi-Exemplar Images and Adaptive Spatial Feature Fusion

CVPR 2020oral

In many real-world face restoration applications, e.g., smartphone photo albums and old films, multiple high-quality (HQ) images of the same person usually are available for a given degraded low-quality (LQ) observation. However, most existing guided face restoration methods are based on single HQ e…

Cited by 129PDFcodeScholar
2020

Learning Flow-based Feature Warping for Face Frontalization with Illumination Inconsistent Supervision

ECCV 2020poster

Despite recent advances in deep learning-based face frontalization methods, photo-realistic and illumination preserving frontal face synthesis is still challenging due to large pose and illumination discrepancy during training. We propose a novel Flow-based Feature Warping Model (FFWM) which can lea…

2020

Towards Photo-Realistic Virtual Try-On by Adaptively Generating-Preserving Image Content

CVPR 2020poster

Image visual try-on aims at transferring a target clothes image onto a reference person, and has become a hot topic in recent years. Prior arts usually focus on preserving the character of a clothes image (e.g. texture, logo, embroidery) when warping it to arbitrary human pose. However, it remains a…

Cited by 337PDFScholar
2020

What Deep CNNs Benefit From Global Covariance Pooling: An Optimization Perspective

CVPR 2020poster

Recent works have demonstrated that global covariance pooling (GCP) has the ability to improve performance of deep convolutional neural networks (CNNs) on visual classification task. Despite considerable advance, the reasons on effectiveness of GCP on deep CNNs have not been well studied. In this pa…

Cited by 30PDFcodeScholar
2019

DAVANet: Stereo Deblurring With View Aggregation

CVPR 2019oral

Nowadays stereo cameras are more commonly adopted in emerging devices such as dual-lens smartphones and unmanned aerial vehicles. However, they also suffer from blurry images in dynamic scenes which leads to visual discomfort and hampers further image processing. Previous works have succeeded in mon…

Cited by 119PDFScholar
2019

Image Inpainting With Learnable Bidirectional Attention Maps

ICCV 2019poster

Most convolutional network (CNN)-based inpainting methods adopt standard convolution to indistinguishably treat valid pixels and holes, making them limited in handling irregular holes and more likely to generate inpainting results with color discrepancy and blurriness. Partial convolution has been s…

Cited by 319PDFcodeScholar
2019

Multispectral and Hyperspectral Image Fusion by MS/HS Fusion Net

CVPR 2019poster

Hyperspectral imaging can help better understand the characteristics of different materials, compared with traditional image systems. However, only high-resolution multispectral (HrMS) and low-resolution hyperspectral (LrHS) images can generally be captured at video rate in practice. In this paper,…

Cited by 291PDFcodeScholar
2019

Perspective-Guided Convolution Networks for Crowd Counting

ICCV 2019poster

In this paper, we propose a novel perspective-guided convolution (PGC) for convolutional neural network (CNN) based crowd counting (i.e. PGCNet), which aims to overcome the dramatic intra-scene scale variations of people due to the perspective effect. While most state-of-the-arts adopt multi-scale o…

Cited by 245PDFcodeScholar
2019

Progressive Image Deraining Networks: A Better and Simpler Baseline

CVPR 2019poster

Along with the deraining performance improvement of deep networks, their structures and learning become more and more complicated and diverse, making it difficult to analyze the contribution of various network modules when developing new deraining networks. To handle this issue, this paper provides…

Cited by 1077PDFcodeScholar
2019

STGAN: A Unified Selective Transfer Network for Arbitrary Image Attribute Editing

CVPR 2019poster

Arbitrary attribute editing generally can be tackled by incorporating encoder-decoder and generative adversarial networks. However, the bottleneck layer in encoder-decoder usually gives rise to blurry and low quality editing result. And adding skip connections improves image quality at the cost of w…

Cited by 427PDFcodeScholar
2019

Spatio-Temporal Filter Adaptive Network for Video Deblurring

ICCV 2019poster

Video deblurring is a challenging task due to the spatially variant blur caused by camera shake, object motions, and depth variations, etc. Existing methods usually estimate optical flow in the blurry video to align consecutive frames or approximate blur kernels. However, they tend to generate artif…

Cited by 247PDFScholar
2018

Deep Cocktail Network: Multi-Source Unsupervised Domain Adaptation With Category Shift

CVPR 2018poster

Most existing unsupervised domain adaptation (UDA) methods are based upon the assumption that source labeled data come from an identical underlying distribution. Whereas in practical scenario, labeled instances are typically collected from diverse sources. Moreover, those sources may not completely…

2018

Deep Non-Blind Deconvolution via Generalized Low-Rank Approximation

NeurIPS 2018poster

In this paper, we present a deep convolutional neural network to capture the inherent properties of image degradation, which can handle different kernels and saturated pixels in a unified framework. The proposed neural network is motivated by the low-rank property of pseudo-inverse kernels. We firs…

Cited by 99SourcePDFScholar
2018

Generative Adversarial Learning Towards Fast Weakly Supervised Detection

CVPR 2018poster

Weakly supervised object detection has attracted extensive research efforts in recent years. Without the need of annotating bounding boxes, the existing methods usually follow a two/multi-stage pipeline with an online compulsive stage to extract object proposals, which is an order of magnitude slowe…

Cited by 93SourcePDFScholar
2018

Global Gated Mixture of Second-order Pooling for Improving Deep Convolutional Neural Networks

NeurIPS 2018poster

In most of existing deep convolutional neural networks (CNNs) for classification, global average (first-order) pooling (GAP) has become a standard module to summarize activations of the last convolution layer as final representation for prediction. Recent researches show integration of higher-order…

2018

Joint Representation and Truncated Inference Learning for Correlation Filter based Tracking

ECCV 2018poster

Correlation filter (CF) based trackers generally include two modules, i.e., feature representation and on-line model adaptation. In existing off-line deep learning models for CF trackers, the model adaptation usually is either abandoned or has closed-form solution to make it feasible to learn deep r…

2018

Learning Convolutional Networks for Content-Weighted Image Compression

CVPR 2018poster

Lossy image compression is generally formulated as a joint rate-distortion optimization problem to learn encoder, quantizer, and decoder. Due to the non-differentiable quantizer and discrete entropy estimation, it is very challenging to develop a convolutional network (CNN)-based image compression…

Cited by 490SourcePDFScholar
2018

Learning Spatial-Temporal Regularized Correlation Filters for Visual Tracking

CVPR 2018poster

Discriminative Correlation Filters (DCF) are efficient in visual tracking but suffer from unwanted boundary effects. Spatially Regularized DCF (SRDCF) has been suggested to resolve this issue by enforcing spatial penalty on DCF coefficients, which, inevitably, improves the tracking performance at th…

2018

Learning Warped Guidance for Blind Face Restoration

ECCV 2018poster

This paper studies the problem of blind face restoration from an unconstrained blurry, noisy, low-resolution, or compressed image (i.e., degraded observation). For better recovery of fine facial details, we modify the problem setting by taking both the degraded observation and a high-quality guided…

2018

Learning a Single Convolutional Super-Resolution Network for Multiple Degradations

CVPR 2018poster

Recent years have witnessed the unprecedented success of deep convolutional neural networks (CNNs) in single image super-resolution (SISR). However, existing CNN-based SISR methods mostly assume that a low-resolution (LR) image is bicubicly downsampled from a high-resolution (HR) image, thus inevita…

2018

Multi-Scale Location-Aware Kernel Representation for Object Detection

CVPR 2018poster

Although Faster R-CNN and its variants have shown promising performance in object detection, they only exploit simple first order representation of object proposals for final classification and regression. Recent classification methods demonstrate that the integration of high order statistics into d…

2018

PM-GANs: Discriminative Representation Learning for Action Recognition Using Partial-modalities

ECCV 2018poster

Data of different modalities generally convey complimentary but heterogeneous information, and a more discriminative representation is often preferred by combining multiple data modalities like the RGB and infrared features. However in reality, obtaining both data channels is challenging due to many…

Cited by 32SourcePDFScholar
2018

Shift-Net: Image Inpainting via Deep Feature Rearrangement

ECCV 2018poster

Deep convolutional networks (CNNs) have exhibited their potential in image inpainting for producing plausible results. However, in most existing methods, e.g., context encoder, the missing parts are predicted by propagating the surrounding convolutional features through a fully connected layer, whic…

2018

VITAL: VIsual Tracking via Adversarial Learning

CVPR 2018poster

The tracking-by-detection framework consists of two stages, i.e., drawing samples around the target object in the first stage and classifying each sample as the target object or as background in the second stage. The performance of existing tracking-by-detection trackers using deep classification ne…

Cited by 654SourcePDFScholar
2018

Weakly-supervised Video Summarization using Variational Encoder-Decoder and Web Prior

ECCV 2018poster

Video summarization is a challenging under-constrained problem because the underlying summary of a single video strongly depends on users' subjective understandings. Data-driven approaches, such as deep neural networks, can deal with the ambiguity inherent in this task to some extent, but it is extr…

2017

Higher-Order Integration of Hierarchical Convolutional Activations for Fine-Grained Visual Categorization

ICCV 2017poster

The success of fine-grained visual categorization (FGVC) extremely relies on the modeling of appearance and interactions of various semantic parts. This makes FGVC very challenging because: (i) part annotation and detection require expert guidance and are very expensive; (ii) parts are of different…

Cited by 240PDFScholar
2017

Is Second-Order Information Helpful for Large-Scale Visual Recognition?

ICCV 2017poster

By stacking layers of convolution and nonlinearity, convolutional networks (ConvNets) effectively learn from low-level to high-level features and discriminative representations. Since the end goal of large-scale recognition is to delineate complex boundaries of thousands of classes, adequate explora…

Cited by 353PDFcodeScholar
2017

Joint Convolutional Analysis and Synthesis Sparse Representation for Single Image Layer Separation

ICCV 2017poster

Analysis sparse representation (ASR) and synthesis sparse representation (SSR) are two representative approaches for sparsity-based image modeling. An image is described mainly by the non-zero coefficients in SSR, while it is characterized by the indices of zeros in ASR. To exploit the complementary…

Cited by 258PDFScholar
2017

Learning Dynamic Guidance for Depth Image Enhancement

CVPR 2017poster

The depth images acquired by consumer depth sensors (e.g., Kinect and ToF) usually are of low resolution and insufficient quality. One natural solution is to incorporate with high resolution RGB camera for exploiting their statistical correlation. However, most existing methods are intuitive and lim…

Cited by 108PDFScholar
2017

Mind the Class Weight Bias: Weighted Maximum Mean Discrepancy for Unsupervised Domain Adaptation

CVPR 2017poster

In domain adaptation, maximum mean discrepancy (MMD) has been widely adopted as a discrepancy metric between the distributions of source and target domains. However, existing MMD-based domain adaptation methods generally ignore the changes of class prior distributions, i.e., class weight bias across…

Cited by 777PDFcodeScholar
2016

A Probabilistic Collaborative Representation Based Approach for Pattern Classification

CVPR 2016poster

Conventional representation based classifiers, ranging from the classical nearest neighbor classifier and nearest subspace classifier to the recently developed sparse representation based classifier (SRC) and collaborative representation based classifier (CRC), are essentially distance based classif…

Cited by 337PDFScholar
2016

Deep Structured Scene Parsing by Learning With Image Descriptions

CVPR 2016oral

This paper addresses the problem of structured scene parsing, i.e., parsing the input scene into a configuration including hierarchical semantic objects with their interaction relations. We propose a deep architecture consisting of two networks: i) a convolutional neural network (CNN) extracting the…

Cited by 40PDFScholar
2016

Dictionary Pair Classifier Driven Convolutional Neural Networks for Object Detection

CVPR 2016poster

Feature representation and object category classification are two key components of most object detection methods. While significant improvements have been achieved for deep feature representation learning, traditional SVM/softmax classifiers remain the dominant methods for final object category cla…

Cited by 53PDFScholar
2016

Joint Learning of Single-Image and Cross-Image Representations for Person Re-Identification

CVPR 2016poster

Person re-identification has been usually solved as either the matching of single-image representation (SIR) or the classification of cross-image representation (CIR). In this work, we exploit the connection between these two categories of methods, and propose a joint learning framework to unify SIR…

Cited by 518PDFScholar
2016

Multispectral Images Denoising by Intrinsic Tensor Sparsity Regularization

CVPR 2016spotlight

Multispectral images (MSI) can help deliver more faithful representation for real scenes than the traditional image system, and enhance the performance of many computer vision tasks. In real cases, however, an MSI is always corrupted by various noises. In this paper, we propose a new tensor-based de…

Cited by 280PDFScholar
2016

RAID-G: Robust Estimation of Approximate Infinite Dimensional Gaussian With Application to Material Recognition

CVPR 2016poster

Infinite dimensional covariance descriptors can provide richer and more discriminative information than their low dimensional counterparts. In this paper, we propose a novel image descriptor, namely, robust approximate infinite dimensional Gaussian (RAID-G). The challenges of RAID-G mainly lie on tw…

Cited by 75PDFScholar
2016

The Solution Path Algorithm for Identity-Aware Multi-Object Tracking

CVPR 2016spotlight

We propose an identity-aware multi-object tracker based on the solution path algorithm. Our tracker not only produces identity-coherent trajectories based on cues such as face recognition, but also has the ability to pinpoint potential tracking errors. The tracker is formulated as a quadratic optimi…

Cited by 51PDFcodeScholar
2015

Convolutional Sparse Coding for Image Super-Resolution

ICCV 2015poster

Sparse coding (SC) plays an important role in versatile computer vision applications such as image super-resolution (SR). Most of the previous SC based SR methods partition the image into overlapped patches, and process each patch separately. These methods, however, ignore the consistency of pixels…

Cited by 448PDFScholar
2015

Discriminative Learning of Iteration-Wise Priors for Blind Deconvolution

CVPR 2015poster

The maximum a posterior (MAP)-based blind deconvolution framework generally involves two stages: blur kernel estimation and non-blind restoration. For blur kernel estimation, sharp edge prediction and carefully designed image priors are vital to the success of MAP. In this paper, we propose a blind…

Cited by 49SourcePDFScholar
2015

Patch Group Based Nonlocal Self-Similarity Prior Learning for Image Denoising

ICCV 2015poster

Patch based image modeling has achieved a great success in low level vision such as image denoising. In particular, the use of image nonlocal self-similarity (NSS) prior, which refers to the fact that a local patch often has many nonlocal similar patches to it across the image, has significantly enh…

Cited by 472PDFScholar
2015

SOLD: Sub-Optimal Low-rank Decomposition for Efficient Video Segmentation

CVPR 2015poster

This paper investigates how to perform robust and efficient unsupervised video segmentation while suppressing the effects of data noises and/or corruptions. We propose a general algorithm, called Sub-Optimal Low-rank Decomposition (SOLD), which pursues the low-rank representation for video segmentat…

Cited by 54SourcePDFScholar