← Search

Ming-Ming Cheng

123 accepted papers

2026

DenoDet V2: Phase-Amplitude Cross Denoising for SAR Object Detection

AAAI 2026technical

One of the primary challenges in Synthetic Aperture Radar (SAR) object detection lies in the pervasive influence of coherent noise. As a common practice, most existing methods, whether handcrafted approaches or deep learning-based methods, employ the analysis or enhancement of object spatial-domain

Cited by 0SourcePDFScholar
2026

Direct 3D-Aware Object Insertion via Decomposed Visual Proxies

ICML 2026poster

Object insertion aims to seamlessly composite a reference object into a specified region of a background image. Recent diffusion-based methods achieve high visual quality but formulate insertion as a simple 2D inpainting task, providing no explicit control over the object’s 3D pose and limiting thei…

Cited by 0SourceScholar
2026

GeoAgent: Learning to Geolocate Everywhere with Reinforced Geographic Characteristics

CVPR 2026

This paper presents GeoAgent, a model capable of reasoning closely with humans and deriving fine-grained address conclusions. Previous RL-based methods have achieved breakthroughs in performance and interpretability but still remain concerns because of their reliance on AI-generated chain-of-thought

Cited by 0SourceScholar
2026

Infinite-World: Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory

ICML 2026poster

We propose **Infinite-World**, a robust interactive world model capable of maintaining coherent visual memory over **1000+ frames** in complex real-world environments. While existing world models can be efficiently optimized on synthetic data with perfect ground-truth, they lack an effective trainin…

Cited by 0SourceScholar
2026

NAIPv2: Debiased Pairwise Learning for Efficient Paper Quality Estimation

ICLR 2026poster

The ability to estimate the quality of scientific papers is central to how both humans and AI systems will advance scientific knowledge in the future. However, existing LLM-based estimation methods suffer from high inference cost, whereas the faster direct score regression approach is limited by sca…

Cited by 0SourcecodeScholar
2026

ORION: Decoupling and Alignment for Unified Autoregressive Understanding and Generation

ICLR 2026poster

Unified multimodal Large Language Models (MLLMs) hold great promise for seamlessly integrating understanding and generation. However, monolithic autoregressive architectures, despite their elegance and conversational fluency, suffer from a fundamental semantic–structural conflict: optimizing for low…

Cited by 0SourceScholar
2026

Predictive Regularization Against Visual Representation Degradation in Multimodal Large Language Models

CVPR 2026

While Multimodal Large Language Models (MLLMs) excel at vision-language tasks, the cost of their language-driven training on internal visual foundational competence remains unclear. In this paper, we conduct a detailed diagnostic analysis to unveil a pervasive issue: visual representation degradatio

Cited by 0SourceScholar
2026

RoadGIE: Towards A Global-Scale Aerial Benchmark for Generalizable Interactive Road Extraction

CVPR 2026

Accurate road segmentation from aerial imagery is fundamental to many geospatial applications. However, existing datasets often suffer from limited scene diversity, low semantic granularity, and poor structural continuity, restricting their generalization across environments. To address these challe

Cited by 0SourcecodeScholar
2026

Rotation Invariant and Symmetry Aware Pixel Difference Network for Remote Sensing Object Detection

CVPR 2026

Recent advancements in remote sensing object detection have predominantly focused on oriented bounding box design and small object feature enhancement, while often overlooking the intrinsic geometric properties of remote sensing images, such as rotation invariance and structural symmetry. Many aeria

Cited by 0SourcecodeScholar
2026

SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection

AAAI 2026technical

With the rapid advancement of remote sensing technology, high-resolution multi-modal imagery is now more widely accessible. Conventional object detection models are trained on a single dataset, often restricted to a specific imaging modality and annotation format. However, such an approach overlooks

Cited by 0SourcePDFScholar
2026

Strip R-CNN: Large Strip Convolution for Remote Sensing Object Detection

AAAI 2026technical

In this paper, we show that current approaches using large square kernels or transformer-based global modeling aggregate contextual information uniformly across spatial dimensions, leading to feature dilution and localization errors for elongated targets. To mitigate this issue, we propose Strip R-C

Cited by 0SourcePDFScholar
2026

The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment

CVPR 2026

Previous works have explored various customized generation tasks given a reference image, but they still face limitations in generating consistent fine-grained details. In this paper, our aim is to solve the inconsistency problem of generated images by applying a reference-guided post-editing approa

Cited by 0SourcecodeScholar
2026

Time-Aware One Step Diffusion Network for Real-World Image Super-Resolution

CVPR 2026

Diffusion-based real-world image super-resolution (Real-ISR) methods have demonstrated impressive performance. To achieve efficient Real-ISR, many works employ Variational Score Distillation (VSD) to distill a pre-trained stable-diffusion (SD) model for one-step SR with a fixed timestep. However, si

Cited by 0SourcecodeScholar
2026

Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted Pretraining

ICML 2026poster

Heterogeneous multi-modal remote sensing object detection aims to accurately detect objects from diverse sensors (e.g., RGB, SAR, Infrared). Existing approaches largely adopt a late alignment paradigm, in which modality alignment and task-specific optimization are entangled during downstream fine-tu…

Cited by 0SourceScholar
2026

WOW-Seg: A Word-free Open World Segmentation Model

ICLR 2026poster

Open world image segmentation aims to achieve precise segmentation and semantic understanding of targets within images by addressing the infinitely open set of object categories encountered in the real world. However, traditional closed-set segmentation approaches struggle to adapt to complex open…

Cited by 0SourceScholar
2025

$InterLCM$: Low-Quality Images as Intermediate States of Latent Consistency Models for Effective Blind Face Restoration

ICLR 2025poster

Diffusion priors have been used for blind face restoration (BFR) by fine-tuning diffusion models (DMs) on restoration datasets to recover low-quality images. However, the naive application of DMs presents several key limitations. (i) The diffusion prior has inferior semantic consistency (e.g., ID,…

Cited by 1SourcePDFScholar
2025

AR-1-to-3: Single Image to Consistent 3D Object via Next-View Prediction

ICCV 2025poster

Novel view synthesis (NVS) is a cornerstone for image-to-3d creation. However, existing works still struggle to maintain consistency between the generated views and the input views, especially when there is a significant camera pose difference, leading to poor-quality 3D geometries and textures. We…

2025

Advancing Textual Prompt Learning with Anchored Attributes

ICCV 2025poster

Textual-based prompt learning methods primarily employ multiple learnable soft prompts and hard class tokens in a cascading manner as text inputs, aiming to align image and text (category) spaces for downstream tasks. However, current training is restricted to aligning images with predefined known c…

2025

Anchor Token Matching: Implicit Structure Locking for Training-free AR Image Editing

ICCV 2025poster

Text-to-image generation has seen groundbreaking advancements with diffusion models, enabling high-fidelity synthesis and precise image editing through cross-attention manipulation. Recently, autoregressive (AR) models have re-emerged as powerful alternatives, leveraging next-token generation to mat…

2025

AngleRoCL: Angle-Robust Concept Learning for Physically View-Invariant Adversarial Patches

NeurIPS 2025poster

Cutting-edge works have demonstrated that text-to-image (T2I) diffusion models can generate adversarial patches that mislead state-of-the-art object detectors in the physical world, revealing detectors' vulnerabilities and risks. However, these methods neglect the T2I patches' attack effectiveness w…

Cited by 0SourcecodeScholar
2025

DFormerv2: Geometry Self-Attention for RGBD Semantic Segmentation

CVPR 2025poster

Recent advances in scene understanding benefit a lot from depth maps because of the 3D geometry information, especially in complex conditions (e.g., low light and overexposed). Existing approaches encode depth maps along with RGB images and perform feature fusion between them to enable more robust p…

2025

DIPO: Dual-State Images Controlled Articulated Object Generation Powered by Diverse Data

NeurIPS 2025poster

We present **DIPO**, a novel framework for the controllable generation of articulated 3D objects from a pair of images: one depicting the object in a resting state and the other in an articulated state. Compared to the single-image approach, our dual-image input imposes only a modest overhead for da…

Cited by 0SourcecodeScholar
2025

DISTA-Net: Dynamic Closely-Spaced Infrared Small Target Unmixing

ICCV 2025poster

Resolving closely-spaced small targets in dense clusters presents a significant challenge in infrared imaging, as the overlapping signals hinder precise determination of their quantity, sub-pixel positions, and radiation intensities. While deep learning has advanced the field of infrared small targe…

2025

DepthVanish: Optimizing Adversarial Interval Structures for Stereo-Depth-Invisible Patches

NeurIPS 2025poster

Stereo depth estimation is a critical task in autonomous driving and robotics, where inaccuracies (such as misidentifying nearby objects as distant) can lead to dangerous situations. Adversarial attacks against stereo depth estimation can help revealing vulnerabilities before deployment. Previous wo…

Cited by 0SourcecodeScholar
2025

From Words to Worth: Newborn Article Impact Prediction with LLM

AAAI 2025technical

Predicting the future impact of newly published articles is pivotal for advancing scientific discovery in an era of unprecedented scholarly expansion. This paper introduces a promising approach, leveraging the capabilities of LLMs to predict the future impact of newborn articles solely based on titl…

Cited by 2SourcePDFScholar
2025

GET: Unlocking the Multi-modal Potential of CLIP for Generalized Category Discovery

CVPR 2025poster

Given unlabelled datasets containing both old and new categories, generalized category discovery (GCD) aims to accurately discover new classes while correctly classifying old classes. Current GCD methods only use a single visual modality of information, resulting in poor classification of visually s…

2025

KAC: Kolmogorov-Arnold Classifier for Continual Learning

CVPR 2025highlight

Continual learning requires models to train continuously across consecutive tasks without forgetting. Most existing methods utilize linear classifiers, which struggle to maintain a stable classification space while learning new tasks. Inspired by the success of Kolmogorov-Arnold Networks (KAN) in pr…

2025

Knowledge Graph Enhanced Generative Multi-modal Models for Class-Incremental Learning

NeurIPS 2025poster

Continual learning in computer vision faces the critical challenge of catastrophic forgetting, where models struggle to retain prior knowledge while adapting to new tasks. Although recent studies have attempted to leverage the generalization capabilities of pre-trained models to mitigate overfitting…

Cited by 0SourceScholar
2025

Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation

ICLR 2025spotlight

Few-shot 3D point cloud segmentation (FS-PCS) aims at generalizing models to segment novel categories with minimal annotated support samples. While existing FS-PCS methods have shown promise, they primarily focus on unimodal point cloud inputs, overlooking the potential benefits of leveraging multim…

2025

OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation

NeurIPS 2025poster

Recent research on representation learning has proved the merits of multi-modal clues for robust semantic segmentation. Nevertheless, a flexible pretrain-and-finetune pipeline for multiple visual modalities remains unexplored. In this paper, we propose a novel multi-modal learning framework, termed…

Cited by 0SourcecodeScholar
2025

One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt

ICLR 2025spotlight

Text-to-image generation models can create high-quality images from input prompts. However, they struggle to support the consistent generation of identity-preserving requirements for storytelling. Existing approaches to this problem typically require extensive training in large datasets or additiona…

2025

RSAR: Restricted State Angle Resolver and Rotated SAR Benchmark

CVPR 2025poster

Rotated object detection has made significant progress in the optical remote sensing. However, advancements in the Synthetic Aperture Radar (SAR) field are laggard behind, primarily due to the absence of a large-scale dataset. Annotating such a dataset is inefficient and costly. A promising solution…

2025

Re-Aligning Language to Visual Objects with an Agentic Workflow

ICLR 2025poster

Language-based object detection (LOD) aims to align visual objects with language expressions. A large amount of paired data is utilized to improve LOD model generalizations. During the training process, recent studies leverage vision-language models (VLMs) to automatically generate human-like expres…

Cited by 0SourcePDFScholar
2025

Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think

NeurIPS 2025oral

REPA and its variants effectively mitigate training challenges in diffusion models by incorporating external visual representations from pretrained models, through alignment between the noisy hidden projections of denoising networks and foundational clean image representations. We argue that the ext…

Cited by 0SourcecodeScholar
2025

Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment

ICCV 2025poster

Semantic segmentation is fundamental to vision systems requiring pixel-level scene understanding, yet deploying it on resource-constrained devices demands efficient architectures. Although existing methods achieve real-time inference through lightweight designs, we reveal their inherent limitation:…

2025

Revisiting End-to-End Learning with Slide-level Supervision in Computational Pathology

NeurIPS 2025poster

Pre-trained encoders for offline feature extraction followed by multiple instance learning (MIL) aggregators have become the dominant paradigm in computational pathology (CPath), benefiting cancer diagnosis and prognosis. However, performance limitations arise from the absence of encoder fine-tuning…

Cited by 0SourcecodeScholar
2025

TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction

ICCV 2025poster

We present TAR3D, a novel framework that consists of a 3D-aware Vector Quantized-Variational AutoEncoder (VQVAE) and a Generative Pre-trained Transformer (GPT) to generate high-quality 3D assets. The core insight of this work is to migrate the multimodal unification and promising learning capabiliti…

2025

TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs

NeurIPS 2025poster

This paper introduces TempSamp-R1, a new reinforcement fine-tuning framework designed to improve the effectiveness of adapting multimodal large language models (MLLMs) to video temporal grounding tasks. We reveal that existing reinforcement learning methods, such as Group Relative Policy Optimizatio…

Cited by 0SourcecodeScholar
2025

Towards RAW Object Detection in Diverse Conditions

CVPR 2025highlight

Existing object detection methods often consider sRGB input, which was compressed from RAW data using ISP originally designed for visualization. However, such compression might lose crucial information for detection, especially under complex light and weather conditions. We introduce the AODRaw data…

2025

Unbiased Region-Language Alignment for Open-Vocabulary Dense Prediction

ICCV 2025poster

Pre-trained vision-language models (VLMs), such as CLIP, have demonstrated impressive zero-shot recognition capability, but still underperform in dense prediction tasks. Self-distillation recently is emerging as a promising approach for fine-tuning VLMs to better adapt to local regions without requi…

2025

VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning

ICCV 2025poster

Recent advances in diffusion models have significantly advanced image generation; however, existing models remain task-specific, limiting their efficiency and generalizability. While universal models attempt to address these limitations, they face critical challenges, including generalizable instruc…

2024

Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation

ICML 2024poster

Pre-trained vision-language models, e.g., CLIP, have been successfully applied to zero-shot semantic segmentation. Existing CLIP-based approaches primarily utilize visual features from the last layer to align with text embeddings, while they neglect the crucial information in intermediate layers tha…

2024

CorrMatch: Label Propagation via Correlation Matching for Semi-Supervised Semantic Segmentation

CVPR 2024poster

This paper presents a simple but performant semi-supervised semantic segmentation approach called CorrMatch. Previous approaches mostly employ complicated training strategies to leverage unlabeled data but overlook the role of correlation maps in modeling the relationships between pairs of locations…

2024

CrossKD: Cross-Head Knowledge Distillation for Object Detection

CVPR 2024poster

Knowledge Distillation (KD) has been validated as an effective model compression technique for learning compact object detectors. Existing state-of-the-art KD methods for object detection are mostly based on feature imitation. In this paper we present a general and effective prediction mimicking dis…

2024

DFormer: Rethinking RGBD Representation Learning for Semantic Segmentation

ICLR 2024poster

We present DFormer, a novel RGB-D pretraining framework to learn transferable representations for RGB-D segmentation tasks. DFormer has two new key innovations: 1) Unlike previous works that encode RGB-D information with RGB pretrained backbone, we pretrain the backbone using image-depth pairs from…

Cited by 58SourcePDFScholar
2024

Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference

NeurIPS 2024poster

One of the main drawback of diffusion models is the slow inference time for image generation. Among the most successful approaches to addressing this problem are distillation methods. However, these methods require considerable computational resources. In this paper, we take another approach to diff…

2024

Fine-Grained Knowledge Selection and Restoration for Non-exemplar Class Incremental Learning

AAAI 2024technical

Non-exemplar class incremental learning aims to learn both the new and old tasks without accessing any training data from the past. This strict restriction enlarges the difficulty of alleviating catastrophic forgetting since all techniques can only be applied to current task data. Considering this c…

2024

Generative Multi-modal Models are Good Class Incremental Learners

CVPR 2024poster

In class incremental learning (CIL) scenarios the phenomenon of catastrophic forgetting caused by the classifier's bias towards the current task has long posed a significant challenge. It is mainly caused by the characteristic of discriminative models. With the growing popularity of the generative m…

2024

Let’s Start Over: Retraining with Selective Samples for Generalized Category Discovery

IJCAI 2024poster

Generalized Category Discovery (GCD) presents a realistic and challenging problem in open-world learning. Given a par- tially labeled dataset, GCD aims to categorize unlabeled data by leveraging visual knowledge from the labeled data, where the unlabeled data includes both known and unknown clas…

2024

OPUS: Occupancy Prediction Using a Sparse Set

NeurIPS 2024poster

Occupancy prediction, aiming at predicting the occupancy status within voxelized 3D environment, is quickly gaining momentum within the autonomous driving community. Mainstream occupancy prediction works first discretize the 3D environment into voxels, then perform classification on such dense grids…

2024

PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding

CVPR 2024poster

Recent advances in text-to-image generation have made remarkable progress in synthesizing realistic human photos conditioned on given text prompts. However existing personalized generation methods cannot simultaneously satisfy the requirements of high efficiency promising identity (ID) fidelity and…

2024

SARDet-100K: Towards Open-Source Benchmark and ToolKit for Large-Scale SAR Object Detection

NeurIPS 2024spotlight

Synthetic Aperture Radar (SAR) object detection has gained significant attention recently due to its irreplaceable all-weather imaging capabilities. However, this research field suffers from both limited public datasets (mostly comprising <2K images with only mono-category objects) and inaccessible…

2024

StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation

NeurIPS 2024spotlight

For recent diffusion-based generative models, maintaining consistent content across a series of generated images, especially those containing subjects and complex details, presents a significant challenge. In this paper, we propose a simple but effective self-attention mechanism, termed Consistent S…

2024

Task-Adaptive Saliency Guidance for Exemplar-free Class Incremental Learning

CVPR 2024poster

Exemplar-free Class Incremental Learning (EFCIL) aims to sequentially learn tasks with access only to data from the current one. EFCIL is of interest because it mitigates concerns about privacy and long-term storage of data while at the same time alleviating the problem of catastrophic forgetting in…

2024

TeMO: Towards Text-Driven 3D Stylization for Multi-Object Meshes

CVPR 2024poster

Recent progress in the text-driven 3D stylization of a single object has been considerably promoted by CLIP-based methods. However the stylization of multi-object 3D scenes is still impeded in that the image-text pairs used for pre-training CLIP mostly consist of an object. Meanwhile the local detai…

Cited by 9SourcePDFScholar
2024

Token Merging for Training-Free Semantic Binding in Text-to-Image Synthesis

NeurIPS 2024poster

Although text-to-image (T2I) models exhibit remarkable generation capabilities, they frequently fail to accurately bind semantically related objects or attributes in the input prompts; a challenge termed semantic binding. Previous approaches either involve intensive fine-tuning of the entire T2I mod…

2024

Towards Stable 3D Object Detection

ECCV 2024poster

"In autonomous driving, the temporal stability of 3D object detection greatly impacts the driving safety. However, the detection stability cannot be accessed by existing metrics such as mAP and MOTA, and consequently is less explored by the community. To bridge this gap, this work proposes (), a new…

2024

Traffic Scene Parsing through the TSP6K Dataset

CVPR 2024poster

Traffic scene perception in computer vision is a critically important task to achieve intelligent cities. To date most existing datasets focus on autonomous driving scenes. We observe that the models trained on those driving datasets often yield unsatisfactory results on traffic monitoring scenes. H…

2023

AMT: All-Pairs Multi-Field Transforms for Efficient Frame Interpolation

CVPR 2023poster

We present All-Pairs Multi-Field Transforms (AMT), a new network architecture for video frame interpolation. It is based on two essential designs. First, we build bidirectional correlation volumes for all pairs of pixels and use the predicted bilateral flows to retrieve correlations for updating bot…

2023

Large Selective Kernel Network for Remote Sensing Object Detection

ICCV 2023poster

Recent research on remote sensing object detection has largely focused on improving the representation of oriented bounding boxes but has overlooked the unique prior knowledge presented in remote sensing scenarios. Such prior knowledge can be useful because tiny remote sensing objects may be mistake…

Cited by 467PDFcodeScholar
2023

Looking Through the Glass: Neural Surface Reconstruction Against High Specular Reflections

CVPR 2023poster

Neural implicit methods have achieved high-quality 3D object surfaces under slight specular highlights. However, high specular reflections (HSR) often appear in front of target objects when we capture them through glasses. The complex ambiguity in these scenes violates the multi-view consistency, th…

2023

Masked Autoencoders are Efficient Class Incremental Learners

ICCV 2023poster

Class Incremental Learning (CIL) aims to sequentially learn new classes while avoiding catastrophic forgetting of previous knowledge. We propose to use Masked Autoencoders (MAEs) as efficient learners for CIL. MAEs were originally designed to learn useful representations through reconstr…

Cited by 17PDFcodeScholar
2023

Masked Diffusion Transformer is a Strong Image Synthesizer

ICCV 2023poster

Despite its success in image synthesis, we observe that diffusion probabilistic models (DPMs) often lack contextual reasoning ability to learn the relations among object parts in an image, leading to a slow learning process. To solve this issue, we propose a Masked Diffusion Transformer (MDT) that i…

Cited by 130PDFcodeScholar
2023

SLAN: Self-Locator Aided Network for Vision-Language Understanding

ICCV 2023poster

Learning fine-grained interplay between vision and language contributes to a more accurate understanding for Vision-Language tasks. However, it remains challenging to extract key image regions according to the texts for semantic alignments. Most existing works are either limited by text-agnostic an…

Cited by 0PDFScholar
2023

SRFormer: Permuted Self-Attention for Single Image Super-Resolution

ICCV 2023poster

Previous works have shown that increasing the window size for Transformer-based image super-resolution models (e.g., SwinIR) can significantly improve the model performance but the computation overhead is also considerable. In this paper, we present SRFormer, a simple but novel method that can enjoy…

Cited by 216PDFcodeScholar
2022

"Restore Globally, Refine Locally: A Mask-Guided Scheme to Accelerate Super-Resolution Networks"

ECCV 2022poster

"Single image super-resolution (SR) has been boosted by deep convolutional neural networks with growing model complexity and computational costs. To deploy existing SR networks onto edge devices, it is necessary to accelerate them for large image (4K) processing. The different areas in an image ofte…

2022

FocusCut: Diving Into a Focus View in Interactive Segmentation

CVPR 2022oral

Interactive image segmentation is an essential tool in pixel-level annotation and image editing. To obtain a high-precision binary segmentation mask, users tend to add interaction clicks around the object details, such as edges and holes, for efficient refinement. Current methods regard these repair…

Cited by 75PDFcodeScholar
2022

Localization Distillation for Dense Object Detection

CVPR 2022poster

Knowledge distillation (KD) has witnessed its powerful capability in learning compact models in object detection. Previous KD methods for object detection mostly focus on imitating deep features within the imitation regions instead of logit mimicking on classification due to the inefficiency in dist…

Cited by 239PDFcodeScholar
2022

Long-Tailed Class Incremental Learning

ECCV 2022poster

"In class incremental learning (CIL) a model must learn new classes in a sequential manner without forgetting old ones. However, conventional CIL methods consider a balanced distribution for each new task, which ignores the prevalence of long-tailed distributions in the real world. In this work we p…

2022

On the Connection between Local Attention and Dynamic Depth-wise Convolution

ICLR 2022spotlight

Vision Transformer (ViT) attains state-of-the-art performance in visual recognition, and the variant, Local Vision Transformer, makes further improvements. The major component in Local Vision Transformer, local attention, performs the attention separately over small local windows. We rephrase local…

2022

Representation Compensation Networks for Continual Semantic Segmentation

CVPR 2022poster

In this work, we study the continual semantic segmentation problem, where the deep neural networks are required to incorporate new classes continually without catastrophic forgetting. We propose to use a structural re-parameterization mechanism, named representation compensation (RC) module, to deco…

Cited by 131PDFcodeScholar
2022

SegNeXt: Rethinking Convolutional Attention Design for Semantic Segmentation

NeurIPS 2022accept

We present SegNeXt, a simple convolutional network architecture for semantic segmentation. Recent transformer-based models have dominated the field of se- mantic segmentation due to the efficiency of self-attention in encoding spatial information. In this paper, we show that convolutional attention…

2022

Towards an End-to-End Framework for Flow-Guided Video Inpainting

CVPR 2022poster

Optical flow, which captures motion information across frames, is exploited in recent video inpainting methods through propagating pixels along its trajectories. However, the hand-crafted flow-based processes in these methods are applied separately to form the whole inpainting pipeline. Thus, they a…

Cited by 186PDFcodeScholar
2022

VQFR: Blind Face Restoration with Vector-Quantized Dictionary and Parallel Decoder

ECCV 2022poster

"Although generative facial prior and geometric prior have recently demonstrated high-quality results for blind face restoration, producing fine-grained facial details faithful to inputs remains a challenging problem. Motivated by the classical dictionary-based methods and the recent vector quantiza…

2021

DOTS: Decoupling Operation and Topology in Differentiable Architecture Search

CVPR 2021poster

Differentiable Architecture Search (DARTS) has attracted extensive attention due to its efficiency in searching for cell structures. DARTS mainly focuses on the operation search and derives the cell topology from the operation weights. However, the operation weights can not indicate the importance o…

Cited by 72PDFcodeScholar
2021

Global2Local: Efficient Structure Search for Video Action Segmentation

CVPR 2021poster

Temporal receptive fields of models play an important role in action segmentation. Large receptive fields facilitate the long-term relations among video clips while small receptive fields help capture the local details. Existing methods construct models with hand-designed receptive fields in layers.…

Cited by 97PDFcodeScholar
2021

Semi-Supervised Learning with Meta-Gradient

AISTATS 2021poster

In this work, we propose a simple yet effective meta-learning algorithm in semi-supervised learning. We notice that most existing consistency-based approaches suffer from overfitting and limited model generalization ability, especially when training with only a small number of labeled data. To allev…

Cited by 11SourcePDFScholar
2021

Structured sparsification with joint optimization of group convolution and channel shuffle

UAI 2021poster

Recent advances in convolutional neural networks (CNNs) usually come with the expense of excessive computational overhead and memory footprint. Network compression aims to alleviate this issue by training compact models with comparable performance. However, existing compression techniques either ent…

2021

Temporal Modulation Network for Controllable Space-Time Video Super-Resolution

CVPR 2021poster

Space-time video super-resolution (STVSR) aims to increase the spatial and temporal resolutions of low-resolution and low-frame-rate videos. Recently, deformable convolution based methods have achieved promising STVSR performance, but they could only infer the intermediate frame pre-defined in the t…

Cited by 113PDFcodeScholar
2021

iNAS: Integral NAS for Device-Aware Salient Object Detection

ICCV 2021poster

Existing salient object detection (SOD) models usually focus on either backbone feature extractors or saliency heads, ignoring their relations. A powerful backbone could still achieve sub-optimal performance with a weak saliency head and vice versa. Moreover, the balance between model performance an…

Cited by 13PDFScholar
2020

Highly Efficient Salient Object Detection with 100K Parameters

ECCV 2020poster

Salient object detection models often demand a considerable amount of computation cost to make precise prediction for each pixel, making them hardly applicable on low-power devices. In this paper, we aim to relieve the contradiction between computation cost and model performance by improving the net…

2020

ICNet: Intra-saliency Correlation Network for Co-Saliency Detection

NeurIPS 2020poster

Intra-saliency and inter-saliency cues have been extensively studied for co-saliency detection (Co-SOD). Model-based methods produce coarse Co-SOD results due to hand-crafted intra- and inter-saliency features. Current data-driven models exploit inter-saliency cues, but undervalue the potential powe…

2020

Improving Convolutional Networks With Self-Calibrated Convolutions

CVPR 2020poster

Recent advances on CNNs are mostly devoted to designing more complex architectures to enhance their representation learning capacity. In this paper, we consider how to improve the basic convolutional feature transformation process of CNNs without tuning the model architectures. To this end, we prese…

Cited by 542PDFcodeScholar
2020

Interactive Image Segmentation With First Click Attention

CVPR 2020poster

In the task of interactive image segmentation, users initially click one point to segment the main body of the target object and then provide more points on mislabeled regions iteratively for a precise segmentation. Existing methods treat all interaction points indiscriminately, ignoring the differe…

Cited by 201PDFScholar
2020

Taking a Deeper Look at Co-Salient Object Detection

CVPR 2020poster

Co-salient object detection (CoSOD) is a newly emerging and rapidly growing branch of salient object detection (SOD), which aims to detect the co-occurring salient objects in multiple images. However, existing CoSOD datasets often have a serious data bias, which assumes that each group of images con…

Cited by 100PDFScholar
2020

VecRoad: Point-Based Iterative Graph Exploration for Road Graphs Extraction

CVPR 2020poster

Extracting road graphs from aerial images automatically is more efficient and costs less than from field acquisition. This can be done by a post-processing step that vectorizes road segmentation predicted by CNN, but imperfect predictions will result in road graphs with low connectivity. On the othe…

Cited by 117PDFScholar
2019

A Simple Pooling-Based Design for Real-Time Salient Object Detection

CVPR 2019poster

We solve the problem of salient object detection by investigating how to expand the role of pooling in convolutional neural networks. Based on the U-shape architecture, we first build a global guidance module (GGM) upon the bottom-up pathway, aiming at providing layers at different feature levels th…

Cited by 1257PDFScholar
2019

An Iterative and Cooperative Top-Down and Bottom-Up Inference Network for Salient Object Detection

CVPR 2019poster

This paper presents a salient object detection method that integrates both top-down and bottom-up saliency inference in an iterative and cooperative manner. The top-down process is used for coarse-to-fine saliency estimation, where high-level saliency is gradually integrated with finer lower-layer f…

Cited by 259PDFScholar
2019

Contrast Prior and Fluid Pyramid Integration for RGBD Salient Object Detection

CVPR 2019poster

The large availability of depth sensors provides valuable complementary information for salient object detection (SOD) in RGBD images. However, due to the inherent difference between RGB and depth information, extracting features from the depth channel using ImageNet pre-trained backbone models and…

Cited by 451PDFScholar
2019

EGNet: Edge Guidance Network for Salient Object Detection

ICCV 2019poster

Fully convolutional neural networks (FCNs) have shown their advantages in the salient object detection task. However, most existing FCNs-based methods still suffer from coarse object boundaries. In this paper, to solve this problem, we focus on the complementarity between salient edge information an…

Cited by 1302PDFScholar
2019

IP102: A Large-Scale Benchmark Dataset for Insect Pest Recognition

CVPR 2019oral

Insect pests are one of the main factors affecting agricultural product yield. Accurate recognition of insect pests facilitates timely preventive measures to avoid economic losses. However, the existing datasets for the visual classification task mainly focus on common objects, e.g., flowers and do…

Cited by 519PDFcodeScholar
2019

Image Inpainting With Learnable Bidirectional Attention Maps

ICCV 2019poster

Most convolutional network (CNN)-based inpainting methods adopt standard convolution to indistinguishably treat valid pixels and holes, making them limited in handling irregular holes and more likely to generate inpainting results with color discrepancy and blurriness. Partial convolution has been s…

Cited by 319PDFcodeScholar
2019

Integral Object Mining via Online Attention Accumulation

ICCV 2019poster

Object attention maps generated by image classifiers are usually used as priors for weakly-supervised segmentation approaches. However, normal image classifiers produce attention only at the most discriminative object parts, which limits the performance of weakly-supervised segmentation task. Theref…

Cited by 279PDFScholar
2019

Joint Acne Image Grading and Counting via Label Distribution Learning

ICCV 2019accepted

Accurate grading of skin disease severity plays a crucial role in precise treatment for patients. Acne vulgaris, the most common skin disease in adolescence, can be graded by evidence-based lesion counting as well as experience-based global estimation in the medical field. However, due to the appear…

2019

Multi-Level Context Ultra-Aggregation for Stereo Matching

CVPR 2019poster

Exploiting multi-level context information to cost volume can improve the performance of learning-based stereo matching methods. In recent years, 3-D Convolution Neural Networks (3-D CNNs) show the advantages in regularizing cost volume but are limited by unary features learning in matching cost com…

Cited by 137PDFScholar
2019

Optimizing the F-Measure for Threshold-Free Salient Object Detection

ICCV 2019poster

Current CNN-based solutions to salient object detection (SOD) mainly rely on the optimization of cross-entropy loss (CELoss). Then the quality of detected saliency maps is often evaluated in terms of F-measure. In this paper, we investigate an interesting issue: can we consistently use the F-measure…

Cited by 90PDFScholar
2019

S4Net: Single Stage Salient-Instance Segmentation

CVPR 2019poster

We consider an interesting problem---salient instance segmentation. Other than producing approximate bounding boxes, our network also outputs high-quality instance-level segments. Taking into account the category-independent property of each target, we design a single stage salient instance segmenta…

Cited by 108PDFcodeScholar
2019

Scoot: A Perceptual Metric for Facial Sketches

ICCV 2019poster

While it is trivial for humans to quickly assess the perceptual similarity between two images, the underlying mechanism are thought to be quite complex. Despite this, the most widely adopted perceptual metrics today, such as SSIM and FSIM, are simple, shallow functions, and fail to consider many fac…

Cited by 57PDFScholar
2019

Shifting More Attention to Video Salient Object Detection

CVPR 2019oral

The last decade has witnessed a growing interest in video salient object detection (VSOD). However, the research community long-term lacked a well-established VSOD dataset representative of real dynamic scenes with high-quality annotations. To address this issue, we elaborately collected a visual-at…

Cited by 561PDFcodeScholar
2019

Unsupervised Scale-consistent Depth and Ego-motion Learning from Monocular Video

NeurIPS 2019poster

Recent work has shown that CNN-based depth and ego-motion estimators can be learned using unlabelled monocular videos. However, the performance is limited by unidentified moving objects that violate the underlying static scene assumption in geometric image reconstruction. More significantly, due to…

2019

Zero-Shot Emotion Recognition via Affective Structural Embedding

ICCV 2019poster

Image emotion recognition attracts much attention in recent years due to its wide applications. It aims to classify the emotional response of humans, where candidate emotion categories are generally defined by specific psychological theories, such as Ekman's six basic emotions. However, with the dev…

Cited by 64PDFScholar
2018

Associating Inter-Image Salient Instances for Weakly Supervised Semantic Segmentation

ECCV 2018poster

Effectively bridging between image level keyword annotations and corresponding image pixels is one of the main challenges in weakly supervised semantic segmentation. In this paper, we use an instance-level salient object detector to automatically generate salient instances (candidate objects) for tr…

Cited by 120SourcePDFScholar
2018

Crowd Counting With Deep Negative Correlation Learning

CVPR 2018poster

Deep convolutional networks (ConvNets) have achieved unprecedented performances on many computer vision tasks. However, their adaptations to crowd counting on single images are still in their infancy and suffer from severe over-fitting. Here we propose a new learning strategy to produce generalizabl…

2018

Revisiting Video Saliency: A Large-Scale Benchmark and a New Model

CVPR 2018poster

In this work, we contribute to video saliency research in two ways. First, we introduce a new benchmark for predicting human eye movements during dynamic scene free-viewing, which is long-time urged in this field. Our dataset, named DHF1K~(Dynamic Human Fixation), consists of 1K high-quality, elabor…

2018

Salient Objects in Clutter: Bringing Salient Object Detection to the Foreground

ECCV 2018poster

We provide a comprehensive evaluation of salient object detection (SOD) models. Our analysis identifies a serious design bias of existing SOD datasets which assumes that each image contains at least one clearly outstanding salient object in low clutter. The design bias has led to a saturated high pe…

Cited by 380SourcePDFScholar
2018

Structured Skip List: A Compact Data Structure for 3D Reconstruction

IROS 2018poster

The model produced by 3D reconstruction algorithm is usually represented by voxels. The management of these voxels is usually divided into two categories: ordered and unordered methods. The ordered method holds too many empty voxels to maintain data order which leads to a low storage efficiency. On…

Cited by 3SourceScholar
2017

Deeply Supervised Salient Object Detection With Short Connections

CVPR 2017poster

Recent progress on saliency detection is substantial, benefiting mostly from the explosive development of Convolutional Neural Networks (CNNs). Semantic segmentation and saliency detection algorithms developed lately have been mostly based on Fully Convolutional Neural Networks (FCNs). There is stil…

Cited by 1892PDFcodeScholar
2017

GMS: Grid-based Motion Statistics for Fast, Ultra-Robust Feature Correspondence

CVPR 2017poster

Incorporating smoothness constraints into feature matching is known to enable ultra-robust matching. However, such formulations are both complex and slow, making them unsuitable for video applications. This paper proposes GMS (Grid-based Motion Statistics), a simple means of encapsulating motion smo…

Cited by 862PDFcodeScholar
2017

Object Region Mining With Adversarial Erasing: A Simple Classification to Semantic Segmentation Approach

CVPR 2017oral

We investigate a principle way to progressively mine discriminative object regions using classification networks to address the weakly-supervised semantic segmentation problems. Classification networks are only responsive to small and sparse discriminative regions from the object of interest, which…

Cited by 1023PDFScholar
2017

Structure-Measure: A New Way to Evaluate Foreground Maps

ICCV 2017spotlight

Foreground map evaluation is crucial for gauging the progress of object segmentation algorithms, in particular in the filed of salient object detection where the purpose is to accurately detect and segment the most salient object in a scene. Several widely-used measures such as Area Under the Curve…

Cited by 1925PDFcodeScholar