← Search

Zheng-Jun Zha

154 accepted papers

2026

Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts

ICML 2026poster

As the computational demands for pre-training Large Language Models (LLMs) continue to surge, the need for efficient training paradigms becomes critical. Despite the vast resources already invested in existing pre-trained checkpoints, these assets often remain under-leveraged due to architectural li…

Cited by 0SourceScholar
2026

E-MaT:Event-oriented Mamba for Egocentric Point Tracking

AAAI 2026technical

Egocentric point tracking aims to localize points on object surfaces from a first-person perspective and serves as a critical step toward embodied intelligence. Recent methods rely on video input, tracking query points through feature matching across consecutive frames. However, these methods strug

Cited by 0SourcePDFScholar
2026

Event-Illumination Collaborative Low-light Image Enhancement with a High-resolution Real-world Dataset

CVPR 2026

Event-based low-light image enhancement (LIE) methods mainly focus on incorporating high dynamic range (HDR) information from events while overlooking the essential global illumination in images and the inherent noise sensitivity of event signals in real-world scenarios. To address these issues, we

Cited by 0SourcecodeScholar
2026

FinPercep-RM: A Fine-grained Reward Model and Co-evolutionary Curriculum for RL-based Real-world Super-Resolution

CVPR 2026

Inspired by the success of Reinforcement Learning with Human Feedback (RLHF) in image generation, recent work has adapted reward-based learning to image super-resolution (ISR) by using Image Quality Assessment (IQA) models as rewards. However, existing IQA models typically output only a single globa

Cited by 0SourcecodeScholar
2026

Gloria: Consistent Character Video Generation via Content Anchors

CVPR 2026

Digital characters are central to modern media, yet generating character videos with long-duration, consistent multi-view appearance and expressive identity remains challenging. Existing approaches either provide insufficient context to preserve identity or leverage non-character-centric information

Cited by 0SourceScholar
2026

Learning to Diversify and Focus: A Reinforcement Framework for Open-Vocabulary HOI Detection

CVPR 2026

Open-Vocabulary Human-Object Interaction (OV-HOI) detection aims to recognize novel HOI categories beyond the training set. Existing OV-HOI detection approaches typically leverage CLIP to extract global visual representations and perform cross-attention between learnable queries and global features

Cited by 0SourceScholar
2026

PSP: Prompt-Guided Self-Training Sampling Policy for Active Prompt Learning

ICLR 2026poster

Active Prompt Learning (APL) using vision-language models (\textit{e.g.}, CLIP) has attracted considerable attention for mitigating the dependence on fully labeled dataset in downstream task adaptation. However, existing methods fail to explicitly leverage prompt to guide sample selection, resulting…

Cited by 0SourcecodeScholar
2026

Pixel to Gaussian: Ultra-Fast Continuous Super-Resolution with 2D Gaussian Modeling

ICLR 2026poster

Arbitrary-scale super-resolution (ASSR) aims to reconstruct high-resolution (HR) images from low-resolution (LR) inputs with arbitrary upsampling factors using a single model, addressing the limitations of traditional SR methods constrained to fixed-scale factors (\textit{e.g.}, $\times$ 2). Recent…

Cited by 0SourcecodeScholar
2026

TOUCH: Text-guided Controllable Generation of Free-Form Hand-Object Interactions

ICLR 2026poster

Hand-object interaction (HOI) is fundamental for humans to express intent. Existing HOI generation research is predominantly confined to fixed grasping patterns, where control is tied to physical priors such as force closure or generic intent instructions, even when expressed through elaborate langu…

Cited by 0SourceScholar
2026

Time-Specialized Event-Image Alignment for Blur-to-Video Decomposition

CVPR 2026

Motion blur is a common degradation in dynamic imaging. Recent studies have moved beyond restoring a single sharp image from a blurred input and instead target blur decomposition: recovering a temporally continuous sharp video sequence from one motion-blurred image. Event cameras, with their microse

Cited by 0SourcecodeScholar
2026

Towards Sequence Modeling Alignment between Tokenizer and Autoregressive Model

ICLR 2026poster

Autoregressive image generation aims to predict the next token based on previous ones. However, this process is challenged by the bidirectional dependencies inherent in conventional image tokenizations, which creates a fundamental misalignment with the unidirectional nature of autoregressive models.…

Cited by 0SourcecodeScholar
2026

Unbiased Gradient Estimation for Event Binning via Functional Backpropagation

ICLR 2026poster

Event-based vision encodes dynamic scenes as asynchronous spatio-temporal spikes called events. To leverage conventional image processing pipelines, events are typically binned into frames. However, binning functions are discontinuous, which truncates gradients at the frame level and forces most eve…

Cited by 0SourcecodeScholar
2026

WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens

CVPR 2026

Recent progress in multimodal large language models (MLLMs) has highlighted the challenge of efficiently bridging pre-trained Vision-Language Models (VLMs) with Diffusion Models. While methods using a fixed number of learnable query tokens offer computational efficiency, they suffer from task genera

Cited by 0SourceScholar
2025

A Lottery Ticket Hypothesis Approach with Sparse Fine-tuning and MAE for Image Forgery Detection and Localization

AAAI 2025technical

The rise in sophisticated image forgery techniques, driven by advancements in image editing and generation, has posed new security challenges. Traditional methods, designed for specific tampering artifacts, struggle with out-of-distribution image forgery detection. In this paper, we propose a shift…

2025

BEVTrack: A Simple and Strong Baseline for 3D Single Object Tracking in Bird's-Eye View

IJCAI 2025

3D Single Object Tracking (SOT) is a fundamental task in computer vision and plays a critical role in applications like autonomous driving. However, existing algorithms often involve complex designs and multiple loss functions, making model training and deployment challenging. Furthermore, their rel

2025

Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning

CVPR 2025poster

Generating detailed captions comprehending text-rich visual content in images has received growing attention for Large Vision-Language Models (LVLMs). However, few studies have developed benchmarks specifically tailored for detailed captions to measure their accuracy and comprehensiveness. In this p…

2025

Boosting Image De-Raining via Central-Surrounding Synergistic Convolution

AAAI 2025technical

Rainy images suffer from quality degradation due to the synergistic effect of rain streaks and accumulation. The rain streaks are anisotropic and show a specific directional arrangement, while the rain accumulation is isotropic and shows a consistent concentration distribution in local regions. This…

Cited by 1SourcePDFScholar
2025

DCTMamba: Advancing JPEG Image Restoration Through Long-Sequence Modeling and Adaptive Frequency Strategy

AAAI 2025technical

Despite the advanced long-sequence modeling of Mamba, which has expanded its applications in image restoration, there remains a lack of exploration combining its strengths with the specific characteristics of JPEG image restoration, where high-frequency components are lost after the Discrete Cosine…

2025

Decouple to Reconstruct: High Quality UHD Restoration via Active Feature Disentanglement and Reversible Fusion

ICCV 2025poster

Ultra-high-definition (UHD) image restoration often faces computational bottlenecks and information loss due to its extremely high resolution. Existing studies based on Variational Autoencoders (VAE) improve efficiency by transferring the image restoration process from pixel space to latent space. H…

Cited by 0SourcePDFScholar
2025

Directing Mamba to Complex Textures: An Efficient Texture-Aware State Space Model for Image Restoration

IJCAI 2025

Image restoration aims to recover details and enhance contrast in degraded images. With the growing demand for high-quality imaging (e.g., 4K and 8K), achieving a balance between restoration quality and computational efficiency has become increasingly critical. Existing methods, primarily based on C

Cited by 0SourcePDFScholar
2025

EF-3DGS: Event-Aided Free-Trajectory 3D Gaussian Splatting

NeurIPS 2025spotlight

Scene reconstruction from casually captured videos has wide real-world applications. Despite recent progress, existing methods relying on traditional cameras tend to fail in high-speed scenarios due to insufficient observations and inaccurate pose estimation. Event cameras, inspired by biological vi…

Cited by 0SourceScholar
2025

EVDM: Event-based Real-world Video Deblurring with Mamba

ICCV 2025poster

Existing event-based video deblurring methods face limitations in extracting and fusing long-range spatiotemporal motion information from events, primarily due to restricted receptive fields or low computational efficiency, resulting in suboptimal deblurring performance.To address these issues, we i…

2025

Enhanced Pansharpening via Quaternion Spatial-Spectral Interactions

ICCV 2025poster

Pansharpening aims to generate high-resolution multispectral (MS) images by fusing panchromatic (PAN) images with corresponding low-resolution MS images. However, many existing methods struggle to fully capture spatial and spectral interactions, limiting their effectiveness. To address this, we prop…

2025

EventMamba: Enhancing Spatio-Temporal Locality with State Space Models for Event-Based Video Reconstruction

AAAI 2025technical

Leveraging its robust linear global modeling capability, Mamba has notably excelled in computer vision. Despite its success, existing Mamba-based vision models have overlooked the nuances of event-driven tasks, especially in video reconstruction. Event-based video reconstruction (EBVR) demands spati…

Cited by 0SourcePDFScholar
2025

Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning

NeurIPS 2025poster

The rapid spread of multimodal misinformation on social media has raised growing concerns, while research on video misinformation detection remains limited due to the lack of large-scale, diverse datasets. Existing methods often overfit to rigid templates and lack deep reasoning over deceptive conte…

Cited by 0SourcecodeScholar
2025

FourierMamba: Fourier Learning Integration with State Space Models for Image Deraining

ICML 2025poster

Image deraining aims to remove rain streaks from rainy images and restore clear backgrounds. Currently, some research that employs the Fourier transform has proved to be effective for image deraining, due to it acting as an effective frequency prior for capturing rain streaks. However, despite there…

Cited by 17SourcePDFScholar
2025

GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding

CVPR 2025poster

Open-Vocabulary 3D object affordance grounding aims to anticipate "action possibilities" regions on 3D objects with arbitrary instructions, which is crucial for robots to generically perceive real scenarios and respond to operational changes. Existing methods focus on combining images or languages t…

2025

HOIMamba: Efficient Mamba-based Disentangled Progressive Learning for HOI Detection

AAAI 2025technical

Human-object interaction (HOI) detection aims to detect the spatial positions of human-object pairs and recognize their interactions. Existing single-branch, two-branch, and three-branch methods are challenging to make an appropriate trade-off on efficiency, multi-task decoupling, and collaborative…

Cited by 0SourcePDFScholar
2025

Hierarchical Knowledge Prompt Tuning for Multi-task Test-Time Adaptation

CVPR 2025poster

Test-time adaptation using vision-language models (such as CLIP) to quickly adjust to distributional shifts of downstream tasks has shown great potential. Despite significant progress, existing methods are still limited to single-task test-time adaptation scenarios and have not effectively explored…

Cited by 0SourcePDFScholar
2025

Improved Video VAE for Latent Video Diffusion Model

CVPR 2025poster

Variational Autoencoder (VAE) aims to compress pixel data into low-dimensional latent space, playing an important role in OpenAI's Sora and other latent video diffusion generation models. While most existing video VAEs inflate a pre-trained image VAE into the 3D causal structure for temporal-spatial…

2025

Latent Harmony: Synergistic Unified UHD Image Restoration via Latent Space Regularization and Controllable Refinement

NeurIPS 2025poster

Ultra-High Definition (UHD) image restoration struggles to balance computational efficiency and detail retention. While Variational Autoencoders (VAEs) offer improved efficiency by operating in the latent space, with the Gaussian variational constraint, this compression preserves semantics but sacri…

Cited by 0SourceScholar
2025

Learnable Frequency Decomposition for Image Forgery Detection and Localization

IJCAI 2025

Concern for image authenticity spurs research in image forgery detection and localization (IFDL). Most deep learning-based methods focus primarily on spatial domain modeling and have not fully explored frequency domain strategies. In this paper, we observe and analyze the frequency characteristic ch

Cited by 0SourcePDFScholar
2025

MATE: Motion-Augmented Temporal Consistency for Event-based Point Tracking

ICCV 2025poster

Tracking Any Point (TAP) plays a crucial role in motion analysis. Video-based approaches rely on iterative local matching for tracking, but they assume linear motion during the blind time between frames, which leads to point loss under large displacements or nonlinear motion. The high temporal resol…

Cited by 0SourcePDFScholar
2025

MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling

CVPR 2025poster

Recent advancements in multi-modal large language models have propelled the development of joint probabilistic models capable of both image understanding and generation. However, we have identified that recent methods suffer from loss of image information during understanding task, due to either ima…

Cited by 11SourcePDFScholar
2025

Neural Fractional Attention Differential Equations

NeurIPS 2025poster

The integration of differential equations with neural networks has created powerful tools for modeling complex dynamics effectively across diverse machine learning applications. While standard integer-order neural ordinary differential equations (ODEs) have shown considerable success, they are limit…

Cited by 0SourcecodeScholar
2025

Optimizing Large Language Model Training Using FP4 Quantization

ICML 2025poster

The growing computational demands of training large language models (LLMs) necessitate more efficient methods. Quantized training presents a promising solution by enabling low-bit arithmetic operations to reduce these costs. While FP8 precision has demonstrated feasibility, leveraging FP4 remains a…

Cited by 8SourcePDFScholar
2025

PAID: Pairwise Angular-Invariant Decomposition for Continual Test-Time Adaptation

NeurIPS 2025poster

Continual Test-Time Adaptation (CTTA) aims to online adapt a pre-trained model to changing environments during inference. Most existing methods focus on exploiting target data, while overlooking another crucial source of information, the pre-trained weights, which encode underutilized domain-invaria…

Cited by 0SourcecodeScholar
2025

PMQ-VE: Progressive Multi-Frame Quantization for Video Enhancement

NeurIPS 2025poster

Multi-frame video enhancement tasks aim to improve the spatial and temporal resolution and quality of video sequences by leveraging temporal information from multiple frames, which are widely used in streaming video processing, surveillance, and generation. Although numerous Transformer-based enhanc…

Cited by 0SourcecodeScholar
2025

QMambaBSR: Burst Image Super-Resolution with Query State Space Model

CVPR 2025poster

Burst super-resolution (BurstSR) aims to reconstruct high-resolution images by fusing subpixel details from multiple low-resolution burst frames. The primary challenge lies in effectively extracting useful information while mitigating the impact of high-frequency noise. Most existing methods rely on…

Cited by 5SourcePDFScholar
2025

Reliable Lifelong Multimodal Editing: Conflict-Aware Retrieval Meets Multi-Level Guidance

NeurIPS 2025poster

The dynamic nature of real-world information demands efficient knowledge editing in multimodal large language models (MLLMs) to ensure continuous knowledge updates. However, existing methods often struggle with precise matching in large-scale knowledge retrieval and lack multi-level guidance for coo…

Cited by 0SourceScholar
2025

SCott: Accelerating Diffusion Models with Stochastic Consistency Distillation

AAAI 2025technical

The iterative sampling procedure employed by diffusion models (DMs) often leads to significant latency. To address this, we propose Stochastic Consistency Distillation (SCott) to enable accelerated text-to-image generation, where high-quality generations can be achieved with just 2-4 sampling steps…

Cited by 2SourcePDFScholar
2025

SIGMAN: Scaling 3D Human Gaussian Generation with Millions of Assets

ICCV 2025poster

3D human digitization has long been a highly pursued yet challenging task. Existing methods aim to generate high-quality 3D digital humans from single or multiple views, but remain primarily constrained by current paradigms and the scarcity of 3D human assets. Specifically, recent approaches fall in…

Cited by 0SourcePDFScholar
2025

Towards Realistic Data Generation for Real-World Super-Resolution

ICLR 2025poster

Existing image super-resolution (SR) techniques often fail to generalize effectively in complex real-world settings due to the significant divergence between training data and practical scenarios. To address this challenge, previous efforts have either manually simulated intricate physical-based deg…

Cited by 14SourcePDFScholar
2025

UHD-processer: Unified UHD Image Restoration with Progressive Frequency Learning and Degradation-aware Prompts

CVPR 2025poster

We introduce UHD-Processor, a unified and robust framework for all-in-one image restoration, which is particularly resource-efficient for Ultra-High-Definition (UHD) images. To address the limitations of traditional all-in-one methods that rely on complex restoration backbones, our strategy employs…

2025

ViewPoint: Panoramic Video Generation with Pretrained Diffusion Models

NeurIPS 2025poster

Panoramic video generation aims to synthesize 360-degree immersive videos, holding significant importance in the fields of VR, world models, and spatial intelligence. Existing works fail to synthesize high-quality panoramic videos due to the inherent modality gap between panoramic data and perspecti…

Cited by 0SourceScholar
2025

WeGen: A Unified Model for Interactive Multimodal Generation as We Chat

CVPR 2025poster

Existing multimodal generative models fall short as qualified design copilots, as they often struggle to generate imaginative outputs once instructions are less detailed or lack the ability to maintain consistency with the provided references. In this work, we introduce WeGen, a model that unifies m…

2024

CCM: Real-Time Controllable Visual Content Creation Using Text-to-Image Consistency Models

ICML 2024poster

Consistency Models (CMs) have showed a promise in creating high-quality images with few steps. However, the way to add new conditional controls to the pre-trained CMs has not been explored. In this paper, we explore the pivotal subject of leveraging the generative capacity and efficiency of consiste…

Cited by 4SourcePDFScholar
2024

Context-aware Difference Distilling for Multi-change Captioning

ACL 2024long

Multi-change captioning aims to describe complex and coupled changes within an image pair in natural language. Compared with single-change captioning, this task requires the model to have higher-level cognition ability to reason an arbitrary number of changes. In this paper, we propose a novel conte…

2024

DreamClean: Restoring Clean Image Using Deep Diffusion Prior

ICLR 2024poster

Image restoration poses a garners substantial interest due to the exponential surge in demands for recovering high-quality images from diverse mobile camera devices, adverse lighting conditions, suboptimal shooting environments, and frequent image compression for efficient transmission purposes. Yet…

Cited by 9SourcePDFScholar
2024

EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric Views

NeurIPS 2024poster

Understanding egocentric human-object interaction (HOI) is a fundamental aspect of human-centric perception, facilitating applications like AR/VR and embodied AI. For the egocentric HOI, in addition to perceiving semantics e.g., ''what'' interaction is occurring, capturing ''where'' the interaction…

Cited by 6SourcePDFScholar
2024

Event-Adapted Video Super-Resolution

ECCV 2024poster

"Introducing event cameras into video super-resolution (VSR) shows great promise. In practice, however, integrating event data as a new modality necessitates a laborious model architecture design. This not only consumes substantial time and effort but also disregards valuable insights from successfu…

Cited by 6SourcePDFScholar
2024

HomoFormer: Homogenized Transformer for Image Shadow Removal

CVPR 2024poster

The spatial non-uniformity and diverse patterns of shadow degradation conflict with the weight sharing manner of dominant models which may lead to an unsatisfactory compromise. To tackle with this issue we present a novel strategy from the view of shadow transformation in this paper: directly homoge…

2024

LEMON: Learning 3D Human-Object Interaction Relation from 2D Images

CVPR 2024poster

Learning 3D human-object interaction relation is pivotal to embodied AI and interaction modeling. Most existing methods approach the goal by learning to predict isolated interaction elements e.g. human contact object affordance and human-object spatial relation primarily from the perspective of eith…

2024

Learning Discriminative Noise Guidance for Image Forgery Detection and Localization

AAAI 2024technical

This study introduces a new method for detecting and localizing image forgery by focusing on manipulation traces within the noise domain. We posit that nearly invisible noise in RGB images carries tampering traces, useful for distinguishing and locating forgeries. However, the advancement of tamperi…

Cited by 16SourcePDFScholar
2024

LoTLIP: Improving Language-Image Pre-training for Long Text Understanding

NeurIPS 2024poster

In this work, we empirically confirm that the key reason causing such an issue is that the training images are usually paired with short captions, leaving certain tokens easily overshadowed by salient tokens. Towards this problem, our initial attempt is to relabel the data with long captions, howeve…

2024

Natural Language-centered Inference Network for Multi-modal Fake News Detection

IJCAI 2024poster

The proliferation of fake news with image and text in the internet has triggered widespread concern. Existing research has made important contributions in cross-modal information interaction and fusion, but fails to fundamentally address the modality gap among news image, text, and news-related exte…

Cited by 4SourcePDFScholar
2024

Neuromorphic Event Signal-Driven Network for Video De-raining

AAAI 2024technical

Convolutional neural networks-based video de-raining methods commonly rely on dense intensity frames captured by CMOS sensors. However, the limited temporal resolution of these sensors hinders the capture of dynamic rainfall information, limiting further improvement in de-raining performance. This s…

Cited by 10SourcePDFScholar
2024

Noise-assisted Prompt Learning for Image Forgery Detection and Localization

ECCV 2024poster

"We present CLIP-IFDL, a novel image forgery detection and localization (IFDL) model that harnesses the power of Contrastive Language Image Pre-Training (CLIP). However, directly incorporating CLIP in forgery detection poses challenges, given its lack of specific prompts and forgery consciousness. T…

Cited by 4SourcePDFScholar
2024

Prompt-Enhanced Multiple Instance Learning for Weakly Supervised Video Anomaly Detection

CVPR 2024poster

Weakly-supervised Video Anomaly Detection (wVAD) aims to detect frame-level anomalies using only video-level labels in training. Due to the limitation of coarse-grained labels Multi-Instance Learning (MIL) is prevailing in wVAD. However MIL suffers from insufficiency of binary supervision to model d…

2024

Revisiting Single Image Reflection Removal In the Wild

CVPR 2024poster

This research focuses on the issue of single-image reflection removal (SIRR) in real-world conditions examining it from two angles: the collection pipeline of real reflection pairs and the perception of real reflection locations. We devise an advanced reflection collection pipeline that is highly ad…

2023

Adaptive Frequency Filters As Efficient Global Token Mixers

ICCV 2023poster

Recent vision transformers, large-kernel CNNs and MLPs have attained remarkable successes in broad vision tasks thanks to their effective information fusion in the global scope. However, their efficient deployments, especially on mobile devices, still suffer from noteworthy challenges due to the hea…

Cited by 69PDFcodeScholar
2023

DreamWaltz: Make a Scene with Complex 3D Animatable Avatars

NeurIPS 2023poster

We present DreamWaltz, a novel framework for generating and animating complex 3D avatars given text guidance and parametric human body prior. While recent methods have shown encouraging results for text-to-3D generation of common objects, creating high-quality and animatable 3D avatars remains chall…

2023

Event-Guided Person Re-Identification via Sparse-Dense Complementary Learning

CVPR 2023poster

Video-based person re-identification (Re-ID) is a prominent computer vision topic due to its wide range of video surveillance applications. Most existing methods utilize spatial and temporal correlations in frame sequences to obtain discriminative person features. However, inevitable degradations, e…

Cited by 17SourcePDFScholar
2023

Exploring Tuning Characteristics of Ventral Stream’s Neurons for Few-Shot Image Classification

AAAI 2023technical

Human has the remarkable ability of learning novel objects by browsing extremely few examples, which may be attributed to the generic and robust feature extracted in the ventral stream of our brain for representing visual objects. In this sense, the tuning characteristics of ventral stream's neurons…

Cited by 12SourcePDFScholar
2023

Grounding 3D Object Affordance from 2D Interactions in Images

ICCV 2023poster

Grounding 3D object affordance seeks to locate objects' "action possibilities" regions in the 3D space, which serves as a link between perception and operation for embodied agents. Existing studies primarily focus on connecting visual affordances with geometry structures, e.g., relying on annotation…

Cited by 34PDFcodeScholar
2023

Learning Cross-Representation Affinity Consistency for Sparsely Supervised Biomedical Instance Segmentation

ICCV 2023poster

Sparse instance-level supervision has recently been explored to address insufficient annotation in biomedical instance segmentation, which is easier to annotate crowded instances and better preserves instance completeness for 3D volumetric datasets compared to common semi-supervision.In this paper,…

Cited by 8PDFcodeScholar
2023

Learning To Dub Movies via Hierarchical Prosody Models

CVPR 2023poster

Given a piece of text, a video clip and a reference audio, the movie dubbing (also known as visual voice clone, V2C) task aims to generate speeches that match the speaker's emotion presented in the video using the desired speaker voice as reference. V2C is more challenging than conventional text-to-…

2023

Neural Dependencies Emerging From Learning Massive Categories

CVPR 2023poster

This work presents two astonishing findings on neural networks learned for large-scale image classification. 1) Given a well-trained model, the logits predicted for some category can be directly obtained by linearly combining the predictions of a few other categories, which we call neural dependency…

2023

Random Shuffle Transformer for Image Restoration

ICML 2023poster

Non-local interactions play a vital role in boosting performance for image restoration. However, local window Transformer has been preferred due to its efficiency for processing high-resolution images. The superiority in efficiency comes at the cost of sacrificing the ability to model non-local inte…

2023

Regularized Mask Tuning: Uncovering Hidden Knowledge in Pre-Trained Vision-Language Models

ICCV 2023poster

Prompt tuning and adapter tuning have shown great potential in transferring pre-trained vision-language models (VLMs) to various downstream tasks. In this work, we design a new type of tuning method, termed as regularized mask tuning, which masks the network parameters through a learnable selection.…

Cited by 12PDFScholar
2023

Self-Organizing Pathway Expansion for Non-Exemplar Class-Incremental Learning

ICCV 2023poster

Non-exemplar class-incremental learning aims to recognize both the old and new classes without access to old class samples. The conflict between old and new class optimization is exacerbated since the shared neural pathways can only be differentiated by the incremental samples. To address this probl…

Cited by 12PDFScholar
2023

Self-supervised Cross-view Representation Reconstruction for Change Captioning

ICCV 2023poster

Change captioning aims to describe the difference between a pair of similar images. Its key challenge is how to learn a stable difference representation under pseudo changes caused by viewpoint change. In this paper, we address this by proposing a self-supervised cross-view representation reconstruc…

Cited by 36PDFcodeScholar
2023

Spatial-Aware Token for Weakly Supervised Object Localization

ICCV 2023poster

Weakly supervised object localization (WSOL) is a challenging task aiming to localize objects with only image-level supervision. Recent works apply visual transformer to WSOL and achieve significant success by exploiting the long-range feature dependency in self-attention mechanism. However, existin…

Cited by 13PDFcodeScholar
2023

Streaming Video Model

CVPR 2023poster

Video understanding tasks have traditionally been modeled by two separate architectures, specially tailored for two distinct tasks. Sequence-based video tasks, such as action recognition, use a video backbone to directly extract spatiotemporal features, while frame-based video tasks, such as multipl…

2023

Text-Driven Generative Domain Adaptation with Spectral Consistency Regularization

ICCV 2023poster

Combined with the generative prior of pre-trained models and the flexibility of text, text-driven generative domain adaptation can generate images from a wide range of target domains. However, current methods still suffer from overfitting and the mode collapse problem. In this paper, we analyze the…

Cited by 8PDFcodeScholar
2022

Automatic Relation-Aware Graph Network Proliferation

CVPR 2022oral

Graph neural architecture search has sparked much attention as Graph Neural Networks (GNNs) have shown powerful reasoning capability in many relational tasks. However, the currently used graph search space overemphasizes learning node features and neglects mining hierarchical relational information.…

Cited by 12PDFcodeScholar
2022

Debiased Batch Normalization via Gaussian Process for Generalizable Person Re-identification

AAAI 2022technical

Generalizable person re-identification aims to learn a model with only several labeled source domains that can perform well on unseen domains. Without access to the unseen domain, the feature statistics of the batch normalization (BN) layer learned from a limited number of source domains is doubtles…

Cited by 36SourcePDFScholar
2022

Degradation-Agnostic Correspondence From Resolution-Asymmetric Stereo

CVPR 2022poster

In this paper, we study the problem of stereo matching from a pair of images with different resolutions, e.g., those acquired with a tele-wide camera system. Due to the difficulty of obtaining ground-truth disparity labels in diverse real-world systems, we start from an unsupervised learning perspec…

Cited by 10PDFScholar
2022

EMScore: Evaluating Video Captioning via Coarse-Grained and Fine-Grained Embedding Matching

CVPR 2022poster

Current metrics for video captioning are mostly based on the text-level comparison between reference and candidate captions. However, they have some insuperable drawbacks, e.g., they cannot handle videos without references, and they may result in biased evaluation due to the one-to-many nature of vi…

Cited by 43PDFcodeScholar
2022

Efficient Model-Driven Network for Shadow Removal

AAAI 2022technical

Deep Convolutional Neural Networks (CNNs) based methods have achieved significant breakthroughs in the task of single image shadow removal. However, the performance of these methods remains limited for several reasons. First, the existing shadow illumination model ignores the spatially variant prope…

2022

Event-driven Video Deblurring via Spatio-Temporal Relation-Aware Network

IJCAI 2022poster

Video deblurring with event information has attracted considerable attention. To help deblur each frame, existing methods usually compress a specific event sequence into a feature tensor with the same size as the corresponding video. However, this strategy neither considers the pixel-level spatial b…

2022

Exploring Figure-Ground Assignment Mechanism in Perceptual Organization

NeurIPS 2022accept

Perceptual organization is a challenging visual task that aims to perceive and group the individual visual element so that it is easy to understand the meaning of the scene as a whole. Most recent methods building upon advanced Convolutional Neural Network (CNN) come from learning discriminative rep…

Cited by 20SourcePDFScholar
2022

Exploring Fourier Prior for Single Image Rain Removal

IJCAI 2022poster

Deep convolutional neural networks (CNNs) have become dominant in the task of single image rain removal. Most of current CNN methods, however, suffer from the problem of overfitting on one single synthetic dataset as they neglect the intrinsic prior of the physical properties of rain streaks. To add…

2022

Few Shot Generative Model Adaption via Relaxed Spatial Structural Alignment

CVPR 2022poster

Training a generative adversarial network (GAN) with limited data has been a challenging task. A feasible solution is to start with a GAN well-trained on a large scale source domain and adapt it to the target domain with a few samples, termed as few shot generative model adaption. However, existing…

Cited by 90PDFcodeScholar
2022

JPEG Artifacts Removal via Contrastive Representation Learning

ECCV 2022poster

"To meet the needs of practical applications, current deep learning-based methods focus on using a single model to handle JPEG images with different compression qualities, while few of them consider the auxiliary effects of the compression quality information. Recently, several methods estimate qual…

2022

Lifelong Unsupervised Domain Adaptive Person Re-Identification With Coordinated Anti-Forgetting and Adaptation

CVPR 2022poster

Unsupervised domain adaptive person re-identification (ReID) has been extensively investigated to mitigate the adverse effects of domain gaps. Those works assume the target domain data can be accessible all at once. However, for the real-world streaming data, this hinders the timely adaptation to ch…

Cited by 42PDFScholar
2022

Modality-Adaptive Mixup and Invariant Decomposition for RGB-Infrared Person Re-identification

AAAI 2022technical

RGB-infrared person re-identification is an emerging cross-modality re-identification task, which is very challenging due to significant modality discrepancy between RGB and infrared images. In this work, we propose a novel modality-adaptive mixup and invariant decomposition (MID) approach for RGB-i…

Cited by 105SourcePDFScholar
2022

Multi-Grained Spatio-Temporal Features Perceived Network for Event-Based Lip-Reading

CVPR 2022poster

Automatic lip-reading (ALR) aims to recognize words using visual information from the speaker's lip movements. In this work, we introduce a novel type of sensing device, event cameras, for the task of ALR. Event cameras have both technical and application advantages over conventional cameras for the…

Cited by 34PDFcodeScholar
2022

Principled Knowledge Extrapolation with GANs

ICML 2022spotlight

Human can extrapolate well, generalize daily knowledge into unseen scenarios, raise and answer counterfactual questions. To imitate this ability via generative models, previous works have extensively studied explicitly encoding Structural Causal Models (SCMs) into architectures of generator networks…

2022

ProgressiveMotionSeg: Mutually Reinforced Framework for Event-Based Motion Segmentation

AAAI 2022technical

Dynamic Vision Sensor (DVS) can asynchronously output the events reflecting apparent motion of objects with microsecond resolution, and shows great application potential in monitoring and other fields. However, the output event stream of existing DVS inevitably contains background activity noise (BA…

Cited by 10SourcePDFScholar
2022

Rank Diminishing in Deep Neural Networks

NeurIPS 2022accept

The rank of neural networks measures information flowing across layers. It is an instance of a key structural condition that applies across broad domains of machine learning. In particular, the assumption of low-rank feature representations led to algorithmic developments in many architectures. For…

2022

S2N: Suppression-Strengthen Network for Event-Based Recognition under Variant Illuminations

ECCV 2022poster

"The emerging event-based sensors have demonstrated out-standing potential in visual tasks thanks to their high speed and high dynamic range. However, the event degradation due to imaging under low illumination obscures the correlation between event signals and brings uncertainty into event represen…

2022

Self-Sustaining Representation Expansion for Non-Exemplar Class-Incremental Learning

CVPR 2022poster

Non-exemplar class-incremental learning is to recognize both the old and new classes when old class samples cannot be saved. It is a challenging task since representation optimization and feature retention can only be achieved under supervision from new classes. To address this problem, we propose a…

Cited by 207PDFScholar
2022

Temporal Complementarity-Guided Reinforcement Learning for Image-to-Video Person Re-Identification

CVPR 2022poster

Image-to-video person re-identification aims to retrieve the same pedestrian as the image-based query from a video-based gallery set. Existing methods treat it as a cross-modality retrieval task and learn the common latent embeddings from image and video modalities, which are both less effective and…

Cited by 17PDFScholar
2022

Unsupervised Coherent Video Cartoonization with Perceptual Motion Consistency

AAAI 2022technical

In recent years, creative content generations like style transfer and neural photo editing have attracted more and more attention. Among these, cartoonization of real-world scenes has promising applications in entertainment and industry. Different from image translations focusing on improving the st…

2022

Weakly Supervised High-Fidelity Clothing Model Generation

CVPR 2022poster

The development of online economics arouses the demand of generating images of models on product clothes, to display new clothes and promote sales. However, the expensive proprietary model images challenge the existing image virtual try-on methods in this scenario, as most of them need to be trained…

Cited by 8PDFcodeScholar
2021

Exploiting Sample Uncertainty for Domain Adaptive Person Re-Identification

AAAI 2021technical

Many unsupervised domain adaptive (UDA) person ReID approaches combine clustering-based pseudo-label prediction with feature fine-tuning. However, because of domain gap, the pseudo-labels are not always reliable and there are noisy/incorrect labels. This would mislead the feature representation lea…

Cited by 190SourcePDFScholar
2021

Group-aware Label Transfer for Domain Adaptive Person Re-identification

CVPR 2021poster

Unsupervised Domain Adaptive (UDA) person re-identification (ReID) aims at adapting the model trained on a labeled source-domain dataset to a target-domain dataset without any further annotations. Most successful UDA-ReID approaches combine clustering-based pseudo-label prediction with representatio…

Cited by 231PDFcodeScholar
2021

Learning Conditional Knowledge Distillation for Degraded-Reference Image Quality Assessment

ICCV 2021poster

An important scenario for image quality assessment (IQA) is to evaluate image restoration (IR) algorithms. The state-of-the-art approaches adopt a full-reference paradigm that compares restored images with their corresponding pristine-quality images. However, pristine-quality images are usually unav…

Cited by 62PDFcodeScholar
2021

Low-Rank Subspaces in GANs

NeurIPS 2021poster

The latent space of a Generative Adversarial Network (GAN) has been shown to encode rich semantics within some subspaces. To identify these subspaces, researchers typically analyze the statistical information from a collection of synthesized data, and the identified subspaces tend to control image a…

2021

Rain Streak Removal via Dual Graph Convolutional Network

AAAI 2021technical

Deep convolutional neural networks (CNNs) have become dominant in the single image de-raining area. However, most deep CNNs-based de-raining methods are designed by stacking vanilla convolutional layers, which can only be used to model local relations. Therefore, long-range contextual information is…

Cited by 146SourcePDFScholar
2021

Rethinking Graph Neural Architecture Search From Message-Passing

CVPR 2021poster

Graph neural networks (GNNs) emerged recently as a standard toolkit for learning from data on graphs. Current GNN designing works depend on immense human expertise to explore different message-passing mechanisms, and require manual enumeration to determine the proper message-passing depth. Inspired…

Cited by 66PDFcodeScholar
2021

Self-Promoted Prototype Refinement for Few-Shot Class-Incremental Learning

CVPR 2021poster

Few-shot class-incremental learning is to recognize the new classes given few samples and not forget the old classes. It is a challenging task since representation optimization and prototype reorganization can only be achieved under little supervision. To address this problem, we propose a novel inc…

Cited by 203PDFcodeScholar
2021

Self-Supervised Visual Representations Learning by Contrastive Mask Prediction

ICCV 2021poster

Advanced self-supervised visual representation learning methods rely on the instance discrimination (ID) pretext task. We point out that the ID task has an implicit semantic consistency (SC) assumption, which may not hold in unconstrained datasets. In this paper, we propose a novel contrastive mask…

Cited by 49PDFcodeScholar
2021

Spatial-Temporal Correlation and Topology Learning for Person Re-Identification in Videos

CVPR 2021poster

Video-based person re-identification aims to match pedestrians from video sequences across non-overlapping camera views. The key factor for video person re-identification is to effectively exploit both spatial and temporal clues from video sequences. In this work, we propose a novel Spatial-Temporal…

Cited by 81PDFScholar
2021

Structured Multi-Level Interaction Network for Video Moment Localization via Language Query

CVPR 2021poster

We address the problem of localizing a specific moment described by a natural language query. Existing works interact the query with either video frame or moment proposal, and neglect the inherent structure of moment construction for both cross-modal understanding and video content comprehension, wh…

Cited by 102PDFScholar
2021

Training Spiking Neural Networks with Accumulated Spiking Flow

AAAI 2021technical

The fast development of neuromorphic hardwares promotes Spiking Neural Networks (SNNs) to a thrilling research avenue. Current SNNs, though much efficient, are less effective compared with leading Artificial Neural Networks (ANNs) especially in supervised learning tasks. Recent efforts further demon…

2021

Uncertainty Principles of Encoding GANs

ICML 2021spotlight

The compelling synthesis results of Generative Adversarial Networks (GANs) demonstrate rich semantic knowledge in their latent codes. To obtain this knowledge for downstream applications, encoding GANs has been proposed to learn encoders, such that real world data can be encoded to latent codes, whi…

Cited by 8SourcePDFScholar
2020

Co-Saliency Spatio-Temporal Interaction Network for Person Re-Identification in Videos

IJCAI 2020poster

Person re-identification aims at identifying a certain pedestrian across non-overlapping camera networks. Video-based person re-identification approaches have gained significant attention recently, expanding image-based approaches by learning features from multiple frames. In this work, we propose a…

Cited by 0SourcePDFScholar
2020

ContourNet: Taking a Further Step Toward Accurate Arbitrary-Shaped Scene Text Detection

CVPR 2020poster

Scene text detection has witnessed rapid development in recent years. However, there still exists two main challenges: 1) many methods suffer from false positives in their text representations; 2) the large scale variance of scene texts makes it hard for network to learn samples. In this paper, we p…

Cited by 273PDFcodeScholar
2020

Domain-Aware Visual Bias Eliminating for Generalized Zero-Shot Learning

CVPR 2020poster

Generalized zero-shot learning aims to recognize images from seen and unseen domains. Recent methods focus on learning a unified semantic-aligned visual representation to transfer knowledge between two domains, while ignoring the effect of semantic-free visual representation in alleviating the biase…

Cited by 201PDFcodeScholar
2020

Hierarchical Granularity Transfer Learning

NeurIPS 2020poster

In the real world, object categories usually have a hierarchical granularity tree. Nowadays, most researchers focus on recognizing categories in a specific granularity, \emph{e.g.,} basic-level or sub(ordinate)-level. Compared with basic-level categories, the sub-level categories provide more valuab…

Cited by 5SourcePDFScholar
2020

JPEG Artifacts Removal via Compression Quality Ranker-Guided Networks

IJCAI 2020poster

Existing deep learning-based image de-blocking methods use only pixel-level loss functions to guide network training. The JPEG compression factor, which reflects the degradation degree, has not been fully utilized. However, due to the non-differentiability, the compression factor cannot be directly…

Cited by 0SourcePDFScholar
2020

Learning Semantic-aware Normalization for Generative Adversarial Networks

NeurIPS 2020spotlight

The recent advances in image generation have been achieved by style-based image generators. Such approaches learn to disentangle latent factors in different image scales and encode latent factors as “style” to control image synthesis. However, existing approaches cannot further disentangle fine-grai…

2020

Learning to Discretely Compose Reasoning Module Networks for Video Captioning

IJCAI 2020poster

Generating natural language descriptions for videos, i.e., video captioning, essentially requires step-by-step reasoning along the generation process. For example, to generate the sentence “a man is shooting a basketball”, we need to first locate and describe the subject “man”, next reason out the m…

2020

Multi-Scale Group Transformer for Long Sequence Modeling in Speech Separation

IJCAI 2020poster

In this paper, we introduce Transformer to the time-domain methods for single-channel speech separation. Transformer has the potential to boost speech separation performance because of its strong sequence modeling capability. However, its computational complexity, which grows quadratically with the…

Cited by 0SourcePDFScholar
2020

Multi-Scale Spatial-Temporal Integration Convolutional Tube for Human Action Recognition

IJCAI 2020poster

Applying multi-scale representations leads to consistent performance improvements on a wide range of image recognition tasks. However, with the addition of the temporal dimension in video domain, directly obtaining layer-wise multi-scale spatial-temporal features will add a lot extra computational c…

Cited by 0SourcePDFScholar
2020

Object Relational Graph With Teacher-Recommended Learning for Video Captioning

CVPR 2020poster

Taking full advantage of the information from both vision and language is critical for the video captioning task. Existing models lack adequate visual representation due to the neglect of interaction between object, and sufficient training for content-related words due to long-tailed problems. In th…

Cited by 387PDFScholar
2020

Parsing-Based View-Aware Embedding Network for Vehicle Re-Identification

CVPR 2020poster

Vehicle Re-Identification is to find images of the same vehicle from various views in the cross-camera scenario. The main challenges of this task are the large intra-instance distance caused by different views and the subtle inter-instance discrepancy caused by similar vehicles. In this paper, we pr…

Cited by 255PDFcodeScholar
2020

Real-World Person Re-Identification via Degradation Invariance Learning

CVPR 2020poster

Person re-identification (Re-ID) in real-world scenarios usually suffers from various degradation factors, e.g., low-resolution, weak illumination, blurring and adverse weather. On the one hand, these degradations lead to severe discriminative information loss, which significantly obstructs identity…

Cited by 89PDFScholar
2020

Self-Supervised Domain-Aware Generative Network for Generalized Zero-Shot Learning

CVPR 2020poster

Generalized Zero-Shot Learning (GZSL) aims at recognizing both seen and unseen classes by constructing correspondence between visual and semantic embedding. However, existing methods have severely suffered from the strong bias problem, where unseen instances in target domain tend to be recognized as…

Cited by 78PDFScholar
2020

State-Relabeling Adversarial Active Learning

CVPR 2020oral

Active learning is to design label-efficient algorithms by sampling the most representative samples to be labeled by an oracle. In this paper, we propose a state relabeling adversarial active learning model (SRAAL), that leverages both the annotation and the labeled/unlabeled state information for d…

Cited by 160PDFScholar
2019

Adaptive Reconstruction Network for Weakly Supervised Referring Expression Grounding

ICCV 2019poster

Weakly supervised referring expression grounding aims at localizing the referential object in an image according to the linguistic query, where the mapping between the referential object and query is unknown in the training stage. To address this problem, we propose a novel end-to-end adaptive recon…

Cited by 110PDFcodeScholar
2019

Adaptive Transfer Network for Cross-Domain Person Re-Identification

CVPR 2019poster

Recent deep learning based person re-identification approaches have steadily improved the performance for benchmarks, however they often fail to generalize well from one domain to another. In this work, we propose a novel adaptive transfer network (ATNet) for effective cross-domain person re-identif…

Cited by 350PDFScholar
2019

JPEG Artifacts Reduction via Deep Convolutional Sparse Coding

ICCV 2019poster

To effectively reduce JPEG compression artifacts, we propose a deep convolutional sparse coding (DCSC) network architecture. We design our DCSC in the framework of classic learned iterative shrinkage-threshold algorithm. To focus on recognizing and separating artifacts only, we sparsely code the fea…

Cited by 139PDFScholar
2019

Learning Deep Bilinear Transformation for Fine-grained Image Representation

NeurIPS 2019poster

Bilinear feature transformation has shown the state-of-the-art performance in learning fine-grained image representations. However, the computational cost to learn pairwise interactions between deep feature channels is prohibitively expensive, which restricts this powerful transformation to be used…

2019

Looking for the Devil in the Details: Learning Trilinear Attention Sampling Network for Fine-Grained Image Recognition

CVPR 2019poster

Learning subtle yet discriminative features (e.g., beak and eyes for a bird) plays a significant role in fine-grained image recognition. Existing attention-based approaches localize and amplify significant parts to learn fine-grained details, which often suffer from a limited number of parts and hea…

Cited by 543PDFcodeScholar
2019

Making History Matter: History-Advantage Sequence Training for Visual Dialog

ICCV 2019poster

We study the multi-round response generation in visual dialog, where a response is generated according to a visually grounded conversational history. Given a triplet: an image, Q&A history, and current question, all the prevailing methods follow a codec (i.e., encoder-decoder) fashion in a supervise…

Cited by 81PDFcodeScholar
2018

MiCT: Mixed 3D/2D Convolutional Tube for Human Action Recognition

CVPR 2018poster

Human actions in videos are three-dimensional (3D) signals. Recent attempts use 3D convolutional neural networks (CNNs) to explore spatio-temporal information for human action recognition. Though promising, 3D CNNs have not achieved high performanceon on this task with respect to their well-establis…

Cited by 295SourcePDFScholar
2016

Comparative Deep Learning of Hybrid Representations for Image Recommendations

CVPR 2016poster

In many image-related tasks, learning expressive and discriminative representations of images is essential, and deep learning has been studied for automating the learning of such representations. Some user-centric tasks, such as image recommendations, call for effective representations of not only i…

Cited by 162PDFScholar