← Search

Xiaopeng Zhang

68 accepted papers

2026

CogniEdit: Dense Gradient Flow Optimization for Fine-Grained Image Editing

CVPR 2026

Instruction-based image editing with diffusion models has achieved impressive results, yet existing methods struggle with fine-grained instructions specifying precise attributes such as colors, positions, and quantities. While recent approaches employ Group Relative Policy Optimization (GRPO) for al

Cited by 0SourcecodeScholar
2026

DehazeGS: Seeing Through Fog with 3D Gaussian Splatting

AAAI 2026technical

Current novel view synthesis methods are typically designed for high-quality and clean input images. However, in foggy scenes, scattering and attenuation can significantly degrade the quality of rendering. Although NeRF-based dehazing approaches have been developed, their reliance on deep fully conn

Cited by 0SourcePDFScholar
2026

Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal Interaction

AAAI 2026technical

Dense video captioning jointly localizes and captions salient events in untrimmed videos. Recent methods primarily focus on leveraging additional prior knowledge and advanced multi-task architectures to achieve competitive performance. However, these pipelines rely on implicit modeling that uses fra

Cited by 0SourcePDFScholar
2026

GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs

ICLR 2026poster

Vision encoders are indispensable for allowing impressive performance of Multimodal Large Language Models (MLLMs) in vision–language tasks such as visual question answering and reasoning. However, existing vision encoders focus on global image representations but overlook fine-grained regional analy…

Cited by 0SourceScholar
2026

LaplacianFormer:Rethinking Linear Attention with Laplacian Kernel

ICLR 2026poster

The quadratic complexity of softmax attention presents a major obstacle for scaling Transformers to high-resolution vision tasks. Existing linear attention variants often replace the softmax with Gaussian kernels to reduce complexity, but such approximations lack theoretical grounding and tend to ov…

Cited by 0SourceScholar
2026

O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing

AAAI 2026technical

Diffusion models have recently advanced video editing, yet controllable editing remains challenging due to the need for precise manipulation of diverse object properties. Current methods require different control signal for diverse editing tasks, which complicates model design and demands significan

Cited by 0SourcePDFScholar
2026

ParaUni: Enhance Generation in Unified Multimodal Model with Reinforcement-driven Hierarchical Parallel Information Interaction

CVPR 2026

Unified multimodal models significantly improve visual generation by combining vision-language models (VLMs) with diffusion models. However, existing methods struggle to fully balance sufficient interaction and flexible implementation due to vast representation difference. Considering abundant and h

Cited by 0SourcecodeScholar
2025

AccidentX: A Large-Scale Multimodal BEV Dataset for Traffic Accident Analysis and Prevention

IROS 2025

With the rapid development and widespread application of autonomous driving technology, the accurate analysis and prevention of traffic accidents have become critical challenges. However, current traffic accident datasets are often constrained by limited scale and diversity, impeding progress in thi

Cited by 0SourceScholar
2025

CLIP-Adapted Region-to-Text Learning for Generative Open-Vocabulary Semantic Segmentation

ICCV 2025poster

In recent years, Open-Vocabulary Semantic Segmentation (OVSS) has been largely advanced. However, existing methods mostly rely on a pre-trained vision-language model (e.g., CLIP) and require a predefined set of classes to guide the semantic segmentation process during the inference. This not only na…

Cited by 0SourcePDFScholar
2025

Diffusion-Driven Progressive Target Manipulation for Source-Free Domain Adaptation

NeurIPS 2025poster

Source-free domain adaptation (SFDA) is a challenging task that tackles domain shifts using only a pre-trained source model and unlabeled target data. Existing SFDA methods are restricted by the fundamental limitation of source-target domain discrepancy. Non-generation SFDA methods suffer from unrel…

Cited by 0SourceScholar
2025

DiffusionIMU: Diffusion-Based Inertial Navigation with Iterative Motion Refinement

IJCAI 2025

Inertial navigation enables self-contained localization using only Inertial Measurement Units (IMUs), making it widely applicable in various domains such as navigation, augmented reality, and robotics. However, existing methods suffer from drift accumulation due to the sensor noise and difficulty ca

Cited by 0SourcePDFScholar
2025

IM-Zero: Instance-level Motion Controllable Video Generation in a Zero-shot Manner

CVPR 2025poster

Controllability of video generation has been recently concerned in addition to the quality of generated videos. The main challenge to controllable video generation is to synthesize videos based on user-specified instance spatial locations and movement trajectories. However, existing methods suffer f…

Cited by 0SourcePDFScholar
2025

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models

ICCV 2025poster

Vision encoders serve as the cornerstone of multimodal understanding. Single-encoder architectures like CLIP exhibit inherent constraints in generalizing across diverse multimodal tasks, while recent multi-encoder fusion methods introduce prohibitive computational overhead to achieve superior perfor…

2025

OmniDraft: A cross-vocabulary, online adaptive drafter for on-device speculative decoding

NeurIPS 2025poster

Speculative decoding generally dictates having a small, efficient draft model that is either pretrained or distilled offline to a particular target model series, for instance, Llama or Qwen models. However, within online deployment settings, there are two major challenges: 1) usage of a target model…

Cited by 0SourceScholar
2025

PanoDiT: Panoramic Videos Generation with Diffusion Transformer

AAAI 2025technical

As immersive experiences become increasingly popular, panoramic video has garnered significant attention in both research and applications. The high cost associated with capturing panoramic video underscores the need for efficient prompt-based generation methods. Although recent text-to-video (T2V)…

Cited by 0SourcePDFScholar
2025

SAM-CP: Marrying SAM with Composable Prompts for Versatile Segmentation

ICLR 2025poster

The Segment Anything model (SAM) has shown a generalized ability to group image pixels into patches, but applying it to semantic-aware segmentation still faces major challenges. This paper presents SAM-CP, a simple approach that establishes two types of composable prompts beyond SAM and composes the…

2025

Segment Any 3D Gaussians

AAAI 2025technical

This paper presents SAGA (Segment Any 3D GAussians), a highly efficient 3D promptable segmentation method based on 3D Gaussian Splatting (3D-GS). Given 2D visual prompts as input, SAGA can segment the corresponding 3D target represented by 3D Gaussians within 4 ms. This is achieved by attaching a sc…

2025

Tackling View-Dependent Semantics in 3D Language Gaussian Splatting

ICML 2025poster

Recent advancements in 3D Gaussian Splatting (3D-GS) enable high-quality 3D scene reconstruction from RGB images. Many studies extend this paradigm for language-driven open-vocabulary scene understanding. However, most of them simply project 2D semantic features onto 3D Gaussians and overlook a fund…

2024

4D Gaussian Splatting for Real-Time Dynamic Scene Rendering

CVPR 2024poster

Representing and rendering dynamic scenes has been an important but challenging task. Especially to accurately model complex motions high efficiency is usually hard to guarantee. To achieve real-time dynamic scene rendering while also enjoying high training and storage efficiency we propose 4D Gauss…

2024

AlignZeg: Mitigating Objective Misalignment for Zero-shot Semantic Segmentation

ECCV 2024poster

"A serious issue that harms the performance of zero-shot visual recognition is named objective misalignment, i.e., the learning objective prioritizes improving the recognition accuracy of seen classes rather than unseen classes, while the latter is the true target to pursue. This issue becomes more…

Cited by 4SourcePDFScholar
2024

BarLeRIa: An Efficient Tuning Framework for Referring Image Segmentation

ICLR 2024spotlight

Pre-training followed by full fine-tuning has gradually been substituted by Parameter-Efficient Tuning (PET) in the field of computer vision. PET has gained popularity, especially in the context of large-scale models, due to its ability to reduce transfer learning costs and conserve hardware resourc…

2024

Bootstrap AutoEncoders With Contrastive Paradigm for Self-supervised Gaze Estimation

ICML 2024poster

Existing self-supervised methods for gaze estimation using the dominant streams of contrastive and generative approaches are restricted to eye images and could fail in general full-face settings. In this paper, we reveal that contrastive methods are ineffective in data augmentation for self-supervis…

Cited by 0SourcePDFScholar
2024

ControlVideo: Training-free Controllable Text-to-video Generation

ICLR 2024poster

Text-driven diffusion models have unlocked unprecedented abilities in image generation, whereas their video counterpart lags behind due to the excessive training cost. To avert the training burden, we propose a training-free ControlVideo to produce high-quality videos based on the provided text prom…

2024

DefFusion: Deformable Multimodal Representation Fusion for 3D Semantic Segmentation

ICRA 2024poster

The complementarity between camera and LiDAR data makes fusion methods a promising approach to improve 3D semantic segmentation performance. Recent transformer-based methods have also demonstrated superiority in segmentation. However, multimodal solutions incorporating transformers are underexplored…

Cited by 8SourceScholar
2024

GaussianDreamer: Fast Generation from Text to 3D Gaussians by Bridging 2D and 3D Diffusion Models

CVPR 2024poster

In recent times the generation of 3D assets from text prompts has shown impressive results. Both 2D and 3D diffusion models can help generate decent 3D objects based on prompts. 3D diffusion models have good 3D consistency but their quality and generalization are limited as trainable 3D data is expe…

2024

GaussianEditor: Editing 3D Gaussians Delicately with Text Instructions

CVPR 2024poster

Recently impressive results have been achieved in 3D scene editing with text instructions based on a 2D diffusion model. However current diffusion models primarily generate images by predicting noise in the latent space and the editing is usually applied to the whole image which makes it challenging…

2024

Hybrid Distillation: Connecting Masked Autoencoders with Contrastive Learners

ICLR 2024poster

As two prominent strategies for representation learning, Contrastive Learning (CL) and Masked Image Modeling (MIM) have witnessed significant progress. Previous studies have demonstrated the advantages of each approach in specific scenarios. CL, resembling supervised pre-training, excels at capturin…

Cited by 3SourcePDFScholar
2024

QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models

ICLR 2024poster

Recently years have witnessed a rapid development of large language models (LLMs). Despite the strong ability in many language-understanding tasks, the heavy computational burden largely restricts the application of LLMs especially when one needs to deploy them onto edge devices. In this paper, we p…

2024

SVDTree: Semantic Voxel Diffusion for Single Image Tree Reconstruction

CVPR 2024poster

Efficiently representing and reconstructing the 3D geometry of biological trees remains a challenging problem in computer vision and graphics. We propose a novel approach for generating realistic tree models from single-view photographs. We cast the 3D information inference problem to a semantic vox…

2024

Spectral Prompt Tuning: Unveiling Unseen Classes for Zero-Shot Semantic Segmentation

AAAI 2024technical

Recently, CLIP has found practical utility in the domain of pixel-level zero-shot segmentation tasks. The present landscape features two-stage methodologies beset by issues such as intricate pipelines and elevated computational costs. While current one-stage approaches alleviate these concerns and…

2024

Stepping Forward on the Last Mile

NeurIPS 2024poster

Continuously adapting pre-trained models to local data on resource constrained edge devices is the \emph{last mile} for model deployment. However, as models increase in size and depth, backpropagation requires a large amount of memory, which becomes prohibitive for edge devices. In addition, most ex…

Cited by 1SourcePDFScholar
2024

UnionFormer: Unified-Learning Transformer with Multi-View Representation for Image Manipulation Detection and Localization

CVPR 2024poster

We present UnionFormer a novel framework that integrates tampering clues across three views by unified learning for image manipulation detection and localization. Specifically we construct a BSFI-Net to extract tampering features from RGB and noise views achieving enhanced responsiveness to boundary…

Cited by 10SourcePDFScholar
2023

Adapting Shortcut With Normalizing Flow: An Efficient Tuning Framework for Visual Recognition

CVPR 2023poster

Pretraining followed by fine-tuning has proven to be effective in visual recognition tasks. However, fine-tuning all parameters can be computationally expensive, particularly for large-scale models. To mitigate the computational and storage demands, recent research has explored Parameter-Efficient F…

2023

AiluRus: A Scalable ViT Framework for Dense Prediction

NeurIPS 2023poster

Vision transformers (ViTs) have emerged as a prevalent architecture for vision tasks owing to their impressive performance. However, their complexity dramatically increases when handling long token sequences, particularly for dense prediction tasks that require high-resolution input. Notably, dense…

2023

Integrally Pre-Trained Transformer Pyramid Networks

CVPR 2023poster

In this paper, we present an integral pre-training framework based on masked image modeling (MIM). We advocate for pre-training the backbone and neck jointly so that the transfer gap between MIM and downstream recognition tasks is minimal. We make two technical contributions. First, we unify the rec…

2023

Progressively Compressed Auto-Encoder for Self-supervised Representation Learning

ICLR 2023poster

As a typical self-supervised learning strategy, Masked Image Modeling (MIM) is driven by recovering all masked patches from visible ones. However, patches from the same image are highly correlated and it is redundant to reconstruct all the masked patches. We find that this redundancy is neglected by…

2023

Prune Spatio-temporal Tokens by Semantic-aware Temporal Accumulation

ICCV 2023poster

Transformers have become the primary backbone of the computer vision community due to their impressive performance. However, the unfriendly computation cost impedes their potential in the video recognition domain. To optimize the speed-accuracy trade-off, we propose Semantic-aware Temporal Accumulat…

Cited by 24PDFcodeScholar
2023

Self Correspondence Distillation for End-to-End Weakly-Supervised Semantic Segmentation

AAAI 2023technical

Efficiently training accurate deep models for weakly supervised semantic segmentation (WSSS) with image-level labels is challenging and important. Recently, end-to-end WSSS methods have become the focus of research due to their high training efficiency. However, current methods suffer from insuffici…

2023

The KFIoU Loss for Rotated Object Detection

ICLR 2023poster

Differing from the well-developed horizontal object detection area whereby the computing-friendly IoU based loss is readily adopted and well fits with the detection metrics, rotation detectors often involve a more complicated loss based on SkewIoU which is unfriendly to gradient-based training. In t…

Cited by 239SourcePDFScholar
2023

Treating Pseudo-labels Generation as Image Matting for Weakly Supervised Semantic Segmentation

ICCV 2023poster

Generating accurate pseudo-labels under the supervision of image categories is a crucial step in Weakly Supervised Semantic Segmentation (WSSS). In this work, we propose a Mat-Label pipeline that provides a fresh way to treat WSSS pseudo-labels generation as an image matting task. By taking a trimap…

Cited by 29PDFcodeScholar
2022

A Transformer-Based Decoder for Semantic Segmentation with Multi-level Context Mining

ECCV 2022poster

"Transformers have recently shown superior performance than CNN on semantic segmentation. However, previous works mostly focus on the deliberate design of the encoder, while seldom considering the decoder part. In this paper, we find that a light weighted decoder counts for segmentation, and propose…

2022

Active Pointly-Supervised Instance Segmentation

ECCV 2022poster

"The requirement of expensive annotations is a major burden for training a well-performed instance segmentation model. In this paper, we present an economic active learning setting, named active pointly-supervised instance segmentation (APIS), which starts with box-level annotations and iteratively…

2022

Bag of Instances Aggregation Boosts Self-supervised Distillation

ICLR 2022poster

Recent advances in self-supervised learning have experienced remarkable progress, especially for contrastive learning based methods, which regard each image as well as its augmentations as an individual class and try to distinguish them from all other images. However, due to the large quantity of ex…

2022

Can Semantic Labels Assist Self-Supervised Visual Representation Learning?

AAAI 2022technical

Recently, contrastive learning has largely advanced the progress of unsupervised visual representation learning. Pre-trained on ImageNet, some self-supervised algorithms reported higher transfer learning performance compared to fully-supervised methods, seeming to deliver the message that human labe…

Cited by 33SourcePDFScholar
2022

DOMAINDESC: Learning Local Descriptors With Domain Adaptation

ICASSP 2022accepted

Robust and efficient local descriptor is crucial in a wide range of applications. In this paper, we propose a novel descriptor DomainDesc which is invariant as much as possible by learning local Descriptor with Domain adaptation. We design the feature-level domain adaptation loss to improve robustne…

Cited by 0SourceScholar
2022

GeoROS: Georeferenced Real-time Orthophoto Stitching with Unmanned Aerial Vehicle

IROS 2022poster

Simultaneous orthophoto stitching during the flight of Unmanned Aerial Vehicles (UAV) can greatly promote the practicability and instantaneity of diverse applications such as emergency disaster rescue, digital agriculture, and cadastral survey, which is of remarkable interest in aerial photogrammetr…

Cited by 3SourceScholar
2022

MSG-Transformer: Exchanging Local Spatial Information by Manipulating Messenger Tokens

CVPR 2022poster

Transformers have offered a new methodology of designing neural networks for visual recognition. Compared to convolutional networks, Transformers enjoy the ability of referring to global features at each stage, yet the attention module brings higher computational overhead that obstructs the applicat…

Cited by 96PDFcodeScholar
2022

MTLDesc: Looking Wider to Describe Better

AAAI 2022technical

Limited by the locality of convolutional neural networks, most existing local features description methods only learn local descriptors with local information and lack awareness of global and surrounding spatial context. In this work, we focus on making local descriptors ``look wider to describe bet…

2022

One-Bit Active Query With Contrastive Pairs

CVPR 2022poster

How to achieve better results with fewer labeling costs remains a challenging task. In this paper, we present a new active learning framework, which for the first time incorporates contrastive learning into recently proposed one-bit supervision. Here one-bit supervision denotes a simple Yes or No qu…

Cited by 9PDFcodeScholar
2022

SdAE: Self-Distillated Masked Autoencoder

ECCV 2022poster

"With the development of generative-based self-supervised learning (SSL) approaches like BeiT and MAE, how to learn good representations by masking random patches of the input image and reconstructing the missing information has grown in concern. However, BeiT and PeCo need a “pre-pretraining” stage…

2022

TAPE: Task-Agnostic Prior Embedding for Image Restoration

ECCV 2022poster

"Learning a generalized prior for natural image restoration is an important yet challenging task. Early methods mostly involved handcrafted priors including normalized sparsity, â„“0 gradients, dark channel priors, etc.. Recently, deep neural networks have been used to learn various image priors but…

Cited by 65SourcePDFScholar
2021

Don’t Miss the Potential Customers! Retrieving Similar Ads to Improve User Targeting

EMNLP 2021finding

User targeting is an essential task in the modern advertising industry: given a package of ads for a particular category of products (e.g., green tea), identify the online users to whom the ad package should be targeted. A (ad package specific) user targeting model is typically trained using histori…

Cited by 1SourcePDFScholar
2021

Rethinking Rotated Object Detection with Gaussian Wasserstein Distance Loss

ICML 2021spotlight

Boundary discontinuity and its inconsistency to the final detection metric have been the bottleneck for rotating detection regression loss design. In this paper, we propose a novel regression loss based on Gaussian Wasserstein distance as a fundamental approach to solve the problem. Specifically, th…

2020

Central Similarity Quantization for Efficient Image and Video Retrieval

CVPR 2020poster

Existing data-dependent hashing methods usually learn hash functions from pairwise or triplet data relationships, which only capture the data similarity locally, and often suffer from low learning efficiency and low collision rate. In this work, we propose a new global similarity metric, termed as c…

Cited by 400PDFcodeScholar
2020

Circumventing Outliers of AutoAugment with Knowledge Distillation

ECCV 2020poster

AutoAugment has been a powerful algorithm that improves the accuracy of many vision tasks, yet it is sensitive to the operator space as well as hyper-parameters, and an improper setting may degenerate network optimization. This paper delves deep into the working mechanism, and reveals that AutoAugme…

Cited by 76SourcePDFScholar
2020

PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture Search

ICLR 2020spotlight

Differentiable architecture search (DARTS) provided a fast solution in finding effective network architectures, but suffered from large memory and computing overheads in jointly training a super-net and searching for an optimal architecture. In this paper, we present a novel approach, namely Partia…

Cited by 920SourcecodeScholar
2019

A Robust Local Spectral Descriptor for Matching Non-Rigid Shapes With Incompatible Shape Structures

CVPR 2019poster

Constructing a robust and discriminative local descriptor for 3D shape is a key component of many computer vision applications. Although existing learning-based approaches can achieve good performance in some specific benchmarks, they usually fail to learn enough information from shapes with differe…

Cited by 25PDFScholar
2018

Learning 3D Keypoint Descriptors for Non-Rigid Shape Matching

ECCV 2018poster

In this paper, we present a novel deep learning framework that derives discriminative local descriptors for 3D surface shapes. In contrast to previous convolutional neural networks (CNNs) that rely on rendering multi-view images or extracting intrinsic shape properties, we parameterize the multi-sca…

Cited by 49SourcePDFScholar
2018

ML-LocNet: Improving Object Localization with Multi-view Learning Network

ECCV 2018poster

This paper addresses Weakly Supervised Object Localization (WSOL) with only image-level supervision. We propose a Multi-view Learning Localization Network (ML-LocNet) by incorporating multi-view learning into a two-phase WSOL model. The multi-view learning would benefit localization due to the compl…

Cited by 29SourcePDFScholar
2017

Hardware-Efficient Guided Image Filtering for Multi-Label Problem

CVPR 2017poster

The Guided Filter (GF) is well-known for its linear complexity. However, when filtering an image with an n-channel guidance, GF needs to invert an n xn matrix for each pixel. To the best of our knowledge existing matrix inverse algorithms are inefficient on current hardwares. This shortcoming limits…

Cited by 11PDFScholar
2016

Picking Deep Filter Responses for Fine-Grained Image Recognition

CVPR 2016poster

Recognizing fine-grained sub-categories such as birds and dogs is extremely challenging due to the highly localized and subtle differences in some specific parts. Most previous works rely on object/part level annotations to build part-based representation, which is demanding in practical application…

Cited by 401PDFScholar
2015

Segment Graph Based Image Filtering: Fast Structure-Preserving Smoothing

ICCV 2015poster

In this paper, we design a new edge-aware structure, named segment graph, to represent the image and we further develop a novel double weighted average image filter (SGF) based on the segment graph. In our SGF, we use the tree distance on the segment graph to define the internal weight function of t…

Cited by 70PDFScholar