← Search

Pingping Zhang

37 accepted papers

2026

CADTrack: Learning Contextual Aggregation with Deformable Alignment for Robust RGBT Tracking

AAAI 2026technical

RGB-Thermal (RGBT) tracking aims to exploit visible and thermal infrared modalities for robust all-weather object tracking. However, existing RGBT trackers struggle to resolve modality discrepancies, which poses great challenges for robust feature representation. This limitation hinders effective cr

Cited by 0SourcePDFScholar
2026

From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training

CVPR 2026

Reinforcement Learning with Verifiable Rewards (RLVR) for Multimodal Large Language Models (MLLMs) is highly dependent on high-quality labeled data, which is often scarce and prone to substantial annotation noise in real-world scenarios. Existing unsupervised RLVR methods, including pure entropy min

Cited by 0SourcecodeScholar
2026

RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented Generation

CVPR 2026

RGB-Thermal (RGBT) tracking aims to achieve robust object localization across diverse environmental conditions by fusing visible and thermal infrared modalities. However, existing RGBT trackers rely solely on initial-frame visual information for target modeling, failing to adapt to appearance variat

Cited by 0SourcecodeScholar
2026

Reinforcing Video Object Segmentation to Think before it Segments

CVPR 2026

Video reasoning segmentation (VRS) endeavors to delineate referred objects in videos guided by implicit instructions that encapsulate human intent and temporal logic. Previous approaches leverage large vision language models (LVLMs) to encode object semantics into \SEG tokens for mask prediction. Ho

Cited by 0SourceScholar
2026

Signal: Selective Interaction and Global-local Alignment for Multi-Modal Object Re-Identification

AAAI 2026technical

Multi-modal object Re-IDentification (ReID) is devoted to retrieving specific objects through the exploitation of complementary multi-modal image information. Existing methods mainly concentrate on the fusion of multi-modal features, yet neglecting the background interference. Besides, current multi

Cited by 0SourcePDFScholar
2026

VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models

AAAI 2026technical

Multimodal Large Language Models (MLLM) have enabled a wide range of advanced vision-language applications, including fine-grained object recognition and contextual understanding. When querying specific regions or objects in an image, human users naturally use "Visual Prompts" (VP) like bounding box

Cited by 0SourcePDFScholar
2026

X-ReID: Multi-granularity Information Interaction for Video-Based Visible-Infrared Person Re-Identification

AAAI 2026technical

Large-scale vision-language models (e.g., CLIP) have recently achieved remarkable performance in retrieval tasks, yet their potential for Video-based Visible-Infrared Person Re-Identification (VVI-ReID) remains largely unexplored. The primary challenges are narrowing the modality gap and leveraging

Cited by 0SourcePDFScholar
2025

CLIMB-ReID: A Hybrid CLIP-Mamba Framework for Person Re-Identification

AAAI 2025technical

Person Re-IDentification (ReID) aims to identify specific persons from non-overlapping cameras. Recently, some works have suggested using large-scale pre-trained vision-language models like CLIP to boost ReID performance. Unfortunately, existing methods still struggle to address two key issues simul…

2025

DeMo: Decoupled Feature-Based Mixture of Experts for Multi-Modal Object Re-Identification

AAAI 2025technical

Multi-modal object Re-IDentification (ReID) aims to retrieve specific objects by combining complementary information from multiple modalities. Existing multi-modal object ReID methods primarily focus on the fusion of heterogeneous features. However, they often overlook the dynamic quality changes in…

2025

Hierarchical Proxy Learning for Cloth-Changing Person Re-Identification

ICASSP 2025accepted

Cloth-Changing person Re-Identification (CC-ReID) depends significantly on learning discriminative features under the cloth-changing scenario. It is quite challenging due to the large intra-person variance and small inter-person variance caused by clothes changing. To address these issues, in this w…

Cited by 0SourceScholar
2025

IDEA: Inverted Text with Cooperative Deformable Aggregation for Multi-modal Object Re-Identification

CVPR 2025poster

Multi-modal object Re-IDentification (ReID) aims to retrieve specific objects by utilizing complementary information from various modalities. However, existing methods focus on fusing heterogeneous visual features, neglecting the potential benefits of text-based semantic information. To address this…

Cited by 4SourcePDFScholar
2025

KG-RAG: Enhancing GUI Agent Decision-Making via Knowledge Graph-Driven Retrieval-Augmented Generation

EMNLP 2025

Despite recent progress, Graphic User Interface (GUI) agents powered by Large Language Models (LLMs) struggle with complex mobile tasks due to limited app-specific knowledge. While UI Transition Graphs (UTGs) offer structured navigation representations, they are underutilized due to poor extraction

Cited by 0SourcePDFScholar
2025

Large Language Models for Lossless Image Compression: Next-Pixel Prediction in Language Space is All You Need

NeurIPS 2025poster

We have recently witnessed that ''Intelligence" and `''Compression" are the two sides of the same coin, where the language large model (LLM) with unprecedented intelligence is a general-purpose lossless compressor for various data modalities. This attribute is particularly appealing to the lossless…

Cited by 0SourcecodeScholar
2025

MambaPro: Multi-Modal Object Re-identification with Mamba Aggregation and Synergistic Prompt

AAAI 2025technical

Multi-modal object Re-IDentification (ReID) aims to retrieve specific objects by utilizing complementary image information from different modalities. Recently, large-scale pre-trained models like CLIP have demonstrated impressive performance in traditional single-modal ReID tasks. However, they rema…

2025

Test-time Adaptation for Image Compression with Distribution Regularization

ICLR 2025poster

Current test- or compression-time adaptation image compression (TTA-IC) approaches, which leverage both latent and decoder refinements as a two-step adaptation scheme, have potentially enhanced the rate-distortion (R-D) performance of learned image compression models on cross-domain compression task…

Cited by 1SourcePDFScholar
2025

The Devil is in Temporal Token: High Quality Video Reasoning Segmentation

CVPR 2025poster

Existing methods for Video Reasoning Segmentation rely heavily on a single special token to represent the object in the keyframe or the entire video, inadequately capturing spatial complexity and inter-frame motion. To overcome these challenges, we propose VRS-HQ, an end-to-end video reasoning segme…

2025

Towards Open-Vocabulary Remote Sensing Image Semantic Segmentation

AAAI 2025technical

Recently, deep learning based methods have revolutionized remote sensing image segmentation. However, these methods usually rely on a predefined semantic class set, thus needing additional image annotation and model training when adapting to new classes. More importantly, they are unable to segment…

2024

Asymmetric Mask Scheme for Self-Supervised Real Image Denoising

ECCV 2024poster

"In recent years, self-supervised denoising methods have gained significant success and become critically important in the field of image restoration. Among them, the blind spot network based methods are the most typical type and have attracted the attentions of a large number of researchers. Althou…

2024

Fantastic Animals and Where to Find Them: Segment Any Marine Animal with Dual SAM

CVPR 2024highlight

As an important pillar of underwater intelligence Marine Animal Segmentation (MAS) involves segmenting animals within marine environments. Previous methods don't excel in extracting long-range contextual features and overlook the connectivity between discrete pixels. Recently Segment Anything Model…

2024

MAS-SAM: Segment Any Marine Animal with Aggregated Features

IJCAI 2024poster

Recently, Segment Anything Model (SAM) shows exceptional performance in generating high-quality object masks and achieving zero-shot image segmentation. However, as a versatile vision model, SAM is primarily trained with large-scale natural light images. In underwater scenes, it exhibits substantial…

2024

Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-Identification

CVPR 2024poster

Single-modal object re-identification (ReID) faces great challenges in maintaining robustness within complex visual scenarios. In contrast multi-modal object ReID utilizes complementary information from diverse modalities showing great potentials for practical applications. However previous methods…

2024

MarvelOVD: Marrying Object Recognition and Vision-Language Models for Robust Open-Vocabulary Object Detection

ECCV 2024poster

"Learning from pseudo-labels that generated with VLMs (Vision Language Models) has been shown as a promising solution to assist open vocabulary detection (OVD) in recent studies. However, due to the domain gap between VLM and vision-detection tasks, pseudo-labels produced by the VLMs are prone to be…

2024

Part Representation Learning with Teacher-Student Decoder for Occluded Person Re-Identification

ICASSP 2024accepted

Occluded person re-identification (ReID) is a very challenging task due to the occlusion disturbance and incomplete target information. Leveraging external cues such as human pose or parsing to locate and align part features has been proven to be very effective in occluded person ReID. Meanwhile, re…

Cited by 0SourceScholar
2024

Safety-First Tracker: A Trajectory Planning Framework for Omnidirectional Robot Tracking

IROS 2024poster

This paper introduces a Safety-First Tracker (SF-Tracker) designed for omnidirectional autonomous tracking robots. The position and orientation of omnidirectional robots are decoupled for stepwise planning to ensure trajectory safety and maintain target visibility. SF-Tracker puts the trajectory saf…

Cited by 0SourcecodeScholar
2024

TF-CLIP: Learning Text-Free CLIP for Video-Based Person Re-identification

AAAI 2024technical

Large-scale language-image pre-trained models (e.g., CLIP) have shown superior performances on many cross-modal retrieval tasks. However, the problem of transferring the knowledge learned from such models to video-based person re-identification (ReID) has barely been explored. In addition, there is…

2024

TOP-ReID: Multi-Spectral Object Re-identification with Token Permutation

AAAI 2024technical

Multi-spectral object Re-identification (ReID) aims to retrieve specific objects by leveraging complementary information from different image spectra. It delivers great advantages over traditional single-spectral ReID in complex visual environment. However, the significant distribution gap among dif…

2023

Learning Progressive Modality-Shared Transformers for Effective Visible-Infrared Person Re-identification

AAAI 2023technical

Visible-Infrared Person Re-Identification (VI-ReID) is a challenging retrieval task under complex modality changes. Existing methods usually focus on extracting discriminative visual features while ignoring the reliability and commonality of visual features between different modalities. In this pape…

2022

URetinex-Net: Retinex-Based Deep Unfolding Network for Low-Light Image Enhancement

CVPR 2022poster

Retinex model-based methods have shown to be effective in layer-wise manipulation with well-designed priors for low-light image enhancement. However, the commonly used hand-crafted priors and optimization-driven solutions lead to the absence of adaptivity and efficiency. To address these issues, in…

Cited by 612PDFcodeScholar
2021

Pyramid Spatial-Temporal Aggregation for Video-Based Person Re-Identification

ICCV 2021poster

Video-based person re-identification aims to associate the video clips of the same person across multiple non-overlapping cameras. Spatial-temporal representations can provide richer and complementary information between frames, which are crucial to distinguish the target person when occlusion occur…

Cited by 107PDFcodeScholar
2021

Sparse-to-Dense Feature Matching: Intra and Inter Domain Cross-Modal Learning in Domain Adaptation for 3D Semantic Segmentation

ICCV 2021poster

Domain adaptation is critical for success when confronting with the lack of annotations in a new domain. As the huge time consumption of labeling process on 3D point cloud, domain adaptation for 3D semantic segmentation is of great expectation. With the rise of multi-modal datasets, large amount of…

Cited by 66PDFcodeScholar
2021

Watching You: Global-Guided Reciprocal Learning for Video-Based Person Re-Identification

CVPR 2021poster

Video-based person re-identification (Re-ID) aims to automatically retrieve video sequences of the same person under non-overlapping cameras. To achieve this goal, it is the key to fully utilize abundant spatial and temporal cues in videos. Existing methods usually focus on the most conspicuous imag…

Cited by 118PDFcodeScholar
2020

Semi-Supervised Crowd Counting via Self-Training on Surrogate Tasks

ECCV 2020poster

Most existing crowd counting systems rely on the availability of the object location annotation which can be expensive to obtain. To reduce the annotation cost, one attractive solution is to leverage a large number of unlabeled images to build a crowd counting model in semi-supervised fashion. This…

Cited by 96SourcePDFScholar
2019

Cascaded Context Pyramid for Full-Resolution 3D Semantic Scene Completion

ICCV 2019oral

Semantic Scene Completion (SSC) aims to simultaneously predict the volumetric occupancy and semantic category of a 3D scene. It helps intelligent devices to understand and interact with the surrounding scenes. Due to the high-memory requirement, current methods only produce low-resolution completion…

Cited by 77PDFScholar
2017

A Stagewise Refinement Model for Detecting Salient Objects in Images

ICCV 2017poster

Deep convolutional neural networks (CNNs) have been successfully applied to a wide variety of problems in computer vision, including salient object detection. To detect and segment salient objects accurately, it is necessary to extract and combine high-level semantic features with low-level fine det…

Cited by 521PDFcodeScholar
2017

Amulet: Aggregating Multi-Level Convolutional Features for Salient Object Detection

ICCV 2017poster

Fully convolutional neural networks (FCNs) have shown outstanding performance in many dense labeling problems. One key pillar of these successes is mining relevant information from features in convolutional layers. However, how to better aggregate multi-level convolutional feature maps for salient o…

Cited by 1024PDFScholar
2017

Learning Uncertain Convolutional Features for Accurate Saliency Detection

ICCV 2017poster

Deep convolutional neural networks (CNNs) have delivered superior performance in many computer vision tasks. In this paper, we propose a novel deep fully convolutional network model for accurate salient object detection. The key contribution of this work is to learn deep uncertain convolutional feat…

Cited by 447PDFScholar