← Search

Tianzhu Zhang

108 accepted papers

2026

Adaptive Agent Selection and Interaction Network for Image-to-Point Cloud Registration

AAAI 2026technical

Typical detection-free methods for image-to-point cloud registration leverage transformer-based architectures to aggregate cross-modal features and establish correspondences. However, they often struggle under challenging conditions, where noise disrupts similarity computation and leads to incorrect

Cited by 0SourcePDFScholar
2026

Adaptive Augmentation-Aware Latent Learning for Robust LiDAR Semantic Segmentation

ICLR 2026poster

Adverse weather conditions significantly degrade the performance of LiDAR point cloud semantic segmentation networks by introducing large distribution shifts. Existing augmentation-based methods attempt to enhance robustness by simulating weather interference during training. However, they struggle…

Cited by 0SourceScholar
2026

Adversarial Attacks Already Tell the Answer: Directional Bias-Guided Test-time Defense for Vision-Language Models

ICLR 2026poster

Vision-Language Models (VLMs), such as CLIP, have shown strong zero-shot generalization but remain highly vulnerable to adversarial perturbations, posing serious risks in real-world applications. Test-time defenses for VLMs have recently emerged as a promising and efficient approach to defend agains…

Cited by 0SourceScholar
2026

Beyond Logits: Coherent Hallucination Mitigation via Attention Contrastive Decoding

ICML 2026poster

Large Vision-Language Models (LVLMs) demonstrate impressive multimodal capabilities, yet suffer from hallucination—generating factually inaccurate content. Contrastive Decoding (CD) mitigates this by contrasting amateur and expert branches at the logit level. However, our investigation reveals that …

Cited by 0SourceScholar
2026

BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration

ICLR 2026poster

Diffusion Transformer has shown remarkable abilities in generating high-fidelity videos, delivering visually coherent frames and rich details over extended durations. However, existing video generation models still fall short in subject-consistent video generation due to an inherent difficulty in pa…

Cited by 0SourceScholar
2026

ComPose: A Unified Completion-Pose Framework for Robust Category-Level Object Pose Estimation

CVPR 2026

Category-level object pose estimation aims to predict the pose and size of arbitrary objects in specific categories. Existing methods struggle with the inherent incompleteness of observed point clouds, which limits their ability to capture complete object shapes for robust pose reasoning. While poin

Cited by 0SourceScholar
2026

ExMesh: EXplicit Mesh Reconstruction with Topology Adaptation

CVPR 2026

Reconstructing surface meshes from multi-view images has remained a core challenge in recent years. Most existing methods, whether implicit or explicit, depend on intermediate representations and post-processing steps like Marching Cubes or TSDF fusion, often resulting in artifacts and fragmented ge

Cited by 0SourceScholar
2026

FS-I2P: A Hierarchical Focus–Sweep Registration Network with Dynamically Allocated Depth

ICML 2026poster

Image-to-point cloud registration is often challenged by viewpoint changes, cross-modal discrepancies, and repetitive textures, which induce scale ambiguity and consequently lead to erroneous correspondences. Recent detection-free methods alleviate this issue by leveraging multi-scale features and t…

Cited by 0SourceScholar
2026

Generalizable Structure-Aware Keypoint Correspondence for Category-Unified 3D Single Object Tracking

CVPR 2026

3D single object tracking (SOT) in point clouds is essential for real-world 3D perception, yet it remains challenging due to data sparsity and large variations in scale and structure across diverse object categories. Most existing methods rely on a category-specific paradigm that trains separate mod

Cited by 0SourceScholar
2026

GeoGuide: Hierarchical Geometric Guidance for Open-Vocabulary 3D Semantic Segmentation

CVPR 2026

Open-vocabulary 3D semantic segmentation aims to segment arbitrary categories beyond the training set. Existing methods predominantly rely on distilling knowledge from 2D open-vocabulary models. However, aligning 3D features to the 2D representation space restricts intrinsic 3D geometric learning an

Cited by 0SourceScholar
2026

Learning Generalized Trackers with Elastic Token Budgets

ICML 2026poster

Visual tracking aims to estimate target states in video sequences, with applications spanning diverse computational requirements. Recent methods optimize trackers using manually pruned image tokens with a fixed budget to reduce computational costs. However, these trackers, once trained, are constrai…

Cited by 0SourceScholar
2026

MeshSplat: Generalizable Sparse-View Surface Reconstruction via Gaussian Splatting

AAAI 2026technical

Surface reconstruction has been widely studied in computer vision and graphics. However, existing surface reconstruction works struggle to recover accurate scene geometry when the input views are extremely sparse. To address this issue, we propose MeshSplat, a generalizable sparse-view surface recon

Cited by 0SourcePDFScholar
2026

PointChain: Learning Generalizable Point Cloud Representations via Structural Chain Modeling

AAAI 2026technical

Recent advances in point cloud analysis have increasingly leveraged large-scale unlabeled data through self-supervised representation learning. Autoregressive models based on next-token prediction have shown strong performance, but they usually model point clouds as linear sequences, ignoring their

Cited by 0SourcePDFScholar
2026

RayI2P: Learning Rays for Image-to-Point Cloud Registration

ICLR 2026poster

Image-to-point cloud registration aims to estimate the 6-DoF camera pose of a query image relative to a 3D point cloud map. Existing methods fall into two categories: matching-free methods regress pose directly using geometric priors, but lack fine-grained supervision and struggle with precise align…

Cited by 0SourceScholar
2026

ReFlow: Self-correction Motion Learning for Dynamic Scene Reconstruction

CVPR 2026

We present ReFlow, a unified framework for monocular dynamic scene reconstruction that learns 3D motion in a novel self-correction manner from raw video. Existing methods often suffer from incomplete scene initialization for dynamic regions, leading to unstable reconstruction and motion estimation,

Cited by 0SourceScholar
2026

Rethinking 2D-3D Registration: A Novel Network for High-Value Zone Selection and Representation Consistency Alignment

CVPR 2026

Both detection-then-match and detection-free methods have been extensively studied for image-to-point cloud registration, yet they still face significant challenges. The detection-then-match approach emphasizes high-quality correspondences but is limited by the availability of repeatable keypoints,

Cited by 0SourceScholar
2026

SCoA: Revisiting Domain Generalized Object Detection with Style-Conditioned Adaptation

ICML 2026poster

Domain generalized object detection (DGOD) aims to train an object detector on a single source domain and generalize it to unseen target domains. Recent advances in DGOD have increasingly exploited vision foundation models (VFMs) via parameter-efficient finetuning strategies. However, existing appro…

Cited by 0SourceScholar
2026

Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models

ICML 2026poster

Vision-Language Models (VLMs) are costly at inference time because they must process long sequences of visual tokens. Existing token pruning methods often degrade under high compression by blindly discarding information, breaking spatial structure or collapsing diversity. We propose SpecFlow, a trai…

Cited by 0SourceScholar
2026

SunFaded: Illumination-Aware Gaussian Splatting for Dark Scenes with Camera-Mounted Active Lighting

CVPR 2026

Gaussian Splatting has emerged as a popular 3D representation technique, but still struggles with appearance inconsistencies, especially in dark scenes that require active illumination (e.g., camera flashes or co-moving light sources) to capture usable images, leading to dramatic local appearance fl

Cited by 0SourceScholar
2025

Alleviate and Mining: Rethinking Unsupervised Domain Adaptation for Mitochondria Segmentation from Pseudo-Label Perspective

AAAI 2025technical

Mitochondria segmentation from electron microscopy (EM) images plays a crucial role in biological and medical research. However, models trained on source domains often suffer from performance degradation when applied to target domains due to domain shift. Unsupervised domain adaptation (UDA) methods…

Cited by 1SourcePDFScholar
2025

Balanced Learning for Domain Adaptive Semantic Segmentation

ICML 2025poster

Unsupervised domain adaptation (UDA) for semantic segmentation aims to transfer knowledge from a labeled source domain to an unlabeled target domain. Despite the effectiveness of self-training techniques in UDA, they struggle to learn each class in a balanced manner due to inherent class imbalance a…

Cited by 0SourcePDFScholar
2025

Beyond Confidence: Exploiting Homogeneous Pattern for Semi-Supervised Semantic Segmentation

ICML 2025poster

The critical challenge of semi-supervised semantic segmentation lies in how to fully exploit a large volume of unlabeled data to improve the model's generalization performance for robust segmentation. Existing methods mainly rely on confidence-based scoring functions in the prediction space to filte…

Cited by 0SourcePDFScholar
2025

Beyond Pixel and Object: Part Feature as Reference for Few-Shot Video Object Segmentation

AAAI 2025technical

Few-Shot Video Object Segmentation (FSVOS) aims to achieve accurate segmentation of video sequences supported by limited annotated images. In this work, we analyze the deficiencies inherent in the use of object prototypes and pixel features as references in previous methods. Then we shed light on th…

Cited by 0SourcePDFScholar
2025

BeyondMix: Leveraging Structural Priors and Long-Range Dependencies for Domain-Invariant LiDAR Segmentation

NeurIPS 2025poster

Domain adaptation for LiDAR semantic segmentation remains challenging due to the complex structural properties of point cloud data. While mix-based paradigms have shown promise, they often fail to fully leverage the rich structural priors inherent in 3D LiDAR point clouds. In this paper, we identify…

Cited by 0SourceScholar
2025

Bridge 2D-3D: Uncertainty-aware Hierarchical Registration Network with Domain Alignment

AAAI 2025technical

The method for image-to-point cloud registration typically determines the rigid transformation using a coarse-to-fine pipeline. However, directly and uniformly matching image patches with point cloud patches may lead to focusing on incorrect noise patches during matching while ignoring key ones. Mor…

Cited by 1SourcePDFScholar
2025

CA-I2P: Channel-Adaptive Registration Network with Global Optimal Selection

ICCV 2025poster

Detection-free methods typically follow a coarse-to-fine pipeline, extracting image and point cloud features for patch-level matching and refining dense pixel-to-point correspondences. However, differences in feature channel attention between images and point clouds may lead to degraded matching res…

Cited by 0SourcePDFScholar
2025

CUBE360: Learning Cubic Field Representation for Monocular Panoramic Depth Estimation

RA-L 2025

Panoramic depth estimation presents significant challenges due to the severe distortion caused by equirectangular projection (ERP) and the limited availability of panoramic RGB-D datasets. Inspired by the recent success of neural rendering, we propose a self-supervised method, named CUBE360, that le

Cited by 0SourceScholar
2025

Diffusion-based Source-biased Model for Single Domain Generalized Object Detection

ICCV 2025poster

Single domain generalized object detection aims to train an object detector on a single source domain and generalize it to any unseen domain. Although existing approaches based on data augmentation exhibit promising results, they overlook domain discrepancies across multiple augmented domains, which…

Cited by 0SourcePDFScholar
2025

Dual-Agent Optimization framework for Cross-Domain Few-Shot Segmentation

CVPR 2025poster

Cross-Domain Few-Shot Segmentation (CD-FSS) extends the generalization ability of Few-Shot Segmentation (FSS) beyond a single domain, enabling more practical applications. However, directly employing conventional FSS methods suffers from severe performance degradation in cross-domain settings, prima…

Cited by 0SourcePDFScholar
2025

EF-3DGS: Event-Aided Free-Trajectory 3D Gaussian Splatting

NeurIPS 2025spotlight

Scene reconstruction from casually captured videos has wide real-world applications. Despite recent progress, existing methods relying on traditional cameras tend to fail in high-speed scenarios due to insufficient observations and inaccurate pose estimation. Event cameras, inspired by biological vi…

Cited by 0SourceScholar
2025

Exploring Semantic Masked Autoencoder for Self-supervised Point Cloud Understanding

IJCAI 2025

Point cloud understanding aims to acquire robust and general feature representations from unlabeled data. Masked point modeling-based methods have recently shown significant performance across various downstream tasks. These pre-training methods rely on random masking strategies to establish the per

Cited by 0SourcePDFScholar
2025

Exploring Vision Semantic Prompt for Efficient Point Cloud Understanding

ICML 2025poster

A series of pre-trained models have demonstrated promising results in point cloud understanding tasks and are widely applied to downstream tasks through fine-tuning. However, full fine-tuning leads to the forgetting of pretrained knowledge and substantial storage costs on edge devices. To address th…

Cited by 0SourcePDFScholar
2025

Exploring the Better Multimodal Synergy Strategy for Vision-Language Models

AAAI 2025technical

Vision-Language models (VLMs) have shown great potential in enhancing open-world visual concept comprehension. Recent researches focus on an optimum multimodal collaboration strategy that significantly advances CLIP-based few-shot tasks. However, existing prompt-based solutions suffer from unidirect…

Cited by 0SourcePDFScholar
2025

Generalized Few-Shot Point Cloud Segmentation via LLM-Assisted Hyper-Relation Matching

ICCV 2025poster

Generalized few-shot point cloud segmentation (GFS-3DSeg) aims to segment objects of both base and novel classes using abundant base class samples and limited novel class samples. Existing GFS-3DSeg methods encounter bottlenecks due to the scarcity of novel class data and inter-class confusion. In t…

Cited by 0SourcePDFScholar
2025

Implicit Correspondence Learning for Image-to-Point Cloud Registration

CVPR 2025highlight

Image-to-point cloud registration aims to estimate the camera pose of a given image within a 3D scene point cloud. In this area, matching-based methods have achieved leading performance by first detecting the overlapping region, then matching point and pixel features learned by neural networks and f…

Cited by 0SourcePDFScholar
2025

Learning Neural Scene Representation from iToF Imaging

ICCV 2025poster

Indirect Time-of-Flight (iToF) cameras are popular for 3D perception because they are cost-effective and easy to deploy. They emit modulated infrared signals to illuminate the scene and process the received signals to generate amplitude and phase images. The depth is calculated from the phase using…

Cited by 0SourcePDFScholar
2025

Learning Shape-Independent Transformation via Spherical Representations for Category-Level Object Pose Estimation

ICLR 2025poster

Category-level object pose estimation aims to determine the pose and size of novel objects in specific categories. Existing correspondence-based approaches typically adopt point-based representations to establish the correspondences between primitive observed points and normalized object coordinates…

Cited by 2SourcePDFScholar
2025

ObjectGS: Object-aware Scene Reconstruction and Scene Understanding via Gaussian Splatting

ICCV 2025poster

3D Gaussian Splatting is renowned for its high-fidelity reconstructions and real-time novel view synthesis, yet its lack of semantic understanding limits object-level perception. In this work, we propose ObjectGS, an object-aware framework that unifies 3D scene reconstruction with semantic understan…

2025

Pamba: Enhancing Global Interaction in Point Clouds via State Space Model

AAAI 2025technical

Transformers have demonstrated impressive results for 3D point cloud semantic segmentation. However, the quadratic complexity of transformer makes computation costs high, limiting the number of points that can be processed simultaneously and impeding the modeling of long-range dependencies between o…

Cited by 0SourcePDFScholar
2025

Rethinking Correspondence-based Category-Level Object Pose Estimation

CVPR 2025poster

Category-level object pose estimation aims to determine the pose and size of arbitrary objects within given categories. Existing two-stage correspondence-based methods first establish correspondences between camera and object coordinates, and then acquire the object pose using a pose fitting algorit…

Cited by 1SourcePDFScholar
2025

Rethinking Noisy Video-Text Retrieval via Relation-aware Alignment

CVPR 2025poster

Video-Text Retrieval (VTR) is a core task in multi-modal understanding, drawing growing attention from both academia and industry in recent years. While numerous VTR methods have achieved success, most of them assume accurate visual-text correspondences during training, which is difficult to ensure…

Cited by 0SourcePDFScholar
2025

SAS: Segment Any 3D Scene with Integrated 2D Priors

ICCV 2025poster

The open vocabulary capability of 3D models is increasingly valued, as traditional methods with models trained with fixed categories fail to recognize unseen objects in complex dynamic 3D scenes. In this paper, we propose a simple yet effective approach, SAS, to integrate the open vocabulary capabil…

Cited by 0SourcePDFScholar
2025

State Space Model Meets Transformer: A New Paradigm for 3D Object Detection

ICLR 2025poster

DETR-based methods, which use multi-layer transformer decoders to refine object queries iteratively, have shown promising performance in 3D indoor object detection. However, the scene point features in the transformer decoder remain fixed, leading to minimal contributions from later decoder layers,…

Cited by 0SourcePDFScholar
2025

Structure-Aware Correspondence Learning for Relative Pose Estimation

CVPR 2025highlight

Relative pose estimation provides a promising way for achieving object-agnostic pose estimation. Despite the success of existing 3D correspondence-based methods, the reliance on explicit feature matching suffers from small overlaps in visible regions and unreliable feature estimation for invisibl…

Cited by 0SourcePDFScholar
2025

Towards Robust Pseudo-Label Learning in Semantic Segmentation: An Encoding Perspective

NeurIPS 2025poster

Pseudo-label learning is widely used in semantic segmentation, particularly in label-scarce scenarios such as unsupervised domain adaptation (UDA) and semi-supervised learning (SSL). Despite its success, this paradigm can generate erroneous pseudo-labels, which are further amplified during training…

Cited by 0SourcecodeScholar
2025

Towards Unsupervised Domain Bridging via Image Degradation in Semantic Segmentation

NeurIPS 2025poster

Semantic segmentation suffers from significant performance degradation when the trained network is applied to a different domain. To address this issue, unsupervised domain adaptation (UDA) has been extensively studied. Despite the effectiveness of selftraining techniques in UDA, they still overlo…

Cited by 0SourcecodeScholar
2024

Aggregation and Purification: Dual Enhancement Network for Point Cloud Few-shot Segmentation

IJCAI 2024poster

Point cloud few-shot semantic segmentation (PC-FSS) aims to segment objects within query samples of new categories given only a handful of annotated support samples. Although PC-FSS demonstrates enhanced category generalization capabilities compared to the fully supervised paradigm, the prevalent…

Cited by 6SourcePDFScholar
2024

BSNet: Box-Supervised Simulation-assisted Mean Teacher for 3D Instance Segmentation

CVPR 2024poster

3D instance segmentation (3DIS) is a crucial task but point-level annotations are tedious in fully supervised settings. Thus using bounding boxes (bboxes) as annotations has shown great potential. The current mainstream approach is a two-step process involving the generation of pseudo-labels from bo…

2024

DN-4DGS: Denoised Deformable Network with Temporal-Spatial Aggregation for Dynamic Scene Rendering

NeurIPS 2024poster

Dynamic scenes rendering is an intriguing yet challenging problem. Although current methods based on NeRF have achieved satisfactory performance, they still can not reach real-time levels. Recently, 3D Gaussian Splatting (3DGS) has garnered researchers' attention due to their outstanding rendering q…

2024

Diff3DETR: Agent-based Diffusion Model for Semi-supervised 3D Object Detection

ECCV 2024poster

"3D object detection is essential for understanding 3D scenes. Contemporary techniques often require extensive annotated training data, yet obtaining point-wise annotations for point clouds is time-consuming and laborious. Recent developments in semi-supervised methods seek to mitigate this problem…

Cited by 6SourcePDFScholar
2024

Electron Microscopy Images as Set of Fragments for Mitochondrial Segmentation

AAAI 2024technical

Automatic mitochondrial segmentation enjoys great popularity with the development of deep learning. However, the coarse prediction raised by the presence of regular 3D grids in previous methods regardless of 3D CNN or the vision transformers suggest a possibly sub-optimal feature arrangement. To mit…

Cited by 8SourcePDFScholar
2024

Exploring Reliable Matching with Phase Enhancement for Night-time Semantic Segmentation

ECCV 2024poster

"Semantic segmentation of night-time images holds significant importance in computer vision, particularly for applications like night environment perception in autonomous driving systems. However, existing methods tend to parse night-time images from a day-time perspective, leaving the inherent chal…

Cited by 4SourcePDFScholar
2024

Free Lunch for Gait Recognition: A Novel Relation Descriptor

ECCV 2024poster

"Gait recognition is to seek correct matches for query individuals by their unique walking patterns. However, current methods focus solely on extracting individual-specific features, overlooking “interpersonal” relationships. In this paper, we propose a novel Relation Descriptor that captures not on…

Cited by 3SourcePDFScholar
2024

Image-to-Image Matching via Foundation Models: A New Perspective for Open-Vocabulary Semantic Segmentation

CVPR 2024poster

Open-vocabulary semantic segmentation (OVS) aims to segment images of arbitrary categories specified by class labels or captions. However most previous best-performing methods whether pixel grouping methods or region recognition methods suffer from false matches between image features and category l…

Cited by 15SourcePDFScholar
2024

Instance-Adaptive and Geometric-Aware Keypoint Learning for Category-Level 6D Object Pose Estimation

CVPR 2024poster

Category-level 6D object pose estimation aims to estimate the rotation translation and size of unseen instances within specific categories. In this area dense correspondence-based methods have achieved leading performance. However they do not explicitly consider the local and global geometric inform…

2024

Localization and Expansion: A Decoupled Framework for Point Cloud Few-shot Semantic Segmentation

ECCV 2024poster

"Point cloud few-shot semantic segmentation (PC-FSS) aims to segment targets of novel categories in a given query point cloud with only a few annotated support samples. The current top-performing prototypical learning methods employ prototypes originating from support samples to direct the classific…

Cited by 5SourcePDFScholar
2024

MotionGS: Exploring Explicit Motion Guidance for Deformable 3D Gaussian Splatting

NeurIPS 2024poster

Dynamic scene reconstruction is a long-term challenge in the field of 3D vision. Recently, the emergence of 3D Gaussian Splatting has provided new insights into this problem. Although subsequent efforts rapidly extend static 3D Gaussian to dynamic scenes, they often lack explicit constraints on obje…

2024

Pay Attention to Target: Relation-Aware Temporal Consistency for Domain Adaptive Video Semantic Segmentation

AAAI 2024technical

Video semantic segmentation has achieved conspicuous achievements attributed to the development of deep learning, but suffers from labor-intensive annotated training data gathering. To alleviate the data-hunger issue, domain adaptation approaches are developed in the hope of adapting the model train…

Cited by 14SourcePDFScholar
2024

RankMatch: Exploring the Better Consistency Regularization for Semi-supervised Semantic Segmentation

CVPR 2024poster

The key lie in semi-supervised semantic segmentation is how to fully exploit substantial unlabeled data to improve the model's generalization performance by resorting to constructing effective supervision signals. Most methods tend to directly apply contrastive learning to seek additional supervisio…

2024

SD2Event:Self-supervised Learning of Dynamic Detectors and Contextual Descriptors for Event Cameras

CVPR 2024poster

Event cameras offer many advantages over traditional frame-based cameras such as high dynamic range and low latency. Therefore event cameras are widely applied in diverse computer vision applications where event-based keypoint detection is a fundamental task. However achieving robust event-based key…

Cited by 6SourcePDFScholar
2024

Task-Adaptive Prompted Transformer for Cross-Domain Few-Shot Learning

AAAI 2024technical

Cross-Domain Few-Shot Learning (CD-FSL) aims at recognizing samples in novel classes from unseen domains that are vastly different from training classes, with few labeled samples. However, the large domain gap between training and novel classes makes previous FSL methods perform poorly. To address t…

2024

Unifying Visual and Vision-Language Tracking via Contrastive Learning

AAAI 2024technical

Single object tracking aims to locate the target object in a video sequence according to the state specified by different modal references, including the initial bounding box (BBOX), natural language (NL), or both (NL+BBOX). Due to the gap between different modalities, most existing trackers are des…

2023

Adaptive Template Transformer for Mitochondria Segmentation in Electron Microscopy Images

ICCV 2023poster

Mitochondria, as tiny structures within the cell, are of significant importance to study cell functions for biological and clinical analysis. And exploring how to automatically segment mitochondria in electron microscopy (EM) images has attracted increasing attention. However, most of existing metho…

Cited by 19PDFScholar
2023

Alignment Before Aggregation: Trajectory Memory Retrieval Network for Video Object Segmentation

ICCV 2023poster

Memory-based methods in semi-supervised video object segmentation task achieve competitive performance by performing dense matching between query and memory frames. However, most of the existing methods neglect the fact that videos carry rich temporal information yet redundant spatial information. I…

Cited by 15PDFScholar
2023

Appearance Prompt Vision Transformer for Connectome Reconstruction

IJCAI 2023poster

Neural connectivity reconstruction aims to understand the function of biological reconstruction and promote basic scientific research. The intricate morphology and densely intertwined branches make it an extremely challenging task. Most previous best-performing methods adopt affinity learning or met…

Cited by 16SourcePDFScholar
2023

Camouflaged Instance Segmentation via Explicit De-Camouflaging

CVPR 2023highlight

Camouflaged Instance Segmentation (CIS) aims at predicting the instance-level masks of camouflaged objects, which are usually the animals in the wild adapting their appearance to match the surroundings. Previous instance segmentation methods perform poorly on this task as they are easily disturbed b…

Cited by 37SourcePDFScholar
2023

D2Former: Jointly Learning Hierarchical Detectors and Contextual Descriptors via Agent-Based Transformers

CVPR 2023highlight

Establishing pixel-level matches between image pairs is vital for a variety of computer vision applications. However, achieving robust image matching remains challenging because CNN extracted descriptors usually lack discriminative ability in texture-less regions and keypoint detectors are only good…

Cited by 10SourcePDFScholar
2023

DAW: Exploring the Better Weighting Function for Semi-supervised Semantic Segmentation

NeurIPS 2023poster

The critical challenge of semi-supervised semantic segmentation lies in how to fully exploit a large volume of unlabeled data to improve the model’s generalization performance for robust segmentation. Existing methods tend to employ certain criteria (weighting function) to select pixel-level pseudo…

Cited by 21SourcePDFScholar
2023

Domain Generalized Stereo Matching via Hierarchical Visual Transformation

CVPR 2023poster

Recently, deep Stereo Matching (SM) networks have shown impressive performance and attracted increasing attention in computer vision. However, existing deep SM networks are prone to learn dataset-dependent shortcuts, which fail to generalize well on unseen realistic datasets. This paper takes a step…

Cited by 26SourcePDFScholar
2023

DualRel: Semi-Supervised Mitochondria Segmentation From a Prototype Perspective

CVPR 2023poster

Automatic mitochondria segmentation enjoys great popularity with the development of deep learning. However, existing methods rely heavily on the labor-intensive manual gathering by experienced domain experts. And naively applying semi-supervised segmentation methods in the natural image field to mit…

Cited by 25SourcePDFScholar
2023

Dynamic Generative Targeted Attacks With Pattern Injection

CVPR 2023poster

Adversarial attacks can evaluate model robustness and have been of great concerns in recent years. Among various attacks, targeted attacks aim at misleading victim models to output adversary-desired predictions, which are more challenging and threatening than untargeted ones. Existing targeted attac…

Cited by 23SourcePDFScholar
2023

Focus on Query: Adversarial Mining Transformer for Few-Shot Segmentation

NeurIPS 2023poster

Few-shot segmentation (FSS) aims to segment objects of new categories given only a handful of annotated samples. Previous works focus their efforts on exploring the support information while paying less attention to the mining of the critical query branch. In this paper, we rethink the importance of…

2023

Foreground-Background Distribution Modeling Transformer for Visual Object Tracking

ICCV 2023poster

Visual object tracking is a fundamental research topic with a broad range of applications. Benefiting from the rapid development of Transformer, pure Transformer trackers have achieved great progress. However, the feature learning of these Transformer-based trackers is easily disturbed by complex ba…

Cited by 38PDFScholar
2023

Multimodal High-order Relation Transformer for Scene Boundary Detection

ICCV 2023poster

Scene boundary detection breaks down long videos into meaningful story-telling units and plays a crucial role in high-level video understanding. Despite significant advancements in this area, this task remains a challenging problem as it requires a comprehensive understanding of multimodal cues and…

Cited by 5PDFScholar
2023

Not Every Side Is Equal: Localization Uncertainty Estimation for Semi-Supervised 3D Object Detection

ICCV 2023poster

Semi-supervised 3D object detection from point cloud aims to train a detector with a small number of labeled data and a large number of unlabeled data. The core of existing methods lies in how to select high-quality pseudo-labels using the designed quality evaluation criterion. However, these method…

Cited by 5PDFcodeScholar
2023

Proposal-Based Multiple Instance Learning for Weakly-Supervised Temporal Action Localization

CVPR 2023poster

Weakly-supervised temporal action localization aims to localize and recognize actions in untrimmed videos with only video-level category labels during training. Without instance-level annotations, most existing methods follow the Segment-based Multiple Instance Learning (S-MIL) framework, where the…

2023

Query Refinement Transformer for 3D Instance Segmentation

ICCV 2023poster

3D instance segmentation aims to predict a set of object instances in a scene and represent them as binary foreground masks with corresponding semantic labels. However, object instances are diverse in shape and category,and point clouds are usually sparse, unordered, and irregular, which leads to a…

Cited by 32PDFScholar
2023

SE-ORNet: Self-Ensembling Orientation-Aware Network for Unsupervised Point Cloud Shape Correspondence

CVPR 2023poster

Unsupervised point cloud shape correspondence aims to obtain dense point-to-point correspondences between point clouds without manually annotated pairs. However, humans and some animals have bilateral symmetry and various orientations, which leads to severe mispredictions of symmetrical parts. Besid…

2022

A Keypoint-Based Global Association Network for Lane Detection

CVPR 2022poster

Lane detection is a challenging task that requires predicting complex topology shapes of lane lines and distinguishing different types of lanes simultaneously. Earlier works follow a top-down roadmap to regress predefined anchors into various shapes of lane lines, which lacks enough flexibility to f…

Cited by 154PDFcodeScholar
2022

Cross-Modality Transformer for Visible-Infrared Person Re-identification

ECCV 2022poster

"Visible-infrared person re-identification (VI-ReID) is a challenging task due to the large cross-modality discrepancies and intra-class variations. Existing works mainly focus on learning modality-shared representations by embedding different modalities into the same feature space. However, these m…

Cited by 98SourcePDFScholar
2022

Motion-Modulated Temporal Fragment Alignment Network for Few-Shot Action Recognition

CVPR 2022poster

While the majority of FSL models focus on image classification, the extension to action recognition is rather challenging due to the additional temporal dimension in videos. To address this issue, we propose an end-to-end Motion-modulated Temporal Fragment Alignment Network (MTFAN) by jointly explor…

Cited by 79PDFScholar
2021

Action Unit Memory Network for Weakly Supervised Temporal Action Localization

CVPR 2021poster

Weakly supervised temporal action localization aims to detect and localize actions in untrimmed videos with only video-level labels during training. However, without frame-level annotations, it is challenging to achieve localization completeness and relieve background interference. In this paper, we…

Cited by 106PDFScholar
2021

Diverse Part Discovery: Occluded Person Re-Identification With Part-Aware Transformer

CVPR 2021poster

Occluded person re-identification (Re-ID) is a challenging task as persons are frequently occluded by various obstacles or other persons, especially in the crowd scenario. To address these issues, we propose a novel end-to-end Part-Aware Transformer (PAT) for occluded person Re-ID through diverse pa…

Cited by 433PDFScholar
2021

Foreground Activation Maps for Weakly Supervised Object Localization

ICCV 2021poster

Weakly supervised object localization (WSOL) aims to localize objects with only image-level labels, which has better scalability and practicability than fully supervised methods in the actual deployment. However, with only image-level labels, learning object classification models tends to activate o…

Cited by 74PDFScholar
2021

Geometry Uncertainty Projection Network for Monocular 3D Object Detection

ICCV 2021poster

Monocular 3D object detection has received increasing attention due to the wide application in autonomous driving. Existing works mainly focus on introducing geometry projection to predict depth priors for each object. Despite their impressive progress, these methods neglect the geometry leverage ef…

Cited by 263PDFcodeScholar
2021

Lesion-Aware Transformers for Diabetic Retinopathy Grading

CVPR 2021poster

Diabetic retinopathy (DR) is the leading cause of permanent blindness in the working-age population. And automatic DR diagnosis can assist ophthalmologists to design tailored treatments for patients, including DR grading and lesion discovery. However, most of existing methods treat DR grading and le…

Cited by 137PDFScholar
2021

Meta-Attack: Class-Agnostic and Model-Agnostic Physical Adversarial Attack

ICCV 2021poster

Modern deep neural networks are often vulnerable to adversarial examples. Most exist attack methods focus on crafting adversarial examples in the digital domain, while only limited works study physical adversarial attack. However, it is more challenging to generate effective adversarial examples in…

Cited by 25PDFScholar
2021

Uncertainty Guided Collaborative Training for Weakly Supervised Temporal Action Detection

CVPR 2021poster

Weakly supervised temporal action detection aims to localize temporal boundaries of actions and identify their categories simultaneously with only video-level category labels during training. Among existing methods, attention-based methods have achieved superior performance by separating action and…

Cited by 105PDFScholar
2020

Cross-Modality Person Re-Identification With Shared-Specific Feature Transfer

CVPR 2020poster

Cross-modality person re-identification (cm-ReID) is a challenging but key technology for intelligent video analysis. Existing works mainly focus on learning modality-shared representation by embedding different modalities into a same feature space, lowering the upper bound of feature distinctivenes…

Cited by 434PDFScholar
2020

Graph Structured Network for Image-Text Matching

CVPR 2020poster

Image-text matching has received growing interest since it bridges vision and language. The key challenge lies in how to learn correspondence between image and text. Existing works learn coarse correspondence based on object co-occurrence statistics, while failing to learn fine-grained phrase corres…

Cited by 306PDFcodeScholar
2020

Multi-Modality Cross Attention Network for Image and Sentence Matching

CVPR 2020poster

The key of image and sentence matching is to accurately measure the visual-semantic similarity between an image and a sentence. However, most existing methods make use of only the intra-modality relationship within each modality or the inter-modality relationship between image regions and sentence w…

Cited by 470PDFScholar
2020

Self-Supervised Domain-Aware Generative Network for Generalized Zero-Shot Learning

CVPR 2020poster

Generalized Zero-Shot Learning (GZSL) aims at recognizing both seen and unseen classes by constructing correspondence between visual and semantic embedding. However, existing methods have severely suffered from the strong bias problem, where unseen instances in target domain tend to be recognized as…

Cited by 78PDFScholar
2019

GCAN: Graph Convolutional Adversarial Network for Unsupervised Domain Adaptation

CVPR 2019poster

To bridge source and target domains for domain adaptation, there are three important types of information including data structure, domain label, and class label. Most existing domain adaptation approaches exploit only one or two types of this information and cannot make them complement and enhance…

Cited by 169PDFcodeScholar
2019

RGB-Infrared Cross-Modality Person Re-Identification via Joint Pixel and Feature Alignment

ICCV 2019poster

RGB-Infrared (IR) person re-identification is an important and challenging task due to large cross-modality variations between RGB and IR images. Most conventional approaches aim to bridge the cross-modality gap with feature alignment by feature representation learning. Different from existing metho…

Cited by 479PDFScholar
2018

Joint Pose and Expression Modeling for Facial Expression Recognition

CVPR 2018poster

Facial expression recognition (FER) is a challenging task due to different expressions under arbitrary poses. Most conventional approaches either perform face frontalization on a non-frontal facial image or learn separate classifiers for each pose. Different from existing methods, in this paper, we…

2015

Structural Sparse Tracking

CVPR 2015poster

Sparse representation has been applied to visual tracking by finding the best target candidate with minimal reconstruction error by use of target templates. However, most sparse representation based trackers only consider holistic or local representations and do not make full use of the intrinsic st…

Cited by 216SourcePDFScholar