← Search

Yachao Zhang

34 accepted papers

2026

BeyondSparse: Facilitating Mamba to Enhance Cross-Domain 3D Semantic Segmentation in Adverse Weather

AAAI 2026technical

Domain generalization (DG) and domain adaptation (DA) for 3D semantic segmentation enable the model to maintain high performance while avoiding labor-intensive and time-consuming annotation of target-domain data. However, under adverse weather conditions, the injection of spatial noise will affect t

Cited by 0SourcePDFScholar
2026

Direct Segmentation without Logits Optimization for Training-Free Open-Vocabulary Semantic Segmentation

CVPR 2026

Open-vocabulary semantic segmentation (OVSS) aims to segment arbitrary category regions in images using open-vocabulary prompts, necessitating that existing methods possess pixel-level vision-language alignment capability. Typically, this capability involves computing the cosine similarity, ie, logi

Cited by 0SourcecodeScholar
2026

PC-CrossDiff: Point-Cluster Dual-Level Cross-Modal Differential Attention for Unified 3D Referring and Segmentation

AAAI 2026technical

3D Visual Grounding (3DVG) aims to localize the referent of natural language referring expressions through two core tasks: Referring Expression Comprehension (3DREC) and Segmentation (3DRES). While existing methods achieve high accuracy in simple, single-object scenes, they suffer from severe perfor

Cited by 0SourcePDFScholar
2026

Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability Perspective

AAAI 2026technical

Open-vocabulary semantic segmentation (OVSS) employs pixel-level vision-language alignment to associate category-related prompts with corresponding pixels. A key challenge is enhancing the multimodal dense prediction capability, specifically this pixel-level multimodal alignment. Although existing m

Cited by 0SourcePDFScholar
2026

UZ3DVG: Unaided Zero-Shot 3D Visual Grounding with Generated Language Conditions

CVPR 2026

Zero-Shot 3D Visual Grounding (Zero-Shot 3DVG) aims to localize target objects in 3D scenes from natural language descriptions without relying on instance-wise description annotations. Existing methods rely on extra 2D images during inference and/or require multi-turn interactions with large languag

Cited by 0SourcecodeScholar
2026

xMHashSeg: Cross-modal Hash Learning for Training-free Unsupervised LiDAR Semantic Segmentation

AAAI 2026technical

3D semantic segmentation serves as a fundamental component in many applications, such as autonomous driving and medical image analysis. Although recent methods have advanced the field, adapting these methods to new environments or object categories without extensive retraining remains a significant

Cited by 0SourcePDFScholar
2025

A Plug-and-Play Physical Motion Restoration Approach for In-the-Wild High-Difficulty Motions

ICCV 2025poster

Extracting physically plausible 3D human motion from videos is a critical task. Although existing simulation-based motion imitation methods can enhance the physical quality of daily motions estimated from monocular video capture, extending this capability to high-difficulty motions remains an open c…

2025

AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision Reward

CVPR 2025poster

Recently, text-to-motion models open new possibilities for creating realistic human motion with greater efficiency and flexibility. However, aligning motion generation with event-level textual descriptions presents unique challenges due to the complex, nuanced relationship between textual prompts an…

Cited by 1SourcePDFScholar
2025

FloorPlan-LLaMa: Aligning Architects’ Feedback and Domain Knowledge in Architectural Floor Plan Generation

ACL 2025long

Floor plans serve as a graphical language through which architects sketch and communicate their design ideas. Actually, in the Architecture, Engineering, and Construction (AEC) design stages, generating floor plans is a complex task requiring domain expertise and alignment with user requirements. Ho…

Cited by 0SourcePDFScholar
2025

Multi-Schema Proximity Network for Composed Image Retrieval

ICCV 2025poster

Composed Image Retrieval (CIR) aims to retrieve a target image using a query that combines a reference image and a textual description, benefiting users to express their intent more effectively. Despite significant advances in CIR methods, two unresolved problems remain: 1) existing methods overlook…

Cited by 0SourcePDFScholar
2025

Omni-Query Active Learning for Source-Free Domain Adaptive Cross-Modality 3D Semantic Segmentation

AAAI 2025technical

Source-Free Domain Adaptation (SFDA) aims to transfer a pre-trained source model to the unlabeled target domain without accessing the source data, thereby effectively solving labeled data dependency and domain shift problems. However, the SFDA setting faces a bottleneck due to the absence of supervi…

2025

Task-Aware Prompt Gradient Projection for Parameter-Efficient Tuning Federated Class-Incremental Learning

ICCV 2025poster

Federated Continual Learning (FCL) has recently garnered significant attention due to its ability to continuously learn new tasks while protecting user privacy. However, existing Data-Free Knowledge Transfer (DFKT) methods require training the entire model, leading to high training and communication…

Cited by 0SourcePDFScholar
2024

Chain of Generation: Multi-Modal Gesture Synthesis via Cascaded Conditional Control

AAAI 2024technical

This study aims to improve the generation of 3D gestures by utilizing multimodal information from human speech. Previous studies have focused on incorporating additional modalities to enhance the quality of generated gestures. However, these methods perform poorly when certain modalities are missing…

Cited by 14SourcePDFScholar
2024

Cross-Modal Match for Language Conditioned 3D Object Grounding

AAAI 2024technical

Language conditioned 3D object grounding aims to find the object within the 3D scene mentioned by natural language descriptions, which mainly depends on the matching between visual and natural language. Considerable improvement in grounding performance is achieved by improving the multimodal fusion…

Cited by 9SourcePDFScholar
2024

Exploring Multi-Modal Control in Music-Driven Dance Generation

ICASSP 2024accepted

Existing music-driven 3D dance generation methods mainly concentrate on high-quality dance generation, but lack sufficient control during the generation process. To address these issues, we propose a unified framework capable of generating high-quality dance movements and supporting multi-modal cont…

Cited by 0SourceScholar
2024

Learning Commonality, Divergence and Variety for Unsupervised Visible-Infrared Person Re-identification

NeurIPS 2024poster

Unsupervised visible-infrared person re-identification (USVI-ReID) aims to match specified persons in infrared images to visible images without annotations, and vice versa. USVI-ReID is a challenging yet underexplored task. Most existing methods address the USVI-ReID through cluster-based contrastiv…

2024

Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance Primitives

CVPR 2024poster

We propose Lodge a network capable of generating extremely long dance sequences conditioned on given music. We design Lodge as a two-stage coarse to fine diffusion architecture and propose the characteristic dance primitives that possess significant expressiveness as intermediate representations bet…

2024

MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space Models

NeurIPS 2024poster

Gesture synthesis is a vital realm of human-computer interaction, with wide-ranging applications across various fields like film, robotics, and virtual reality. Recent advancements have utilized the diffusion model to improve gesture synthesis. However, the high computational complexity of these t…

2024

Multi-Memory Matching for Unsupervised Visible-Infrared Person Re-Identification

ECCV 2024poster

"Unsupervised visible-infrared person re-identification (USL-VI-ReID) is a promising yet highly challenging retrieval task. The key challenges in USL-VI-ReID are to accurately generate pseudo-labels and establish pseudo-label correspondences across modalities without relying on any prior annotations…

2024

One-Stage Training Generative Paradigm for Generalized Zero-Shot Learning

ICASSP 2024accepted

Zero-shot learning image classification aims to identify unseen classes not present during training. Generalized zero-shot learning (GZSL) is more in line with realistic scenarios due to its ability of recognizing both seen and unseen classes. Current GZSL methods mostly utilize generative adversari…

Cited by 0SourceScholar
2024

Strategic Preys Make Acute Predators: Enhancing Camouflaged Object Detectors by Generating Camouflaged Objects

ICLR 2024poster

Camouflaged object detection (COD) is the challenging task of identifying camouflaged objects visually blended into surroundings. Albeit achieving remarkable success, existing COD detectors still struggle to obtain precise results in some challenging cases. To handle this problem, we draw inspiratio…

2024

Text2Avatar: Text to 3d Human Avatar Generation with Codebook-Driven Body Controllable Attribute

ICASSP 2024accepted

Generating 3D human models directly from text helps reduce the cost and time of character modeling. However, achieving multi-attribute controllable and realistic 3D human avatar generation is still challenging due to feature coupling and the scarcity of realistic 3D human avatar datasets. To address…

Cited by 0SourceScholar
2024

UniDSeg: Unified Cross-Domain 3D Semantic Segmentation via Visual Foundation Models Prior

NeurIPS 2024poster

3D semantic segmentation using an adapting model trained from a source domain with or without accessing unlabeled target-domain data is the fundamental task in computer vision, containing domain adaptation and domain generalization. The essence of simultaneously solving cross-domain tasks is to enha…

2023

BEV-DG: Cross-Modal Learning under Bird's-Eye View for Domain Generalization of 3D Semantic Segmentation

ICCV 2023poster

Cross-modal Unsupervised Domain Adaptation (UDA) aims to exploit the complementarity of 2D-3D data to overcome the lack of annotation in a new domain. However, UDA methods rely on access to the target domain during training, meaning the trained model only works in a specific target domain. In light…

Cited by 17PDFScholar
2023

Camouflaged Object Detection With Feature Decomposition and Edge Reconstruction

CVPR 2023poster

Camouflaged object detection (COD) aims to address the tough issue of identifying camouflaged objects visually blended into the surrounding backgrounds. COD is a challenging task due to the intrinsic similarity of camouflaged objects with the background, as well as their ambiguous boundaries. Existi…

Cited by 260SourcePDFScholar
2023

Dual Pseudo-Labels Interactive Self-Training for Semi-Supervised Visible-Infrared Person Re-Identification

ICCV 2023poster

Visible-infrared person re-identification (VI-ReID) aims to match a specific person from a gallery of images captured from non-overlapping visible and infrared cameras. Most works focus on fully supervised VI-ReID, which requires substantial cross-modality annotation that is more expensive than the…

Cited by 42PDFcodeScholar
2023

Efficient Converted Spiking Neural Network for 3D and 2D Classification

ICCV 2023poster

Spiking Neural Networks (SNNs) have attracted enormous research interest due to their low-power and biologically plausible nature. Existing ANN-SNN conversion methods can achieve lossless conversion by converting a well-trained Artificial Neural Network (ANN) into an SNN. However, converted SNN requ…

Cited by 16PDFScholar
2023

FineDance: A Fine-grained Choreography Dataset for 3D Full Body Dance Generation

ICCV 2023poster

Generating full-body and multi-genre dance sequences from given music is a challenging task, due to the limitations of existing datasets and the inherent complexity of the fine-grained hand motion and dance genres. To address these problems, we propose FineDance, which contains 14.6 hours of music-…

Cited by 58PDFcodeScholar
2023

VS-Boost: Boosting Visual-Semantic Association for Generalized Zero-Shot Learning

IJCAI 2023poster

Unlike conventional zero-shot learning (CZSL) which only focuses on the recognition of unseen classes by using the classifier trained on seen classes and semantic embeddings, generalized zero-shot learning (GZSL) aims at recognizing both the seen and unseen classes, so it is more challenging due to…

Cited by 17SourcePDFScholar
2023

Weakly Supervised 3D Segmentation via Receptive-Driven Pseudo Label Consistency and Structural Consistency

AAAI 2023technical

As manual point-wise label is time and labor-intensive for fully supervised large-scale point cloud semantic segmentation, weakly supervised method is increasingly active. However, existing methods fail to generate high-quality pseudo labels effectively, leading to unsatisfactory results. In this p…

Cited by 11SourcePDFScholar
2023

Weakly-Supervised Concealed Object Segmentation with SAM-based Pseudo Labeling and Multi-scale Feature Grouping

NeurIPS 2023poster

Weakly-Supervised Concealed Object Segmentation (WSCOS) aims to segment objects well blended with surrounding environments using sparsely-annotated data for model training. It remains a challenging task since (1) it is hard to distinguish concealed objects from the background due to the intrinsic s…

Cited by 132SourcePDFScholar
2021

Perturbed Self-Distillation: Weakly Supervised Large-Scale Point Cloud Semantic Segmentation

ICCV 2021poster

Large-scale point cloud semantic segmentation has wide applications. Current popular researches mainly focus on fully supervised learning which demands expensive and tedious manual point-wise annotation. Weakly supervised learning is an alternative way to avoid this exhausting annotation. However, f…

Cited by 162PDFScholar
2021

Weakly Supervised Semantic Segmentation for Large-Scale Point Cloud

AAAI 2021technical

Existing methods for large-scale point cloud semantic segmentation require expensive, tedious and error-prone manual point-wise annotation. Intuitively, weakly supervised training is a direct solution to reduce the labeling costs. However, for weakly supervised large-scale point cloud semantic segme…