← Search

Yong Zhou

20 accepted papers

2026

CLIPDet3D: Vision-Language Collaborative Distillation for 3D Object Detection

AAAI 2026technical

Multi-view 3D object detection plays a vital role in autonomous driving systems due to its ability to perceive complex scenes accurately. However, real-world driving data often exhibits a long-tailed distribution, causing significant drops in detection accuracy for rare categories in existing method

Cited by 0SourcePDFScholar
2026

Causal Decoupling Domain Generalization for Remote Sensing Change Detection

AAAI 2026technical

While current state-of-the-art Remote Sensing Change Detection (RSCD) methods can achieve impressive results on individual datasets, they become unreliable in unseen environments and imaging conditions, with performance metrics declining by as much as 60% to 80%. Simultaneously, variable environment

Cited by 0SourcePDFScholar
2026

DTTNet: Improving Video Shadow Detection via Dark-Aware Guidance and Tokenized Temporal Modeling

AAAI 2026technical

Video shadow detection confronts two entwined difficulties: distinguishing shadows from complex backgrounds and modeling dynamic shadow deformations under varying illumination. To address shadow-background ambiguity, we leverage linguistic priors through the proposed Vision-language Match Module (VM

Cited by 0SourcePDFScholar
2026

GRASP: Enhancing Zero-Shot Multi-Modal Anomaly Detection via Geometric Refinement and Semantic Prompting

IJCAI 2026

Zero-shot multimodal anomaly detection is critical for identifying structural defects that are often invisible to traditional RGB images, particularly in scenarios lacking target domain samples. However, existing methodologies face two significant impediments: the prohibitive computational overhead

Cited by 0Scholar
2026

Unified Representation Causal Prompt Distillation for Re-Inference-Free Lifelong Person Re-Identification

AAAI 2026technical

Lifelong person re-identification (LReID) aims to retrieve the target person from sequentially collected data. Due to significant domain gaps between datasets and the continuous increase of training data from different scenarios, weak inter-domain generalization and catastrophic forgetting issues ha

Cited by 0SourcePDFScholar
2025

Beyond Individual and Point: Next POI Recommendation via Region-aware Dynamic Hypergraph with Dual-level Modeling

IJCAI 2025

Next POI recommendation contributes to the prosperity of various intelligent location-based services. Existing studies focus on exploring sequential patterns and POI interactions using sequential and graph-based methods to enhance recommendation performance. However, they don't effectively exploit g

Cited by 0SourcePDFScholar
2025

Counterfactual Knowledge Maintenance for Unsupervised Domain Adaptation

IJCAI 2025

Traditional unsupervised domain adaptation (UDA) struggles to extract rich semantics due to backbone limitations. Recent large-scale pre-trained visual-language models (VLMs) have shown strong zero-shot learning capabilities in UDA tasks. However, directly using VLMs results in a mixture of semantic

2025

GSDet: Gaussian Splatting for Oriented Object Detection

IJCAI 2025

Oriented object detection has advanced with the development of convolutional neural networks (CNNs) and transformers. However, modern detectors still rely on predefined object candidates, such as anchors in CNN-based methods or queries in transformer-based methods, which struggle to capture spatial

2025

Modality-Guided Dynamic Graph Fusion and Temporal Diffusion for Self-Supervised RGB-T Tracking

IJCAI 2025

To reduce the reliance on large-scale annotations, self-supervised RGB-T tracking approaches have garnered significant attention. However, the omission of the object region by erroneous pseudo-label or the introduction of background noise affects the efficiency of modality fusion, while pseudo-label

2025

ReDiffDet: Rotation-equivariant Diffusion Model for Oriented Object Detection

CVPR 2025poster

The diffusion model has been successfully applied to various detection tasks. However, it still faces several challenges when used for oriented object detection: objects that are arbitrarily rotated require the diffusion model to encode their orientation information; uncontrollable random boxes inac…

2025

Structured IB: Improving Information Bottleneck with Structured Feature Learning

AAAI 2025technical

The Information Bottleneck (IB) principle has emerged as a promising approach for enhancing the generalization, robustness, and interpretability of deep neural networks, demonstrating efficacy across image segmentation, document clustering, and semantic communication. Among IB implementations, the I…

2024

Dual Encoder: Exploiting the Potential of Syntactic and Semantic for Aspect Sentiment Triplet Extraction

COLING 2024main

Aspect Sentiment Triple Extraction (ASTE) is an emerging task in fine-grained sentiment analysis. Recent studies have employed Graph Neural Networks (GNN) to model the syntax-semantic relationships inherent in triplet elements. However, they have yet to fully tap into the vast potential of syntactic…

2024

GMM-ResNet2: Ensemble of Group Resnet Networks for Synthetic Speech Detection

ICASSP 2024accepted

Deep learning models are widely used for speaker recognition and spoofing speech detection. We propose the GMM-ResNet2 for synthesis speech detection. Compared with the previous GMM-ResNet model, GMM-ResNet2 has four improvements. Firstly, the different order GMMs have different capabilities to form…

Cited by 0SourceScholar
2024

Gmm-Resnext: Combining Generative and Discriminative Models for Speaker Verification

ICASSP 2024accepted

With the development of deep learning, many different network architectures have been explored in speaker verification. However, most network architectures rely on a single deep learning architecture, and hybrid networks combining different architectures have been little studied in ASV task. In this…

Cited by 0SourceScholar
2024

LRANet: Towards Accurate and Efficient Scene Text Detection with Low-Rank Approximation Network

AAAI 2024technical

Recently, regression-based methods, which predict parameterized text shapes for text localization, have gained popularity in scene text detection. However, the existing parameterized text shape methods still have limitations in modeling arbitrary-shaped texts due to ignoring the utilization of text-…

2022

Gan-Based Joint Activity Detection and Channel Estimation for Grant-Free Random Access

ICASSP 2022accepted

Joint activity detection and channel estimation (JADCE) for grant-free random access is a critical issue that needs to be addressed to support massive connectivity in IoT networks. However, the existing model-free learning method can only achieve either activity detection or channel estimation, but…

Cited by 0SourceScholar
2022

Pixel-Level and Affinity-Level Knowledge Distillation for Unsupervised Segmentation of Covid-19 Lesions

ICASSP 2022accepted

Automatic segmentation of COVID-19 lesions is essential for computer-aided diagnosis. However, this task remains challenging because widely-used supervised based methods require large-scale annotated data that is difficult to obtain. Although an unsupervised method based on anomaly detection has sho…

Cited by 0SourceScholar
2022

Show, Deconfound and Tell: Image Captioning With Causal Inference

CVPR 2022poster

The transformer-based encoder-decoder framework has shown remarkable performance in image captioning. However, most transformer-based captioning methods ever overlook two kinds of elusive confounders: the visual confounder and the linguistic confounder, which generally lead to harmful bias, induce t…

Cited by 67PDFcodeScholar