← Search

Wentao Fan

9 accepted papers

2026

4D Local Modeling Toward Dynamic Global Perception for Ambiguity-free Rotation-Invariant Point Cloud Analysis

CVPR 2026

Rotation invariance remains a core challenge in point cloud analysis, where existing methods often struggle with structural ambiguities and insufficient global context. Most rotation-invariant (RI) representations are derived from local coordinate systems, which inherently suffer from point-pair amb

Cited by 0SourcecodeScholar
2026

DRFGD: Disentangled Representation-Focused Generative Defense for Attack-Tolerant Cross-Modal Hashing

AAAI 2026technical

With the widespread deployment of cross-modal retrieval in real-world scenarios, ensuring robustness against adversarial attacks is increasingly critical. Remarkably, deep cross-modal hashing is highly vulnerable to adversarial attacks due to its discrete nature and low-dimensional hash codes, while

Cited by 0SourcePDFScholar
2026

Enhancing Rotation-Invariant 3D Learning with Global Pose Awareness and Attention Mechanisms

AAAI 2026technical

Recent advances in rotation-invariant (RI) learning for 3D point clouds typically replace raw coordinates with handcrafted RI features to ensure robustness under arbitrary rotations. However, these approaches often suffer from the loss of global pose information, making them incapable of distinguish

Cited by 0SourcePDFScholar
2026

HyperXRec: Unifying Preference Clusters and LLM Experts for Robust Explainable Recommendations

IJCAI 2026

Explainable recommendation is crucial for building user trust, yet producing natural-language rationales that faithfully reflect the underlying decision process remains challenging. Most LLM-based explainable recommenders incorporate collaborative signals through shallow prompting or lightweight ada

Cited by 0Scholar
2026

Semantic-Aware Feature Enhancement for Partial Label Learning

AAAI 2026technical

Partial label learning (PLL) aims to learn from the data where each instance is associated with a candidate label set, with only one being valid. Most existing approaches are designed to eliminate noisy labels and use the remaining reliable ones for model training, following a label-centric learning

Cited by 0SourcePDFScholar
2025

How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?

ICCV 2025poster

Audio-visual semantic segmentation (AVSS) represents an extension of the audio-visual segmentation (AVS) task, necessitating a semantic understanding of audio-visual scenes beyond merely identifying sound-emitting objects at the visual pixel level. Contrary to a previous methodology, by decomposing…

Cited by 0SourcePDFScholar
2025

Spiking Generative Models Based on Variational Autoencoder and Adversarial Training

ICASSP 2025accepted

Deep neural networks (DNNs) have demonstrated exceptional performance across a variety of applications, yet they require substantial computing and power resources. In contrast, Spiking Neural Networks (SNNs) offer significant potential for energy-efficient computing due to their binary, event-driven…

Cited by 0SourceScholar
2017

A hierarchical Dirichlet process mixture of GID Distributions with feature selection for spatio-temporal video modeling and segmentation

ICASSP 2017accepted

In this paper, a hierarchical Dirichlet process (HDP) mixture model of generalized inverted Dirichlet (GID) distributions with an unsupervised feature selection scheme is developed. The proposed model is learned via a principled variational framework and then deployed for video modeling and segmenta…

Cited by 0SourceScholar