← Search

Wei Qian

16 accepted papers

2026

Face-Guided Sentiment Boundary Enhancement for Weakly-Supervised Temporal Sentiment Localization

CVPR 2026

Point-level weakly-supervised temporal sentiment localization (P-WTSL) aims to detect sentiment-relevant segments in untrimmed multimodal videos using timestamp sentiment annotations, which greatly reduces the costly frame-level labeling. To tackle the intrinsic challenges of imprecise sentiment bou

Cited by 0SourcecodeScholar
2026

Towards Benchmarking Privacy Vulnerabilities in Selective Forgetting with Large Language Models

AAAI 2026technical

The rapid advancements in artificial intelligence (AI) have primarily focused on the process of learning from data to acquire knowledgeable learning systems. As these systems are increasingly deployed in critical areas, ensuring their privacy and alignment with human values is paramount. Recently, s

Cited by 0SourcePDFScholar
2025

Dense Audio-Visual Event Localization Under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration

AAAI 2025technical

In the field of audio-visual learning, most research tasks focus exclusively on short videos. This paper focuses on the more practical Dense Audio-Visual Event Localization (DAVEL) task, advancing audio-visual scene understanding for longer, untrimmed videos. This task seeks to identify and temporal…

2025

Neuron Explanations for Conformal Prediction (Student Abstract)

AAAI 2025technical

Conformal prediction (CP) has gained prominence as a popular technique for uncertainty quantification in deep neural networks (DNNs), providing statistically rigorous uncertainty sets. However, existing CP methods fail to clarify the origins of predictive uncertainties. While neuron-level interpreta…

Cited by 0SourcePDFScholar
2025

PhysDiff: Physiology-based Dynamicity Disentangled Diffusion Model for Remote Physiological Measurement

AAAI 2025technical

Recent works on remote PhotoPlethysmoGraphy (rPPG) estimation typically use techniques like CNNs and Transformers to encode implicit features from facial videos for prediction. These methods learn to directly map facial videos to the static values of rPPG signals, overlooking the inherent dynamic ch…

2024

Automated Natural Language Explanation of Deep Visual Neurons with Large Models (Student Abstract)

AAAI 2024technical

Interpreting deep neural networks through examining neurons offers distinct advantages when it comes to exploring the inner workings of Deep Neural Networks. Previous research has indicated that specific neurons within deep vision networks possess semantic meaning and play pivotal roles in model per…

Cited by 0SourcePDFScholar
2024

Data Poisoning Attacks against Conformal Prediction

ICML 2024poster

The efficient and theoretically sound uncertainty quantification is crucial for building trust in deep learning models. This has spurred a growing interest in conformal prediction (CP), a powerful technique that provides a model-agnostic and distribution-free method for obtaining conformal predictio…

Cited by 4SourcePDFScholar
2024

Learning Occupancy for Monocular 3D Object Detection

CVPR 2024poster

Monocular 3D detection is a challenging task due to the lack of accurate 3D information. Existing approaches typically rely on geometry constraints and dense depth estimates to facilitate the learning but often fail to fully exploit the benefits of three-dimensional feature extraction in frustum and…

2024

Regulating Intermediate 3D Features for Vision-Centric Autonomous Driving

AAAI 2024technical

Multi-camera perception tasks have gained significant attention in the field of autonomous driving. However, existing frameworks based on Lift-Splat-Shoot (LSS) in the multi-camera setting cannot produce suitable dense 3D features due to the projection nature and uncontrollable densification process…

2024

Rethinking Adversarial Robustness in the Context of the Right to be Forgotten

ICML 2024poster

The past few years have seen an intense research interest in the practical needs of the "right to be forgotten", which has motivated researchers to develop machine unlearning methods to unlearn a fraction of training data and its lineage. While existing machine unlearning methods prioritize the prot…

Cited by 5SourcePDFScholar
2024

Towards Modeling Uncertainties of Self-Explaining Neural Networks via Conformal Prediction

AAAI 2024technical

Despite the recent progress in deep neural networks (DNNs), it remains challenging to explain the predictions made by DNNs. Existing explanation methods for DNNs mainly focus on post-hoc explanations where another explanatory model is employed to provide explanations. The fact that post-hoc methods…

Cited by 6SourcePDFScholar
2023

MonoNeRD: NeRF-like Representations for Monocular 3D Object Detection

ICCV 2023poster

In the field of monocular 3D detection, it is common practice to utilize scene geometric clues to enhance the detector's performance. However, many existing works adopt these clues explicitly such as estimating a depth map and back-projecting it into 3D space. This explicit methodology induces spars…

Cited by 35PDFcodeScholar
2023

Static and Sequential Malicious Attacks in the Context of Selective Forgetting

NeurIPS 2023poster

With the growing demand for the right to be forgotten, there is an increasing need for machine learning models to forget sensitive data and its impact. To address this, the paradigm of selective forgetting (a.k.a machine unlearning) has been extensively studied, which aims to remove the impact of re…

Cited by 19SourcePDFScholar
2022

Graph R-CNN: Towards Accurate 3D Object Detection with Semantic-Decorated Local Graph

ECCV 2022poster

"Two-stage detectors have gained much popularity in 3D object detection. Most two-stage 3D detectors utilize grid points, voxel grids, or sampled keypoints for RoI feature extraction in the second stage. Such methods, however, are inefficient in handling unevenly distributed and sparse outdoor point…

2019

Global Convergence of Least Squares EM for Demixing Two Log-Concave Densities

NeurIPS 2019poster

This work studies the location estimation problem for a mixture of two rotation invariant log-concave densities. We demonstrate that Least Squares EM, a variant of the EM algorithm, converges to the true location parameter from a randomly initialized point. Moreover, we establish the explicit conver…