← Search

Yidi Li

7 accepted papers

2025

Empowering Vector Graphics with Consistently Arbitrary Viewing and View-dependent Visibility

CVPR 2025highlight

This work presents a novel text-to-vector graphics generation approach, Dream3DVG, allowing for arbitrary viewpoint viewing, progressive detail optimization, and view-dependent occlusion awareness. Our approach is a dual-branch optimization framework, consisting of an auxiliary 3D Gaussian Splattin…

2025

Multi-Stage Multimodal Distillation for Audio-Visual Speaker Tracking

ICASSP 2025accepted

Speaker tracking plays a crucial role in various human-robot interaction applications. Recently, leveraging multimodal information, such as audio and visual signals, has become an important strategy for enhancing the robustness of the tracking system. However, current methods face challenges in effe…

Cited by 0SourceScholar
2024

Adaptive Fourier Decomposition Based Signal Extraction on Weak Electromagnetic Field

ICASSP 2024accepted

Shaft-rate electromagnetic (EM) field is a critical feature in the detection of ships and underwater vehicles. However, the signal-to-noise ratio of the shaft-rate EM field is greatly reduced due to the presence of the static EM field, whose main energy is concentrated in the low-frequency section.…

Cited by 0SourceScholar
2024

AttA-NET: Attention Aggregation Network for Audio-Visual Emotion Recognition

ICASSP 2024accepted

In video-based emotion recognition, effective multi-modal fusion techniques are essential to leverage the complementary relationship between audio and visual modalities. Recent attention-based fusion methods are widely leveraged for capturing modal-shared properties. However, they often ignore the m…

Cited by 0SourceScholar
2023

Boosting Person Re-Identification with Viewpoint Contrastive Learning and Adversarial Training

ICASSP 2023accepted

Person re-identification (ReID) aims at retrieving a person of interest across multiple cameras. Despite significant progress in person ReID, viewpoint variation remains an obstacle to extracting discriminative features for retrieval. To address this problem, we propose a Viewpoint-Robust Network (V…

Cited by 0SourceScholar
2023

Cascade RDN: Towards Accurate Localization in Industrial Visual Anomaly Detection With Structural Anomaly Generation

RA-L 2023

Unsupervised visual anomaly detection uses only anomaly-free images to detect anomalous patterns, whose recent methods mainly focus on the anomaly classification sub-task but neglect to localize anomalies accurately. Existing reconstruction-based and representation-based methods yield anomaly score

Cited by 3SourceScholar
2022

Multi-Modal Perception Attention Network with Self-Supervised Learning for Audio-Visual Speaker Tracking

AAAI 2022technical

Multi-modal fusion is proven to be an effective method to improve the accuracy and robustness of speaker tracking, especially in complex scenarios. However, how to combine the heterogeneous information and exploit the complementarity of multi-modal signals remains a challenging issue. In this paper,…