← Search

Yifang Yin

12 accepted papers

2025

Causal-Inspired Multitask Learning for Video-Based Human Pose Estimation

AAAI 2025technical

Video-based human pose estimation has long been a fundamental yet challenging problem in computer vision. Previous studies focus on spatio-temporal modeling through the enhancement of architecture design and optimization strategies. However, they overlook the causal relationships in the joints, lead…

Cited by 1SourcePDFScholar
2025

Few-Shot Incremental Learning via Foreground Aggregation and Knowledge Transfer for Audio-Visual Semantic Segmentation

AAAI 2025technical

Audio-Visual Semantic Segmentation (AVSS) has gained significant attention in the multi-modal domain, aiming to segment video objects that produce specific sounds in the corresponding audio. Despite notable progress, existing methods still struggle to handle new classes not included in the original…

Cited by 0SourcePDFScholar
2025

HVIS: A Human-like Vision and Inference System for Human Motion Prediction

AAAI 2025technical

Grasping the intricacies of human motion, which involve perceiving spatio-temporal dependence and multi-scale effects, is essential for predicting human motion. While humans inherently possess the requisite skills to navigate this issue, it proves to be markedly more challenging for machines to emul…

Cited by 1SourcePDFScholar
2025

Priority Guided Explanation for Knowledge Tracing with Dual Ranking and Similarity Consistency

IJCAI 2025

Knowledge tracing plays a pivotal role in enabling personalized learning on online platforms. While deep learning-based approaches have achieved impressive predictive performance, their limited interpretability poses a significant barrier to practical adoption. Existing explanation methods primarily

Cited by 0SourcePDFScholar
2025

Self-Perturbed Anomaly-Aware Graph Dynamics for Multivariate Time-Series Anomaly Detection

NeurIPS 2025spotlight

Detecting anomalies in multivariate time-series data is an essential task across various domains, yet there are unresolved challenges such as (1) severe class imbalance between normal and anomalous data due to rare anomaly availability in the real world; (2) limited adaptability of the static graph-…

Cited by 0SourceScholar
2024

Expressiveness is Effectiveness: Self-supervised Fashion-aware CLIP for Video-to-Shop Retrieval

IJCAI 2024poster

The rise of online shopping and social media has spurred the Video-to-Shop Retrieval (VSR) task, which involves identifying fashion items (e.g., clothing) in videos and matching them with identical products provided by stores. In real-world scenarios, human movement in dynamic video scenes can cause…

Cited by 1SourcePDFScholar
2024

Rethinking Human Motion Prediction with Symplectic Integral

CVPR 2024poster

Long-term and accurate forecasting is the long-standing pursuit of the human motion prediction task. Existing methods typically suffer from dramatic degradation in prediction accuracy with the increasing prediction horizon. It comes down to two reasons:1? Insufficient numerical stability.Unforeseen…

Cited by 2SourcePDFScholar
2024

SOGDet: Semantic-Occupancy Guided Multi-View 3D Object Detection

AAAI 2024technical

In the field of autonomous driving, accurate and comprehensive perception of the 3D environment is crucial. Bird's Eye View (BEV) based methods have emerged as a promising solution for 3D object detection using multi-view images as input. However, existing 3D object detection methods often ignore th…

2023

CrossMatch: Source-Free Domain Adaptive Semantic Segmentation via Cross-Modal Consistency Training

ICCV 2023poster

Source-free domain adaptive semantic segmentation has gained increasing attention recently. It eases the requirement of full data access to the source domain by transferring knowledge only from a well-trained source model. However, reducing the uncertainty of the target pseudo labels becomes inevita…

Cited by 15PDFScholar
2021

A Spatial Regulated Patch-Wise Approach for Cervical Dysplasia Diagnosis

AAAI 2021technical

Cervical dysplasia diagnosis via visual investigation is a challenging problem. Recent approaches use deep learning techniques to extract features and require the downsampling of high-resolution cervical screening images to smaller sizes for training. Such a reduction may result in the loss of visua…

Cited by 8SourcePDFScholar
2021

Enhanced Audio Tagging via Multi- to Single-Modal Teacher-Student Mutual Learning

AAAI 2021technical

Recognizing ongoing events based on acoustic clues has been a critical yet challenging problem that has attracted significant research attention in recent years. Joint audio-visual analysis can improve the event detection accuracy but may not always be feasible as under many circumstances only audio…

Cited by 16SourcePDFScholar
2020

Mt-Gcn For Multi-Label Audio Tagging With Noisy Labels

ICASSP 2020accepted

Multi-label audio tagging is the task of predicting the types of sounds occurring in an audio clip. Recently, large-scale audio datasets such as Google's AudioSet, have allowed researchers to use deep learning techniques for this task but this comes at the cost of label noise in the datasets. Audio…

Cited by 0SourceScholar