← Search

Jiahui Pan

15 accepted papers

2025

An End-to-End Graph-Guided Spatiotemporal Model for Adaptive Frame-Level Facial Affect Analysis in the Wild

ICASSP 2025accepted

Human emotional states in real life are varied and complex. It is difficult for existing methods to capture robust facial expression features dynamically, especially in a large head pose and occlusion. In this paper, a novel end-to-end graph-guided spatiotemporal convolutional network (GSTCN) is pro…

Cited by 0SourceScholar
2025

Enhancing Multi-Channel Speech with Limited Microphones via Spherical Harmonic Transform

ICASSP 2025accepted

The performance of traditional beamforming algorithms is influenced by the number of microphones, with performance improving as the number increases. However, in practice, the number of microphones is often limited. In this paper, we propose a novel virtual microphone estimation method that combines…

Cited by 0SourceScholar
2025

PlanningArena: A Modular Benchmark for Multidimensional Evaluation of Planning and Tool Learning

ACL 2025long

One of the research focuses of large language models (LLMs) is the ability to generate action plans. Recent studies have revealed that the performance of LLMs can be significantly improved by integrating external tools. Based on this, we propose a benchmark framework called PlanningArena, which aims…

Cited by 0SourcePDFScholar
2025

Reading Your Heart: Learning ECG Words and Sentences via Pre-training ECG Language Model

ICLR 2025poster

Electrocardiogram (ECG) is essential for the clinical diagnosis of arrhythmias and other heart diseases, but deep learning methods based on ECG often face limitations due to the need for high-quality annotations. Although previous ECG self-supervised learning (eSSL) methods have made significant pro…

2025

VisualEDU: A Benchmark for Assessing Coding and Visual Comprehension through Educational Problem-Solving Video Generation

EMNLP 2025

Generating logically coherent video from text (T2V) for reasoning-intensive tasks like mathematical problem-solving presents a significant challenge for Vision-Language Models (VLMs). Therefore, we introduce VisualEDU, a benchmark based on Manim package to rigorously evaluate VLM capabilities in pro

2024

A Density-driven Iterative Prototype Optimization for Transductive Few-shot Learning

IJCAI 2024poster

Few-shot learning (FSL) poses a considerable challenge since it aims to improve the model generalization ability with limited labeled data. Previous works usually attempt to construct class-specific prototypes and then predict novel classes using these prototypes. However, the feature distribution r…

2024

CariesXrays: Enhancing Caries Detection in Hospital-Scale Panoramic Dental X-rays via Feature Pyramid Contrastive Learning

AAAI 2024technical

Dental caries has been widely recognized as one of the most prevalent chronic diseases in the field of public health. Despite advancements in automated diagnosis across various medical domains, it remains a substantial challenge for dental caries detection due to its inherent variability and intrica…

2024

Efficient Multi-Channel Speech Enhancement with Spherical Harmonics Injection for Directional Encoding

ICASSP 2024accepted

Multi-channel speech enhancement extracts speech using multiple microphones that capture spatial cues. Effectively utilizing directional information is therefore key. Deep learning shows great potential on multi-channel speech enhancement and often takes short-time Fourier Transform (STFT) as inputs…

Cited by 0SourceScholar
2024

Innovative Directional Encoding in Speech Processing: Leveraging Spherical Harmonics Injection for Multi-Channel Speech Enhancement

IJCAI 2024poster

Multi-channel speech enhancement leverages multiple microphones to extract target speech signals amid background noise. Effectively utilizing directional cues is key for robust enhancement. While deep learning shows promise for multi-channel speech processing, most methods operate on short-time Four…

2024

Medical Vision-Language Representation Learning with Cross-Modal Multi-Teacher Contrastive Distillation

ICASSP 2024accepted

Medical vision-language representation learning has garnered considerable attention owing to its applicability to extracting generic representations from the image and text modality. However, it still remains challenging to acquire a more comprehensive understanding of intra- and inter-modal semanti…

Cited by 0SourceScholar
2023

Two-Phase Prototypical Contrastive Domain Generalization for Cross-Subject EEG-Based Emotion Recognition

ICASSP 2023accepted

EEG signals of different individuals belong to different domains and have different data distributions because great individual distinctions exist in EEG signals. The existing methods on EEG-based emotion recognition often ignore this property and need to collect extensive EEG data for new subjects…

Cited by 0SourceScholar
2022

A Robust Deep Audio Splicing Detection Method via Singularity Detection Feature

ICASSP 2022accepted

There are many methods for detecting forged audio produced by conversion and synthesis. However, as a simpler method of forgery, splicing has not attracted widespread attention. Based on the characteristic that the tampering operation will cause singularities at high-frequency components, we propose…

Cited by 0SourceScholar
2022

Joint Temporal Convolutional Networks and Adversarial Discriminative Domain Adaptation for EEG-Based Cross-Subject Emotion Recognition

ICASSP 2022accepted

Cross-subject emotion recognition is one of the most challenging tasks in electroencephalogram (EEG)-based emotion recognition. To guarantee the constancy of feature representations across domains and to eliminate differences between domains, we explored the feasibility of combining temporal convolu…

Cited by 0SourceScholar
2018

Deep Bilinear Learning for RGB-D Action Recognition

ECCV 2018poster

In this paper, we focus on exploring modality-temporal mutual information for RGB-D action recognition. In order to learn time-varying information and multi-modal features jointly, we propose a novel deep bilinear learning framework. In the framework, we propose bilinear blocks that consist of two l…

Cited by 116SourcePDFScholar