← Search

Jing Han

26 accepted papers

2026

MindTracker: Unveiling Implicit Emotions in Long-Horizon Dialogues

IJCAI 2026

Affective computing has achieved notable success in recognizing explicit emotions from short, isolated dialogue segments. However, human emotions are often implicitly expressed, internally regulated, and dynamically evolve over extended interactions. Existing models struggle to disentangle internal

Cited by 0Scholar
2026

PocketLLM: Ultimate Compression of Large Language Models via Meta Networks

AAAI 2026technical

As Large Language Models (LLMs) continue to grow in size, storing and transmitting them on edge devices becomes increasingly challenging. Traditional methods like quantization and pruning struggle to achieve extreme compression of LLMs without sacrificing accuracy. In this paper, we introduce Pocket

Cited by 0SourcePDFScholar
2026

SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs

AAAI 2026technical

Neural Speech Codecs face a fundamental trade-off at low bitrates: preserving acoustic fidelity often compromises semantic richness. To address this, we introduce SACodec, a novel codec built upon an asymmetric dual-quantizer that employs our proposed Semantic Anchoring mechanism. This design strate

Cited by 0SourcePDFScholar
2025

DepMGNN: Matrixial Graph Neural Network for Video-based Automatic Depression Assessment

AAAI 2025technical

Depression can be reflected by long-term human spatio-temporal facial behaviours. While human face videos recorded in real-world usually have long and variable lengths, existing video-based depression assessment approaches frequently re-sample/down-sample such videos to short and equal-length videos…

2025

DiC: Rethinking Conv3x3 Designs in Diffusion Models

CVPR 2025poster

Diffusion models have shown exceptional performance in visual generation tasks. Recently, these models have shifted from traditional U-Shaped CNN-Attention hybrid structures to fully transformer-based isotropic architectures. While these transformers exhibit strong scalability and performance, their…

2025

Heart Sounds for High Blood Pressure Prediction

ICASSP 2025accepted

Hypertension, a major risk factor for cardiovascular diseases, often goes undetected due to its asymptomatic nature. This study explores a novel approach to detecting elevated blood pressure using heart sounds, aiming to provide a non-invasive, potentially continuous monitoring solution. We evaluate…

Cited by 0SourceScholar
2025

MoRAgent: Parameter Efficient Agent Tuning with Mixture-of-Roles

ICML 2025poster

Despite recent advancements of fine-tuning large language models (LLMs) to facilitate agent tasks, parameter-efficient fine-tuning (PEFT) methodologies for agent remain largely unexplored. In this paper, we introduce three key strategies for PEFT in agent tasks: 1) Inspired by the increasingly domin…

2025

Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling

ICASSP 2025accepted

The lack of labeled data is a common challenge in speech classification tasks, particularly those requiring extensive subjective assessment, such as cognitive state classification. In this work, we propose a Semi-Supervised Learning (SSL) framework, introducing a novel multi-view pseudo-labeling met…

Cited by 0SourceScholar
2025

SpeCache: Speculative Key-Value Caching for Efficient Generation of LLMs

ICML 2025poster

Transformer-based large language models (LLMs) have already achieved remarkable results on long-text tasks, but the limited GPU memory (VRAM) resources struggle to accommodate the linearly growing demand for key-value (KV) cache as the sequence length increases, which has become a bottleneck for the…

Cited by 0SourcePDFScholar
2025

UNICL-SAM: Uncertainty-Driven In-Context Segmentation with Part Prototype Discovery

CVPR 2025poster

Recent advancements in in-context segmentation generalists have demonstrated significant success in performing various image segmentation tasks using a limited number of labeled example images. However, real-world applications present challenges due to the variability of support examples, which ofte…

Cited by 0SourcePDFScholar
2024

Customising General Large Language Models for Specialised Emotion Recognition Tasks

ICASSP 2024accepted

The advent of large language models (LLMs) has gained tremendous attention over the past year. Previous studies have shown the astonishing performance of LLMs not only in other tasks but also in emotion recognition in terms of accuracy, universality, explanation, robustness, few/zero-shot learning,…

Cited by 0SourceScholar
2024

Domain Separation Graph Neural Networks for Saliency Object Ranking

CVPR 2024poster

Saliency object ranking (SOR) has attracted significant attention recently. Previous methods usually failed to explicitly explore the saliency degree-related relationships between objects. In this paper we propose a novel Domain Separation Graph Neural Network (DSGNN) which starts with separately ex…

2024

HAFFormer: A Hierarchical Attention-Free Framework for Alzheimer's Disease Detection From Spontaneous Speech

ICASSP 2024accepted

Automatically detecting Alzheimer’s Disease (AD) from spontaneous speech plays an important role in its early diagnosis. Recent approaches highly rely on the Transformer architectures due to its efficiency in modelling long-range context dependencies. However, the quadratic increase in computational…

Cited by 0SourceScholar
2024

Intelligent Cardiac Auscultation for Murmur Detection via Parallel-Attentive Models with Uncertainty Estimation

ICASSP 2024accepted

Heart murmurs are a common manifestation of cardiovascular diseases and can provide crucial clues to early cardiac abnormalities. While most current research methods primarily focus on the accuracy of models, they often overlook other important aspects such as the interpretability of machine learnin…

Cited by 0SourceScholar
2024

Towards Open Respiratory Acoustic Foundation Models: Pretraining and Benchmarking

NeurIPS 2024poster

Respiratory audio, such as coughing and breathing sounds, has predictive power for a wide range of healthcare applications, yet is currently under-explored. The main problem for those applications arises from the difficulty in collecting large labeled task-specific data for model development. Genera…

2023

Cross-Device Federated Learning for Mobile Health Diagnostics: A First Study on COVID-19 Detection

ICASSP 2023accepted

Federated learning (FL) aided health diagnostic models can incorporate data from a large number of personal edge devices (e.g., mobile phones) while keeping the data local to the originating devices, largely ensuring privacy. However, such a cross-device FL approach for health diagnostics still impo…

Cited by 0SourceScholar
2021

COVID-19 Sounds: A Large-Scale Audio Dataset for Digital Respiratory Screening

NeurIPS 2021poster

Audio signals are widely recognised as powerful indicators of overall health status, and there has been increasing interest in leveraging sound for affordable COVID-19 screening through machine learning. However, there has also been scepticism regarding the initial efforts, due to perhaps the lack o…

Cited by 84SourceScholar
2021

Exploring Automatic COVID-19 Diagnosis via Voice and Symptoms from Crowdsourced Data

ICASSP 2021accepted

The development of fast and accurate screening tools, which could facilitate testing and prevent more costly clinical tests, is key to the current pandemic of COVID-19. In this context, some initial work shows promise in detecting diagnostic signals of COVID-19 from audio sounds. In this paper, we p…

Cited by 0SourceScholar
2020

Generating and Protecting Against Adversarial Attacks for Deep Speech-Based Emotion Recognition Models

ICASSP 2020accepted

The development of deep learning models for speech emotion recognition has become a popular area of research. Adversarially generated data can cause false predictions, and in an endeavor to ensure model robustness, defense methods against such attacks should be addressed. With this in mind, in this…

Cited by 0SourceScholar
2019

Attention-based Atrous Convolutional Neural Networks: Visualisation and Understanding Perspectives of Acoustic Scenes

ICASSP 2019accepted

The goal of Acoustic Scene Classification (ASC) is to recognise the environment in which an audio waveform has been recorded. Recently, deep neural networks have been applied to ASC and have achieved state-of-the-art performance. However, few works have investigated how to visualise and understand w…

Cited by 0SourceScholar
2019

Compact Convolutional Recurrent Neural Networks via Binarization for Speech Emotion Recognition

ICASSP 2019accepted

Despite the great advances, most of the recently developed automatic speech recognition systems focus on working in a server-client manner, and thus often require a high computational cost, such as the storage size and memory accesses. This, however, does not satisfy the increasing demand for a succ…

Cited by 0SourceScholar
2019

Implicit Fusion by Joint Audiovisual Training for Emotion Recognition in Mono Modality

ICASSP 2019accepted

Despite significant advances in emotion recognition from one individual modality, previous studies fail to take advantage of other modalities to train models in mono-modal scenarios. In this work, we propose a novel joint training model which implicitly fuses audio and visual information in the trai…

Cited by 0SourceScholar
2018

Towards Conditional Adversarial Training for Predicting Emotions from Speech

ICASSP 2018accepted

Motivated by the encouraging results recently obtained by generative adversarial networks in various image processing tasks, we propose a conditional adversarial training framework to predict dimensional representations of emotion, i. e., arousal and valence, from speech signals. The framework consi…

Cited by 0SourceScholar
2017

Prediction-based learning for continuous emotion recognition in speech

ICASSP 2017accepted

In this paper, a prediction-based learning framework is proposed for a continuous prediction task of emotion recognition from speech, which is one of the key components of affective computing in multimedia. The main goal of this framework is to utmost exploit the individual advantages of different r…

Cited by 0SourceScholar
2017

Reconstruction-error-based learning for continuous emotion recognition in speech

ICASSP 2017accepted

To advance the performance of continuous emotion recognition from speech, we introduce a reconstruction-error-based (RE-based) learning framework with memory-enhanced Recurrent Neural Networks (RNN). In the framework, two successive RNN models are adopted, where the first model is used as an autoenc…

Cited by 0SourceScholar
2016

Cross lingual speech emotion recognition using canonical correlation analysis on principal component subspace

ICASSP 2016accepted

This paper proposes an analytical approach based on Kernel Canonical Correlation Analysis (KCCA) for domain adaptation. To generate paired instances for KCCA, we mapped source and target data onto both source and target principal components. We performed pair-wise domain adaptation between four emot…

Cited by 0SourceScholar