← Search

Jie Zhu

26 accepted papers

2026

CARE: COGNITIVE-REASONING AUGMENTED REINFORCEMENT FOR EMOTIONAL SUPPORT CONVERSATION

ICASSP 2026poster

Emotional Support Conversation (ESC) plays a vital role in alleviating psychological stress and providing emotional value through dialogue. While recent studies have largely focused on data augmentation and synthetic corpus construction, they often overlook the deeper cognitive reasoning processes t…

Cited by 0SourcePDFScholar
2026

Evaluating, Synthesizing, and Enhancing for Customer Support Conversation

AAAI 2026technical

Effective customer support requires not only accurate problem-solving but also structured and empathetic communication aligned with professional standards. However, existing dialogue datasets often lack strategic guidance, and real-world service data is difficult to access and annotate. To address t

Cited by 0SourcePDFScholar
2026

FINMCP-BENCH: BENCHMARKING LLM AGENTS FOR REAL-WORLD FINANCIAL TOOL USE UNDER THE MODEL CONTEXT PROTOCOL

ICASSP 2026poster

This paper introduces \textbf{FinMCP-Bench}, a novel benchmark for evaluating large language models (LLMs) in solving real-world financial problems through tool invocation of financial model context protocols. FinMCP-Bench contains 613 samples spanning 10 main scenarios and 33 sub-scenarios, featuri…

Cited by 0SourcePDFScholar
2026

Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models

IJCAI 2026

Process Reward Models (PRMs) supervise intermediate reasoning steps in large language models (LLMs), but existing PRMs are mainly trained on general-domain data and struggle with the structured, symbolic, and fact-sensitive nature of financial reasoning. Financial tasks require not only correct fina

Cited by 0Scholar
2026

FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human Recognition

CVPR 2026

Model fusion is a key strategy for robust recognition in unconstrained scenarios, as different models provide complementary strengths. This is especially important for whole-body human recognition, where biometric cues such as face, gait, and body shape vary across samples and are typically integrat

Cited by 7SourcecodeScholar
2026

Plan and Budget: Effective and Efficient Test-Time Scaling on Reasoning Large Language Models

ICLR 2026poster

Large Language Models (LLMs) have achieved remarkable success in complex reasoning tasks, but their inference remains computationally inefficient. We observe a common failure mode in many prevalent LLMs, overthinking, where models generate verbose and tangential reasoning traces even for simple quer…

Cited by 0SourceScholar
2025

A Quality-Guided Mixture of Score-Fusion Experts Framework for Human Recognition

ICCV 2025poster

Whole-body biometric recognition is a challenging multi-modal task that integrates various biometric modalities, including face, gait, and body. This integration is essential for overcoming the limitations of unimodal systems. Traditionally, whole-body recognition involves deploying different models…

2025

Dual-calibrated Co-training Framework for Personalized Federated Semi-Supervised Medical Image Segmentation

AAAI 2025technical

Federated Semi-Supervised Learning (FSSL) has emerged as a crucial topic in medical image analysis, allowing multiple medical institutions to collaboratively train a global model using limited labeled data. However, existing FSSL methods focus solely on an effective combination of federated learning…

2025

Heterogeneous Graph Convolutional Neural Networks for EEG-fNIRS Bimodal Emotion Recognition

ICASSP 2025accepted

Leveraging multimodal brain signals, such as electroencephalogram (EEG) and functional near-infrared spectroscopy (fNIRS), for the objective detection of brain activity is regarded as a promising approach for affective brain-computer interface. Existing EEG-fNIRS bimodal methods primarily focus on d…

Cited by 0SourceScholar
2025

MFinMeeting: A Multilingual, Multi-Sector, and Multi-Task Financial Meeting Understanding Evaluation Dataset

ACL 2025finding

Recent breakthroughs in large language models (LLMs) have led to the development of new benchmarks for evaluating their performance in the financial domain. However, current financial benchmarks often rely on news articles, earnings reports, or announcements, making it challenging to capture the rea…

2025

M²N: A Progressive Macro-to-Micro 3D Modeling Scheme for Unveiling Drug-Target Affinity

AAAI 2025technical

Accurate drug-target affinity (DTA) prediction holds significant potential in the field of artificial intelligence (AI)-based drug discovery. However, existing methods primarily operate at a single scale, specifically at the macro (residue) scale for target proteins and the micro (atom) scale for dr…

Cited by 0SourcePDFScholar
2025

SimulPL: Aligning Human Preferences in Simultaneous Machine Translation

ICLR 2025poster

Simultaneous Machine Translation (SiMT) generates translations while receiving streaming source inputs. This requires the SiMT model to learn a read/write policy, deciding when to translate and when to wait for more source input. Numerous linguistic studies indicate that audiences in SiMT scenarios…

2025

Speech Enhancement with Overlapped-Frame Information Fusion and Causal Self-Attention

ICASSP 2025accepted

For time-frequency (TF) domain speech enhancement (SE) methods, the overlap-and-add operation in the inverse TF transformation inevitably leads to an algorithmic delay equal to the window size. However, typical causal SE systems fail to utilize the future speech information within this inherent dela…

Cited by 0SourceScholar
2024

Benchmarking Large Language Models on CFLUE - A Chinese Financial Language Understanding Evaluation Dataset

ACL 2024findings

In light of recent breakthroughs in large language models (LLMs) that have revolutionized natural language processing (NLP), there is an urgent need for new benchmarks to keep pace with the fast development of LLMs. In this paper, we propose CFLUE, the Chinese Financial Language Understanding Evalua…

2024

MoLE: Enhancing Human-centric Text-to-image Diffusion via Mixture of Low-rank Experts

NeurIPS 2024poster

Text-to-image diffusion has attracted vast attention due to its impressive image-generation capabilities. However, when it comes to human-centric text-to-image generation, particularly in the context of faces and hands, the results often fall short of naturalness due to insufficient training priors.…

Cited by 2SourcePDFScholar
2024

Progressively Learning from Macro-Expressions for Micro-Expression Recognition

ICASSP 2024accepted

Micro-expression (ME) recognition is challenging due to the low-intensity facial motions. An idea to overcome this is learning assisted by macro-expressions (MaEs). However, the intensity gap between MaE and ME is so huge that related works fail to effectively leverage MaE’s assistance in overcoming…

Cited by 0SourceScholar
2023

E-CRF: Embedded Conditional Random Field for Boundary-caused Class Weights Confusion in Semantic Segmentation

ICLR 2023poster

Modern semantic segmentation methods devote much effect to adjusting image feature representations to improve the segmentation performance in various ways, such as architecture design, attention mechnism, etc. However, almost all those methods neglect the particularity of class weights (in the class…

2023

Framewise Multiple Sound Source Localization and Counting Using Binaural Spatial Audio Signals

ICASSP 2023accepted

Sound source localization is the problem of estimating the positions of one or several sound sources. In terms of binaural audio, localization is a paramount perceptual characteristic which can be assessed subjectively or objectively. For objective evaluation of binaural sound localization, typical…

Cited by 0SourceScholar
2020

Regularized Beamformer for the Spherical Microphone Array to Cope with the White Noise Amplification

ICASSP 2020accepted

Spherical microphone arrays with compact aperture and maximum directivity factor have been one of the popular research fields but are usually accompanied by the white noise amplification problem, which hinders them for practical applications. This paper presents two regularization methods to optimiz…

Cited by 0SourceScholar
2020

Spatial Attentional Bilinear 3D Convolutional Network for Video-Based Autism Spectrum Disorder Detection

ICASSP 2020accepted

Video-based Autism Spectrum Disorder (ASD) detection is a challenge to most video classification networks due to the high degree of similarity between categories. Bilinear pooling is a second-order method, which is widely used in fine-grained visual recognition. However, the average summation in bil…

Cited by 0SourceScholar
2015

Investigating bias in non-parametric mutual information estimation

ICASSP 2015accepted

In this paper, our aim is to investigate the control of bias accumulation when estimating mutual information from nearest neighbors non-parametric approach with continuously distributed random data. Using a multidimensional Taylor series expansion, a general relationship between the estimation bias…

Cited by 0SourceScholar