← Search

Xianfeng Wang

3 accepted papers

2026

LifeEval: A Multimodal Benchmark for Assistive AI in Egocentric Daily Life Tasks

CVPR 2026

The rapid progress of Multimodal Large Language Models (MLLMs) marks a significant step toward artificial general intelligence, offering great potential for augmenting human capabilities. However, their ability to provide effective assistance in dynamic, real-world environments remains largely under

Cited by 0SourceScholar
2021

A Novel end-to-end Speech Emotion Recognition Network with Stacked Transformer Layers

ICASSP 2021accepted

Speech emotion recognition (SER) aims to automatically recognize emotional category for a given speech utterance. The performance of a SER system heavily relies on the effectiveness of global representation expressed at utterance level. To effectively extract such a global feature, the mainstream of…

Cited by 0SourceScholar
2021

Unimodal and Crossmodal Refinement Network for Multimodal Sequence Fusion

EMNLP 2021main

Effective unimodal representation and complementary crossmodal representation fusion are both important in multimodal representation learning. Prior works often modulate one modal feature to another straightforwardly and thus, underutilizing both unimodal and crossmodal representation refinements, w…