← Search

Jiaming Zhou

25 accepted papers

2026

DIFFA: Large Language Diffusion Models Can Listen and Understand

AAAI 2026technical

Recent advances in large language models (LLMs) have shown remarkable capabilities across textual and multimodal domains. In parallel, large language diffusion models have emerged as a promising alternative to the autoregressive paradigm, offering improved controllability, bidirectional context mode

Cited by 0SourcePDFScholar
2026

EgoTraj-Bench: Towards Robust Trajectory Prediction under Ego-View Noisy Observations

ICRA 2026poster

Reliable trajectory prediction from an ego-centric perspective is crucial for robotic navigation in human-centric environments. However, existing methods typically assume noiseless observation histories, failing to account for the perceptual artifacts inherent in first-person vision, such as occlusi…

2026

FVAR: Next-Focus Prediction for Visual Autoregressive Modeling

CVPR 2026

Visual autoregressive models achieve remarkable generation quality through next-scale predictions across multi-scale token pyramids. However, the conventional method uses uniform scale downsampling to build these pyramids, leading to aliasing artifacts that compromise fine details and introduce unwa

Cited by 0SourceScholar
2026

Reason with Thumbnails, Answer with Focus: An Efficient and Effective Paradigm for Multimodal Grounded Visual Reasoning

ICML 2026poster

To enhance the interpretability of multimodal large language models' outputs, recent efforts explored Grounded Visual Reasoning (GVR), in which the model is trained to select relevant image regions before answering the question. However, the multi-round ``ground-then-answer'' and reasoning nature of…

Cited by 0SourceScholar
2026

TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models

AAAI 2026technical

Text-to-Audio (TTA) generation has made rapid progress, but current evaluation methods remain narrow, focusing mainly on perceptual quality while overlooking robustness, generalization, and ethical concerns. We present TTA-Bench, a comprehensive benchmark for evaluating TTA models across functional

Cited by 0SourcePDFScholar
2025

ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5

ACL 2025long

Automatic speech recognition (ASR) systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0 and HuBERT. However, developing robust ASR models for young children’s speech remains challenging due to differences in pronunciation, tone,…

2025

Emotion-Preserving Prosody Anonymization Network for Voice Privacy Protection

ICASSP 2025accepted

Balancing emotion preservation and privacy protection in voice anonymization presents a significant challenge, particularly due to the difficulty of effectively handling prosody, a key feature in speech. While preserving prosodic features in anonymized speech enhances emotional expression, it also i…

Cited by 0SourceScholar
2025

Enhancing Emotion Recognition in Incomplete Data: A Novel Cross-Modal Alignment, Reconstruction, and Refinement Framework

ICASSP 2025accepted

Multimodal emotion recognition systems rely heavily on the full availability of modalities, suffering significant performance declines when modal data is incomplete. To tackle this issue, we present the Cross-Modal Alignment, Reconstruction, and Refinement (CM-ARR) framework, an innovative approach…

Cited by 0SourceScholar
2025

Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment

ICASSP 2025accepted

Multimodal emotion recognition (MER), leveraging speech and text, has emerged as a pivotal domain within human-computer interaction, demanding sophisticated methods for effective multimodal integration. The challenge of aligning features across these modalities is significant, with most existing app…

Cited by 0SourceScholar
2025

Exploring the Limits of Vision-Language-Action Manipulation in Cross-task Generalization

NeurIPS 2025poster

The generalization capabilities of vision-language-action (VLA) models to unseen tasks are crucial to achieving general-purpose robotic manipulation in open-world settings. However, the cross-task generalization capabilities of existing VLA models remain significantly underexplored. To address this…

Cited by 0SourceScholar
2025

GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation

CoRL 2025poster

Learning manipulation skills from human demonstration videos offers a promising path toward generalizable and interpretable robotic intelligence—particularly through the lens of *actionable affordances*. However, transferring such knowledge remains challenging due to: 1) a lack of large-scale data…

Cited by 0SourceScholar
2025

Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores

ICASSP 2025accepted

The kNN-CTC model has proven to be effective for monolingual automatic speech recognition (ASR). However, its direct application to multilingual scenarios like code-switching, presents challenges. Although there is potential for performance improvement, a kNN-CTC model utilizing a single bilingual d…

Cited by 0SourceScholar
2025

M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper

ICASSP 2025accepted

State-of-the-art models like OpenAI’s Whisper exhibit strong performance in multilingual automatic speech recognition (ASR), but they still face challenges in accurately recognizing diverse subdialects. In this paper, we propose M2R-Whisper, a novel multi-stage and multi-scale retrieval augmentation…

Cited by 0SourceScholar
2025

Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation

CVPR 2025poster

Learning generalizable visual representations across different embodied environments is essential for effective robotic manipulation in real-world scenarios. However, the limited scale and diversity of robot demonstration data pose a significant challenge. Recent research has explored leveraging lar…

Cited by 8SourcePDFScholar
2025

MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation

ICASSP 2025accepted

The technology for generating music from textual descriptions has seen rapid advancements. However, evaluating text-to-music (TTM) systems remains a significant challenge, primarily due to the difficulty of balancing performance and cost with existing objective and subjective evaluation methods. In…

Cited by 0SourceScholar
2025

Omni-Perception: Omnidirectional Collision Avoidance of Legged Robots in Dynamic Environments

CoRL 2025oral

Agile locomotion in complex 3D environments requires robust spatial awareness to safely avoid diverse obstacles such as aerial clutter, uneven terrain, and dynamic agents. Depth-based perception approaches often struggle with sensor noise, lighting variability, computational overhead from intermedia…

Cited by 0SourceScholar
2025

SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for Super-Aged Seniors

NeurIPS 2025poster

While voice technologies increasingly serve aging populations, current systems exhibit significant performance gaps due to inadequate training data capturing elderly-specific vocal characteristics like presbyphonia and dialectal variations. The limited data available on super-aged individuals in exi…

Cited by 0SourcecodeScholar
2024

CIF-T: A Novel CIF-Based Transducer Architecture for Automatic Speech Recognition

ICASSP 2024accepted

RNN-T models are widely used in ASR, which rely on the RNN-T loss to achieve length alignment between input audio and target sequence. However, the implementation complexity and the alignment-based optimization target of RNN-T loss lead to computational redundancy and a reduced role for predictor ne…

Cited by 0SourceScholar
2024

CKGConv: General Graph Convolution with Continuous Kernels

ICML 2024poster

The existing definitions of graph convolution, either from spatial or spectral perspectives, are inflexible and not unified. Defining a general convolution operator in the graph domain is challenging due to the lack of canonical coordinates, the presence of irregular structures, and the properties o…

2024

Contrastive Imitation Learning for Language-guided Multi-Task Robotic Manipulation

CoRL 2024poster

Developing robots capable of executing various manipulation tasks, guided by natural language instructions and visual observations of intricate real-world environments, remains a significant challenge in robotics. Such robot agents need to understand linguistic commands and distinguish between the…

Cited by 11SourceScholar
2024

KNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels

ICASSP 2024accepted

The success of retrieval-augmented language models in various natural language processing (NLP) tasks has been constrained in automatic speech recognition (ASR) applications due to challenges in constructing fine-grained audio-text datastores. This paper presents kNN-CTC, a novel approach that overc…

Cited by 0SourceScholar
2023

Diversifying Spatial-Temporal Perception for Video Domain Generalization

NeurIPS 2023poster

Video domain generalization aims to learn generalizable video classification models for unseen target domains by training in a source domain. A critical challenge of video domain generalization is to defend against the heavy reliance on domain-specific cues extracted from the source domain when reco…

2023

MADI: Inter-Domain Matching and Intra-Domain Discrimination for Cross-Domain Speech Recognition

ICASSP 2023accepted

End-to-end automatic speech recognition (ASR) usually suffers from performance degradation when applied to a new domain due to domain shift. Unsupervised domain adaptation (UDA) aims to improve the performance on the unlabeled target domain by transferring knowledge from the source to the target dom…

Cited by 0SourceScholar
2021

Graph-Based High-Order Relation Modeling for Long-Term Action Recognition

CVPR 2021poster

Long-term actions involve many important visual concepts, e.g., objects, motions, and sub-actions, and there are various relations among these concepts, which we call basic relations. These basic relations will jointly affect each other during the temporal evolution of long-term actions, which forms…

Cited by 67PDFScholar