← Search

Jianbo Ma

10 accepted papers

2026

BEYOND VIDEO-TO-SFX: VIDEO TO AUDIO SYNTHESIS WITH ENVIRONMENTALLY AWARE SPEECH

ICASSP 2026poster

The generation of realistic, context-aware audio is important in real-world applications such as video game development. While existing video-to-audio (V2A) methods mainly focus on Foley sound generation, they struggle to produce intelligible speech. Meanwhile, current environmental speech synthesis…

Cited by 0SourcePDFScholar
2026

Occlusion-Aware SORT: Observing Occlusion for Robust Multi-Object Tracking

CVPR 2026

Multi-object tracking (MOT) involves analyzing object trajectories and counting the number of objects in video sequences. However, 2D MOT faces challenges due to positional cost confusion arising from partial occlusion. To address this issue, we present the novel Occlusion-Aware SORT (OA-SORT) frame

Cited by 0SourceScholar
2026

Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance

CVPR 2026

Vision-Language Models (VLMs) have enhanced traditional LLMs with visual capabilities through the integration of vision encoders. While recent works have explored various combinations of vision encoders and LLMs, there still lacks a principled understanding of what makes a vision encoder suitable fo

Cited by 0SourceScholar
2026

Tracking the Unstable: Appearance-Guided Motion Modeling for Robust Multi-Object Tracking in UAV-Captured Videos

AAAI 2026technical

Multi-object tracking (MOT) aims to track multiple objects while maintaining consistent identities across frames of a given video. In unmanned aerial vehicle (UAV) recorded videos, frequent viewpoint changes and complex UAV-ground relative motion dynamics pose significant challenges, which often lea

Cited by 0SourcePDFScholar
2025

Data-Driven White Noise Gain Constrained Robust Superdirective Beamformer for Speech Enhancement

ICASSP 2025accepted

Superdirective beamformers are highly effective at suppressing directional interference and diffuse noise, but their practical use is often constrained by the problem of white noise amplification. Robust superdirective beamforming methods typically address this by imposing a constraint on the white…

Cited by 0SourceScholar
2025

Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models

ICASSP 2025accepted

Spatial audio is a crucial component in creating immersive experiences. Traditional simulation-based approaches to generate spatial audio rely on expertise, have limited scalability, and assume independence between semantic and spatial information. To address these issues, we explore end-to-end spat…

Cited by 0SourceScholar
2025

Rethinking Mamba in Speech Processing by Self-Supervised Models

ICASSP 2025accepted

The Mamba-based model has demonstrated outstanding performance across tasks in computer vision, natural language processing, and speech processing. However, in the realm of speech processing, the Mamba-based model’s performance varies across different tasks. For instance, in tasks such as speech enh…

Cited by 0SourceScholar
2024

V2A-Mapper: A Lightweight Solution for Vision-to-Audio Generation by Connecting Foundation Models

AAAI 2024technical

Building artificial intelligence (AI) systems on top of a set of foundation models (FMs) is becoming a new paradigm in AI research. Their representative and generative abilities learnt from vast amounts of data can be easily adapted and transferred to a wide range of downstream tasks without extra t…

2018

Speaker-Phonetic Vector Estimation for Short Duration Speaker Verification

ICASSP 2018accepted

Phonetic variability is one of the primary challenges in short duration speaker verification. This paper proposes a novel method that modifies the standard normal distribution prior in the total variability model to use a mixture of Gaussians as the prior distribution. The proposed speaker-phonetic…

Cited by 0SourceScholar