← Search

Ming Cheng

37 accepted papers

2026

DiffCrossGait: Trajectory-Level Alignment for 2D-3D Cross-Modal Gait Recognition via Latent Diffusion

ICML 2026poster

Cross-modal 2D–3D gait recognition is impeded by inherent domain discrepancies between 2D silhouette and 3D point cloud distributions. While prior methods align only final embeddings, we propose DiffCrossGait, which enforces trajectory-level alignment by driving both modalities with shared noise in …

Cited by 0SourceScholar
2026

ENHANCING SPEAKER VERIFICATION WITH W2V-BERT 2.0 AND KNOWLEDGE DISTILLATION GUIDED STRUCTURED PRUNING

ICASSP 2026oral

Large-scale self-supervised Pre-Trained Models (PTMs) have shown significant improvements in the speaker verification (SV) task by providing rich feature representations. In this paper, we utilize w2v-BERT 2.0, a model with approximately 600 million parameters trained on 4.5 million hours of unlabel…

Cited by 0SourcePDFScholar
2026

One-Turn Knockout: Traceable and Editable Proxy Unlearning Under Asymmetric Access Constraints

IJCAI 2026

Machine unlearning (MUL) aims to remove the influence of specific data from a trained model for data privacy and model adaptability. Existing MUL methods mostly assume the internal parameters and the training data of the target model are accessible. Nevertheless, in most practical scenarios, the mod

Cited by 0Scholar
2026

Physically-Based LiDAR Smoke Simulation for Robust 3D Object Detection

AAAI 2026technical

3D object detection in adverse weather is crucial for autonomous driving, especially in smoke where LiDAR data becomes sparse and noisy. Due to the lack of real smoke data, this paper introduces a physics-based simulation framework to generate realistic LiDAR point clouds of smoke and augment large-

Cited by 0SourcePDFScholar
2026

Walking Further: Semantic-Aware Multimodal Gait Recognition Under Long-Range Conditions

AAAI 2026technical

Gait recognition is an emerging biometric technology that enables non-intrusive and hard-to-spoof human identification. However, most existing methods are confined to short-range, unimodal settings and fail to generalize to long-range and cross-distance scenarios under real-world conditions. To addr

Cited by 0SourcePDFScholar
2025

A Generalizable Rhetorical Strategy Annotation Model Using LLM-based Debate Simulation and Labelling

EMNLP 2025

Rhetorical strategies are central to persuasive communication, from political discourse and marketing to legal argumentation. However, analysis of rhetorical strategies has been limited by reliance on human annotation, which is costly, inconsistent, difficult to scale. Their associated datasets are

Cited by 0SourcePDFScholar
2025

GS-CPR: Efficient Camera Pose Refinement via 3D Gaussian Splatting

ICLR 2025poster

We leverage 3D Gaussian Splatting (3DGS) as a scene representation and propose a novel test-time camera pose refinement (CPR) framework, GS-CPR. This framework enhances the localization accuracy of state-of-the-art absolute pose regression and scene coordinate regression methods. The 3DGS model rend…

Cited by 7SourcePDFScholar
2025

Learning Sparsity for Effective and Efficient Music Performance Question Answering

ACL 2025short

Music performances, characterized by dense and continuous audio as well as seamless audio-visual integration, present unique challenges for multimodal scene understanding and reasoning. Recent Music Performance Audio-Visual Question Answering (Music AVQA) datasets have been proposed to reflect these…

Cited by 0SourcePDFScholar
2025

Point-Cache: Test-time Dynamic and Hierarchical Cache for Robust and Generalizable Point Cloud Analysis

CVPR 2025poster

This paper proposes a general solution to enable point cloud recognition models to handle distribution shifts at test time. Unlike prior methods, which rely heavily on training data (often inaccessible during online inference) and are limited to recognizing a fixed set of point cloud classes predefi…

2025

ProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering

EMNLP 2025

Visual Question Answering (VQA) is increasingly used in diverse applications ranging from general visual reasoning to safety-critical domains such as medical imaging and autonomous systems, where models must provide not only accurate answers but also explanations that humans can easily understand an

Cited by 0SourcePDFScholar
2025

Role-aware Multi-agent Reinforcement Learning for Coordinated Emergency Traffic Control

NeurIPS 2025poster

Emergency traffic control presents an increasingly critical challenge, requiring seamless coordination among emergency vehicles, regular vehicles, and traffic lights to ensure efficient passage for all vehicles. Existing models primarily only focus on traffic light control, leaving emergency and reg…

Cited by 0SourceScholar
2025

Sci-LoRA: Mixture of Scientific LoRAs for Cross-Domain Lay Paraphrasing

ACL 2025finding

Lay paraphrasing aims to make scientific information accessible to audiences without technical backgrounds. However, most existing studies focus on a single domain, such as biomedicine. With the rise of interdisciplinary research, it is increasingly necessary to comprehend knowledge spanning multipl…

2025

Temporal Working Memory: Query-Guided Segment Refinement for Enhanced Multimodal Understanding

NAACL 2025findings

Multimodal foundation models (MFMs) have demonstrated significant success in tasks such as visual captioning, question answering, and image-text retrieval. However, these models face inherent limitations due to their finite internal capacity, which restricts their ability to process extended tempora…

2025

VTechAGP: An Academic-to-General-Audience Text Paraphrase Dataset and Benchmark Models

NAACL 2025long

Existing text simplification or paraphrase datasets mainly focus on sentence-level text generation in a general domain. These datasets are typically developed without using domain knowledge. In this paper, we release a novel dataset, VTechAGP, which is the first academic-to-general-audience text par…

2025

Visual Zero-Shot E-Commerce Product Attribute Value Extraction

NAACL 2025industry

Existing zero-shot product attribute value (aspect) extraction approaches in e-Commerce industry rely on uni-modal or multi-modal models, where the sellers are asked to provide detailed textual inputs (product descriptions) for the products. However, manually providing (typing) the product descripti…

2024

Bridging LiDAR Gaps: A Multi-LiDARs Domain Adaptation Dataset for 3D Semantic Segmentation

IJCAI 2024poster

We focus on the domain adaptation problem for 3D semantic segmentation, addressing the challenge of data variability in point clouds collected by different LiDARs. Existing benchmarks often mix different types of datasets, which blurs and complicates segmentation evaluations. Here, we introduce a Mu…

2024

Density-guided Translator Boosts Synthetic-to-Real Unsupervised Domain Adaptive Segmentation of 3D Point Clouds

CVPR 2024poster

3D synthetic-to-real unsupervised domain adaptive segmentation is crucial to annotating new domains. Self-training is a competitive approach for this task but its performance is limited by different sensor sampling patterns (i.e. variations in point density) and incomplete training strategies. In th…

2024

Design of A Rigid-soft Hybrid Robotic Glove with Force Sensing Function

ICRA 2024poster

Soft robotic gloves can not only provide timely, effective, safe and cheap rehabilitation training for patients with impaired movement function of hand, but also assist in completing daily grasping activities. However, most soft robotic gloves are completely composed of flexible structures. Although…

Cited by 1SourceScholar
2024

DiffLoc: Diffusion Model for Outdoor LiDAR Localization

CVPR 2024poster

Absolute pose regression (APR) estimates global pose in an end-to-end manner achieving impressive results in learn-based LiDAR localization. However compared to the top-performing methods reliant on 3D-3D correspondence matching APR's accuracy still has room for improvement. We recognize APR's lack…

2024

Efficient Personal Voice Activity Detection with Wake Word Reference Speech

ICASSP 2024accepted

Personal voice activity detection (PVAD) is gradually used in speech assistants. Traditional PVAD schemes extract the target speaker’s embedding from existing query reference speech through a pre-trained speaker verification model. Consequently, the performance of the PVAD model may suffer if the qu…

Cited by 0SourceScholar
2024

Joint Inference of Speaker Diarization and ASR with Multi-Stage Information Sharing

ICASSP 2024accepted

In this paper, we introduce a novel approach that unifies Automatic Speech Recognition (ASR) and speaker diarization in a cohesive framework. Utilizing the synergies between the two tasks, our method effectively extracts speaker-specific information from the lower layers of a pretrained Conformer-ba…

Cited by 0SourceScholar
2024

Learning Musical Representations for Music Performance Question Answering

EMNLP 2024finding

Music performances are representative scenarios for audio-visual modeling. Unlike common scenarios with sparse audio, music performances continuously involve dense audio signals throughout. While existing multimodal learning methods on the audio-video QA demonstrate impressive capabilities on genera…

2024

Optimal Containment Control of Multiple Quadrotors via Reinforcement Learning*

ICRA 2024poster

This paper explores the optimal containment control problem for nonlinear and underactuated quadrotors with multiple team leaders governed by nonlinear dynamics, employing the reinforcement learning. A cascade controller is formulated, comprising a position control component to ensure containment ac…

Cited by 0SourceScholar
2024

Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual Conformer

ICASSP 2024accepted

In recent years, neural network-based Wake Word Spotting achieves good performance on clean audio samples but struggles in noisy environments. Audio-Visual Wake Word Spotting (AVWWS) receives lots of attention because visual lip movement information is not affected by complex acoustic scenes. Previo…

Cited by 0SourceScholar
2024

Viiat-Hand: A Reach-and-Grasp Restoration System Integrating Voice Interaction, Computer Vision, Auditory and Tactile Feedback for Non-Sighted Amputees

RA-L 2024

For non-sighted and visually impaired (BVI) amputees, the combined loss of vision and grasping abilities turns the seemingly simple task of reaching and grasping into a significant challenge. This letter introduces a novel multi-sensory prosthesis system designed for BVI amputees to assist in percep

Cited by 9SourceScholar
2024

Voxblink: A Large Scale Speaker Verification Dataset on Camera

ICASSP 2024accepted

In this paper, we introduce a large-scale and high-quality audiovisual speaker verification dataset, named VoxBlink. We propose an innovative and robust automatic audio-visual data mining pipeline to curate this dataset, which contains 1.45M utterances from 38K speakers. Due to the inherent nature o…

Cited by 0SourceScholar
2023

DiffuseStyleGesture: Stylized Audio-Driven Co-Speech Gesture Generation with Diffusion Models

IJCAI 2023poster

The art of communication beyond speech there are gestures. The automatic co-speech gesture generation draws much attention in computer animation. It is a challenging task due to the diversity of gestures and the difficulty of matching the rhythm and semantics of the gesture to the corresponding spee…

2023

Target-Speaker Voice Activity Detection Via Sequence-to-Sequence Prediction

ICASSP 2023accepted

Target-speaker voice activity detection is currently a promising approach for speaker diarization in complex acoustic environments. This paper presents a novel Sequence-to-Sequence Target-Speaker Voice Activity Detection (Seq2Seq-TSVAD) method that can efficiently address the joint modeling of large…

Cited by 0SourceScholar
2023

The DKU Post-Challenge Audio-Visual Wake Word Spotting System for the 2021 MISP Challenge: Deep Analysis

ICASSP 2023accepted

This paper further explores our previous wake word spotting system ranked 2-nd in Track 1 of the MISP Challenge 2021. First, we investigate a robust unimodal approach based on 3D and 2D convolution and adopt the simple attention module (SimAM) for our system to improve performance. Second, we explor…

Cited by 0SourceScholar
2023

The WHU-Alibaba Audio-Visual Speaker Diarization System for the MISP 2022 Challenge

ICASSP 2023accepted

This paper describes the system developed by the WHU-Alibaba team for the Multimodal Information Based Speech Processing (MISP) 2022 Challenge. We extend the Sequence-to-Sequence Target-Speaker Voice Activity Detection framework to simultaneously detect multiple speakers’ voice activities from audio…

Cited by 0SourceScholar
2022

Multi-Graph Fusion Networks for Urban Region Embedding

IJCAI 2022poster

Learning the embeddings for urban regions from human mobility data can reveal the functionality of regions, and then enables the correlated but distinct tasks such as crime prediction. Human mobility data contains rich but abundant information, which yields to the comprehensive region embeddings for…

2022

The DKU Audio-Visual Wake Word Spotting System for the 2021 MISP Challenge

ICASSP 2022accepted

This paper describes the system developed by the DKU team for the MISP Challenge 2021. We present a two-stage approach consisting of end-to-end neural networks for the audio-visual wake word spotting task. We first process audio and video data to give them a similar structure and then train two unim…

Cited by 0SourceScholar
2021

A Model-Free Synchronous Control of Humanoid Robot Finger

ICRA 2021poster

For a multi-fingered robot hand, the individual control over single joints cannot guarantee their fine collaboration. For achieving a high-precision synchronization, a theory of synchronous control is introduced to multi-fingered robot hands. This paper introduced a new model-free and cross-coupling…

Cited by 2SourceScholar
2021

A Triplet Appearance Parsing Network for Person Re-Identification

ICASSP 2021accepted

As one of the specific vision tasks, person re-identification has become a prevalent research topic in the field of multimedia and computer vision. However, existing feature extraction methods, originating from the quality of the bounding boxes which could cause the inhomogeneity and incoherence of…

Cited by 0SourceScholar
2019

RF-Net: An End-To-End Image Matching Network Based on Receptive Field

CVPR 2019poster

This paper proposes a new end-to-end trainable matching network based on receptive field, RF-Net, to compute sparse correspondence between images. Building end-to-end trainable matching framework is desirable and challenging. The very recent approach, LF-Net, successfully embeds the entire feature e…

Cited by 128PDFScholar