← Search

Jingdong Chen

98 accepted papers

2026

ACTIVE-o3 : Empowering MLLMs with Active Perception via Pure Reinforcement Learning

ICML 2026poster

Active vision, also known as active perception, refers to actively selecting where and how to look in order to gather task-relevant information. It is a critical component of efficient perception and decision-making in humans and advanced embodied agents. With the rise of Multimodal Large Language M…

Cited by 0SourceScholar
2026

FORWARD CONVOLUTIVE PREDICTION FOR FRAME ONLINE MONAURAL SPEECH DEREVERBERATION BASED ON KRONECKER PRODUCT DECOMPOSITION

ICASSP 2026poster

Dereverberation has long been a crucial research topic in speech processing, aiming to alleviate the adverse effects of reverberation in voice communication and speech interaction systems. Among existing approaches, forward convolutional prediction (FCP) has recently attracted attention. It typicall…

Cited by 0SourcePDFScholar
2026

HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses Through Reasoning MLLMs

AAAI 2026technical

While Multimodal Large Language Models (MLLMs) show immense promise for achieving truly human-like interactions, progress is hindered by the lack of fine-grained evaluation frameworks for human-centered scenarios, encompassing both the understanding of complex human intentions and the provision of e

Cited by 0SourcePDFScholar
2026

ONLINE NEURAL FUSION OF DISTORTIONLESS DIFFERENTIAL BEAMFORMERS FOR ROBUST SPEECH ENHANCEMENT

ICASSP 2026poster

Fixed beamforming is widely used in practice since it does not depend on the estimation of noise statistics and provides relatively stable performance. However, a single beamformer cannot adapt to varying acoustic conditions, which limits its interference suppression capability. To address this, ada…

Cited by 0SourcePDFScholar
2026

ROBUST ONLINE OVERDETERMINED INDEPENDENT VECTOR ANALYSIS BASED ON BILINEAR DECOMPOSITION

ICASSP 2026oral

Online blind source separation is essential for both speech communication and human-machine interaction. Among existing approaches, overdetermined independent vector analysis (OverIVA) delivers strong performance by exploiting the statistical independence of source signals and the orthogonality betw…

Cited by 0SourcePDFScholar
2026

SCAN: Self-Calibrated AutoregressioN for High-Quality Visual Generation

AAAI 2026technical

Human artists can continuously refine their coarse sketches during artistic creation. This is quite different from existing autoregressive generation, where a token is determined once sampled. Aiming to flexibly refine the generated contents, this paper presents a Self-Calibrated AutoregressioN (SCA

Cited by 0SourcePDFScholar
2026

Selecting Samples on Graphs: A Unified Dataset Pruning Framework for Lossless Training Acceleration

ICML 2026poster

The rapid growth of modern training datasets has significantly increased computational cost, motivating dataset pruning(DP) methods which retain only a subset of informative samples to reduce training cost. Existing pruning criteria typically rely on either intrinsic signals that assess samples inde…

Cited by 0SourceScholar
2026

SkySense-VITA: Towards Universal In-context Segmentation of Multi-modal Remote Sensing Imagery

CVPR 2026

While recent foundation models for remote sensing segmentation have shown notable progress, they still fall short in processing diverse multi-modal inputs, synergizing complementary prompt types, and leveraging semantic hierarchies. To address these limitations, we introduce SkySense-VITA, a unified

Cited by 0SourceScholar
2026

TriC-Motion: Tri-Domain Causal Modeling Grounded Text-to-Motion Generation

ICLR 2026poster

Text-to-motion generation, a rapidly evolving field in computer vision, aims to produce realistic and text-aligned motion sequences. Current methods primarily focus on spatial-temporal modeling or independent frequency domain analysis, lacking a unified framework for joint optimization across spatia…

Cited by 0SourcecodeScholar
2026

UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception

AAAI 2026technical

The remarkable success of diffusion models in text-to-image generation has sparked growing interest in expanding their capabilities to a variety of multi-modal tasks, including image understanding, manipulation, and perception. These tasks require advanced semantic comprehension across both visual a

Cited by 0SourcePDFScholar
2026

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation

CVPR 2026

The rapid advancement of AIGC-based video generation has underscored the critical need for comprehensive evaluation frameworks that go beyond traditional generation quality metrics to encompass aesthetic appeal. However, existing benchmarks remain largely focused on technical fidelity, leaving a sig

Cited by 0SourceScholar
2026

VGA-BenchV2: An Expanded Unified Benchmark and Multi-Model Framework for Evaluating Video Aesthetics and Generation Quality

IJCAI 2026

The rapid advancement of AIGC video generation calls for evaluation frameworks that move beyond technical fidelity and incorporate human-centered aesthetic assessment. Existing benchmarks often overlook fine-grained perceptual qualities such as visual aesthetics, artistic style, and human preference

Cited by 0Scholar
2025

ARGenSeg: Image Segmentation with Autoregressive Image Generation Model

NeurIPS 2025poster

We propose a novel AutoRegressive Generation-based paradigm for image Segmentation (ARGenSeg), achieving multimodal understanding and pixel-level perception within a unified framework. Prior works integrating image segmentation into multimodal large language models (MLLMs) typically employ either b…

Cited by 0SourceScholar
2025

Advances in Microphone Array Processing and Multichannel Speech Enhancement

ICASSP 2025accepted

This paper reviews pioneering works in microphone array processing and multichannel speech enhancement, highlighting historical achievements, technological evolution, commercialization aspects, and key challenges. It provides valuable insights into the progression and future direction of these areas…

Cited by 0SourceScholar
2025

Animate-X: Universal Character Image Animation with Enhanced Motion Representation

ICLR 2025poster

Character image animation, which generates high-quality videos from a reference image and target pose sequence, has seen significant progress in recent years. However, most existing methods only apply to human figures, which usually do not generalize well on anthropomorphic characters commonly used…

Cited by 14SourcePDFScholar
2025

CasP: Improving Semi-Dense Feature Matching Pipeline Leveraging Cascaded Correspondence Priors for Guidance

ICCV 2025poster

Semi-dense feature matching methods have shown strong performance in challenging scenarios. However, the existing pipeline relies on a global search across the entire feature map to establish coarse matches, limiting further improvements in accuracy and efficiency. Motivated by this limitation, we p…

2025

DOA Estimation Based on Enhanced SRP-MVDR Using Kronecker Product Decomposition for Large Rectangular Microphone Arrays

ICASSP 2025accepted

Direction-of-arrival (DOA) estimation is a key process in microphone array systems. The steered response power-based minimum variance distortionless response (SRP-MVDR) method performs very well in challenging acoustic environments but suffers from exponential complexity as the number of microphones…

Cited by 0SourceScholar
2025

Data-Driven White Noise Gain Constrained Robust Superdirective Beamformer for Speech Enhancement

ICASSP 2025accepted

Superdirective beamformers are highly effective at suppressing directional interference and diffuse noise, but their practical use is often constrained by the problem of white noise amplification. Robust superdirective beamforming methods typically address this by imposing a constraint on the white…

Cited by 0SourceScholar
2025

Design and Optimization of Superdirective Beamforming and Post-Filtering for Speech Enhancement

ICASSP 2025accepted

Superdirective beamformers, used with small microphone arrays, are highly attractive due to their high directivity and frequency-invariant beampatterns, making them well-suited for processing broadband acoustic and speech signals. However, these beamformers are very sensitive to array imperfections…

Cited by 0SourceScholar
2025

Design of Robust Differential Beamformers with Microphone Arrays of Arbitrary Planar Geometry

ICASSP 2025accepted

Differential microphone arrays (DMAs) have garnered significant attention in recent research and development due to their high directivity and frequency-invariant beampatterns. However, DMAs frequently encounter substantial white noise amplification, which limits their practical applications. This p…

Cited by 0SourceScholar
2025

Exploring Spectral Signatures of Chinese liquor using Machine Learning and SHapley Additive exPlanations

ICASSP 2025accepted

Chinese liquor holds great cultural and economic significance globally. The accurate classification of aroma types and alcohol content is crucial for quality control in Chinese liquor production. To address limitations such as subjectivity and sensor drift in current methods, this study introduces a…

Cited by 0SourceScholar
2025

HomoMatcher: Achieving Dense Feature Matching with Semi-Dense Efficiency by Homography Estimation

AAAI 2025technical

Feature matching between image pairs is a fundamental problem in computer vision that drives many applications, such as SLAM. Recently, semi-dense matching approaches have achieved substantial performance enhancements and established a widely-accepted coarse-to-fine paradigm. However, the majority…

Cited by 0SourcePDFScholar
2025

Microphone Array Beamforming for Speech Enhancement Based on Dynamic Mode Decomposition

ICASSP 2025accepted

Microphone array beamforming is widely used to extract desired speech signals from noisy environments. While most research in this area focuses on utilizing spatial information, less attention is given to the intrinsic physical mechanisms underlying microphone array observations. This paper aims to…

Cited by 4SourceScholar
2025

Mimir: Improving Video Diffusion Models for Precise Text Understanding

CVPR 2025poster

Text serves as the key control signal in video generation due to its narrative nature. To render text descriptions into video clips, current video diffusion models borrow features from text encoders yet struggle with limited text comprehension. The recent success of large language models (LLMs) show…

Cited by 4SourcePDFScholar
2025

MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation

CVPR 2025poster

The image-to-video (I2V) generation is conditioned on the static image, which has been enhanced recently by the motion intensity as an additional control signal. These motion-aware models are appealing to generate diverse motion patterns, yet there lacks a reliable motion estimator for training such…

Cited by 4SourcePDFScholar
2025

On the Design of Low-Rank Differential Beamformers with Nonuniform Linear Microphone Arrays

ICASSP 2025accepted

Kronecker product beamforming is an effective technique for designing beamformers with nonuniform linear arrays (NULAs). However, current techniques are restricted to NULAs with specific configurations, where the steering vector of the array is represented as a Kronecker product of steering vectors…

Cited by 0SourceScholar
2025

Radiation and Directivity Analysis of a Vibrating Dome-Shaped Radiator Mounted on an Infinite Baffle

ICASSP 2025accepted

Accurate modeling and analysis of a radiator mounted on an infinite baffle are crucial for understanding its acoustic radiation characteristics. This paper investigates the radiation behavior of a convex dome-shaped radiator in such a condition, showing that, under the far-field approximation, the p…

Cited by 0SourceScholar
2025

Reversing Flow for Image Restoration

CVPR 2025poster

Image restoration aims to recover high-quality (HQ) images from degraded low-quality (LQ) ones by reversing the effects of degradation. Existing generative models for image restoration, including diffusion and score-based models, often treat the degradation process as a stochastic transformation, wh…

Cited by 0SourcePDFScholar
2025

SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing

ICCV 2025poster

The multi-modal remote sensing foundation model (MM-RSFM) has significantly advanced various Earth observation tasks, such as urban planning, environmental monitoring, and natural disaster management. However, most existing approaches generally require the training of separate backbone networks for…

Cited by 0SourcePDFScholar
2025

SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language Modeling

CVPR 2025poster

Open-world interpretation aims to accurately localize and recognize all objects within images by vision-language models (VLMs). While substantial progress has been made in this task for natural images, the advancements for remote sensing (RS) images still remain limited, primarily due to these two c…

2025

VADB: A Large-Scale Video Aesthetic Database with Professional and Multi-Dimensional Annotations

NeurIPS 2025poster

Video aesthetic assessment, a vital area in multimedia computing, integrates computer vision with human cognition. Its progress is limited by the lack of standardized datasets and robust models, as the temporal dynamics of video and multimodal fusion challenges hinder direct application of image-bas…

Cited by 0SourcecodeScholar
2025

VideoMAR: Autoregressive Video Generation with Continuous Tokens

NeurIPS 2025poster

Masked-based autoregressive models have demonstrated promising image generation capability in continuous space. However, their potential for video generation remains under-explored. Masked-based autoregressive models have demonstrated promising image generation capability in continuous space. Howev…

Cited by 0SourceScholar
2025

When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning

ICCV 2025poster

Efficient vision-language understanding of large Remote Sensing Images (RSIs) is meaningful but challenging. Current Large Vision-Language Models (LVLMs) typically employ limited pre-defined grids to process images, leading to information loss when handling gigapixel RSIs. Conversely, using unlimite…

2024

A Computationally Efficient Semi-Blind Source Separation Approach for Nonlinear Echo Cancellation Based on an Element-Wise Iterative Source Steering

ICASSP 2024accepted

While the semi-blind source separation-based acoustic echo cancellation (SBSS-AEC) has received much research attention due to its promising performance during double-talk compared to the traditional adaptive algorithms, it suffers from system latency and nonlinear distortions. To circumvent these d…

Cited by 0SourceScholar
2024

A Steered Response Power Approach with Bilinear Prediction-Based Trade-Off Prewhitening for Speaker Localization

ICASSP 2024accepted

This paper studies the problem of acoustic source localization in room environments. It presents an improved steered response power (SRP) approach with low-complexity and trade-off prewhitening. This method consists of two steps. In the first one, the linear predictor that is used to model the speec…

Cited by 0SourceScholar
2024

Accelerating Pre-training of Multimodal LLMs via Chain-of-Sight

NeurIPS 2024poster

This paper introduces Chain-of-Sight, a vision-language bridge module that accelerates the pre-training of Multimodal Large Language Models (MLLMs). Our approach employs a sequence of visual resamplers that capture visual details at various spacial scales. This architecture not only leverages globa…

Cited by 3SourcePDFScholar
2024

Beamforming Through Online Convex Combination of Differential Beamformers

ICASSP 2024accepted

Thanks to their high directivity, compact size, and reliable performance, differential microphone arrays (DMAs) have attracted great interest from both industry and academia as they have demonstrated great potential to be used in a wide range of applications for high-fidelity speech acquisition. Nev…

Cited by 0SourceScholar
2024

Differential Beamforming with Null Constraints for Spherical Microphone Arrays

ICASSP 2024accepted

Differential microphone arrays (DMAs) can measure both the acoustic pressure field and the differential acoustic pressure fields, which gives them great advantages in a wide range of applications for acoustic and speech signal acquisition. The core component of DMAs is the so-called differential bea…

Cited by 0SourceScholar
2024

Directional Gain Based Noise Covariance Matrix Estimation for MVDR Beamforming

ICASSP 2024accepted

This paper is devoted to the problem of noise covariance matrix (NCM) estimation. It proposes a time-frequency masking based approach. We first present an optimal mask function based on the mean-squared error criterion. To estimate this mask, we employ the recently developed directional gain method…

Cited by 0SourceScholar
2024

EcoMatcher: Efficient Clustering Oriented Matcher for Detector-free Image Matching

ECCV 2024poster

"Detector-free local feature matching methods have demonstrated significant performance improvements since leveraging the power of Transformer architecture. The global receptive field allows for simultaneous interaction among all elements, proving particularly beneficial in regions with low texture…

Cited by 1SourcePDFScholar
2024

Learning Dynamic Tetrahedra for High-Quality Talking Head Synthesis

CVPR 2024poster

Recent works in implicit representations such as Neural Radiance Fields (NeRF) have advanced the generation of realistic and animatable head avatars from video sequences. These implicit methods are still confronted by visual artifacts and jitters since the lack of explicit geometric constraints pose…

2024

LogicMP: A Neuro-symbolic Approach for Encoding First-order Logic Constraints

ICLR 2024poster

Integrating first-order logic constraints (FOLCs) with neural networks is a crucial but challenging problem since it involves modeling intricate correlations to satisfy the constraints. This paper proposes a novel neural layer, LogicMP, which performs mean-field variational inference over a Markov L…

2024

On the Design of Planar Differential Microphone Arrays with Specified Beamwidth or Sidelobe Level

ICASSP 2024accepted

This paper investigates the problem of designing differential beam-formers with planar microphone arrays to achieve not only the desired target directivity pattern but also control the beamwidth (BW) or sidelobe level (SLL). We first discuss the target directivity patterns and express the Dolph-Cheb…

Cited by 0SourceScholar
2024

POA: Pre-training Once for Models of All Sizes

ECCV 2024poster

"Large-scale self-supervised pre-training has paved the way for one foundation model to handle many different vision tasks. Most pre-training methodologies train a single model of a certain size at one time. Nevertheless, various computation or storage constraints in real-world scenarios require sub…

2024

SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery

CVPR 2024poster

Prior studies on Remote Sensing Foundation Model (RSFM) reveal immense potential towards a generic model for Earth Observation. Nevertheless these works primarily focus on a single modality without temporal and geo-context modeling hampering their capabilities for diverse tasks. In this study we pre…

Cited by 140SourcePDFScholar
2024

Stereophonic Music Source Separation with Spatially-Informed Bridging Band-Split Network

ICASSP 2024accepted

Stereophonic music source separation (MSS) is a problem of extracting individual source tracks, e.g. bass, drums, vocals, from a stereo music recording. Deep neural network (DNN) based MSS systems have demonstrated great promise though spatial panning cues and time-frequency spectral structures in s…

Cited by 0SourceScholar
2024

StyleTokenizer: Defining Image Style by a Single Instance for Controlling Diffusion Models

ECCV 2024poster

"Despite the burst of innovative methods for controlling the diffusion process, effectively controlling image styles in text-to-image generation remains a challenging task. Many adapter-based methods impose image representation conditions on the denoising process to accomplish image control. However…

2024

The Multimodal Information Based Speech Processing (MISP) 2023 Challenge: Audio-Visual Target Speaker Extraction

ICASSP 2024accepted

Previous Multimodal Information based Speech Processing (MISP) challenges mainly focused on audio-visual speech recognition (AVSR) with commendable success. However, the most advanced back-end recognition systems often hit performance limits due to the complex acoustic environments. This has prompte…

Cited by 0SourceScholar
2024

Towards Better Vision-Inspired Vision-Language Models

CVPR 2024poster

Vision-language (VL) models have achieved unprecedented success recently in which the connection module is the key to bridge the modality gap. Nevertheless the abundant visual clues are not sufficiently exploited in most existing methods. On the vision side most existing approaches only use the last…

Cited by 2SourcePDFScholar
2023

A Frequency-Domain Recursive Least-Squares Adaptive Filtering Algorithm Based On A Kronecker Product Decomposition

ICASSP 2023accepted

This paper proposes a frequency-domain recursive least-squares (RLS) adaptive filtering algorithm for identifying time-varying acoustic systems in noisy environments. The Kronecker product (KP) is employed to decompose the model filter of the acoustic channel impulse response into two sets of short…

Cited by 0SourceScholar
2023

On Multiple-Input/Binaural-Output Antiphasic Speaker Signal Extraction

ICASSP 2023accepted

This paper studies the problem of target speaker signal exaction and antiphasic rendering with an array of microphones in the scenarios where there are two active speakers. Based on the important findings achieved in the psychoacoustic field as well as our recent works on single-channel speech enhan…

Cited by 0SourceScholar
2023

Simultaneously Short- and Long-Term Temporal Modeling for Semi-Supervised Video Semantic Segmentation

CVPR 2023poster

In order to tackle video semantic segmentation task at a lower cost, e.g., only one frame annotated per video, lots of efforts have been devoted to investigate the utilization of those unlabeled frames by either assigning pseudo labels or performing feature enhancement. In this work, we propose a no…

Cited by 12SourcePDFScholar
2023

Spatially Informed Independent vector analysis for Source Extraction based on the convolutive Transfer Function Model

ICASSP 2023accepted

Spatial information can help improve source separation performance. Numerous spatially informed source extraction methods based on the independent vector analysis (IVA) have been developed, which can achieve reasonably good performance in non- or weakly reverberant environments. However, the perform…

Cited by 0SourceScholar
2023

Summary on the Multimodal Information Based Speech Processing (MISP) 2022 Challenge

ICASSP 2023accepted

The Multimodal Information based Speech Processing (MISP) 2022 challenge aimed to enhance speech processing performance in harsh acoustic environments by leveraging additional modalities such as video or text. The challenge included two tracks: audio-visual speaker diarization (AVSD) and audio-visua…

Cited by 0SourceScholar
2023

Switching Kronecker Product Linear Filtering for Multispeaker Adaptive Speech Dereverberation

ICASSP 2023accepted

Dereverberation, a process to mitigate or eliminate the reverberation effect, plays an important role in hands-free speech communication and human-machine interfaces. Tremendous efforts have been devoted to this problem and various methods have been developed over the last three decades. Those metho…

Cited by 0SourceScholar
2023

The Multimodal Information Based Speech Processing (Misp) 2022 Challenge: Audio-Visual Diarization And Recognition

ICASSP 2023accepted

The Multi-modal Information based Speech Processing (MISP) challenge aims to extend the application of signal processing technology in specific scenarios by promoting the research into wake-up words, speaker diarization, speech recognition, and other technologies. The MISP2022 challenge has two trac…

Cited by 0SourceScholar
2023

Uncertainty-guided Learning for Improving Image Manipulation Detection

ICCV 2023poster

Image manipulation detection (IMD) is of vital importance as faking images and spreading misinformation can be malicious and harm our daily life. IMD is the core technique to solve these issues and poses challenges in two main aspects: (1) Data Uncertainty, i.e., the manipulated artifacts are often…

Cited by 18PDFcodeScholar
2022

CMUA-Watermark: A Cross-Model Universal Adversarial Watermark for Combating Deepfakes

AAAI 2022technical

Malicious applications of deepfakes (i.e., technologies generating target facial attributes or entire faces from facial images) have posed a huge threat to individuals' reputation and security. To mitigate these threats, recent studies have proposed adversarial watermarks to combat deepfake models,…

2022

Hierarchical Memory Learning for Fine-Grained Scene Graph Generation

ECCV 2022poster

"Regarding Scene Graph Generation (SGG), coarse and fine predicates mix in the dataset due to the crowd-sourced labeling, and the long-tail problem is also pronounced. Given this tricky situation, many existing SGG methods treat the predicates equally and learn the model under the supervision of mix…

Cited by 31SourcePDFScholar
2022

Robust Pressure Matching with ATF Perturbation Constraints for Sound Field Control

ICASSP 2022accepted

Sound field control systems deployed in room acoustic environments require knowing the acoustic channel impulse responses between the loudspeakers and matching microphones, which are challenging to estimate accurately due to perturbations caused by such factors as temperature changes and sensors’ po…

Cited by 0SourceScholar
2022

SimAN: Exploring Self-Supervised Representation Learning of Scene Text via Similarity-Aware Normalization

CVPR 2022poster

Recently self-supervised representation learning has drawn considerable attention from the scene text recognition community. Different from previous studies using contrastive learning, we tackle the issue from an alternative perspective, i.e., by formulating the representation learning scheme in a g…

Cited by 37PDFcodeScholar
2022

Study of the Null Directions on The Performance of Differential Beamformers

ICASSP 2022accepted

Null directions are important parameters for differential beamformers, which play an important role on the beamforming performance. In this paper, we investigate the performance of differential beamformers as a function of the null directions. We first derive the directivity factor (DF) as an explic…

Cited by 0SourceScholar
2022

The First Multimodal Information Based Speech Processing (Misp) Challenge: Data, Tasks, Baselines And Results

ICASSP 2022accepted

In this paper we discuss the rational of the Multi-model Information based Speech Processing (MISP) Challenge, and provide a detailed description of the data recorded, the two evaluation tasks and the corresponding baselines, followed by a summary of submitted systems and evaluation results. The MIS…

Cited by 0SourceScholar
2022

Training Object Detectors From Scratch: An Empirical Study in the Era of Vision Transformer

CVPR 2022poster

Modeling in computer vision has long been dominated by convolutional neural networks (CNNs). Recently, in light of the excellent performances of self-attention mechanism in the language field, transformers tailored for visual data have drawn numerous attention and triumphed CNNs in various vision ta…

Cited by 16PDFScholar
2021

A Conceptual Approach of Passive Human-Intention-Orientated Variable Admittance Control using Power Envelope

IROS 2021poster

Two main challenges that need to be addressed in physical human-robot interaction (pHRI) are efficient recognition of human intention and interaction safety. In this paper, a general human intention framework was summarized, firstly, according to the robot's roles: a passive follower and a compliant…

Cited by 4SourceScholar
2021

A Simplified Wiener Beamformer Based on Covariance Matrix Modelling

ICASSP 2021accepted

This paper is devoted to the problem of adaptive beamforming with small-spaced microphone arrays. In this context, the Wiener filter is an optimal beamformer in the mean-squared error (MSE) sense. However, it requires good estimates of the covariance matrices of the speech signal of interest and noi…

Cited by 0SourceScholar
2021

Combined Differential Beamforming With Uniform Linear Microphone Arrays

ICASSP 2021accepted

While differential beamformers have been widely used in voice communication and human-machine speech interface systems to enhance speech signals of interest, how to design such beamformers that on the one hand can achieve the highest possible directivity factor (DF) and on the other hand are able to…

Cited by 0SourceScholar
2021

LPSNet: A Lightweight Solution for Fast Panoptic Segmentation

CVPR 2021poster

Panoptic segmentation is a challenging task aiming to simultaneously segment objects (things) at instance level and background contents (stuff) at semantic level. Existing methods mostly utilize two-stage detection network to attain instance segmentation results, and fully convolutional network to p…

Cited by 45PDFScholar
2021

MatchVIE: Exploiting Match Relevancy between Entities for Visual Information Extraction

IJCAI 2021poster

Visual Information Extraction (VIE) task aims to extract key information from multifarious document images (e.g., invoices and purchase receipts). Most previous methods treat the VIE task simply as a sequence labeling problem or classification problem, which requires models to carefully identify eac…

Cited by 33SourcePDFScholar
2021

On the Design of Square Differential Microphone Arrays with a Multistage Structure

ICASSP 2021accepted

This paper studies the problem of designing square differential microphone arrays (SDMAs). It presents a multistage approach, which first divides an SDMA composed of M <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> microphones into (M − 1) <sup…

Cited by 0SourceScholar
2021

Robust Recursive Least M-Estimate Adaptive Filter for the Identification of Low-Rank Acoustic Systems

ICASSP 2021accepted

To identify acoustic systems (which are low-rank in nature) in non-Gaussian and Gaussian noise, a robust recursive least M-estimate adaptive filtering algorithm is developed in this paper by applying the nearest Kronecker product to decompose the acoustic impulse response. Two M-estimators, i.e., th…

Cited by 0SourceScholar
2021

Robust Steerable Differential Beamformers with Null Constraints for Concentric Circular Microphone Arrays

ICASSP 2021accepted

Differential beamformers with concentric circular microphone arrays (CCMAs) are desirable for use in various applications since they can form frequency-invariant spatial responses, have better beam steering flexibility than linear arrays, and suffer less with beampattern irregularity and white noise…

Cited by 0SourceScholar
2020

An Improved Solution to the Frequency-Invariant Beamforming with Concentric Circular Microphone Arrays

ICASSP 2020accepted

Frequency-invariant beamforming with circular microphone arrays (CMAs) has drawn a significant amount of attention for its steering flexibility and high directivity. However, frequency-invariant beam-forming with CMAs often suffers from the so-called null problem, which is caused by the zeros of the…

Cited by 0SourceScholar
2020

Partial AUC Optimization Based Deep Speaker Embeddings with Class-Center Learning for Text-Independent Speaker Verification

ICASSP 2020accepted

Deep embedding based text-independent speaker verification has demonstrated superior performance to traditional methods in many challenging scenarios. Its loss functions can be generally categorized into two classes, i.e., verification and identification. The verification loss functions match the pi…

Cited by 0SourceScholar
2020

Proximal Multitask Learning Over Distributed Networks with Jointly Sparse Structure

ICASSP 2020accepted

Modeling relations between local optimum parameter vectors in multitask networks has attracted much attention over the last years. This work considers a distributed optimization problem for parameter vectors with a jointly sparse structure among nodes, that is, the parameter vectors share the same s…

Cited by 0SourceScholar
2020

Robust Frequency-Domain Recursive Least M-Estimate Adaptive Filter For Acoustic System Identification

ICASSP 2020accepted

To identify acoustic systems in non-Gaussian and Gaussian noises, a robust frequency-domain recursive least M-estimate (FRLM) adaptive filtering algorithm is proposed. The cost function of the adaptive filter is defined by using a robust time-domain M-estimator, while its update equation is derived…

Cited by 0SourceScholar
2020

Robust and steerable kronecker product differential beamforming With rectangular microphone arrays

ICASSP 2020accepted

Differential microphone arrays (DMAs), a class of welldesigned small-size arrays combined with differential beamforming, are very useful for processing broadband acoustic, audio, and speech signals in a wide range of applications. However, most efforts in the literature so far have been devoted to l…

Cited by 0SourceScholar
2019

AUC Optimization for Deep Learning Based Voice Activity Detection

ICASSP 2019accepted

Voice activity detection (VAD) based on deep neural networks (DNN) has demonstrated good performance in adverse acoustic environments. Current DNN based VAD optimizes a surrogate function, e.g. minimum cross-entropy or minimum squared error, at a given decision threshold. However, VAD usually works…

Cited by 0SourceScholar
2019

Design of Optimal Linear Differential Microphone Arrays Based Array Geometry Optimization

ICASSP 2019accepted

This paper presents a method to design optimal linear differential microphone arrays (DMAs) by optimizing the array geometry. By constraining the DMA beamformer to achieve a given target value of the directivity factor (DF) with a specified target frequency-invariant beampattern while achieving also…

Cited by 0SourceScholar
2019

On the Design of Flexible Kronecker Product Beamformers with Linear Microphone Arrays

ICASSP 2019accepted

This paper proposes a method for the design of flexible Kronecker product beamformers based on the decomposition of the steering vector of a physical array as a Kronecker product of steering vectors of two smaller virtual arrays. With this decomposition, the global beamforming filter is designed by…

Cited by 0SourceScholar
2019

Properties and Limits of the Minimum-norm Differential Beamformers with Circular Microphone Arrays

ICASSP 2019accepted

Small aperture circular microphone arrays (CMAs) have been widely used in many applications such as teleconferencing, smartspeakers, and robotics. A critical component of such arrays is the differential beamformer, which can achieve relatively high spatial gains with the same beampatterns at most fr…

Cited by 0SourceScholar
2018

A Single-Channel Noise Reduction Filtering/Smoothing Technique in the Time Domain

ICASSP 2018accepted

In this paper, we present a single-channel smoothing-and-filtering technique for noise reduction in the time domain. Unlike traditional noise reduction methods, which directly apply a noise reduction filter to the noisy signal, the developed technique achieves noise reduction in two steps. It first…

Cited by 0SourceScholar
2018

On Speech Enhancement Using Microphone Arrays in the Presence of Co-Directional Interference

ICASSP 2018accepted

Beamforming using microphone arrays has been widely used for enhancing speech signals of interest and suppressing noise and interference in a wide range of applications. In order to make it work, beamforming generally assumes that the speech source of interest and the interference source are inciden…

Cited by 5SourceScholar
2018

On the Design of Robust Steerable Frequency-Invariant Beampatterns with Concentric Circular Microphone Arrays

ICASSP 2018accepted

This paper studies the problem of frequency-invariant beamforming with concentric circular microphone arrays (CCMAs). We develop a beamforming algorithm based on an optimal approximation of the beamformer's beampattern with the Jacobi-Anger expansion. In comparison with the existing frequency-invari…

Cited by 0SourceScholar
2017

A minimum variance partially distortionless response filter for single-channel noise reduction

ICASSP 2017accepted

This paper deals with the problem of single-channel noise reduction. Thanks to the eigenvalue decomposition, we arrange the eigenvalues of the speech correlation matrix in such a way that all the spectral mode signal-to-noise ratios (SNRs) of the noisy speech are ordered in a descending manner. By m…

Cited by 0SourceScholar
2017

Robust multichannel TDOA estimation for speaker localization using the impulsive characteristics of speech spectrum

ICASSP 2017accepted

Time delay estimation (TDE) plays an important role in localizing and tracking radiating acoustic sources. Although many efforts have been devoted to this problem in the literature, the robustness of TDE with respect to noise and reverberation remains a great challenge for practical systems. In this…

Cited by 0SourceScholar
2017

Study of the frequency-domain multichannel noise reduction problem with the householder transformation

ICASSP 2017accepted

This paper presents an approach to the multichannel noise reduction problem. It first transforms the multichannel noisy speech signals into the frequency domain. A Householder transformation is then constructed, which converts the multichannel coefficients in each frequency bin into two components:…

Cited by 0SourceScholar
2016

A single-channel noise cancelation filter in the short-time-fourier-transform domain

ICASSP 2016accepted

This paper develops a single-channel noise cancelation filter in the short-time Fourier transform (STFT) domain by combining the subspace method and the optimal filtering technique via joint diagonalization of the desired clean speech and noise signal correlation matrices. This filter is shown to be…

Cited by 0SourceScholar
2016

Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin

ICML 2016poster

We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech–two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning allows us to handle a diverse variety of s…

2016

On time delay estimation based on multichannel spatiotemporal sparse linear prediction

ICASSP 2016accepted

Noise and reverberation can significantly affect the performance of time delay estimation (TDE) in room acoustic environments. The multichannel cross-correlation coefficient (MCCC) algorithm, which extends the traditional cross-correlation method from two to multiple channels, can exploit the spatia…

Cited by 0SourceScholar
2016

Subspace superdirective beamformers based on joint diagonalization

ICASSP 2016accepted

Although they have been intensively studied and used in many applications due to their high directivity factor (DF), superdirective beamformers are sensitive to sensor noise and mismatch between sensors. This paper studies the problem of superdirective beamforming combined with the joint diagonaliza…

Cited by 0SourceScholar
2015

Investigation of a parametric gain approach to single-channel speech enhancement

ICASSP 2015accepted

This paper investigates a parametric gain approach to single-channel noise reduction in the frequency domain. In comparison with the traditional parametric Wiener gain, the major novelty of this presented approach is that the parametric gain is formulated to estimate the noise by using the mean-squa…

Cited by 0SourceScholar
2015

Optimal design of directivity patterns for endfire linear microphone arrays

ICASSP 2015accepted

Directivity pattern or beampattern is an important performance measure in all fixed beamformers. Given a microphone array, how to design the beamforming filter so that the resulting directivity pattern is close to the desired one is a critical issue. In this paper, we study the design of such patter…

Cited by 0SourceScholar
2015

Optimal single-channel noise reduction filtering matrices from the pearson correlation coefficient perspective

ICASSP 2015accepted

This paper studies the problem of single-channel noise reduction in the time domain, where an estimate of a vector of the desired clean speech is achieved by filtering a frame of the noisy signal with a rectangular filtering matrix. The core issue with this problem formulation is then the estimation…

Cited by 0SourceScholar