← Search

Qing Wang

91 accepted papers

2026

A UNIFIED SPOKEN LANGUAGE MODEL WITH INJECTED EMOTIONAL-ATTRIBUTION THINKING FOR HUMAN-LIKE INTERACTION

ICASSP 2026poster

This paper presents a unified spoken language model for emotional intelligence, enhanced by a novel data construction strategy termed Injected Emotional-Attribution Thinking (IEAT). IEAT incorporates user emotional states and their underlying causes into the model's internal reasoning process, enabl…

Cited by 0SourcePDFScholar
2026

Adversarial Attack on Black-Box Multi-Agent by Adaptive Perturbation

AAAI 2026technical

Evaluating security and reliability for multi-agent systems (MAS) is urgent as they become increasingly prevalent in various applications. As an evaluation technique, existing adversarial attack frameworks face certain limitations, e.g., impracticality due to the requirement of white-box information

Cited by 0SourcePDFScholar
2026

Enhancing DPSGD via Per-Sample Momentum and Low-Pass Filtering

AAAI 2026technical

Differentially Private Stochastic Gradient Descent (DPSGD) is widely used to train deep neural networks with formal privacy guarantees. However, the addition of differential privacy (DP) often degrades model accuracy by introducing both noise and bias. Existing techniques typically address only one

Cited by 0SourcePDFScholar
2026

Joint-GCG: Unified Gradient-Based Poisoning Attacks on Retrieval-Augmented Generation Systems

AAAI 2026technical

Retrieval-Augmented Generation (RAG) systems enhance Large Language Models (LLMs) by retrieving relevant documents from external corpora before generating responses. This approach significantly expands LLM capabilities by leveraging vast, up-to-date external knowledge. However, this reliance on exte

Cited by 0SourcePDFScholar
2026

ManifoldNeuS: Manifold-aware View Optimizability for Pose-Free Neural Surface Reconstruction

CVPR 2026

Jointly optimizing camera poses and object geometry from unposed images is a challenging task in neural surface reconstruction. Existing methods often suffer from pose drift and geometric distortion, stemming from the easy-view bias --- uniform view optimization favors easy-to-optimize views with ab

Cited by 0SourceScholar
2026

Many Minds, One Path: LLM-Augmented Consensus Decision for Distributed Control in Multi-Agent Collaborative Stable Scenarios

AAAI 2026technical

Distributed multi-agent systems are increasingly deployed in dynamic and high-stakes environments such as power grids, intelligent traffic systems, and collaborative robotics. In these systems, long-term stability, the ability to maintain coherent and safe system behavior over time, is critical but

Cited by 0SourcePDFScholar
2026

Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning

ICASSP 2026oral

We present Task 5 of the DCASE 2025 Challenge: an Audio Question Answering (AQA) benchmark spanning multiple domains of sound understanding. This task defines three QA subsets (Bioacoustics, Temporal Soundscapes, and Complex QA) to test audio-language models on interactive question-answering over di…

Cited by 0SourcePDFScholar
2026

OrchJail: Jailbreaking Tool-Calling Text-to-Image Agents by Orchestration-Guided Fuzzing

ICML 2026poster

Tool-calling text-to-image (T2I) agents can plan and execute multi-step tool chains to accomplish complex generation and editing queries. However, this capability introduces a new safety attack surface: harmful outputs may arise from tool orchestration, where individually benign steps combine into u…

Cited by 0SourceScholar
2026

SMoFi: Step-wise Momentum Fusion for Split Federated Learning on Heterogeneous Data

AAAI 2026technical

Split Federated Learning is a system-efficient federated learning paradigm that leverages the rich computing resources at a central server to train model partitions. Data heterogeneity across silos, however, presents a major challenge undermining the convergence speed and accuracy of the global mode

Cited by 0SourcePDFScholar
2026

Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks

ICML 2026poster

The safety alignment of Large Language Models (LLMs) remains vulnerable to Harmful Fine-tuning (HFT). While existing defenses impose constraints on parameters, gradients, or internal representations, we observe that they can be effectively circumvented under persistent HFT. Our analysis traces this …

Cited by 0SourceScholar
2026

WENETSPEECH-CHUAN: A LARGE-SCALE SICHUANESE CORPUS WITH RICH ANNOTATION FOR DIALECTAL SPEECH PROCESSING

ICASSP 2026poster

The scarcity of large-scale, open-source data for dialects severely hinders progress in speech technology, a challenge particularly acute for the widely spoken Sichuanese dialects of Chinese. To address this critical gap, we introduce WenetSpeech-Chuan, a 10,000-hour, richly annotated corpus constru…

Cited by 0SourcePDFScholar
2025

An Experimental Study on Joint Modeling for Sound Event Localization and Detection with Source Distance Estimation

ICASSP 2025accepted

In traditional sound event localization and detection (SELD) tasks, the focus is typically on sound event detection (SED) and direction-of-arrival (DOA) estimation, but they fall short of providing full spatial information about the sound source. The 3D SELD task addresses this limitation by integra…

Cited by 0SourceScholar
2025

Bright-NeRF: Brightening Neural Radiance Field with Color Restoration from Low-Light RAW Images

AAAI 2025technical

Neural Radiance Fields (NeRF) have demonstrated prominent performance in novel view synthesis tasks. However, their input heavily relies on image acquisition under normal light conditions, making it challenging to learn accurate scene contents in low-light environments where images typically exhibit…

Cited by 0SourcePDFScholar
2025

DU-PMVS: Learned Patchmatch Multi-View Stereo Based on Deformable Feature Pyramid and Uncertainty Awareness Modeling

ICASSP 2025accepted

Multi-View Stereo is widely utilized for reconstructing the dense geometric structure of objects from multiple viewpoints. Recently, learning-based PatchMatch MVS methods have attracted significant attention due to their high efficiency and accuracy. However, existing methods neglect the constraints…

Cited by 0SourceScholar
2025

DeepSN: A Sheaf Neural Framework for Influence Maximization

AAAI 2025technical

Influence maximization is a key topic in data mining, with broad applications in social network analysis and viral marketing. In recent years, researchers have increasingly turned to machine learning techniques to address this problem. By learning the underlying diffusion processes from data, these…

2025

DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification

ICASSP 2025accepted

Being a form of biometric identification, the security of the speaker identification (SID) system is of utmost importance. To better understand the robustness of SID systems, we aim to perform more realistic attacks in SID, which are challenging for humans and machines to detect. In this study, we p…

Cited by 0SourceScholar
2025

Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation Learning

ICLR 2025poster

Recovering a spectrum of diverse policies from a set of expert trajectories is an important research topic in imitation learning. After determining a latent style for a trajectory, previous diverse polices recovering methods usually employ a vanilla behavioral cloning learning objective conditioned…

Cited by 0SourcePDFScholar
2025

Flick: Empowering Federated Learning with Commonsense Knowledge

NeurIPS 2025poster

Federated Learning (FL) has emerged as a privacy-preserving framework for training models on data generated at the edge. However, the heterogeneity of data silos (e.g., label skew and domain shift) often leads to inconsistent learning objectives and suboptimal model performance. Inspired by the data…

Cited by 0SourceScholar
2025

From Allies to Adversaries: Manipulating LLM Tool-Calling through Adversarial Injection

NAACL 2025long

Tool-calling has changed Large Language Model (LLM) applications by integrating external tools, significantly enhancing their functionality across diverse tasks. However, this integration also introduces new security vulnerabilities, particularly in the tool scheduling mechanisms of LLM, which have…

2025

GLiM: Integrating Graph Transformer and LLM for Document-Level Biomedical Relation Extraction with Incomplete Labeling

ACL 2025finding

Document-level relation extraction (DocRE) identifies relations between entities across an entire document. However, as the number and complexity of entities and entity-pair relations grow, the problem space expands quadratically, causing incomplete annotations and frequent false negatives, especial…

2025

Indicators to Measure and Analyze Healthcare Accessibility in Metropolitan Cities in China

RA-L 2025

Healthcare accessibility is crucial in public medical services, and has become one of center issues of healthcare services. Enabling fair and quick access to healthcare services is of critical importance for all populations, particularly for the elderly people in metropolitan cities. In this paper,

Cited by 0SourceScholar
2025

Investigating Context Faithfulness in Large Language Models: The Roles of Memory Strength and Evidence Style

ACL 2025finding

Retrieval-augmented generation (RAG) improves Large Language Models (LLMs) by incorporating external information into the response generation process. However, how context-faithful LLMs are and what factors influence LLMs’ context faithfulness remain largely unexplored. In this study, we investigate…

2025

MVANet: Multi-Stage Video Attention Network for Sound Event Localization and Detection with Source Distance Estimation

ICASSP 2025accepted

Sound event localization and detection with source distance estimation (3D SELD) involves not only identifying the sound category and its direction-of-arrival (DOA) but also predicting the source's distance, aiming to provide full information about the sound position. This paper proposes a multi-sta…

Cited by 0SourceScholar
2025

Material Anything: Generating Materials for Any 3D Object via Diffusion

CVPR 2025highlight

We present **Material Anything**, a fully-automated, unified diffusion framework designed to generate physically-based materials for 3D objects. Unlike existing methods that rely on complex pipelines or case-specific optimizations, Material Anything offers a robust, end-to-end solution adaptable to…

Cited by 4SourcePDFScholar
2025

Mimicking the Familiar: Dynamic Command Generation for Information Theft Attacks in LLM Tool-Learning System

ACL 2025long

Information theft attacks pose a significant risk to Large Language Model (LLM) tool-learning systems. Adversaries can inject malicious commands through compromised tools, manipulating LLMs to send sensitive information to these tools, which leads to potential privacy breaches. However, existing att…

2025

MultiPL-MoE: Multi-Programming-Lingual Extension of Large Language Models through Hybrid Mixture-of-Experts

EMNLP 2025

Despite LLMs’ excellent code creation capabilities, multilingual code generation remains extremely challenging. To address this, we intent to improve the multi-programming-lingual (MultiPL) performance of the base LLMs while retaining the most popular ones using restricted computational resources. W

2025

NTR-Gaussian: Nighttime Dynamic Thermal Reconstruction with 4D Gaussian Splatting Based on Thermodynamics

CVPR 2025poster

Thermal infrared imaging enables a non-invasive measurement of the surface temperature of objects with all-weather applicability. Leveraging such techniques for 3D reconstruction can accurately reflect the temperature distribution of a scene, thereby supporting applications such as building monitori…

Cited by 1SourcePDFScholar
2025

One Shot Dominance: Knowledge Poisoning Attack on Retrieval-Augmented Generation Systems

EMNLP 2025

Large Language Models (LLMs) enhanced with Retrieval-Augmented Generation (RAG) have shown improved performance in generating accurate responses. However, the dependence on external knowledge bases introduces potential security vulnerabilities, particularly when these knowledge bases are publicly ac

Cited by 0SourcePDFScholar
2025

Re-Examine Distantly Supervised NER: A New Benchmark and a Simple Approach

COLING 2025main

Distantly-Supervised Named Entity Recognition (DS-NER) uses knowledge bases or dictionaries for annotations, reducing manual efforts but rely on large human labeled validation set. In this paper, we introduce a real-life DS-NER dataset, QTL, where the training data is annotated using domain dictiona…

2025

STEM-POM: Evaluating Language Models Math-Symbol Reasoning in Document Parsing

ACL 2025finding

Advances in large language models (LLMs) have spurred research into enhancing their reasoning capabilities, particularly in math-rich STEM (Science, Technology, Engineering, and Mathematics) documents.While LLMs can generate equations or solve math-related queries, their ability to fully understand…

2025

STaR: Seamless Spatial-Temporal Aware Motion Retargeting with Penetration and Consistency Constraints

ICCV 2025poster

Motion retargeting seeks to faithfully replicate the spatio-temporal motion characteristics of a source character onto a target character with a different body shape. Apart from motion semantics preservation, ensuring geometric plausibility and maintaining temporal consistency are also crucial for e…

2025

Towards Bridging Generalization and Expressivity of Graph Neural Networks

ICLR 2025poster

Expressivity and generalization are two critical aspects of graph neural networks (GNNs). While significant progress has been made in studying the expressivity of GNNs, much less is known about their generalization capabilities, particularly when dealing with the inherent complexity of graph-structu…

Cited by 1SourcePDFScholar
2025

Towards a More Generalized Approach in Open Relation Extraction

ACL 2025long

Open Relation Extraction (OpenRE) seeks to identify and extract novel relational facts between named entities from unlabeled data without pre-defined relation schemas. Traditional OpenRE methods typically assume that the unlabeled data consists solely of novel relations or is pre-divided into known…

2025

Understanding Individual Agent Importance in Multi-Agent System via Counterfactual Reasoning

AAAI 2025technical

Explaining multi-agent systems (MAS) is urgent as these systems become increasingly prevalent in various applications. Previous work has provided explanations for the actions or states of agents, yet falls short in understanding the blackboxed agent’s importance within a MAS and the overall team str…

Cited by 0SourcePDFScholar
2024

3-D Near-Field Localization by Jointly Exploiting Spatial and Temporal Information Based on a Nonuniform Cross Array

ICASSP 2024accepted

In this paper, an underdetermined three-dimensional (3-D) near-field source localization method is proposed, based on a two-dimensional (2-D) symmetric nonuniform cross array. Firstly, the fourth-order cumulant of the near-field observations with multiple delay lags is exploited to construct virtual…

Cited by 0SourceScholar
2024

Feature Mixing-Based Active Learning for Multi-Label Text Classification

ICASSP 2024accepted

Active learning (AL) aims to reduce labeling costs by selecting the most valuable samples to annotate from a set of unlabeled data. However, recognizing these samples is particularly challenging in multi-label text classification tasks due to the high dimensionality but sparseness of label spaces. E…

Cited by 0SourceScholar
2024

FedTrans: Client-Transparent Utility Estimation for Robust Federated Learning

ICLR 2024poster

Federated Learning (FL) is an important privacy-preserving learning paradigm that plays an important role in the Intelligent Internet of Things. Training a global model in FL, however, is vulnerable to the noise in the heterogeneous data across the clients. In this paper, we introduce **FedTrans**,…

Cited by 0SourcePDFScholar
2024

GenDecider: Integrating “None of the Candidates” Judgments in Zero-Shot Entity Linking Re-ranking

NAACL 2024short

We introduce GenDecider, a novel re-ranking approach for Zero-Shot Entity Linking (ZSEL), built on the Llama model. It innovatively detects scenarios where the correct entity is not among the retrieved candidates, a common oversight in existing re-ranking methods. By autoregressively generating outp…

2024

How to Evaluate the Generalization of Detection? A Benchmark for Comprehensive Open-Vocabulary Detection

AAAI 2024technical

Object detection (OD) in computer vision has made significant progress in recent years, transitioning from closed-set labels to open-vocabulary detection (OVD) based on large-scale vision-language pre-training (VLP). However, current evaluation methods and datasets are limited to testing generalizat…

2024

HumanNorm: Learning Normal Diffusion Model for High-quality and Realistic 3D Human Generation

CVPR 2024poster

Recent text-to-3D methods employing diffusion models have made significant advancements in 3D human generation. However these approaches face challenges due to the limitations of text-to-image diffusion models which lack an understanding of 3D structures. Consequently these methods struggle to achie…

2024

Play Guessing Game with LLM: Indirect Jailbreak Attack with Implicit Clues

ACL 2024findings

With the development of LLMs, the security threats of LLMs are getting more and more attention. Numerous jailbreak attacks have been proposed to assess the security defense of LLMs. Current jailbreak attacks primarily utilize scenario camouflage techniques. However their explicitly mention of malici…

Cited by 42SourcePDFScholar
2024

Repairing Catastrophic-Neglect in Text-to-Image Diffusion Models via Attention-Guided Feature Enhancement

EMNLP 2024finding

Text-to-Image Diffusion Models (T2I DMs) have garnered significant attention for their ability to generate high-quality images from textual descriptions.However, these models often produce images that do not fully align with the input prompts, resulting in semantic inconsistencies.The most prominent…

2024

SRFUND: A Multi-Granularity Hierarchical Structure Reconstruction Benchmark in Form Understanding

NeurIPS 2024poster

Accurately identifying and organizing textual content is crucial for the automation of document processing in the field of form understanding. Existing datasets, such as FUNSD and XFUND, support entity classification and relationship prediction tasks but are typically limited to local and entity-lev…

2023

$\mathscr{N}$-WL: A New Hierarchy of Expressivity for Graph Neural Networks

ICLR 2023poster

The expressive power of Graph Neural Networks (GNNs) is fundamental for understanding their capabilities and limitations, i.e., what graph properties can or cannot be learnt by a GNN. Since standard GNNs have been characterised to be upper-bounded by the Weisfeiler-Lehman (1-WL) algorithm, recent a…

Cited by 19SourcePDFScholar
2023

An Experimental Study on Sound Event Localization and Detection Under Realistic Testing Conditions

ICASSP 2023accepted

We study four data augmentation (DA) techniques and two model architectures on realistic data for sound event localization and detection (SELD). First, based on ResNet-Conformer (RC), we compare the four DA approaches on the realistic DCASE 2022 SELD test set which is often not easy to handle due to…

Cited by 0SourceScholar
2023

CNVid-3.5M: Build, Filter, and Pre-Train the Large-Scale Public Chinese Video-Text Dataset

CVPR 2023poster

Owing to well-designed large-scale video-text datasets, recent years have witnessed tremendous progress in video-text pre-training. However, existing large-scale video-text datasets are mostly English-only. Though there are certain methods studying the Chinese video-text pre-training, they pre-train…

2023

Distance-Based Weight Transfer for Fine-Tuning From Near-Field to Far-Field Speaker Verification

ICASSP 2023accepted

The scarcity of labeled far-field speech is a constraint for training superior far-field speaker verification systems. In general, fine-tuning the model pre-trained on large-scale near- field speech through a small amount of far-field speech substantially outperforms training from scratch. However,…

Cited by 0SourceScholar
2023

Distinguishable Speaker Anonymization Based on Formant and Fundamental Frequency Scaling

ICASSP 2023accepted

Speech data on the Internet are proliferating exponentially because of the emergence of social media, and the sharing of such personal data raises obvious security and privacy concerns. One solution to mitigate these concerns involves concealing speaker identities before sharing speech data, also re…

Cited by 0SourceScholar
2023

Fault Injection Based Interventional Causal Learning for Distributed Applications

AAAI 2023technical

We apply the machinery of interventional causal learning with programmable interventions to the domain of applications management. Modern applications are modularized into interdependent components or services (e.g. microservices) for ease of development and management. The communication graph among…

2023

Improving Unsupervised Relation Extraction by Augmenting Diverse Sentence Pairs

EMNLP 2023long main

Unsupervised relation extraction (URE) aims to extract relations between named entities from raw text without requiring manual annotations or pre-existing knowledge bases. In recent studies of URE, researchers put a notable emphasis on contrastive learning strategies for acquiring relation represen…

Cited by 0SourcecodeScholar
2023

Incorporating Lip Features into Audio-Visual Multi-Speaker DOA Estimation by Gated Fusion

ICASSP 2023accepted

The audio-visual direction of arrival (DOA) estimation has demonstrated superior performance recently. In this paper, we present a novel audio-visual multi-speaker DOA estimation network, which for the first time incorporates multi-speaker lip features to adapt the complex overlapping and noisy scen…

Cited by 0SourceScholar
2023

Inverting the Imaging Process by Learning an Implicit Camera Model

CVPR 2023poster

Representing visual signals with implicit coordinate-based neural networks, as an effective replacement of the traditional discrete signal representation, has gained considerable popularity in computer vision and graphics. In contrast to existing implicit neural representations which focus on modell…

Cited by 14SourcePDFScholar
2023

Large Language Models are Complex Table Parsers

EMNLP 2023long main

With the Generative Pre-trained Transformer 3.5 (GPT-3.5) exhibiting remarkable reasoning and comprehension abilities in Natural Language Processing (NLP), most Question Answering (QA) research has primarily centered around general QA tasks based on GPT, neglecting the specific challenges posed by C…

Cited by 0SourceScholar
2023

Local Implicit Ray Function for Generalizable Radiance Field Representation

CVPR 2023poster

We propose LIRF (Local Implicit Ray Function), a generalizable neural rendering approach for novel view rendering. Current generalizable neural radiance fields (NeRF) methods sample a scene with a single ray per pixel and may therefore render blurred or aliased views when the input views and rendere…

Cited by 30SourcePDFScholar
2023

Loss Function Design for DNN-Based Sound Event Localization and Detection on Low-Resource Realistic Data

ICASSP 2023accepted

This study focuses on the design of a loss function for a deep neural network (DNN)-based model with two branches, which is used to solve sound event localization and detection (SELD) on low-resource realistic data. To this end, we employ a secondary network for audio classification, which provides…

Cited by 0SourceScholar
2023

Neural Ideal Large Eddy Simulation: Modeling Turbulence with Neural Stochastic Differential Equations

NeurIPS 2023poster

We introduce a data-driven learning framework that assimilates two powerful ideas: ideal large eddy simulation (LES) from turbulence closure modeling and neural stochastic differential equations (SDE) for stochastic modeling. The ideal LES models the LES flow by treating each full-order trajectory a…

Cited by 8SourcePDFScholar
2023

Preserving Background Sound in Noise-Robust Voice Conversion Via Multi-Task Learning

ICASSP 2023accepted

Background sound is an informative form of art that is helpful in providing a more immersive experience in real-application voice conversion (VC) scenarios. However, prior research about VC, mainly focusing on clean voices, pay rare attention to VC with background sound. The critical problem for pre…

Cited by 0SourceScholar
2023

Restructuring Graph for Higher Homophily via Adaptive Spectral Clustering

AAAI 2023technical

While a growing body of literature has been studying new Graph Neural Networks (GNNs) that work on both homophilic and heterophilic graphs, little has been done on adapting classical GNNs to less-homophilic graphs. Although the ability to handle less-homophilic graphs is restricted, classical GNNs s…

2023

The NERCSLIP-USTC System for the L3DAS23 Challenge Task2: 3D Sound Event Localization and Detection (SELD)

ICASSP 2023accepted

Sound event localization and detection (SELD) aims at identifying the temporal activities of a known set of sound event classes and estimating their locations. It remains challenging especially when there are overlapped acoustic events. In this work, a robust network architecture with data augmentat…

Cited by 0SourceScholar
2022

A Large-Scale Comprehensive Dataset and Copy-Overlap Aware Evaluation Protocol for Segment-Level Video Copy Detection

CVPR 2022poster

In this paper, we introduce VCSL (Video Copy Segment Localization), a new comprehensive segment-level annotated video copy dataset. Compared with existing copy detection datasets restricted by either video-level annotation or small-scale, VCSL not only has two orders of magnitude more segment-level…

Cited by 18PDFcodeScholar
2022

Modeling and Analysis of Operating Room Workflow in a Tertiary A Hospital

RA-L 2022

Operating room (OR) is one of the most critical units in a hospital. Managing surgical processes in ORs for better utilization of medical resources, safe delivery of surgical cases, improving patient outcome, and reducing cost, is of significant importance. In this paper, a discrete-event simulation

Cited by 6SourceScholar
2022

MuiDial: Improving Dialogue Disentanglement with Intent-Based Mutual Learning

IJCAI 2022poster

The main goal of dialogue disentanglement is to separate the mixed utterances from a chat slice into independent dialogues. Existing models often utilize either an utterance-to-utterance (U2U) prediction to determine whether two utterances that have the “reply-to” relationship belong to one dialogue…

Cited by 2SourcePDFScholar
2021

Dialogue Disentanglement in Software Engineering: How Far are We?

IJCAI 2021poster

Despite the valuable information contained in software chat messages, disentangling them into distinct conversations is an essential prerequisite for any in-depth analyses that utilize this information. To provide a better understanding of the current state-of-the-art, we evaluate five popular dialo…

2021

Speech Enhancement Autoencoder with Hierarchical Latent Structure

ICASSP 2021accepted

A new hierarchical convolutional neural network-based autoencoder architecture called SEHAE (Speech Enhancement Hierarchical AutoEncoder) is introduced, in which the latent representation is decomposed into several parts that correspond to different scales. The model consists of three functionally d…

Cited by 0SourceScholar
2020

Emotion Classification by Jointly Learning to Lexiconize and Classify

COLING 2020main

Emotion lexicons have been shown effective for emotion classification (Baziotis et al., 2018). Previous studies handle emotion lexicon construction and emotion classification separately. In this paper, we propose an emotional network (EmNet) to jointly learn sentence emotions and construct emotion l…

2020

Geometry Constrained Progressive Learning for Lstm-Based Speech Enhancement

ICASSP 2020accepted

In our previous work, a progressive learning framework for long short-term memory (LSTM)-based speech enhancement was proposed to improve the performance in low SNR environment, where each LSTM layer is guided to learn an intermediate target with a specific SNR gain via the MMSE criterion. However,…

Cited by 0SourceScholar
2019

DFNets: Spectral CNNs for Graphs with Feedback-Looped Filters

NeurIPS 2019poster

We propose a novel spectral convolutional neural network (CNN) model on graph structured data, namely Distributed Feedback-Looped Networks (DFNets). This model is incorporated with a robust class of spectral graph filters, called feedback-looped filters, to provide better localization on vertices, w…

2019

Grid-Wise Control for Multi-Agent Reinforcement Learning in Video Game AI

ICML 2019oral

We consider the problem of multi-agent reinforcement learning (MARL) in video game AI, where the agents are located in a spatial grid-world environment and the number of agents varies both within and across episodes. The challenge is to flexibly control an arbitrary number of agents while achieving…

Cited by 72SourcePDFScholar
2019

Gridless Super-resolution Doa Estimation with Unknown Mutual Coupling

ICASSP 2019accepted

In this paper, a gridless super-resolution direction-of-arrival (DOA) estimation method with unknown mutual coupling is proposed. A new clean steering vector is obtained based on the banded symmetric Toeplitz structure of the mutual coupling matrix (MCM). Further, atomic norms associated with the ar…

Cited by 0SourceScholar
2018

Exponentially Weighted Imitation Learning for Batched Historical Data

NeurIPS 2018poster

We consider deep policy learning with only batched historical trajectories. The main challenge of this problem is that the learner no longer has a simulator or ``environment oracle'' as in most reinforcement learning settings. To solve this problem, we propose a monotonic advantage reweighted imitat…

2018

Unsupervised Domain Adaptation via Domain Adversarial Training for Speaker Recognition

ICASSP 2018accepted

The i-vector approach to speaker recognition has achieved good performance when the domain of the evaluation dataset is similar to that of the training dataset. However, in realworld applications, there is always a mismatch between the training and evaluation datasets, that leads to performance degr…

Cited by 0SourceScholar