← Search

Yuexian Zou

94 accepted papers

2026

IC-Custom: Diverse Image Customization via In-Context Learning

ICLR 2026poster

Image customization, a crucial technique for industrial media production, aims to generate content that is consistent with reference images. However, current approaches conventionally separate image customization into position-aware and position-free customization paradigms and lack a universal fram…

Cited by 0SourcecodeScholar
2026

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation

AAAI 2026technical

Large vision-language models (LVLMs) have demonstrated impressive capabilities across diverse multimodal tasks, yet they remain highly susceptible to visual hallucinations (VH), often producing confident but inaccurate descriptions of visual content. Building on the insight that not all tokens and a

Cited by 0SourcePDFScholar
2025

ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by Reducing Data Distribution Errors

ACL 2025long

Multilingual audio-text retrieval (ML-ATR) is a challenging task that aims to retrieve audio clips or multilingual texts from databases. However, existing ML-ATR schemes suffer from inconsistencies for instance similarity matching across languages. To address the inconsistency issue in multilingual…

2025

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding

ACL 2025finding

Vision-language Models (VLMs) have shown remarkable capabilities in advancing general artificial intelligence, yet the irrational encoding of visual positions persists in inhibiting the models’ comprehensive perception performance across different levels of granularity. In this work, we propose Pyra…

2025

Audio-Faces Intra-Frame Alignment with Graph Attention Networks for Active Speaker Detection

ICASSP 2025accepted

Audio-Visual Active Speaker Detection(ASD) is the task of identifying, at any given moment, who is actively speaking in a multi-person scene by using audio and visual cues. Current main stream ASD methods separately encode audio and facial features, then adopt post-feature fusion approach where the…

Cited by 0SourceScholar
2025

Image Conductor: Precision Control for Interactive Video Synthesis

AAAI 2025technical

Filmmaking and animation production often require sophisticated techniques for coordinating camera transitions and object movements, typically involving labor-intensive real-world capturing. Despite advancements in generative AI for video creation, achieving precise control over motion for interacti…

2025

UniCoTT: A Unified Framework for Structural Chain-of-Thought Distillation

ICLR 2025poster

Chains of thought (CoTs) have achieved success in enhancing the reasoning capabilities of large language models (LLMs), while their effectiveness is predominantly observed in LLMs. Existing solutions methods adopt distillation to inject chain-of-thought capabilities into small models (SLMs). Howeve…

2025

VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token Sparsification

CVPR 2025poster

Large Vision-Language Models (LVLMs) may produce outputs that are unfaithful to reality, also known as visual hallucinations (VH), which significantly impedes their real-world usage. To alleviate VH, various decoding strategies have been proposed to enhance visual information. However, many of these…

2024

Aligner²: Enhancing Joint Multiple Intent Detection and Slot Filling via Adjustive and Forced Cross-Task Alignment

AAAI 2024technical

Multi-intent spoken language understanding (SLU) has garnered growing attention due to its ability to handle multiple intent utterances, which closely mirrors practical scenarios. Unlike traditional SLU, each intent in multi-intent SLU corresponds to its designated scope for slots, which occurs in…

2024

AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

AAAI 2024technical

Large language models (LLMs) have exhibited remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. Despite the recent success, current LLMs are not capable of processing complex audio information or conducting spoken conversations (lik…

2024

Code-Switching Can be Better Aligners: Advancing Cross-Lingual SLU through Representation-Level and Prediction-Level Alignment

ACL 2024short

Zero-shot cross-lingual spoken language understanding (SLU) can promote the globalization application of dialog systems, which has attracted increasing attention. While current code-switching based cross-lingual SLU frameworks have shown promising results, they (i) predominantly utilize contrastive…

2024

Cyclical Contrastive Learning Based on Geodesic for Zero-shot Cross-lingual Spoken Language Understanding

ACL 2024findings

Owing to the scarcity of labeled training data, Spoken Language Understanding (SLU) is still a challenging task in low-resource languages. Therefore, zero-shot cross-lingual SLU attracts more and more attention. Contrastive learning is widely applied to explicitly align representations of similar se…

2024

Dual-oriented Disentangled Network with Counterfactual Intervention for Multimodal Intent Detection

EMNLP 2024main

Multimodal intent detection is designed to leverage diverse modalities for a comprehensive understanding of user intentions in real-world scenarios, thus playing a critical role in modern task-oriented dialogue systems. Existing methods have made great progress in modal alignment and fusion, however…

Cited by 1SourcePDFScholar
2024

Embracing Language Inclusivity and Diversity in CLIP through Continual Language Learning

AAAI 2024technical

While vision-language pre-trained models (VL-PTMs) have advanced multimodal research in recent years, their mastery in a few languages like English restricts their applicability in broader communities. To this end, there is an increasing interest in developing multilingual VL models via a joint-lear…

2024

Exploiting Auxiliary Caption for Video Grounding

AAAI 2024technical

Video grounding aims to locate a moment of interest matching the given query sentence from an untrimmed video. Previous works ignore the sparsity dilemma in video annotations, which fails to provide the context information between potential events and query sentences in the dataset. In this paper, w…

Cited by 17SourcePDFScholar
2024

Game on Tree: Visual Hallucination Mitigation via Coarse-to-Fine View Tree and Game Theory

EMNLP 2024main

Large Vision-Language Models (LVLMs) may produce outputs that are unfaithful to reality, also known as visual hallucinations (VH), which hinders their application in multimodal understanding and decision-making. In this work, we introduce a novel plug-and-play train-free decoding algorithm named Gam…

2024

KDProR: A Knowledge-Decoupling Probabilistic Framework for Video-Text Retrieval

ECCV 2024poster

"Existing video-text retrieval methods predominantly focus on designing diverse cross-modal interaction mechanisms between captions and videos. However, those approaches diverge from human learning paradigms, where humans possess the capability to seek and associate knowledge from an open set, rathe…

Cited by 8SourcePDFScholar
2024

Knowledge-enhanced Prompt Tuning for Dialogue-based Relation Extraction with Trigger and Label Semantic

COLING 2024main

Dialogue-based relation extraction (DRE) aims to determine the semantic relation of a given pair of arguments from a piece of dialogue, which has received increasing attention. Due to the low information density of dialogue text, it is difficult for the model to focus on key information. To this end…

2024

Learning to Match Representations is Better for End-to-End Task-Oriented Dialog System

EMNLP 2024finding

Due to the rapid development with pre-trained language models, fully end-to-end Task-Oriented Dialogue (TOD) systems exhibit superior performance. How to achieve the ability to efficiently retrieve entities in cross-domain large-scale databases is a key issue. Most existing end-to-end Task-Oriented…

Cited by 0SourcePDFScholar
2024

MaCSC: Towards Multimodal-augmented Pre-trained Language Models via Conceptual Prototypes and Self-balancing Calibration

NAACL 2024long

Pre-trained language models (PLMs) that rely solely on textual data may exhibit limitations in multimodal semantics comprehension. Existing solutions attempt to alleviate this issue by incorporating explicit image retrieval or generation techniques.However, these methods: (1) focus exclusively on th…

2024

MoE-SLU: Towards ASR-Robust Spoken Language Understanding via Mixture-of-Experts

ACL 2024findings

As a crucial task in the task-oriented dialogue systems, spoken language understanding (SLU) has garnered increasing attention. However, errors from automatic speech recognition (ASR) often hinder the performance of understanding. To tackle this problem, we propose MoE-SLU, an ASR-Robust SLU framewo…

Cited by 2SourcePDFScholar
2024

On the Worst Prompt Performance of Large Language Models

NeurIPS 2024poster

The performance of large language models (LLMs) is acutely sensitive to the phrasing of prompts, which raises significant concerns about their reliability in real-world scenarios. Existing studies often divide prompts into task-level instructions and case-level inputs and primarily focus on evaluati…

Cited by 8SourcePDFScholar
2024

PCAD: Towards ASR-Robust Spoken Language Understanding via Prototype Calibration and Asymmetric Decoupling

ACL 2024long

Spoken language understanding (SLU) inevitably suffers from error propagation from automatic speech recognition (ASR) in actual scenarios. Some recent works attempt to alleviate this issue through contrastive learning. However, they (1) sample negative pairs incorrectly in pre-training; (2) only foc…

2024

Relevance Is a Guiding Light: Relevance-aware Adaptive Learning for End-to-end Task-oriented Dialogue System

EMNLP 2024main

Retrieving accurate domain knowledge and providing helpful information are crucial in developing an effective end-to-end task-oriented dialogue system (E2ETOD). The field has witnessed numerous methods following the retrieve-then-generate paradigm and training their systems on one specific domain. H…

Cited by 1SourcePDFScholar
2024

Soul-Mix: Enhancing Multimodal Machine Translation with Manifold Mixup

ACL 2024long

Multimodal machine translation (MMT) aims to improve the performance of machine translation with the help of visual information, which has received widespread attention recently. It has been verified that visual information brings greater performance gains when the textual information is limited. Ho…

2024

Towards Explainable Joint Models via Information Theory for Multiple Intent Detection and Slot Filling

AAAI 2024technical

Recent joint models for multi-intent detection and slot filling have obtained promising results through modeling the unidirectional or bidirectional guidance between intent and slot. However, existing works design joint models heuristically and lack some theoretical exploration, including (1) theore…

2024

Towards Multi-Intent Spoken Language Understanding via Hierarchical Attention and Optimal Transport

AAAI 2024technical

Multi-Intent spoken language understanding (SLU) can handle complicated utterances expressing multiple intents, which has attracted increasing attention from researchers. Although existing models have achieved promising performance, most of them still suffer from two leading problems: (1) each inten…

2024

Towards Multi-modal Sarcasm Detection via Disentangled Multi-grained Multi-modal Distilling

COLING 2024main

Multi-modal sarcasm detection aims to identify whether a given sample with multi-modal information (i.e., text and image) is sarcastic, which has received increasing attention due to the rapid growth of multi-modal posts on modern social media. However, mainstream models process the input of each mo…

2024

What are the Generator Preferences for End-to-end Task-Oriented Dialog System?

EMNLP 2024main

Fully end-to-end task-oriented dialogue (EToD) systems have shown excellent performance, which requires the ability to retrieve entities accurately for generation. Existing methods improve the accuracy of entity retrieval and construct data flows between retrieval results and response generator, ach…

Cited by 0SourcePDFScholar
2023

A Dynamic Graph Interactive Framework with Label-Semantic Injection for Spoken Language Understanding

ICASSP 2023accepted

Multi-intent detection and slot filling joint models are gaining increasing traction since they are closer to complicated real-world scenarios. However, existing approaches (1) focus on identifying implicit correlations between utterances and one-hot encoded labels in both tasks while ignoring expli…

Cited by 0SourceScholar
2023

Accelerating Multiple Intent Detection and Slot Filling via Targeted Knowledge Distillation

EMNLP 2023long findings

Recent non-autoregressive Spoken Language Understanding (SLU) models attracts increasing attention owing to the high inference speed. However, most of them still (1) suffer from the multi-modality problem since the prior knowledge about the reference is relatively poor during inference; (2) fail to…

Cited by 0SourceScholar
2023

Enhancing Code-Switching for Cross-lingual SLU: A Unified View of Semantic and Grammatical Coherence

EMNLP 2023short main

Despite the success of spoken language understanding (SLU) in high-resource languages, achieving similar performance in low-resource settings, such as zero-shot scenarios, remains challenging due to limited labeled training data. To improve zero-shot cross-lingual SLU, recent studies have explored c…

Cited by 0SourceScholar
2023

FTM: A Frame-Level Timeline Modeling Method for Temporal Graph Representation Learning

AAAI 2023technical

Learning representations for graph-structured data is essential for graph analytical tasks. While remarkable progress has been made on static graphs, researches on temporal graphs are still in its beginning stage. The bottleneck of the temporal graph representation learning approach is the neighborh…

2023

FiTs: Fine-Grained Two-Stage Training for Knowledge-Aware Question Answering

AAAI 2023technical

Knowledge-aware question answering (KAQA) requires the model to answer questions over a knowledge base, which is essential for both open-domain QA and domain-specific QA, especially when language models alone cannot provide all the knowledge needed. Despite the promising result of recent KAQA system…

2023

G2L: Semantically Aligned and Uniform Video Grounding via Geodesic and Game Theory

ICCV 2023oral

The recent video grounding works attempt to introduce vanilla contrastive learning into video grounding. However, we claim that this naive solution is suboptimal. Contrastive learning requires two key properties: (1) alignment of features of similar samples, and (2) uniformity of the induced distrib…

Cited by 52PDFScholar
2023

Improving Retrieval-Based Dialogue System Via Syntax-Informed Attention

ICASSP 2023accepted

Multi-turn response selection is a challenging task due to its high demands on efficient extraction of the matching features from abundant information provided by context utterances. Since incorporating syntactic information like dependency structures into neural models can promote a better understa…

Cited by 0SourceScholar
2023

Improving Text-Audio Retrieval by Text-Aware Attention Pooling and Prior Matrix Revised Loss

ICASSP 2023accepted

In text-audio retrieval (TAR) tasks, due to the heterogeneity of contents between text and audio, the semantic information contained in the text is only similar to certain frames within the audio. Yet, existing works aggregate the entire audio without considering the text, such as mean-pooling over…

Cited by 0SourceScholar
2023

Improving Weakly Supervised Sound Event Detection with Causal Intervention

ICASSP 2023accepted

Existing weakly supervised sound event detection (WSSED) work has not explored both types of co-occurrences simultaneously, i.e., some sound events often co-occur, and their occurrences are usually accompanied by specific background sounds, so they would be inevitably entangled, causing misclassific…

Cited by 0SourceScholar
2023

Iterative Proposal Refinement for Weakly-Supervised Video Grounding

CVPR 2023poster

Weakly-Supervised Video Grounding (WSVG) aims to localize events of interest in untrimmed videos with only video-level annotations. To date, most of the state-of-the-art WSVG methods follow a two-stage pipeline, i.e., firstly generating potential temporal proposals and then grounding with these prop…

2023

M3ST: Mix at Three Levels for Speech Translation

ICASSP 2023accepted

How to solve the data scarcity problem for end-to-end speech-to-text translation (ST)? It’s well known that data augmentation is an efficient method to improve performance for many tasks by enlarging the dataset. In this paper, we propose Mix at three levels for Speech Translation (M <sup xmlns:mml=…

Cited by 0SourceScholar
2023

ML-LMCL: Mutual Learning and Large-Margin Contrastive Learning for Improving ASR Robustness in Spoken Language Understanding

ACL 2023findings

Spoken language understanding (SLU) is a fundamental task in the task-oriented dialogue systems. However, the inevitable errors from automatic speech recognition (ASR) usually impair the understanding performance and lead to error propagation. Although there are some attempts to address this problem…

2023

MRRL: Modifying the Reference via Reinforcement Learning for Non-Autoregressive Joint Multiple Intent Detection and Slot Filling

EMNLP 2023long findings

With the rise of non-autoregressive approach, some non-autoregressive models for joint multiple intent detection and slot filling have obtained the promising inference speed. However, most existing SLU models (1) suffer from the multi-modality problem that leads to reference intents and slots may no…

Cited by 0SourceScholar
2023

MultiCapCLIP: Auto-Encoding Prompts for Zero-Shot Multilingual Visual Captioning

ACL 2023long

Supervised visual captioning models typically require a large scale of images or videos paired with descriptions in a specific language (i.e., the vision-caption pairs) for training. However, collecting and labeling large-scale datasets is time-consuming and expensive for many scenarios and language…

2023

Multimodal Prompt Learning for Product Title Generation with Extremely Limited Labels

ACL 2023findings

Generating an informative and attractive title for the product is a crucial task for e-commerce. Most existing works follow the standard multimodal natural language generation approaches, e.g., image captioning, and employ the large scale of human-labelled datasets to train desirable models. However…

Cited by 6SourcePDFScholar
2023

SSVMR: Saliency-Based Self-Training for Video-Music Retrieval

ICASSP 2023accepted

With the rise of short videos, the demand for selecting appropriate background music (BGM) for a video has increased significantly, video-music retrieval (VMR) task gradually draws much attention by research community. As other cross-modal learning tasks, existing VMR approaches usually attempt to m…

Cited by 0SourceScholar
2023

Towards Unified Spoken Language Understanding Decoding via Label-aware Compact Linguistics Representations

ACL 2023findings

Joint intent detection and slot filling models have shown promising success in recent years due to the high correlations between the two tasks. However, previous works independently decode the two tasks, which could result in misaligned predictions for both tasks. To address this shortcoming, we pro…

2023

Unify, Align and Refine: Multi-Level Semantic Alignment for Radiology Report Generation

ICCV 2023poster

Automatic radiology report generation has attracted enormous research interest due to its practical value in reducing the workload of radiologists. However, simultaneously establishing global correspondences between the image (e.g., Chest X-ray) and its related report and local alignments between im…

Cited by 42PDFScholar
2022

A Mutual Learning Framework for Few-Shot Sound Event Detection

ICASSP 2022accepted

Although prototypical network (ProtoNet) has proved to be an effective method for few-shot sound event detection, two problems still exist. Firstly, the small-scaled support set is insufficient so that the class prototypes may not represent the class center accurately. Secondly, the feature extracto…

Cited by 0SourceScholar
2022

A Transformer-based Threshold-Free Framework for Multi-Intent NLU

COLING 2022main

Multi-intent natural language understanding (NLU) has recently gained attention. It detects multiple intents in an utterance, which is better suited to real-world scenarios. However, the state-of-the-art joint NLU models mainly detect multiple intents on threshold-based strategy, resulting in one ma…

2022

Consistent Training and Decoding for End-to-End Speech Recognition Using Lattice-Free MMI

ICASSP 2022accepted

Recently, End-to-End (E2E) frameworks have achieved remarkable results on various Automatic Speech Recognition (ASR) tasks. However, Lattice-Free Maximum Mutual Information (LF-MMI), as one of the discriminative training criteria that show superior performance in hybrid ASR systems, is rarely adopte…

Cited by 0SourceScholar
2022

End-to-end Spoken Conversational Question Answering: Task, Dataset and Model

NAACL 2022findings

In spoken question answering, the systems are designed to answer questions from contiguous text spans within the related speech transcripts. However, the most natural way that human seek or test their knowledge is via human conversations. Therefore, we propose a new Spoken Conversational Question An…

Cited by 37SourcePDFScholar
2022

Improving Dual-Microphone Speech Enhancement by Learning Cross-Channel Features with Multi-Head Attention

ICASSP 2022accepted

Hand-crafted spatial features, such as inter-channel intensity difference (IID) and inter-channel phase difference (IPD), play a fundamental role in recent deep learning based dual-microphone speech enhancement (DMSE) systems. However, learning the mutual relationship between artificially designed s…

Cited by 0SourceScholar
2022

Learning Decoupling Features Through Orthogonality Regularization

ICASSP 2022accepted

Keyword spotting (KWS) and speaker verification (SV) are two important tasks in speech applications. Research shows that the state-of-art KWS and SV models are trained independently using different datasets since they expect to learn distinctive acoustic features. However, humans can distinguish lan…

Cited by 0SourceScholar
2022

LocVTP: Video-Text Pre-training for Temporal Localization

ECCV 2022poster

"Video-Text Pre-training (VTP) aims to learn transferable representations for various downstream tasks from large-scale web videos. To date, almost all existing VTP methods are limited to retrieval-based downstream tasks, e.g., video retrieval, whereas their transfer potentials on localization-based…

2022

Towards Joint Intent Detection and Slot Filling via Higher-order Attention

IJCAI 2022poster

Recently, attention-based models for joint intent detection and slot filling have achieved state-of-the-art performance. However, we think the conventional attention can only capture the first-order feature interaction between two tasks and is insufficient. To address this issue, we propose a unifie…

2022

Unsupervised Pre-Training for Temporal Action Localization Tasks

CVPR 2022poster

Unsupervised video representation learning has made remarkable achievements in recent years. However, most existing methods are designed and optimized for video classification. These pre-trained models can be sub-optimal for temporal localization tasks due to the inherent discrepancy between video-l…

Cited by 64PDFcodeScholar
2021

Adaptive Bi-Directional Attention: Exploring Multi-Granularity Representations for Machine Reading Comprehension

ICASSP 2021accepted

Recently, the attention-enhanced multi-layer encoder, such as Transformer, has been extensively studied in Machine Reading Comprehension (MRC). To predict the answer, it is common practice to employ a predictor to draw information only from the final encoder layer which generates the coarse-grained…

Cited by 0SourceScholar
2021

Audio-Oriented Multimodal Machine Comprehension via Dynamic Inter- and Intra-modality Attention

AAAI 2021technical

While Machine Comprehension (MC) has attracted extensive research interests in recent years, existing approaches mainly belong to the category of Machine Reading Comprehension task which mines textual inputs (paragraphs and questions) to predict the answers (choices or text spans). However, there ar…

Cited by 29SourcePDFScholar
2021

CoLA: Weakly-Supervised Temporal Action Localization With Snippet Contrastive Learning

CVPR 2021poster

Weakly-supervised temporal action localization (WS-TAL) aims to localize actions in untrimmed videos with only video-level labels. Most existing models follow the "localization by classification" procedure: locate temporal regions contributing most to the video-level classification. Generally, they…

Cited by 184PDFcodeScholar
2021

Contrastive Self-Supervised Learning for Text-Independent Speaker Verification

ICASSP 2021accepted

Current speaker verification models rely on supervised training with massive annotated data. But the collection of labeled utterances from multiple speakers is expensive and facing privacy issues. To open up an opportunity for utilizing massive unlabeled utterance data, our work exploits a contrasti…

Cited by 0SourceScholar
2021

Exploring and Distilling Posterior and Prior Knowledge for Radiology Report Generation

CVPR 2021poster

Automatically generating radiology reports can improve current clinical practice in diagnostic radiology. On one hand, it can relieve radiologists from the heavy burden of report writing; On the other hand, it can remind radiologists of abnormalities and avoid the misdiagnosis and missed diagnosis.…

Cited by 393PDFScholar
2021

FWB-Net: Front White Balance Network for Color Shift Correction in Single Image Dehazing Via Atmospheric Light Estimation

ICASSP 2021accepted

In recent years, single image dehazing deep models based on Atmospheric Scattering Model (ASM) have achieved remarkable results. But the dehazing outputs of those models suffer from color shift. Analyzing the ASM model shows that the atmospheric light factor (ALF) is set as a scalar which indicates…

Cited by 0SourceScholar
2021

MRD-Net: Multi-Modal Residual Knowledge Distillation for Spoken Question Answering

IJCAI 2021poster

Spoken question answering (SQA) has recently drawn considerable attention in the speech community. It requires systems to find correct answers from the given spoken passages simultaneously. The common SQA systems consist of the automatic speech recognition (ASR) module and text-based question answer…

Cited by 37SourcePDFScholar
2021

On Pursuit of Designing Multi-modal Transformer for Video Grounding

EMNLP 2021main

Video grounding aims to localize the temporal segment corresponding to a sentence query from an untrimmed video. Almost all existing video grounding methods fall into two frameworks: 1) Top-down model: It predefines a set of segment candidates and then conducts segment classification and regression.…

Cited by 91SourcePDFScholar
2021

RR-Net: Injecting Interactive Semantics in Human-Object Interaction Detection

IJCAI 2021poster

Human-Object Interaction (HOI) detection devotes to learn how humans interact with surrounding objects. Latest end-to-end HOI detectors are short of relation reasoning, which leads to inability to learn HOI-specific interactive semantics for predictions. In this paper, we therefore propose novel rel…

Cited by 4SourcePDFScholar
2021

SRF-Net: Selective Receptive Field Network for Anchor-Free Temporal Action Detection

ICASSP 2021accepted

Temporal action detection (TAD) is a challenging task which aims to temporally localize and recognize the human action in untrimmed videos. Current mainstream one-stage TAD approaches localize and classify action proposals relying on pre-defined anchors, where the location and scale for action insta…

Cited by 0SourceScholar
2021

Self-supervised Contrastive Cross-Modality Representation Learning for Spoken Question Answering

EMNLP 2021finding

Spoken question answering (SQA) requires fine-grained understanding of both spoken documents and questions for the optimal answer prediction. In this paper, we propose novel training schemes for spoken question answering with a self-supervised training stage and a contrastive representation learning…

Cited by 61SourcePDFScholar
2021

Sentiment Injected Iteratively Co-Interactive Network for Spoken Language Understanding

ICASSP 2021accepted

Spoken Language Understanding (SLU) is an essential part of the spoken dialogue system, which typically consists of intent detection (ID) and slot filling (SF) tasks. During the conversation, most utterances of people contain rich sentimental information, which is helpful for performing the ID and S…

Cited by 0SourceScholar
2020

Enhancing End-to-End Multi-Channel Speech Separation Via Spatial Feature Learning

ICASSP 2020accepted

Hand-crafted spatial features (e.g., inter-channel phase difference, IPD) play a fundamental role in recent deep learning based multi-channel speech separation (MCSS) methods. However, these manually designed spatial features are hard to incorporate into the end-to-end optimized MCSS framework. In t…

Cited by 0SourceScholar
2020

Prophet Attention: Predicting Attention with Future Attention

NeurIPS 2020poster

Recently, attention based models have been used extensively in many sequence-to-sequence learning systems. Especially for image captioning, the attention based models are expected to ground correct image regions with proper generated words. However, for each time step in the decoding process, the at…

Cited by 77SourcePDFScholar
2020

Rethinking Skip Connection with Layer Normalization

COLING 2020main

Skip connection is a widely-used technique to improve the performance and the convergence of deep neural networks, which is believed to relieve the difficulty in optimization due to non-linearity by propagating a linear component through the neural network layers. However, from another point of view…

Cited by 0SourcePDFScholar
2020

Semanticgan: Generative Adversarial Networks For Semantic Image To Photo-Realistic Image Translation

ICASSP 2020accepted

Generative Adversarial Networks (GANs) have shown remarkable success in Semantic label map to Photo-realistic image Translation (S2PT) task. However, the results of the state-of-the-art approaches are often limited to blurriness and artifacts, and still far from realistic, since these methods lack e…

Cited by 0SourceScholar
2020

Weakly Labelled Audio Tagging Via Convolutional Networks with Spatial and Channel-Wise Attention

ICASSP 2020accepted

Multiple instance learning (MIL) with convolutional neural networks (CNNs) has been proposed recently for weakly labelled audio tagging. However, features from the various CNN filtering channels and spatial regions are often treated equally, which may limit its performance in event prediction. In th…

Cited by 0SourceScholar
2019

Semantic Super-resolution for Extremely Low-resolution Vehicle License Plate

ICASSP 2019accepted

Vehicle license plate (VLP) super-resolution (SR) is of great demand in intelligent traffic systems. Super-Resolution for extremely low-resolution VLP remains challenging and the state-of-the-art SR methods hardly provide satisfying results for low-resolution (LR) VLPs. In this study, from a new per…

Cited by 0SourceScholar
2018

Inverse Atmoshperic Scattering Modeling with Convolutional Neural Networks for Single Image Dehazing

ICASSP 2018accepted

Single image dehazing is an ill-posed problem. Most existing works use the atmospheric scattering model (ASM) [1] and some natural priors to dehazing. Recently, DehazeNet [2] was developed using deep learning approach achieves the state-of-the-art results on many test hazy images, which motivates us…

Cited by 0SourceScholar
2018

Multi-Scale Object Detection with Feature Fusion and Region Objectness Network

ICASSP 2018accepted

Though tremendous progresses have been made in object detection due to the deep convolutional networks, one of the remaining challenges is the multi-scale object detection(MOD). To improve the performance of MOD task, we take Faster region-based CNN (Faster R-CNN) framework and work on two specific…

Cited by 0SourceScholar
2017

Example-based Visual Object Counting for complex background with a local low-rank constraint

ICASSP 2017accepted

Visual object counting (VOC) is important in many real-world applications. Our previous work approximated sparsity-constrain example-based VOC (ASE-VOC) works well with insufficient training data. It assumes that image patches share the similar local geometry with counterpart density maps, and then…

Cited by 0SourceScholar
2017

Robust speaker DOA estimation based on the inter-sensor data ratio model and binary mask estimation in the bispectrum domain

ICASSP 2017accepted

When noise is directional instead of diffuse, the majority of conventional direction of arrival (DOA) estimation techniques suffer from performance degradation because of mismatched noise models. In this paper, a novel robust DOA estimation algorithm is developed as an initial investigation into DOA…

Cited by 0SourceScholar
2015

A parametric modeling approach for wireless capsule endoscopy hazy image restoration

ICASSP 2015accepted

Wireless capsule endoscopy (WCE) is an innovative solution for gastrointestinal disease detection. The image quality of WCE is not satisfactory for medical applications since some of them are dark or hazy. For the purpose of improving WCE image quality, we take a new way to establish a parametric im…

Cited by 0SourceScholar