← Search

Xiaodong He

71 accepted papers

2026

Biologically plausible heavy-tailed connectivity enhances generalizations on cognitive tasks in recurrent neural networks

ICML 2026poster

While heavy-tailed synaptic weight distributions are pervasive in biological neural networks, their computational role---particularly in relation to generalization---remains poorly understood. To address this, we develop a novel optimal-transport-based optimization algorithm that incorporates key bi…

Cited by 0SourceScholar
2026

Curvature-Constrained Vector Field for Motion Planning of Nonholonomic Robots

ICRA 2026poster

Vector fields handle nonholonomic motion planning as they provide reference orientation for robots. However, additionally incorporating curvature constraints becomes challenging, due to the interconnection between the design of the curvature-bounded vector field and the tracking controller under und…

2025

Comet: Dialog Context Fusion Mechanism for End-to-End Task-Oriented Dialog with Multi-task Learning

COLING 2025main

Existing end-to-end task-oriented dialog systems often encounter challenges arising from implicit information, coreference, and the presence of noisy and irrelevant data within the dialog context. These issues hinder the system’s ability to fully comprehend critical information and lead to inaccurat…

2025

HOIGen-1M: A Large-scale Dataset for Human-Object Interaction Video Generation

CVPR 2025poster

Text-to-video (T2V) generation has made tremendous progress in generating complicated scenes based on texts. However, human-object interaction (HOI) often cannot be precisely generated by current T2V models due to the lack of large-scale videos with accurate captions for HOI. To address this issue,…

2025

Scaling Down Text Encoders of Text-to-Image Diffusion Models

CVPR 2025poster

Text encoders in diffusion models have rapidly evolved, transitioning from CLIP to T5-XXL. Although this evolution has significantly enhanced the models' ability to understand complex prompts and generate text, it also leads to a substantial increase in the number of parameters. Despite T5 series en…

2025

UME: Upcycling Mixture-of-Experts for Scalable and Efficient Automatic Speech Recognition

ICASSP 2025accepted

Recent advancements in scaling up models have significantly improved performance in Automatic Speech Recognition (ASR) tasks. However, training large ASR models from scratch remains costly. To address this issue, we introduce UME, a novel method that efficiently Upcycles pretrained dense ASR checkpo…

Cited by 0SourceScholar
2024

Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorld

CVPR 2024poster

While large language models (LLMs) excel in a simulated world of texts they struggle to interact with the more realistic world without perceptions of other modalities such as visual or audio signals. Although vision-language models (VLMs) integrate LLM modules (1) aligned with static image features…

2024

MuEP: A Multimodal Benchmark for Embodied Planning with Foundation Models

IJCAI 2024poster

Foundation models have demonstrated significant emergent abilities, holding great promise for enhancing embodied agents' reasoning and planning capacities. However, the absence of a comprehensive benchmark for evaluating embodied agents with multimodal observations in complex environments remains a…

2024

POCE: Primal Policy Optimization with Conservative Estimation for Multi-constraint Offline Reinforcement Learning

CVPR 2024poster

Multi-constraint offline reinforcement learning (RL) promises to learn policies that satisfy both cumulative and state-wise costs from offline datasets. This arrangement provides an effective approach for the widespread application of RL in high-risk scenarios where both cumulative and state-wise co…

2024

TICOP: Time-Critical Coordinated Planning for Fixed-Wing UAVs in Unknown Unstructured Environments

RA-L 2024

Safe coordination of fixed-wing UAVs in unstructured environments poses challenges due to the intricate coupling of UAV cooperation, obstacle avoidance, and motion constraints. One task is time-critical coordination, which means that all UAVs can safely reach their destinations simultaneously. Exist

Cited by 3SourceScholar
2023

AUGUST: an Automatic Generation Understudy for Synthesizing Conversational Recommendation Datasets

ACL 2023findings

High-quality data is essential for conversational recommendation systems and serves as the cornerstone of the network architecture development and training strategy design. Existing works contribute heavy human efforts to manually labeling or designing and extending recommender dialogue templates. H…

2023

Dialog-Post: Multi-Level Self-Supervised Objectives and Hierarchical Model for Dialogue Post-Training

ACL 2023long

Dialogue representation and understanding aim to convert conversational inputs into embeddings and fulfill discriminative tasks. Compared with free-form text, dialogue has two important characteristics, hierarchical semantic structure and multi-facet attributes. Therefore, directly applying the pret…

2023

DiffusEmp: A Diffusion Model-Based Framework with Multi-Grained Control for Empathetic Response Generation

ACL 2023long

Empathy is a crucial factor in open-domain conversations, which naturally shows one’s caring and understanding to others. Though several methods have been proposed to generate empathetic responses, existing works often lead to monotonous empathy that refers to generic and safe expressions. In this p…

Cited by 20SourcePDFScholar
2023

Improving Disfluency Detection with Multi-Scale Self Attention and Contrastive Learning

ICASSP 2023accepted

Disfluency detection aims to recognize disfluencies in sentences. Existing works usually adopt a sequence labeling model to tackle this task. They also attempt to integrate into models the feature that the disfluencies are similar to the correct phrase, the so-called "rough copy". However, they heav…

Cited by 0SourceScholar
2023

MNER-QG: An End-to-End MRC Framework for Multimodal Named Entity Recognition with Query Grounding

AAAI 2023technical

Multimodal named entity recognition (MNER) is a critical step in information extraction, which aims to detect entity spans and classify them to corresponding entity types given a sentence-image pair. Existing methods either (1) obtain named entities with coarse-grained visual clues from attention me…

Cited by 53SourcePDFScholar
2023

Mars: Modeling Context & State Representations with Contrastive Learning for End-to-End Task-Oriented Dialog

ACL 2023findings

Traditional end-to-end task-oriented dialog systems first convert dialog context into belief state and action state before generating the system response. The system response performance is significantly affected by the quality of the belief state and action state. We first explore what dialog conte…

2023

MoNET: Tackle State Momentum via Noise-Enhanced Training for Dialogue State Tracking

ACL 2023findings

Dialogue state tracking (DST) aims to convert the dialogue history into dialogue states which consist of slot-value pairs. As condensed structural information memorizes all history information, the dialogue state in the previous turn is typically adopted as the input for predicting the current state…

Cited by 9SourcePDFScholar
2023

SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation

ICML 2023poster

Recently, the contrastive language-image pre-training, e.g., CLIP, has demonstrated promising results on various downstream tasks. The pre-trained model can capture enriched visual concepts for images by learning from a large scale of text-image data. However, transferring the learned visual knowled…

2023

UFO2: A Unified Pre-Training Framework for Online and Offline Speech Recognition

ICASSP 2023accepted

In this paper, we propose a Unified pre-training Framework for Online and Offline (UFO2) Automatic Speech Recognition (ASR), which 1) simplifies the two separate training workflows for online and offline modes into one process, and 2) improves the Word Error Rate (WER) performance with limited utter…

Cited by 0SourceScholar
2022

BORT: Back and Denoising Reconstruction for End-to-End Task-Oriented Dialog

NAACL 2022findings

A typical end-to-end task-oriented dialog system transfers context into dialog state, and upon which generates a response, which usually faces the problem of error propagation from both previously generated inaccurate dialog states and responses, especially in low-resource scenarios. To alleviate th…

2022

Building Robust Spoken Language Understanding by Cross Attention Between Phoneme Sequence and ASR Hypothesis

ICASSP 2022accepted

Building Spoken Language Understanding (SLU) robust to Automatic Speech Recognition (ASR) errors is an essential issue for various voice-enabled virtual assistants. Considering that most ASR errors are caused by phonetic confusion between similar-sounding expressions, intuitively, leveraging the pho…

Cited by 0SourceScholar
2022

Correctable-DST: Mitigating Historical Context Mismatch between Training and Inference for Improved Dialogue State Tracking

EMNLP 2022main

Recently proposed dialogue state tracking (DST) approaches predict the dialogue state of a target turn sequentially based on the previous dialogue state. During the training time, the ground-truth previous dialogue state is utilized as the historical context. However, only the previously predicted d…

Cited by 5SourcePDFScholar
2022

Don’t Take It Literally: An Edit-Invariant Sequence Loss for Text Generation

NAACL 2022long

Neural text generation models are typically trained by maximizing log-likelihood with the sequence cross entropy (CE) loss, which encourages an exact token-by-token match between a target sequence with a generated sequence. Such training objective is sub-optimal when the target sequence is not perfe…

2022

Few-Shot Table Understanding: A Benchmark Dataset and Pre-Training Baseline

COLING 2022main

Few-shot table understanding is a critical and challenging problem in real-world scenario as annotations over large amount of tables are usually costly. Pre-trained language models (PLMs), which have recently flourished on tabular data, have demonstrated their effectiveness for table understanding t…

2022

Fine- and Coarse-Granularity Hybrid Self-Attention for Efficient BERT

ACL 2022long

Transformer-based pre-trained models, such as BERT, have shown extraordinary success in achieving state-of-the-art results in many natural language processing applications. However, deploying these models can be prohibitively costly, as the standard self-attention mechanism of the Transformer suffer…

2022

Gated Multimodal Fusion with Contrastive Learning for Turn-Taking Prediction in Human-Robot Dialogue

ICASSP 2022accepted

Turn-taking, aiming to decide when the next speaker can start talking, is an essential component in building human-robot spoken dialogue systems. Previous studies indicate that multi-modal cues can facilitate this challenging task. However, due to the paucity of public multimodal datasets, current m…

Cited by 0SourceScholar
2022

JDDC 2.1: A Multimodal Chinese Dialogue Dataset with Joint Tasks of Query Rewriting, Response Generation, Discourse Parsing, and Summarization

EMNLP 2022main

The popularity of multimodal dialogue has stimulated the need for a new generation of dialogue agents with multimodal interactivity.When users communicate with customer service, they may express their requirements by means of text, images, or even videos. Visual information usually acts as discrimin…

2022

LUNA: Learning Slot-Turn Alignment for Dialogue State Tracking

NAACL 2022long

Dialogue state tracking (DST) aims to predict the current dialogue state given the dialogue history. Existing methods generally exploit the utterances of all dialogue turns to assign value for each slot. This could lead to suboptimal results due to the information introduced from irrelevant utteranc…

2022

Learning to Generate Poetic Chinese Landscape Painting with Calligraphy

IJCAI 2022poster

In this paper, we present a novel system (denoted as Polaca) to generate poetic Chinese landscape painting with calligraphy. Unlike previous single image-to-image painting generation, Polaca takes the classic poetry as input and outputs the artistic landscape painting image with the corresponding ca…

Cited by 10SourcePDFScholar
2022

MuGER2: Multi-Granularity Evidence Retrieval and Reasoning for Hybrid Question Answering

EMNLP 2022finding

Hybrid question answering (HQA) aims to answer questions over heterogeneous data, including tables and passages linked to table cells. The heterogeneous data can provide different granularity evidence to HQA models, e.t., column, row, cell, and link. Conventional HQA models usually retrieve coarse-…

2022

OPERA: Operation-Pivoted Discrete Reasoning over Text

NAACL 2022long

Machine reading comprehension (MRC) that requires discrete reasoning involving symbolic operations, e.g., addition, sorting, and counting, is a challenging task. According to this nature, semantic parsing-based methods predict interpretable but complex logical forms. However, logical form generation…

2022

P3LM: Probabilistically Permuted Prophet Language Modeling for Generative Pre-Training

EMNLP 2022finding

Conventional autoregressive left-to-right (L2R) sequence generation faces two issues during decoding: limited to unidirectional target sequence modeling, and constrained on strong local dependencies.To address the aforementioned problem, we propose P3LM, a probabilistically permuted prophet language…

Cited by 0SourcePDFScholar
2022

PRINCE: Prefix-Masked Decoding for Knowledge Enhanced Sequence-to-Sequence Pre-Training

EMNLP 2022main

Pre-trained Language Models (PLMs) have shown effectiveness in various Natural Language Processing (NLP) tasks. Denoising autoencoder is one of the most successful pre-training frameworks, learning to recompose the original text given a noise-corrupted one. The existing studies mainly focus on injec…

2022

Tracking Satisfaction States for Customer Satisfaction Prediction in E-commerce Service Chatbots

COLING 2022main

Due to the increasing use of service chatbots in E-commerce platforms in recent years, customer satisfaction prediction (CSP) is gaining more and more attention. CSP is dedicated to evaluating subjective customer satisfaction in conversational service and thus helps improve customer service experien…

2022

UniRPG: Unified Discrete Reasoning over Table and Text as Program Generation

EMNLP 2022main

Question answering requiring discrete reasoning, e.g., arithmetic computing, comparison, and counting, over knowledge is a challenging task.In this paper, we propose UniRPG, a semantic-parsing-based approach advanced in interpretability and scalability, to perform Unified discrete Reasoning over het…

2021

Conversational Query Rewriting with Self-Supervised Learning

ICASSP 2021accepted

Context modeling plays a critical role in building multi-turn dialogue systems. Conversational Query Rewriting (CQR) aims to simplify the multi-turn dialogue modeling into a single-turn problem by explicitly rewriting the conversational query into a self-contained utterance. However, existing approa…

Cited by 0SourceScholar
2021

Dian: Duration Informed Auto-Regressive Network for Voice Cloning

ICASSP 2021accepted

In this paper, we propose a novel end-to-end speech synthesis approach, Duration Informed Auto-regressive Network (DIAN), which consists of an acoustic model and a separate duration model. Un-like other auto-regressive TTS methods, the duration information of phonemes is provided as part of the inpu…

Cited by 0SourceScholar
2021

Graph Ensemble Learning over Multiple Dependency Trees for Aspect-level Sentiment Classification

NAACL 2021long

Recent work on aspect-level sentiment classification has demonstrated the efficacy of incorporating syntactic structures such as dependency trees with graph neural networks (GNN), but these approaches are usually vulnerable to parsing errors. To better leverage syntactic information in the face of u…

Cited by 65SourcePDFScholar
2021

Improving Prosody Modelling with Cross-Utterance Bert Embeddings for End-to-End Speech Synthesis

ICASSP 2021accepted

Although speech prosody is related to the linguistic information up to the discourse structure, most text-to-speech (TTS) systems only take into account the information within each sentence. This makes it challenging when converting a paragraph of text into natural and expressive speech. In this pap…

Cited by 0SourceScholar
2021

K-PLUG: Knowledge-injected Pre-trained Language Model for Natural Language Understanding and Generation in E-Commerce

EMNLP 2021finding

Existing pre-trained language models (PLMs) have demonstrated the effectiveness of self-supervised learning for a broad range of natural language processing (NLP) tasks. However, most of them are not explicitly aware of domain-specific knowledge, which is essential for downstream tasks in many domai…

2021

Learn to Copy from the Copying History: Correlational Copy Network for Abstractive Summarization

EMNLP 2021main

The copying mechanism has had considerable success in abstractive summarization, facilitating models to directly copy words from the input text to the output summary. Existing works mostly employ encoder-decoder attention, which applies copying at each time step independently of the former ones. How…

2021

Neural Kalman Filtering for Speech Enhancement

ICASSP 2021accepted

Conventional learning-based speech enhancement methods usually utilize existing building blocks to design the deep neural networks (DNNs), while how to effectively integrate the statistical signal processing based schemes, which are expert-knowledge driven and could ameliorate the over-fitting probl…

Cited by 0SourceScholar
2021

RoR: Read-over-Read for Long Document Machine Reading Comprehension

EMNLP 2021finding

Transformer-based pre-trained models, such as BERT, have achieved remarkable results on machine reading comprehension. However, due to the constraint of encoding length (e.g., 512 WordPiece tokens), a long document is usually split into multiple chunks that are independently read. It results in the…

2021

SGG: Learning to Select, Guide, and Generate for Keyphrase Generation

NAACL 2021long

Keyphrases, that concisely summarize the high-level topics discussed in a document, can be categorized into present keyphrase which explicitly appears in the source text and absent keyphrase which does not match any contiguous subsequence but is highly semantically related to the source. Most existi…

2020

Group Contextual Encoding for 3D Point Clouds

NeurIPS 2020poster

Global context is crucial for 3D point cloud scene understanding tasks. In this work, we extended the contextual encoding layer that was originally designed for 2D tasks to 3D Point Cloud scenarios. The encoding layer learns a set of code words in the feature space of the 3D point cloud to characte…

2020

Learning to Decouple Relations: Few-Shot Relation Classification with Entity-Guided Attention and Confusion-Aware Training

COLING 2020main

This paper aims to enhance the few-shot relation classification especially for sentences that jointly describe multiple relations. Due to the fact that some relations usually keep high co-occurrence in the same context, previous few-shot relation classifiers struggle to distinguish them with few ann…

Cited by 49SourcePDFScholar
2020

Multimodal Sentence Summarization via Multimodal Selective Encoding

COLING 2020main

This paper studies the problem of generating a summary for a given sentence-image pair. Existing multimodal sequence-to-sequence approaches mainly focus on enhancing the decoder by visual signals, while ignoring that the image can improve the ability of the encoder to identify highlights of a news e…

Cited by 41SourcePDFScholar
2020

On the Faithfulness for E-commerce Product Summarization

COLING 2020main

In this work, we present a model to generate e-commerce product summaries. The consistency between the generated summary and the product attributes is an essential criterion for the ecommerce product summarization task. To enhance the consistency, first, we encode the product attribute table to guid…

2019

Aligning Visual Regions and Textual Concepts for Semantic-Grounded Image Representations

NeurIPS 2019poster

In vision-and-language grounding problems, fine-grained representations of the image are considered to be of paramount importance. Most of the current systems incorporate visual features and textual concepts as a sketch of an image. However, plainly inferred representations are usually undesirable i…

2019

Collaborative Learning of Semi-Supervised Segmentation and Classification for Medical Images

CVPR 2019poster

Medical image analysis has two important research areas: disease grading and fine-grained lesion segmentation. Although the former problem often relies on the latter, the two are usually studied separately. Disease severity grading can be treated as a classification problem, which only requires imag…

Cited by 327PDFScholar
2019

Deep Speaker Embedding Learning with Multi-level Pooling for Text-independent Speaker Verification

ICASSP 2019accepted

This paper aims to improve the widely used deep speaker embedding x-vector model. We propose the following improvements: (1) a hybrid neural network structure using both time delay neural network (TDNN) and long short-term memory neural networks (LSTM) to generate complementary speaker information a…

Cited by 0SourceScholar
2019

Object-Driven Text-To-Image Synthesis via Adversarial Training

CVPR 2019poster

In this paper, we propose Object-driven Attentive Generative Adversarial Newtorks (Obj-GANs) that allow attention-driven, multi-stage refinement for synthesizing complex images from text descriptions. With a novel object-driven attentive generative network, the Obj-GAN can synthesize salient objects…

Cited by 386PDFScholar
2018

AttnGAN: Fine-Grained Text to Image Generation With Attentional Generative Adversarial Networks

CVPR 2018poster

In this paper, we propose an Attentional Generative Adversarial Network (AttnGAN) that allows attention-driven, multi-stage refinement for fine-grained text-to-image generation. With a novel attentional generative network, the AttnGAN can synthesize fine-grained details at different sub-regions of…

2018

Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering

CVPR 2018poster

Top-down visual attention mechanisms have been used extensively in image captioning and visual question answering (VQA) to enable deeper image understanding through fine-grained analysis and even multiple steps of reasoning. In this work, we propose a combined bottom-up and top-down attention mechan…

2018

CleanNet: Transfer Learning for Scalable Image Classifier Training With Label Noise

CVPR 2018poster

In this paper, we study the problem of learning image classification models with label noise. Existing approaches depending on human supervision are generally not scalable as manually identifying correct or incorrect labels is time-consuming, whereas approaches not relying on human supervision are s…

2018

Constrained Convolutional-Recurrent Networks to Improve Speech Quality with Low Impact on Recognition Accuracy

ICASSP 2018accepted

For a speech-enhancement algorithm, it is highly desirable to simultaneously improve perceptual quality and recognition rate. Thanks to computational costs and model complexities, it is challenging to train a model that effectively optimizes both metrics at the same time. In this paper, we propose a…

Cited by 10SourceScholar
2018

On the Discrimination-Generalization Tradeoff in GANs

ICLR 2018poster

Generative adversarial training can be generally understood as minimizing certain moment matching loss defined by a set of discriminator functions, typically neural networks. The discriminator set should be large enough to be able to uniquely identify the true distribution (discriminative), and als…

Cited by 0SourcePDFScholar
2018

Stacked Cross Attention for Image-Text Matching

ECCV 2018poster

In this paper, we study the problem of image-text matching. Inferring the latent semantic alignment between objects or other salient stuff (e.g. snow, sky, lawn) and the corresponding words in sentences allows to capture fine-grained interplay between vision and language, and makes image-text matchi…

2018

Tips and Tricks for Visual Question Answering: Learnings From the 2017 Challenge

CVPR 2018poster

This paper presents a state-of-the-art model for visual question answering (VQA), which won the first place in the 2017 VQA Challenge. VQA is a task of significant importance for research in artificial intelligence, given its multimodal nature, clear evaluation protocol, and potential real-world app…

Cited by 500SourcePDFScholar
2017

Adversarial Ranking for Language Generation

NeurIPS 2017poster

Generative adversarial networks (GANs) have great successes on synthesizing data. However, the existing GANs restrict the discriminator to be a binary classifier, and thus limit their learning capacity for tasks that need to synthesize output with rich structures such as natural language description…

2017

Character-level deep conflation for business data analytics

ICASSP 2017accepted

Connecting different text attributes associated with the same entity (conflation) is important in business data analytics since it could help merge two different tables in a database to provide a more comprehensive profile of an entity. However, the conflation task is challenging because two text st…

Cited by 0SourceScholar
2017

Deep Learning With Low Precision by Half-Wave Gaussian Quantization

CVPR 2017spotlight

The problem of quantizing the activations of a deep neural network is considered. An examination of the popular binary quantization approach shows that this consists of approximating a classical non-linearity, the hyperbolic tangent, by two functions: a piecewise constant sign function, which is use…

Cited by 632PDFcodeScholar
2017

Semantic Compositional Networks for Visual Captioning

CVPR 2017spotlight

A Semantic Compositional Network (SCN) is developed for image captioning, in which semantic concepts (i.e., tags) are detected from the image, and the probability of each tag is used to compose the parameters in a long short-term memory (LSTM) network. The SCN extends each weight matrix of the LSTM…

Cited by 561PDFcodeScholar
2016

Interpreting the prediction process of a deep network constructed from supervised topic models

ICASSP 2016accepted

In this paper, we propose an approach to interpret the prediction process of the BP-sLDA model, which is a supervised Latent Dirichlet Allocation model trained by Back Propagation over a deep architecture. The model is shown to achieve state-of-the-art prediction performance on several large-scale t…

Cited by 0SourceScholar
2016

Zero-shot learning of intent embeddings for expansion by convolutional deep structured semantic models

ICASSP 2016accepted

The recent surge of intelligent personal assistants motivates spoken language understanding of dialogue systems. However, the domain constraint along with the inflexible intent schema remains a big issue. This paper focuses on the task of intent expansion, which helps remove the domain limit and mak…

Cited by 0SourceScholar
2015

End-to-end Learning of LDA by Mirror-Descent Back Propagation over a Deep Architecture

NeurIPS 2015poster

We develop a fully discriminative learning approach for supervised Latent Dirichlet Allocation (LDA) model using Back Propagation (i.e., BP-sLDA), which maximizes the posterior probability of the prediction variable given the input document. Different from traditional variational learning or Gibbs s…

2015

From Captions to Visual Concepts and Back

CVPR 2015poster

This paper presents a novel approach for automatically generating image descriptions: visual detectors, language models, and multimodal similarity models learnt directly from a dataset of image captions. We use multiple instance learning to train visual detectors for words that commonly occur in cap…