← Search

Dawei Zhang

18 accepted papers

2026

Exploiting All Mamba Fusion for Efficient RGB-D Tracking

AAAI 2026technical

Despite the progress made through deep learning, existing Visual Object Tracking (VOT) frameworks struggle with real-world challenges. Recent approaches incorporate additional modalities like Depth, Thermal Infrared, and Language to enhance the robustness of VOT, particularly with the improvement of

Cited by 0SourcePDFScholar
2026

IGIANet: Illumination Guided Implicit Alignment Network for Infrared–Visible UAV Detection

AAAI 2026technical

Visible-Infrared (RGB-IR) Unmanned Aerial Vehicle (UAV) object detection integrates complementary cues from visible and infrared sensors, offering broad application potential. However, due to sensor parallax, it still faces the challenge of weak spatial misalignment, which significantly limits its p

Cited by 0SourcePDFScholar
2023

Dynamic TF-TDNN: Dynamic Time Delay Neural Network Based on Temporal-Frequency Attention for Dialect Recognition

ICASSP 2023accepted

Dialect recognition aims to recognize dialect categories in utterances, which has been applied in many audio applications. Recently, various Time Delayed Neural Network (TDNN) based AI models are proposed to solve dialect recognition problems, such as D-TDNN, DMC-TDNN, and ECAPA-TDNN, however, most…

Cited by 0SourceScholar
2022

KSAM: Infusing Multi-Source Knowledge into Dialogue Generation via Knowledge Source Aware Multi-Head Decoding

ACL 2022findings

Knowledge-enhanced methods have bridged the gap between human beings and machines in generating dialogue responses. However, most previous works solely seek knowledge from a single source, and thus they often fail to obtain available knowledge because of the insufficient coverage of a single knowled…

Cited by 6SourcePDFScholar
2022

MOBA-E2C: Generating MOBA Game Commentaries via Capturing Highlight Events from the Meta-Data

EMNLP 2022finding

MOBA (Multiplayer Online Battle Arena) games such as Dota2 are currently one of the most popular e-sports gaming genres. Following professional commentaries is a great way to understand and enjoy a MOBA game. However, massive game competitions lack commentaries because of the shortage of professiona…

2022

Section-Aware Commonsense Knowledge-Grounded Dialogue Generation with Pre-trained Language Model

COLING 2022main

In knowledge-grounded dialogue generation, pre-trained language models (PLMs) can be expected to deepen the fusing of dialogue context and knowledge because of their superior ability of semantic understanding. Unlike adopting the plain text knowledge, it is thorny to leverage the structural commonse…

2021

Haptic Feedback Improves Human-Robot Agreement and User Satisfaction in Shared-Autonomy Teleoperation

ICRA 2021poster

Shared autonomy teleoperation can guarantee safety, but does so by reducing the human operator’s control authority, which can lead to reduced levels of human-robot agreement and user satisfaction. This paper presents a novel haptic shared autonomy teleoperation paradigm that uses haptic feedback to…

Cited by 26SourceScholar
2021

Knowledge-Aware Dialogue Generation via Hierarchical Infobox Accessing and Infobox-Dialogue Interaction Graph Network

IJCAI 2021poster

Due to limited knowledge carried by queries, traditional dialogue systems often face the dilemma of generating boring responses, leading to poor user experience. To alleviate this issue, this paper proposes a novel infobox knowledge-aware dialogue generation approach, HITA-Graph, with three unique f…

2021

More is Better: Enhancing Open-Domain Dialogue Generation via Multi-Source Heterogeneous Knowledge

EMNLP 2021main

Despite achieving remarkable performance, previous knowledge-enhanced works usually only use a single-source homogeneous knowledge base of limited knowledge coverage. Thus, they often degenerate into traditional methods because not all dialogues can be linked with knowledge entries. This paper propo…

2021

Multi Path Training Framework for Data-Driven Open-Domain Conversation System

ICASSP 2021accepted

Nowadays, web data is often used to train a dialogue system. However, noises in web data can disturb the training process, as well as can impact the performance. Consequently, dialogue models tend to be brittle when receiving noisy inputs during the inference. This paper proposes a novel framework,…

Cited by 0SourceScholar
2021

Visual Tracking via Hierarchical Deep Reinforcement Learning

AAAI 2021technical

Visual tracking has achieved great progress due to numerous different algorithms. However, deep trackers based on classification or Siamese network still have their specific limitations. In this work, we show how to teach machines to track a generic object in videos like humans, who can use a few se…

Cited by 35SourcePDFScholar
2020

TopicKA: Generating Commonsense Knowledge-Aware Dialogue Responses Towards the Recommended Topic Fact

IJCAI 2020poster

Insufficient semantic understanding of dialogue always leads to the appearance of generic responses, in generative dialogue systems. Recently, high-quality knowledge bases have been introduced to enhance dialogue understanding, as well as to reduce the prevalence of boring responses. Although such k…

2017

A novel pitch extraction based on jointly trained deep BLSTM Recurrent Neural Networks with bottleneck features

ICASSP 2017accepted

Pitch is an important characteristic of speech and is useful for many applications. However, it is still challenging to estimate pitch in strong noise. In this paper, we propose a joint training approach to determinate pitch. First, a Bidirectional Long Short-Term Memory Recurrent Neural Networks (B…

Cited by 0SourceScholar
2016

Extraction of tongue contour in real-time magnetic resonance imaging sequences

ICASSP 2016accepted

Real-time magnetic resonance imaging (rtMRI) is becoming a practical tool in speech production research and language pathology observation. It is still a challenge to extract the tongue contour accurately in rtMRI sequences, since tongue is a soft tissue and often touches other organs such as lips a…

Cited by 0SourceScholar