← Search

Ao Zhang

13 accepted papers

2026

PegasusFlow: Parallel Rolling-Denoising Score Sampling for Robot Diffusion Planner Flow Matching

ICRA 2026poster

Diffusion models offer powerful generative capabilities for robot trajectory planning, yet their practical deployment on robots is hindered by a critical bottleneck: reliance on imitation learning from expert demonstrations. This paradigm is often impractical for specialized robots where data is sca…

2025

DARNet: A Dual Attention Residual Network for Medical Image Classification

ICASSP 2025accepted

In the field of medical image analysis, accurate classification of images is crucial for diagnosing diseases and formulating treatment plans. Many studies have shown that global features and local features help reduce noise interference in medical images. Due to the fixed receptive field size of the…

Cited by 0SourceScholar
2025

MLSDET: Multi-LLM Statistical Deep Ensemble for Chinese AI-Generated Text Detection

ICASSP 2025accepted

With the rapid advancements in pre-trained large language models like ChatGPT, the surge of AI-generated text, particularly in Chinese, has presented significant challenges to existing detection systems due to its increasing realism and complexity. To address this, we introduce MLSDET: a groundbreak…

Cited by 0SourceScholar
2024

NExT-Chat: An LMM for Chat, Detection and Segmentation

ICML 2024poster

The development of large language models (LLMs) has greatly advanced the field of multimodal understanding, leading to the emergence of large multimodal models (LMMs). In order to enhance visual comprehension, recent studies have equipped LMMs with region-level understanding capabilities by represen…

2023

The NPU-ASLP System for Audio-Visual Speech Recognition in MISP 2022 Challenge

ICASSP 2023accepted

This paper describes our NPU-ASLP system for the Audio-Visual Diarization and Recognition (AVDR) task in the Multi-modal Information based Speech Processing (MISP) 2022 Challenge. Specifically, the weighted prediction error (WPE) and guided source separation (GSS) techniques are used to reduce rever…

Cited by 0SourceScholar
2023

VE-KWS: Visual Modality Enhanced End-to-End Keyword Spotting

ICASSP 2023accepted

The performance of the keyword spotting (KWS) system based on audio modality, commonly measured in false alarms and false rejects, degrades significantly under the far field and noisy conditions. Therefore, audio-visual keyword spotting, which leverages complementary relationships over multiple moda…

Cited by 0SourceScholar
2023

VPGTrans: Transfer Visual Prompt Generator across LLMs

NeurIPS 2023poster

Since developing a new multimodal LLM (MLLM) by pre-training on tremendous image-text pairs from scratch can be exceedingly resource-consuming, connecting an existing LLM with a comparatively lightweight visual prompt generator (VPG) becomes a feasible paradigm. However, further tuning the VPG compo…

2023

Visually Grounded Commonsense Knowledge Acquisition

AAAI 2023technical

Large-scale commonsense knowledge bases empower a broad range of AI applications, where the automatic extraction of commonsense knowledge (CKE) is a fundamental and challenging problem. CKE from text is known for suffering from the inherent sparsity and reporting bias of commonsense in text. Visual…

2022

Fine-Grained Scene Graph Generation with Data Transfer

ECCV 2022poster

"Scene graph generation (SGG) is designed to extract (subject, predicate, object) triplets in images. Recent works have made a steady progress on SGG, and provide useful tools for high-level vision and language understanding. However, due to the data distribution problems including long-tail distrib…

2022

PEVL: Position-enhanced Pre-training and Prompt Tuning for Vision-language Models

EMNLP 2022main

Vision-language pre-training (VLP) has shown impressive performance on a wide range of cross-modal tasks, where VLP models without reliance on object detectors are becoming the mainstream due to their superior computation efficiency and competitive performance. However, the removal of object detecto…

2022

Prompt Tuning for Discriminative Pre-trained Language Models

ACL 2022findings

Recent works have shown promising results of prompt tuning in stimulating pre-trained language models (PLMs) for natural language processing (NLP) tasks. However, to the best of our knowledge, existing works focus on prompt-tuning generative PLMs that are pre-trained to generate target tokens, such…

2021

Data Augmentation with Hierarchical SQL-to-Question Generation for Cross-domain Text-to-SQL Parsing

EMNLP 2021main

Data augmentation has attracted a lot of research attention in the deep learning era for its ability in alleviating data sparseness. The lack of labeled data for unseen evaluation databases is exactly the major challenge for cross-domain text-to-SQL parsing. Previous works either require human inter…

2021

Visual Distant Supervision for Scene Graph Generation

ICCV 2021poster

Scene graph generation aims to identify objects and their relations in images, providing structured image representations that can facilitate numerous applications in computer vision. However, scene graph models usually require supervised learning on large quantities of labeled data with intensive h…

Cited by 51PDFcodeScholar