← Search

Zongheng Tang

5 accepted papers

2026

AerialVLA: A Vision-Language-Action Model for Aerial Navigation with Online Dialogue

AAAI 2026technical

Visual Dialogue Navigation (VDN) aims to enable agents to reach target locations through dialogue with humans. The integration of VDN into Unmanned Aerial Vehicle (UAV) systems enhances human-machine interaction by enabling intuitive, hands-free operation, thereby unlocking vast applications. Howeve

Cited by 5SourcePDFScholar
2025

CoST: Efficient Collaborative Perception From Unified Spatiotemporal Perspective

ICCV 2025poster

Collaborative perception shares information among different agents and helps solving problems that individual agents may face, e.g., occlusions and small sensing range. Prior methods usually separate the multi-agent fusion and multi-time fusion into two consecutive steps. In contrast, this paper pro…

2025

Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation

AAAI 2025technical

In this paper, we propose an Audio-Language-Referenced SAM 2 (AL-Ref-SAM 2) pipeline to explore the training-free paradigm for audio and language-referenced video object segmentation, namely AVS and RVOS tasks. The intuitive solution leverages GroundingDINO to identify the target object from a singl…

2023

DETR With Additional Global Aggregation for Cross-Domain Weakly Supervised Object Detection

CVPR 2023poster

This paper presents a DETR-based method for cross-domain weakly supervised object detection (CDWSOD), aiming at adapting the detector from source to target domain through weak supervision. We think DETR has strong potential for CDWSOD due to an insight: the encoder and the decoder in DETR are both b…

Cited by 17SourcePDFScholar