← Search

Jirong Zha

4 accepted papers

2026

AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning

AAAI 2026technical

Multimodal Large Language Models (MLLMs) have shown promise in single-agent vision tasks, yet benchmarks for evaluating multi-agent collaborative perception remain scarce. This gap is critical, as multi-drone systems provide enhanced coverage, robustness, and collaboration compared to single-sensor

Cited by 0SourcePDFScholar
2025

How to Enable LLM with 3D Capacity? A Survey of Spatial Reasoning in LLM

IJCAI 2025

3D spatial understanding is essential in real-world applications such as robotics, autonomous vehicles, virtual reality, and medical imaging. Recently, Large Language Models (LLMs), having demonstrated remarkable success across various domains, have been leveraged to enhance 3D understanding tasks,

Cited by 0SourcePDFScholar
2025

UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces

ACL 2025long

Large multimodal models exhibit remarkable intelligence, yet their embodied cognitive abilities during motion in open-ended urban aerial spaces remain to be explored. We introduce a benchmark to evaluate whether video-large language models (Video-LLMs) can naturally process continuous first-person v…

Cited by 0SourcePDFScholar