← Search

Jinyu Chen

10 accepted papers

2026

AerialVLA: A Vision-Language-Action Model for Aerial Navigation with Online Dialogue

AAAI 2026technical

Visual Dialogue Navigation (VDN) aims to enable agents to reach target locations through dialogue with humans. The integration of VDN into Unmanned Aerial Vehicle (UAV) systems enhances human-machine interaction by enabling intuitive, hands-free operation, thereby unlocking vast applications. Howeve

Cited by 0SourcePDFScholar
2025

CoST: Efficient Collaborative Perception From Unified Spatiotemporal Perspective

ICCV 2025poster

Collaborative perception shares information among different agents and helps solving problems that individual agents may face, e.g., occlusions and small sensing range. Prior methods usually separate the multi-agent fusion and multi-time fusion into two consecutive steps. In contrast, this paper pro…

2025

LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

CVPR 2025poster

Recent advancements in multimodal large language models (MLLMs) have shown promising results, yet existing approaches struggle to effectively handle both temporal and spatial localization simultaneously. This challenge stems from two key issues: first, incorporating spatial-temporal localization int…

2025

Towards Realistic UAV Vision-Language Navigation: Platform, Benchmark, and Methodology

ICLR 2025poster

Developing agents capable of navigating to a target location based on language instructions and visual information, known as vision-language navigation (VLN), has attracted widespread interest. Most research has focused on ground-based agents, while UAV-based VLN remains relatively underexplored. Re…

Cited by 12SourcePDFScholar
2024

Controllable Navigation Instruction Generation with Chain of Thought Prompting

ECCV 2024poster

"Instruction generation is a vital and multidisciplinary research area with broad applications. Existing instruction generation models are limited to generating instructions in a single style from a particular dataset, and the style and content of generated instructions cannot be controlled. Moreove…

2024

Overcome Modal Bias in Multi-modal Federated Learning via Balanced Modality Selection

ECCV 2024poster

"Selecting proper clients to participate in each federated learning (FL) round is critical to effectively harness a broad range of distributed data. Existing client selection methods simply consider the mining of distributed uni-modal data, yet, their effectiveness may diminish in multi-modal FL (MF…

2023

FedDWA: Personalized Federated Learning with Dynamic Weight Adjustment

IJCAI 2023poster

Different from conventional federated learning, personalized federated learning (PFL) is able to train a customized model for each individual client according to its unique requirement. The mainstream approach is to adopt a kind of weighted aggregation method to generate personalized models, in whic…

2023

Omnidirectional Information Gathering for Knowledge Transfer-Based Audio-Visual Navigation

ICCV 2023poster

Audio-visual navigation is an audio-targeted wayfinding task where a robot agent is entailed to travel a never-before-seen 3D environment towards the sounding source. In this article, we present ORAN, an omnidirectional audio-visual navigator based on cross-task navigation skill transfer. In particu…

Cited by 8PDFcodeScholar
2022

Reinforced Structured State-Evolution for Vision-Language Navigation

CVPR 2022poster

Vision-and-language Navigation (VLN) task requires an embodied agent to navigate to a remote location following a natural language instruction. Previous methods usually adopt a sequence model (e.g., Transformer and LSTM) as the navigator. In such a paradigm, the sequence model predicts action at eac…

Cited by 47PDFcodeScholar
2021

Room-and-Object Aware Knowledge Reasoning for Remote Embodied Referring Expression

CVPR 2021poster

The Remote Embodied Referring Expression (REVERIE) is a recently raised task that requires an agent to navigate to and localise a referred remote object according to a high-level language instruction. Different from related VLN tasks, the key to REVERIE is to conduct goal-oriented exploration instea…

Cited by 91PDFcodeScholar