← Search

Chen Gao

55 accepted papers

2026

AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning

AAAI 2026technical

Multimodal Large Language Models (MLLMs) have shown promise in single-agent vision tasks, yet benchmarks for evaluating multi-agent collaborative perception remain scarce. This gap is critical, as multi-drone systems provide enhanced coverage, robustness, and collaboration compared to single-sensor

Cited by 0SourcePDFScholar
2026

Contrastive Weak-to-Strong Generalization

ICML 2026poster

Weak-to-strong generalization provides a promising paradigm for scaling large language models (LLMs) by training stronger models on samples from aligned weaker ones, without requiring human feedback or explicit reward modeling. However, its robustness and generalization are hindered by the noise and…

Cited by 0SourceScholar
2026

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation

ICML 2026poster

While Vision-Language-Action (VLA) models offer broad general capabilities, deploying them on specific hardware requires real-world adaptation to bridge the embodiment gap. Since robot demonstrations are costly, this adaptation must often occur under a strict data budget. In this work, we identify a…

Cited by 0SourceScholar
2026

LF-BVN: Blind-View Network for Self-Supervised Light Field Denoising

CVPR 2026

Recent advances in learning-based Light Field (LF) image denoising have achieved impressive results. However, these methods rely heavily on large-scale noisy-clean image pairs and often fail to generalize to unseen or complex noise.In this work, we observe that the inherent multi-view consistency of

Cited by 0SourcecodeScholar
2026

Semantic Audio-Visual Navigation in Continuous Environments

CVPR 2026

Audio-visual navigation enables embodied agents to navigate toward sound-emitting targets by leveraging both auditory and visual cues. However, most existing approaches rely on precomputed room impulse responses (RIRs) for binaural audio rendering, restricting agents to discrete grid positions and l

Cited by 0SourcecodeScholar
2026

SmartAgent: Chain-of-User-Thought for Embodied Personalized Agent in Cyber World

AAAI 2026technical

Recent advances in embodied agents with multimodal perception and reasoning capabilities based on large vision-language models (LVLMs), excel in autonomously interacting either real or cyber worlds, helping people make intelligent decisions in complex environments. However, the current works are nor

Cited by 0SourcePDFScholar
2026

Towards Autonomous UAV Visual Object Search in City Space: Benchmark and Agentic Methodology

AAAI 2026technical

Aerial Visual Object Search (AVOS) tasks in urban environments require Unmanned Aerial Vehicles (UAVs) to autonomously search for and identify target objects based on visual inputs without external guidance. Existing approaches struggle in complex urban environments due to redundant semantic process

Cited by 0SourcePDFScholar
2026

Training-Free Sparse Attention for Fast Video Generation via Offline Layer-Wise Sparsity Profiling and Online Bidirectional Co-Clustering

ICML 2026poster

Diffusion Transformers (DiTs) achieve strong video generation quality but suffer from high inference cost due to dense 3D attention, leading to the development of sparse attention technologies to improve efficiency. However, existing training-free sparse attention methods in video generation still f…

Cited by 0SourceScholar
2026

iWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation Framework

ICML 2026poster

Achieving Artificial General Intelligence (AGI) requires agents that learn and interact adaptively, with interactive world models providing scalable environments for perception, reasoning, and action. Yet current research still lacks large-scale datasets and unified benchmarks to evaluate their phys…

Cited by 0SourceScholar
2025

Analyzing and Modeling LLM Response Lengths with Extreme Value Theory: Anchoring Effects and Hybrid Distributions

EMNLP 2025

We present a statistical framework for modeling and controlling large language model (LLM) response lengths using extreme value theory. Analyzing 14,301 GPT-4o responses across temperature and prompting conditions, with cross-validation on Qwen and DeepSeek architectures, we demonstrate that verbosi

Cited by 0SourcePDFScholar
2025

Balanced Token Pruning: Accelerating Vision Language Models Beyond Local Optimization

NeurIPS 2025poster

Large Vision-Language Models (LVLMs) have shown impressive performance across multi-modal tasks by encoding images into thousands of tokens. However, the large number of image tokens results in significant computational overhead, and the use of dynamic high-resolution inputs further increases this b…

Cited by 0SourcecodeScholar
2025

CityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City Space

EMNLP 2025

Embodied Question Answering (EQA) has primarily focused on indoor environments, leaving the complexities of urban settings—spanning environment, action, and perception—largely unexplored. To bridge this gap, we introduce CityEQA, a new task where an embodied agent answers open-vocabulary questions t

2025

CityNavAgent: Aerial Vision-and-Language Navigation with Hierarchical Semantic Planning and Global Memory

ACL 2025long

Aerial vision-and-language navigation (VLN) — requiring drones to interpret natural language instructions and navigate complex urban environments — emerges as a critical embodied AI challenge that bridges human-robot interaction, 3D spatial reasoning, and real-world deployment. Although existing gro…

2025

Defining and Evaluating Visual Language Models’ Basic Spatial Abilities: A Perspective from Psychometrics

ACL 2025long

The Theory of Multiple Intelligences underscores the hierarchical nature of cognitive capabilities. To advance Spatial Artificial Intelligence, we pioneer a psychometric framework defining five Basic Spatial Abilities (BSAs) in Visual Language Models (VLMs): Spatial Perception, Spatial Relation, Spa…

Cited by 0SourcePDFScholar
2025

Epipolar Consistent Attention Aggregation Network for Unsupervised Light Field Disparity Estimation

ICCV 2025poster

Disparity estimation is an essential step in processing and analyzing Light Field (LF) images. Recent methods construct the cost volume to exploit the correspondence of the LFs over the preset maximum disparity, limiting them to process the large parallax scenes. Different from constructing cost vol…

Cited by 0SourcePDFScholar
2025

Exploring View Consistency for Scene-Adaptive Low-Light Light Field Image Enhancement

ICCV 2025poster

Light Field (LF) images captured under low illumination conditions typically exhibit low quality. Recent learning-based methods for low-light LF enhancement are generally tailored to specific illumination inputs, limiting their performance in real-world scenes. Moreover, how to maintain the inherent…

Cited by 0SourcePDFScholar
2025

How to Enable LLM with 3D Capacity? A Survey of Spatial Reasoning in LLM

IJCAI 2025

3D spatial understanding is essential in real-world applications such as robotics, autonomous vehicles, virtual reality, and medical imaging. Recently, Large Language Models (LLMs), having demonstrated remarkable success across various domains, have been leveraged to enhance 3D understanding tasks,

Cited by 0SourcePDFScholar
2025

Iterative Sparse Attention for Long-sequence Recommendation

AAAI 2025technical

Longer historical behaviors often improve recommendation accuracy but bring efficient problems. As sequences get longer, the following two main challenges have not been addressed: (1) efficient modeling under increasing sequence length and (2) interest drifting within historical items. In this paper…

2025

MIA-Tuner: Adapting Large Language Models as Pre-training Text Detector

AAAI 2025technical

The increasing parameters and expansive dataset of large lan- guage models (LLMs) highlight the urgent demand for a technical solution to audit the underlying privacy risks and copyright issues associated with LLMs. Existing studies have partially addressed this need through an exploration of the pr…

2025

Open-Set Living Need Prediction with Large Language Models

ACL 2025finding

Living needs are the needs people generate in their daily lives for survival and well-being. On life service platforms like Meituan, user purchases are driven by living needs, making accurate living need predictions crucial for personalized service recommendations. Traditional approaches treat this…

Cited by 0SourcePDFScholar
2025

PANDA: Towards Generalist Video Anomaly Detection via Agentic AI Engineer

NeurIPS 2025poster

Video anomaly detection (VAD) is a critical yet challenging task due to the complex and diverse nature of real-world scenarios. Previous methods typically rely on domain-specific training data and manual adjustments when applying to new scenarios and unseen anomaly types, suffering from high labor c…

Cited by 0SourcecodeScholar
2025

PychoAgent: Psychology-driven LLM Agents for Explainable Panic Prediction on Social Media during Sudden Disaster Events

EMNLP 2025

Accurately predicting public panic sentiment on social media is crucial for proactive governance and crisis management. Current efforts on this problem face three main challenges: lack of finely annotated data hinders emotion prediction studies, unmodeled risk perception causes prediction inaccuraci

2025

RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation

NeurIPS 2025poster

Recent advances in vision-language models (VLMs) have enabled instruction-conditioned robotic systems with improved generalization. However, most existing work focuses on reactive System 1 policies, underutilizing VLMs’ strengths in semantic reasoning and long-horizon planning. These System 2 capabi…

Cited by 0SourceScholar
2025

RoboScape: Physics-informed Embodied World Model

NeurIPS 2025spotlight

World models have become indispensable tools for embodied intelligence, serving as powerful simulators capable of generating realistic robotic videos while addressing critical data scarcity challenges. However, current embodied world models exhibit limited physical awareness, particularly in modelin…

Cited by 0SourcecodeScholar
2025

Textured Gaussians for Enhanced 3D Scene Appearance Modeling

CVPR 2025poster

3D Gaussian Splatting (3DGS) has recently emerged as a state-of-the-art 3D reconstruction and rendering technique due to its high-quality results and fast training and rendering time. However, pixels covered by the same Gaussian are always shaded in the same color up to a Gaussian falloff scaling fa…

Cited by 3SourcePDFScholar
2025

Towards Realistic Earth-Observation Constellation Scheduling: Benchmark and Methodology

NeurIPS 2025poster

Agile Earth Observation Satellites (AEOSs) constellations offer unprecedented flexibility for monitoring the Earth’s surface, but their scheduling remains challenging under large-scale scenarios, dynamic environments, and stringent constraints. Existing methods often simplify these complexities,…

Cited by 0SourcecodeScholar
2025

UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces

ACL 2025long

Large multimodal models exhibit remarkable intelligence, yet their embodied cognitive abilities during motion in open-ended urban aerial spaces remain to be explored. We introduce a benchmark to evaluate whether video-large language models (Video-LLMs) can naturally process continuous first-person v…

Cited by 0SourcePDFScholar
2024

Aerial Image-based Inter-day Registration for Precision Agriculture

ICRA 2024poster

Satellite imagery has traditionally been used to collect crop statistics, but its low resolution and registration accuracy limit agricultural analytics to plant stand levels and large areas. Precision agriculture seeks analytic tools at near single plant level, and this work explores how to improve…

Cited by 4SourceScholar
2024

EconAgent: Large Language Model-Empowered Agents for Simulating Macroeconomic Activities

ACL 2024long

The advent of artificial intelligence has led to a growing emphasis on data-driven modeling in macroeconomics, with agent-based modeling (ABM) emerging as a prominent bottom-up simulation paradigm. In ABM, agents (*e.g.*, households, firms) interact within a macroeconomic environment, collectively g…

2024

Eliminating Cross-modal Conflicts in BEV Space for LiDAR-Camera 3D Object Detection

ICRA 2024poster

Recent 3D object detectors typically utilize multi-sensor data and unify multi-modal features in the shared bird’s-eye view (BEV) representation space. However, our empirical findings indicate that previous methods have limitations in generating fusion BEV features free from cross-modal conflicts. T…

Cited by 12SourcecodeScholar
2024

Global-Local Collaborative Inference with LLM for Lidar-Based Open-Vocabulary Detection

ECCV 2024poster

"Open-Vocabulary Detection (OVD) is the task of detecting all interesting objects in a given scene without predefined object classes. Extensive work has been done to deal with the OVD for 2D RGB images, but the exploration of 3D OVD is still limited. Intuitively, lidar point clouds provide 3D inform…

2024

MHPS: Multimodality-Guided Hierarchical Policy Search for Knowledge Graph Reasoning

ICASSP 2024accepted

Recently, path inference-based knowledge graph reasoning (KGR) methods have attracted great attention due to their good performance and interpretability. However, as the number of hops increases, the search space grows exponentially, making the reward sparse and the process of reasoning difficult. T…

Cited by 0SourceScholar
2024

Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration

NeurIPS 2024poster

Membership Inference Attacks (MIA) aim to infer whether a target data record has been utilized for model training or not. Existing MIAs designed for large language models (LLMs) can be bifurcated into two types: reference-free and reference-based attacks. Although reference-based attacks appear prom…

2024

SpecNeRF: Gaussian Directional Encoding for Specular Reflections

CVPR 2024highlight

Neural radiance fields have achieved remarkable performance in modeling the appearance of 3D scenes. However existing approaches still struggle with the view-dependent appearance of glossy surfaces especially under complex lighting of indoor environments. Unlike existing methods which typically assu…

Cited by 9SourcePDFScholar
2023

Adaptive Zone-Aware Hierarchical Planner for Vision-Language Navigation

CVPR 2023poster

The task of Vision-Language Navigation (VLN) is for an embodied agent to reach the global goal according to the instruction. Essentially, during navigation, a series of sub-goals need to be adaptively set and achieved, which is naturally a hierarchical navigation process. However, previous methods l…

2023

OmnimatteRF: Robust Omnimatte with 3D Background Modeling

ICCV 2023poster

Video matting has broad applications, from adding interesting effects to casually captured movies to assisting video production professionals. Matting with associated effects such as shadows and reflections has also attracted increasing research activity, and methods like Omnimatte have been propos…

Cited by 7PDFcodeScholar
2023

Progressively Optimized Local Radiance Fields for Robust View Synthesis

CVPR 2023poster

We present an algorithm for reconstructing the radiance field of a large-scale scene from a single casually captured video. The task poses two core challenges. First, most existing radiance field reconstruction approaches rely on accurate pre-estimated camera poses from Structure-from-Motion algorit…

Cited by 106SourcePDFScholar
2023

Robust Dynamic Radiance Fields

CVPR 2023poster

Dynamic radiance field reconstruction methods aim to model the time-varying structure and appearance of a dynamic scene. Existing methods, however, assume that accurate camera poses can be reliably estimated by Structure from Motion (SfM) algorithms. These methods, thus, are unreliable as SfM algori…

2022

3D-SPS: Single-Stage 3D Visual Grounding via Referred Point Progressive Selection

CVPR 2022oral

3D visual grounding aims to locate the referred target object in 3D point cloud scenes according to a free-form language description. Previous methods mostly follow a two-stage paradigm, i.e., language-irrelevant detection and cross-modal matching, which is limited by the isolated architecture. In s…

Cited by 69PDFcodeScholar
2022

Reinforced Structured State-Evolution for Vision-Language Navigation

CVPR 2022poster

Vision-and-language Navigation (VLN) task requires an embodied agent to navigate to a remote location following a natural language instruction. Previous methods usually adopt a sequence model (e.g., Transformer and LSTM) as the navigator. In such a paradigm, the sequence model predicts action at eac…

Cited by 47PDFcodeScholar
2021

Language-Guided Global Image Editing via Cross-Modal Cyclic Mechanism

ICCV 2021poster

Editing an image automatically via a linguistic request can significantly save laborious manual work and is friendly to photography novice. In this paper, we focus on the task of language-guided global image editing. Existing works suffer from imbalanced data distribution of real-world datasets and…

Cited by 29PDFScholar
2021

Mining the Benefits of Two-stage and One-stage HOI Detection

NeurIPS 2021poster

Two-stage methods have dominated Human-Object Interaction~(HOI) detection for several years. Recently, one-stage HOI detection methods have become popular. In this paper, we aim to explore the essential pros and cons of two-stage and one-stage methods. With this as the goal, we find that conventiona…

2021

Progressive Feature Interaction Search for Deep Sparse Network

NeurIPS 2021poster

Deep sparse networks (DSNs), of which the crux is exploring the high-order feature interactions, have become the state-of-the-art on the prediction task with high-sparsity features. However, these models suffer from low computation efficiency, including large model size and slow model inference, whi…

Cited by 16SourcePDFScholar
2021

Room-and-Object Aware Knowledge Reasoning for Remote Embodied Referring Expression

CVPR 2021poster

The Remote Embodied Referring Expression (REVERIE) is a recently raised task that requires an agent to navigate to and localise a referred remote object according to a high-level language instruction. Different from related VLN tasks, the key to REVERIE is to conduct goal-oriented exploration instea…

Cited by 91PDFcodeScholar
2020

AdversarialNAS: Adversarial Neural Architecture Search for GANs

CVPR 2020poster

Neural Architecture Search (NAS) that aims to automate the procedure of architecture design has achieved promising results in many computer vision fields. In this paper, we propose an AdversarialNAS method specially tailored for Generative Adversarial Networks (GANs) to search for a superior generat…

Cited by 114PDFcodeScholar
2020

DRG: Dual Relation Graph for Human-Object Interaction Detection

ECCV 2020poster

We tackle the challenging problem of human-object interaction (HOI) detection. Existing methods either recognize the interaction of each human-object pair in isolation or perform joint inference based on complex appearance-based features. In this paper, we leverage an abstract spatial-semantic repre…

2020

NAS-DIP: Learning Deep Image Prior with Neural Architecture Search

ECCV 2020poster

Recent work has shown that the structure of deep convolutional neural networks can be used as a structured image prior for solving various inverse image restoration tasks. Instead of using hand-designed architectures, we propose to search for neural architectures that capture stronger image priors.…

2020

PSGAN: Pose and Expression Robust Spatial-Aware GAN for Customizable Makeup Transfer

CVPR 2020oral

In this paper, we address the makeup transfer task, which aims to transfer the makeup from a reference image to a source image. Existing methods have achieved promising progress in constrained scenarios, but transferring between images with large pose and expression differences is still challenging.…

Cited by 181PDFcodeScholar
2019

Why Can't I Dance in the Mall? Learning to Mitigate Scene Bias in Action Recognition

NeurIPS 2019poster

Human activities often occur in specific scene contexts, e.g., playing basketball on a basketball court. Training a model using existing video datasets thus inevitably captures and leverages such bias (instead of using the actual discriminative cues). The learned representation may not generalize we…