← Search

Feng Gao

68 accepted papers

2026

CamGeo: Sparse Camera-Conditioned Image-to-Video Generation with 3D Geometry Priors

ICML 2026poster

Sparse camera-conditioned image-to-video generation presents a pivotal challenge: synthesizing geometrically consistent 3D motion from minimal pose cues. Existing methods, which largely rely on dense supervision or naive interpolation, suffer from severe pose drift and motion discontinuities due to …

Cited by 0SourceScholar
2026

ClipGStream: Clip-Stream Gaussian Splatting for Any Length and Any Motion Multi-View Dynamic Scene Reconstruction

CVPR 2026

Dynamic 3D scene reconstruction is essential for immersive media such as VR, MR, and XR, yet remains challenging for long multi-view sequences with large-scale motion. Existing dynamic Gaussian approaches are either Frame-Stream, offering scalability but poor temporal stability, or Clip, achieving l

Cited by 0SourceScholar
2026

DCA-LUT: Deep Chromatic Alignment with 5D LUT for Purple Fringing Removal

AAAI 2026technical

Purple fringing, a persistent artifact caused by Longitudinal Chromatic Aberration (LCA) in camera lenses, has long degraded the clarity and realism of digital imaging. Traditional solutions rely on complex and expensive apochromatic (APO) lens hardware and the extraction of handcrafted features, ig

Cited by 0SourcePDFScholar
2026

FlightBench: Benchmarking Learning-Based Methods for Ego-Vision-Based Quadrotors Navigation

ICRA 2026poster

Ego-vision-based navigation in cluttered environments is crucial for mobile systems, particularly agile quadrotors. While learning-based methods have shown promise recently, head-to-head comparisons with cutting-edge optimization-based approaches are scarce, leaving open the question of where and to…

2026

Hysteresis-Aware Neural Network Modeling and Whole-Body Reinforcement Learning Control of Soft Robots

ICRA 2026poster

Soft robots are inherently compliant and safe, making them suitable for humaninteractive applications such as surgery. However, their nonlinear and hysteretic behavior poses significant challenges for accurate modeling and control. We present a soft robotic system and propose a hysteresis-aware whol…

2026

Intrinsic Geometry-Appearance Consistency Optimization for Sparse-View Gaussian Splatting

CVPR 2026

3D Gaussian Splatting (3DGS) represents scenes through primitives with coupled intrinsic properties: geometric attributes (position, covariance, opacity) and appearance attributes (view-dependent color). Faithful reconstruction requires intrinsic geometry-appearance consistency, where geometry accur

Cited by 0SourceScholar
2026

Joint Learning of General and Diverse Patterns with Mixture of Memory Experts for Weakly-Supervised Video Anomaly Detection

CVPR 2026

Weakly-supervised Video Anomaly Detection (wVAD) aims to detect abnormal events using only binary labels, making it challenging to capture both the diversity of anomalies and their shared semantic cues. Existing methods either focus on a generic anomaly pattern, achieving strong generalization but w

Cited by 0SourceScholar
2026

JuggleRL: Mastering Ball Juggling with a Quadrotor Via Deep Reinforcement Learning

ICRA 2026poster

Aerial robots interacting with objects must perform precise, contact-rich maneuvers under uncertainty. In this paper, we study the problem of aerial ball juggling using a quadrotor equipped with a racket, a task that demands accurate timing, stable control, and continuous adaptation. We propose Jugg…

2026

MLM: Learning Multi-Task Loco-Manipulation Whole-Body Control for Quadruped Robot With Arm

RA-L 2026

Whole-body loco-manipulation for quadruped robots with arms remains a challenging problem, particularly in achieving multi-task control. To address this, we propose MLM, a reinforcement learning framework driven by both real-world and simulation data. It enables a six-DoF robotic arm–equipped quadru

Cited by 4SourceScholar
2026

MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling

ICLR 2026poster

Existing video generation models predominantly emphasize appearance fidelity while exhibiting limited ability to synthesize complex human motions, such as whole-body movements, long-range dynamics, and fine-grained human–environment interactions. This often leads to unrealistic or physically implaus…

Cited by 0SourceScholar
2026

Neural Internal Model Control: Learning a Robust Control Policy Via Predictive Error Feedback

ICRA 2026poster

Accurate motion control in the face of disturbances within complex environments remains a major challenge in robotics. Classical model-based approaches often struggle with nonlinearities and unstructured disturbances, while reinforcement learning (RL)-based methods can be fragile when encountering u…

2026

OpenDance: Multimodal Controllable 3D Dance Generation with Large-scale Internet Data

CVPR 2026

Music-driven 3D dance generation offers significant creative potential, yet practical applications demand versatile and multimodal control. Given the highly dynamic and complex human motion covering various styles and genres, dance generation requires satisfying diverse conditions beyond just music

Cited by 0SourceScholar
2026

Preserving Forgery Artifacts: AI-Generated Video Detection at Native Scale

ICLR 2026poster

The rapid advancement of video generation models has enabled the creation of highly realistic synthetic media, raising significant societal concerns regarding the spread of misinformation. However, current detection methods suffer from critical limitations. They often rely on preprocessing operation…

Cited by 0SourceScholar
2026

SCAN: Self-Calibrated AutoregressioN for High-Quality Visual Generation

AAAI 2026technical

Human artists can continuously refine their coarse sketches during artistic creation. This is quite different from existing autoregressive generation, where a token is determined once sampled. Aiming to flexibly refine the generated contents, this paper presents a Self-Calibrated AutoregressioN (SCA

Cited by 0SourcePDFScholar
2026

SEA-PACE: Semi-Supervised Underwater Image Enhancement via Gaussian Process–Assisted Self-Paced Learning

AAAI 2026technical

The scarcity of paired data severely limits the performance and generalization of learning-based underwater image enhancement (UIE) methods. This challenge is particularly prominent in scenes with complex degradations. Semi-supervised learning has emerged as a promising solution by enabling the util

Cited by 0SourcePDFScholar
2026

U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation

CVPR 2026

Full-stack multimodal interaction in real-time is a central goal in building intelligent embodied agents capable of natural, dynamic communication. However, existing systems are either limited to unimodal generation or suffer from degraded reasoning and poor cross-modal alignment, preventing coheren

Cited by 0SourcecodeScholar
2026

UAV-CB: A Complex-Background RGB-T Dataset and Local Frequency Bridge Network for UAV Detection

CVPR 2026

Detecting Unmanned Aerial Vehicles (UAVs) in low-altitude environments is essential for perception and defense systems but remains highly challenging due to complex backgrounds, camouflage, and multimodal interference. In real-world scenarios, UAVs are frequently visually blended with surrounding st

Cited by 0SourcecodeScholar
2026

UQ-Bench: A Benchmark for Evaluating Multimodal LLMs on Underwater Image Quality Assessment

AAAI 2026technical

Despite the rapid progress of multimodal large language models (MLLMs), their capacity for low-level visual perception in underwater environments remains underexplored. To address this gap, we present UQ-Bench, the first systematically designed benchmark for evaluating the ability of MLLMs to percei

Cited by 0SourcePDFScholar
2026

USER: A Unified and Extensible System for Online Real-World Policy Learning in Embodied AI

RSS 2026poster

Online policy learning directly in the physical world is a promising yet challenging direction for embodied intelligence. Unlike simulation, real-world systems cannot be arbitrarily accelerated, cheaply reset, or massively replicated, which makes scalable data collection, heterogeneous deployment, a…

Cited by 0SourceScholar
2026

What Matters in Learning a Zero-Shot Sim-To-Real RL Policy for Quadrotor Control? a Comprehensive Study

ICRA 2026poster

Precise and agile flight maneuvers are essential for quadrotor applications, yet traditional control methods are limited by their reliance on flat trajectories or computationally intensive optimization. Reinforcement learning (RL)-based policies offer a promising alternative by directly mapping obse…

2026

Zero-Shot VISUAL GROUNDING in 3D Gaussians via View Retrieval

ICASSP 2026poster

3D Visual Grounding (3DVG) aims to locate objects in 3D scenes based on text prompts, which is essential for applications such as robotics. However, existing 3DVG methods encounter two main challenges: first, they struggle to handle the implicit representation of spatial textures in 3D Gaussian Spla…

Cited by 0SourcePDFScholar
2025

ARM: Appearance Reconstruction Model for Relightable 3D Generation

CVPR 2025highlight

Recent image-to-3D reconstruction models have greatly advanced geometry generation, but they still struggle to faithfully generate realistic appearance. To address this, we introduce ARM, a novel method that reconstructs high-quality 3D meshes and realistic appearance from sparse-view images. The co…

2025

Aligning Human Motion Generation with Human Perceptions

ICLR 2025poster

Human motion generation is a critical task with a wide spectrum of applications. Achieving high realism in generated motions requires naturalness, smoothness, and plausibility. However, current evaluation metrics often rely on simple heuristics or distribution distances and do not align well with hu…

2025

Both Supply and Precision: Sample Debias and Ranking Consistency Joint Learning for Large Scale Pre-Ranking System

AAAI 2025technical

Cascade ranking architecture, composed of matching, pre-ranking, ranking and re-ranking stages, is usually adopted to balance the efficiency and effectiveness in real-world recommendation system (RS). As the middle stage of RS, pre-ranking aims to quickly filter out the low-quality items selected at…

Cited by 0SourcePDFScholar
2025

FlightBench: Benchmarking Learning-Based Methods for Ego-Vision-Based Quadrotors Navigation

RA-L 2025

Ego-vision-based navigation in cluttered environments is crucial for mobile systems, particularly agile quadrotors. While learning-based methods have shown promise recently, head-to-head comparisons with cutting-edge optimization-based approaches are scarce, leaving open the question of where and to

Cited by 3SourcecodeScholar
2025

Hysteresis-Aware Neural Network Modeling and Whole-Body Reinforcement Learning Control of Soft Robots

RA-L 2025

Soft robots are inherently compliant and safe, making them suitable for human-interactive applications such as surgery. However, their nonlinear and hysteretic behavior, arising from the properties of soft materials, presents substantial challenges for accurate modeling and control. In this study, w

Cited by 2SourceScholar
2025

Learning Natural and Robust Hexapod Locomotion over Complex Terrains via Motion Priors based on Deep Reinforcement Learning

IROS 2025

Multi-legged robots offer enhanced stability to navigate complex terrains with their multiple legs interacting with the environment. However, how to effectively coordinate the multiple legs in a larger action exploration space to generate natural and robust movements is a key issue. In this paper, w

Cited by 0SourceScholar
2025

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network

ICML 2025poster

Reinforcement learning (RL) for continuous control often requires large amounts of online interaction data. Value-based RL methods can mitigate this burden by offering relatively high sample efficiency. Some studies further enhance sample efficiency by incorporating offline demonstration data to “…

Cited by 0SourcePDFScholar
2025

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

NeurIPS 2025poster

Audio-driven human animation methods, such as talking head and talking body generation, have made remarkable progress in generating synchronized facial movements and appealing visual quality videos. However, existing methods primarily focus on single human animation and struggle with multi-stream au…

Cited by 0SourcecodeScholar
2025

M-LLM Based Video Frame Selection for Efficient Video Understanding

CVPR 2025poster

Recent advances in Multi-Modal Large Language Models (M-LLMs) show promising results in video reasoning. Popular Multi-Modal Large Language Model (M-LLM) frameworks usually apply naive uniform sampling to reduce the number of video frames that are fed into an M-LLM, particularly for long context vid…

Cited by 3SourcePDFScholar
2025

Mastering Multi-Drone Volleyball through Hierarchical Co-Self-Play Reinforcement Learning

CoRL 2025poster

In this paper, we tackle the problem of learning to play 3v3 multi-drone volleyball, a new embodied competitive task that requires both high-level strategic coordination and low-level agile control. The task is turn-based, multi-agent, and physically grounded, posing significant challenges due to it…

Cited by 0SourceScholar
2025

Multi-UAV Formation Control with Static and Dynamic Obstacle Avoidance via Reinforcement Learning

IROS 2025

This paper tackles the challenging task of maintaining formation among multiple unmanned aerial vehicles (UAVs) while avoiding both static and dynamic obstacles during directed flight. The complexity of the task arises from its multi-objective nature, the large exploration space, and the sim-to-real

Cited by 7SourceScholar
2025

Neural Internal Model Control: Learning a Robust Control Policy Via Predictive Error Feedback

RA-L 2025

Accurate motion control in the face of disturbances within complex environments remains a major challenge in robotics. Classical model-based approaches often struggle with nonlinearities and unstructured disturbances, while reinforcement learning (RL)-based methods can be fragile when encountering u

Cited by 5SourcecodeScholar
2025

OmniGuard: Hybrid Manipulation Localization via Augmented Versatile Deep Image Watermarking

CVPR 2025poster

With the rapid growth of generative AI and its widespread application in image editing, new risks have emerged regarding the authenticity and integrity of digital content. Existing versatile watermarking approaches suffer from trade-offs between tamper localization precision and visual quality. Cons…

Cited by 4SourcePDFScholar
2025

Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance

EMNLP 2025

Vision-Language-Action (VLA) models have made substantial progress by leveraging the robust capabilities of Visual Language Models (VLMs). However, VLMs’ significant parameter size and autoregressive (AR) decoding nature impose considerable computational demands on VLA models. While Speculative Deco

2025

TWLHex: A Biologically Inspired Multi-Morphology Transformable Wheel-Legged Hexapod

RA-L 2025

This paper presents a novel biologically inspired transformable wheel-legged hexapod, TWLHex, which is capable of actively switching among three morphologies: leg, wheel, and spoke. The mechanical design and kinematic models for transformable mechanism (TranMech) and wheel-spoke leg mechanism (WSLeg

Cited by 1SourceScholar
2025

VolleyBots: A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic Play

NeurIPS 2025poster

Robot sports, characterized by well-defined objectives, explicit rules, and dynamic interactions, present ideal scenarios for demonstrating embodied intelligence. In this paper, we present VolleyBots, a novel robot sports testbed where multiple drones cooperate and compete in the sport of volleybal…

Cited by 0SourcecodeScholar
2025

What Can RL Bring to VLA Generalization? An Empirical Study

NeurIPS 2025poster

Large Vision-Language Action (VLA) models have shown significant potential for embodied AI. However, their predominant training via supervised fine-tuning (SFT) limits generalization due to susceptibility to compounding errors under distribution shifts. Reinforcement learning (RL) offers a path to…

Cited by 0SourcecodeScholar
2025

What Matters in Learning a Zero-Shot Sim-to-Real RL Policy for Quadrotor Control? A Comprehensive Study

RA-L 2025

Precise and agile flight maneuvers are essential for quadrotor applications, yet traditional control methods are limited by their reliance on flat trajectories or computationally intensive optimization. Reinforcement learning (RL)-based policies offer a promising alternative by directly mapping obse

Cited by 13SourceScholar
2024

Adaptive-Gradient Policy Optimization: Enhancing Policy Learning in Non-Smooth Differentiable Simulations

ICML 2024poster

Recent advancements in differentiable simulators highlight the potential of policy optimization using simulation gradients. Yet, these approaches are largely contingent on the continuity and smoothness of the simulation, which precludes the use of certain simulation engines, such as Mujoco. To tackl…

Cited by 2SourcePDFScholar
2024

Atlas3D: Physically Constrained Self-Supporting Text-to-3D for Simulation and Fabrication

NeurIPS 2024poster

Existing diffusion-based text-to-3D generation methods primarily focus on producing visually realistic shapes and appearances, often neglecting the physical constraints necessary for downstream tasks. Generated models frequently fail to maintain balance when placed in physics-based simulations or 3D…

Cited by 5SourcePDFScholar
2024

Exploring Cross-Domain Few-Shot Classification via Frequency-Aware Prompting

IJCAI 2024poster

Cross-Domain Few-Shot Learning has witnessed great stride with the development of meta-learning. However, most existing methods pay more attention to learning domain-adaptive inductive bias (meta-knowledge) through feature-wise manipulation or task diversity improvement while neglecting the phenomen…

2024

Flow Priors for Linear Inverse Problems via Iterative Corrupted Trajectory Matching

NeurIPS 2024poster

Generative models based on flow matching have attracted significant attention for their simplicity and superior performance in high-resolution image synthesis. By leveraging the instantaneous change-of-variables formula, one can directly compute image likelihoods from a learned flow, making them ent…

2024

Obstacle Avoidance Strategy for a Novel Skiing Robot in Unknown Snow Environments

RA-L 2024

Obstacle avoidance is a significant part of robot autonomous navigation. Compared with traditional legged and actively driven wheeled robots, the skiing robot is a nonholonomic system, which cannot precisely control the speed. This paper presents an obstacle avoidance method based on the risk region

Cited by 1SourceScholar
2024

OmniDrones: An Efficient and Flexible Platform for Reinforcement Learning in Drone Control

RA-L 2024

In this work, we introduce <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">OmniDrones</i> , an efficient and flexible platform tailored for reinforcement learning in drone control, built on Nvidia's Omniverse Isaac Sim. It employs a bottom-up design

Cited by 53SourcecodeScholar
2024

PKU-DyMVHumans: A Multi-View Video Benchmark for High-Fidelity Dynamic Human Modeling

CVPR 2024poster

High-quality human reconstruction and photo-realistic rendering of a dynamic scene is a long-standing problem in computer vision and graphics. Despite considerable efforts invested in developing various capture systems and reconstruction algorithms recent advancements still struggle with loose or ov…

2024

Skews in the Phenomenon Space Hinder Generalization in Text-to-Image Generation

ECCV 2024poster

"The literature on text-to-image generation is plagued by issues of faithfully composing entities with relations. But there lacks a formal understanding of how entity-relation compositions can be effectively learned. Moreover, the underlying phenomenon space that meaningfully reflects the problem st…

2023

CL-MVSNet: Unsupervised Multi-View Stereo with Dual-Level Contrastive Learning

ICCV 2023poster

Unsupervised Multi-View Stereo (MVS) methods have achieved promising progress recently. However, previous methods primarily depend on the photometric consistency assumption, which may suffer from two limitations: indistinguishable regions and view-dependent effects, e.g., low-textured areas and refl…

Cited by 16PDFcodeScholar
2023

GIVL: Improving Geographical Inclusivity of Vision-Language Models With Pre-Training Methods

CVPR 2023poster

A key goal for the advancement of AI is to develop technologies that serve the needs not just of one group but of all communities regardless of their geographical region. In fact, a significant proportion of knowledge is locally shared by people from certain regions but may not apply equally in othe…

2023

Learning non-Markovian Decision-Making from State-only Sequences

NeurIPS 2023poster

Conventional imitation learning assumes access to the actions of demonstrators, but these motor signals are often non-observable in naturalistic settings. Additionally, sequential decision-making behaviors in these settings can deviate from the assumptions of a standard Markov Decision Process (MDP)…

Cited by 9SourcePDFScholar
2023

Learning-Based Distortion Compensation for a Hybrid Simulator of Space Docking

RA-L 2023

By effectively utilizing the fidelity of a physical simulation and the flexibility of a numerical simulation, the hybrid simulation is applicable to test the complicated docking contact process of various kinds of spacecraft. However, the hybrid simulation of space docking often has a divergence or

Cited by 5SourceScholar
2022

Efficient One-Pass Multi-View Subspace Clustering with Consensus Anchors

AAAI 2022technical

Multi-view subspace clustering (MVSC) optimally integrates multiple graph structure information to improve clustering performance. Recently, many anchor-based variants are proposed to reduce the computational complexity of MVSC. Though achieving considerable acceleration, we observe that most of the…

2022

Is MultiWOZ a Solved Task? An Interactive TOD Evaluation Framework with User Simulator

EMNLP 2022finding

Task-Oriented Dialogue (TOD) systems are drawing more and more attention in recent studies.Current methods focus on constructing pre-trained models or fine-tuning strategies while the evaluation of TOD is limited by a policy mismatch problem.That is, during evaluation, the user utterances are from t…

2022

Learning from Students: Online Contrastive Distillation Network for General Continual Learning

IJCAI 2022poster

The goal of General Continual Learning (GCL) is to preserve learned knowledge and learn new knowledge with constant memory from an infinite data stream where task boundaries are blurry. Distilling the model's response of reserved samples between the old and the new models is an effective way to achi…

2022

Meta Convolutional Neural Networks for Single Domain Generalization

CVPR 2022poster

In single domain generalization, models trained with data from only one domain are required to perform well on many unseen domains. In this paper, we propose a new model, termed meta convolutional neural network, to solve the single domain generalization problem in image recognition. The key idea is…

Cited by 60PDFScholar
2022

Transform-Retrieve-Generate: Natural Language-Centric Outside-Knowledge Visual Question Answering

CVPR 2022poster

Outside-knowledge visual question answering (OK-VQA) requires the agent to comprehend the image, make use of relevant knowledge from the entire web, and digest all the information to answer the question. Most previous works address the problem by first fusing the image and question in the multi-moda…

Cited by 113PDFcodeScholar
2021

Design and soft-landing control of a six-legged mobile repetitive lander for lunar exploration

ICRA 2021poster

The autonomous robots consisting of an immovable lander and a rover are widely deployed to explore extraterrestrial planets. However, these robots have two main limitations: (1) the separate design for lander and rover respectively results in heavy mass and big volume of the whole system, which incr…

Cited by 8SourceScholar
2021

Simultaneous Control of Terrain Adaptation and Wheel Speed Allocation for a Planetary Rover With an Active Suspension System

RA-L 2021

Active suspensions are important features of many recent planetary rovers. For such a rover, the control strategy is crucial to its performance. This letter presents the control method for a planetary rover equipped with an active suspension system. The control algorithm is based on the estimation o

Cited by 9SourceScholar
2021

Stair Climbing Capability-Based Dimensional Synthesis for the Multi-legged Robot

ICRA 2021poster

Staircase is a typical obstacle for the legged robot to overcome in buildings. This paper studies the stair climbing capability-based dimensional synthesis for a hexapod legged robot, i.e., exploring how to determine the leg length and the longitudinal body length concerning the target staircase in…

Cited by 10SourceScholar
2020

Attentional Separation-and-Aggregation Network for Self-supervised Depth-Pose Learning in Dynamic Scenes

CoRL 2020

Learning depth and ego-motion from unlabeled videos via self-supervision from epipolar projection can improve the robustness and accuracy of the 3D perception and localization of vision-based robots. However, the rigid projection computed by ego-motion cannot represent all scene points, such as poin

Cited by 0SourcePDFScholar
2019

Learning Perceptual Inference by Contrasting

NeurIPS 2019spotlight

“Thinking in pictures,” [1] i.e., spatial-temporal reasoning, effortless and instantaneous for humans, is believed to be a significant ability to perform logical induction and a crucial factor in the intellectual history of technology development. Modern Artificial Intelligence (AI), fueled by massi…

2019

RAVEN: A Dataset for Relational and Analogical Visual REasoNing

CVPR 2019poster

Dramatic progress has been witnessed in basic vision tasks involving low-level perception, such as object recognition, detection, and tracking. Unfortunately, there is still enormous performance gap between artificial vision systems and human intelligence in terms of higher-level vision problems, es…

Cited by 357PDFScholar
2017

A glove-based system for studying hand-object manipulation via joint pose and force sensing

IROS 2017poster

We present a design of an easy-to-replicate glove-based system that can reliably perform simultaneous hand pose and force sensing in real time, for the purpose of collecting human hand data during fine manipulative actions. The design consists of a sensory glove that is capable of jointly collecting…

Cited by 70SourceScholar
2017

Feeling the force: Integrating force and pose for fluent discovery through imitation learning to open medicine bottles

IROS 2017poster

Learning complex robot manipulation policies for real-world objects is challenging, often requiring significant tuning within controlled environments. In this paper, we learn a manipulation model to execute tasks with multiple stages and variable structure, which typically are not suitable for most…

Cited by 78SourceScholar