← Search

Yuwei Wu

60 accepted papers

2026

Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense Charts

CVPR 2026

Multimodal large language models (MLLMs) have shown considerable potential in chart understanding and reasoning tasks. However, they still struggle with high information density (HID) charts characterized by multiple subplots, legends, and dense annotations due to three major challenges: (1) limited

Cited by 0SourcecodeScholar
2026

Composition-Incremental Learning for Compositional Generalization

AAAI 2026technical

Compositional generalization has achieved substantial progress in computer vision on pre-collected training data. Nonetheless, real-world data continually emerges, with possible compositions being nearly infinite, long-tailed, and not entirely visible. Thus, an ideal model is supposed to gradually i

Cited by 0SourcePDFScholar
2026

LongSplat: Online Generalizable 3D Gaussian Splatting from Long Sequence Images

AAAI 2026technical

3D Gaussian Splatting (3DGS) achieves high-fidelity novel view synthesis, but its application in online long-sequence scenarios is still restricted. Existing methods either rely on slow per-scene optimization or lack efficient frame-wise 3DGS updates, making them unsuitable for online long-sequence

Cited by 0SourcePDFScholar
2026

Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds

ICLR 2026poster

Modality alignment is critical for vision-language models (VLMs) to effectively integrate information across modalities. However, existing methods extract hierarchical features from text while representing each image with a single feature, leading to asymmetric and suboptimal alignment. To address t…

Cited by 0SourceScholar
2026

Rethinking 3D Shape Generation: Diffusion over Superquadrics

ICML 2026poster

Diffusion models have advanced 3D shape generation, yet most methods still denoise in high-cardinality spaces (e.g., voxel/SDF grids, meshes, or point clouds), which is computationally and memory intensive and makes it difficult to scale in terms of both higher resolution and stronger controllabilit…

Cited by 0SourceScholar
2026

Towards Optimizing a Convex Cover of Collision-Free Space for Trajectory Generation

ICRA 2026poster

We propose an online iterative algorithm to optimize a convex cover to under-approximate the free space for autonomous navigation to delineate Safe Flight Corridors (SFC). The convex cover consists of a set of polytopes such that the union of the polytopes represents obstacle-free space, allowing us…

2025

Beyond the Seen: Bounded Distribution Estimation for Open-Vocabulary Learning

NeurIPS 2025poster

Open-vocabulary learning requires modeling the data distribution in open environments, which consists of both seen-class and unseen-class data. Existing methods estimate the distribution in open environments using seen-class data, where the absence of unseen classes makes the estimation error inhe…

Cited by 0SourceScholar
2025

Consistency of Compositional Generalization Across Multiple Levels

AAAI 2025technical

Compositional generalization is the capability of a model to understand novel compositions composed of seen concepts. There are multiple levels of novel compositions including phrase-phrase level, phrase-word level, and word-word level. Existing methods achieve promising compositional generalization…

2025

Diving into the Fusion of Monocular Priors for Generalized Stereo Matching

ICCV 2025poster

The matching formulation makes it naturally hard for the stereo matching to handle ill-posed regions like occlusions and non-Lambertian surfaces. Fusing monocular priors has been proven helpful for ill-posed matching, but the biased monocular prior learned from small stereo datasets constrains the g…

2025

Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning

NeurIPS 2025poster

Multimodal agents, which integrate a controller (e.g., a vision language model) with external tools, have demonstrated remarkable capabilities in tackling complex multimodal tasks. Existing approaches for training these agents, both supervised fine-tuning and reinforcement learning, depend on extens…

Cited by 0SourceScholar
2025

Large Language Models are Demonstration Pre-Selectors for Themselves

ICML 2025poster

In-context learning with large language models (LLMs) delivers strong few-shot performance by choosing few-shot demonstrations from the entire training dataset. However, previous few-shot in-context learning methods, which calculate similarity scores for choosing demonstrations, incur high computati…

Cited by 0SourcePDFScholar
2025

Multi-Sourced Compositional Generalization in Visual Question Answering

IJCAI 2025

Compositional generalization is the ability of generalizing novel compositions from seen primitives, and has received much attention in vision-and-language (V&L) recently. Due to the multi-modal nature of V&L tasks, the primitives composing compositions source from different modalities, resulting in

2025

Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage

ICLR 2025spotlight

The advancement of large language models (LLMs) prompts the development of multi-modal agents, which are used as a controller to call external tools, providing a feasible way to solve practical tasks. In this paper, we propose a multi-modal agent tuning method that automatically generates multi-moda…

Cited by 5SourcePDFScholar
2025

Resilient Multi-Robot Target Tracking with Sensing and Communication Danger Zones

IROS 2025

Multi-robot collaboration for target tracking in adversarial environments poses significant challenges, including system failures, dynamic priority shifts, and other unpredictable factors. These challenges become even more pronounced when the environment is unknown. In this paper, we propose a resil

Cited by 1SourceScholar
2025

Sekai: A Video Dataset towards World Exploration

NeurIPS 2025poster

Video generation techniques have made remarkable progress, promising to be the foundation of interactive world exploration. However, existing video generation datasets are not well-suited for world exploration training as they suffer from some limitations: limited locations, short duration, static s…

Cited by 0SourceScholar
2025

Towards Optimizing a Convex Cover of Collision-Free Space for Trajectory Generation

RA-L 2025

We propose an online iterative algorithm to optimize a convex cover to under-approximate the free space for autonomous navigation to delineate Safe Flight Corridors (SFC). The convex cover consists of a set of polytopes such that the union of the polytopes represents obstacle-free space, allowing us

Cited by 10SourceScholar
2025

Vision Transformers for End-to-End Vision-Based Quadrotor Obstacle Avoidance

ICRA 2025

We demonstrate the capabilities of an attentionbased end-to-end approach for high-speed vision-based quadrotor obstacle avoidance in dense, cluttered environments, with comparison to various state-of-the-art learning architectures. Quadrotor unmanned aerial vehicles (UAVs) have tremendous maneuverab

Cited by 22SourceScholar
2025

World Knowledge-Enhanced Reasoning Using Instruction-Guided Interactor in Autonomous Driving

AAAI 2025technical

The Multi-modal Large Language Models (MLLMs) with extensive world knowledge have revitalized autonomous driving, particularly in reasoning tasks within perceivable regions. However, when faced with perception-limited areas (dynamic or static occlusion regions), MLLMs struggle to effectively integra…

Cited by 2SourcePDFScholar
2024

3D Affordance Keypoint Detection for Robotic Manipulation

IROS 2024poster

This paper presents a novel approach for affordance-informed robotic manipulation by introducing 3D keypoints to enhance the understanding of object parts’ functionality. The proposed approach provides direct information about what the potential use of objects is, as well as guidance on where and ho…

Cited by 0SourceScholar
2024

Design and Evaluation of Motion Planners for Quadrotors in Environments with Varying Complexities

ICRA 2024poster

Motion planning techniques for quadrotors have advanced significantly over the past decade. Most successful planners have two stages: a front-end that determines a path that incorporates geometric (or kinematic or input) constraints and specifies the homotopy class of the trajectory, and a back-end…

Cited by 5SourcecodeScholar
2024

FIRE: A Dataset for Feedback Integration and Refinement Evaluation of Multimodal Models

NeurIPS 2024poster

Vision language models (VLMs) have achieved impressive progress in diverse applications, becoming a prevalent research direction. In this paper, we build FIRE, a feedback-refinement dataset, consisting of 1.1M multi-turn conversations that are derived from 27 source datasets, empowering VLMs to spon…

Cited by 4SourcePDFScholar
2024

I Get the Hang of It! A Learning-Free Method to Predict Hanging Poses for Previously Unseen Objects

RA-L 2024

The action of hanging previously unseen objects remains a challenge for robots due to the multitude of object shapes and the limited number of stable hanging arrangements. This paper proposes a learning-free framework that enables robots to infer stable relative poses between the object being hung (

Cited by 1SourceScholar
2024

In-Context Compositional Generalization for Large Vision-Language Models

EMNLP 2024main

Recent work has revealed that in-context learning for large language models exhibits compositional generalization capacity, which can be enhanced by selecting in-context demonstrations similar to test cases to provide contextual information. However, how to exhibit in-context compositional generaliz…

Cited by 2SourcePDFScholar
2024

SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge

NeurIPS 2024poster

Large vision-language models (LVLMs) are ignorant of the up-to-date knowledge, such as LLaVA series, because they cannot be updated frequently due to the large amount of resources required, and therefore fail in many cases. For example, if a LVLM was released on January 2024, and it wouldn't know th…

Cited by 2SourcePDFScholar
2024

Trajectory Optimization with Global Yaw Parameterization for Field-of-View Constrained Autonomous Flight

IROS 2024poster

Trajectory generation for quadrotors with limited field-of-view sensors has numerous applications such as aerial exploration, coverage, inspection, videography, and target tracking. Most previous works simplify the task of optimizing yaw trajectories by either aligning the heading of the robot with…

Cited by 3SourceScholar
2023

Exploring the Effect of Primitives for Compositional Generalization in Vision-and-Language

CVPR 2023poster

Compositionality is one of the fundamental properties of human cognition (Fodor & Pylyshyn, 1988). Compositional generalization is critical to simulate the compositional capability of humans, and has received much attention in the vision-and-language (V&L) community. It is essential to understand th…

2023

Fast-StrucTexT: An Efficient Hourglass Transformer with Modality-guided Dynamic Token Merge for Document Understanding

IJCAI 2023poster

Transformers achieve promising performance in document understanding because of their high effectiveness and still suffer from quadratic computational complexity dependency on the sequence length. General efficient transformers are challenging to be directly adapted to model document. They are unabl…

Cited by 8SourcePDFScholar
2023

Learning-Free Grasping of Unknown Objects Using Hidden Superquadrics

RSS 2023poster

Robotic grasping is an essential and fundamental task and has been studied extensively over the past several decades. Traditional work analyzes physical models of the objects and computes force-closure grasps. Such methods require pre-knowledge of the complete 3D model of an object, which can be har…

Cited by 4SourcePDFScholar
2023

Marching-Primitives: Shape Abstraction From Signed Distance Function

CVPR 2023highlight

Representing complex objects with basic geometric primitives has long been a topic in computer vision. Primitive-based representations have the merits of compactness and computational efficiency in higher-level tasks such as physics simulation, collision checking, and robotic manipulation. Unlike pr…

2023

SEER: Safe Efficient Exploration for Aerial Robots using Learning to Predict Information Gain

ICRA 2023poster

We address the problem of efficient 3-D exploration in indoor environments for micro aerial vehicles with limited sensing capabilities and payload/power constraints. We develop an indoor exploration framework that uses learning to predict the occupancy of unseen areas, extracts semantic features, sa…

Cited by 46SourcecodeScholar
2022

Efficient Riemannian Meta-Optimization by Implicit Differentiation

AAAI 2022technical

To solve optimization problems with nonlinear constrains, the recently developed Riemannian meta-optimization methods show promise, which train neural networks as an optimizer to perform optimization on Riemannian manifolds. A key challenge is the heavy computational and memory burdens, because com…

2022

Hyperbolic Feature Augmentation via Distribution Estimation and Infinite Sampling on Manifolds

NeurIPS 2022accept

Learning in hyperbolic spaces has attracted growing attention recently, owing to their capabilities in capturing hierarchical structures of data. However, existing learning algorithms in the hyperbolic space tend to overfit when limited data is given. In this paper, we propose a hyperbolic feature a…

Cited by 12SourcePDFScholar
2022

Learning the Dynamics of Visual Relational Reasoning via Reinforced Path Routing

AAAI 2022technical

Reasoning is a dynamic process. In cognitive theories, the dynamics of reasoning refers to reasoning states over time after successive state transitions. Modeling the cognitive dynamics is of utmost importance to simulate human reasoning capability. In this paper, we propose to learn the reasoning d…

Cited by 8SourcePDFScholar
2022

Maintaining Reasoning Consistency in Compositional Visual Question Answering

CVPR 2022poster

A compositional question refers to a question that contains multiple visual concepts (e.g., objects, attributes, and relationships) and requires compositional reasoning to answer. Existing VQA models can answer a compositional question well, but cannot work well in terms of reasoning consistency in…

Cited by 29PDFcodeScholar
2022

Primitive-Based Shape Abstraction via Nonparametric Bayesian Inference

ECCV 2022poster

"3D shape abstraction has drawn great interest over the years. Apart from low-level representations such as meshes and voxels, researchers also seek to semantically abstract complex objects with basic geometric primitives. Recent deep learning methods rely heavily on datasets, with limited generalit…

Cited by 26SourcePDFScholar
2022

Primitive3D: 3D Object Dataset Synthesis From Randomly Assembled Primitives

CVPR 2022poster

Numerous advancements of deep learning can be attributed to access to large-scale and well-annotated datasets. However, such a dataset is prohibitively expensive in 3D computer vision due to the substantial collection cost. To alleviate this issue, we propose a cost-effective method for automaticall…

Cited by 5PDFScholar
2022

Robust and Accurate Superquadric Recovery: A Probabilistic Approach

CVPR 2022oral

Interpreting objects with basic geometric primitives has long been studied in computer vision. Among geometric primitives, superquadrics are well known for their ability to represent a wide range of shapes with few parameters. However, as the first and foremost step, recovering superquadrics accurat…

Cited by 53PDFcodeScholar
2021

Learning a Gradient-free Riemannian Optimizer on Tangent Spaces

AAAI 2021technical

A principal way of addressing constrained optimization problems is to model them as problems on Riemannian manifolds. Recently, Riemannian meta-optimization provides a promising way for solving constrained optimization problems by learning optimizers on Riemannian manifolds in a data-driven fashion,…

2020

On Isometry Robustness of Deep 3D Point Cloud Models Under Adversarial Attacks

CVPR 2020poster

While deep learning in 3D domain has achieved revolutionary performance in many tasks, the robustness of these models has not been sufficiently studied or explored. Regarding the 3D adversarial samples, most existing works focus on manipulation of local points, which may fail to invoke the global ge…

Cited by 95PDFcodeScholar
2017

HOPE: Hierarchical Object Prototype Encoding for Efficient Object Instance Search in Videos

CVPR 2017poster

This paper tackles the problem of efficient and effective object instance search in videos. To effectively capture the relevance between a query and video frames and precisely localize the particular object, we leverage the object proposals to improve the quality of object instance search in videos.…

Cited by 16PDFScholar