← Search

Guangyao Shi

11 accepted papers

2026

First Frame Is the Place to Go for Video Content Customization

CVPR 2026

What role does the first frame play in video generation models? Traditionally, it's viewed as the spatial-temporal starting point of a video, merely a seed for subsequent animation. In this work, we reveal a fundamentally different perspective: video models implicitly treat the first frame as a conc

Cited by 0SourcecodeScholar
2025

VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding

NeurIPS 2025poster

Vision Language models (VLMs) have achieved remarkable success in video understanding tasks. Yet, a key question remains: Do they comprehend visual information or merely learn superficial mappings between visual and textual patterns? Understanding visual cues, particularly those related to physics…

Cited by 0SourcecodeScholar
2024

AG-Cvg: Coverage Planning with a Mobile Recharging UGV and an Energy-Constrained UAV

ICRA 2024poster

In this paper, we present an approach for coverage path planning for a team of an energy-constrained Unmanned Aerial Vehicle (UAV) and an Unmanned Ground Vehicle (UGV). Both the UAV and the UGV have predefined areas that they have to cover. The goal is to perform complete coverage by both robots whi…

Cited by 8SourceScholar
2024

Inverse Submodular Maximization with Application to Human-in-the-Loop Multi-Robot Multi-Objective Coverage Control

IROS 2024poster

We consider a new type of inverse combinatorial optimization, Inverse Submodular Maximization (ISM), for human-in-the-loop multi-robot coordination. Forward combinatorial optimization - solving a combinatorial problem given the reward (cost)-related parameters - is widely used in multi-robot coordin…

Cited by 2SourceScholar
2024

LAVA: Long-horizon Visual Action based Food Acquisition

IROS 2024poster

Robotic Assisted Feeding (RAF) addresses the fundamental need for individuals with mobility impairments to regain autonomy in feeding themselves. The goal of RAF is to use a robot arm to acquire and transfer food to individuals from the table. Existing RAF methods primarily focus on solid foods, lea…

Cited by 8SourceScholar
2023

Data-Driven Distributionally Robust Optimal Control with State-Dependent Noise

IROS 2023poster

Distributionally Robust Optimal Control (DROC) is a technique that enables robust control in a stochastic setting when the true distribution is not known. Traditional DROC approaches require given ambiguity sets or a KL divergence bound to represent the distributional uncertainty. These may not be k…

Cited by 7SourcecodeScholar
2023

Decision-Oriented Learning with Differentiable Submodular Maximization for Vehicle Routing Problem

IROS 2023poster

We study the problem of learning a function that maps context observations (input) to parameters of a submodular function (output). Our motivating case study is a specific type of vehicle routing problem, in which a team of Unmanned Ground Vehicles (UGVs) can serve as mobile charging stations to rec…

Cited by 3SourceScholar
2023

Risk-aware Recharging Rendezvous for a Collaborative Team of UAVs and UGVs

ICRA 2023poster

We introduce and investigate the recharging rendezvous problem for a collaborative team of Unmanned Aerial Vehicles (UAVs) and Unmanned Ground Vehicles (UGVs), in which UAVs with limited battery capacity and UGVS persistently monitor an area. The UGVs also act as mobile recharging stations for the U…

Cited by 18SourceScholar
2022

Interactive Multi-Robot Aerial Cinematography Through Hemispherical Manifold Coverage

IROS 2022poster

This paper presents a distributed interactive framework to provide high-level position instructions for multi-robot aerial cinematography based on coverage over a hemisphere. The control strategy based on optimization of the coverage functional and geometric relationships over a hemisphere is presen…

Cited by 6SourceScholar
2021

Communication-Aware Multi-robot Coordination with Submodular Maximization

ICRA 2021poster

Submodular maximization has been widely used in many multi-robot task planning problems including information gathering, exploration, and target tracking. However, the interplay between submodular maximization and communication is rarely explored in the multi-robot setting. In many cases, maximizing…

Cited by 20SourceScholar
2020

Robust Multiple-Path Orienteering Problem: Securing Against Adversarial Attacks

RSS 2020poster

The multiple-path orienteering problem asks for paths for a team of robots that maximize the total reward collected while satisfying budget constraints on the path length. This problem models many multi-robot routing tasks such as exploring unknown environments and information gathering for env…

Cited by 31SourcePDFScholar