← Search

Zihao Deng

10 accepted papers

2026

Adapting Execution-Time Objectives for Multi-Robot Policies via Collaborative Flow Policy Guidance

RSS 2026poster

Multi-robot teams are increasingly gaining attention due to their ability to scale up in terms of task workloads and complexities. However, existing approaches struggle with three key limitations: the inability of unimodal policies to capture multi-modal joint strategies, the rigidity of fixed polic…

Cited by 0SourceScholar
2026

Collaborative Planning with Concurrent Synchronization for Operationally Constrained UAV-UGV Teams

ICRA 2026poster

Collaborative planning under operational constraints is an essential capability for heterogeneous robot teams tackling complex large-scale real-world tasks. Unmanned Aerial Vehicles (UAVs) offer rapid environmental coverage, but flight time is often limited by energy constraints, whereas Unmanned Gr…

2025

Coordinated Multi-Robot Navigation with Formation Adaptation

ICRA 2025

Coordinated multi-robot navigation is an essential ability for a team of robots operating in diverse environments. Robot teams often need to maintain specific formations, such as wedge formations, to enhance visibility, positioning, and efficiency during fast movement. However, complex environments

Cited by 4SourceScholar
2025

Subteaming and Adaptive Formation Control for Coordinated Multi-Robot Navigation

CoRL 2025poster

Coordinated multi-robot navigation is essential for robots to operate as a team in diverse environments. During navigation, robot teams usually need to maintain specific formations, such as circular formations to protect human teammates at the center. However, in complex scenarios such as narrow c…

Cited by 0SourceScholar
2024

MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response

NAACL 2024findings

Large Language Models (LLMs) have shown immense potential in multimodal applications, yet the convergence of textual and musical domains remains not well-explored. To address this gap, we present MusiLingo, a novel system for music caption generation and music-related query responses. MusiLingo empl…

2023

Factorized Contrastive Learning: Going Beyond Multi-view Redundancy

NeurIPS 2023poster

In a wide range of multimodal tasks, contrastive learning has become a particularly appealing approach since it can successfully learn representations from abundant unlabeled data with only pairing information (e.g., image-caption or video-audio pairs). Underpinning these approaches is the assumptio…

2023

MultiViz: Towards Visualizing and Understanding Multimodal Models

ICLR 2023poster

The promise of multimodal models for real-world applications has inspired research in visualizing and understanding their internal mechanics with the end goal of empowering stakeholders to visualize model behavior, perform model debugging, and promote trust in machine learning models. However, moder…

2023

Quantifying & Modeling Multimodal Interactions: An Information Decomposition Framework

NeurIPS 2023poster

The recent explosion of interest in multimodal applications has resulted in a wide selection of datasets and methods for representing and integrating information from different modalities. Despite these empirical advances, there remain fundamental research questions: How can we quantify the interact…

2022

Polynomial Time Reinforcement Learning in Factored State MDPs with Linear Value Functions

AISTATS 2022poster

Many reinforcement learning (RL) environments in practice feature enormous state spaces that may be described compactly by a "factored" structure, that may be modeled by Factored Markov Decision Processes (FMDPs). We present the first polynomial time algorithm for RL in Factored State MDPs (generali…

Cited by 4SourcePDFScholar