← Search

YuHao Li

11 accepted papers

2026

Adapting Execution-Time Objectives for Multi-Robot Policies via Collaborative Flow Policy Guidance

RSS 2026poster

Multi-robot teams are increasingly gaining attention due to their ability to scale up in terms of task workloads and complexities. However, existing approaches struggle with three key limitations: the inability of unimodal policies to capture multi-modal joint strategies, the rigidity of fixed polic…

Cited by 0SourceScholar
2026

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks

ICLR 2026poster

Deep reasoning is fundamental for solving complex tasks, especially in vision-centric scenarios that demand sequential, multimodal understanding. However, existing benchmarks typically evaluate agents with fully synthetic, single-turn queries, limited visual modalities, and lack a framework to asses…

Cited by 0SourcecodeScholar
2026

Hard-Constrained Graph Generation with Discrete-Projection Diffusion

ICML 2026poster

Diffusion models have achieved remarkable success in graph generation, but enforcing hard constraints on generated graphs remains challenging, limiting their deployment in constraint-critical applications. Existing approaches either fail to guarantee strict constraint satisfaction or are limited to …

Cited by 0SourceScholar
2025

A Culturally-diverse Multilingual Multimodal Video Benchmark & Model

EMNLP 2025

Large multimodal models (LMMs) have recently gained attention due to their effectiveness to understand and generate descriptions of visual content. Most existing LMMs are in English language. While few recent works explore multilingual image LMMs, to the best of our knowledge, moving beyond the Engl

Cited by 0SourcePDFScholar
2025

An Effective and Secure Federated Multi-View Clustering Method with Information-Theoretic Perspective

ICML 2025poster

Recently, federated multi-view clustering (FedMVC) has gained attention for its ability to mine complementary clustering structures from multiple clients without exposing private data. Existing methods mainly focus on addressing the feature heterogeneity problem brought by views on different clients…

Cited by 0SourcePDFScholar
2025

DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding

IROS 2025

While large multimodal models (LMMs) have demonstrated strong performance across various Visual Question Answering (VQA) tasks, certain challenges require complex multi-step reasoning to reach accurate answers. One particularly challenging task is autonomous driving, which demands thorough cognitive

Cited by 32SourcecodeScholar
2025

LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

ACL 2025finding

Step-by-step reasoning is crucial for solving complex visual tasks, yet existing approaches lack a comprehensive framework for evaluating this capability and do not emphasize step-wise problem-solving. To this end, we propose a comprehensive framework for advancing multi-step visual reasoning in lar…

2025

Reinforcement Learning Within the Classical Robotics Stack: A Case Study in Robot Soccer

ICRA 2025

Robot decision-making in partially observable, real-time, dynamic, and multi-agent environments remains a difficult and unsolved challenge. Model-free reinforcement learning (RL) is a promising approach to learning decisionmaking in such domains, however, end-to-end RL in complex environments is oft

Cited by 6SourceScholar
2025

VSNet: Focusing on the Linguistic Characteristics of Sign Language

CVPR 2025poster

Sign language is a visual language expressed through complex movements of the upper body. The human skeleton plays a critical role in sign language recognition due to its good separation from the video background. However, mainstream skeleton-based sign language recognition models often overly focus…

2024

CoFiI2P: Coarse-to-Fine Correspondences-Based Image to Point Cloud Registration

RA-L 2024

Image-to-point cloud (I2P) registration is a fundamental task for robots and autonomous vehicles to achieve cross-modality data fusion and localization. Current I2P registration methods primarily focus on estimating correspondences at the point or pixel level, often neglecting global alignment. As a

Cited by 17SourceScholar