← Search

Shuo Liu

31 accepted papers

2026

EcoVLA: Environment-Aware Adaptive Pruning with Interleaved Inference Orchestration for Vision-Language-Action Models

ICML 2026spotlight

While Vision-Language-Action (VLA) models hold promise in embodied intelligence, their large parameter counts lead to substantial inference latency that hinders real-time manipulation, motivating parameter sparsification. However, as the environment evolves during VLA execution, the optimal sparsity…

Cited by 0SourceScholar
2026

GeoReward: Mitigating Contextual Variable Overestimation in Vision-Language Models for Cross-Market Preference Prediction

ICML 2026poster

Vision-language models (VLMs) excel in many multimodal tasks but remain prone to a subtle yet impactful failure mode: they tend to overestimate dominant visual-textual cues while underestimating sparse but decision-critical contextual variables. This issue, which we term Contextual Variable Overesti…

Cited by 0SourceScholar
2026

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic

ICML 2026poster

Recent work has explored optimizing LLM collaboration through Multi-Agent Reinforcement Learning (MARL). However, most MARL fine-tuning approaches rely on predefined execution protocols, which often require centralized execution. Decentralized LLM collaboration is more appealing in practice, as agen…

Cited by 0SourceScholar
2026

MolmoAct: Action Reasoning Models That Can Reason in Space

ICRA 2026poster

Reasoning is essential for purposeful action, yet most robotic foundation models map perception and instructions directly to control, limiting adaptability, generalization, and semantic grounding. We introduce Action Reasoning Models (ARMs), which integrate perception, planning, and control through …

2026

MolmoSpaces: Large-Scale Open Ecosystem for Robot Manipulation and Navigation

RSS 2026poster

Deploying robots at scale demands robustness to the long tail of everyday situations. The countless variations in scene layout, object geometry, and task specifications that characterize real environments are vast and underrepresented in existing robot benchmarks. Measuring this level of generalizat…

Cited by 0SourceScholar
2026

Sonar Mapping and Obstacle Avoidance for Autonomous Underwater Vehicles in Unknown Marine Environments

RA-L 2026

This letter proposes a safe navigation framework for autonomous underwater vehicles (AUVs) that integrates sonar mapping and motion planning operating in unknown environments. The challenge is that sparse sonar data causes incomplete obstacle boundaries, posing risks to the vehicle's safe navigation

Cited by 0SourceScholar
2025

A Fast-Adaptive Cognitive Diagnosis Framework for Computerized Adaptive Testing Systems

IJCAI 2025

Computerized Adaptive Testing (CAT) measures student ability by iteratively selecting informative questions, with core components being the Cognitive Diagnosis Model (CDM) and selection strategy. Current research focuses on optimizing the selection strategy, assuming relatively accurate CDM results.

2025

Constrained Offline Black-Box Optimization via Risk Evaluation and Management

AAAI 2025technical

Offline black-box optimization aims to identify the optimal solution of a black-box objective function under the guidance of a surrogate model constructed solely from a pre-collected dataset. It is commonly used in industrial scenarios, which often involve constraints, i.e., constrained offline opti…

Cited by 1SourcePDFScholar
2025

Effective and Efficient Representation Learning for Flight Trajectories

AAAI 2025technical

Flight trajectory data plays a vital role in the traffic management community, especially for downstream tasks such as trajectory prediction, flight recognition, and anomaly detection. Existing works often utilize handcrafted features and design models for different tasks individually, which heavily…

2025

Human-Robot Cooperative Heavy Payload Manipulation based on Whole-Body Model Predictive Control

IROS 2025

Human-robot collaborative manipulation with mobile, multiple manipulators is crucial for expanding robotic applications, requiring precise handling of coupled force-position constraints between partners. Current systems, however, exhibit end-effector oscillations and instability during dynamic inter

Cited by 1SourceScholar
2025

MVINS: A Magnetism&vision Aided Inertial Navigation System for Autonomous Underwater Vehicles

RA-L 2025

We present a robust underwater navigation system that integrates magnetic, visual, and inertial measurements from commercial off-the-shelf sensors. Visual Inertial Navigation Systems (VINS) face challenges when used for Autonomous Underwater Vehicle (AUV) localization in perceptually degraded enviro

Cited by 3SourceScholar
2025

OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation

CVPR 2025poster

Multimodal Large Language Models (MLLMs) have made significant strides in visual understanding and generation tasks. However, generating interleaved image-text content remains a challenge, which requires integrated multimodal understanding and generation abilities. While the progress in unified mode…

2025

Relation-Augmented Dueling Bayesian Optimization via Preference Propagation

IJCAI 2025

In black-box optimization, when directly evaluating the function values of solutions is very costly or infeasible, access to the objective function is often limited to comparing pairs of solutions, which yields dueling black-box optimization. Dueling optimization is solely based on pairwise preferen

2025

SOO-Bench: Benchmarks for Evaluating the Stability of Offline Black-Box Optimization

ICLR 2025poster

Black-box optimization aims to find the optima through building a model close to the black-box objective function based on function value evaluation. However, in many real-world tasks, such as the design of molecular formulas and mechanical structures, it is perilous, costly, or even infeasible to e…

2025

VIMS: A Visual-Inertial-Magnetic-Sonar SLAM System in Underwater Environments

IROS 2025

In this study, we present a novel simultaneous localization and mapping (SLAM) system, VIMS, designed for underwater navigation. Conventional visual-inertial state estimators encounter significant practical challenges in perceptually degraded underwater environments, particularly in scale estimation

Cited by 0SourceScholar
2024

A Simple yet Scalable Granger Causal Structural Learning Approach for Topological Event Sequences

NeurIPS 2024poster

In modern telecommunication networks, faults manifest as alarms, generating thousands of events daily. Network operators need an efficient method to identify the root causes of these alarms to mitigate potential losses. This task is challenging due to the increasing scale of telecommunication networ…

Cited by 0SourcePDFScholar
2024

ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Ablation Capability for Large Vision-Language Models

NeurIPS 2024spotlight

Multi-turn visual conversation is an important ability of real-world AI assistants. However, the related evaluation benchmark is missed. This paper presents ConvBench, a multi-turn conversation benchmark with hierarchical capabilities ablation evaluation for Large Vision-Language Models (LVLMs). Co…

2024

MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

ICML 2024poster

Large Vision-Language Models (LVLMs) show significant strides in general-propose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited number of multimodal tasks testing rudimentary capabilities, falling short in t…

Cited by 84SourcePDFScholar
2024

Needle In A Multimodal Haystack

NeurIPS 2024poster

With the rapid advancement of multimodal large language models (MLLMs), their evaluation has become increasingly comprehensive. However, understanding long multimodal content, as a foundational ability for real-world applications, remains underexplored. In this work, we present Needle In A Multimoda…

2024

SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge

NeurIPS 2024poster

Large vision-language models (LVLMs) are ignorant of the up-to-date knowledge, such as LLaVA series, because they cannot be updated frequently due to the large amount of resources required, and therefore fail in many cases. For example, if a LVLM was released on January 2024, and it wouldn't know th…

Cited by 2SourcePDFScholar
2022

Bi-Directional Modality Fusion Network For Audio-Visual Event Localization

ICASSP 2022accepted

Audio and visual signals stimulate many audio-visual sensory neurons of persons to generate audio-visual contents, helping humans perceive the world. Most of the existing audio-visual event localization approaches focus on generating audio-visual features by fusing the audio and visual modalities fo…

Cited by 0SourceScholar
2022

Convoluational Transformer With Adaptive Position Embedding For Covid-19 Detection From Cough Sounds

ICASSP 2022accepted

Covid-19 has caused a huge health crisis worldwide in the past two years. Although an early detection of the virus through nucleic acid screening can considerably reduce its spread, the efficiency of this diagnostic process is limited by its complexity and costs. Hence, an effective and inexpensive…

Cited by 7SourceScholar
2021

A Novel Attention-Based Gated Recurrent Unit and its Efficacy in Speech Emotion Recognition

ICASSP 2021accepted

Notwithstanding the significant advancements in the field of deep learning, the basic long short-term memory (LSTM) or Gated Recurrent Unit (GRU) units have largely remained unchanged and unexplored. There are several possibilities in advancing the state-of-art by rightly adapting and enhancing the…

Cited by 0SourceScholar
2017

Grasp quality evaluation and planning for objects with negative curvature

ICRA 2017poster

We consider the problem of grasping concave objects, i.e., objects whose surface includes regions with negative curvature. When a multifingered hand is used to restrain these objects, these areas can be advantageously used to determine grasps capable of more robustly resisting to external disturbanc…

Cited by 3SourceScholar