← Search

Ziyu Liu

16 accepted papers

2026

ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning

CVPR 2026

Reward models are critical for aligning vision-language systems with human preferences, yet current approaches suffer from hallucination, weak visual grounding, and an inability to use tools for verification, limiting their reliability on complex multimodal reasoning tasks.We present **ARM-Thinker**

Cited by 0SourcecodeScholar
2026

Exploiting Vulnerabilities: Universal Adversarial Attacks on Vision-Language-Action Models in Robotics

ICRA 2026poster

近年来,视觉-语言-行动(VLA)模型通过无缝整合视觉感知、语言理解和动作生成,在端到端的学习框架中彻底革新了机器人作。然而,由于这些模型设计为直接与物理世界和人类交互,其安全性至关重要,即使是小漏洞也可能导致灾难性故障。在本研究中,我们提出了通用对抗对象,这是一种表面纹理优化的球体,当置于机器人视野内时,任务成功率会显著降低。具体来说,我们的方法引入了一个多层次攻击框架,能够共同干扰轨迹规划、任务执行和动作控制。我们在模拟和现实机器人环境中验证了我们的方法。实验结果表明,对抗对象在两种代表性VLA模型(Pi0和RDT&#

Cited by 0Scholar
2026

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning

CVPR 2026

When video reasoning requires external knowledge, many systems with large multimodal models (LMMs) adopt retrieval augmentation to supply the missing context. Appending textual or multi-clip evidence, however, forces heterogeneous signals into a single attention space. We observe diluted attention a

Cited by 0SourceScholar
2026

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning

CVPR 2026

Video reasoning has advanced with large multimodal models (LMMs), yet their inference is often a single pass that returns an answer without verifying whether the reasoning is evidence-aligned. We introduce **Reinforce to Learn, Elect to Reason (RLER)**, a dual paradigm that decouples learning to pro

Cited by 0SourceScholar
2026

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience

ICML 2026poster

Repurposing large vision-language models (LVLMs) as computer use agents (CUAs) has led to substantial breakthroughs, primarily driven by human-labeled data. However, these models often struggle with novel and specialized software, particularly in scenarios lacking human annotations. To address this …

Cited by 0SourceScholar
2025

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

ACL 2025finding

Despite the promising performance of Large Vision Language Models (LVLMs) in visual understanding, they occasionally generate incorrect outputs. While reward models (RMs) with reinforcement learning or test-time scaling offer the potential for improving generation quality, a critical gap remains: pu…

2025

MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models

ICLR 2025poster

Visual preference alignment involves training Large Vision-Language Models (LVLMs) to predict human preferences between visual inputs. This is typically achieved by using labeled datasets of chosen/rejected pairs and employing optimization algorithms like direct preference optimization (DPO). Existi…

2025

Swin-VasMamba: A Topologically Constrained Model For 3D Vascular Segmentation

ICASSP 2025accepted

Accurate 3D vascular segmentation is essential for diagnosing and treating vascular diseases. This task remains challenging due to the complexity of the 3D data and the morphological diversity of blood vessels. In recent years, state space models (SSMs) have received a great attention for its good p…

Cited by 0SourceScholar
2024

MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

NeurIPS 2024poster

Generating natural and meaningful responses to communicate with multi-modal human inputs is a fundamental capability of Large Vision-Language Models (LVLMs). While current open-source LVLMs demonstrate promising performance in simplified scenarios such as single-turn single-image input, they fall sh…

2024

MMLONGBENCH-DOC: Benchmarking Long-context Document Understanding with Visualizations

NeurIPS 2024spotlight

Understanding documents with rich layouts and multi-modal components is a long-standing and practical task. Recent Large Vision-Language Models (LVLMs) have made remarkable strides in various tasks, particularly in single-page document understanding (DU). However, their abilities on long-context DU…

Cited by 27SourcePDFScholar
2021

The Distributed Discrete Gaussian Mechanism for Federated Learning with Secure Aggregation

ICML 2021spotlight

We consider training models on private data that are distributed across user devices. To ensure privacy, we add on-device noise and use secure aggregation so that only the noisy sum is revealed to the server. We present a comprehensive end-to-end system, which appropriately discretizes the data and…

2021

The Skellam Mechanism for Differentially Private Federated Learning

NeurIPS 2021poster

We introduce the multi-dimensional Skellam mechanism, a discrete differential privacy mechanism based on the difference of two independent Poisson random variables. To quantify its privacy guarantees, we analyze the privacy loss distribution via a numerical evaluation and provide a sharp bound on th…

2021

Transfer Learning via Optimal Transportation for Integrative Cancer Patient Stratification

IJCAI 2021poster

The Stratification of early-stage cancer patients for the prediction of clinical outcome is a challenging task since cancer is associated with various molecular aberrations. A single biomarker often cannot provide sufficient information to stratify early-stage patients effectively. Understanding the…

Cited by 5SourcePDFScholar
2020

Disentangling and Unifying Graph Convolutions for Skeleton-Based Action Recognition

CVPR 2020oral

Spatial-temporal graphs have been widely used by skeleton-based action recognition algorithms to model human action dynamics. To capture robust movement patterns from these graphs, long-range and multi-scale context aggregation and spatial-temporal dependency modeling are critical aspects of a power…

Cited by 1312PDFcodeScholar
2020

Learning Implicit Credit Assignment for Cooperative Multi-Agent Reinforcement Learning

NeurIPS 2020poster

We present a multi-agent actor-critic method that aims to implicitly address the credit assignment problem under fully cooperative settings. Our key motivation is that credit assignment among agents may not require an explicit formulation as long as (1) the policy gradients derived from a centralize…

2020

TLPG-Tracker: Joint Learning of Target Localization and Proposal Generation for Visual Tracking

IJCAI 2020poster

Target localization and proposal generation are two essential subtasks in generic visual tracking, and it is a challenge to address both the two efficiently. In this paper, we propose an efficient two-stage architecture which makes full use of the complementarity of two subtasks to achieve robust lo…

Cited by 0SourcePDFScholar