← Search

Zhiyuan Feng

11 accepted papers

2026

AD-MIR: Bridging the Gap from Perception to Persuasion in Advertising Video Understanding via Structured Reasoning

ICML 2026poster

Multimodal understanding of advertising videos is essential for interpreting the intricate relationship between visual storytelling and abstract persuasion strategies. However, despite excelling at general search, existing agents often struggle to bridge the cognitive gap between pixel-level percept…

Cited by 0SourceScholar
2026

From Human Videos to Robot Manipulation: A Survey on Action-Relevant Representation Transfer for Scalable Vision-Language-Action Learning

IJCAI 2026

Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision–Language–Action (VLA) models. However, most existing approaches rely on large collections of robot demonstrations, which are costly to obtain and tightly coupled to specific embodiments. Human vide

Cited by 0Scholar
2026

HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains. However, current benchmarks for evaluating MLLMs primarily emphasize general knowledge and vertical step-by-step reasoning typical of STEM disciplines, while overlooking the distinct…

Cited by 0SourcecodeScholar
2026

HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models

CVPR 2026

Achieving human-like spatial intelligence for vision-language models (VLMs) requires inferring 3D structures from 2D observations, recognizing object properties and relations in 3D space, and performing high-level spatial reasoning. In this paper, we propose a principled hierarchical framework that

Cited by 0SourcecodeScholar
2026

MULTIMODAL MULTI-AGENT EMPOWERED LEGAL JUDGMENT PREDICTION

ICASSP 2026poster

Legal Judgment Prediction (LJP) aims to predict the outcomes of legal cases based on factual descriptions, serving as a fundamental task to advance the development of legal systems. Traditional methods often rely on statistical analyses or role-based simulations but face challenges with multiple all…

Cited by 0SourcePDFScholar
2026

Revisual-R1: Advancing Multimodal Reasoning From Optimized Cold Start to Staged Reinforcement Learning

ICLR 2026poster

Inspired by the remarkable reasoning capabilities of Deepseek-R1 in complex textual tasks, many works attempt to incentivize similar capabilities in Multimodal Large Language Models (MLLMs) by directly applying reinforcement learning (RL). However, they still struggle to activate complex reasoning.…

Cited by 0SourcecodeScholar
2026

Scalable Vision-Language-Action Model Pretraining for Robotic Dexterous Manipulation with Real-Life Human Activity Videos

ICRA 2026poster

This paper presents an approach for pretraining robotic manipulation Vision-Language-Action (VLA) models using a large corpus of unscripted real-life video recordings of human hand activities. Treating human hand as dexterous robot end-effector, we show that "in-the-wild" egocentric human videos wit…

Cited by 0Scholar
2026

Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes

ICLR 2026poster

Vision-language models (VLMs) are essential to Embodied AI, enabling robots to perceive, reason, and act in complex environments. They also serve as the foundation for the recent Vision-Language-Action (VLA) models. Yet, most evaluations of VLMs focus on single-view settings, leaving their ability t…

Cited by 0SourcecodeScholar
2025

MIRA: Medical Time Series Foundation Model for Real-World Health Data

NeurIPS 2025poster

A unified foundation model for medical time series—pretrained on open access and ethically reviewed medical corpora—offers the potential to reduce annotation burdens, minimize model customization, and enable robust transfer across clinical institutions, modalities, and tasks, particularly in data-sc…

Cited by 0SourceScholar
2025

TransDiff: Diffusion-Based Method for Manipulating Transparent Objects Using a Single RGB-D Image

ICRA 2025

Manipulating transparent objects presents significant challenges due to the complexities introduced by their reflection and refraction properties, which considerably hinder the accurate estimation of their 3D shapes. To address these challenges, we propose a single-view RGB-D-based depth completion

Cited by 3SourcecodeScholar
2024

The Continuous Jump Control of a Locust-Inspired Robot With Omnidirectional Trajectory Adjustment

RA-L 2024

Jumping is an effective way for small robots to overcome obstacles. After years of development, many miniature jumping robots have been proposed with various mechanisms, and they have achieved jump trajectory control, fall recovery, and even continuous jumps. However, most miniature jumping robots d

Cited by 7SourceScholar