← Search

Yifan Hou

30 accepted papers

2026

Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal Reasoning

ICLR 2026poster

Multimodal large language models (MLLMs) promise enhanced reasoning by integrating diverse inputs such as text, vision, and audio. Yet, despite their perceptual strengths, their reasoning ability across modalities remains underexplored, with conflicting reports on whether additional modalities help…

Cited by 0SourceScholar
2026

DexMachina: Functional Retargeting for Bimanual Dexterous Manipulation

ICML 2026poster

We study the problem of functional retargeting: learning dexterous manipulation policies to track object states from human hand-object demonstrations. We focus on long-horizon, bimanual tasks with articulated objects, which are challenging due to large action space, spatiotemporal discontinuities, a…

Cited by 0SourcecodeScholar
2026

Diversity Matters: Revisiting Test-Time Compute in Vision-Language Models

ICML 2026poster

Test-time compute (TTC) strategies have emerged as a lightweight approach to boost reasoning in large language models, but their applicability to vision-language models (VLMs) remains unclear. We present a systematic study of TTC for visual reasoning across seven open-source VLMs and six benchmarks,…

Cited by 0SourceScholar
2026

In-The-Wild Compliant Manipulation with UMI-FT

ICRA 2026poster

Many manipulation tasks require careful force modulation. With insufficient force the task may fail, while excessive force could cause damage. The high cost, bulky size and fragility of commercial force/torque (F/T) sensors have limited large-scale, force-aware policy learning. We introduce UMI-FT, …

2026

Unveiling the Visual Counting Bottleneck in Vision-Language Models

ICML 2026poster

While Large Vision-Language Models (VLMs) excel at interpolation, they suffer catastrophic failures in systematic generalization, most notably in visual counting beyond training distributions. In this work, we investigate this extrapolation bottleneck by deconstructing visual counting into three cog…

Cited by 0SourceScholar
2025

Adaptive Compliance Policy: Learning Approximate Compliance for Diffusion Guided Control

ICRA 2025

Compliance plays a crucial role in manipulation, as it balances between the concurrent control of position and force under uncertainties. Yet compliance is often overlooked by today's visuomotor policies that solely focus on position control. This paper introduces Adaptive Compliance Policy (ACP), a

Cited by 57SourcecodeScholar
2025

Can Vision-Language Models Solve Visual Math Equations?

EMNLP 2025

Despite strong performance in visual understanding and language-based reasoning, Vision-Language Models (VLMs) struggle with tasks requiring integrated perception and symbolic computation. We study this limitation through visual equation solving, where mathematical equations are embedded in images,

Cited by 0SourcePDFScholar
2025

Compliant Residual DAgger: Improving Real-World Contact-Rich Manipulation with Human Corrections

NeurIPS 2025poster

We address key challenges in Dataset Aggregation (DAgger) for real-world contact- rich manipulation: how to collect informative human correction data and how to effectively update policies with this new data. We introduce Compliant Residual DAgger (CR-DAgger), which contains two novel components: 1)…

Cited by 0SourceScholar
2025

DexUMI: Using Human Hand as the Universal Manipulation Interface for Dexterous Manipulation

CoRL 2025oral

We present DexUMI - a data collection and policy learning framework that uses the human hand as the natural interface to transfer dexterous manipulation skills to various robot hands. DexUMI incorporates hardware and software adaptations to minimize the embodiment gap between the human hand and vari…

Cited by 0SourceScholar
2025

Do Vision-Language Models Really Understand Visual Language?

ICML 2025poster

Visual language is a system of communication that conveys information through symbols, shapes, and spatial arrangements. Diagrams are a typical example of a visual language depicting complex concepts and their relationships in the form of an image. The symbolic nature of diagrams presents significan…

Cited by 2SourcePDFScholar
2025

Explore the Reasoning Capability of LLMs in the Chess Testbed

NAACL 2025short

Reasoning is a central capability of human intelligence. In recent years, with the advent of large-scale datasets, pretrained large language models have emerged with new capabilities, including reasoning. However, these models still struggle with long-term, complex reasoning tasks, such as playing c…

Cited by 1SourcePDFScholar
2025

HealthCards: Exploring Text-to-Image Generation as Visual Aids for Healthcare Knowledge Democratizing and Education

EMNLP 2025

The evolution of text-to-image (T2I) generation techniques has introduced new capabilities for information visualization, with the potential to advance knowledge democratization and education. In this paper, we investigate how T2I models can be adapted to generate educational health knowledge conten

2025

Vision in Action: Learning Active Perception from Human Demonstrations

CoRL 2025poster

We present Vision in Action (ViA), an active perception system for bimanual robot manipulation. ViA learns task-relevant active perceptual strategies (e.g., searching, tracking, and focusing) directly from human demonstrations. On the hardware side, ViA employs a simple yet effective 6-DoF robotic n…

Cited by 0SourceScholar
2024

Unveiling the Art of Heading Design: A Harmonious Blend of Summarization, Neology, and Algorithm

ACL 2024findings

Crafting an appealing heading is crucial for attracting readers and marketing work or products. A popular way is to summarize the main idea with a refined description and a memorable acronym. However, there lacks a systematic study and a formal benchmark including datasets and metrics. Motivated by…

Cited by 1SourcePDFScholar
2024

What Do Language Models Learn in Context? The Structured Task Hypothesis.

ACL 2024long

Large language models (LLMs) exhibit an intriguing ability to learn a novel task from in-context examples presented in a demonstration, termed in-context learning (ICL). Understandably, a swath of research has been dedicated to uncovering the theories underpinning ICL. One popular hypothesis explain…

2023

Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Models

EMNLP 2023long main

Recent work has shown that language models (LMs) have strong multi-step (i.e., procedural) reasoning capabilities. However, it is unclear whether LMs perform these tasks by cheating with answers memorized from pretraining corpus, or, via a multi-step reasoning mechanism. In this paper, we try to ans…

Cited by 0SourcecodeScholar
2022

Adapters for Enhanced Modeling of Multilingual Knowledge and Text

EMNLP 2022finding

Large language models appear to learn facts from the large text corpora they are trained on. Such facts are encoded implicitly within their many parameters, making it difficult to verify or manipulate what knowledge has been learned. Language models have recently been extended to multilingual langua…

2022

Contact Mode Guided Motion Planning for Quasidynamic Dexterous Manipulation in 3D

ICRA 2022poster

This paper presents Contact Mode Guided Manipulation Planning (CMGMP) for 3D quasistatic and quasi-dynamic rigid body motion planning in dexterous manipulation. The CMGMP algorithm generates hybrid motion plans including both continuous state transitions and discrete contact mode switches, without t…

Cited by 62SourceScholar
2022

Extrinsic Dexterous Manipulation with a Direct-drive Hand: A Case Study

IROS 2022poster

This paper explores a novel approach to dexterous manipulation, aimed at levels of speed, precision, robustness, and simplicity suitable for practical deployment. The enabling technology is a Direct-drive Hand (DDHand) comprising two fingers, two DOFs each, that exhibit high speed and a light touch.…

Cited by 6SourceScholar
2021

Bird’s Eye: Probing for Linguistic Graph Structures with a Simple Information-Theoretic Approach

ACL 2021long

NLP has a rich history of representing our prior understanding of language in the form of graphs. Recent work on analyzing contextualized text representations has focused on hand-designed probe models to understand how and to what extent do these representations encode a particular linguistic phenom…

2021

Contact Mode Guided Sampling-Based Planning for Quasistatic Dexterous Manipulation in 2D

ICRA 2021poster

The discontinuities and multi-modality introduced by contacts make manipulation planning challenging. Many previous works avoid this problem by pre-designing a set of high-level motion primitives like grasping and pushing. However, such motion primitives are often not adequate to describe dexterous…

Cited by 49SourceScholar
2020

Measuring and Improving the Use of Graph Information in Graph Neural Networks

ICLR 2020poster

Graph neural networks (GNNs) have been widely used for representation learning on graph data. However, there is limited understanding on how much performance GNNs actually gain from graph data. This paper introduces a context-surrounding GNN framework and proposes two smoothness metrics to measure t…

Cited by 0SourcecodeScholar