← Search

Zhi Hou

9 accepted papers

2026

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy

AAAI 2026technical

Vision-Language-Action (VLA) models frequently encounter challenges in generalizing to real-world environments due to inherent discrepancies between observation and action spaces. Although training data are collected from diverse camera perspectives, the models typically predict end-effector poses w

Cited by 0SourcePDFScholar
2026

Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning

ICLR 2026poster

While significant research has focused on developing embodied reasoning capabilities using Vision-Language Models (VLMs) or integrating advanced VLMs into Vision-Language-Action (VLA) models for end-to-end robot control, few studies directly address the critical gap between upstream VLM-based reason…

Cited by 0SourcecodeScholar
2025

Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

ICCV 2025poster

While recent vision-language-action models trained on diverse robot datasets exhibit promising generalization capabilities with limited in-domain data, their reliance on compact action heads to predict discretized or continuous actions constrains adaptability to heterogeneous action spaces. We prese…

Cited by 0SourcePDFScholar
2025

TimeFormer: Capturing Temporal Relationships of Deformable 3D Gaussians for Robust Reconstruction

ICCV 2025poster

Dynamic scene reconstruction is a long-term challenge in 3D vision. Recent methods extend 3D Gaussian Splatting to dynamic scenes via additional deformation fields and apply explicit constraints like motion flow to guide the deformation. However, they learn motion changes from individual timestamps…

Cited by 0SourcePDFScholar
2022

BatchFormer: Learning To Explore Sample Relationships for Robust Representation Learning

CVPR 2022poster

Despite the success of deep neural networks, there are still many challenges in deep representation learning due to the data scarcity issues such as data imbalance, unseen distribution, and domain shift. To address the above-mentioned issues, a variety of methods have been devised to explore the sam…

Cited by 101PDFcodeScholar
2022

Discovering Human-Object Interaction Concepts via Self-Compositional Learning

ECCV 2022poster

"A comprehensive understanding of human-object interaction (HOI) requires detecting not only a small portion of predefined HOI concepts (or categories) but also other reasonable HOI concepts, while current approaches usually fail to explore a huge portion of unknown HOI concepts (i.e., unknown but r…

2021

Affordance Transfer Learning for Human-Object Interaction Detection

CVPR 2021poster

Reasoning the human-object interactions (HOI) is essential for deeper scene understanding, while object affordances (or functionalities) are of great importance for human to discover unseen HOIs with novel objects. Inspired by this, we introduce an affordance transfer learning approach to jointly de…

Cited by 139PDFcodeScholar
2021

Detecting Human-Object Interaction via Fabricated Compositional Learning

CVPR 2021poster

Human-Object Interaction (HOI) detection, inferring the relationships between human and objects from images/videos, is a fundamental task for high-level scene understanding. However, HOI detection usually suffers from the open long-tailed nature of interactions with objects, while human has extremel…

Cited by 118PDFcodeScholar
2020

Visual Compositional Learning for Human-Object Interaction Detection

ECCV 2020poster

Human-Object interaction (HOI) detection aims to localize and infer relationships between human and objects in an image. It is challenging because an enormous number of possible combinations of objects and verbs types forms a long-tail distribution. We devise a deep Visual Compositional Learning (VC…