← Search

Ryo Hachiuma

19 accepted papers

2026

4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation

CVPR 2026

Despite advances in Multimodal LLMs (MLLMs), their ability to reason over 3D structures and temporal dynamics remains limited, constrained by weak 4D perception and temporal understanding. Existing 3D and 4D Video Question Answering (VQA) benchmarks also emphasize static scenes and lack region-level

Cited by 0SourcecodeScholar
2026

Interpretable Debiasing of Vision-Language Models for Social Fairness

CVPR 2026

The rapid advancement of Vision-Language models (VLMs) has raised growing concerns that their black-box reasoning processes could lead to unintended forms of social bias. Current debiasing approaches focus on mitigating surface-level bias signals through post-hoc learning or test-time algorithms, wh

Cited by 0SourceScholar
2026

Learning from Synthetic Data via Provenance-Based Input Gradient Guidance

CVPR 2026

Learning methods using synthetic data have attracted attention as an effective approach for increasing the diversity of training data while reducing collection costs, thereby improving the robustness of model discrimination. However, many existing methods improve robustness only indirectly through t

Cited by 0SourcecodeScholar
2026

V2V-GoT: Vehicle-To-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models and Graph-Of-Thoughts

ICRA 2026poster

Current state-of-the-art autonomous vehicles could face safety critical situations when their local sensors are occluded by large objects on the road nearby. Vehicle-to-vehicle (V2V) cooperative autonomous driving is proposed to address this problem. More recent work further adopts a new approach th…

2026

V2V-LLM: Vehicle-To-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models

ICRA 2026poster

Current autonomous driving vehicles rely mainly on their individual sensors to understand surrounding scenes and plan for future trajectories, which can be unreliable when the sensors are malfunctioning or occluded. To address this problem, cooperative perception methods via vehicle-to-vehicle (V2V)…

2025

Bias in Gender Bias Benchmarks: How Spurious Features Distort Evaluation

ICCV 2025poster

Gender bias in vision-language foundation models (VLMs) raises concerns about their safe deployment and is typically evaluated using benchmarks with gender annotations on real-world images. However, as these benchmarks often contain spurious correlations between gender and non-gender features, such…

Cited by 0SourcePDFScholar
2025

Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks

CVPR 2025poster

We present Omni-RGPT, a multimodal large language model designed to facilitate region-level comprehension for both images and videos. To achieve consistent region representation across spatio-temporal dimensions, we introduce Token Mark, a set of tokens highlighting the target regions within the vis…

Cited by 2SourcePDFScholar
2025

SANER: Annotation-free Societal Attribute Neutralizer for Debiasing CLIP

ICLR 2025poster

Large-scale vision-language models, such as CLIP, are known to contain societal bias regarding protected attributes (e.g., gender, age). This paper aims to address the problems of societal bias in CLIP. Although previous studies have proposed to debias societal bias through adversarial learning or t…

Cited by 2SourcePDFScholar
2025

Unified Reinforcement and Imitation Learning for Vision-Language Models

NeurIPS 2025poster

Vision-Language Models (VLMs) have achieved remarkable progress, yet their large scale often renders them impractical for resource-constrained environments. This paper introduces Unified Reinforcement and Imitation Learning (RIL), a novel and efficient training algorithm designed to create powerful,…

Cited by 0SourceScholar
2025

VLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language Models

CVPR 2025poster

The recent surge in high-quality visual instruction tuning samples from closed-source vision-language models (VLMs) such as GPT-4V has accelerated the release of open-source VLMs across various model sizes. However, scaling VLMs to improve performance using larger models brings significant computati…

Cited by 0SourcePDFScholar
2024

From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment

EMNLP 2024main

Large language models (LLMs) have enhanced the capacity of vision-language models to caption visual text. This generative approach to image caption enrichment further makes textual captions more descriptive, improving alignment with the visual context. However, while many studies focus on the benefi…

Cited by 3SourcePDFScholar
2024

Multimodal Cross-Domain Few-Shot Learning for Egocentric Action Recognition

ECCV 2024poster

"We address a novel cross-domain few-shot learning task (CD-FSL) with multimodal input and unlabeled target data for egocentric action recognition. This paper simultaneously tackles two critical challenges associated with egocentric action recognition in CD-FSL settings: (1) the extreme domain gap i…

Cited by 6SourcePDFScholar
2023

Prompt-Guided Zero-Shot Anomaly Action Recognition Using Pretrained Deep Skeleton Features

CVPR 2023poster

This study investigates unsupervised anomaly action recognition, which identifies video-level abnormal-human-behavior events in an unsupervised manner without abnormal samples, and simultaneously addresses three limitations in the conventional skeleton-based approaches: target domain-dependent DNN t…

2023

Unified Keypoint-Based Action Recognition Framework via Structured Keypoint Pooling

CVPR 2023poster

This paper simultaneously addresses three limitations associated with conventional skeleton-based action recognition; skeleton detection and tracking errors, poor variety of the targeted actions, as well as person-wise and frame-wise action recognition. A point cloud deep-learning paradigm is introd…

Cited by 55SourcePDFScholar
2021

Dynamics-regulated kinematic policy for egocentric pose estimation

NeurIPS 2021poster

We propose a method for object-aware 3D egocentric pose estimation that tightly integrates kinematics modeling, dynamics modeling, and scene object information. Unlike prior kinematics or dynamics-based approaches where the two components are used disjointly, we synergize the two approaches via dyna…