← Search

Chan-Wei Hu

4 accepted papers

2025

Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization

EMNLP 2025

The emergence of large Vision Language Models (VLMs) has broadened the scope and capabilities of single-modal Large Language Models (LLMs) by integrating visual modalities, thereby unlocking transformative cross-modal applications in a variety of real-world scenarios. Despite their impressive perfor

2018

Liquid Pouring Monitoring via Rich Sensory Inputs

ECCV 2018poster

Humans have the amazing ability to perform very subtle manipulation task using a closed-loop control system with imprecise mechanics (i.e., our body parts) but rich sensory information (e.g., vision, tactile, etc.). In the closed-loop system, the ability to monitor the state of the task via rich sen…

Cited by 9SourcePDFScholar
2018

Omnidirectional CNN for Visual Place Recognition and Navigation

ICRA 2018poster

Visual place recognition is challenging, especially when only a few place exemplars are given. To mitigate the challenge, we consider place recognition method using omnidirectional cameras and propose a novel Omnidirectional Convolutional Neural Network (O-CNN) to handle severe camera pose variation…

Cited by 88SourceScholar
2017

Anticipating Daily Intention Using On-Wrist Motion Triggered Sensing

ICCV 2017spotlight

Anticipating human intention by observing one's actions has many applications. For instance, picking up a cellphone, then a charger (actions) implies that one wants to charge the cellphone (intention). By anticipating the intention, an intelligent system can guide the user to the closest power outle…

Cited by 32PDFcodeScholar