← Search

Yura Choi

3 accepted papers

2026

Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering

CVPR 2026

Understanding and answering questions based on a user's pointing gesture is essential for next-generation egocentric AI assistants. However, current Multimodal Large Language Models (MLLMs) struggle with such tasks due to the lack of gesture-rich data and their limited ability to infer fine-grained

Cited by 0SourceScholar
2025

ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO

AAAI 2025technical

Iterative self-improvement, a concept extending beyond personal growth, has found powerful applications in machine learning, particularly in transforming weak models into strong ones. While recent advances in natural language processing have shown its efficacy through iterative preference optimizati…

Cited by 0SourcePDFScholar
2024

Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback

ACL 2024long

Recent advancements in large language models have influenced the development of video large multimodal models (VLMMs). Previous approaches for VLMMs involve Supervised Fine-Tuning (SFT) with instruction-tuned datasets, integrating LLM with visual encoders, and additional learnable parameters. Here,…