← Search

Miso Choi

4 accepted papers

2026

Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding

CVPR 2026

Large Vision-Language Models (LVLMs) have shown strong performance across various multimodal tasks by leveraging the reasoning capabilities of Large Language Models (LLMs). However, processing visually complex and information-rich images, such as infographics or document layouts, requires these mode

Cited by 0SourcecodeScholar
2026

The Truth Stays in the Family: Enhancing Contextual Truthfulness via Inherited Heads in Model Lineages

ICML 2026poster

Recent advances in large language models (LLMs) have led to the emergence of specialized multimodal LLMs (MLLMs), forming distinct model families that share a common foundation language models. Despite this evolutionary trend, it remains unexplored whether a fundamental behavioral link exists betwee…

Cited by 0SourceScholar
2026

Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization

AAAI 2026technical

Vision-Language Models (VLMs) have been widely used in various visual recognition tasks due to their remarkable generalization capabilities. As these models grow in size and complexity, fine-tuning becomes costly, emphasizing the need to reuse adaptation knowledge from

Cited by 0SourcePDFScholar
2023

Open-vocabulary Video Question Answering: A New Benchmark for Evaluating the Generalizability of Video Question Answering Models

ICCV 2023poster

Video Question Answering (VideoQA) is a challenging task that entails complex multi-modal reasoning. In contrast to multiple-choice VideoQA which aims to predict the answer given several options, the goal of open-ended VideoQA is to answer questions without restricting candidate answers. However, th…

Cited by 7PDFcodeScholar