← Search

Yi-Chia Chen

2 accepted papers

2026

VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction

AAAI 2026technical

Recent advances in multimodal large language models (MLLMs) have significantly enhanced video understanding capabilities, opening new possibilities for practical applications. Yet current video benchmarks focus largely on indoor scenes or short-range outdoor activities, leaving the challenges associ

Cited by 0SourcePDFScholar
2024

SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentation

ECCV 2024poster

"We introduce SAM4MLLM, an innovative approach which integrates the Segment Anything Model (SAM) with Multi-Modal Large Language Models (MLLMs) for pixel-aware tasks. Our method enables MLLMs to learn pixel-level location information without requiring excessive modifications to the existing model ar…