← Search

Ming-Chang Chiu

4 accepted papers

2026

MegaCoin: Enhancing Medium-Grained Color Perception for Vision-Language Models

AAAI 2026technical

In vision-language models (VLMs), the ability to perceive and interpret color and physical environment is crucial for achieving contextually accurate understanding and interaction. However, despite advances in multimodal modeling, there remains a significant lack of specialized datasets that rigorou

Cited by 0SourcePDFScholar
2024

DDI-CoCo: A Dataset for Understanding the Effect of Color Contrast in Machine-Assisted Skin Disease Detection

ICASSP 2024accepted

Skin tone as a demographic bias and inconsistent human labeling poses challenges in dermatology AI. We take another angle to investigate color contrast’s impact, beyond skin tones, on malignancy detection in skin disease datasets: We hypothesize that in addition to skin tones, the color difference b…

Cited by 0SourceScholar
2024

VideoPoet: A Large Language Model for Zero-Shot Video Generation

ICML 2024oral

We present VideoPoet, a language model capable of synthesizing high-quality video from a large variety of conditioning signals. VideoPoet employs a decoder-only transformer architecture that processes multimodal inputs -- including images, videos, text, and audio. The training protocol follows that…

Cited by 257SourcePDFScholar
2023

Better May Not Be Fairer: A Study on Subgroup Discrepancy in Image Classification

ICCV 2023poster

In this paper, we provide 20,000 non-trivial human annotations on popular datasets as a first step to bridge gap to studying how natural semantic spurious features affect image classification, as prior works often study datasets mixing low-level features due to limitations in accessing realistic dat…

Cited by 5PDFcodeScholar