← Search

G. Thomas Hudson

2 accepted papers

2025

Everything is a Video: Unifying Modalities through Next-Frame Prediction

ICCV 2025poster

Multimodal learning, which involves integrating information from various modalities such as text, images, audio, and video, is pivotal for numerous complex tasks like visual question answering, cross-modal retrieval, and caption generation. Traditional approaches rely on modality-specific encoders a…

2023

Length is a Curse and a Blessing for Document-level Semantics

EMNLP 2023long main

In recent years, contrastive learning (CL) has been extensively utilized to recover sentence and document-level encoding capability from pre-trained language models. In this work, we question the length generalizability of CL-based models, i.e., their vulnerability towards length-induced semantic sh…

Cited by 0SourcecodeScholar