← Search

Zineng Tang

9 accepted papers

2024

CoDi-2: In-Context Interleaved and Interactive Any-to-Any Generation

CVPR 2024highlight

We present CoDi-2 a Multimodal Large Language Model (MLLM) for learning in-context interleaved multimodal representations. By aligning modalities with language for both encoding and generation CoDi-2 empowers Large Language Models (LLMs) to understand modality-interleaved instructions and in-context…

Cited by 54SourcePDFScholar
2023

Any-to-Any Generation via Composable Diffusion

NeurIPS 2023poster

We present Composable Diffusion (CoDi), a novel generative model capable of generating any combination of output modalities, such as language, image, video, or audio, from any combination of input modalities. Unlike existing generative AI systems, CoDi can generate multiple modalities in parallel an…

2023

Paxion: Patching Action Knowledge in Video-Language Foundation Models

NeurIPS 2023spotlight

Action knowledge involves the understanding of textual, visual, and temporal aspects of actions. We introduce the **Action Dynamics Benchmark (ActionBench)** containing two carefully designed probing tasks: Action Antonym and Video Reversal, which targets multimodal alignment capabilities and tempor…

2023

Unifying Vision, Text, and Layout for Universal Document Processing

CVPR 2023highlight

We propose Universal Document Processing (UDOP), a foundation Document AI model which unifies text, image, and layout modalities together with varied task formats, including document understanding and generation. UDOP leverages the spatial correlation between textual content and document image to mo…

2021

DeCEMBERT: Learning from Noisy Instructional Videos via Dense Captions and Entropy Minimization

NAACL 2021long

Leveraging large-scale unlabeled web videos such as instructional videos for pre-training followed by task-specific finetuning has become the de facto approach for many video-and-language tasks. However, these instructional videos are very noisy, the accompanying ASR narrations are often incomplete,…

2021

VidLanKD: Improving Language Understanding via Video-Distilled Knowledge Transfer

NeurIPS 2021poster

Since visual perception can give rich information beyond text descriptions for world understanding, there has been increasing interest in leveraging visual grounding for language learning. Recently, vokenization (Tan and Bansal, 2020) has attracted attention by using the predictions of a text-to-ima…