← Search

Yong Cheng

12 accepted papers

2024

Language Model Beats Diffusion - Tokenizer is key to visual generation

ICLR 2024poster

While Large Language Models (LLMs) are the dominant models for generative tasks in language, they do not perform as well as diffusion models on image and video generation. To effectively use LLMs for visual generation, one crucial component is the visual tokenizer that maps pixel-space inputs to dis…

Cited by 296SourcePDFScholar
2024

VideoPoet: A Large Language Model for Zero-Shot Video Generation

ICML 2024oral

We present VideoPoet, a language model capable of synthesizing high-quality video from a large variety of conditioning signals. VideoPoet employs a decoder-only transformer architecture that processes multimodal inputs -- including images, videos, text, and audio. The training protocol follows that…

Cited by 257SourcePDFScholar
2023

MAGVIT: Masked Generative Video Transformer

CVPR 2023highlight

We introduce the MAsked Generative VIdeo Transformer, MAGVIT, to tackle various video synthesis tasks with a single model. We introduce a 3D tokenizer to quantize a video into spatial-temporal visual tokens and propose an embedding method for masked video token modeling to facilitate multi-task lear…

2023

Mu$^2$SLAM: Multitask, Multilingual Speech and Language Models

ICML 2023oral

We present Mu$^2$SLAM, a multilingual sequence-to-sequence model pre-trained jointly on unlabeled speech, unlabeled text and supervised data spanning Automatic Speech Recognition (ASR), Automatic Speech Translation (AST) and Machine Translation (MT), in over 100 languages. By leveraging a quantized…

Cited by 21SourcePDFScholar
2023

SPAE: Semantic Pyramid AutoEncoder for Multimodal Generation with Frozen LLMs

NeurIPS 2023spotlight

In this work, we introduce Semantic Pyramid AutoEncoder (SPAE) for enabling frozen LLMs to perform both understanding and generation tasks involving non-linguistic modalities such as images or videos. SPAE converts between raw pixels and interpretable lexical tokens (or words) extracted from the LLM…

Cited by 59SourcePDFScholar
2022

Examining Scaling and Transfer of Language Model Architectures for Machine Translation

ICML 2022spotlight

Natural language understanding and generation models follow one of the two dominant architectural paradigms: language models (LMs) that process concatenated sequences in a single stack of layers, and encoder-decoder models (EncDec) that utilize separate layer stacks for input and output processing.…

Cited by 21SourcePDFScholar
2022

Multilingual Mix: Example Interpolation Improves Multilingual Neural Machine Translation

ACL 2022long

Multilingual neural machine translation models are trained to maximize the likelihood of a mix of examples drawn from multiple language pairs. The dominant inductive bias applied to these models is a shared vocabulary and a shared set of parameters across languages; the inputs and labels correspondi…

Cited by 17SourcePDFScholar
2021

Self-supervised and Supervised Joint Training for Resource-rich Machine Translation

ICML 2021spotlight

Self-supervised pre-training of text representations has been successfully applied to low-resource Neural Machine Translation (NMT). However, it usually fails to achieve notable gains on resource-rich NMT. In this paper, we propose a joint training approach, F2-XEnDec, to combine self-supervised and…

Cited by 18SourcePDFScholar
2018

Development and Error Compensation of a Flexible Multi-Joint Manipulator Applied in Nuclear Fusion Environment

IROS 2018poster

Experimental Advanced Superconducting Tokamak (EAST) is the world's first fully superconducting tokamak fusion device with non-circular cross-section which was built in China The EAST articulated maintenance arm (EAMA) system is developed for real-time detection and rapid repair operations to damage…

Cited by 4SourceScholar
2017

Hybrid beamforming for large-scale MIMO systems using uplink-downlink duality

ICASSP 2017accepted

We consider the problem of designing hybrid analog-digital beamformers in a downlink multi-user large-scale MIMO system. The objective is to minimize the total transmit power, while fulfilling SINR targets of all users. A dual virtual uplink problem is formulated for the original downlink problem ba…

Cited by 0SourceScholar
2016

Optimal resource block allocation and muting in heterogeneous networks

ICASSP 2016accepted

In this paper, we investigate user association and resource block (RB) allocation in downlink heterogeneous cellular networks. Our goal is to jointly optimize user association and RB allocation to maximize network throughput while taking into account fairness among users. To effectively control inte…

Cited by 0SourceScholar