← Search

Liangliang Cao

17 accepted papers

2025

Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention

ICML 2025poster

In recent years there have been remarkable breakthroughs in image-to-video generation. However, the 3D consistency and camera controllability of generated frames have remained unsolved. Recent studies have attempted to incorporate camera control into the generation process, but their results are oft…

Cited by 8SourcePDFScholar
2025

Model Steering: Learning with a Reference Model Improves Generalization Bounds and Scaling Laws

ICML 2025spotlight

This paper formalizes an emerging learning paradigm that uses a trained model as a reference to guide and enhance the training of a target model through strategic data selection or weighting, named **model steering**. While ad-hoc methods have been used in various contexts, including the training of…

2024

Efficient-3Dim: Learning a Generalizable Single-image Novel-view Synthesizer in One Day

ICLR 2024poster

The task of novel view synthesis aims to generate unseen perspectives of an object or scene from a limited set of input images. Nevertheless, synthesizing novel views from a single image remains a significant challenge. Previous approaches tackle this problem by adopting mesh prediction, multi-plane…

Cited by 0SourcePDFScholar
2024

Ferret: Refer and Ground Anything Anywhere at Any Granularity

ICLR 2024spotlight

We introduce Ferret, a new Multimodal Large Language Model (MLLM) capable of understanding spatial referring of any shape or granularity within an image and accurately grounding open-vocabulary descriptions. To unify referring and grounding in the LLM paradigm, Ferret employs a novel and powerful hy…

2023

STAIR: Learning Sparse Text and Image Representation in Grounded Tokens

EMNLP 2023long main

Image and text retrieval is one of the foundational tasks in the vision and language domain with multiple real-world applications. State-of-the-art contrastive approaches, e.g. CLIP, ALIGN, represent images and texts as dense embeddings and calculate the similarity in the dense embedding space as th…

Cited by 0SourceScholar
2022

Improving Confidence Estimation on Out-of-Domain Data for End-to-End Speech Recognition

ICASSP 2022accepted

As end-to-end automatic speech recognition (ASR) models reach promising performance, various downstream tasks rely on good confidence estimators for these systems. Recent research has shown that model-based confidence estimators have a significant advantage over using the output softmax probabilitie…

Cited by 16SourceScholar
2021

Confidence Estimation for Attention-Based Sequence-to-Sequence Models for Speech Recognition

ICASSP 2021accepted

For various speech-related tasks, confidence scores from a speech recogniser are a useful measure to assess the quality of transcriptions. In traditional hidden Markov model-based automatic speech recognition (ASR) systems, confidence scores can be reliably obtained from word posteriors in decoding…

Cited by 0SourceScholar
2021

Improving Streaming Automatic Speech Recognition with Non-Streaming Model Distillation on Unsupervised Data

ICASSP 2021accepted

Streaming end-to-end automatic speech recognition (ASR) models are widely used on smart speakers and on-device applications. Since these models are expected to transcribe speech with minimal latency, they are constrained to be causal with no future context, compared to their non-streaming counterpar…

Cited by 0SourceScholar
2021

Learning Word-Level Confidence for Subword End-To-End ASR

ICASSP 2021accepted

We study the problem of word-level confidence estimation in subword-based end-to-end (E2E) models for automatic speech recognition (ASR). Although prior works have proposed training auxiliary confidence models for ASR systems, they do not extend naturally to systems that operate on word-pieces (WP)…

Cited by 0SourceScholar
2020

Label-Efficient Learning on Point Clouds using Approximate Convex Decompositions

ECCV 2020poster

The problems of shape classification and part segmentation from 3D point clouds have garnered increasing attention in the last few years. Both of these problems, however, suffer from relatively small training sets, creating the need for statistically efficient methods to learn 3D shape representatio…

2020

Speech Sentiment Analysis via Pre-Trained Features from End-to-End ASR Models

ICASSP 2020accepted

In this paper, we propose to use pre-trained features from end-to-end ASR models to solve speech sentiment analysis as a down-stream task. We show that end-to-end ASR features, which integrate both acoustic and text information from speech, achieve promising results. We use RNN with self-attention a…

Cited by 0SourceScholar
2019

Automatic Adaptation of Object Detectors to New Domains Using Self-Training

CVPR 2019poster

This work addresses the unsupervised adaptation of an existing object detector to a new target domain. We assume that a large number of unlabeled videos from this domain are readily available. We automatically obtain labels on the target data by using high-confidence detections from the existing det…

Cited by 184PDFScholar
2018

Focal Visual-Text Attention for Visual Question Answering

CVPR 2018poster

Recent insights on language and vision with neural networks have been successfully applied to simple single-image visual question answering. However, to tackle real-life question answering problems on multimedia collections such as personal photos, we have to look at whole collections with sequences…

2018

Lip2Audspec: Speech Reconstruction from Silent Lip Movements Video

ICASSP 2018accepted

In this study, we propose a deep neural network for reconstructing intelligible speech from silent lip movement videos. We use auditory spectrogram as spectral representation of speech and its corresponding sound generation method resulting in a more natural sounding reconstructed speech. Our propos…

Cited by 0SourceScholar
2016

TGIF: A New Dataset and Benchmark on Animated GIF Description

CVPR 2016spotlight

With the recent popularity of animated GIFs on social media, there is need for ways to index them with rich metadata. To advance research on animated GIF understanding, we collected a new dataset, Tumblr GIF (TGIF), with 100K animated GIFs from Tumblr and 120K natural language descriptions obtained…

Cited by 330PDFcodeScholar