← Search

Long Mai

22 accepted papers

2025

GaussianVideo: Efficient Video Representation via Hierarchical Gaussian Splatting

ICCV 2025poster

Efficient neural representations for dynamic video scenes are critical for applications ranging from video compression to interactive simulations. Yet, existing methods often face challenges related to high memory usage, lengthy training times, and temporal consistency. To address these issues, we i…

Cited by 0SourcePDFScholar
2025

Progressive Growing of Video Tokenizers for Temporally Compact Latent Spaces

ICCV 2025poster

Video tokenizers are essential for latent video diffusion models, converting raw video data into spatiotemporally compressed latent spaces for efficient training. However, extending state-of-the-art video tokenizers to achieve a temporal compression ratio beyond 4x without increasing channel capacit…

2025

REGEN: Learning Compact Video Embedding with (Re-)Generative Decoder

ICCV 2025poster

We present a novel perspective on learning video embedders for generative modeling: rather than requiring an exact reproduction of an input video, an effective embedder should focus on synthesizing visually plausible reconstructions. This relaxed criterion enables substantial improvements in compres…

Cited by 0SourcePDFScholar
2025

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models

ICCV 2025poster

Multi-head self-attention (MHSA) is a key component of Transformers, a widely popular architecture in both language and vision. Multiple heads intuitively enable different parallel processes over the same input. Yet, they also obscure the attribution of each input patch to the output of a model. We…

2024

Enhancing Conversation Smoothness in Language Learning Chatbots: An Evaluation of GPT4 for ASR Error Correction

ICASSP 2024accepted

The integration of natural language processing (NLP) technologies into educational applications has shown promising results, particularly in the language learning domain. Many spoken open-domain chatbots have been used as speaking partners, helping language learners improve their language skills. Ho…

Cited by 0SourceScholar
2024

MVDream: Multi-view Diffusion for 3D Generation

ICLR 2024poster

We introduce MVDream, a diffusion model that is able to generate consistent multi-view images from a given text prompt. Learning from both 2D and 3D data, a multi-view diffusion model can achieve the generalizability of 2D diffusion models and the consistency of 3D renderings. We demonstrate that su…

Cited by 630SourcePDFScholar
2021

Boosting Monocular Depth Estimation Models to High-Resolution via Content-Adaptive Multi-Resolution Merging

CVPR 2021poster

Neural networks have shown great abilities in estimating depth from a single image. However, the inferred depth maps are well below one-megapixel resolution and often lack fine-grained details, which limits their practicality. Our method builds on our analysis on how the input resolution and the sce…

Cited by 230PDFcodeScholar
2021

Learning To Recover 3D Scene Shape From a Single Image

CVPR 2021poster

Despite significant progress in monocular depth estimation in the wild, recent state-of-the-art methods cannot be used to recover accurate 3D scene shape due to an unknown depth shift induced by shift-invariant reconstruction losses used in mixed-data depth prediction training, and possible unknown…

Cited by 284PDFcodeScholar
2020

Active Speakers in Context

CVPR 2020poster

Current methods for active speaker detection focus on modeling audiovisual information from a single speaker. This strategy can be adequate for addressing single-speaker scenarios, but it prevents accurate detection when the task is to identify who of many candidate speakers are talking. This paper…

Cited by 104PDFcodeScholar
2020

BlockGAN: Learning 3D Object-aware Scene Representations from Unlabelled Images

NeurIPS 2020poster

We present BlockGAN, an image generative model that learns object-aware 3D scene representations directly from unlabelled 2D images. Current work on scene representation learning either ignores scene background or treats the whole scene as one object. Meanwhile, work that considers scene composition…

2020

Context-Aware Group Captioning via Self-Attention and Contrastive Features

CVPR 2020poster

While image captioning has progressed rapidly, existing works focus mainly on describing single images. In this paper, we introduce a new task, context-aware group captioning, which aims to describe a group of target images in the context of another group of related reference images. Context-aware g…

Cited by 49PDFScholar
2020

Structure-Guided Ranking Loss for Single Image Depth Prediction

CVPR 2020poster

Single image depth prediction is a challenging task due to its ill-posed nature and challenges with capturing ground truth for supervision. Large-scale disparity data generated from stereo photos and 3D videos is a promising source of supervision, however, such disparity data can only approximate th…

Cited by 214PDFcodeScholar
2019

An Internal Learning Approach to Video Inpainting

ICCV 2019poster

We propose a novel video inpainting algorithm that simultaneously hallucinates missing appearance and motion (optical flow) information, building upon the recent 'Deep Image Prior' (DIP) that exploits convolutional network architectures to enforce plausible texture in static images. In extending DIP…

Cited by 101PDFcodeScholar
2019

MultiSeg: Semantically Meaningful, Scale-Diverse Segmentations From Minimal User Input

ICCV 2019poster

Existing deep learning-based interactive image segmentation approaches typically assume the target-of-interest is always a single object and fail to account for the potential diversity in user expectations, thus requiring excessive user input when it comes to segmenting an object part or a group of…

Cited by 45PDFcodeScholar
2019

Strike (With) a Pose: Neural Networks Are Easily Fooled by Strange Poses of Familiar Objects

CVPR 2019poster

Despite excellent performance on stationary test sets, deep neural networks (DNNs) can fail to generalize to out-of-distribution (OoD) inputs, including natural, non-adversarial ones, which are common in real-world settings. In this paper, we present a framework for discovering DNN failures that har…

Cited by 392PDFcodeScholar
2018

Interactive Boundary Prediction for Object Selection

ECCV 2018poster

Interactive image segmentation is critical for many image editing tasks. While recent advanced methods on interactive segmentation focus on the region-based paradigm, more traditional boundary-based methods such as Intelligent Scissor are still popular in practice as they allow users to have active…

Cited by 65SourcePDFScholar
2017

Spatial-Semantic Image Search by Visual Feature Synthesis

CVPR 2017spotlight

The performance of image retrieval has been improved tremendously in recent years through the use of deep feature representations. Most existing methods, however, aim to retrieve images that are visually similar or semantically relevant to the query, irrespective of spatial configuration. In this pa…

Cited by 52PDFcodeScholar