← Search

Zechen Bai

18 accepted papers

2026

Demo2Tutorial: From Human Experience to Multimodal Software Tutorials

CVPR 2026

Human experience in digital environments offers a vast, underexplored resource of authentic, untrimmed interactions that contain rich procedural knowledge. We introduce Demo2Tutorial, a framework that transforms this experience captured via screen recordings and interaction logs into structured, mul

Cited by 0SourcecodeScholar
2026

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation

ICML 2026poster

While Vision-Language-Action (VLA) models offer broad general capabilities, deploying them on specific hardware requires real-world adaptation to bridge the embodiment gap. Since robot demonstrations are costly, this adaptation must often occur under a strict data budget. In this work, we identify a…

Cited by 0SourceScholar
2025

Bridging Information Asymmetry in Text-video Retrieval: A Data-centric Approach

ICLR 2025poster

As online video content rapidly grows, the task of text-video retrieval (TVR) becomes increasingly important. A key challenge in TVR is the information asymmetry between video and text: videos are inherently richer in information, while their textual descriptions often capture only fragments of this…

Cited by 0SourcePDFScholar
2025

Impossible Videos

ICML 2025poster

Synthetic videos nowadays is widely used to complement data scarcity and diversity of real-world videos. Current synthetic datasets primarily replicate real-world scenarios, leaving impossible, counterfactual and anti-reality video concepts underexplored. This work aims to answer two questions: 1) C…

Cited by 1SourcePDFScholar
2025

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

ICLR 2025poster

We present a unified transformer, i.e., Show-o, that unifies multimodal understanding and generation. Unlike fully autoregressive models, Show-o unifies autoregressive and (discrete) diffusion modeling to adaptively handle inputs and outputs of various and mixed modalities. The unified model flexibl…

Cited by 164SourcePDFScholar
2025

ShowUI: One Vision-Language-Action Model for GUI Visual Agent

CVPR 2025poster

Building Graphical User Interface (GUI) assistants holds significant promise for enhancing human workflow productivity. While most agents are language-based, relying on closed-source API with text-rich meta-information (e.g., HTML or accessibility tree), they show limitations in perceiving UI visual…

2025

VaporTok: RL-Driven Adaptive Video Tokenizer with Prior & Task Awareness

NeurIPS 2025poster

Recent advances in visual tokenizers have demonstrated their effectiveness for multimodal large language models and autoregressive generative models. However, most existing visual tokenizers rely on a fixed downsampling rate at a given visual resolution, and consequently produce a constant number of…

Cited by 0SourceScholar
2025

You Only Communicate Once: One-shot Federated Low-Rank Adaptation of MLLM

NeurIPS 2025poster

Multimodal Large Language Models (MLLMs) with Federated Learning (FL) can quickly adapt to privacy-sensitive tasks, but face significant challenges such as high communication costs and increased attack risks, due to their reliance on multi-round communication. To address this, One-shot FL (OFL) has…

Cited by 0SourcecodeScholar
2024

Adaptive Slot Attention: Object Discovery with Dynamic Slot Number

CVPR 2024poster

Object-centric learning (OCL) extracts the representation of objects with slots offering an exceptional blend of flexibility and interpretability for abstracting low-level perceptual features. A widely adopted method within OCL is slot attention which utilizes attention mechanisms to iteratively ref…

2024

AssistGUI: Task-Oriented PC Graphical User Interface Automation

CVPR 2024poster

Graphical User Interface (GUI) automation holds significant promise for assisting users with complex tasks thereby boosting human productivity. Existing works leveraging Large Language Model (LLM) or LLM-based AI agents have shown capabilities in automating tasks on Android and Web platforms. Howeve…

Cited by 7SourcePDFScholar
2024

DoFIT: Domain-aware Federated Instruction Tuning with Alleviated Catastrophic Forgetting

NeurIPS 2024poster

Federated Instruction Tuning (FIT) advances collaborative training on decentralized data, crucially enhancing model's capability and safeguarding data privacy. However, existing FIT methods are dedicated to handling data heterogeneity across different clients (i.e., client-aware data heterogeneity),…

2024

LOVA3: Learning to Visual Question Answering, Asking and Assessment

NeurIPS 2024poster

Question answering, asking, and assessment are three innate human traits crucial for understanding the world and acquiring knowledge. By enhancing these capabilities, humans can more effectively utilize data, leading to better comprehension and learning outcomes. However, current Multimodal Large La…

2024

One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos

NeurIPS 2024poster

We introduce VideoLISA, a video-based multimodal large language model designed to tackle the problem of language-instructed reasoning segmentation in videos. Leveraging the reasoning capabilities and world knowledge of large language models, and augmented by the Segment Anything Model, VideoLISA gen…

2023

Unsupervised Open-Vocabulary Object Localization in Videos

ICCV 2023poster

In this paper, we show that recent advances in video representation learning and pre-trained vision-language models allow for substantial improvements in self-supervised video object localization. We propose a method that first localizes objects in videos via a slot attention approach and then assig…

Cited by 7PDFcodeScholar
2021

Explain Me the Painting: Multi-Topic Knowledgeable Art Description Generation

ICCV 2021poster

Have you ever looked at a painting and wondered what is the story behind it? This work presents a framework to bring art closer to people by generating comprehensive descriptions of fine-art paintings. Generating informative descriptions for artworks, however, is extremely challenging, as it require…

Cited by 54PDFcodeScholar
2021

Unsupervised Multi-Source Domain Adaptation for Person Re-Identification

CVPR 2021poster

Unsupervised domain adaptation (UDA) methods for person re-identification (re-ID) aim at transferring re-ID knowledge from labeled source data to unlabeled target data. Among these methods, the pseudo-label-based branch has achieved great success, whereas most of them only use limited data from a si…

Cited by 112PDFScholar