← Search

Yiqi Lin

8 accepted papers

2026

Demo2Tutorial: From Human Experience to Multimodal Software Tutorials

CVPR 2026

Human experience in digital environments offers a vast, underexplored resource of authentic, untrimmed interactions that contain rich procedural knowledge. We introduce Demo2Tutorial, a framework that transforms this experience captured via screen recordings and interaction logs into structured, mul

Cited by 0SourcecodeScholar
2026

GAHMN: A Generative Approach for High-Dimensional Mediation Analysis

AAAI 2026technical

High-dimensional mediation analysis (HMA) seeks to uncover complex causal mechanisms involving numerous mediators and plays a crucial role in scientific and social sciences. In this work, we introduce the Generative Adversarial High-dimensional Mediation Network (GAHMN), a novel, scalable structured

Cited by 0SourcePDFScholar
2025

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

CVPR 2025poster

Recent video large language models (Video LLMs) often depend on costly human annotations or proprietary APIs (e.g., GPT-4o) to produce training data, which limits their training at scale. In this paper, we explore large-scale training for Video LLM with cheap automatic speech recognition (ASR) trans…

2024

Leveraging Visual Tokens for Extended Text Contexts in Multi-Modal Learning

NeurIPS 2024poster

Training models with longer in-context lengths is a significant challenge for multimodal machine learning due to substantial GPU memory and computational costs. This exploratory study does not present state-of-the-art models; rather, it introduces an innovative method designed to increase in-context…

2024

Parrot Captions Teach CLIP to Spot Text

ECCV 2024oral

"Despite CLIP [?] being the foundation model in numerous vision-language applications, CLIP suffers from a severe text spotting bias. Such bias causes CLIP models to ‘Parrot’ the visual text embedded within images while disregarding the authentic visual semantics. We uncover that in the most popular…

2023

SEPT: Towards Scalable and Efficient Visual Pre-training

AAAI 2023technical

Recently, the self-supervised pre-training paradigm has shown great potential in leveraging large-scale unlabeled data to improve downstream task performance. However, increasing the scale of unlabeled pre-training data in real-world scenarios requires prohibitive computational costs and faces the c…

Cited by 1SourcePDFScholar
2021

Removing the Background by Adding the Background: Towards Background Robust Self-Supervised Video Representation Learning

CVPR 2021poster

Self-supervised learning has shown great potentials in improving the video representation ability of deep neural networks by getting supervision from the data itself. However, some of the current methods tend to cheat from the background, i.e., the prediction is highly dependent on the video backgro…

Cited by 112PDFcodeScholar