← Search

Xi Yin

22 accepted papers

2026

UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning

AAAI 2026technical

The recent DeepSeek-R1 has showcased the emergence of reasoning capabilities in large language models (LLMs) through reinforcement learning (RL) with rule-based rewards. Despite its success in language tasks, its application in multimodal domains, particularly in graphic user interface (GUI) agent t

Cited by 0SourcePDFScholar
2025

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution

CVPR 2025highlight

Diffusion models, and their generalization, flow matching, have had a remarkable impact on the field of media generation. Here, the conventional approach is to learn the complex mapping from a simple source distribution of Gaussian noise to the target media distribution. For cross-modal tasks such a…

Cited by 0SourcePDFScholar
2025

Generating Multi-Image Synthetic Data for Text-to-Image Customization

ICCV 2025poster

Customization of text-to-image models enables users to insert new concepts or objects and generate them in unseen settings. Existing methods either rely on comparatively expensive test-time optimization or train encoders on single-image datasets without multi-image supervision, which can limit image…

2025

MotiF: Making Text Count in Image Animation with Motion Focal Loss

CVPR 2025poster

Text-Image-to-Video (TI2V) generation aims to generate a video from an image following a text description, which is also referred to as text-guided image animation. Most existing methods struggle to generate videos that align well with the text prompts, particularly when motion is specified. To over…

2024

Factorizing Text-to-Video Generation by Explicit Image Conditioning

ECCV 2024poster

"We present , a text-to-video generation model that factorizes the generation into two steps: first generating an image conditioned on the text, and then generating a video conditioned on the text and the generated image. We identify critical design decisions–adjusted noise schedules for diffusion,…

Cited by 84SourcePDFScholar
2023

CCEval: A Representative Evaluation Benchmark for the Chinese-centric Multilingual Machine Translation

EMNLP 2023short findings

The Chinese-centric Multilingual Machine Translation (MMT) has gained more importance recently due to increasing demands from international business development and cross-cultural exchanges. However, an important factor that limits the progress of this area is the lack of highly representative and…

Cited by 0SourceScholar
2023

MaLP: Manipulation Localization Using a Proactive Scheme

CVPR 2023poster

Advancements in the generation quality of various Generative Models (GMs) has made it necessary to not only perform binary manipulation detection but also localize the modified pixels in an image. However, prior works termed as passive for manipulation localization exhibit poor generalization perfor…

2023

Make-A-Video: Text-to-Video Generation without Text-Video Data

ICLR 2023poster

We propose Make-A-Video -- an approach for directly translating the tremendous recent progress in Text-to-Image (T2I) generation to Text-to-Video (T2V). Our intuition is simple: learn what the world looks like and how it is described from paired text-image data, and learn how the world moves from un…

Cited by 1412SourcePDFScholar
2023

SpaText: Spatio-Textual Representation for Controllable Image Generation

CVPR 2023poster

Recent text-to-image diffusion models are able to generate convincing results of unprecedented quality. However, it is nearly impossible to control the shapes of different regions/objects or their layout in a fine-grained fashion. Previous attempts to provide such controls were hindered by their rel…

Cited by 226SourcePDFScholar
2022

Long Video Generation with Time-Agnostic VQGAN and Time-Sensitive Transformer

ECCV 2022poster

"Videos are created to express emotion, exchange information, and share experiences. Video synthesis has intrigued researchers for a long time. Despite the rapid progress driven by advances in visual synthesis, most existing studies focus on improving the frames’ quality and the transitions between…

2022

MUGEN: A Playground for Video-Audio-Text Multimodal Understanding and GENeration

ECCV 2022poster

"Multimodal video-audio-text understanding and generation can benefit from datasets that are narrow but rich. The narrowness allows bite-sized challenges that the research community can make progress on. The richness ensures we are making progress along the core challenges. To this end, we present a…

2021

A Multiplexed Network for End-to-End, Multilingual OCR

CVPR 2021poster

Recent advances in OCR have shown that an end-to-end (E2E) training pipeline that includes both detection and recognition leads to the best results. However, many existing methods focus primarily on Latin-alphabet languages, often even only case-insensitive English characters. In this paper, we prop…

Cited by 61PDFcodeScholar
2021

TAP: Text-Aware Pre-Training for Text-VQA and Text-Caption

CVPR 2021poster

In this paper, we propose Text-Aware Pre-training (TAP) for Text-VQA and Text-Caption tasks. These two tasks aim at reading and understanding scene text in images for question answering and image caption generation, respectively. In contrast to the conventional vision-language pre-training that fail…

Cited by 192PDFcodeScholar
2021

VIVO: Visual Vocabulary Pre-Training for Novel Object Captioning

AAAI 2021technical

It is highly desirable yet challenging to generate image captions that can describe novel objects which are unseen in caption-labeled training data, a capability that is evaluated in the novel object captioning challenge (nocaps). In this challenge, no additional image-caption training data, other t…

Cited by 72SourcePDFScholar
2021

img2pose: Face Alignment and Detection via 6DoF, Face Pose Estimation

CVPR 2021poster

We propose real-time, six degrees of freedom (6DoF), 3D face pose estimation without face detection or landmark localization. We observe that estimating the 6DoF rigid transformation of a face is a simpler problem than facial landmark detection, often used for 3D face alignment. In addition, 6DoF of…

Cited by 171PDFcodeScholar
2020

Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks

ECCV 2020poster

Large-scale pre-training methods of learning cross-modal representations on image-text pairs are becoming popular for vision-language tasks. While existing methods simply concatenate image region features and text features as input to the model to be pre-trained and use self-attention to learn image…

2019

Feature Transfer Learning for Face Recognition With Under-Represented Data

CVPR 2019poster

Despite the large volume of face recognition datasets, there is a significant portion of subjects, of which the samples are insufficient and thus under-represented. Ignoring such significant portion results in insufficient training data. Training with under-represented data leads to biased classifie…

Cited by 396PDFScholar
2019

Gait Recognition via Disentangled Representation Learning

CVPR 2019oral

Gait, the walking pattern of individuals, is one of the most important biometrics modalities. Most of the existing gait recognition methods take silhouettes or articulated body models as the gait features. These methods suffer from degraded recognition performance when handling confounding variables…

Cited by 324PDFScholar