← Search

Xin Cai

10 accepted papers

2026

Controllable First-Frame-Guided Video Editing via Mask-Aware LoRA Fine-Tuning

ICLR 2026poster

Video editing using diffusion models has achieved remarkable results in generating high-quality edits for videos. However, current methods often rely on large-scale pretraining, limiting flexibility for specific edits. First-frame-guided editing provides control over the first frame, but lacks fine-…

Cited by 0SourcecodeScholar
2026

FlashVSR: Towards Real-time Diffusion-Based Streaming Video Super Resolution

CVPR 2026

Diffusion models have recently advanced video restoration, but applying them to real-world and AIGC-generated video super-resolution (VSR) remains challenging due to high latency, prohibitive computation, and poor generalization to ultra-high resolutions. Our goal in this work is to make diffusion-b

Cited by 0SourcecodeScholar
2026

MindDriver: Introducing Progressive Multimodal Reasoning for Autonomous Driving

CVPR 2026

Vision-Language Models (VLM) exhibit strong reasoning capabilities, showing promise for end-to-end autonomous driving systems. Chain-of-Thought (CoT), as VLM's widely used reasoning strategy, is facing critical challenges. Existing textual CoT has a large gap between text semantic space and trajecto

Cited by 0SourcecodeScholar
2026

Twins: Learn to Predict Unified Representations with Focal Loss

ICML 2026poster

Unified multimodal models seek a shared visual token space that supports both multimodal understanding and image generation. Discrete methods unify the interface via a shared codebook, whereas continuous pipelines often rely on two disparate representations—semantic features (e.g., ViT) for understa…

Cited by 0SourceScholar
2025

Teaching Large Language Models to Regress Accurate Image Quality Scores Using Score Distribution

CVPR 2025poster

With the rapid advancement of Multi-modal Large Language Models (MLLMs), MLLM-based Image Quality Assessment (IQA) methods have shown promising performance in linguistic quality description. However, current methods still fall short in accurately scoring image quality. In this work, we aim to levera…

2025

UltraFusion: Ultra High Dynamic Imaging using Exposure Fusion

CVPR 2025highlight

Capturing high dynamic range (HDR) scenes is one of the most important issues in camera design. Majority of cameras use exposure fusion, which fuses images captured by different exposure levels, to increase dynamic range. However, this approach can only handle images with limited exposure difference…

2024

PhoCoLens: Photorealistic and Consistent Reconstruction in Lensless Imaging

NeurIPS 2024spotlight

Lensless cameras offer significant advantages in size, weight, and cost compared to traditional lens-based systems. Without a focusing lens, lensless cameras rely on computational algorithms to recover the scenes from multiplexed measurements. However, current algorithms struggle with inaccurate for…

Cited by 7SourcePDFScholar
2023

Source-Free Adaptive Gaze Estimation by Uncertainty Reduction

CVPR 2023poster

Gaze estimation across domains has been explored recently because the training data are usually collected under controlled conditions while the trained gaze estimators are used in real and diverse environments. However, due to privacy and efficiency concerns, simultaneous access to annotated source…