← Search

Yiran Xu

12 accepted papers

2026

DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance

CVPR 2026

Recent vision-language model (VLM)-based approaches have achieved impressive results on SVG generation. However, because they generate only text and lack visual signals during decoding, they often struggle with complex semantics and fail to produce visually appealing or geometrically coherent SVGs.

Cited by 0SourcecodeScholar
2026

Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual Generation

CVPR 2026

Achieving precise alignment between user intent and generated visuals remains a central challenge in text-to-visual generation, as a single attempt often fails to produce the desired output. To handle this, prior approaches mainly scale the visual generation process (e.g., increasing sampling steps

Cited by 0SourceScholar
2026

Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO

ICML 2026poster

We identify a new dimension for enhancing rollout diversity in Group Relative Policy Optimization (GRPO) for LLMs. While GRPO relies on diverse rollouts, prevailing strategies primarily increase diversity by injecting more token-level randomness, which may introduce step-wise noise and leads to inco…

Cited by 0SourceScholar
2025

Beyond Generation: A Diffusion-based Low-level Feature Extractor for Detecting AI-generated Images

CVPR 2025poster

The prevalence of AI-generated images has evoked concerns regarding the potential misuse of image generation technologies. In response, numerous detection methods aim to identify AI-generated images by analyzing generative artifacts. Unfortunately, most detectors quickly become obsolete with the dev…

Cited by 0SourcePDFScholar
2025

Cropper: Vision-Language Model for Image Cropping through In-Context Learning

CVPR 2025poster

The goal of image cropping is to identify visually appealing crops in an image. Conventional methods are trained on specific datasets and fail to adapt to new requirements. Recent breakthroughs in large vision-language models (VLMs) enable visual in-context learning without explicit training. Howeve…

Cited by 2SourcePDFScholar
2025

Fine-grained Prompt Screening: Defending Against Backdoor Attack on Text-to-Image Diffusion Models

IJCAI 2025

Text-to-image (T2I) diffusion models exhibit impressive generation capabilities in recently studies. However, they are vulnerable to backdoor attacks, where model outputs are manipulated by malicious triggers. In this paper, we propose a novel input-level defense method, called Fine-grained Prompt S

Cited by 0SourcePDFScholar
2025

VideoGigaGAN: Towards Detail-rich Video Super-Resolution

CVPR 2025poster

Video super-resolution (VSR) models achieve temporal consistency but often produce blurrier results than their image-based counterparts due to limited generative capacity. This prompts the question: can we adapt a generative image upsampler for VSR while preserving temporal consistency? We introduce…

Cited by 17SourcePDFScholar
2024

Flash-Splat: 3D Reflection Removal with Flash Cues and Gaussian Splats

ECCV 2024poster

"We introduce a simple yet effective approach for separating transmitted and reflected light. Our key insight is that the powerful novel view synthesis capabilities provided by modern inverse rendering methods (e.g., 3D Gaussian splatting) allow one to perform flash/no-flash reflection separation us…

Cited by 8SourcePDFScholar
2024

In-N-Out: Faithful 3D GAN Inversion with Volumetric Decomposition for Face Editing

CVPR 2024poster

3D-aware GANs offer new capabilities for view synthesis while preserving the editing functionalities of their 2D counterparts. GAN inversion is a crucial step that seeks the latent code to reconstruct input images or videos subsequently enabling diverse editing tasks through manipulation of this lat…

Cited by 3SourcePDFScholar
2020

Explainable Object-Induced Action Decision for Autonomous Vehicles

CVPR 2020poster

A new paradigm is proposed for autonomous driving. The new paradigm lies between the end-to-end and pipelined approaches, and is inspired by how humans solve the problem. While it relies on scene understanding, the latter only considers objects that could originate hazard. These are denoted as actio…

Cited by 147PDFcodeScholar