← Search

YiLin Wang

51 accepted papers

2026

3DrawAgent: Teaching LLM to Draw in 3D with Early Contrastive Experience

CVPR 2026

Sketching in 3D space enables expressive reasoning about shape, structure, and spatial relationships, yet generating 3D sketches through natural language remains a major challenge. In this work, we introduce 3DrawAgent, a training-free, language-driven framework for 3D sketch generation that leverag

Cited by 0SourceScholar
2026

Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing

AAAI 2026technical

Recent diffusion-based image editing methods have made great strides in text-guided tasks but often struggle with complex, indirect instructions. Additionally, current models frequently exhibit poor identity preservation, unintended edits, or rely on manual masks. To overcome these limitations, we i

Cited by 0SourcePDFScholar
2026

MDS-VQA: Model-Informed Data Selection for Video Quality Assessment

CVPR 2026

Learning-based video quality assessment (VQA) has advanced rapidly, yet progress is increasingly constrained by a disconnect between model design and dataset curation. Model-centric approaches often iterate on fixed benchmarks, while data-centric efforts collect new human labels without systematical

Cited by 0SourcecodeScholar
2026

OR-PRM: A Process Reward Model for Algorithmic Problem in Operations Research

ICLR 2026poster

Large language models (LLMs) with Process Reward Models (PRMs) have shown strong reasoning ability, yet their potential in Operations Research (OR) remains unexplored. We present the first PRM tailored for OR, but find that directly training on mainstream datasets yields surprisingly weak performanc…

Cited by 0SourceScholar
2026

SCALING AUDIO-VISUAL QUALITY ASSESSMENT DATASET VIA CROWDSOURCING

ICASSP 2026oral

Audio-visual quality assessment (AVQA) research has been stalled by limitations of existing datasets: they are typically small in scale, with insufficient diversity in content and quality, and annotated only with overall scores. These shortcomings provide limited support for model development and mu…

Cited by 0SourcePDFScholar
2026

Seeing Beyond 8bits: Subjective and Objective Quality Assessment of HDR-UGC Videos

CVPR 2026

High Dynamic Range (HDR) user-generated (UGC) videos are rapidly proliferating across social platforms, yet most perceptual video quality assessment (VQA) systems remain tailored to Standard Dynamic Range (SDR). HDR's higher bit depth, wide color gamut, and elevated luminance range expose distortion

Cited by 0SourcecodeScholar
2026

Stepwise Credit Assignment for GRPO on Flow-Matching Models

CVPR 2026

Flow-GRPO successfully applies reinforcement learning to flow models, but uses uniform credit assignment across all steps. This ignores the temporal structure of diffusion generation: early steps determine composition and content (low-frequency structure), while late steps resolve details and textur

Cited by 0SourceScholar
2026

TextGround4M: A Prompt-Aligned Dataset for Layout-Aware Text Rendering

AAAI 2026technical

Despite recent advances in text-to-image (T2I) generation, models still struggle to accurately render prompt-specified text with correct spatial layout—especially in multi-span, structured settings. This challenge is driven not only by the lack of datasets that align prompts with the exact text and

Cited by 0SourcePDFScholar
2025

An Ensemble Approach to Short-form Video Quality Assessment Using Multimodal LLM

ICASSP 2025accepted

The rise of short-form videos, characterized by diverse content, editing styles, and artifacts, poses substantial challenges for learning-based blind video quality assessment (BVQA) models. Multimodal large language models (MLLMs), renowned for their superior generalization capabilities, present a p…

Cited by 0SourceScholar
2025

Continuously Steering LLMs Sensitivity to Contextual Knowledge with Proxy Models

EMNLP 2025

In Large Language Models (LLMs) generation, there exist knowledge conflicts, and scenarios where parametric knowledge contradicts knowledge provided in the context. Previous works studied tuning, decoding algorithms, or locating and editing context-aware neurons to adapt LLMs to be faithful to new c

2025

EventRAG: Enhancing LLM Generation with Event Knowledge Graphs

ACL 2025long

Retrieval-augmented generation (RAG) systems often struggle with narrative-rich documents and event-centric reasoning, particularly when synthesizing information across multiple sources. We present EventRAG, a novel framework that enhances text generation through structured event representations. We…

Cited by 0SourcePDFScholar
2025

FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity

CVPR 2025poster

The advent of large Vision-Language Models (VLMs) has significantly advanced multimodal tasks, enabling more sophisticated and accurate integration of visual and textual information across various applications, including image and video captioning, visual question answering, and cross-modal retrieva…

Cited by 6SourcePDFScholar
2025

ITFormer: Bridging Time Series and Natural Language for Multi-Modal QA with Large-Scale Multitask Dataset

ICML 2025poster

Time-series data are critical in diverse applications, such as industrial monitoring, medical diagnostics, and climate research. However, effectively integrating these high-dimensional temporal signals with natural language for dynamic, interactive tasks remains a significant challenge. To address t…

2025

Language-Specific Layer Matters: Efficient Multilingual Enhancement for Large Vision-Language Models

EMNLP 2025

Large vision-language models (LVLMs) have demonstrated exceptional capabilities in understanding visual information with human languages but also exhibit an imbalance in multilingual capabilities. In this work, we delve into the multilingual working pattern of LVLMs and identify a salient correlatio

2025

MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked Transformer

ICLR 2025poster

Generative masked transformer have demonstrated remarkable success across various content generation tasks, primarily due to their ability to effectively model large-scale dataset distributions with high consistency. However, in the animation domain, large datasets are not always available. Applying…

Cited by 0SourcePDFScholar
2025

OmniStyle: Filtering High Quality Style Transfer Data at Scale

CVPR 2025poster

In this paper, we introduce OmniStyle-1M, a large-scale paired style transfer dataset comprising over one million content-style-stylized image triplets across 1,000 diverse style categories, each enhanced with textual descriptions and instruction prompts. We show that OmniStyle-1M can not only impro…

Cited by 0SourcePDFScholar
2025

SLAM: Towards Efficient Multilingual Reasoning via Selective Language Alignment

COLING 2025main

Despite the significant improvements achieved by large language models (LLMs) in English reasoning tasks, these models continue to struggle with multilingual reasoning. Recent studies leverage a full-parameter and two-stage training paradigm to teach models to first understand non-English questions…

2025

SigStyle: Signature Style Transfer via Personalized Text-to-Image Models

AAAI 2025technical

Style transfer enables the seamless integration of artistic styles from a style image into a content image, resulting in visually striking and aesthetically enriched outputs. Despite numerous advances in this field, existing methods did not explicitly focus on the signature style, which represents…

Cited by 1SourcePDFScholar
2025

UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics

CVPR 2025highlight

We introduce UniReal, a unified framework designed to address various image generation and editing tasks. Existing solutions often vary by tasks, yet share fundamental principles: preserving consistency between inputs and outputs while capturing visual variations. Inspired by recent video generation…

2024

Amodal Scene Analysis via Holistic Occlusion Relation Inference and Generative Mask Completion

AAAI 2024technical

Amodal scene analysis entails interpreting the occlusion relationship among scene elements and inferring the possible shapes of the invisible parts. Existing methods typically frame this task as an extended instance segmentation or a pair-wise object de-occlusion problem. In this work, we propose a…

2024

Bayesian Diffusion Models for 3D Shape Reconstruction

CVPR 2024poster

We present Bayesian Diffusion Models (BDM) a prediction algorithm that performs effective Bayesian inference by tightly coupling the top-down (prior) information with the bottom-up (data-driven) procedure via joint diffusion processes. We demonstrate the application of BDM on the 3D shape reconstruc…

2024

Dolfin: Diffusion Layout Transformers without Autoencoder

ECCV 2024poster

"In this paper, we introduce a new generative model, Diffusion Layout Transformers without Autoencoder (Dolfin), that attains significantly improved modeling capability and transparency over the existing approaches. Dolfin employs a Transformer-based diffusion process to model layout generation. In…

Cited by 17SourcePDFScholar
2024

Enhanced KPI Anomaly Detection: An Unsupervised Hybrid Model with Dynamic Threshold

ICASSP 2024accepted

Anomaly detection based on key performance indicator (KPI) is an important topic in the field of intelligent operation and maintenance. The problem of insufficient annotated samples is widespread in the industrial Internet, and it severely impairs the performance of data-driven anomaly detection. Pr…

Cited by 0SourceScholar
2024

GSD: View-Guided Gaussian Splatting Diffusion for 3D Reconstruction

ECCV 2024poster

"We present GSD, a diffusion model approach based on Gaussian Splatting (GS) representation for 3D object reconstruction from a single view. Prior works suffer from inconsistent 3D geometry or mediocre rendering quality due to improper representations. We take a step towards resolving these shortcom…

Cited by 7SourcePDFScholar
2024

KC-GenRe: A Knowledge-constrained Generative Re-ranking Method Based on Large Language Models for Knowledge Graph Completion

COLING 2024main

The goal of knowledge graph completion (KGC) is to predict missing facts among entities. Previous methods for KGC re-ranking are mostly built on non-generative language models to obtain the probability of each candidate. Recently, generative large language models (LLMs) have shown outstanding perfor…

2024

Learning Mutually Informed Representations for Characters and Subwords

NAACL 2024findings

Most pretrained language models rely on subword tokenization, which processes text as a sequence of subword tokens. However, different granularities of text, such as characters, subwords, and words, can contain different kinds of information. Previous studies have shown that incorporating multiple i…

2024

TokenCompose: Text-to-Image Diffusion with Token-level Supervision

CVPR 2024poster

We present TokenCompose a Latent Diffusion Model for text-to-image generation that achieves enhanced consistency between user-specified text prompts and model-generated images. Despite its tremendous success the standard denoising process in the Latent Diffusion Model takes text prompts as condition…

2024

UniHuman: A Unified Model For Editing Human Images in the Wild

CVPR 2024poster

Human image editing includes tasks like changing a person's pose their clothing or editing the image according to a text prompt. However prior work often tackles these tasks separately overlooking the benefit of mutual reinforcement from learning them jointly. In this paper we propose UniHuman a uni…

2023

A Canonicalization-Enhanced Known Fact-Aware Framework For Open Knowledge Graph Link Prediction

IJCAI 2023poster

Open knowledge graph (OpenKG) link prediction aims to predict missing factual triples in the form of (head noun phrase, relation phrase, tail noun phrase). Since triples are not canonicalized, previous methods either focus on canonicalizing noun phrases (NPs) to reduce graph sparsity, or utilize tex…

2023

Interactive Portrait Harmonization

ICLR 2023poster

Current image harmonization methods consider the entire background as the guidance for harmonization. However, this may limit the capability for user to choose any specific object/person in the background to guide the harmonization. To enable flexible interaction between user and harmonization, we i…

Cited by 17SourcePDFScholar
2023

LightPainter: Interactive Portrait Relighting With Freehand Scribble

CVPR 2023poster

Recent portrait relighting methods have achieved realistic results of portrait lighting effects given a desired lighting representation such as an environment map. However, these methods are not intuitive for user interaction and lack precise lighting control. We introduce LightPainter, a scribble-b…

Cited by 14SourcePDFScholar
2023

PHOTOSWAP: Personalized Subject Swapping in Images

NeurIPS 2023poster

In an era where images and visual content dominate our digital landscape, the ability to manipulate and personalize these images has become a necessity. Envision seamlessly substituting a tabby cat lounging on a sunlit window sill in a photograph with your own playful puppy, all while preserving the…

Cited by 36SourcePDFScholar
2022

Lite Vision Transformer With Enhanced Self-Attention

CVPR 2022poster

Despite the impressive representation capacity of vision transformer models, current light-weight vision transformer models still suffer from inconsistent and incorrect dense predictions at local regions. We suspect that the power of their self-attention mechanism is limited in shallower and thinner…

Cited by 151PDFcodeScholar
2021

Mask Guided Matting via Progressive Refinement Network

CVPR 2021poster

We propose Mask Guided (MG) Matting, a robust matting framework that takes a general coarse mask as guidance. MG Matting leverages a network (PRN) design which encourages the matting model to provide self-guidance to progressively refine the uncertain regions through the decoding process. A series o…

Cited by 153PDFcodeScholar
2021

Multimodal Contrastive Training for Visual Representation Learning

CVPR 2021poster

We develop an approach to learning visual representations that embraces multimodal data, driven by a combination of intra- and inter-modal similarity preservation objectives. Unlike existing visual pre-training methods, which solve a proxy prediction task in a single domain, our method exploits intr…

Cited by 215PDFcodeScholar
2021

Regression or classification? New methods to evaluate no-reference picture and video quality models

ICASSP 2021accepted

Video and image quality assessment has long been projected as a regression problem, which requires predicting a continuous quality score given an input stimulus. However, recent efforts have shown that accurate quality score regression on real-world user-generated content (UGC) is a very challenging…

Cited by 0SourceScholar
2021

Rich Features for Perceptual Quality Assessment of UGC Videos

CVPR 2021poster

Video quality assessment for User Generated Content (UGC) is an important topic in both industry and academia. Most existing methods only focus on one aspect of the perceptual quality assessment, such as technical quality or compression artifacts. In this paper, we create a large scale dataset to co…

Cited by 105PDFScholar
2021

SSH: A Self-Supervised Framework for Image Harmonization

ICCV 2021poster

Image harmonization aims to improve the quality of image compositing by matching the "appearance"" (e.g., color tone, brightness and contrast) between foreground and background images. However, collecting large-scale annotated datasets for this task requires complex professional retouching. Instead,…

Cited by 96PDFcodeScholar
2020

A Feasible Level Proximal Point Method for Nonconvex Sparse Constrained Optimization

NeurIPS 2020poster

Nonconvex sparse models have received significant attention in high-dimensional machine learning. In this paper, we study a new model consisting of a general convex or nonconvex objectives and a variety of continuous nonconvex sparsity-inducing constraints. For this constrained model, we propose a n…

Cited by 11SourcePDFScholar
2020

BBAND INDEX: A NO-REFERENCE BANDING ARTIFACT PREDICTOR

ICASSP 2020accepted

Banding artifact, or false contouring, is a common video compression impairment that tends to appear on large flat regions in encoded videos. These staircase-shaped color bands can be very noticeable in high-definition videos. Here we study this artifact, and propose a new distortion-specific no-ref…

Cited by 0SourceScholar
2020

Incorporating Reinforced Adversarial Learning in Autoregressive Image Generation

ECCV 2020poster

Autoregressive models recently achieved comparable results versus state-of-the-art Generative Adversarial Networks (GANs) with the help of Vector Quantized Variational AutoEncoders (VQ-VAE). However, autoregressive models have several limitations such as exposure bias and their training objective do…

Cited by 17SourcePDFScholar
2020

Shape Adaptor: A Learnable Resizing Module

ECCV 2020poster

We present a novel resizing module for neural networks: shape adaptor, a drop-in enhancement built on top of traditional resizing layers, such as pooling, bilinear sampling, and strided convolution. Whilst traditional resizing layers have fixed and deterministic reshaping factors, our module allows…

2018

Generalizing Graph Matching beyond Quadratic Assignment Model

NeurIPS 2018poster

Graph matching has received persistent attention over decades, which can be formulated as a quadratic assignment problem (QAP). We show that a large family of functions, which we define as Separable Functions, can approximate discrete graph matching in the continuous domain asymptotically by varying…

Cited by 54SourcePDFScholar
2016

PPP: Joint Pointwise and Pairwise Image Label Prediction

CVPR 2016accepted

Pointwise label and Pairwise label are both widely used in computer vision tasks. For example, supervised image classification and annotation approaches use pointwise label, while attribute-based image relative learning often adopts pairwise labels. These two types of labels are often considered ind…

Cited by 38SourcePDFScholar