← Search

Yancheng Bai

8 accepted papers

2026

From Scale to Speed: Adaptive Test-Time Scaling for Image Editing

CVPR 2026

Image Chain-of-Thought (Image-CoT) is a test-time scaling paradigm that improves image generation by extending inference time. Most Image-CoT methods focus on text-to-image (T2I) generation. Unlike T2I generation, image editing is goal-directed: the solution space is constrained by the source image

Cited by 0SourceScholar
2026

Layer-wise Instance Binding for Regional and Occlusion Control in Text-to-Image Diffusion Transformers

CVPR 2026

Region-instructed layout control in text-to-image generation is highly practical, yet existing methods suffer from limitations: (i) training-based approaches inherit data bias and often degrade image quality, and (ii) current techniques struggle with occlusion order, limiting real-world usability. T

Cited by 0SourcecodeScholar
2026

SCALAR: Scale-wise Controllable Visual Autoregressive Learning

AAAI 2026technical

Controllable image synthesis, which enables fine-grained control over generated outputs, has emerged as a key focus in visual generative modeling. However, controllable generation remains challenging for Visual Autoregressive (VAR) models due to their hierarchical, next-scale prediction style. Exist

Cited by 0SourcePDFScholar
2026

Semantic Context Matters: Improving Conditioning for Autoregressive Models

CVPR 2026

Recently, autoregressive (AR) models have shown strong potential in image generation, offering better scalability and easier integration with unified multi-modal models compared to diffusion methods.However, extending AR models to controllable image editing remains challenging due to weak and ineffi

Cited by 0SourcecodeScholar
2024

A Study on the Adverse Impact of Synthetic Speech on Speech Recognition

ICASSP 2024accepted

The high-quality synthetic speech by TTS has been widely used in the field of human-computer interaction, bringing users better experience. However, synthetic speech is prone to be mixed with real human speech as part of the noise and recorded by the microphone, which leads to performance decrease f…

Cited by 0SourceScholar
2018

Finding Tiny Faces in the Wild With Generative Adversarial Network

CVPR 2018poster

Face detection techniques have been developed for decades, and one of remaining open challenges is detecting small faces in unconstrained conditions. The reason is that tiny faces are often lacking detailed information and blurring. In this paper, we proposed an algorithm to directly generate a clea…

Cited by 251SourcePDFScholar
2018

SOD-MTGAN: Small Object Detection via Multi-Task Generative Adversarial Network

ECCV 2018poster

Object detection is a fundamental and important problem in computer vision. Although impressive results have been achieved on large/medium sized objects on large-scale detection benchmarks (e.g. the COCO dataset), the performance on small objects is far from satisfaction. The reason is that small ob…

2018

W2F: A Weakly-Supervised to Fully-Supervised Framework for Object Detection

CVPR 2018poster

Weakly-supervised object detection has attracted much attention lately, since it does not require bounding box annotations for training. Although significant progress has also been made, there is still a large gap in performance between weakly-supervised and fully-supervised object detection. Recent…

Cited by 150SourcePDFScholar