← Search

Yahui Liu

14 accepted papers

2026

All-in-One Slider for Attribute Manipulation in Diffusion Models

CVPR 2026

Text-to-image (T2I) diffusion models have made significant strides in generating high-quality images. However, progressively manipulating certain attributes of generated images to meet the desired user expectations remains challenging, particularly for content with rich details, such as human faces.

Cited by 0SourcecodeScholar
2026

Evaluating Text Creativity across Diverse Domains: a Dataset and Large Language Model Evaluator

ICLR 2026poster

Creativity evaluation remains a challenging frontier for large language models (LLMs). Current evaluations heavily rely on inefficient and costly human judgments, hindering progress in enhancing machine creativity. While automated methods exist, ranging from psychological testing to heuristic- or pr…

Cited by 0SourcecodeScholar
2026

Not Just What’s There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-Tuning

AAAI 2026technical

Vision-Language Models (VLMs) like CLIP struggle to understand negation, often embedding affirmatives and negatives similarly (e.g., matching "no dog" with dog images). Existing methods refine negation understanding via fine-tuning CLIP’s text encoder, risking overfitting. In this work, we propose C

Cited by 0SourcePDFScholar
2026

The Better You Learn, the Smarter You Prune: Towards Efficient Vision-Language-Action Models Via Differentiable Token Pruning

ICRA 2026poster

We present LightVLA, a simple yet effective differentiable token pruning framework for vision-language-action (VLA) models. While VLA models have shown impressive capability in executing real-world robotic tasks, their deployment on resource-constrained platforms is often bottlenecked by the heavy a…

2025

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search

ACL 2025long

Video captioning can be used to assess the video understanding capabilities of Multimodal Large Language Models (MLLMs).However, existing benchmarks and evaluation protocols suffer from crucial issues, such as inadequate or homogeneous creation of key points, exorbitant cost of data creation, and li…

2025

LTOS: Layout-controllable Text-object Synthesis via Adaptive Cross-attention Fusions

ICASSP 2025accepted

Controllable text-to-image generation synthesizes visual text and objects in images with certain conditions. However, existing visual text rendering and layout-to-image generation tasks focus on single modality generation or rendering, leaving yet-to-be-bridged gaps between the approaches correspond…

Cited by 0SourceScholar
2025

What Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought Reasoning

EMNLP 2025

Recent advances in reasoning with large language models (LLMs) have popularized Long Chain-of-Thought (LCoT), a strategy that encourages deliberate and step-by-step reasoning before producing a final answer. While LCoTs have enabled expert-level performance in complex tasks, how the internal structu

Cited by 0SourcePDFScholar
2024

Filamentary Convolution for Spoken Language Identification: A Brain-Inspired Approach

ICASSP 2024accepted

Spoken language identification (SLI) by human beings relies on the hierarchical understanding of one or a few words within the voice signal, encapsulated within the corresponding time windows. Concurrently, frequency-domain features play a crucial role in enhancing identification. The short-time Fou…

Cited by 0SourceScholar
2023

Graph-Based Scenario-Adaptive Lane-Changing Trajectory Planning for Autonomous Driving

RA-L 2023

Trajectory planning is one of the key challenges to the rapid and large-scale deployment of autonomous driving. The lane-changing trajectory planning algorithm for autonomous driving is typically formulated as a optimization process of a cost function, which can be challenging to manually tune for d

Cited by 13SourceScholar
2023

Masked Jigsaw Puzzle: A Versatile Position Embedding for Vision Transformers

CVPR 2023poster

Position Embeddings (PEs), an arguably indispensable component in Vision Transformers (ViTs), have been shown to improve the performance of ViTs on many vision tasks. However, PEs have a potentially high risk of privacy leakage since the spatial information of the input patches is exposed. This cave…

2022

Investigating Data Variance in Evaluations of Automatic Machine Translation Metrics

ACL 2022findings

Current practices in metric evaluation focus on one single dataset, e.g., Newstest dataset in each year’s WMT Metrics Shared Task. However, in this paper, we qualitatively and quantitatively show that the performances of metrics are sensitive to data. The ranking of metrics varies when the evaluatio…

Cited by 4SourcePDFScholar
2022

MuCPAD: A Multi-Domain Chinese Predicate-Argument Dataset

NAACL 2022long

During the past decade, neural network models have made tremendous progress on in-domain semantic role labeling (SRL). However, performance drops dramatically under the out-of-domain setting. In order to facilitate research on cross-domain SRL, this paper presents MuCPAD, a multi-domain Chinese pred…

2021

Efficient Training of Visual Transformers with Small Datasets

NeurIPS 2021poster

Visual Transformers (VTs) are emerging as an architectural paradigm alternative to Convolutional networks (CNNs). Differently from CNNs, VTs can capture global relations between image elements and they potentially have a larger representation capacity. However, the lack of the typical convolutional…

2021

Smoothing the Disentangled Latent Style Space for Unsupervised Image-to-Image Translation

CVPR 2021poster

Image-to-Image (I2I) multi-domain translation models are usually evaluated also using the quality of their semantic interpolation results. However, state-of-the-art models frequently show abrupt changes in the image appearance during interpolation, and usually perform poorly in interpolations across…

Cited by 58PDFScholar