← Search

Yumeng Zhang

10 accepted papers

2026

Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation

ICML 2026poster

Latent diffusion models have enabled high-quality video synthesis, yet their inference remains costly and time-consuming. As diffusion transformers become increasingly efficient, the latency bottleneck inevitably shifts to VAE decoders. To reduce their latency while maintaining quality, we propose a…

Cited by 0SourceScholar
2026

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection

CVPR 2026

Recent advances in Vision-Language Models (VLMs) have benefited from Reinforcement Learning (RL) for enhanced reasoning. However, existing methods still face critical limitations, including the lack of low-level visual information and effective visual feedback. To address these problems, this paper

Cited by 0SourceScholar
2025

Stephanie: Step-by-Step Dialogues for Mimicking Human Interactions in Social Conversations

NAACL 2025findings

In the rapidly evolving field of natural language processing, dialogue systems primarily employ a single-step dialogue paradigm. Although this paradigm is commonly adopted, it lacks the depth and fluidity of human interactions and does not appear natural. We introduce a novel **Step**-by-Step Dialog…

Cited by 2SourcePDFScholar
2025

U-ViLAR: Uncertainty-Aware Visual Localization for Autonomous Driving via Differentiable Association and Registration

ICCV 2025poster

Accurate localization using visual information is a critical yet challenging task, especially in urban environments where nearby buildings and construction sites significantly degrade GNSS (Global Navigation Satellite System) signal quality. This issue underscores the importance of visual localizati…

Cited by 0SourcePDFScholar
2024

CLEAN–EVAL: Clean Evaluation on Contaminated Large Language Models

NAACL 2024findings

We are currently in an era of fierce competition among various large language models (LLMs), continuously pushing the boundaries of benchmark performance. However, genuinely assessing the capabilities of these LLMs has become a challenging and critical issue due to potential data contamination. In t…

Cited by 17SourcePDFScholar
2024

Practical Anytime Algorithms for Judicious Partitioning of Active Directory Attack Graphs

IJCAI 2024poster

Given a directed graph, a set of source nodes, a target node and a budget, we study the problem of maximizing the number of source nodes disconnected from the target node by removing edges not exceeding the budget. Our model is mainly motivated by a cyber security use case where we need to minimize…

2024

Unveiling the Generalization Power of Fine-Tuned Large Language Models

NAACL 2024long

While Large Language Models (LLMs) have demonstrated exceptional multitasking abilities, fine-tuning these models on downstream, domain-specific datasets is often necessary to yield superior performance on test sets compared to their counterparts without fine-tuning. However, the comprehensive effec…

2023

Forward Flow for Novel View Synthesis of Dynamic Scenes

ICCV 2023oral

This paper proposes a neural radiance field (NeRF) approach for novel view synthesis of dynamic scenes using forward warping. Existing methods often adopt a static NeRF to represent the canonical space, and render dynamic images at other time steps by mapping the sampled 3D points back to the canoni…

Cited by 48PDFcodeScholar
2022

Waveform Optimization for Wireless Power Transfer with Power Amplifier and Energy Harvester Non-linearities

ICASSP 2022accepted

Waveform optimization has recently been shown to be a key technique to boost the efficiency and range of far-field wireless power transfer (WPT). Current research has optimized transmit waveform adaptive to channel state information (CSI) and accounting for energy harvester (EH)’s non-linearity but…

Cited by 0SourceScholar