← Search

Feiyi Wang

3 accepted papers

2025

Modulated Diffusion: Accelerating Generative Modeling with Modulated Quantization

ICML 2025poster

Diffusion models have emerged as powerful generative models, but their high computation cost in iterative sampling remains a significant bottleneck. In this work, we present an in-depth and insightful study of state-of-the-art acceleration techniques for diffusion models, including caching and quant…

2025

REALM: Recursive Relevance Modeling for LLM-based Document Re-Ranking

EMNLP 2025

Large Language Models (LLMs) have shown strong capabilities in document re-ranking, a key component in modern Information Retrieval (IR) systems. However, existing LLM-based approaches face notable limitations, including ranking uncertainty, unstable top- k recovery, and high token cost due to token

2024

ProTransformer: Robustify Transformers via Plug-and-Play Paradigm

NeurIPS 2024poster

Transformer-based architectures have dominated various areas of machine learning in recent years. In this paper, we introduce a novel robust attention mechanism designed to enhance the resilience of transformer-based architectures. Crucially, this technique can be integrated into existing transforme…