← Search

Yifei Xia

8 accepted papers

2026

EchoAttention: Exploiting Token-Pair Redundancy and Frame-Block Similarity for Efficient Long Video Generation

ICML 2026poster

Diffusion Transformers (DiTs) are increasingly adopted for long-video generation, yet inference is dominated by the quadratic cost of 3D full attention. Sparse attention mitigates this bottleneck by exploiting *token-pair redundancy* and pruning query-key interactions. Nevertheless, its effectivenes…

Cited by 0SourceScholar
2025

Dense Metric Depth Estimation via Event-based Differential Focus Volume Prompting

NeurIPS 2025poster

Dense metric depth estimation has witnessed great developments in recent years. While single-image-based methods have demonstrated commendable performance in certain circumstances, they may encounter challenges regarding scale ambiguities and visual illusions in real world. Traditional depth-from-fo…

Cited by 0SourcecodeScholar
2025

PanoWan: Lifting Diffusion Video Generation Models to 360$^\circ$ with Latitude/Longitude-aware Mechanisms

NeurIPS 2025poster

Panoramic video generation enables immersive 360$^\circ$ content creation, valuable in applications that demand scene-consistent world exploration. However, existing panoramic video generation models struggle to leverage pre-trained generative priors from conventional text-to-video models for high-q…

Cited by 0SourceScholar
2025

PhyS-EdiT: Physics-aware Semantic Image Editing with Text Description

CVPR 2025poster

Achieving joint control over material properties, lighting, and high-level semantics in images is essential for applications in digital media, advertising, and interactive design. Existing methods often isolate these properties, lacking a cohesive approach to manipulating materials, lighting, and se…

Cited by 0SourcePDFScholar
2025

PlaNet: Learning to Mitigate Atmospheric Turbulence in Planetary Images

AAAI 2025technical

Obtaining planetary images with good visual quality is not an easy task since they are usually degenerated by atmospheric turbulence during the imaging procedure. Existing atmospheric turbulence mitigation methods designed for conventional images cannot be applied to planetary images, since the obje…

Cited by 0SourcePDFScholar
2025

Training-free and Adaptive Sparse Attention for Efficient Long Video Generation

ICCV 2025poster

Generating high-quality long videos with Diffusion Transformers (DiTs) faces significant latency due to computationally intensive attention mechanisms. For instance, generating an 8s 720p video (110K tokens) with HunyuanVideo requires around 600 PFLOPs, with attention computations consuming about 50…

Cited by 0SourcePDFScholar
2024

Efficient Multi-task LLM Quantization and Serving for Multiple LoRA Adapters

NeurIPS 2024poster

With the remarkable achievements of large language models (LLMs), the demand for fine-tuning and deploying LLMs in various downstream tasks has garnered widespread interest. Parameter-efficient fine-tuning techniques represented by LoRA and model quantization techniques represented by GPTQ and AWQ a…

Cited by 3SourcePDFScholar
2024

NB-GTR: Narrow-Band Guided Turbulence Removal

CVPR 2024poster

The removal of atmospheric turbulence is crucial for long-distance imaging. Leveraging the stochastic nature of atmospheric turbulence numerous algorithms have been developed that employ multi-frame input to mitigate the turbulence. However when limited to a single frame existing algorithms face sub…

Cited by 2SourcePDFScholar