← Search

Pengtao Chen

7 accepted papers

2026

ReasonEdit: Towards Reasoning-Enhanced Image Editing Models

CVPR 2026

Recent advances in image editing models have shown remarkable progress. A common architectural design couples a multimodal large language model (MLLM) encoder with a diffusion decoder, as seen in systems such as Step1X-Edit and Qwen-Image-Edit, where the MLLM encodes both the reference image and the

Cited by 0SourcecodeScholar
2026

RegionE: Adaptive Region-Aware Generation for Efficient Image Editing

ICLR 2026poster

Recently, instruction-based image editing (IIE) has received widespread attention. In practice, IIE often modifies only specific regions of an image, while the remaining areas largely remain unchanged. Although these two types of regions differ significantly in generation difficulty and computationa…

Cited by 0SourceScholar
2026

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers

AAAI 2026technical

While Diffusion Transformers (DiTs) have achieved breakthroughs in video generation, this long sequence generation task remains constrained by the quadratic complexity of attention mechanisms, resulting in significant inference latency. Through detailed analysis of attention maps in Video Diffusion

Cited by 0SourcePDFScholar
2025

DiTFastAttnV2: Head-wise Attention Compression for Multi-Modality Diffusion Transformers

ICCV 2025poster

Text-to-image generation models, especially Multimodal Diffusion Transformers (MMDiT), have shown remarkable progress in generating high-quality images. However, these models often face significant computational bottlenecks, particularly in attention mechanisms, which hinder their scalability and ef…

2025

FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding

NeurIPS 2025poster

Multimodal Large Language Models (MLLMs) have shown impressive video content understanding capabilities but struggle with fine-grained motion comprehension. To comprehensively assess the motion understanding ability of existing MLLMs, we introduce FAVOR-Bench, which comprises 1,776 videos from both…

Cited by 0SourceScholar
2025

Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuning

CVPR 2025poster

Model quantization reduces the bit-width of weights and activations, improving memory efficiency and inference speed in diffusion models. However, achieving 4-bit quantization remains challenging. Existing methods, primarily based on integer quantization and post-training quantization fine-tuning, s…

Cited by 0SourcePDFScholar
2024

A Manta Ray-Inspired Fast-Swimming Soft Electrohydraulic Robotic Fish

RA-L 2024

Underwater soft robots inspired by marine life have shown great potential in ocean exploration, monitoring, scientific research, etc., due to their excellent safety, compatibility and adaptability when interacting with underwater environments. However, most of their soft actuators suffer performance

Cited by 11SourceScholar