← Search

Alberto Gil Couto Pimentel Ramos

4 accepted papers

2026

NanoFLUX: Distillation-Driven Compression of Large Text-to-Image Generation Models for Mobile Devices

ICML 2026poster

While large-scale text-to-image diffusion models continue to improve in visual quality, their increasing scale has widened the gap between state-of-the-art models and on-device solutions. To address this gap, we introduce NanoFLUX, a **2.4B** text-to-image flow-matching model distilled from **17B** …

Cited by 0SourceScholar
2026

RFDM: Residual Flow Diffusion Models for Video Editing

CVPR 2026

Instructional video editing applies edits to an input video using only text prompts, enabling intuitive natural-language control. Despite the rapid progress, most methods still require fixed-length inputs and substantial compute. Meanwhile, autoregressive video generation enables efficient variable-

Cited by 0SourcecodeScholar
2025

Upcycling Text-to-Image Diffusion Models for Multi-Task Capabilities

ICML 2025poster

Text-to-image synthesis has witnessed remarkable advancements in recent years. Many attempts have been made to adopt text-to-image models to support multiple tasks. However, existing approaches typically require resource-intensive re-training or additional parameters to accommodate for the new tasks…

Cited by 0SourcePDFScholar
2022

Conditioning Sequence-to-sequence Networks with Learned Activations

ICLR 2022poster

Conditional neural networks play an important role in a number of sequence-to-sequence modeling tasks, including personalized sound enhancement (PSE), speaker dependent automatic speech recognition (ASR), and generative modeling such as text-to-speech synthesis. In conditional neural networks, the o…

Cited by 13SourcePDFScholar