← Search

Kevin Duarte

10 accepted papers

2025

CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation

ICCV 2025poster

In text-to-image (T2I) generation, achieving fine-grained control over attributes - such as age or smile - remains challenging, even with detailed text prompts. Slider-based methods offer a solution for precise control of image attributes.Existing approaches typically train individual adapter for ea…

Cited by 0SourcePDFScholar
2024

Plug-and-Play Diffusion Distillation

CVPR 2024poster

Diffusion models have shown tremendous results in image generation. However due to the iterative nature of the diffusion process and its reliance on classifier-free guidance inference times are slow. In this paper we propose a new distillation approach for guided diffusion models in which an externa…

Cited by 10SourcePDFScholar
2021

Found a Reason for me? Weakly-supervised Grounded Visual Question Answering using Capsules

CVPR 2021poster

The problem of grounding VQA tasks has seen an increased attention in the research community recently, with most attempts usually focusing on solving this task by using pretrained object detectors. However, pre-trained object detectors require bounding box annotations for detecting relevant objects…

Cited by 46PDFcodeScholar
2021

In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised Learning

ICLR 2021poster

The recent research in semi-supervised learning (SSL) is mostly dominated by consistency regularization based methods which achieve strong performance. However, they heavily rely on domain-specific data augmentations, which are not easy to generate for all data modalities. Pseudo-labeling (PL) is a…

2021

Modeling Multi-Label Action Dependencies for Temporal Action Localization

CVPR 2021poster

Real world videos contain many complex actions with inherent relationships between action classes. In this work, we propose an attention-based architecture that model these action relationships for the task of temporal action localization in untrimmed videos. As opposed to previous works which lever…

Cited by 82PDFcodeScholar
2021

Multimodal Clustering Networks for Self-Supervised Learning From Unlabeled Videos

ICCV 2021poster

Multimodal self-supervised learning is getting more and more attention as it allows not only to train large networks without human supervision but also to search and retrieve data across various modalities. In this context, this paper proposes a framework that, starting from a pre-trained backbone,…

Cited by 110PDFcodeScholar
2021

Reformulating Zero-shot Action Recognition for Multi-label Actions

NeurIPS 2021poster

The goal of zero-shot action recognition (ZSAR) is to classify action classes which were not previously seen during training. Traditionally, this is achieved by training a network to map, or regress, visual inputs to a semantic space where a nearest neighbor classifier is used to select the closest…

Cited by 25SourcePDFScholar
2019

CapsuleVOS: Semi-Supervised Video Object Segmentation Using Capsule Routing

ICCV 2019poster

In this work we propose a capsule-based approach for semi-supervised video object segmentation. Current video object segmentation methods are frame-based and often require optical flow to capture temporal consistency across frames which can be difficult to compute. To this end, we propose a video ba…

Cited by 84PDFcodeScholar