← Search

Hieu Le

25 accepted papers

2026

MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed Videos

CVPR 2026

Temporal Action Detection (TAD) in untrimmed videos poses significant challenges, particularly for Activities of Daily Living (ADL) requiring models to (1) process long-duration videos, (2) capture temporal variations in actions, and (3) simultaneously detect dense overlapping actions. Existing CNN

Cited by 0SourcecodeScholar
2026

Personalized Image Descriptions from Attention Sequences

CVPR 2026

People can view the same image differently: they focus on different regions, objects, and details in varying orders and describe them in distinct linguistic styles. This leads to substantial variability in image descriptions. However, existing models for personalized image description focus on lingu

Cited by 0SourcecodeScholar
2025

Enhancing Compositional Text-to-Image Generation with Reliable Random Seeds

ICLR 2025spotlight

Text-to-image diffusion models have demonstrated remarkable capability in generating realistic images from arbitrary text prompts. However, they often produce inconsistent results for compositional prompts such as "two dogs" or "a penguin on the right of a bowl". Understanding these inconsistencies…

Cited by 1SourcePDFScholar
2025

Few-shot Personalized Scanpath Prediction

CVPR 2025poster

A personalized model for scanpath prediction provides insights into the visual preferences and attention patterns of individual subjects. However, existing methods for training scanpath prediction models are data-intensive and cannot be effectively personalized to new individuals with only a few ava…

2025

Importance-Based Token Merging for Efficient Image and Video Generation

ICCV 2025poster

Token merging can effectively accelerate various vision systems by processing groups of similar tokens only once and sharing the results across them. However, existing token grouping methods are often ad hoc and random, disregarding the actual content of the samples. We show that preserving high-inf…

Cited by 0SourcePDFScholar
2024

Beyond Pixels: Semi-Supervised Semantic Segmentation with a Multi-scale Patch-based Multi-Label Classifier

ECCV 2024poster

"Incorporating pixel contextual information is critical for accurate segmentation. In this paper, we show that an effective way to incorporate contextual information is through a patch-based classifier. This patch classifier is trained to identify classes present within an image region, which facili…

2024

Enabling Uncertainty Estimation in Iterative Neural Networks

ICML 2024poster

Turning pass-through network architectures into iterative ones, which use their own output as input, is a well-known approach for boosting performance. In this paper, we argue that such architectures offer an additional benefit: The convergence rate of their successive outputs is highly correlated w…

2024

Weighting Pseudo-Labels via High-Activation Feature Index Similarity and Object Detection for Semi-Supervised Segmentation

ECCV 2024poster

"Semi-supervised semantic segmentation methods leverage unlabeled data by pseudo-labeling them. Thus the success of these methods hinges on the reliability of the pseudo-labels. Existing methods mostly choose high-confidence pixels in an effort to avoid erroneous pseudo-labels. However, high confide…

2023

Generating Features With Increased Crop-Related Diversity for Few-Shot Object Detection

CVPR 2023poster

Two-stage object detectors generate object proposals and classify them to detect objects in images. These proposals often do not perfectly contain the objects but overlap with them in many possible ways, exhibiting great variability in the difficulty levels of the proposals. Training a robust classi…

Cited by 45SourcePDFScholar
2022

Semi-supervised Adversarial Text Generation based on Seq2Seq models

EMNLP 2022industry

To improve deep learning models’ robustness, adversarial training has been frequently used in computer vision with satisfying results. However, adversarial perturbation on text have turned out to be more challenging due to the discrete nature of text. The generated adversarial text might not sound n…

Cited by 5SourcePDFScholar
2021

Variational Feature Disentangling for Fine-Grained Few-Shot Classification

ICCV 2021poster

Data augmentation is an intuitive step towards solving the problem of few-shot classification. However, ensuring both discriminability and diversity in the augmented samples is challenging. To address this, we propose a feature disentanglement framework that allows us to augment features with random…

Cited by 76PDFcodeScholar
2018

A+D Net: Training a Shadow Detector with Adversarial Shadow Attenuation

ECCV 2018poster

We propose a novel GAN-based framework for detecting shadows in images, in which a shadow detection network (D-Net) is trained together with a shadow attenuation network (A-Net) that generates adversarial training examples. The A-Net modifies the original training images constrained by a simplified…

Cited by 142SourcePDFScholar
2015

Efficient Video Segmentation Using Parametric Graph Partitioning

ICCV 2015poster

Video segmentation is the task of grouping similar pixels in the spatio-temporal domain, and has become an important preprocessing step for subsequent video analysis. Most video segmentation and supervoxel methods output a hierarchy of segmentations, but while this provides useful multiscale informa…

Cited by 35PDFScholar