← Search

Ngan Le

24 accepted papers

2026

DuFal: Dual-Frequency-Aware Learning for High-Fidelity Extremely Sparse-view CBCT Reconstruction

ICML 2026poster

Sparse-view Cone-Beam Computed Tomography reconstruction from limited X-ray projections remains a challenging problem in medical imaging due to the inherent undersampling of fine-grained anatomical details, which correspond to high-frequency components. Conventional CNN-based methods often struggle …

Cited by 0SourceScholar
2026

Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective

AAAI 2026technical

As embodied agents operate in increasingly complex environments, the ability to perceive, track, and reason about individual object instances over time becomes essential, especially in tasks requiring sequenced interactions with visually similar objects. In non-Markovian settings, critical decision

Cited by 0SourcePDFScholar
2026

SIGMA: A Physics-Based Benchmark for Gas Chimney Understanding in Seismic Images

CVPR 2026

Seismic images reconstruct subsurface reflectivity from field recordings, guiding exploration and reservoir monitoring. Gas chimneys are vertical anomalies caused by subsurface fluid migration. Understanding these phenomena is crucial for assessing hydrocarbon potential and avoiding drilling hazards

Cited by 0SourcecodeScholar
2026

SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection

CVPR 2026

Camera-only 3D object detection has emerged as a cost-effective and scalable alternative to LiDAR for autonomous driving, yet existing methods primarily prioritize overall performance while overlooking the severe long-tail imbalance inherent in real-world datasets. In practice, many rare but safety-

Cited by 0SourceScholar
2025

CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling

ICCV 2025poster

Understanding radiologists' eye movement during Computed Tomography (CT) reading is crucial for developing effective interpretable computer-aided diagnosis systems. However, CT research in this area has been limited by the lack of publicly available eye-tracking datasets and the three-dimensional co…

2025

EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba

ICCV 2025poster

Estimating human dance motion is a challenging task with various industrial applications. Recently, many efforts have focused on predicting human dance motion using either egocentric video or music as input. However, the task of jointly estimating human motion from both egocentric video and music re…

Cited by 0SourcePDFScholar
2025

Robotic-CLIP: Fine-Tuning CLIP on Action Data for Robotic Applications

ICRA 2025

Vision language models have played a key role in extracting meaningful features for various robotic applications. Among these, Contrastive Language-Image Pretraining (CLIP) is widely used in robotic tasks that require both vision and natural language understanding. However, CLIP was trained solely o

Cited by 11SourceScholar
2024

Language-Conditioned Affordance-Pose Detection in 3D Point Clouds

ICRA 2024poster

Affordance detection and pose estimation are of great importance in many robotic applications. Their combination helps the robot gain an enhanced manipulation capability, in which the generated pose can facilitate the corresponding affordance task. Previous methods for affodance-pose joint learning…

Cited by 17SourcecodeScholar
2024

Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance

ECCV 2024oral

"6-DoF grasp detection has been a fundamental and challenging problem in robotic vision. While previous works have focused on ensuring grasp stability, they often do not consider human intention conveyed through natural language, hindering effective collaboration between robots and users in complex…

2024

Language-driven Grasp Detection with Mask-guided Attention

IROS 2024poster

Grasp detection is an essential task in robotics with various industrial applications. However, traditional methods often struggle with occlusions and do not utilize language for grasping. Incorporating natural language into grasp detection remains a challenging task and largely unexplored. To addre…

Cited by 1SourceScholar
2024

Lightweight Language-driven Grasp Detection using Conditional Consistency Model

IROS 2024

Language-driven grasp detection is a fundamental yet challenging task in robotics with various industrial applications. This work presents a new approach for language-driven grasp detection that leverages lightweight diffusion models to achieve fast inference time. By integrating diffusion processes

Cited by 12SourceScholar
2024

Open-Fusion: Real-time Open-Vocabulary 3D Mapping and Queryable Scene Representation

ICRA 2024poster

Precise 3D environmental mapping with semantics is essential in robotics. Existing methods often rely on pre-defined concepts during training or are time-intensive when generating semantic maps. This paper presents Open-Fusion, an approach for real-time open-vocabulary 3D mapping and queryable scene…

Cited by 31SourcecodeScholar
2024

Open-Vocabulary Affordance Detection using Knowledge Distillation and Text-Point Correlation

ICRA 2024poster

Affordance detection presents intricate challenges and has a wide range of robotic applications. Previous works have faced limitations such as the complexities of 3D object shapes, the wide range of potential affordances on real-world objects, and the lack of open-vocabulary support for affordance u…

Cited by 10SourcecodeScholar
2024

WAVER: Writing-Style Agnostic Text-Video Retrieval Via Distilling Vision-Language Models Through Open-Vocabulary Knowledge

ICASSP 2024accepted

Text-video retrieval, a prominent sub-field within the domain of multimodal information retrieval, has witnessed remarkable growth in recent years. However, existing methods assume video scenes are consistent with unbiased descriptions. These limitations fail to align with real-world scenarios since…

Cited by 0SourceScholar
2024

Z-GMOT: Zero-shot Generic Multiple Object Tracking

NAACL 2024findings

Despite recent significant progress, Multi-Object Tracking (MOT) faces limitations such as reliance on prior knowledge and predefined categories and struggles with unseen objects. To address these issues, Generic Multiple Object Tracking (GMOT) has emerged as an alternative approach, requiring less…

2023

FREDOM: Fairness Domain Adaptation Approach to Semantic Scene Understanding

CVPR 2023poster

Although Domain Adaptation in Semantic Scene Segmentation has shown impressive improvement in recent years, the fairness concerns in the domain adaptation have yet to be well defined and addressed. In addition, fairness is one of the most critical aspects when deploying the segmentation models into…

2023

Open-Vocabulary Affordance Detection in 3D Point Clouds

IROS 2023poster

Affordance detection is a challenging problem with a wide variety of robotic applications. Traditional affordance detection methods are limited to a predefined set of affordance labels, hence potentially restricting the adaptability of intelligent robots in complex and dynamic environments. In this…

Cited by 33SourcecodeScholar
2023

VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph Captioning

AAAI 2023technical

Video Paragraph Captioning aims to generate a multi-sentence description of an untrimmed video with multiple temporal event locations in a coherent storytelling. Following the human perception process, where the scene is effectively understood by decomposing it into visual (e.g. human, animal) and…

2021

Agent-Environment Network for Temporal Action Proposal Generation

ICASSP 2021accepted

Temporal action proposal generation is an essential and challenging task that aims at localizing temporal intervals containing human actions in untrimmed videos. Most of existing approaches are unable to follow the human cognitive process of understanding the video context due to lack of attention m…

Cited by 0SourceScholar
2021

BiMaL: Bijective Maximum Likelihood Approach to Domain Adaptation in Semantic Scene Segmentation

ICCV 2021poster

Semantic segmentation aims to predict pixel-level labels. It has become a popular task in various computer vision applications. While fully supervised segmentation methods have achieved high accuracy on large-scale vision datasets, they are unable to generalize on a new test environment or a new dom…

Cited by 43PDFcodeScholar
2021

The Right To Talk: An Audio-Visual Transformer Approach

ICCV 2021poster

Turn-taking has played an essential role in structuring the regulation of a conversation. The task of identifying the main speaker (who is properly taking his/her turn of speaking) and the interrupters (who are interrupting or reacting to the main speaker's utterances) remains a challenging task. Al…

Cited by 46PDFcodeScholar
2019

Automatic Face Aging in Videos via Deep Reinforcement Learning

CVPR 2019poster

This paper presents a novel approach for synthesizing automatically age-progressed facial images in video sequences using Deep Reinforcement Learning. The proposed method models facial structures and the longitudinal face-aging process of given subjects coherently across video frames. The approach i…

Cited by 46PDFScholar
2017

Temporal Non-Volume Preserving Approach to Facial Age-Progression and Age-Invariant Face Recognition

ICCV 2017oral

Modeling the long-term facial aging process is extremely challenging due to the presence of large and non-linear variations during the face development stages. In order to efficiently address the problem, this work first decomposes the aging process into multiple short-term stages. Then, a novel gen…

Cited by 97PDFScholar