← Search

Anh Nguyen

71 accepted papers

2026

AeroScene: Progressive Scene Synthesis for Aerial Robotics

ICRA 2026poster

Generative models have shown substantial impact across multiple domains, their potential for scene synthesis remains underexplored in robotics. This gap is more evident in drone simulators, where simulation environments still rely heavily on manual efforts, which are time-consuming to create and dif…

2026

AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers

CVPR 2026

Affordance learning is a complex challenge in many applications, where existing approaches primarily focus on the geometric structures, visual knowledge, and affordance labels of objects to determine interactable regions. However, extending this learning capability to a scene is significantly more c

Cited by 0SourcecodeScholar
2026

One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual Learning

ICLR 2026poster

Prompt-based methods have recently gained prominence in Continual Learning (CL) due to their strong performance and memory efficiency. A prevalent strategy in this paradigm assigns a dedicated subset of prompts to each task, which, while effective, incurs substantial computational overhead and cause…

Cited by 0SourcecodeScholar
2026

Provably Data-driven Lagrangian Relaxation for Mixed Integer Linear Programming

ICML 2026poster

Lagrangian Relaxation (LR) is a powerful technique for solving large-scale Mixed Integer Linear Programming (MILP), particularly those with decomposable structures like Vehicle Routing or Unit Commitment. By relaxing coupling constraints, LR enables parallel solving of subproblems and frequently yie…

Cited by 0SourceScholar
2026

Provably Data-driven Multiple Hyper-parameter Tuning with Structured Loss Function

ICML 2026poster

Data-driven algorithm design automates hyperparameter tuning, but its statistical foundations remain limited because model performance can depend on hyperparameters in implicit and highly non-smooth ways. Existing guarantees focus on the simple case of a one-dimensional (scalar) hyperparameter. This…

Cited by 0SourceScholar
2026

Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective

AAAI 2026technical

As embodied agents operate in increasingly complex environments, the ability to perceive, track, and reason about individual object instances over time becomes essential, especially in tasks requiring sequenced interactions with visually similar objects. In non-Markovian settings, critical decision

Cited by 0SourcePDFScholar
2026

Revisit Visual Prompt Tuning: The Expressiveness of Prompt Experts

ICLR 2026poster

Visual Prompt Tuning (VPT) has proven effective for parameter-efficient adaptation of pre-trained vision models to downstream tasks by inserting task-specific learnable prompt tokens. Despite its empirical success, a comprehensive theoretical understanding of VPT remains an active area of research.…

Cited by 0SourcecodeScholar
2026

SIGMA: A Physics-Based Benchmark for Gas Chimney Understanding in Seismic Images

CVPR 2026

Seismic images reconstruct subsurface reflectivity from field recordings, guiding exploration and reservoir monitoring. Gas chimneys are vertical anomalies caused by subsurface fluid migration. Understanding these phenomena is crucial for assessing hydrocarbon potential and avoiding drilling hazards

Cited by 0SourcecodeScholar
2026

SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection

CVPR 2026

Camera-only 3D object detection has emerged as a cost-effective and scalable alternative to LiDAR for autonomous driving, yet existing methods primarily prioritize overall performance while overlooking the severe long-tail imbalance inherent in real-world datasets. In practice, many rare but safety-

Cited by 0SourceScholar
2026

Sequential Information Bottleneck Fusion: Towards Robust and Generalizable Multi-Modal Brain Tumor Segmentation

ICLR 2026poster

Brain tumor segmentation in multi-modal MRIs poses significant challenges when one or more modalities are missing. Recent approaches commonly employ parallel fusion strategies; however, these methods often risk losing crucial shared information across modalities, which can degrade segmentation perfo…

Cited by 0SourceScholar
2026

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

ICML 2026poster

Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children or animals. Despite this, they attain high scores on many popular visual benchmarks, with headroom rapidly eroded by surging model progress. To address …

Cited by 0SourceScholar
2025

Are Spatial-Temporal Graph Convolution Networks for Human Action Recognition Over-Parameterized?

CVPR 2025poster

Spatial-temporal graph convolutional networks (ST-GCNs) showcase impressive performance in skeleton-based human action recognition (HAR). However, despite the development of numerous models, their recognition performance does not differ significantly after aligning the input settings. With this obse…

2025

CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling

ICCV 2025poster

Understanding radiologists' eye movement during Computed Tomography (CT) reading is crucial for developing effective interpretable computer-aided diagnosis systems. However, CT research in this area has been limited by the lack of publicly available eye-tracking datasets and the three-dimensional co…

2025

EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba

ICCV 2025poster

Estimating human dance motion is a challenging task with various industrial applications. Recently, many efforts have focused on predicting human dance motion using either egocentric video or music as input. However, the task of jointly estimating human motion from both egocentric video and music re…

Cited by 0SourcePDFScholar
2025

FedEFM: Federated Endovascular Foundation Model with Unseen Data

ICRA 2025

In endovascular surgery, the precise identification of catheters and guidewires in X-ray images is essential for reducing intervention risks. However, accurately segmenting catheter and guidewire structures is challenging due to the limited availability of labeled data. Foundation models offer a pro

Cited by 3SourceScholar
2025

Fractal Calibration for Long-tailed Object Detection

CVPR 2025poster

Real-world datasets follow an imbalanced distribution, which poses significant challenges in rare-category object detection. Recent studies tackle this problem by developing re-weighting and re-sampling methods, that utilise the class frequencies of the dataset. However, these techniques focus solel…

2025

GraspMAS: Zero-Shot Language-driven Grasp Detection with Multi-Agent System

IROS 2025

Language-driven grasp detection has the potential to revolutionize human-robot interaction by allowing robots to understand and execute grasping tasks based on natural language commands. However, existing approaches face two key challenges. First, they often struggle to interpret complex text instru

Cited by 0SourcecodeScholar
2025

GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature Learning

IROS 2025

Grasp detection is a fundamental robotic task critical to the success of many industrial applications. However, current language-driven models for this task often struggle with cluttered images, lengthy textual descriptions, or slow inference speed. We introduce GraspMamba, a new language-driven gra

Cited by 5SourceScholar
2025

Hybrid Gripper with Passive Pneumatic Soft Joints for Grasping Deformable Thin Objects

ICRA 2025

Grasping a variety of objects remains a key challenge in the development of versatile robotic systems. The human hand is remarkably dexterous, capable of grasping and manipulating objects with diverse shapes, mechanical properties, and textures. Inspired by how humans use two fingers to pick up thin

Cited by 4SourceScholar
2025

Improved Training Technique for Shortcut Models

NeurIPS 2025poster

Shortcut models represent a promising, non-adversarial paradigm for generative modeling, uniquely supporting one-step, few-step, and multi-step sampling from a single trained network. However, their widespread adoption has been stymied by critical performance bottlenecks. This paper tackles the five…

Cited by 0SourceScholar
2025

LP-Diff: Towards Improved Restoration of Real-World Degraded License Plate

CVPR 2025highlight

License plate (LP) recognition is crucial in intelligent traffic management systems. However, factors such as long distances and poor camera quality often lead to severe degradation of captured LP images, posing challenges to accurate recognition. The design of License Plate Image Restoration (LPIR)…

2025

Lightweight Temporal Transformer Decomposition for Federated Autonomous Driving

IROS 2025

Traditional vision-based autonomous driving systems often face difficulties in navigating complex environments when relying solely on single-image inputs. To overcome this limitation, incorporating temporal data such as past image frames or steering sequences, has proven effective in enhancing robus

Cited by 0SourcecodeScholar
2025

Modeling The States of Liquid Phase Change Pouch Actuators by Reservoir Computing

IROS 2025

Liquid phase change pouch actuators (liquid pouch motors) hold great promise for a wide range of robotic applications, from artificial organs to pneumatic manipulators for dexterous manipulation. However, the usability of liquid pouch motors remains challenging due to the nonlinear intrinsic propert

Cited by 0SourcecodeScholar
2025

More Reliable Pseudo-labels, Better Performance: A Generalized Approach to Single Positive Multi-label Learning

ICCV 2025poster

Multi-label learning is a challenging computer vision task that requires assigning multiple categories to each image. However, fully annotating large-scale datasets is often impractical due to high costs and effort, motivating the study of learning from partially annotated data. In the extreme case…

Cited by 0SourcePDFScholar
2025

NUMINA: A Natural Understanding Benchmark for Multi-dimensional Intelligence and Numerical Reasoning Abilities

EMNLP 2025

Recent advancements in 2D multimodal large language models (MLLMs) have significantly improved performance in vision-language tasks. However, extending these capabilities to 3D environments remains a distinct challenge due to the complexity of spatial reasoning. Nevertheless, existing 3D benchmarks

2025

Online Trajectory Replanner for Dynamically Grasping Irregular Objects

ICRA 2025

This paper presents a new trajectory replanner for grasping irregular objects. Unlike conventional grasping tasks where the object's geometry is assumed simple, we aim to achieve a “dynamic grasp” of the irregular objects, which requires continuous adjustment during the grasping process. To effectiv

Cited by 0SourceScholar
2025

Robotic-CLIP: Fine-Tuning CLIP on Action Data for Robotic Applications

ICRA 2025

Vision language models have played a key role in extracting meaningful features for various robotic applications. Among these, Contrastive Language-Image Pretraining (CLIP) is widely used in robotic tasks that require both vision and natural language understanding. However, CLIP was trained solely o

Cited by 11SourceScholar
2025

SplineFormer: An Explainable Transformer Network for Autonomous Endovascular Navigation

IROS 2025

Robot-assisted endovascular navigation provides significant advantages, including reduced radiation exposure for surgeons and improved patient safety. However, a major challenge is to control curvilinear instruments like guidewires precisely for smooth and accurate navigation while adapting to anato

Cited by 0SourceScholar
2025

Supercharged One-step Text-to-Image Diffusion Models with Negative Prompts

ICCV 2025poster

The escalating demand for real-time image synthesis has driven significant advancements in one-step diffusion models, which inherently offer expedited generation speeds compared to traditional multi-step methods. However, this enhanced efficiency is frequently accompanied by a compromise in the cont…

Cited by 0SourcePDFScholar
2025

Towards Autonomous Wood-Log Grasping with a Forestry Crane: Simulator and Benchmarking

ICRA 2025

Forestry machines operated in forest production environments face challenges when performing manipulation tasks, especially regarding the complicated dynamics of underactuated crane systems and the heavy weight of logs to be grasped. This study investigates the feasibility of using reinforcement lea

Cited by 4SourceScholar
2025

Towards a Universal 3D Medical Multi-modality Generalization via Learning Personalized Invariant Representation

ICCV 2025poster

Variations in medical imaging modalities and individual anatomical differences pose challenges to cross-modality generalization in multi-modal tasks. Existing methods often concentrate exclusively on common anatomical patterns, thereby neglecting individual differences and consequently limiting thei…

2025

Weakly-Supervised Learning via Multi-Lateral Decoder Branching for Tool Segmentation in Robot-Assisted Cardiovascular Catheterization

ICRA 2025

Robot-assisted catheterization has garnered a good attention for its potentials in treating cardiovascular diseases. However, advancing surgeon-robot collaboration still requires further research, particularly on task-specific automation. For instance, automated tool segmentation can assist surgeons

Cited by 0SourceScholar
2024

Dynamic Semantic-Based Spatial Graph Convolution Network for Skeleton-Based Human Action Recognition

AAAI 2024technical

Graph convolutional networks (GCNs) have attracted great attention and achieved remarkable performance in skeleton-based action recognition. However, most of the previous works are designed to refine skeleton topology without considering the types of different joints and edges, making them infeasibl…

2024

GlitchBench: Can Large Multimodal Models Detect Video Game Glitches?

CVPR 2024poster

Large multimodal models (LMMs) have evolved from large language models (LLMs) to integrate multiple input modalities such as visual inputs. This integration augments the capacity of LLMs for tasks requiring visual comprehension and reasoning. However the extent and limitations of their enhanced abil…

2024

HabiCrowd: A High Performance Simulator for Crowd-Aware Visual Navigation

IROS 2024poster

Visual navigation, a foundational aspect of Embodied AI (E-AI) and robotics has been extensively studied in the past few years. While many 3D simulators have been introduced for the visual navigation tasks, scarcely works have combined human dynamics, creating the gap between simulation and real-wor…

Cited by 3SourcecodeScholar
2024

Interpret Your Decision: Logical Reasoning Regularization for Generalization in Visual Classification

NeurIPS 2024spotlight

Vision models excel in image classification but struggle to generalize to unseen data, such as classifying images from unseen domains or discovering novel categories. In this paper, we explore the relationship between logical reasoning and deep learning generalization in visual classification. A log…

2024

Language-Conditioned Affordance-Pose Detection in 3D Point Clouds

ICRA 2024poster

Affordance detection and pose estimation are of great importance in many robotic applications. Their combination helps the robot gain an enhanced manipulation capability, in which the generated pose can facilitate the corresponding affordance task. Previous methods for affodance-pose joint learning…

Cited by 17SourcecodeScholar
2024

Language-driven Grasp Detection with Mask-guided Attention

IROS 2024poster

Grasp detection is an essential task in robotics with various industrial applications. However, traditional methods often struggle with occlusions and do not utilize language for grasping. Incorporating natural language into grasp detection remains a challenging task and largely unexplored. To addre…

Cited by 1SourceScholar
2024

Lightweight Language-driven Grasp Detection using Conditional Consistency Model

IROS 2024

Language-driven grasp detection is a fundamental yet challenging task in robotics with various industrial applications. This work presents a new approach for language-driven grasp detection that leverages lightweight diffusion models to achieve fast inference time. By integrating diffusion processes

Cited by 12SourceScholar
2024

Open-Fusion: Real-time Open-Vocabulary 3D Mapping and Queryable Scene Representation

ICRA 2024poster

Precise 3D environmental mapping with semantics is essential in robotics. Existing methods often rely on pre-defined concepts during training or are time-intensive when generating semantic maps. This paper presents Open-Fusion, an approach for real-time open-vocabulary 3D mapping and queryable scene…

Cited by 31SourcecodeScholar
2024

Open-Vocabulary Affordance Detection using Knowledge Distillation and Text-Point Correlation

ICRA 2024poster

Affordance detection presents intricate challenges and has a wide range of robotic applications. Previous works have faced limitations such as the complexities of 3D object shapes, the wide range of potential affordances on real-world objects, and the lack of open-vocabulary support for affordance u…

Cited by 10SourcecodeScholar
2024

PEEB: Part-based Image Classifiers with an Explainable and Editable Language Bottleneck

NAACL 2024findings

CLIP-based classifiers rely on the prompt containing a class name that is known to the text encoder. Therefore, they perform poorly on new classes or the classes whose names rarely appear on the Internet (e.g., scientific names of birds). For fine-grained classification, we propose PEEB – an explain…

2024

Reducing Non-IID Effects in Federated Autonomous Driving with Contrastive Divergence Loss

ICRA 2024poster

Federated learning has been widely applied in autonomous driving since it enables training a learning model among vehicles without sharing users’ data. However, data from autonomous vehicles usually suffer from the non-independent-and-identically-distributed (non-IID) problem, which may cause negati…

Cited by 0SourcecodeScholar
2024

WAVER: Writing-Style Agnostic Text-Video Retrieval Via Distilling Vision-Language Models Through Open-Vocabulary Knowledge

ICASSP 2024accepted

Text-video retrieval, a prominent sub-field within the domain of multimodal information retrieval, has witnessed remarkable growth in recent years. However, existing methods assume video scenes are consistent with unbiased descriptions. These limitations fail to align with real-world scenarios since…

Cited by 0SourceScholar
2023

CodeT: Code Generation with Generated Tests

ICLR 2023poster

The task of generating code solutions for a given programming problem can benefit from the use of pre-trained language models such as Codex, which can produce multiple diverse samples. However, a major challenge for this task is to select the most appropriate solution from the multiple samples gener…

2023

Language-driven Scene Synthesis using Multi-conditional Diffusion Model

NeurIPS 2023poster

Scene synthesis is a challenging problem with several industrial applications. Recently, substantial efforts have been directed to synthesize the scene using human motions, room layouts, or spatial graphs as the input. However, few studies have addressed this problem from multiple modalities, especi…

2023

Music-Driven Group Choreography

CVPR 2023poster

Music-driven choreography is a challenging problem with a wide variety of industrial applications. Recently, many methods have been proposed to synthesize dance motions from music for a single dancer. However, generating dance motion for a group remains an open problem. In this paper, we present AIO…

2023

Open-Vocabulary Affordance Detection in 3D Point Clouds

IROS 2023poster

Affordance detection is a challenging problem with a wide variety of robotic applications. Traditional affordance detection methods are limited to a predefined set of affordance labels, hence potentially restricting the adaptability of intelligent robots in complex and dynamic environments. In this…

Cited by 33SourcecodeScholar
2023

Reducing Training Time in Cross-Silo Federated Learning Using Multigraph Topology

ICCV 2023poster

Federated learning is an active research topic since it enables several participants to jointly train a model without sharing local data. Currently, cross-silo federated learning is a popular training setting that utilizes a few hundred reliable data silos with high-speed access links to training a…

Cited by 3PDFcodeScholar
2022

DeepFace-EMD: Re-Ranking Using Patch-Wise Earth Mover's Distance Improves Out-of-Distribution Face Identification

CVPR 2022poster

Face identification (FI) is ubiquitous and drives many high-stake decisions made by the law enforcement. State-of-the-art FI approaches compare two images by taking the cosine similarity between their image embeddings. Yet, such approach suffers from poor out-of-distribution (OOD) generalization to…

Cited by 33PDFcodeScholar
2022

Long-Tailed Instance Segmentation Using Gumbel Optimized Loss

ECCV 2022poster

"Major advancements have been made in the field of object detection and segmentation recently. However, when it comes to rare categories, the state-of-the-art methods fail to detect them, resulting in a significant performance gap between rare and frequent categories. In this paper, we identify that…

2021

Speech Emotion Recognition Using Semantic Information

ICASSP 2021accepted

Speech emotion recognition is a crucial problem manifesting in a multitude of applications such as human computer interaction and education. Although several advancements have been made in the recent years, especially with the advent of Deep Neural Networks (DNN), most of the studies in the literatu…

Cited by 0SourceScholar
2020

Autonomous Navigation in Complex Environments with Deep Multimodal Fusion Network

IROS 2020poster

Autonomous navigation in complex environments is a crucial task in time-sensitive scenarios such as disaster response or search and rescue. However, complex environments pose significant challenges for autonomous platforms to navigate due to their challenging properties: constrained narrow passages,…

Cited by 52SourceScholar
2020

Collaborative Robot-Assisted Endovascular Catheterization with Generative Adversarial Imitation Learning

ICRA 2020poster

Master-slave systems for endovascular catheterization have brought major clinical benefits including reduced radiation doses to the operators, improved precision and stability of the instruments, as well as reduced procedural duration. Emerging deep reinforcement learning (RL) technologies could pot…

Cited by 116SourceScholar
2020

End-to-End Real-time Catheter Segmentation with Optical Flow-Guided Warping during Endovascular Intervention

ICRA 2020poster

Accurate real-time catheter segmentation is an important pre-requisite for robot-assisted endovascular intervention. Most of the existing learning-based methods for catheter segmentation and tracking are only trained on smallscale datasets or synthetic data due to the difficulties of ground-truth an…

Cited by 38SourceScholar
2019

Strike (With) a Pose: Neural Networks Are Easily Fooled by Strange Poses of Familiar Objects

CVPR 2019poster

Despite excellent performance on stationary test sets, deep neural networks (DNNs) can fail to generalize to out-of-distribution (OoD) inputs, including natural, non-adversarial ones, which are common in real-world settings. In this paper, we present a framework for discovering DNN failures that har…

Cited by 392PDFcodeScholar
2018

AffordanceNet: An End-to-End Deep Learning Approach for Object Affordance Detection

ICRA 2018poster

We propose AffordanceNet, a new deep learning approach to simultaneously detect multiple objects and their affordances from RGB images. Our AffordanceNet has two branches: an object detection branch to localize and classify the object, and an affordance detection branch to assign each pixel in the o…

Cited by 366SourcecodeScholar
2018

Translating Videos to Commands for Robotic Manipulation with Deep Recurrent Neural Networks

ICRA 2018poster

We present a new method to translate videos to commands for robotic manipulation using Deep Recurrent Neural Networks (RNN). Our framework first extracts deep features from the input video frames with a deep Convolutional Neural Networks (CNN). Two RNN layers with an encoder-decoder architecture are…

Cited by 84SourceScholar
2017

Object-based affordances detection with Convolutional Neural Networks and dense Conditional Random Fields

IROS 2017poster

We present a new method to detect object affordances in real-world scenes using deep Convolutional Neural Networks (CNN), an object detector and dense Conditional Random Fields (CRF). Our system first trains an object detector to generate bounding box candidates from the images. A deep CNN is then u…

Cited by 209SourceScholar
2017

Plug & Play Generative Networks: Conditional Iterative Generation of Images in Latent Space

CVPR 2017spotlight

Generating high-resolution, photo-realistic images has been a long-standing goal in machine learning. Recently, Nguyen et al. 2016 showed one interesting way to synthesize novel images by performing gradient descent in the latent space of a generator network to maximize the activations of one or mul…

Cited by 1035PDFScholar
2016

Detecting object affordances with Convolutional Neural Networks

IROS 2016poster

We present a novel and real-time method to detect object affordances from RGB-D images. Our method trains a deep Convolutional Neural Network (CNN) to learn deep features from the input data in an end-to-end manner. The CNN has an encoder-decoder architecture in order to obtain smooth label predicti…

Cited by 229SourceScholar
2016

Preparatory object reorientation for task-oriented grasping

IROS 2016poster

This paper describes a new task-oriented grasping method to reorient a rigid object to its nominal pose, which is defined as the configuration that it needs to be grasped from, in order to successfully execute a particular manipulation task. Our method combines two key insights: (1) a visual 6 Degre…

Cited by 27SourceScholar
2016

Synthesizing the preferred inputs for neurons in neural networks via deep generator networks

NeurIPS 2016poster

Deep neural networks (DNNs) have demonstrated state-of-the-art results on many pattern recognition tasks, especially vision classification problems. Understanding the inner workings of such computational brains is both fascinating basic science that is interesting in its own right---similar to why w…

Cited by 907SourcePDFScholar
2015

Deep Neural Networks Are Easily Fooled: High Confidence Predictions for Unrecognizable Images

CVPR 2015poster

Deep neural networks (DNNs) have recently been achieving state-of-the-art performance on a variety of pattern-recognition tasks, most notably visual classification problems. Given that DNNs are now able to classify objects in images with near-human-level performance, questions naturally arise as to…