← Search

Toan Nguyen

9 accepted papers

2026

Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective

AAAI 2026technical

As embodied agents operate in increasingly complex environments, the ability to perceive, track, and reason about individual object instances over time becomes essential, especially in tasks requiring sequenced interactions with visually similar objects. In non-Markovian settings, critical decision

Cited by 0SourcePDFScholar
2025

h-Edit: Effective and Flexible Diffusion-Based Editing via Doob's h-Transform

CVPR 2025poster

We introduce a theoretical framework for diffusion-based image editing by formulating it as a reverse-time bridge modeling problem. This approach modifies the backward process of a pretrained diffusion model to construct a bridge that converges to an implicit distribution associated with the editing…

2024

Crossing Linguistic Horizons: Finetuning and Comprehensive Evaluation of Vietnamese Large Language Models

NAACL 2024findings

Recent advancements in large language models (LLMs) have underscored their importance in the evolution of artificial intelligence. However, despite extensive pretraining on multilingual datasets, available open-sourced LLMs exhibit limited effectiveness in processing Vietnamese. The challenge is exa…

2024

HabiCrowd: A High Performance Simulator for Crowd-Aware Visual Navigation

IROS 2024poster

Visual navigation, a foundational aspect of Embodied AI (E-AI) and robotics has been extensively studied in the past few years. While many 3D simulators have been introduced for the visual navigation tasks, scarcely works have combined human dynamics, creating the gap between simulation and real-wor…

Cited by 3SourcecodeScholar
2024

Language-Conditioned Affordance-Pose Detection in 3D Point Clouds

ICRA 2024poster

Affordance detection and pose estimation are of great importance in many robotic applications. Their combination helps the robot gain an enhanced manipulation capability, in which the generated pose can facilitate the corresponding affordance task. Previous methods for affodance-pose joint learning…

Cited by 17SourcecodeScholar
2024

Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance

ECCV 2024oral

"6-DoF grasp detection has been a fundamental and challenging problem in robotic vision. While previous works have focused on ensuring grasp stability, they often do not consider human intention conveyed through natural language, hindering effective collaboration between robots and users in complex…

2024

Open-Vocabulary Affordance Detection using Knowledge Distillation and Text-Point Correlation

ICRA 2024poster

Affordance detection presents intricate challenges and has a wide range of robotic applications. Previous works have faced limitations such as the complexities of 3D object shapes, the wide range of potential affordances on real-world objects, and the lack of open-vocabulary support for affordance u…

Cited by 10SourcecodeScholar
2023

Open-Vocabulary Affordance Detection in 3D Point Clouds

IROS 2023poster

Affordance detection is a challenging problem with a wide variety of robotic applications. Traditional affordance detection methods are limited to a predefined set of affordance labels, hence potentially restricting the adaptability of intelligent robots in complex and dynamic environments. In this…

Cited by 33SourcecodeScholar
2021

Hogwild! over Distributed Local Data Sets with Linearly Increasing Mini-Batch Sizes

AISTATS 2021poster

Hogwild! implements asynchronous Stochastic Gradient Descent (SGD) where multiple threads in parallel access a common repository containing training data, perform SGD iterations and update shared state that represents a jointly learned (global) model. We consider big data analysis where training dat…

Cited by 6SourcePDFScholar