← Search

Khanh Nguyen

7 accepted papers

2026

Catching the First Light of Tomorrow: A Hackathon-Based Framework for Introducing High School Students to AI Agents

AAAI 2026technical

Artificial Intelligence (AI), particularly in the form of intelligent AI agents, is transforming education, industry, and everyday life. These agents extend the capabilities of Large Language Models (LLMs) by integrating planning, decision-making, tool use, and multi-agent collaboration, enabling sy

Cited by 0SourcePDFScholar
2026

RI-Mamba: Rotation-Invariant Mamba for Robust Text-to-Shape Retrieval

CVPR 2026

3D assets have rapidly expanded in quantity and diversity due to the growing popularity of virtual reality and gaming. As a result, text-to-shape retrieval has become essential in facilitating intuitive search within large repositories. However, existing methods require canonical poses and support f

Cited by 0SourcecodeScholar
2025

DocMIA: Document-Level Membership Inference Attacks against DocVQA Models

ICLR 2025poster

Document Visual Question Answering (DocVQA) has introduced a new paradigm for end-to-end document understanding, and quickly became one of the standard benchmarks for multimodal LLMs. Automating document processing workflows, driven by DocVQA models, presents significant potential for many business…

2025

Occlusion-aware Text-Image-Point Cloud Pretraining for Open-World 3D Object Recognition

CVPR 2025poster

Recent open-world representation learning approaches have leveraged CLIP to enable zero-shot 3D object recognition. However, performance on real point clouds with occlusions still falls short due to unrealistic pretraining settings. Additionally, these methods incur high inference costs because they…

2023

Define, Evaluate, and Improve Task-Oriented Cognitive Capabilities for Instruction Generation Models

ACL 2023findings

Recent work studies the cognitive capabilities of language models through psychological tests designed for humans. While these studies are helpful for understanding the general capabilities of these models, there is no guarantee that a model possessing sufficient capabilities to pass those tests wou…

2023

Show, Interpret and Tell: Entity-Aware Contextualised Image Captioning in Wikipedia

AAAI 2023technical

Humans exploit prior knowledge to describe images, and are able to adapt their explanation to specific contextual information given, even to the extent of inventing plausible explanations when contextual information and images do not match. In this work, we propose the novel task of captioning Wikip…

2019

Vision-Based Navigation With Language-Based Assistance via Imitation Learning With Indirect Intervention

CVPR 2019poster

We present Vision-based Navigation with Language-based Assistance (VNLA), a grounded vision-language task where an agent with visual perception is guided via language to find objects in photorealistic indoor environments. The task emulates a real-world scenario in that (a) the requester may not know…

Cited by 144PDFcodeScholar