← Search

Khoa Vo

9 accepted papers

2026

Domain Expansion: A Latent Space Construction Framework for Multi-Task Learning

ICLR 2026poster

Training a single network with multiple objectives often leads to conflicting gradients that degrade shared representations, forcing them into a compromised state that is suboptimal for any single task—a problem we term latent representation collapse. We introduce Domain Expansion, a framework that…

Cited by 0SourceScholar
2026

Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective

AAAI 2026technical

As embodied agents operate in increasingly complex environments, the ability to perceive, track, and reason about individual object instances over time becomes essential, especially in tasks requiring sequenced interactions with visually similar objects. In non-Markovian settings, critical decision

Cited by 0SourcePDFScholar
2026

SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection

CVPR 2026

Camera-only 3D object detection has emerged as a cost-effective and scalable alternative to LiDAR for autonomous driving, yet existing methods primarily prioritize overall performance while overlooking the severe long-tail imbalance inherent in real-world datasets. In practice, many rare but safety-

Cited by 0SourceScholar
2025

CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling

ICCV 2025poster

Understanding radiologists' eye movement during Computed Tomography (CT) reading is crucial for developing effective interpretable computer-aided diagnosis systems. However, CT research in this area has been limited by the lack of publicly available eye-tracking datasets and the three-dimensional co…

2024

HENASY: Learning to Assemble Scene-Entities for Interpretable Egocentric Video-Language Model

NeurIPS 2024poster

Current video-language models (VLMs) rely extensively on instance-level alignment between video and language modalities, which presents two major limitations: (1) visual reasoning disobeys the natural perception that humans do in first-person perspective, leading to a lack of reasoning interpretatio…

2024

Open-Fusion: Real-time Open-Vocabulary 3D Mapping and Queryable Scene Representation

ICRA 2024poster

Precise 3D environmental mapping with semantics is essential in robotics. Existing methods often rely on pre-defined concepts during training or are time-intensive when generating semantic maps. This paper presents Open-Fusion, an approach for real-time open-vocabulary 3D mapping and queryable scene…

Cited by 31SourcecodeScholar
2023

VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph Captioning

AAAI 2023technical

Video Paragraph Captioning aims to generate a multi-sentence description of an untrimmed video with multiple temporal event locations in a coherent storytelling. Following the human perception process, where the scene is effectively understood by decomposing it into visual (e.g. human, animal) and…

2021

Offboard 3D Object Detection From Point Cloud Sequences

CVPR 2021poster

While current 3D object recognition research mostly focuses on the real-time, onboard scenario, there are many offboard use cases of perception that are largely under-explored, such as using machines to automatically generate high-quality 3D labels. Existing 3D object detectors fail to satisfy the h…

Cited by 226PDFScholar