← Search

Minh Tran

11 accepted papers

2025

A Domain Adaptation Framework for Speech Recognition Systems with Only Synthetic data

ICASSP 2025accepted

We introduce DAS (Domain Adaptation with Synthetic data), a novel domain adaptation framework for pre-trained ASR model, designed to efficiently adapt to various language-defined domains without requiring any real data. In particular, DAS first prompts large language models (LLMs) to generate domain…

Cited by 0SourceScholar
2025

CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling

ICCV 2025poster

Understanding radiologists' eye movement during Computed Tomography (CT) reading is crucial for developing effective interpretable computer-aided diagnosis systems. However, CT research in this area has been limited by the lack of publicly available eye-tracking datasets and the three-dimensional co…

2025

DiTaiListener: Controllable High Fidelity Listener Video Generation with Diffusion

ICCV 2025poster

Generating naturalistic and nuanced listener motions for extended interactions remains an open problem. Existing methods often rely on low-dimensional motion codes for facial behavior generation followed by photorealistic rendering, limiting both visual fidelity and expressive richness. To address t…

2025

Head2Body: Body Pose Generation from Multi-sensory Head-mounted Inputs

ICCV 2025poster

Generating body pose from head-mounted, egocentric inputs is essential for immersive VR/AR and assistive technologies, as it supports more natural interactions. However, the task is challenging due to limited visibility of body parts in first-person views and the sparseness of sensory data, with onl…

Cited by 0SourcePDFScholar
2024

HENASY: Learning to Assemble Scene-Entities for Interpretable Egocentric Video-Language Model

NeurIPS 2024poster

Current video-language models (VLMs) rely extensively on instance-level alignment between video and language modalities, which presents two major limitations: (1) visual reasoning disobeys the natural perception that humans do in first-person perspective, leading to a lack of reasoning interpretatio…

2024

Open-Fusion: Real-time Open-Vocabulary 3D Mapping and Queryable Scene Representation

ICRA 2024poster

Precise 3D environmental mapping with semantics is essential in robotics. Existing methods often rely on pre-defined concepts during training or are time-intensive when generating semantic maps. This paper presents Open-Fusion, an approach for real-time open-vocabulary 3D mapping and queryable scene…

Cited by 31SourcecodeScholar
2020

Towards A Friendly Online Community: An Unsupervised Style Transfer Framework for Profanity Redaction

COLING 2020main

Offensive and abusive language is a pressing problem on social media platforms. In this work, we propose a method for transforming offensive comments, statements containing profanity or offensive language, into non-offensive ones. We design a Retrieve, Generate and Edit unsupervised style transfer p…

Cited by 35SourcePDFScholar
2019

A Lightweight, Efficient Fully Powered Knee Prosthesis With Actively Variable Transmission

RA-L 2019

Amputation at the above-knee level severely impairs the ability of an individual to ambulate. As ambulation requires power generation and active control of movements, the passive nature of most available leg prostheses is a major cause of the observed deficits. Powered prostheses aim to address this

Cited by 77SourceScholar
2017

NREL-Exo: A 4-DoFs wearable hip exoskeleton for walking and balance assistance in locomotion

IROS 2017poster

In this paper, we presented a high-power, self-balancing, passively and software-controlled active compliant, and wearable hip exoskeleton to provide walking and balance assistance. The device features powered hip abduction/adduction (HAA) and hip flexion/extension (HFE) modules to provide assistanc…

Cited by 36SourceScholar