← Search

Tuan Tran

8 accepted papers

2026

FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation

ICML 2026poster

Vision–Language–Action (VLA) models enable general-purpose robotic control via large-scale multimodal pretraining, yet their effectiveness under few-shot imitation learning remains limited. We conduct a systematic stress test of state-of-the-art VLA models and show that performance degrades sharply …

Cited by 0SourceScholar
2026

How Good is Post-Hoc Watermarking With Language Model Rephrasing?

ICML 2026poster

Generation-time text watermarking embeds statistical signals into text for traceability of AI-generated content. We explore post-hoc watermarking where an LLM rewrites existing text while applying generation-time watermarking, to protect copyrighted documents, or detect their use in training or RAG …

Cited by 0SourceScholar
2026

Learning to Watermark in the Latent Space of Generative Models

ICML 2026poster

Existing approaches for watermarking AI-generated images often rely on post-hoc methods applied in pixel space, introducing computational overhead and potential visual artifacts. In this work, we explore latent space watermarking and introduce DistSeal, a unified approach for latent watermarking tha…

Cited by 0SourceScholar
2024

Proactive Detection of Voice Cloning with Localized Watermarking

ICML 2024poster

In the rapidly evolving field of speech generative models, there is a pressing need to ensure audio authenticity against the risks of voice cloning. We present AudioSeal, the first audio watermarking technique designed specifically for localized detection of AI-generated speech. AudioSeal employs a…

2023

IPCC-TP: Utilizing Incremental Pearson Correlation Coefficient for Joint Multi-Agent Trajectory Prediction

CVPR 2023poster

Reliable multi-agent trajectory prediction is crucial for the safe planning and control of autonomous systems. Compared with single-agent cases, the major challenge in simultaneously processing multiple agents lies in modeling complex social interactions caused by various driving intentions and road…

Cited by 19SourcePDFScholar
2022

COSM2IC: Optimizing Real-Time Multi-Modal Instruction Comprehension

RA-L 2022

Supporting real-time, on-device execution of <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">multi-modal referring instruction comprehension</i> models is an important challenge to be tackled in embodied Human-Robot Interaction. However, state-of-the

Cited by 11SourceScholar
2021

Feasible and Adaptive Multimodal Trajectory Prediction with Semantic Maneuver Fusion

ICRA 2021poster

Predicting trajectories of participating vehicles is a crucial task towards full and safe autonomous driving. General unconstrained machine learning methods often report unrealistic predictions, and need to be combined with different motion constraints. Existing work either defines some shallow mane…

Cited by 12SourceScholar