← Search

Thao Minh Le

7 accepted papers

2026

Rethinking Deep Alignment Through the Lens of Incomplete Safety Learning

AAAI 2026technical

Large language models exhibit systematic vulnerabilities to adversarial attacks despite extensive safety alignment through supervised fine-tuning and reinforcement learning from human feedback. These vulnerabilities manifest as differential safety behavior across token positions, with safety modific

Cited by 0SourcePDFScholar
2026

Universal Multi-Domain Translation via Diffusion Routers

ICLR 2026poster

Multi-domain translation (MDT) aims to learn translations between multiple domains, yet existing approaches either require fully aligned tuples or can only handle domain pairs seen in training, limiting their practicality and excluding many cross-domain mappings. We introduce universal MDT (UMDT), a…

Cited by 0SourcecodeScholar
2025

Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models

AAAI 2025technical

Existing Large Vision-Language Models (LVLMs) excel at matching concepts across multi-modal inputs but struggle with compositional concepts and high-level relationships between entities. This paper introduces Progressive multi-granular Vision-Language alignments (PromViL), a novel framework to enhan…

2022

Video Dialog As Conversation about Objects Living in Space-Time

ECCV 2022poster

"It would be a technological feat to be able to create a system that can hold a meaningful conversation with humans about what they watch. A setup toward that goal is presented as a video dialog task, where the system is asked to generate natural utterances in response to a question in an ongoing di…

2021

Hierarchical Object-oriented Spatio-Temporal Reasoning for Video Question Answering

IJCAI 2021poster

Video Question Answering (Video QA) is a powerful testbed to develop new AI capabilities. This task necessitates learning to reason about objects, relations, and events across visual and linguistic domains in space-time. High-level reasoning demands lifting from associative visual pattern recognitio…

2020

Hierarchical Conditional Relation Networks for Video Question Answering

CVPR 2020oral

Video question answering (VideoQA) is challenging as it requires modeling capacity to distill dynamic visual artifacts and distant relations and to associate them with linguistic concepts. We introduce a general-purpose reusable neural unit called Conditional Relation Network (CRN) that serves as a…

Cited by 334PDFcodeScholar