← Search

Ngan Hoang Le

4 accepted papers

2025

Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models

AAAI 2025technical

Existing Large Vision-Language Models (LVLMs) excel at matching concepts across multi-modal inputs but struggle with compositional concepts and high-level relationships between entities. This paper introduces Progressive multi-granular Vision-Language alignments (PromViL), a novel framework to enhan…

2024

Accelerating Transformers with Spectrum-Preserving Token Merging

NeurIPS 2024poster

Increasing the throughput of the Transformer architecture, a foundational component used in numerous state-of-the-art models for vision and language tasks (e.g., GPT, LLaVa), is an important problem in machine learning. One recent and effective strategy is to merge token representations within Trans…

2024

DINTR: Tracking via Diffusion-based Interpolation

NeurIPS 2024poster

Object tracking is a fundamental task in computer vision, requiring the localization of objects of interest across video frames. Diffusion models have shown remarkable capabilities in visual generation, making them well-suited for addressing several requirements of the tracking problem. This work pr…

Cited by 0SourcePDFScholar
2024

HENASY: Learning to Assemble Scene-Entities for Interpretable Egocentric Video-Language Model

NeurIPS 2024poster

Current video-language models (VLMs) rely extensively on instance-level alignment between video and language modalities, which presents two major limitations: (1) visual reasoning disobeys the natural perception that humans do in first-person perspective, leading to a lack of reasoning interpretatio…