← Search

Lan Wang

12 accepted papers

2026

Causally-Grounded Dual-Path Attention Intervention for Object Hallucination Mitigation in LVLMs

AAAI 2026technical

Object hallucination remains a critical challenge in Large Vision-Language Models (LVLMs), where models generate content inconsistent with visual inputs. Existing language-decoder based mitigation approaches often regulate visual or textual attention independently, overlooking their interaction as t

Cited by 0SourcePDFScholar
2025

Multimodal Transformers are Hierarchical Modal-wise Heterogeneous Graphs

ACL 2025long

Multimodal Sentiment Analysis (MSA) is a rapidly developing field that integrates multimodal information to recognize sentiments, and existing models have made significant progress in this area. The central challenge in MSA is multimodal fusion, which is predominantly addressed by Multimodal Transfo…

2025

SEAL: Semantic Attention Learning for Long Video Representation

CVPR 2025poster

Long video understanding presents challenges due to the inherent high computational complexity and redundant temporal information. An effective representation for long videos must efficiently process such redundancy while preserving essential contents for downstream tasks. This paper introduces **S…

Cited by 0SourcePDFScholar
2024

An Audio-Textual Diffusion Model for Converting Speech Signals into Ultrasound Tongue Imaging Data

ICASSP 2024accepted

Acoustic-to-articulatory inversion (AAI) is to convert audio into articulator movements, such as ultrasound tongue imaging (UTI) data. An issue of existing AAI methods is only using the personalized acoustic information to derive the general patterns of tongue motions, and thus the quality of genera…

Cited by 0SourceScholar
2024

FairerCLIP: Debiasing CLIP's Zero-Shot Predictions using Functions in RKHSs

ICLR 2024poster

Large pre-trained vision-language models such as CLIP provide compact and general-purpose representations of text and images that are demonstrably effective across multiple downstream zero-shot prediction tasks. However, owing to the nature of their training process, these models have the potential…

Cited by 15SourcePDFScholar
2023

ProTeGe: Untrimmed Pretraining for Video Temporal Grounding by Video Temporal Grounding

CVPR 2023poster

Video temporal grounding (VTG) is the task of localizing a given natural language text query in an arbitrarily long untrimmed video. While the task involves untrimmed videos, all existing VTG methods leverage features from video backbones pretrained on trimmed videos. This is largely due to the lack…

Cited by 15SourcePDFScholar
2019

BLHUC: Bayesian Learning of Hidden Unit Contributions for Deep Neural Network Speaker Adaptation

ICASSP 2019accepted

Speaker adaptation techniques play a key role in reducing the mismatch between speech recognition systems and target users. In order to robustly learn speaker-dependent adaptation parameters, model based DNN adaptation techniques often require a significant amount of data. For example, in the common…

Cited by 0SourceScholar
2018

PM-GANs: Discriminative Representation Learning for Action Recognition Using Partial-modalities

ECCV 2018poster

Data of different modalities generally convey complimentary but heterogeneous information, and a more discriminative representation is often preferred by combining multiple data modalities like the RGB and infrared features. However in reality, obtaining both data channels is challenging due to many…

Cited by 32SourcePDFScholar
2016

Intelligible enhancement of 3D articulation animation by incorporating airflow information

ICASSP 2016accepted

The 3D talking head has been developed fast, in which both external and internal articulators were demonstrated. For Mandarin pronunciation, the aspiration airflow is crucial to discriminate confusable Mandarin consonants. In this paper, we present a 3D talking head system for articulatory and aspir…

Cited by 0SourceScholar