← Search

Bo Fang

7 accepted papers

2026

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders

ICLR 2026poster

Employing Multimodal Large Language Models (MLLMs) for long video understanding remains a challenging problem due to the dilemma between the substantial number of video frames (i.e., visual tokens) versus the limited context length of language models. Traditional uniform sampling often leads to sele…

Cited by 0SourcecodeScholar
2025

Can Large Language Models Understand Intermediate Representations in Compilers?

ICML 2025poster

Intermediate Representations (IRs) play a critical role in compiler design and program analysis, yet their comprehension by *Large Language Models* (LLMs) remains underexplored. In this paper, we present an explorative empirical study evaluating the capabilities of six state-of-the-art LLMs—GPT-4,…

2025

DistinctAD: Distinctive Audio Description Generation in Contexts

CVPR 2025highlight

Audio Descriptions (ADs) aim to provide a narration of a movie in text form, describing non-dialogue-related narratives, such as characters, actions, or scene establishment. Automatic generation of ADs remains challenging due to: i) the domain gap between movie-AD data and existing data used to trai…

Cited by 2SourcePDFScholar
2023

Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval?

CVPR 2023highlight

Most existing text-video retrieval methods focus on cross-modal matching between the visual content of videos and textual query sentences. However, in real-world scenarios, online videos are often accompanied by relevant text information such as titles, tags, and even subtitles, which can be utilize…

2023

LABANet: Lead-Assisting Backbone Attention Network for Oral Multi-Pathology Segmentation

ICASSP 2023accepted

This paper presents a Lead-Assisting Backbone Attention Network (LABANet), which is able to perform multi-pathology instance segmentation of dental panoramic X-rays. A Lead-Assisting Attention Backbone (LAAB), containing two Swin-Transformers, is first developed for feature extraction. The following…

Cited by 0SourceScholar
2023

UATVR: Uncertainty-Adaptive Text-Video Retrieval

ICCV 2023poster

With the explosive growth of web videos and emerging large-scale vision-language pre-training models, e.g., CLIP, retrieving videos of interest with text instructions has attracted increasing attention. A common practice is to transfer text-video pairs to the same embedding space and craft cross-mod…

Cited by 64PDFcodeScholar
2022

Combining Multiple Style Transfer Networks and Transfer Learning For LGE-CMR Segmentation

ICASSP 2022accepted

This paper presents an algorithm for segmenting late gadolinium enhancement cardiac magnetic resonance (LGE-CMR) in the absence of labeled training data. The proposed method includes a data augmentation part and a segmentation network. Multiple style transfer networks are employed for data augmentat…

Cited by 0SourceScholar