← Search

Yimu Wang

15 accepted papers

2025

DREAM: Improving Video-Text Retrieval Through Relevance-Based Augmentation Using Large Foundation Models

NAACL 2025long

Recent progress in video-text retrieval has been driven largely by advancements in model architectures and training strategies. However, the representation learning capabilities of video-text retrieval models remain constrained by low-quality and limited training data annotations. To address this is…

Cited by 0SourcePDFScholar
2025

Hawaii: Hierarchical Visual Knowledge Transfer for Efficient Vision-Language Models

NeurIPS 2025poster

Improving the visual understanding ability of vision-language models (VLMs) is crucial for enhancing their performance across various tasks. While using multiple pretrained visual experts has shown great promise, it often incurs significant computational costs during training and inference. To addre…

Cited by 0SourceScholar
2025

LEO-MINI: An Efficient Multimodal Large Language Model using Conditional Token Reduction and Mixture of Multi-Modal Experts

EMNLP 2025

Redundancy of visual tokens in multi-modal large language models (MLLMs) significantly reduces their computational efficiency. Recent approaches, such as resamplers and summarizers, have sought to reduce the number of visual tokens, but at the cost of visual reasoning ability. To address this, we pr

Cited by 0SourcePDFScholar
2025

NBDESCRIB: A Dataset for Text Description Generation from Tables and Code in Jupyter Notebooks with Guidelines

ACL 2025finding

Generating cell-level descriptions for Jupyter Notebooks, which is a major resource consisting of codes, tables, and descriptions, has been attracting increasing research attention. However, existing methods for Jupyter Notebooks mostly focus on generating descriptions from code snippets or table ou…

Cited by 0SourcePDFScholar
2025

OV-SCAN: Semantically Consistent Alignment for Novel Object Discovery in Open-Vocabulary 3D Object Detection

ICCV 2025poster

Open-vocabulary 3D object detection for autonomous driving aims to detect novel objects beyond the predefined training label sets in point cloud scenes. Existing approaches achieve this by connecting traditional 3D object detectors with vision-language models (VLMs) to regress 3D bounding boxes for…

2024

Lost Domain Generalization Is a Natural Consequence of Lack of Training Domains

AAAI 2024technical

We show a hardness result for the number of training domains required to achieve a small population error in the test domain. Although many domain generalization algorithms have been developed under various domain-invariance assumptions, there is significant evidence to indicate that out-of-distribu…

Cited by 3SourcePDFScholar
2024

Multiobjective Lipschitz Bandits under Lexicographic Ordering

AAAI 2024technical

This paper studies the multiobjective bandit problem under lexicographic ordering, wherein the learner aims to simultaneously maximize ? objectives hierarchically. The only existing algorithm for this problem considers the multi-armed bandit model, and its regret bound is O((KT)^(2/3)) under a metri…

Cited by 3SourcePDFScholar
2023

Balance Act: Mitigating Hubness in Cross-Modal Retrieval with Query and Gallery Banks

EMNLP 2023long main

In this work, we present a post-processing solution to address the hubness problem in cross-modal retrieval, a phenomenon where a small number of gallery data points are frequently retrieved, resulting in a decline in retrieval performance. We first theoretically demonstrate the necessity of incorpo…

Cited by 0SourcecodeScholar
2023

Cooperation or Competition: Avoiding Player Domination for Multi-Target Robustness via Adaptive Budgets

CVPR 2023poster

Despite incredible advances, deep learning has been shown to be susceptible to adversarial attacks. Numerous approaches were proposed to train robust networks both empirically and certifiably. However, most of them defend against only a single type of attack, while recent work steps forward at defen…

Cited by 2SourcePDFScholar
2023

Efficient Algorithms for Generalized Linear Bandits with Heavy-tailed Rewards

NeurIPS 2023poster

This paper investigates the problem of generalized linear bandits with heavy-tailed rewards, whose $(1+\epsilon)$-th moment is bounded for some $\epsilon\in (0,1]$. Although there exist methods for generalized linear bandits, most of them focus on bounded or sub-Gaussian rewards and are not well-sui…

Cited by 4SourcePDFScholar
2023

Multimodal Federated Learning via Contrastive Representation Ensemble

ICLR 2023poster

With the increasing amount of multimedia data on modern mobile systems and IoT infrastructures, harnessing these rich multimodal data without breaching user privacy becomes a critical issue. Federated learning (FL) serves as a privacy-conscious alternative to centralized machine learning. However, e…

2021

Deep Unified Cross-Modality Hashing by Pairwise Data Alignment

IJCAI 2021poster

With the increasing amount of multimedia data, cross-modality hashing has made great progress as it achieves sub-linear search time and low memory space. However, due to the huge discrepancy between different modalities, most existing cross-modality hashing methods cannot learn unified hash codes an…

Cited by 21SourcePDFScholar
2020

Nearly Optimal Regret for Stochastic Linear Bandits with Heavy-Tailed Payoffs

IJCAI 2020poster

In this paper, we study the problem of stochastic linear bandits with finite action sets. Most of existing work assume the payoffs are bounded or sub-Gaussian, which may be violated in some scenarios such as financial markets. To settle this issue, we analyze the linear bandits with heavy-tailed pay…

Cited by 0SourcePDFScholar