← Search

Jiaxuan Zhao

6 accepted papers

2026

Compositional Attribute Imbalance in Vision Datasets

AAAI 2026technical

Visual attribute imbalance is a common yet underexplored issue in image classification, significantly impacting model performance and generalization. In this work, we first define the first-level and second-level attributes of images and then introduce a CLIP-based framework to construct a visual at

Cited by 0SourcePDFScholar
2026

See First, Reason Later: Mutual Information-Guided Reinforcement Learning for Vision-Language Models

ICML 2026poster

Vision-Language Models (VLMs) frequently suffer from visual perception errors and hallucinations that compromise answer accuracy in complex reasoning tasks. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising solution by optimizing policies using answer correctness signals. Desp…

Cited by 0SourceScholar
2026

Task-free Adaptive Meta Black-box Optimization

ICLR 2026oral

Handcrafted optimizers become prohibitively inefficient for complex black-box optimization (BBO) tasks. MetaBBO addresses this challenge by meta-learning to automatically configure optimizers for low-level BBO tasks, thereby eliminating heuristic dependencies. However, existing methods typically req…

Cited by 0SourceScholar
2025

CBP-Tuning: Efficient Local Customization for Black-box Large Language Models

EMNLP 2025

The high costs of customizing large language models (LLMs) fundamentally limit their adaptability to user-specific needs. Consequently, LLMs are increasingly offered as cloud-based services, a paradigm that introduces critical limitations: providers struggle to support personalized customization at

2025

DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts

ACL 2025long

Large language models (LLMs) with the Mixture-of-Experts (MoE) architecture achieve high cost-efficiency by selectively activating a subset of the parameters. Despite the inference efficiency of MoE LLMs, the training of extensive experts from scratch incurs substantial overhead, whereas reconstruct…

2021

Multi-Scale Progressive Attention Network for Video Question Answering

ACL 2021short

Understanding the multi-scale visual information in a video is essential for Video Question Answering (VideoQA). Therefore, we propose a novel Multi-Scale Progressive Attention Network (MSPAN) to achieve relational reasoning between cross-scale video information. We construct clips of different leng…

Cited by 23SourcePDFScholar