← Search

Sikai Bai

5 accepted papers

2026

What You Think is What You See: Driving Exploration in VLM Agents via Visual-Linguistic Curiosity

ICML 2026spotlight

To navigate partially observable visual environments, recent VLM agents increasingly internalize world modeling capabilities directly into their policies via explicit CoT reasoning with reinforcement learning (RL). However, mere passive exploitation of reasoning on visited states is insufficient for…

Cited by 0SourceScholar
2025

Causally Motivated Sycophancy Mitigation for Large Language Models

ICLR 2025poster

Incorporating user preferences into large language models (LLMs) can enhance the personalization and reliability of model outputs and facilitate the application of LLMs to real-world scenarios. However, leveraging user preferences can be a double-edged sword. Recent studies have found that improper…

Cited by 0SourcePDFScholar
2025

DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert Pruning

NeurIPS 2025poster

Despite the significant breakthrough of Mixture-of-Experts (MoE), the increasing scale of these MoE models presents huge memory and storage challenges. Existing MoE pruning methods, which involve reducing parameter size with a uniform sparsity across all layers, often lead to suboptimal outcomes and…

Cited by 0SourceScholar
2024

Combating Data Imbalances in Federated Semi-supervised Learning with Dual Regulators

AAAI 2024technical

Federated learning has become a popular method to learn from decentralized heterogeneous data. Federated semi-supervised learning (FSSL) emerges to train models from a small fraction of labeled data due to label scarcity on decentralized clients. Existing FSSL methods assume independent and identica…

Cited by 8SourcePDFScholar
2024

DiPrompT: Disentangled Prompt Tuning for Multiple Latent Domain Generalization in Federated Learning

CVPR 2024poster

Federated learning (FL) has emerged as a powerful paradigm for learning from decentralized data and federated domain generalization further considers the test dataset (target domain) is absent from the decentralized training data (source domains). However most existing FL methods assume that domain…

Cited by 19SourcePDFScholar