← Search

Bing Hu

9 accepted papers

2026

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model

ICML 2026oral

Vision-Language-Action (VLA) models often suffer from performance degradation under distribution shifts, as they struggle to learn generalized behavior representations across varying environments. While existing approaches attempt to construct behavior representations through action-centric latent v…

Cited by 0SourceScholar
2026

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation

CVPR 2026

Hierarchical Vision-Language-Action (VLA) models have rapidly become a dominant paradigm for robotic manipulation. It typically comprising a Vision-Language backbone for perception and understanding, together with a generative policy for action generation. However, its performance is increasingly bo

Cited by 0SourcecodeScholar
2026

Multiplayer Nash Preference Optimization

ICLR 2026oral

Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models (LLMs) with human preferences. However, reward-based methods built on the Bradley–Terry assumption struggle to capture the non-transitive and heterogeneous nature of real-world p…

Cited by 0SourcecodeScholar
2026

PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding

ICML 2026poster

Speculative decoding can significantly accelerate LLM inference, especially given that its cloud-edge collaborative deployment offers cloud workload offloading, offline robustness, and privacy enhancement. However, existing collaborative inference frameworks with speculative decoding are constrained…

Cited by 0SourceScholar
2026

STRUCTURE-TO-IMAGE: ZERO-SHOT DEPTH ESTIMATION IN COLONOSCOPY VIA HIGH-FIDELITY SIM-TO-REAL ADAPTATION

ICASSP 2026poster

Monocular depth estimation (MDE) for colonoscopy is hampered by the domain gap between simulated and real-world images. Existing image-to-image translation methods, which use depth as a posterior constraint, often produce structural distortions and specular highlights by failing to balance realism w…

Cited by 0SourcePDFScholar
2026

Seizure-Semiology-Suite($S^3$): A Clinically Multimodal Dataset, Benchmark, and Models for Seizure Semiology Understanding

ICML 2026spotlight

While Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in general video understanding, their capacity to interpret involuntary, and spatio-temporally evolving pathologic motor behaviors such as seizure semiology remains largely untested. To address this gap, we intro…

Cited by 0SourceScholar
2025

FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training

NeurIPS 2025poster

The parameter size of modern large language models (LLMs) can be scaled up to the trillion-level via the sparsely-activated Mixture-of-Experts (MoE) technique to avoid excessive increase of the computational costs. To further improve training efficiency, pipelining computation and communication has…

Cited by 0SourceScholar
2025

LION-FS: Fast & Slow Video-Language Thinker as Online Video Assistant

CVPR 2025poster

First-person video assistants are highly anticipated to enhance our daily life through online video dialogue. However, existing online video assistants often sacrifice assistant efficacy for real-time efficiency by processing low-frame-rate videos with coarse-grained visual features. To overcome the…

2024

Deep Variational Incomplete Multi-View Clustering: Exploring Shared Clustering Structures

AAAI 2024technical

Incomplete multi-view clustering (IMVC) aims to reveal shared clustering structures within multi-view data, where only partial views of the samples are available. Existing IMVC methods primarily suffer from two issues: 1) Imputation-based methods inevitably introduce inaccurate imputations, which in…

Cited by 16SourcePDFScholar