← Search

Minchen Yu

2 accepted papers

2026

Think in Cloud, Look at Edges: Semantic-Driven Query Decomposition for Efficient Video Reasoning

ICML 2026spotlight

Long video understanding faces a critical dilemma: cloud-based Large Multimodal Models (LMMs) offer superior reasoning but suffer from prohibitive bandwidth costs and latency, while edge-based solutions sacrifice perception accuracy for speed. Current collaborative approaches attempt to bridge this …

Cited by 0SourceScholar
2026

UniScale: Adaptive Unified Inference Scaling via Online Joint Optimization of Model Routing and Test-Time Scaling

ICML 2026poster

In real-world deployments of large language models (LLMs), balancing inference quality and computational cost has become a central challenge. Existing approaches tackle this trade-off along two largely independent dimensions: model routing, which switches among models of different scales to match re…

Cited by 0SourceScholar