← Search

Yiman Hu

4 accepted papers

2026

E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs

ICML 2026poster

E-commerce short videos represent a high-revenue segment of the online video industry characterized by a goal-driven format and dense multi-modal signals. Current models often struggle with these videos because existing benchmarks focus primarily on general-purpose tasks and neglect the reasoning of…

Cited by 1SourceScholar
2026

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling

ICML 2026poster

Multimodal Large Language Models (MLLMs) have made substantial progress on visual understanding tasks, yet they still perform poorly on high-resolution images. Prior work often attributes this limitation to perceptual constraints, arguing that MLLMs fail to recognize small objects and therefore rely…

Cited by 0SourceScholar
2024

Flatten Long-Range Loss Landscapes for Cross-Domain Few-Shot Learning

CVPR 2024poster

Cross-domain few-shot learning (CDFSL) aims to acquire knowledge from limited training data in the target domain by leveraging prior knowledge transferred from source domains with abundant training samples. CDFSL faces challenges in transferring knowledge across dissimilar domains and fine-tuning mo…

2024

Generate Universal Adversarial Perturbations for Few-Shot Learning

NeurIPS 2024poster

Deep networks are known to be vulnerable to adversarial examples which are deliberately designed to mislead the trained model by introducing imperceptible perturbations to input samples. Compared to traditional perturbations crafted specifically for each data point, Universal Adversarial Perturbatio…

Cited by 0SourcePDFScholar