← Search

Zhong Yang

4 accepted papers

2026

SGP4SR: Separated-Modality Guided User Preference Learning for Multimodal Sequential Recommendation

AAAI 2026technical

With the booming development of multimodal data (e.g., image, text) on internet platforms, multimodal sequential recommendation methods continue to emerge. Most existing methods incorporate item modal features as auxiliary information, typically concatenating them to learn unified user representatio

Cited by 0SourcePDFScholar
2026

Seeing What Matters: A Training-Free Self-Guided Framework for Multimodal Detail Perception and Reasoning

CVPR 2026

Multimodal large language models (MLLMs) have achieved remarkable success on diverse visual-language tasks. However, fixed-resolution models face challenges in perceiving fine-grained visual details, particularly due to *distracted attention* and *blurry vision*. To address these issues, we propose

Cited by 0SourceScholar
2025

EventLens: Enhancing Visual Commonsense Reasoning by Leveraging Event-Aware Pretraining and Cross-modal Linking

ICASSP 2025accepted

Visual Commonsense Reasoning (VCR) is a cognitive task, challenging models to answer visual questions, and to explain the rationale behind their answers. While Large Language Models (LLMs) offer potential for this task, VCR’s complex scenes require specialized approaches to activate their commonsens…

Cited by 0SourceScholar
2019

ChevBot – An Untethered Microrobot Powered by Laser for Microfactory Applications

ICRA 2019poster

In this paper, we introduce a new class of submillimeter robot (ChevBot) for microfactory applications in dry environments, powered by a 532 nm laser beam. ChevBot is an untethered microrobot propelled by a thermal Micro Electro Mechanical (MEMS) actuator upon exposure to the laser light. Novel mode…

Cited by 24SourceScholar