← Search

Zihan Song

4 accepted papers

2026

AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation

ICML 2026poster

While Vision-Language-Action (VLA) models have achieved remarkable success in ground-based embodied intelligence, their application to Aerial Manipulation Systems (AMS) remains a largely unexplored frontier. The inherent characteristics of AMS, including floating-base dynamics, strong coupling betwe…

Cited by 0SourceScholar
2025

Modularized Self-Reflected Video Reasoner for Multimodal LLM with Application to Video Question Answering

ICML 2025poster

Multimodal Large Language Models (Multimodal LLMs) have shown their strength in Video Question Answering (VideoQA). However, due to the black-box nature of end-to-end training strategies, existing approaches based on Multimodal LLMs suffer from the lack of interpretability for VideoQA: they can neit…

Cited by 0SourcePDFScholar
2021

Dual Sparse Attention Network For Session-based Recommendation

AAAI 2021technical

Session-based Recommendations recommend the next possible item for the user with anonymous sessions, whose challenge is that the user’s behavioral preference can only be analyzed in a limited sequence to meet their need. Recent advances evaluate the effectiveness of the attention mechanism in the se…