← Search

Yansong Li

5 accepted papers

2026

LensWalk: Agentic Video Understanding by Planning How You See in Videos

CVPR 2026

The dense, temporal nature of video presents a profound challenge for automated analysis. Despite the use of powerful Vision-Language Models, prevailing methods for video understanding are limited by the inherent disconnect between reasoning and perception: they rely on static, pre-processed informa

Cited by 0SourceScholar
2025

A Context-Aware Contrastive Learning Framework for Hateful Meme Detection and Segmentation

NAACL 2025findings

Amidst the rise of Large Multimodal Models (LMMs) and their widespread application in generating and interpreting complex content, the risk of propagating biased and harmful memes remains significant. Current safety measures often fail to detect subtly integrated hateful content within “Confounder M…

Cited by 0SourcePDFScholar
2024

Generalizing End-To-End Autonomous Driving In Real-World Environments Using Zero-Shot LLMs

CoRL 2024poster

Traditional autonomous driving methods adopt modular design, decomposing tasks into sub-tasks, including perception, prediction, planning, and control. In contrast, end-to-end autonomous driving directly outputs actions from raw sensor data, avoiding error accumulation. However, training an end-to-e…

Cited by 5SourceScholar