← Search

Shengsheng Qian

8 accepted papers

2026

QueryStream: Advancing Streaming Video Understanding with Query-Aware Pruning and Proactive Response

ICLR 2026poster

The increasing demand for real-time interaction in online video scenarios necessitates a new class of efficient streaming video understanding models. However, existing approaches often rely on a flawed, query-agnostic ``change-is-important'' principle, which conflates visual dynamics with semantic r…

Cited by 0SourceScholar
2026

SoMe: A Realistic Benchmark for LLM-based Social Media Agents

AAAI 2026technical

Intelligent agents powered by large language models (LLMs) have recently demonstrated impressive capabilities and gained increasing popularity on social media platforms. While LLM agents are reshaping the ecology of social media, there exists a current gap in conducting a comprehensive evaluation of

Cited by 0SourcePDFScholar
2025

LiveStar: Live Streaming Assistant for Real-World Online Video Understanding

NeurIPS 2025poster

Despite significant progress in Video Large Language Models (Video-LLMs) for offline video understanding, existing online Video-LLMs typically struggle to simultaneously process continuous frame-by-frame inputs and determine optimal response timing, often compromising real-time responsiveness and na…

Cited by 0SourcecodeScholar
2025

SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding

ICLR 2025spotlight

Despite the significant advancements of Large Vision-Language Models (LVLMs) on established benchmarks, there remains a notable gap in suitable evaluation regarding their applicability in the emerging domain of long-context streaming video understanding. Current benchmarks for video understanding ty…

2024

BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents

ACL 2024long

With the prosperity of large language models (LLMs), powerful LLM-based intelligent agents have been developed to provide customized services with a set of user-defined tools. State-of-the-art methods for constructing LLM agents adopt trained LLMs and further fine-tune them on data for the agent tas…

2023

Variational Causal Inference Network for Explanatory Visual Question Answering

ICCV 2023poster

Explanatory Visual Question Answering (EVQA) is a recently proposed multimodal reasoning task that requires answering visual questions and generating multimodal explanations for the reasoning processes. Unlike traditional Visual Question Answering (VQA) which focuses solely on answering, EVQA aims t…

Cited by 16PDFcodeScholar
2021

Dual Adversarial Graph Neural Networks for Multi-label Cross-modal Retrieval

AAAI 2021technical

Cross-modal retrieval has become an active study field with the expanding scale of multimodal data. To date, most existing methods transform multimodal data into a common representation space where semantic similarities between items can be directly measured across different modalities. However, the…

Cited by 70SourcePDFScholar
2021

HiT: Hierarchical Transformer With Momentum Contrast for Video-Text Retrieval

ICCV 2021poster

Video-Text Retrieval has been a hot research topic with the growth of multimedia data on the internet. Transformer for video-text learning has attracted increasing attention due to its promising performance. However, existing cross-modal transformer approaches typically suffer from two major limitat…

Cited by 191PDFScholar