← Search

Yibo Shi

5 accepted papers

2026

Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents

AAAI 2026technical

Smartphones bring significant convenience to users but also enable devices to extensively record various types of personal information. Existing smartphone agents powered by Multimodal Large Language Models (MLLMs) have achieved remarkable performance in automating different tasks. However, as the c

Cited by 0SourcePDFScholar
2025

RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video

NeurIPS 2025poster

Multimodal Large Language Models (MLLMs) increasingly excel at perception,understanding, and reasoning. However, current benchmarks inadequately evaluate their ability to perform these tasks continuously in dynamic, real-world environments. To bridge this gap, we introduce RT V-Bench, a fine-grained…

Cited by 0SourcecodeScholar
2024

Neural Rate Control for Learned Video Compression

ICLR 2024poster

The learning-based video compression method has made significant progress in recent years, exhibiting promising compression performance compared with traditional video codecs. However, prior works have primarily focused on advanced compression architectures while neglecting the rate control techniqu…

Cited by 6SourcePDFScholar