← Search

Liu He

7 accepted papers

2026

Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting

AAAI 2026technical

In this work, we introduce CAT-V (Caption Anything in Video), a training-free framework for fine-grained object-centric video captioning of user-selected instances. CAT-V combines (i) a SAMURAI-based Segmenter for precise object masks across frames, (ii) a TRACE-Uni Temporal Analyzer for event bound

Cited by 0SourcePDFScholar
2025

DocAgent: An Agentic Framework for Multi-Modal Long-Context Document Understanding

EMNLP 2025

Recent advances in large language models (LLMs) have demonstrated significant promise in document understanding and question-answering. Despite the progress, existing approaches can only process short documents due to limited context length or fail to fully leverage multi-modal information. In this

2025

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models

NeurIPS 2025poster

Recent multimodal image generators such as GPT-4o, Gemini 2.0 Flash, and Gemini 2.5 Pro excel at following complex instructions, editing images and maintaining concept consistency. However, they are still evaluated by disjoint toolkits: text-to-image (T2I) benchmarks that lacks multi-modal condition…

Cited by 0SourceScholar
2025

Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic Alignment

ICLR 2025poster

Personalized image generation has emerged from the recent advancements in generative models. However, these generated personalized images often suffer from localized artifacts such as incorrect logos, reducing fidelity and fine-grained identity details of the generated results. Furthermore, there is…

Cited by 1SourcePDFScholar
2025

Unified Planning Framework With Drivable Area Attention Extraction for Autonomous Driving in Urban Scenarios

RA-L 2025

The diversity of urban traffic scenarios poses challenges in stability and generalization for autonomous driving. To tackle this issue, this paper proposes a hierarchical decision-making and planning framework based on reinforcement learning, which employs a unified drivable area cross-attention ext

Cited by 1SourcecodeScholar
2019

Decentralized Full Coverage of Unknown Areas by Multiple Robots With Limited Visibility Sensing

RA-L 2019

This letter addresses the full coverage problem of unknown convex and concave two-dimensional (2-D) areas by multiple robots with limited visibility sensing and communication range. The areas are initially unknown to the multiple robots, and the number of robots is not predefined. In order to accomp

Cited by 13SourceScholar