← Search

Yingsen Zeng

4 accepted papers

2026

ViType: High-Fidelity Visual Text Rendering via Glyph-Aware Multimodal Diffusion

AAAI 2026technical

Current text-to-image models face challenges in visual text rendering: text encoders like CLIP and T5 lack glyph-level understanding and often struggle to distinguish between the specific words to be rendered and their intended semantic meaning within prompts. In addition, inconsistencies between th

Cited by 0SourcePDFScholar
2025

DisTime: Distribution-based Time Representation for Video Large Language Models

ICCV 2025poster

Despite advances in general video understanding, Video Large Language Models (Video-LLMs) face challenges in precise temporal localization due to discrete time representations and limited temporally aware datasets. Existing methods for temporal expression either conflate time with text-based numeric…

2025

InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models

ICCV 2025poster

Boosted by Multi-modal Large Language Models (MLLMs), text-guided universal segmentation models for the image and video domains have made rapid progress recently. However, these methods are often developed separately for specific domains, overlooking the similarities in task settings and solutions a…

2024

UniMD: Towards Unifying Moment Retrieval and Temporal Action Detection

ECCV 2024poster

"Temporal Action Detection (TAD) focuses on detecting pre-defined actions, while Moment Retrieval (MR) aims to identify the events described by open-ended natural language within untrimmed videos. Despite that they focus on different events, we observe they have a significant connection. For instanc…