← Search

Jialin Gao

11 accepted papers

2026

IdentityStory: Taming Your Identity-Preserving Generator for Human-Centric Story Generation

AAAI 2026technical

Recent visual generative models enable story generation with consistent characters from text, but human-centric story generation faces additional challenges, such as maintaining detailed and diverse human face consistency and coordinating multiple characters across different images. This paper prese

Cited by 0SourcePDFScholar
2026

PositionIC: Unified Position and Identity Consistency for Image Customization

CVPR 2026

Recent subject-driven image customization excels in fidelity, yet fine-grained instance-level spatial control remains an elusive challenge, hindering real-world applications. This limitation stems from two factors: a scarcity of scalable, position-annotated datasets, and the entanglement of identity

Cited by 0SourcecodeScholar
2026

PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework

ICLR 2026poster

Generating aesthetic posters is more challenging than simple design images: it requires not only precise text rendering but also the seamless integration of abstract artistic content, striking layouts, and overall stylistic harmony. To address this, we propose PosterCraft, a unified framework that a…

Cited by 0SourcecodeScholar
2026

PosterOmni: Generalized Artistic Poster Creation via Task Distillation and Unified Reward Feedback

CVPR 2026

Image-to-poster generation is a high-demand task requiring not only local adjustments but also high-level design understanding. Models must generate text, layout, style, and visual elements while preserving semantic fidelity and aesthetic coherence. The process spans two regimes: local editing, wher

Cited by 0SourcecodeScholar
2026

PosterReward: Unlocking Accurate Evaluation for High-Quality Graphic Design Generation

CVPR 2026

Recent advancements in the text-rendering capabilities of image generation models have made the end-to-end creation of graphic design content, such as posters, increasingly feasible. However, existing reward models fall short of accurately assessing design quality, as they primarily focus on global

Cited by 0SourcecodeScholar
2026

Rethinking Intermediate Representation for VLM-based Robot Manipulation

CVPR 2026

Vision-Language Model (VLM) is now an important component to enable robust robot manipulation. Yet, using it to translate human instructions into an action-resolvable intermediate representation often needs a tradeoff between VLM-comprehensibility and generalizability. Inspired by context-free gramm

Cited by 0SourceScholar
2025

LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

CVPR 2025poster

Recent advancements in multimodal large language models (MLLMs) have shown promising results, yet existing approaches struggle to effectively handle both temporal and spatial localization simultaneously. This challenge stems from two key issues: first, incorporating spatial-temporal localization int…

2025

SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems

ACL 2025finding

The rapid advancement of Large Multi-modal Models (LMMs) has enabled their application in scientific problem-solving, yet their fine-grained capabilities remain under-explored. In this paper, we introduce SciVerse, a multi-modal scientific evaluation benchmark to thoroughly assess LMMs across 5,735…

2024

Boundary Denoising for Video Activity Localization

ICLR 2024poster

Video activity localization aims at understanding the semantic content in long, untrimmed videos and retrieving actions of interest. The retrieved action with its start and end locations can be used for highlight generation, temporal action detection, etc. Unfortunately, learning the exact boundary…

2024

From 2D to 3D: AISG-SLA Visual Localization Challenge

IJCAI 2024poster

Research in 3D mapping is crucial for smart city applications, yet the cost of acquiring 3D data often hinders progress. Visual localization, particularly monocular camera position estimation, offers a solution by determining the camera's pose solely through visual cues. However, this task is challe…

Cited by 0SourcePDFScholar
2021

Relation-aware Video Reading Comprehension for Temporal Language Grounding

EMNLP 2021main

Temporal language grounding in videos aims to localize the temporal span relevant to the given query sentence. Previous methods treat it either as a boundary regression task or a span extraction task. This paper will formulate temporal language grounding into video reading comprehension and propose…