← Search

Zhimin Li

7 accepted papers

2026

Re-Align: Structured Reasoning-guided Alignment for In-Context Image Generation and Editing

CVPR 2026

In-context image generation and editing (ICGE) enables users to specify visual concepts through interleaved image-text prompts, demanding precise understanding and faithful execution of user intent. Although recent unified multimodal models exhibit promising understanding capabilities, these strengt

Cited by 0SourceScholar
2025

TRKT: Weakly Supervised Dynamic Scene Graph Generation with Temporal-enhanced Relation-aware Knowledge Transferring

ICCV 2025poster

Dynamic Scene Graph Generation (DSGG) aims to create a scene graph for each video frame by detecting objects and predicting their relationships. Weakly Supervised DSGG (WS-DSGG) reduces annotation workload by using an un- localized scene graph from a single frame per video for training. Existing WS-…

2025

Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning

NeurIPS 2025poster

Recent advances in multimodal Reward Models (RMs) have shown significant promise in delivering reward signals to align vision models with human preferences. However, current RMs are generally restricted to providing direct responses or engaging in shallow reasoning processes with limited depth, ofte…

Cited by 0SourceScholar
2024

Digital Avatars: Framework Development and Their Evaluation

IJCAI 2024poster

We present a novel prompting strategy for artificial intelligence driven digital avatars. To better quantify how our prompting strategy affects anthropomorphic features like humor, authenticity, and favorability we present Crowd Vote - an adaptation of Crowd Score that allows for judges to elect a l…

Cited by 0SourcePDFScholar
2024

OED: Towards One-stage End-to-End Dynamic Scene Graph Generation

CVPR 2024poster

Dynamic Scene Graph Generation (DSGG) focuses on identifying visual relationships within the spatial-temporal domain of videos. Conventional approaches often employ multi-stage pipelines which typically consist of object detection temporal association and multi-relation classification. However these…

2022

Category-Aware Transformer Network for Better Human-Object Interaction Detection

CVPR 2022poster

Human-Object Interactions (HOI) detection, which aims to localize a human and a relevant object while recognizing their interaction, is crucial for understanding a still image. Recently, tranformer-based models have significantly advanced the progress of HOI detection. However, the capability of the…

Cited by 46PDFScholar
2022

Improving Human-Object Interaction Detection via Phrase Learning and Label Composition

AAAI 2022technical

Human-Object Interaction (HOI) detection is a fundamental task in high-level human-centric scene understanding. We propose PhraseHOI, containing a HOI branch and a novel phrase branch, to leverage language prior and improve relation expression. Specifically, the phrase branch is supervised by semant…

Cited by 44SourcePDFScholar