← Search

Zhenchao Jin

10 accepted papers

2026

AssetFormer: Modular 3D Assets Generation with Autoregressive Transformer

ICLR 2026poster

The digital industry demands high-quality, diverse modular 3D assets, especially for user-generated content (UGC). In this work, we introduce AssetFormer, an autoregressive Transformer-based model designed to generate modular 3D assets from textual descriptions. Our pilot study leverages real-world…

Cited by 0SourcecodeScholar
2026

Bridging Facial Understanding and Animation via Language Models

CVPR 2026

Text-guided human body animation has advanced rapidly, yet facial animation lags due to the scarcity of well-annotated, text-paired facial corpora. To close this gap, we leverage foundation generative models to synthesize a large, balanced corpus of facial behavior. We design prompts suite covering

Cited by 0SourceScholar
2026

GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision

CVPR 2026

Multimodal large reasoning models (MLRMs) are increasingly deployed for vision-language tasks that produce explicit intermediate rationales. However, reasoning traces can contain unsafe content even when the final answer is non-harmful, creating deployment risks. Existing multimodal safety guards pr

Cited by 0SourcecodeScholar
2025

Large Images Are Gaussians: High-Quality Large Image Representation with Levels of 2D Gaussian Splatting

AAAI 2025technical

While Implicit Neural Representations (INRs) have demonstrated significant success in image representation, they are often hindered by large training memory and slow decoding speed. Recently, Gaussian Splatting (GS) has emerged as a promising solution in 3D reconstruction due to its highquality nove…

2023

Emotional Listener Portrait: Neural Listener Head Generation with Emotion

ICCV 2023poster

Listener head generation centers on generating non-verbal behaviors (e.g., smile) of a listener in reference to the information delivered by a speaker. A significant challenge when generating such responses is the non-deterministic nature of fine-grained facial expressions during a conversation, whi…

Cited by 9PDFScholar
2023

IDRNet: Intervention-Driven Relation Network for Semantic Segmentation

NeurIPS 2023poster

Co-occurrent visual patterns suggest that pixel relation modeling facilitates dense prediction tasks, which inspires the development of numerous context modeling paradigms, \emph{e.g.}, multi-scale-driven and similarity-driven context schemes. Despite the impressive results, these existing paradigms…

2022

Adaptive Face Forgery Detection in Cross Domain

ECCV 2022poster

"It is necessary to develop effective face forgery detection methods with constantly evolving technologies in synthesizing realistic faces which raises serious risks on malicious face tampering. A large and growing body of literature has investigated deep learning-based approaches, especially those…

2021

ISNet: Integrate Image-Level and Semantic-Level Context for Semantic Segmentation

ICCV 2021poster

Co-occurrent visual pattern makes aggregating contextual information a common paradigm to enhance the pixel representation for semantic image segmentation. The existing approaches focus on modeling the context from the perspective of the whole image, i.e., aggregating the image-level contextual info…

Cited by 91PDFcodeScholar
2021

Mining Contextual Information Beyond Image for Semantic Segmentation

ICCV 2021poster

This paper studies the context aggregation problem in semantic image segmentation. The existing researches focus on improving the pixel representations by aggregating the contextual information within individual images. Though impressive, these methods neglect the significance of the representations…

Cited by 105PDFcodeScholar