← Search

Zhu Xu

6 accepted papers

2025

Apply Hierarchical-Chain-of-Generation to Complex Attributes Text-to-3D Generation

CVPR 2025poster

Recent text-to-3D generation models have demonstrated remarkable abilities in producing high-quality 3D assets. Despite their great advancements, current models struggle to generate satisfying 3D objects with complex attributes. The difficulty for such complex attributes 3D generation arises from tw…

2025

Balancing Preservation and Modification: A Region and Semantic Aware Metric for Instruction-Based Image Editing

ICML 2025poster

Instruction-based image editing, which aims to modify the image faithfully towards instruction while preserving irrelevant content unchanged, has made advanced progresses. However, there still lacks a comprehensive metric for assessing the editing quality. Existing metrics either require high costs…

2025

Enhancing Character-Level Understanding in LLMs through Token Internal Structure Learning

ACL 2025long

Tokenization methods like Byte-Pair Encoding (BPE) enhance computational efficiency in large language models (LLMs) but often obscure internal character structures within tokens. This limitation hinders LLMs’ ability to predict precise character positions, which is crucial in tasks like Chinese Spel…

2025

TRKT: Weakly Supervised Dynamic Scene Graph Generation with Temporal-enhanced Relation-aware Knowledge Transferring

ICCV 2025poster

Dynamic Scene Graph Generation (DSGG) aims to create a scene graph for each video frame by detecting objects and predicting their relationships. Weakly Supervised DSGG (WS-DSGG) reduces annotation workload by using an un- localized scene graph from a single frame per video for training. Existing WS-…

2024

3D Vision and Language Pretraining with Large-Scale Synthetic Data

IJCAI 2024poster

3D Vision-Language Pre-training (3D-VLP) aims to provide a pre-train model which can bridge 3D scenes with natural language, which is an important technique for embodied intelligence. However, current 3D-VLP datasets are hindered by limited scene-level diversity and insufficient fine-grained annot…