← Search

Wenting Xu

4 accepted papers

2025

ASANet: Scene Text Recognition With Alternate Self-Attention

ICASSP 2025accepted

Text recognition in complex scenes is a challenging task in Computer Vision. In this paper, we propose an innovative framework for scene text recognition, ASANet, which features an Alternate Attention Enhancement Encoder and a Masked Dual-modal Decoder. The encoder incorporates a 12-layer Alternatin…

Cited by 0SourceScholar
2025

Storynizor: Consistent Story Generation via Inter-Frame Synchronized and Shuffled ID Injection

AAAI 2025technical

Recent advances in text-to-image diffusion models have spurred significant interest in continuous story image generation. In this paper, we introduce Storynizor, a model capable of generating coherent stories with strong inter-frame character consistency, effective foreground-background separation,…

Cited by 1SourcePDFScholar
2025

TB-HSU: Hierarchical 3D Scene Understanding with Contextual Affordances

AAAI 2025technical

The concept of function and affordance is a critical aspect of 3D scene understanding and supports task-oriented objectives. In this work, we develop a model that learns to structure and vary functional affordance across a 3D hierarchical scene graph representing the spatial organization of a scene.…

2023

MvCo-DoT: Multi-View Contrastive Domain Transfer Network for Medical Report Generation

ICASSP 2023accepted

In clinical scenarios, multiple medical images with different views are usually generated at the same time, and they have high semantic consistency. However, the existing medical report generation methods cannot exploit the rich multi-view mutual information of medical images. Therefore, in this wor…

Cited by 9SourceScholar