← Search

Zhen Xiong

8 accepted papers

2026

Unveiling the Potential of Diffusion Large Language Model in Controllable Generation

ICLR 2026poster

Controllable generation is a fundamental task in NLP with many applications, providing a basis for function calling to agentic communication. However, even state-of-the-art autoregressive Large Language Models (LLMs) today exhibit unreliability when required to generate structured output. Inspired b…

Cited by 0SourcecodeScholar
2025

Enhancing Image Generation Fidelity via Progressive Prompts

ICASSP 2025accepted

Diffusion transformer (DiT) architecture catches much attention in image generation, which achieves better fidelity, performance, and diversity. However, most existing DiT-based image generation methods are global-aware synthesis and regional prompt control is less explored. In this paper, we propos…

Cited by 0SourceScholar
2025

GraphNarrator: Generating Textual Explanations for Graph Neural Networks

ACL 2025long

Graph representation learning has garnered significant attention due to its broad applications in various domains, such as recommendation systems and social network analysis. Despite advancements in graph learning methods, challenges still remain in explainability when graphs are associated with sem…

Cited by 0SourcePDFScholar
2025

Rethinking Decoding in Multi-intent Spoken Language Understanding

ICASSP 2025accepted

Multi-intent spoken language understanding (SLU) can handle multiple intent utterances in real-world scenarios, which has gained increasing research attention. Despite promising results achieved by existing joint models, they (1) perform utterance-level or token-level intent detection, resulting in…

Cited by 0SourceScholar
2025

Robust and Efficient 3D Gaussian Splatting for Urban Scene Reconstruction

ICCV 2025poster

We present a framework that enables fast reconstruction and real-time rendering of urban-scale scenes while maintaining robustness against appearance variations across multi-view captures. Our approach begins with scene partitioning for parallel training, employing a visibility-based image selection…

2025

Towards Zero-shot Cross-lingual SLU with Syntax-aware Multi-view Contrastive Learning

ICASSP 2025accepted

Recent state-of-the-art zero-shot cross-lingual spoken language understanding (SLU) models utilize contrastive learning to achieve multilingual semantics alignment between the original utterance and code-switched counterpart. Despite achieving promising results, we discover that they still suffer fr…

Cited by 0SourceScholar
2025

Vulnerability of LLMs to Vertically Aligned Text Manipulations

ACL 2025long

Vertical text input is commonly encountered in various real-world applications, such as mathematical computations and word-based Sudoku puzzles. While current large language models (LLMs) have excelled in natural language tasks, they remain vulnerable to variations in text formatting.Recent research…

Cited by 0SourcePDFScholar