← Search

YUE JIANG

20 accepted papers

2026

AutoSchA: Automatic Hierarchical Music Representations via Multi-Relational Node Isolation

AAAI 2026technical

Hierarchical representations provide powerful and principled approaches for analyzing many musical genres. Such representations have been broadly studied in music theory, for instance via Schenkerian analysis (SchA). Hierarchical music analyses, however, are highly cost-intensive; the analysis of a

Cited by 0SourcePDFScholar
2026

MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-turn Dialogue

ICML 2026poster

Multimodal Large Language Models (MLLMs) demonstrate remarkable visual understanding, yet their reliability in interactive settings is severely undermined by {hallucination snowballing}: a phenomenon where initial errors amplify across conversational turns, leading to a collapse in coherence. This f…

Cited by 0SourceScholar
2026

SatireDecoder: Visual Cascaded Decoupling for Enhancing Satirical Image Comprehension

AAAI 2026technical

Satire, a form of artistic expression combining humor with implicit critique, holds significant social value by illuminating societal issues. Despite its cultural and societal significance, satire comprehension, particularly in purely visual forms, remains a challenging task for current vision-langu

Cited by 0SourcePDFScholar
2025

A Unified Supervised and Unsupervised Dialogue Topic Segmentation Framework Based on Utterance Pair Modeling

NAACL 2025long

The Dialogue Topic Segmentation task aims to divide a dialogue into different topic paragraphs in order to better understand the structure and content of the dialogue. Due to the short sentences, serious references and non-standard language in the dialogue, it is difficult to determine the boundarie…

Cited by 0SourcePDFScholar
2025

BloomScene: Lightweight Structured 3D Gaussian Splatting for Crossmodal Scene Generation

AAAI 2025technical

With the widespread use of virtual reality applications, 3D scene generation has become a new challenging research frontier. 3D scenes have highly complex structures and need to ensure that the output is dense, coherent, and contains all necessary structures. Many current 3D scene generation methods…

2025

CoMT: Chain-of-Medical-Thought Reduces Hallucination in Medical Report Generation

ICASSP 2025accepted

Automatic medical report generation (MRG), which possesses significant research value as it can aid radiologists in clinical diagnosis and report composition, has garnered increasing attention. Despite recent progress, generating accurate reports remains arduous due to the requirement for precise cl…

Cited by 0SourceScholar
2025

DanmakuTPPBench: A Multi-modal Benchmark for Temporal Point Process Modeling and Understanding

NeurIPS 2025poster

We introduce DanmakuTPPBench, a comprehensive benchmark designed to advance multi-modal Temporal Point Process (TPP) modeling in the era of Large Language Models (LLMs). While TPPs have been widely studied for modeling temporal event sequences, existing datasets are predominantly unimodal, hinderin…

Cited by 0SourcecodeScholar
2025

FSTLLM: Spatio-Temporal LLM for Few Shot Time Series Forecasting

ICML 2025poster

Time series forecasting fundamentally relies on accurately modeling complex interdependencies and shared patterns within time series data. Recent advancements, such as Spatio-Temporal Graph Neural Networks (STGNNs) and Time Series Foundation Models (TSFMs), have demonstrated promising results by eff…

2025

Image-level Memorization Detection via Inversion-based Inference Perturbation

ICLR 2025poster

Recent studies have discovered that widely used text-to-image diffusion models can replicate training samples during image generation, a phenomenon known as memorization. Existing detection methods primarily focus on identifying memorized prompts. However, in real-world scenarios, image owners may n…

Cited by 0SourcePDFScholar
2025

MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image Generation

CVPR 2025poster

Diffusion models have shown excellent performance in text-to-image generation. However, existing methods often suffer from performance bottlenecks when dealing with complex prompts involving multiple objects, characteristics, and relations. Therefore, we propose a Multi-agent Collaboration-based Co…

Cited by 0SourcePDFScholar
2025

Offline Reinforcement Learning with Koopman Operators for Control of Soft Robots

IROS 2025

Soft robots are promising to offer flexibility in environmental interaction tasks through compliant deformations. However, the infinite degrees of freedom and high nonlinearity of dynamics pose significant challenges in dynamic modeling and control in soft robots. While online reinforcement learning

Cited by 0SourceScholar
2024

DreamStruct: Understanding Slides and User Interfaces via Synthetic Data Generation

ECCV 2024poster

"Enabling machines to understand structured visuals like slides and user interfaces is essential for making them accessible to people with disabilities. However, achieving such understanding computationally has required manual data collection and annotation, which is time-consuming and labor-intensi…

2024

H-LegalKI: A Hierarchical Legal Knowledge Integration Framework for Legal Community Question Answering

EMNLP 2024finding

Legal question answering (LQA) aims to bridge the gap between the limited availability of legal professionals and the high demand for legal assistance. Traditional LQA approaches typically either select the optimal answers from an answer set or extract answers from law texts. However, they often str…

2024

PediatricsGPT: Large Language Models as Chinese Medical Assistants for Pediatric Applications

NeurIPS 2024poster

Developing intelligent pediatric consultation systems offers promising prospects for improving diagnostic efficiency, especially in China, where healthcare resources are scarce. Despite recent advances in Large Language Models (LLMs) for Chinese medicine, their performance is sub-optimal in pediatri…

2024

SC2: Towards Enhancing Content Preservation and Style Consistency in Long Text Style Transfer

ACL 2024long

Text style transfer (TST) aims to vary the style polarity of text while preserving the semantic content. Although recent advancements have demonstrated remarkable progress in short TST, it remains a relatively straightforward task with limited practical applications. The more comprehensive long TST…

2024

Toward Robust Incomplete Multimodal Sentiment Analysis via Hierarchical Representation Learning

NeurIPS 2024poster

Multimodal Sentiment Analysis (MSA) is an important research area that aims to understand and recognize human sentiment through multiple modalities. The complementary information provided by multimodal fusion promotes better sentiment analysis compared to utilizing only a single modality. Neverthele…

Cited by 1SourcePDFScholar
2024

UrbanLLM: Autonomous Urban Activity Planning and Management with Large Language Models

EMNLP 2024finding

Location-based services play an critical role in improving the quality of our daily lives. Despite the proliferation of numerous specialized AI models within spatio-temporal context of location-based services, these models struggle to autonomously tackle problems regarding complex urban planing and…

2023

Enhancing Cross-lingual Prompting with Dual Prompt Augmentation

ACL 2023findings

Prompting shows promising results in few-shot scenarios. However, its strength for multilingual/cross-lingual problems has not been fully exploited. hao and Schütze (2021) made initial explorations in this direction by presenting that cross-lingual prompting outperforms cross-lingual finetuning. In…

2023

Multivariate Time-series Imputation with Disentangled Temporal Representations

ICLR 2023poster

Multivariate time series often faces the problem of missing value. Many time series imputation methods have been developed in the literature. However, these methods all rely on an entangled representation to model dynamics of time series, which may fail to fully exploit the multiple factors (e.g., p…

Cited by 37SourcePDFScholar
2020

SDFDiff: Differentiable Rendering of Signed Distance Fields for 3D Shape Optimization

CVPR 2020oral

We propose SDFDiff, a novel approach for image-based shape optimization using differentiable rendering of 3D shapes represented by signed distance functions (SDFs). Compared to other representations, SDFs have the advantage that they can represent shapes with arbitrary topology, and that they guaran…

Cited by 271PDFcodeScholar