← Search

Siqi Pan

3 accepted papers

2026

AR&D: A Framework for Retrieving and Describing Concepts for Interpreting AudioLLMs

ICASSP 2026oral

Despite strong performance in audio perception tasks, large audio-language models (AudioLLMs) remain opaque to interpretation. A major factor behind this lack of interpretability is that individual neurons in these models frequently activate in response to several unrelated concepts. We introduce th…

Cited by 0SourcePDFScholar
2026

PROMPTMINER: Black-Box Prompt Stealing against Text-to-Image Generative Models via Reinforcement Learning and VLM-Guided Optimization

CVPR 2026

Text-to-image (T2I) generative models such as Stable Diffusion and FLUX can synthesize realistic, high-quality images directly from textual prompts. The resulting image quality depends critically on well-crafted prompts that specify both subjects and stylistic modifiers, which have become valuable d

Cited by 0SourcecodeScholar
2025

Aligning What Matters: Masked Latent Adaptation for Text-to-Audio-Video Generation

NeurIPS 2025poster

Text-to-Audio-Video (T2AV) generation aims to produce temporally and semantically aligned visual and auditory content from natural language descriptions. While recent progress in text-to-audio and text-to-video models has improved generation quality within each modality, jointly modeling them remain…

Cited by 0SourceScholar