← Search

Xiao Fu

57 accepted papers

2026

CHASE: Contextual History for Adaptive and Simple Exploitation in Large Language Model Jailbreaking

AAAI 2026technical

We propose Contextual History for Adaptive and Simple Exploitation (CHASE), a novel multi-turn method for Large Language Model (LLM) jailbreaking. Rather than directly attack an LLM that may be difficult to jailbreak, CHASE first collects jailbroken histories from an easy-to-jailbreak LLM and then t

Cited by 0SourcePDFScholar
2026

DuoGen: Towards Autonomous Interleaved Multimodal Generation

CVPR 2026

Interleaved multimodal generation enables capabilities beyond unimodal generation models, such as step-by-step instructional guides, visual planning, and generating visual drafts for reasoning. However, the quality of existing interleaved generation models under general instructions remains limited

Cited by 0SourceScholar
2026

Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control

ICLR 2026poster

Recent advances in video diffusion models shows promise for generating robotic decision-making data, with trajectory conditions further enabling fine-grained control. However, existing methods primarily focus on individual object motion and struggle to capture multi-object interaction crucial in com…

Cited by 0SourcecodeScholar
2025

3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation

ICLR 2025poster

This paper aims to manipulate multi-entity 3D motions in video generation. Previous methods on controllable video generation primarily leverage 2D control signals to manipulate object motions and have achieved remarkable synthesis results. However, 2D control signals are inherently limited in expres…

2025

Content-Style Learning from Unaligned Domains: Identifiability under Unknown Latent Dimensions

ICLR 2025poster

Understanding identifiability of latent content and style variables from unaligned multi-domain data is essential for tasks such as domain translation and data generation. Existing works on content-style identification were often developed under somewhat stringent conditions, e.g., that all latent c…

Cited by 0SourcePDFScholar
2025

Diverse Influence Component Analysis: A Geometric Approach to Nonlinear Mixture Identifiability

NeurIPS 2025poster

Latent component identification from unknown *nonlinear* mixtures is a foundational challenge in machine learning, with applications in tasks such as self-supervised learning and causal representation learning. Prior work in *nonlinear independent component analysis* (nICA) has shown that auxiliary…

Cited by 0SourceScholar
2025

Enhancing Multimodal Model Robustness Under Missing Modalities via Memory-Driven Prompt Learning

IJCAI 2025

Existing multimodal models typically assume the availability of all modalities, leading to significant performance degradation when certain modalities are missing. Recent methods have introduced prompt learning to adapt pretrained models to incomplete data, achieving remarkable performance when the

2025

Injecting Visual Features into Whisper for Parameter-Efficient Noise-Robust Audio-Visual Speech Recognition

ICASSP 2025accepted

Audio-visual speech recognition (AVSR) aims to enhance the robustness of an automatic speech recognition (ASR) systems by incorporating visual information from lip movements, especially in challenging noisy environments. Nevertheless, most current approaches either involve training from scratch or f…

Cited by 0SourceScholar
2025

Integrated Interpolation and Matrix Completion for Radio Map Estimation: A Convex Optimization Approach

ICASSP 2025accepted

Radio map estimation (RME) is crucial for effective planning and optimization of wireless networks. Traditional approaches such as interpolation excel at capturing local smoothness in densely populated data but struggle with sparse or irregular data. Conversely, matrix completion (MC) approaches uti…

Cited by 0SourceScholar
2025

Open-Modality Latent Modality Interaction Maximization for Audio-Visual Learning

ICASSP 2025accepted

The utilization of multimodal cues enhances the effectiveness of specific cognitive tasks in audio-visual learning. However, on the one hand, designing a unified model for multimodal learning poses challenges due to the presence of information redundancy and modality noise. On the other hand, existi…

Cited by 0SourceScholar
2025

ReCamMaster: Camera-Controlled Generative Rendering from A Single Video

ICCV 2025poster

Camera control has been actively studied in text or image conditioned video generation tasks. However, altering camera trajectories of a given video remains under-explored, despite its importance in the field of video creation. It is non-trivial due to the extra constraints of maintaining multiple-f…

2025

Under-Counted Matrix Completion Without Detection Features

ICASSP 2025accepted

Under-counted matrix completion (UC-MC) has many important applications, especially in epidemiology and ecology where the observed data are often smaller than the actual numbers. Existing works model the under-counting effects using entry-wise miss detection probabilities, which are usually formulat…

Cited by 0SourceScholar
2024

A Smoothed Bregman Proximal Gradient Algorithm for Decentralized Nonconvex Optimization

ICASSP 2024accepted

Decentralized computation has received considerable research interest lately, due to its wide applications in information processing systems. However, one key requirement to establish convergence for almost all decentralized algorithms, for convex and non-convex problems alike, is that the loss func…

Cited by 0SourceScholar
2024

Identifiable Shared Component Analysis of Unpaired Multimodal Mixtures

NeurIPS 2024poster

A core task in multi-modal learning is to integrate information from multiple feature spaces (e.g., text and audio), offering modality-invariant essential representations of data. Recent research showed that, classical tools such as canonical correlation analysis (CCA) provably identify the shared c…

2024

Noisy Label Learning with Instance-Dependent Outliers: Identifiability via Crowd Wisdom

NeurIPS 2024spotlight

The generation of label noise is often modeled as a process involving a probability transition matrix (also interpreted as the _annotator confusion matrix_) imposed onto the label distribution. Under this model, learning the ``ground-truth classifier''---i.e., the classifier that can be learned if n…

Cited by 1SourcePDFScholar
2024

Towards Identifiable Unsupervised Domain Translation: A Diversified Distribution Matching Approach

ICLR 2024poster

Unsupervised domain translation (UDT) aims to find functions that convert samples from one domain (e.g., sketches) to another domain (e.g., photos) without changing the high-level semantic meaning (also referred to as "content"). The translation functions are often sought by probability distribution…

Cited by 3SourcePDFScholar
2024

Transparent and Scrutable Recommendations Using Natural Language User Profiles

ACL 2024long

Recent state-of-the-art recommender systems predominantly rely on either implicit or explicit feedback from users to suggest new items. While effective in recommending novel options, many recommender systems often use uninterpretable embeddings to represent user preferences. This lack of transparenc…

2023

Deep Clustering with Incomplete Noisy Pairwise Annotations: A Geometric Regularization Approach

ICML 2023poster

The recent integration of deep learning and pairwise similarity annotation-based constrained clustering---i.e., deep constrained clustering (DCC)---has proven effective for incorporating weak supervision into massive data clustering: Less than 1% of pair similarity annotations can often substantiall…

2023

Deep Learning From Crowdsourced Labels: Coupled Cross-Entropy Minimization, Identifiability, and Regularization

ICLR 2023poster

Using noisy crowdsourced labels from multiple annotators, a deep learning-based end-to-end (E2E) system aims to learn the label correction mechanism and the neural classifier simultaneously. To this end, many E2E systems concatenate the neural classifier with multiple annotator-specific label confus…

2023

OmniObject3D: Large-Vocabulary 3D Object Dataset for Realistic Perception, Reconstruction and Generation

CVPR 2023poster

Recent advances in modeling 3D objects mostly rely on synthetic datasets due to the lack of large-scale real-scanned 3D databases. To facilitate the development of 3D perception, reconstruction, and generation in the real world, we propose OmniObject3D, a large vocabulary 3D object dataset with mass…

Cited by 214SourcePDFScholar
2023

Towards Efficient and Optimal Joint Beamforming and Antenna Selection: A Machine Learning Approach

ICASSP 2023accepted

This work revisits the joint transmit beamforming and antenna selection problem. Existing approaches find approximate solutions to this NP-hard problem via various heuristics, e.g., convex/nonconvex relaxation, greedy method, and (deep) supervised learning. However, optimality (or even feasibility)…

Cited by 0SourceScholar
2023

Under-Counted Tensor Completion with Neural Incorporation of Attributes

ICML 2023poster

Systematic under-counting effects are observed in data collected across many disciplines, e.g., epidemiology and ecology. Under-counted tensor completion (UC-TC) is well-motivated for many data analytics tasks, e.g., inferring the case numbers of infectious diseases at unobserved locations from unde…

2022

Communication-Efficient Distributed MAX-VAR Generalized CCA via Error Feedback-Assisted Quantization

ICASSP 2022accepted

Generalized canonical correlation analysis (GCCA) aims to learn common low-dimensional representations from multiple "views" of the data (e.g., audio and video of the same event). In the era of big data, GCCA computation encounters many new challenges. In particular, distributed optimization for GCC…

Cited by 0SourceScholar
2022

On Finite-Sample Identifiability of Contrastive Learning-Based Nonlinear Independent Component Analysis

ICML 2022spotlight

Nonlinear independent component analysis (nICA) aims at recovering statistically independent latent components that are mixed by unknown nonlinear functions. Central to nICA is the identifiability of the latent components, which had been elusive until very recently. Specifically, Hyvärinen et al. ha…

Cited by 11SourcePDFScholar
2022

Understanding Latent Correlation-Based Multiview Learning and Self-Supervision: An Identifiability Perspective

ICLR 2022spotlight

Multiple views of data, both naturally acquired (e.g., image and audio) and artificially produced (e.g., via adding different noise to data samples), have proven useful in enhancing representation learning. Natural views are often handled by multiview analysis tools, e.g., (deep) canonical correlati…

Cited by 39SourcePDFScholar
2021

Crowdsourcing via Annotator Co-occurrence Imputation and Provable Symmetric Nonnegative Matrix Factorization

ICML 2021oral

Unsupervised learning of the Dawid-Skene (D&S) model from noisy, incomplete and crowdsourced annotations has been a long-standing challenge, and is a critical step towards reliably labeling massive data. A recent work takes a coupled nonnegative matrix factorization (CNMF) perspective, and shows app…

Cited by 15SourcePDFScholar
2021

Deep Generative Model Learning For Blind Spectrum Cartography with NMF-Based Radio Map Disaggregation

ICASSP 2021accepted

Spectrum cartography (SC) aims at estimating the multi-aspect (e.g., space, frequency, and time) interference level caused by multiple emitters from limited measurements. Early SC approaches rely on model assumptions about the radio map, e.g., sparsity and smoothness, which may be grossly violated u…

Cited by 0SourceScholar
2021

Fiber-Sampled Stochastic Mirror Descent for Tensor Decomposition with β-Divergence

ICASSP 2021accepted

Canonical polyadic decomposition (CPD) has been a workhorse for multimodal data analytics. This work puts forth a stochastic algorithmic framework for CPD under β-divergence, which is well-motivated in statistical learning—where the Euclidean distance is typically not preferred. Despite the existenc…

Cited by 0SourceScholar
2021

Learning Mixed Membership from Adjacency Graph Via Systematic Edge Query: Identifiability and Algorithm

ICASSP 2021accepted

Graph clustering is a core technique for network analysis problems, e.g., community detection. This work puts forth a node clustering approach for largely incomplete adjacency graphs. Under the considered scenario, instead of having access to the complete graph, only a small amount of queries about…

Cited by 0SourceScholar
2021

Learning to Continuously Optimize Wireless Resource in Episodically Dynamic Environment

ICASSP 2021accepted

There has been a growing interest in developing data-driven, in particular deep neural network (DNN) based methods for modern communication tasks. For a few popular tasks such as power control, beamforming, and MIMO detection, these methods achieve state-of-the-art performance while requiring less c…

Cited by 0SourceScholar
2021

StatEcoNet: Statistical Ecology Neural Networks for Species Distribution Modeling

AAAI 2021technical

This paper focuses on a core task in computational sustainability and statistical ecology: species distribution modeling (SDM). In SDM, the occurrence pattern of a species on a landscape is predicted by environmental features based on observations at a set of locations. At first, SDM may appear to b…

2019

Block-randomized Stochastic Proximal Gradient for Constrained Low-rank Tensor Factorization

ICASSP 2019accepted

This work focuses on canonical polyadic decomposition (CPD) for large-scale tensors. Many prior works rely on data sparsity to develop scalable CPD algorithms, which are not suitable for handling dense tensor, while dense tensors often arise in applications such as image and video processing. As an…

Cited by 0SourceScholar
2019

Crowdsourcing via Pairwise Co-occurrences: Identifiability and Algorithms

NeurIPS 2019poster

The data deluge comes with high demands for data labeling. Crowdsourcing (or, more generally, ensemble learning) techniques aim to produce accurate labels via integrating noisy, non-expert labeling from annotators. The classic Dawid-Skene estimator and its accompanying expectation maximization (EM)…

Cited by 44SourcePDFScholar
2019

Detecting Overlapping and Correlated Communities without Pure Nodes: Identifiability and Algorithm

ICML 2019oral

Many machine learning problems come in the form of networks with relational data between entities, and one of the key unsupervised learning tasks is to detect communities in such a network. We adopt the mixed-membership stochastic blockmodel as the underlying probabilistic model, and give conditions…

Cited by 30SourcePDFScholar
2019

Regular Sampling of Tensor Signals: Theory and Application to FMRI

ICASSP 2019accepted

Sampling lies at the heart of signal processing. The celebrated Shan-non - Nyquist theorem states that in order to reconstruct a continuous or discrete time signal from uniform samples one must sample at a rate twice the highest frequency present in the signal. Numerous signals and images of interes…

Cited by 0SourceScholar
2018

Hi, Bcd! Hybrid Inexact Block Coordinate Descent for Hyperspectral Super-Resolution

ICASSP 2018accepted

Hyperspectral super-resolution (HSR) is a problem of recovering a high-spectral-spatial-resolution image from a multispectral measurement and a hyperspectral measurement, which have low spectral and spatial resolutions, respectively. We consider a low-rank structured matrix factorization formulation…

Cited by 0SourceScholar
2018

Hyperspectral Super-Resolution Via Coupled Tensor Factorization: Identifiability and Algorithms

ICASSP 2018accepted

This work focuses on the problem of fusing a hyperspectral image (HSI) and a multispectral image (MSI) to produce a super-resolution image that admits high spatial and spectral resolutions. Existing algorithms are mostly based on joint low-rank factorization of the ma-tricized HSI and MSI. This fram…

Cited by 0SourceScholar
2018

Large-Scale Regularized Sumcor GCCA via Penalty-Dual Decomposition

ICASSP 2018accepted

The sum-of-correlations (SUMCOR) generalized canonical correlation analysis (GCCA) aims at producing low-dimensional representations of multiview data via enforcing pairwise similarity of the reduced-dimension views. SUMCOR has been applied to a large variety of applications including blind separati…

Cited by 0SourceScholar
2018

Learning Hidden Markov Models from Pairwise Co-occurrences with Application to Topic Modeling

ICML 2018oral

We present a new algorithm for identifying the transition and emission probabilities of a hidden Markov model (HMM) from the emitted data. Expectation-maximization becomes computationally prohibitive for long observation records, which are often required for identification. The new algorithm is part…

Cited by 27SourcePDFScholar
2018

Shared Human-Machine Control for Self-Aware Prostheses

ICASSP 2018accepted

This paper presents a framework for shared, human-machine control of a prosthetic arm. The method employs electromyogram and peripheral neural signals to decode motor intent, and incorporates a higher-level goal in the controller to augment human effort. The controller derivation employs Markov Deci…

Cited by 0SourceScholar
2018

Tensor-Based Parameter Estimation of Double Directional Massive Mimo Channel with Dual-Polarized Antennas

ICASSP 2018accepted

The 3GPP suggests to combine dual polarized (DP) antenna arrays with the double directional (DD) channel model for downlink channel estimation. This combination strikes a good balance between high-capacity communications and parsimonious channel modeling, and also brings limited feedback schemes for…

Cited by 0SourceScholar
2017

A stochastic maximum-likelihood framework for simplex structured matrix factorization

ICASSP 2017accepted

Consider a structured matrix factorizaton (SMF) whose coefficient vectors are constrained to lie in the unit simplex. This kind of simplex SMF (SSMF) has received growing attention and has found many applications such as hyperspectral unmixing in remote sensing, text mining in machine learning, and…

Cited by 0SourceScholar
2017

Scalable and flexible Max-Var generalized canonical correlation analysis via alternating optimization

ICASSP 2017accepted

Unlike dimensionality reduction (DR) tools for single-view data, e.g., principal component analysis (PCA), canonical correlation analysis (CCA) and generalized CCA (GCCA) are able to integrate information from multiple feature spaces of data. This is critical in multi-modal data fusion and analytics…

Cited by 0SourceScholar
2017

Towards K-means-friendly Spaces: Simultaneous Deep Learning and Clustering

ICML 2017poster

Most learning approaches treat dimensionality reduction (DR) and clustering separately (i.e., sequentially), but recent research has shown that optimizing the two tasks jointly can substantially improve the performance of both. The premise behind the latter genre is that the data samples are obtaine…

2016

Anchor-Free Correlated Topic Modeling: Identifiability and Algorithm

NeurIPS 2016poster

In topic modeling, many algorithms that guarantee identifiability of the topics have been developed under the premise that there exist anchor words -- i.e., words that only appear (with positive probability) in one topic. Follow-up work has resorted to three or higher-order statistics of the data co…

Cited by 81SourcePDFScholar
2016

Robust volume minimization-based matrix factorization via alternating optimization

ICASSP 2016accepted

This paper focuses on volume minimization (VolMin)-based structured matrix factorization (SMF), which factors a data matrix into a full-column rank basis and a coefficient matrix whose columns reside in the unit simplex. The VolMin criterion achieves this goal via finding a minimum-volume enclosing…

Cited by 0SourceScholar