← Search

Yuanyuan Liu

32 accepted papers

2026

EmWorld: Emotion World Model with Latent State Evolution for Scenario-Incremental Dynamic Facial Expression Recognition

ICML 2026poster

Dynamic Facial Expression Recognition (DFER) models the temporal evolution of facial expressions in videos. In real-world deployments, changing scenarios distort expression trajectories over time, making it difficult for existing methods to maintain performance. While most current approaches address…

Cited by 0SourceScholar
2026

FedAdamW: A Communication-Efficient Optimizer with Convergence and Generalization Guarantees for Federated Large Models

AAAI 2026technical

AdamW has become one of the most effective optimizers for training large-scale models. We have also observed its effectiveness in the context of federated learning (FL). However, directly applying AdamW in federated learning settings poses significant challenges: (1) due to data heterogeneity, AdamW

Cited by 0SourcePDFScholar
2026

PLA-MGRA: Multi-Granularity and Relation-Aware Learning for Efficient and Generalizable Protein-Ligand Binding Affinity Prediction

AAAI 2026technical

Protein-Ligand Affinity (PLA) prediction quantifies the interaction strength to guide rational drug design. Existing approaches typically analyze interaction at a single granularity and overlook tightly coupled relationships between protein and ligand in both structure and functionality, consequentl

Cited by 0SourcePDFScholar
2026

SGAT: Learning Feature Matching with Singularity-enhanced Graph Attention Network

AAAI 2026technical

The task of image feature matching aims to establish correct correspondences between images from two different views. While approaches based on attention mechanisms have demonstrated remarkable advancements in image feature matching, they still encounter substantial limitations. Specifically, curren

Cited by 0SourcePDFScholar
2026

View-on-Graph: Zero-Shot 3D Visual Grounding via Vision-Language Reasoning on Scene Graphs

AAAI 2026technical

3D visual grounding (3DVG) identifies objects in 3D scenes from language descriptions. Existing zero-shot approaches leverage 2D vision–language models (VLMs) by converting 3D spatial information (SI) into forms amenable to VLM processing, typically as composite inputs such as specified-view renderi

Cited by 0SourcePDFScholar
2026

When Genes Speak: A Semantic-Guided Framework for Spatially Resolved Transcriptomics Data Clustering

AAAI 2026technical

Spatial transcriptomics enables gene expression profiling with spatial context, offering unprecedented insights into the tissue microenvironment. However, most computational models treat genes as isolated numerical features, ignoring the rich biological semantics encoded in their symbols. This preve

Cited by 0SourcePDFScholar
2026

World-Model Inspired Emotion-aware Token Refinement for Training-Free Multimodal Emotion Recognition

ICML 2026spotlight

Multimodal Large Language Models (MLLMs) show promise for Multimodal Emotion Recognition (MER) but often remain unreliable because sparse emotional cues could be easily overwhelmed and affected by redundant context. While fine-tuning is effective, it is usually costly when using large models. Traini…

Cited by 0SourceScholar
2025

Improving Generalization in Federated Learning with Highly Heterogeneous Data via Momentum-Based Stochastic Controlled Weight Averaging

ICML 2025poster

For federated learning (FL) algorithms such as FedSAM, their generalization capability is crucial for real-word applications. In this paper, we revisit the generalization problem in FL and investigate the impact of data heterogeneity on FL generalization. We find that FedSAM usually performs worse t…

Cited by 0SourcePDFScholar
2025

LRGR: Self-Supervised Incomplete Multi-View Clustering via Local Refinement and Global Realignment

IJCAI 2025

Incomplete Multi-View Clustering (IMVC) aims to explore comprehensive representations from multiple views with missing samples. Recent studies have revealed that IMVC methods benefit from Graph Convolutional Network (GCN) in achieving robust feature imputation and effective representation learning.

Cited by 0SourcePDFScholar
2025

Spatially Resolved Transcriptomics Data Clustering with Tailored Spatial-scale Modulation

IJCAI 2025

Spatial transcriptomics, comprising spatial location and high-throughput gene expression information, provides revolutionary insights into disease discovery and cellular evolution. Spatial transcriptomic clustering, which pinpoints distinct spatial domains within tissues, reveals cellular interactio

Cited by 0SourcePDFScholar
2025

Tight High-Probability Bounds for Nonconvex Heavy-Tailed Scenario under Weaker Assumptions

NeurIPS 2025poster

Gradient clipping is increasingly important in centralized learning (CL) and federated learning (FL). Many works focus on its optimization properties under strong assumptions involving Gaussian noise and standard smoothness. However, practical machine learning tasks often only satisfy weaker conditi…

Cited by 0SourceScholar
2025

Unsupervised Degradation Representation Aware Transform for Real-World Blind Image Super-Resolution

AAAI 2025technical

Blind image super-resolution (blind SR) aims to restore a high-resolution (HR) image from a low-resolution (LR) image with unknown degradation. Many existing methods explicitly estimate degradation information from various LR images. However, in most cases, image degradations are independent of imag…

2024

Robust and Faster Zeroth-Order Minimax Optimization: Complexity and Applications

NeurIPS 2024poster

Many zeroth-order (ZO) optimization algorithms have been developed to solve nonconvex minimax problems in machine learning and computer vision areas. However, existing ZO minimax algorithms have high complexity and rely on some strict restrictive conditions for ZO estimations. To address these issue…

Cited by 0SourcePDFScholar
2024

SAVSR: Arbitrary-Scale Video Super-Resolution via a Learned Scale-Adaptive Network

AAAI 2024technical

Deep learning-based video super-resolution (VSR) networks have gained significant performance improvements in recent years. However, existing VSR networks can only support a fixed integer scale super-resolution task, and when we want to perform VSR at multiple scales, we need to train several models…

2024

VS: Reconstructing Clothed 3D Human from Single Image via Vertex Shift

CVPR 2024poster

Various applications require high-fidelity and artifact-free 3D human reconstructions. However current implicit function-based methods inevitably produce artifacts while existing deformation methods are difficult to reconstruct high-fidelity humans wearing loose clothing. In this paper we propose a…

2023

A Single-Loop Accelerated Extra-Gradient Difference Algorithm with Improved Complexity Bounds for Constrained Minimax Optimization

NeurIPS 2023oral

In this paper, we propose a novel extra-gradient difference acceleration algorithm for solving constrained nonconvex-nonconcave (NC-NC) minimax problems. In particular, we design a new extra-gradient difference step to obtain an important quasi-cocoercivity property, which plays a key role to signif…

Cited by 1SourcePDFScholar
2023

Adaptive Non-Local Generative Adversarial Networks for Low-Dose CT Image Denoising

ICASSP 2023accepted

Low-dose computed tomography (CT) has been widely used in medical diagnosis and treatment. Many deep networks have been proposed for low-dose CT denoising. The local receptive field of the convolution affects the network performance. For different input images, conventional neural networks always ad…

Cited by 0SourceScholar
2023

Boosting Adversarial Transferability by Achieving Flat Local Maxima

NeurIPS 2023poster

Transfer-based attack adopts the adversarial examples generated on the surrogate model to attack various models, making it applicable in the physical world and attracting increasing interest. Recently, various adversarial attacks have emerged to boost adversarial transferability from different persp…

2023

Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis

EMNLP 2023long main

Though Multimodal Sentiment Analysis (MSA) proves effective by utilizing rich information from multiple sources (*e.g.,* language, video, and audio), the potential sentiment-irrelevant and conflicting information across modalities may hinder the performance from being further improved. To alleviate…

Cited by 0SourcecodeScholar
2023

Pose-Disentangled Contrastive Learning for Self-Supervised Facial Representation

CVPR 2023poster

Self-supervised facial representation has recently attracted increasing attention due to its ability to perform face understanding without relying on large-scale annotated datasets heavily. However, analytically, current contrastive-based self-supervised learning (SSL) still performs unsatisfactoril…

2022

HNO: High-Order Numerical Architecture for ODE-Inspired Deep Unfolding Networks

AAAI 2022technical

Recently, deep unfolding networks (DUNs) based on optimization algorithms have received increasing attention, and their high efficiency has been confirmed by many experimental and theoretical results. Since this type of networks combines model-based traditional optimization algorithms, they have hi…

Cited by 0SourcePDFScholar
2022

Kill a Bird with Two Stones: Closing the Convergence Gaps in Non-Strongly Convex Optimization by Directly Accelerated SVRG with Double Compensation and Snapshots

ICML 2022spotlight

Recently, some accelerated stochastic variance reduction algorithms such as Katyusha and ASVRG-ADMM achieve faster convergence than non-accelerated methods such as SVRG and SVRG-ADMM. However, there are still some gaps between the oracle complexities and their lower bounds. To fill in these gaps, th…

Cited by 2SourcePDFScholar
2021

Behavior Mimics Distribution: Combining Individual and Group Behaviors for Federated Learning

IJCAI 2021poster

Federated Learning (FL) has become an active and promising distributed machine learning paradigm. As a result of statistical heterogeneity, recent studies clearly show that the performance of popular FL methods (e.g., FedAvg) deteriorates dramatically due to the client drift caused by local updates.…

Cited by 19SourcePDFScholar
2021

Large Motion Video Super-Resolution with Dual Subnet and Multi-Stage Communicated Upsampling

AAAI 2021technical

Video super-resolution (VSR) aims at restoring a video in low-resolution (LR) and improving it to higher-resolution (HR). Due to the characteristics of video tasks, it is very important that motion information among frames should be well concerned, summarized and utilized for guidance in a VSR algor…

Cited by 30SourcePDFScholar
2021

Learned Extragradient ISTA with Interpretable Residual Structures for Sparse Coding

AAAI 2021technical

Recently, the study on learned iterative shrinkage thresholding algorithm (LISTA) has attracted increasing attentions. A large number of experiments as well as some theories have proved the high efficiency of LISTA for solving sparse coding problems. However, existing LISTA methods are all serial co…

Cited by 12SourcePDFScholar
2021

Principal component analysis in the stochastic differential privacy model

UAI 2021poster

In this paper, we study the differentially private Principal Component Analysis (PCA) problem in stochastic optimization settings. We first propose a new stochastic gradient perturbation PCA mechanism (DP-SPCA) for the calculation of the right singular subspace to achieve $(\epsilon,\delta)$-differe…

Cited by 6SourcePDFScholar
2020

Don't Hit Me! Glass Detection in Real-World Scenes

CVPR 2020poster

Glass is very common in our daily life. Existing computer vision systems neglect it and thus may have severe consequences, e.g., a robot may crash into a glass wall. However, sensing the presence of glass is not straightforward. The key challenge is that arbitrary objects/scenes can appear behind th…

Cited by 165PDFScholar
2018

Guaranteed Sufficient Decrease for Stochastic Variance Reduced Gradient Optimization

AISTATS 2018poster

In this paper, we propose a novel sufficient decrease technique for stochastic variance reduced gradient descent methods such as SVRG and SAGA. In order to make sufficient decrease for stochastic optimization, we design a new sufficient decrease criterion, which yields sufficient decrease versions o…

Cited by 0SourcePDFScholar
2017

Accelerated First-order Methods for Geodesically Convex Optimization on Riemannian Manifolds

NeurIPS 2017poster

In this paper, we propose an accelerated first-order method for geodesically convex optimization, which is the generalization of the standard Nesterov's accelerated method from Euclidean space to nonlinear Riemannian space. We first derive two equations and obtain two nonlinear operators for geodesi…

Cited by 101SourcePDFScholar
2016

Automatic speech recognition for acoustical analysis and assessment of cantonese pathological voice and speech

ICASSP 2016accepted

This paper describes the application of state-of-the-art automatic speech recognition (ASR) systems to objective assessment of voice and speech disorders. Acoustical analysis of speech has long been considered a promising approach to non-invasive and objective assessment of people. In the past the t…

Cited by 0SourceScholar
2016

Tractable and Scalable Schatten Quasi-Norm Approximations for Rank Minimization

AISTATS 2016poster

The Schatten quasi-norm was introduced to bridge the gap between the trace norm and rank function. However, existing algorithms are too slow or even impractical for large-scale problems. Motivated by the equivalence relation between the trace norm and its bilinear spectral penalty, we define two tra…

Cited by 47SourcePDFScholar