← Search

Meng Liu

55 accepted papers

2026

Content-style Disentanglement Guided Representation Learning for Deep Incomplete Multi-view Clustering

IJCAI 2026

Deep incomplete multi-view clustering methods can mine patterns of incomplete multi-view data without labels, gaining great attention in various domains. However, current methods are obsessed with aligning view-specific representations from available samples with complete views to learn view-invaria

Cited by 0Scholar
2026

D2MoRA: Diversity-Regulated Asymmetric MoE-LoRA Decomposition for Efficient Multi-Task Adaptation

AAAI 2026technical

Low-Rank Adaptation (LoRA) has emerged as a powerful parameter-efficient fine-tuning method for adapting large language models to downstream tasks. Recent studies have leveraged Mixture-of-Experts (MoE) mechanism to effectively integrate multiple LoRA modules, facilitating efficient parameter adapta

Cited by 0SourcePDFScholar
2026

Dual Optimal Transport for Multi-Concept Composition: Structural Alignment and Texture Injection in Diffusion Models

ICML 2026poster

Diffusion models have shown impressive capabilities in text-to-image synthesis, but multi-concept personalized generation remains challenging, particularly in aligning multiple reference concepts while preserving fidelity. In this work, we propose a novel framework that addresses this challenge with…

Cited by 0SourceScholar
2026

Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding

AAAI 2026technical

AI personal assistants, deployed through robots or wearables, require embodied understanding to collaborate effectively with humans. However, current Multimodal Large Language Models (MLLMs) primarily focus on third-person (exocentric) vision, overlooking the unique challenges of first-person (egoce

Cited by 0SourcePDFScholar
2026

Exploring Synthesizable Chemical Space with Iterative Pathway Refinements

ICLR 2026oral

A well-known pitfall of molecular generative models is that they are not guaranteed to generate synthesizable molecules. Existing solutions for this problem often struggle to effectively navigate exponentially large combinatorial space of synthesizable molecules and suffer from poor coverage. To add…

Cited by 0SourcecodeScholar
2026

Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action Anticipation

AAAI 2026technical

Long-term action anticipation from egocentric video is critical for applications such as human-computer interaction and assistive technologies, where anticipating user intent enables proactive and context-aware AI assistance. However, existing approaches suffer from three key limitations: 1) underut

Cited by 0SourcePDFScholar
2026

MV-FGAD: Towards Efficient and Effective Federated Graph Anomaly Detection via Multi-view Learning

ICML 2026oral

Federated graph anomaly detection (GAD) aims to identify abnormal nodes in distributed subgraphs through collaborative learning. However, existing methods suffer from two limitations. 1) Their reliance on neighborhood aggregation assumes that anomalous information can be sufficiently captured, which…

Cited by 0SourceScholar
2026

PRIME: A Decoupled Multi-agent Actor-Critic for Multi-view Clustering

IJCAI 2026

Deep multi-view clustering draws plentiful attention in various domains, owing to remarkable performance in learning patterns from complementary information of multi-view data. However, previous methods encounter two challenges. They utilize a single pre-defined clustering strategy to perceive diver

Cited by 0Scholar
2026

PhenoBrain: Phenotype-Conditioned Long-Range Communication for Multi-Modal Brain Network Analysis

ICML 2026oral

Multi-modal brain network analysis aims to predict neuropsychiatric status from functional connectomes with heterogeneous phenotypes. However, most existing methods treat phenotypes as auxiliary features and perform late fusion, implicitly assuming that the connectome representation should be learne…

Cited by 0SourceScholar
2026

ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video Retrieval

AAAI 2026technical

With the rapid growth of video data, Composed Video Retrieval (CVR) has emerged as a novel paradigm in video retrieval and is receiving increasing attention from researchers. Unlike unimodal video retrieval methods, the CVR task takes a multi-modal query consisting of a reference video and a piece o

Cited by 0SourcePDFScholar
2026

Refining Few-Step Text-to-Multiview Diffusion via Reinforcement Learning

CVPR 2026

Text-to-multiview (T2MV) diffusion models have shown great promise in generating multiple views of a scene from a single text prompt. While few-step backbones enable real-time T2MV generation, they often compromise key aspects of generation quality, such as per-view fidelity and cross-view consisten

Cited by 0SourcecodeScholar
2026

Semantic Audio-Visual Navigation in Continuous Environments

CVPR 2026

Audio-visual navigation enables embodied agents to navigate toward sound-emitting targets by leveraging both auditory and visual cues. However, most existing approaches rely on precomputed room impulse responses (RIRs) for binaural audio rendering, restricting agents to discrete grid positions and l

Cited by 0SourcecodeScholar
2026

TIME: Temporal-Sensitive Multi-Dimensional Instruction Tuning and Robust Benchmarking for Video-LLMs

AAAI 2026technical

Video large language models have achieved remarkable performance in tasks such as video question answering, however, their temporal understanding remains suboptimal. To address this limitation, we curate a dedicated instruction fine-tuning dataset that focuses on enhancing temporal comprehension acr

Cited by 0SourcePDFScholar
2026

Task-Aware 3D Affordance Segmentation via 2D Guidance and Geometric Refinement

AAAI 2026technical

Understanding 3D scene-level affordances from natural language instructions is essential for enabling embodied agents to interact meaningfully in complex environments. However, this task remains challenging due to the need for semantic reasoning and spatial grounding. Existing methods mainly focus o

Cited by 0SourcePDFScholar
2026

VLAD-Grasp: Vision-Language Adaptive Depth Grasping for Part-Specific Manipulation

RA-L 2026

In household service robot grasping tasks, precisely identifying and grasping specific object parts are essential for ensuring safety and execution efficiency. Although existing Vision–Language Models (VLMs) support open-vocabulary understanding, they still struggle with fine-grained semantic instru

Cited by 0SourceScholar
2026

VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos

ICML 2026poster

In long-video understanding, conventional uniform frame sampling often fails to capture key visual evidence, leading to degraded performance and increased hallucinations. To address this, recent agentic thinking-with-videos paradigms have emerged, adopting a localize–clip–answer pipeline in which th…

Cited by 2SourceScholar
2026

When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?

AAAI 2026technical

Can Multimodal Large Language Models (MLLMs) discern confused objects that are visually present but audio-absent? To study this, we introduce a new benchmark, AV-ConfuseBench, which simulates an “Audio-Visual Confusion” scene by modifying the corresponding sound of an object in the video, e.g., mute

Cited by 0SourcePDFScholar
2025

Dynamic-static Feature Fusion with Multi-scale Attention for Continuous Blood Glucose Prediction

ICASSP 2025accepted

Accurate continuous blood glucose prediction is an effective and direct method for treating type 2 diabetes mellitus. However, current methods are commonly single-domain single-scale blood glucose prediction models. That is, they only learn time correlations within constant time steps of continuous…

Cited by 0SourceScholar
2025

Enhanced then Progressive Fusion with View Graph for Multi-View Clustering

CVPR 2025poster

Multi-view clustering aims to improve clustering accuracy by effectively integrating complementary information from multiple perspectives. However, existing methods often encounter challenges such as feature conflicts between views and insufficient enhancement of individual view features, which hind…

Cited by 0SourcePDFScholar
2025

GenMol: A Drug Discovery Generalist with Discrete Diffusion

ICML 2025poster

Drug discovery is a complex process that involves multiple stages and tasks. However, existing molecular generative models can only tackle some of these tasks. We present *Generalist Molecular generative model* (GenMol), a versatile framework that uses only a *single* discrete diffusion model to han…

Cited by 3SourcePDFScholar
2025

Object-Shot Enhanced Grounding Network for Egocentric Video

CVPR 2025poster

Egocentric video grounding is a crucial task for embodied intelligence applications, distinct from exocentric video moment localization. Existing methods primarily focus on the distributional differences between egocentric and exocentric videos but often neglect key characteristics of egocentric vid…

2025

On the Adversarial Robustness of Multi-Kernel Clustering

ICML 2025poster

Multi-kernel clustering (MKC) has emerged as a powerful method for capturing diverse data patterns, offering robust and generalized representations of data structures. However, the increasing deployment of MKC in real-world applications raises concerns about its vulnerability to adversarial perturba…

Cited by 0SourcePDFScholar
2025

SAINT: Sequence-Aware Integration for Spatial Transcriptomics Multi-View Clustering

NeurIPS 2025poster

Spatial transcriptomics (ST) technologies provide gene expression measurements with spatial resolution, enabling the dissection of tissue structure and function. A fundamental challenge in ST analysis is clustering spatial spots into coherent functional regions. While existing models effectively int…

Cited by 0SourceScholar
2025

Spatial Understanding from Videos: Structured Prompts Meet Simulation Data

NeurIPS 2025spotlight

Visual-spatial understanding, the ability to infer object relationships and layouts from visual input, is fundamental to downstream tasks such as robotic navigation and embodied interaction. However, existing methods face spatial uncertainty and data scarcity, limiting the 3D spatial reasoning capab…

Cited by 0SourceScholar
2024

A Fast and High-quality Text-to-Speech Method with Compressed Auxiliary Corpus and Limited Target Speaker Corpus

COLING 2024main

With an auxiliary corpus (non-target speaker corpus) for model pre-training, Text-to-Speech (TTS) methods can generate high-quality speech with a limited target speaker corpus. However, this approach comes with expensive training costs. To overcome the challenge, a high-quality TTS method is propose…

Cited by 0SourcePDFScholar
2024

Hawkes-Enhanced Spatial-Temporal Hypergraph Contrastive Learning Based on Criminal Correlations

AAAI 2024technical

Crime prediction is a crucial yet challenging task within urban computing, which benefits public safety and resource optimization. Over the years, various models have been proposed, and spatial-temporal hypergraph learning models have recently shown outstanding performances. However, three correlati…

Cited by 7SourcePDFScholar
2024

MINES: Message Intercommunication for Inductive Relation Reasoning over Neighbor-Enhanced Subgraphs

AAAI 2024technical

GraIL and its variants have shown their promising capacities for inductive relation reasoning on knowledge graphs. However, the uni-directional message-passing mechanism hinders such models from exploiting hidden mutual relations between entities in directed graphs. Besides, the enclosing subgraph e…

Cited by 38SourcePDFScholar
2024

Molecule Generation with Fragment Retrieval Augmentation

NeurIPS 2024poster

Fragment-based drug discovery, in which molecular fragments are assembled into new molecules with desirable biochemical properties, has achieved great success. However, many fragment-based molecule generation methods show limited exploration beyond the existing fragments in the database as they only…

Cited by 3SourcePDFScholar
2024

Multi-Factor Adaptive Vision Selection for Egocentric Video Question Answering

ICML 2024poster

The challenge of interpreting the world from a human perspective in Artificial Intelligence (AI) is particularly evident in egocentric video question answering, which grapples with issues like small object recognition, noise suppression, and spatial-temporal reasoning. To address these challenges, w…

2024

On the Markov Property of Neural Algorithmic Reasoning: Analyses and Methods

ICLR 2024spotlight

Neural algorithmic reasoning is an emerging research direction that endows neural networks with the ability to mimic algorithmic executions step-by-step. A common paradigm in existing designs involves the use of historical embeddings in predicting the results of future execution steps. Our observati…

2023

Cross-Modal Audio-Visual Co-Learning for Text-Independent Speaker Verification

ICASSP 2023accepted

Visual speech (i.e., lip motion) is highly related to auditory speech due to the co-occurrence and synchronization in speech production. This paper investigates this correlation and proposes a cross-modal speech co-learning paradigm. The primary motivation of our cross-modal co-learning method is mo…

Cited by 0SourceScholar
2023

FedVMR: A New Federated Learning Method for Video Moment Retrieval

ICASSP 2023accepted

Despite the great success achieved, existing video moment retrieval (VMR) methods are developed under the assumption that data are centralizedly stored. However, in real-world applications, due to the inherent nature of data generation and privacy concerns, data are often distributed on different si…

Cited by 0SourceScholar
2023

Gradient-Guided Importance Sampling for Learning Binary Energy-Based Models

ICLR 2023poster

Learning energy-based models (EBMs) is known to be difficult especially on discrete data where gradient-based learning strategies cannot be applied directly. Although ratio matching is a sound method to learn discrete EBMs, it suffers from expensive computation and excessive memory requirements, the…

2023

Improving Domain Generalization for Prompt-Aware Essay Scoring via Disentangled Representation Learning

ACL 2023long

Automated Essay Scoring (AES) aims to score essays written in response to specific prompts. Many AES models have been proposed, but most of them are either prompt-specific or prompt-adaptive and cannot generalize well on “unseen” prompts. This work focuses on improving the generalization ability of…

Cited by 14SourcePDFScholar
2023

Joint Learning of Label and Environment Causal Independence for Graph Out-of-Distribution Generalization

NeurIPS 2023poster

We tackle the problem of graph out-of-distribution (OOD) generalization. Existing graph OOD algorithms either rely on restricted assumptions or fail to exploit environment information in training data. In this work, we propose to simultaneously incorporate label and environment causal independence (…

2023

Leveraging Positional-Related Local-Global Dependency for Synthetic Speech Detection

ICASSP 2023accepted

Automatic speaker verification (ASV) systems are vulnerable to spoofing attacks. As synthetic speech exhibits local and global artifacts compared to natural speech, incorporating local-global dependency would lead to better anti-spoofing performance. To this end, we propose the Rawformer that levera…

Cited by 0SourceScholar
2023

Noise-Disentanglement Metric Learning for Robust Speaker Verification

ICASSP 2023accepted

Automatic speaker verification (ASV) suffers from performance degradation in noisy environments. To solve this problem, we propose the noise-disentanglement metric learning to reduce the speaker-irrelevant noisy components and build a noise-invariant embedding space. Specifically, the disentanglemen…

Cited by 0SourceScholar
2023

QH9: A Quantum Hamiltonian Prediction Benchmark for QM9 Molecules

NeurIPS 2023poster

Supervised machine learning approaches have been increasingly used in accelerating electronic structure prediction as surrogates of first-principle computational methods, such as density functional theory (DFT). While numerous quantum chemistry datasets focus on chemical properties and atomic forces…

2023

Self-Supervised Audio-Visual Speaker Representation with Co-Meta Learning

ICASSP 2023accepted

In self-supervised speaker verification, the quality of pseudo labels determines the upper bound of its performance and it is not uncommon to end up with massive amount of unreliable pseudo labels. We observe that the complementary information in different modalities ensures a robust supervisory sig…

Cited by 0SourceScholar
2023

Video Timeline Modeling For News Story Understanding

NeurIPS 2023spotlight

In this paper, we present a novel problem, namely video timeline modeling. Our objective is to create a video-associated timeline from a set of videos related to a specific topic, thereby facilitating the content and structure understanding of the story being told. This problem has significant poten…

2022

Generating 3D Molecules for Target Protein Binding

ICML 2022oral

A fundamental problem in drug discovery is to design molecules that bind to specific proteins. To tackle this problem using machine learning methods, here we propose a novel and effective framework, known as GraphBP, to generate 3D molecules that bind to given proteins by placing atoms of specific t…

2022

GraphFM: Improving Large-Scale GNN Training via Feature Momentum

ICML 2022spotlight

Training of graph neural networks (GNNs) for large-scale node classification is challenging. A key difficulty lies in obtaining accurate hidden node representations while avoiding the neighborhood explosion problem. Here, we propose a new technique, named feature momentum (FM), that uses a momentum…

2022

Learning Domain-Invariant Transformation for Speaker Verification

ICASSP 2022accepted

Automatic speaker verification (ASV) faces domain shift caused by the mismatch of intrinsic and extrinsic factors such as recording device and speaking style in real-world applications, which leads to unsatisfactory performance. To this end, we propose the meta generalized transformation via meta-le…

Cited by 0SourceScholar
2022

Spherical Message Passing for 3D Molecular Graphs

ICLR 2022poster

We consider representation learning of 3D molecular graphs in which each atom is associated with a spatial position in 3D. This is an under-explored area of research, and a principled message passing framework is currently lacking. In this work, we conduct analyses in the spherical coordinate system…

Cited by 228SourcePDFScholar
2021

Meta-Learning for Cross-Channel Speaker Verification

ICASSP 2021accepted

Automatic speaker verification (ASV) has been successfully deployed for identity recognition. With increasing use of ASV technology in real-world applications, channel mismatch caused by the recording devices and environments severely degrade its performance, especially in the case of unseen channel…

Cited by 0SourceScholar
2021

Multi-Modal Relational Graph for Cross-Modal Video Moment Retrieval

CVPR 2021poster

Given an untrimmed video and a query sentence, cross-modal video moment retrieval aims to rank a video moment from pre-segmented video moment candidates that best matches the query sentence. Pioneering work typically learns the representations of the textual and visual content separately and then ob…

Cited by 86PDFcodeScholar
2021

Replay-Attack Detection Using Features With Adaptive Spectro-Temporal Resolution

ICASSP 2021accepted

Variable-resolution processing aims to improve the feature representation ability by enlarging the local discriminative details. In previous anti-spoofing studies, different phones and frequency regions were both proven to have various levels of sensitivity to replay distortion. In this paper, an ad…

Cited by 0SourceScholar
2020

Strongly local p-norm-cut algorithms for semi-supervised learning and local graph clustering

NeurIPS 2020poster

Graph based semi-supervised learning is the problem of learning a labeling function for the graph nodes given a few example nodes, often called seeds, usually under the assumption that the graph’s edges indicate similarity of labels. This is closely related to the local graph clustering or communit…

2019

Replay Attack Detection Using Magnitude and Phase Information with Attention-based Adaptive Filters

ICASSP 2019accepted

Automatic Speech Verification (ASV) systems are highly vulnerable to spoofing attacks, and replay attack poses the greatest threat among various spoofing attacks. In this paper, we propose a novel multi-channel feature extraction method with attention-based adaptive filters (AAF). Original phase inf…

Cited by 0SourceScholar