← Search

Zhiyong Wang

59 accepted papers

2026

Detect All-Type Deepfake Audio: Wavelet Prompt Tuning for Enhanced Auditory Perception

AAAI 2026technical

The rapid advancement of audio generation technologies has escalated the risks of malicious deepfake audio across speech, sound, singing voice, and music, threatening multimedia security and trust. While existing countermeasures (CMs) perform well in single-type audio deepfake detection (ADD), their

Cited by 0SourcePDFScholar
2026

DuoCast: Duo-Probabilistic Diffusion for Precipitation Nowcasting

AAAI 2026technical

Accurate short-term precipitation forecasting is critical for weather-sensitive decision-making in agriculture, transportation, and disaster response. Existing deep learning approaches often struggle to balance global structural consistency with local detail preservation, especially under complex me

Cited by 0SourcePDFScholar
2026

Expanding the Chaos: Neural Operator for Stochastic (Partial) Differential Equations

ICML 2026poster

Stochastic differential equations (SDEs) and stochastic partial differential equations (SPDEs) are fundamental for modeling stochastic dynamics across the natural sciences and modern machine learning. Learning their solution operators with deep learning models promises fast solvers and new perspecti…

Cited by 0SourceScholar
2026

FAKE SPEECH WILD: DETECTING DEEPFAKE SPEECH ON SOCIAL MEDIA PLATFORM

ICASSP 2026poster

The rapid advancement of speech generation technology has led to the widespread proliferation of deepfake speech across social media platforms. While deepfake audio countermeasures (CMs) achieve promising results on public datasets, their performance degrades significantly in cross-domain scenarios.…

Cited by 0SourcePDFScholar
2026

HKAFER: Achieve Visual Parameter-Efficient Fine-Tuning via Heterogeneous Kronecker Adaptation for Facial Expression Recognition

AAAI 2026technical

Facial Expression Recognition (FER) seeks to classify affective states from facial images, which remains a challenging problem due to variations in real-world conditions. FER task becomes particularly complex when handling unconstrained environments characterized by partial occlusions, different hea

Cited by 0SourcePDFScholar
2026

LaDy: Lagrangian-Dynamic Informed Network for Skeleton-based Action Segmentation via Spatial-Temporal Modulation

CVPR 2026

Skeleton-based Temporal Action Segmentation (STAS) aims to densely parse untrimmed skeletal sequences into frame-level action categories. However, existing methods, while proficient at capturing spatio-temporal kinematics, neglect the underlying physical dynamics that govern human motion. This overs

Cited by 0SourcecodeScholar
2026

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks

ICML 2026poster

RL-based post-training with GRPO is widely used to improve large language models on individual reasoning tasks. However, real-world deployment requires reliable performance across diverse tasks. A straightforward multi-task adaptation of GRPO often leads to imbalanced outcomes, with some tasks domin…

Cited by 0SourceScholar
2026

Pb4U-GNet: Resolution-Adaptive Garment Simulation via Propagation-before-Update Graph Network

AAAI 2026technical

Garment simulation is fundamental to various applications in computer vision and graphics, from virtual try-on to digital human modelling. However, conventional physics-based methods remain computationally expensive, hindering their application in time-sensitive scenarios. While graph neural network

Cited by 0SourcePDFScholar
2026

PhenoYieldNet: Learning Crop-Aware Phenological Responses for Multi-Crop Yield Prediction

CVPR 2026

Accurate crop yield prediction is crucial for sustainable agriculture and global food security. While existing methods are predominantly developed for single-crop prediction, they often struggle to generalize across diverse crop types, without addressing the unique crop phenological responses that a

Cited by 0SourcecodeScholar
2026

Spectral Scalpel: Amplifying Adjacent Action Discrepancy via Frequency-Selective Filtering for Skeleton-Based Action Segmentation

CVPR 2026

Skeleton-based Temporal Action Segmentation (STAS) seeks to densely segment and classify diverse actions within long, untrimmed skeletal motion sequences. However, existing STAS methodologies face challenges of limited inter-class discriminability and blurred segmentation boundaries, primarily due t

Cited by 0SourcecodeScholar
2026

T-POP: Test-Time Personalization with Online Preference Feedback

ICML 2026poster

Personalizing large language models (LLMs) to individual user preferences is a critical step beyond generating generically helpful responses. However, current personalization methods are ill-suited for new users, as they typically require either slow, resource-intensive fine-tuning or a substantial …

Cited by 0SourceScholar
2026

UrbanVLA: A Vision-Language-Action Model for Urban Micromobility

ICRA 2026poster

Urban micromobility applications, such as delivery robots, demand reliable navigation across large-scale urban environments while following long-horizon route instructions. This task is particularly challenging due to the dynamic and unstructured nature of real-world city areas, yet most existing na…

2025

A Plug-and-Play Diffusion-Styled Conversion Model for Domain Discrepancies in Medical Image Segmentation

ICASSP 2025accepted

Accurate segmentation is a crucial step in medical image analysis. However, models trained on one dataset often suffer from performance degradation when directly applied to a different domain with a different data distribution due to domain discrepancies. To address this issue, we introduce a novel…

Cited by 0SourceScholar
2025

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens

ICCV 2025poster

Recently, Vision Large Language Models (VLLMs) with integrated vision encoders have shown promising performance in vision understanding. They encode visual content into sequences of visual tokens, enabling joint processing of visual and textual data. However, understanding videos, especially long vi…

2025

DC-PCN: Point Cloud Completion Network with Dual-Codebook Guided Quantization

AAAI 2025technical

Point cloud completion aims to reconstruct complete 3D shapes from partial 3D point clouds. With advancements in deep learning techniques, various methods for point cloud completion have been developed. Despite achieving encouraging results, a significant issue remains: these methods often overlook…

2025

DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech

ICASSP 2025accepted

In recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have also gained attention. However, current DiT speech models treat Mel spectrograms as general images, which overlooks the…

Cited by 0SourceScholar
2025

Diffusing to the Top: Boost Graph Neural Networks with Minimal Hyperparameter Tuning

ICLR 2025poster

Graph Neural Networks (GNNs) are proficient in graph representation learning and achieve promising performance on versatile tasks such as node classification and link prediction. Usually, a comprehensive hyperparameter tuning is essential for fully unlocking GNN's top performance, especially for com…

2025

Federated In-Context Learning: Iterative Refinement for Improved Answer Quality

ICML 2025poster

For question-answering (QA) tasks, in-context learning (ICL) enables language models (LMs) to generate responses without modifying their parameters by leveraging examples provided in the input. However, the effectiveness of ICL heavily depends on the availability of high-quality examples, which are…

Cited by 0SourcePDFScholar
2025

HGCLIP: Exploring Vision-Language Models with Graph Representations for Hierarchical Understanding

COLING 2025main

Object categories are typically organized into a multi-granularity taxonomic hierarchy. When classifying categories at different hierarchy levels, traditional uni-modal approaches focus primarily on image features, revealing limitations in complex scenarios. Recent studies integrating Vision-Languag…

2025

InstructHOI: Context-Aware Instruction for Multi-Modal Reasoning in Human-Object Interaction Detection

NeurIPS 2025spotlight

Recently, Large Foundation Models (LFMs), e.g., CLIP and GPT, have significantly advanced the Human-Object Interaction (HOI) detection, due to their superior generalization and transferability. Prior HOI detectors typically employ single- or multi-modal prompts to generate discriminative representat…

Cited by 0SourceScholar
2025

Learning Across the Gap: Hybrid Multi-armed Bandits with Heterogeneous Offline and Online Data

NeurIPS 2025poster

The multi-armed bandit (MAB) is a fundamental online decision-making framework that has been extensively studied over the past two decades. To mitigate the high cost and slow convergence of purely online learning, modern MAB approaches have explored _hybrid_ paradigms that leverage offline data to w…

Cited by 0SourceScholar
2025

Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0

ICASSP 2025accepted

Speech synthesis technology has posed a serious threat to speaker verification systems. Currently, the most effective fake audio detection methods utilize pretrained models, and integrating features from various layers of pretrained model further enhances detection performance. However, most of the…

Cited by 0SourceScholar
2025

Model-based RL as a Minimalist Approach to Horizon-Free and Second-Order Bounds

ICLR 2025poster

Learning a transition model via Maximum Likelihood Estimation (MLE) followed by planning inside the learned model is perhaps the most standard and simplest Model-based Reinforcement Learning (RL) framework. In this work, we show that such a simple Model-based RL scheme, when equipped with optimistic…

Cited by 5SourcePDFScholar
2025

Online Clustering of Dueling Bandits

ICML 2025poster

The contextual multi-armed bandit (MAB) is a widely used framework for problems requiring sequential decision-making under uncertainty, such as recommendation systems. In applications involving a large number of users, the performance of contextual MAB can be significantly improved by facilitating c…

Cited by 0SourcePDFScholar
2025

PUMPS: Skeleton-Agnostic Point-based Universal Motion Pre-Training for Synthesis in Human Motion Tasks

ICCV 2025poster

Motion skeletons drive 3D character animation by transforming bone hierarchies, but differences in proportions or structure make motion data hard to transfer across skeletons, posing challenges for data-driven motion synthesis. Temporal Point Clouds (TPCs) offer an unstructured, cross-compatible mot…

2025

Provable Zero-Shot Generalization in Offline Reinforcement Learning

ICML 2025poster

In this work, we study offline reinforcement learning (RL) with zero-shot generalization property (ZSG), where the agent has access to an offline dataset including experiences from different environments, and the goal of the agent is to train a policy over the training environments which performs we…

Cited by 0SourcePDFScholar
2025

RI-MAE: Rotation-Invariant Masked AutoEncoders for Self-Supervised Point Cloud Representation Learning

AAAI 2025technical

Masked point modeling methods have recently achieved great success in self-supervised learning for point cloud data. However, these methods are sensitive to rotations and often exhibit sharp performance drops when encountering rotational variations. In this paper, we propose a novel Rotation-Invaria…

2025

ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks

EMNLP 2025

Multi-agent systems have emerged as a promising approach for enhancing the reasoning capabilities of large language models in complex problem-solving. However, current MAS frameworks are limited by poor flexibility and scalability, with underdeveloped optimization strategies. To address these challe

2025

RobAVA: A Large-scale Dataset and Baseline Towards Video based Robotic Arm Action Understanding

ICCV 2025poster

Understanding the behaviors of robotic arms is essential for various robotic applications such as logistics management, precision agriculture, and automated manufacturing. However, the lack of large-scale and diverse datasets significantly hinders progress in video-based robotic arm action understan…

2025

VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior

ICCV 2025accepted

Video diffusion models (VDMs) have advanced significantly in recent years, enabling the generation of highly realistic videos and drawing the attention of the community in their potential as world simulators. However, despite their capabilities, VDMs often fail to produce physically plausible videos…

2025

Variance-Dependent Regret Bounds for Nonstationary Linear Bandits

AISTATS 2025poster

We investigate the non-stationary stochastic linear bandit problem where the reward distribution evolves each round. Existing algorithms characterize the non-stationarity by the total variation budget $B_K$, which is the summation of the change of the consecutive feature vectors of the linear bandit…

Cited by 0SourceScholar
2025

When Graph Neural Networks Meet Dynamic Mode Decomposition

ICLR 2025poster

Graph Neural Networks (GNNs) have emerged as fundamental tools for a wide range of prediction tasks on graph-structured data. Recent studies have drawn analogies between GNN feature propagation and diffusion processes, which can be interpreted as dynamical systems. In this paper, we delve deeper int…

Cited by 0SourcePDFScholar
2024

AEDNet: Adaptive Embedding and Multiview-Aware Disentanglement for Point Cloud Completion

ECCV 2024poster

"Point cloud completion involves inferring missing parts of 3D objects from incomplete point cloud data. It requires a model that understands the global structure of the object and reconstructs local details. To this end, we propose a global perception and local attention network, termed AEDNet, for…

Cited by 1SourcePDFScholar
2024

Autoregressive Omni-Aware Outpainting for Open-Vocabulary 360-Degree Image Generation

AAAI 2024technical

A 360-degree (omni-directional) image provides an all-encompassing spherical view of a scene. Recently, there has been an increasing interest in synthesising 360-degree images from conventional narrow field of view (NFoV) images captured by digital cameras and smartphones, for providing immersive ex…

2024

Bridging the Gap: Sketch to Color Diffusion Model with Semantic Prompt Learning

ICASSP 2024accepted

Automatic anime sketch colorization aims to generate a color image from a sketch image, which is challenging due to limited structure and semantic understanding, leading to constrained style, and semantic color inconsistency. In this paper, we introduce a sketch to color diffusion model with semanti…

Cited by 0SourceScholar
2024

Combinatorial Multivariant Multi-Armed Bandits with Applications to Episodic Reinforcement Learning and Beyond

ICML 2024poster

We introduce a novel framework of combinatorial multi-armed bandits (CMAB) with multivariant and probabilistically triggering arms (CMAB-MT), where the outcome of each arm is a $d$-dimensional multivariant random variable and the feedback follows a general arm triggering process. Compared with exist…

Cited by 4SourcePDFScholar
2024

Exploring Self- and Cross-Triplet Correlations for Human-Object Interaction Detection

AAAI 2024technical

Human-Object Interaction (HOI) detection plays a vital role in scene understanding, which aims to predict the HOI triplet in the form of . Existing methods mainly extract multi-modal features (e.g., appearance, object semantics, human pose) and then fuse them together to directly predict HOI triplet…

Cited by 5SourcePDFScholar
2024

Federated Contextual Cascading Bandits with Asynchronous Communication and Heterogeneous Users

AAAI 2024technical

We study the problem of federated contextual combinatorial cascading bandits, where agents collaborate under the coordination of a central server to provide tailored recommendations to users. Existing works consider either a synchronous framework, necessitating full agent participation and global sy…

Cited by 6SourcePDFScholar
2024

Identity-Consistent Diffusion Network for Grading Knee Osteoarthritis Progression in Radiographic Imaging

ECCV 2024poster

"Knee osteoarthritis (KOA), a common form of arthritis that causes physical disability, has become increasingly prevalent in society. Employing computer-aided techniques to automatically assess the severity and progression of KOA can greatly benefit KOA treatment and disease management. Particularly…

Cited by 1SourcePDFScholar
2024

Intelligent Fish Detection System with Similarity-Aware Transformer

IROS 2024poster

Fish detection in water-land transfer has significantly contributed to the fishery. However, manual fish detection in crowd-collaboration performs inefficiently and expensively, involving insufficient accuracy. To further enhance the water-land transfer efficiency, improve detection accuracy, and re…

Cited by 0SourcecodeScholar
2024

MLPER: Multi-Level Prompts for Adaptively Enhancing Vision-Language Emotion Recognition

IROS 2024poster

In the field of robotics, vision-based Emotion Recognition (ER) has achieved significant progress, but it still faces the challenge of poor generalization ability under unconstrained conditions (e.g., occlusions and pose variations). In this work, we propose MLPER model, which introduces Vision-Lang…

Cited by 1SourceScholar
2024

Motion Keyframe Interpolation for Any Human Skeleton using Point Cloud-based Human Motion Data Homogenisation

ECCV 2024poster

"In the character animation field, modern supervised keyframe interpolation models have demonstrated exceptional performance in constructing natural human motions from sparse pose definitions. As supervised models, large motion datasets are necessary to facilitate the learning process; however, sinc…

Cited by 0SourcePDFScholar
2024

SurgicalSAM: Efficient Class Promptable Surgical Instrument Segmentation

AAAI 2024technical

The Segment Anything Model (SAM) is a powerful foundation model that has revolutionised image segmentation. To apply SAM to surgical instrument segmentation, a common approach is to locate precise points or boxes of instruments and then use them as prompts for SAM in a zero-shot manner. However, we…

2024

Terrain Diffusion Network: Climatic-Aware Terrain Generation with Geological Sketch Guidance

AAAI 2024technical

Sketch-based terrain generation seeks to create realistic landscapes for virtual environments in various applications such as computer games, animation and virtual reality. Recently, deep learning based terrain generation has emerged, notably the ones based on generative adversarial networks (GAN).…

2023

Adversarial Attacks on Online Learning to Rank with Click Feedback

NeurIPS 2023poster

Online learning to rank (OLTR) is a sequential decision-making problem where a learning agent selects an ordered list of items and receives feedback through user clicks. Although potential attacks against OLTR algorithms may cause serious losses in real-world applications, there is limited knowledge…

Cited by 6SourcePDFScholar
2023

Continuous Intermediate Token Learning With Implicit Motion Manifold for Keyframe Based Motion Interpolation

CVPR 2023poster

Deriving sophisticated 3D motions from sparse keyframes is a particularly challenging problem, due to continuity and exceptionally skeletal precision. The action features are often derivable accurately from the full series of keyframes, and thus, leveraging the global context with transformers has b…

2023

Efficient Explorative Key-Term Selection Strategies for Conversational Contextual Bandits

AAAI 2023technical

Conversational contextual bandits elicit user preferences by occasionally querying for explicit feedback on key-terms to accelerate learning. However, there are aspects of existing approaches which limit their performance. First, information gained from key-term-level conversations and arm-level re…

2023

LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark

NeurIPS 2023poster

Large language models have emerged as a promising approach towards achieving general-purpose AI agents. The thriving open-source LLM community has greatly accelerated the development of agents that support human-machine dialogue interaction through natural language processing. However, human interac…

2023

Multi-Scale Control Signal-Aware Transformer for Motion Synthesis without Phase

AAAI 2023technical

Synthesizing controllable motion for a character using deep learning has been a promising approach due to its potential to learn a compact model without laborious feature engineering. To produce dynamic motion from weak control signals such as desired paths, existing methods often require auxiliary…

Cited by 10SourcePDFScholar
2023

Online Clustering of Bandits with Misspecified User Models

NeurIPS 2023poster

The contextual linear bandit is an important online learning problem where given arm features, a learning agent selects an arm at each round to maximize the cumulative rewards in the long run. A line of works, called the clustering of bandits (CB), utilize the collaborative effect over user preferen…

Cited by 13SourcePDFScholar
2023

Online Corrupted User Detection and Regret Minimization

NeurIPS 2023poster

In real-world online web systems, multiple users usually arrive sequentially into the system. For applications like click fraud and fake reviews, some users can maliciously perform corrupted (disrupted) behaviors to trick the system. Therefore, it is crucial to design efficient online learning algor…

Cited by 9SourcePDFScholar
2023

Online Visual SLAM Adaptation against Catastrophic Forgetting with Cycle-Consistent Contrastive Learning

ICRA 2023poster

Visual SLAM (Simultaneous Localisation and Mapping) aims to simultaneously estimate camera poses and depth maps from navigation videos captured. While recent deep learning based methods have achieved great success on this task, they tend to work well on source domain data and suffer from performance…

Cited by 3SourceScholar
2023

VAPCNet: Viewpoint-Aware 3D Point Cloud Completion

ICCV 2023poster

Most existing learning-based 3D point cloud completion methods ignore the fact that the completion process is highly coupled with the viewpoint of a partial scan. However, the various viewpoints of incompletely scanned objects in real-world applications are normally unknown and directly estimating t…

Cited by 12PDFcodeScholar
2022

LiDARCap: Long-Range Marker-Less 3D Human Motion Capture With LiDAR Point Clouds

CVPR 2022poster

Existing motion capture datasets are largely short-range and cannot yet fit the need of long-range applications. We propose LiDARHuman26M, a new human motion capture dataset captured by LiDAR at a much longer range to overcome this limitation. Our dataset also includes the ground truth human motions…

Cited by 62PDFScholar
2022

OTExtSum: Extractive Text Summarisation with Optimal Transport

NAACL 2022findings

Extractive text summarisation aims to select salient sentences from a document to form a short yet informative summary. While learning-based methods have achieved promising results, they have several limitations, such as dependence on expensive training and lack of interpretability. Therefore, in th…

2020

Disentangling and Unifying Graph Convolutions for Skeleton-Based Action Recognition

CVPR 2020oral

Spatial-temporal graphs have been widely used by skeleton-based action recognition algorithms to model human action dynamics. To capture robust movement patterns from these graphs, long-range and multi-scale context aggregation and spatial-temporal dependency modeling are critical aspects of a power…

Cited by 1312PDFcodeScholar