← Search

Ning Yang

16 accepted papers

2026

Band Together: Untargeted Adversarial Training with Multimodal Coordination Against Evasion-Based Promotion Attacks

IJCAI 2026

Multimodal recommender systems exploit visual and textual signals to alleviate data sparsity, but this also makes them more vulnerable to evasion-based promotion attacks. Existing defenses are largely limited to single-modal settings and mainly focus on poisoning-based threats, leaving evasion-based

Cited by 0Scholar
2026

CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict Resolution

AAAI 2026technical

Multimodal Large Language Models are primarily trained and evaluated on aligned image-text pairs, which leaves their ability to detect and resolve real-world inconsistencies largely unexplored. In open-domain applications visual and textual cues often conflict, requiring models to perform structured

Cited by 0SourcePDFScholar
2026

Proactive Constrained Policy Optimization with Preemptive Penalty

AAAI 2026technical

Safe Reinforcement Learning (RL) often faces significant issues such as constraint violations and instability, necessitating the use of constrained policy optimization, which seeks optimal policies while ensuring adherence to specific constraints like safety. Typically, constrained optimization prob

Cited by 0SourcePDFScholar
2026

SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning

ICML 2026poster

Self-reflection is a powerful mechanism for credit assignment in human learning, converting sparse outcome feedback into actionable guidance. However, its potential for post-training Large Language Models (LLMs) remains underexplored. We propose Self-Reflective Policy Optimization (SRPO), a framewor…

Cited by 0SourceScholar
2026

TW-CRL: Time-Weighted Contrastive Reward Learning for Efficient Inverse Reinforcement Learning

AAAI 2026technical

Episodic tasks in Reinforcement Learning (RL) often pose challenges due to sparse reward signals and high-dimensional state spaces, which hinder efficient learning. Additionally, these tasks often feature hidden “trap states”—irreversible failures that prevent task completion but do not provide expl

Cited by 0SourcePDFScholar
2026

Token-Importance Guided Direct Preference Optimization

ICLR 2026oral

Aligning Large Language Models (LLMs) with human preferences is crucial for safe and effective AI interactions. While popular methods like Direct Preference Optimization (DPO) have simplified alignment, they remain sensitive to data noise and overlook the differential importance of individual tokens…

Cited by 0SourcecodeScholar
2025

Learn to Swim: Data-Driven LSTM Hydrodynamic Model for Quadruped Robot Gait Optimization

ICRA 2025

This paper presents a Long Short-Term Memory network-based Fluid Experiment Data-Driven model (FED-LSTM) for predicting unsteady, nonlinear hydrodynamic forces on the underwater quadruped robot we constructed. Trained on experimental data from leg force and body drag tests conducted in both a recirc

Cited by 2SourceScholar
2025

MuKA: Multimodal Knowledge Augmented Visual Information-Seeking

COLING 2025main

The visual information-seeking task aims to answer visual questions that require external knowledge, such as “On what date did this building officially open?”. Existing methods using retrieval-augmented generation framework primarily rely on textual knowledge bases to assist multimodal large languag…

2025

Tree-of-Code: A Self-Growing Tree Framework for End-to-End Code Generation and Execution in Complex Tasks

ACL 2025finding

Solving complex reasoning tasks is a key real-world application of agents. Thanks to the pretraining of Large Language Models (LLMs) on code data, recent approaches like CodeAct successfully use code as LLM agents’ action, achieving good results. However, CodeAct greedily generates the next action’s…

2025

Weakly Supervised Contrastive Adversarial Training for Learning Robust Features from Semi-supervised Data

CVPR 2025poster

Existing adversarial training (AT) methods often suffer from incomplete perturbation, meaning that not all non-robust features are perturbed when generating adversarial examples (AEs). This results in residual correlations between non-robust features and labels, leading to suboptimal learning of rob…

2024

Token-level Direct Preference Optimization

ICML 2024poster

Fine-tuning pre-trained Large Language Models (LLMs) is essential to align them with human values and intentions. This process often utilizes methods like pairwise comparisons and KL divergence against a reference LLM, focusing on the evaluation of full answers generated by the models. However, the…

2023

Robust Log-Based Anomaly Detection with Hierarchical Contrastive Learning

ICASSP 2023accepted

Logs are widely employed in modern systems to record critical information and serve as an important source for anomaly detection, which has attracted increasing research interests. However, logs usually suffer from perturbations and it makes the existing log-based anomaly detection methods unstable.…

Cited by 0SourceScholar
2021

Drumming Arm: an Upper-limb Prosthetic System to Restore Grip Control for a Transradial Amputee Drummer

ICRA 2021poster

This paper describes a quasi-passive transradial prosthesis designed to restore drumstick grip control for amputee drummers. A compact motor with low rotor resistance driven by Electromyography (EMG) is implemented in the prosthesis to support real-time drum performances. A variety of grip profiles…

Cited by 6SourceScholar