← Search

Wei Yao

21 accepted papers

2026

FAB: A First-Order AB-based Gradient Algorithm for Distributed Bilevel Optimization over Time-Varying Directed Graphs

ICML 2026poster

Distributed optimization over time-varying directed graphs has shown promising performance in addressing challenges posed by complex communication constraints in real-world scenarios. In many practical settings, however, the direct application of distributed optimization algorithms encounters additi…

Cited by 0SourceScholar
2026

HOSIG: Full-Body Human-Object-Scene Interaction Generation with Hierarchical Scene Perception

AAAI 2026technical

Generating high-fidelity full-body human interactions with dynamic objects and static scenes remains a critical challenge in computer graphics and animation. Existing methods for human-object interaction often neglect scene context, leading to implausible penetrations, while human-scene interaction

Cited by 0SourcePDFScholar
2026

Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn LLM Agents

ICLR 2026poster

Large language model (LLM)–based agents are increasingly trained with reinforcement learning (RL) to enhance their ability to interact with external environments through tool use, particularly in search-based settings that require multi-turn reasoning and knowledge acquisition. However, existing app…

Cited by 0SourcecodeScholar
2025

An Online System Identification Algorithm for Spherical Robot Using the Koopman Theory

RA-L 2025

This letter proposes a novel linear online identification framework for the spherical robot to address the modeling difficulties posed by nonlinearity and time-varying characteristics. Firstly, the Koopman theory is applied to the spherical robot to build a linear model to approximate the nonlineari

Cited by 4SourceScholar
2025

Bilevel Optimization for Adversarial Learning Problems: Sharpness, Generation, and Beyond

NeurIPS 2025poster

Adversarial learning is a widely used paradigm in machine learning, often formulated as a min-max optimization problem where the inner maximization imposes adversarial constraints to guide the outer learner toward more robust solutions. This framework underlies methods such as Sharpness-Aware Minimi…

Cited by 0SourceScholar
2025

DeepWell-Adol: A Scalable Expert-Based Dialogue Corpus for Adolescent Positive Mental Health and Wellbeing Promotion

EMNLP 2025

Promoting positive mental health and well-being, especially in adolescents, is a critical yet underexplored area in natural language processing (NLP). Most existing NLP research focuses on clinical therapy or psychological counseling for the general population, which does not adequately address the

Cited by 0SourcePDFScholar
2025

Efficient Curvature-Aware Hypergradient Approximation for Bilevel Optimization

ICML 2025poster

Bilevel optimization is a powerful tool for many machine learning problems, such as hyperparameter optimization and meta-learning. Estimating hypergradients (also known as implicit gradients) is crucial for developing gradient-based methods for bilevel optimization. In this work, we propose a comput…

Cited by 0SourcePDFScholar
2025

Overcoming Lower-Level Constraints in Bilevel Optimization: A Novel Approach with Regularized Gap Functions

ICLR 2025poster

Constrained bilevel optimization tackles nested structures present in constrained learning tasks like constrained meta-learning, adversarial learning, and distributed bilevel optimization. However, existing bilevel optimization methods mostly are typically restricted to specific constraint settings…

2025

Revisiting Weak-to-Strong Generalization in Theory and Practice: Reverse KL vs. Forward KL

ACL 2025finding

As large language models advance toward superhuman performance, ensuring their alignment with human values and abilities grows increasingly complex. Weak-to-strong generalization offers a promising approach by leveraging predictions from weaker models to guide stronger systems, but its effectiveness…

Cited by 0SourcePDFScholar
2025

Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization

ICLR 2025poster

Superalignment, where humans act as weak supervisors for superhuman models, has become a crucial problem with the rapid development of Large Language Models (LLMs). Recent work has preliminarily studied this problem by using weak models to supervise strong models, and discovered that weakly supervis…

2024

Constrained Bi-Level Optimization: Proximal Lagrangian Value Function Approach and Hessian-free Algorithm

ICLR 2024spotlight

This paper presents a new approach and algorithm for solving a class of constrained Bi-Level Optimization (BLO) problems in which the lower-level problem involves constraints coupling both upper-level and lower-level variables. Such problems have recently gained significant attention due to their br…

Cited by 17SourcePDFScholar
2024

Moreau Envelope for Nonconvex Bi-Level Optimization: A Single-Loop and Hessian-Free Solution Strategy

ICML 2024poster

This work focuses on addressing two major challenges in the context of large-scale nonconvex Bi-Level Optimization (BLO) problems, which are increasingly applied in machine learning due to their ability to model nested structures. These challenges involve ensuring computational efficiency and provid…

Cited by 9SourcePDFScholar
2024

SPABA: A Single-Loop and Probabilistic Stochastic Bilevel Algorithm Achieving Optimal Sample Complexity

ICML 2024poster

While stochastic bilevel optimization methods have been extensively studied for addressing large-scale nested optimization problems in machine learning, it remains an open question whether the optimal complexity bounds for solving bilevel optimization are the same as those in single-level optimizati…

Cited by 4SourcePDFScholar
2024

Towards Tracing Trustworthiness Dynamics: Revisiting Pre-training Period of Large Language Models

ACL 2024findings

Ensuring the trustworthiness of large language models (LLMs) is crucial. Most studies concentrate on fully pre-trained LLMs to better understand and improve LLMs’ trustworthiness. In this paper, to reveal the untapped potential of pre-training, we pioneer the exploration of LLMs’ trustworthiness dur…

2023

Averaged Method of Multipliers for Bi-Level Optimization without Lower-Level Strong Convexity

ICML 2023poster

Gradient methods have become mainstream techniques for Bi-Level Optimization (BLO) in learning fields. The validity of existing works heavily rely on either a restrictive Lower- Level Strong Convexity (LLSC) condition or on solving a series of approximation subproblems with high accuracy or both. In…

2023

Fair Scratch Tickets: Finding Fair Sparse Networks Without Weight Training

CVPR 2023poster

Recent studies suggest that computer vision models come at the risk of compromising fairness. There are extensive works to alleviate unfairness in computer vision using pre-processing, in-processing, and post-processing methods. In this paper, we lead a novel fairness-aware learning paradigm for in-…

2022

Task Autonomous Medical Robot for Both Incision Stapling and Staples Removal

RA-L 2022

Surgical incision is a pervasive procedure in medical environments. Stapling is an incision closure method that is comparable to stitching. While incision closure is performed immediately after a surgery, staples are removed several weeks after the closure. The workload of surgeons and the probabili

Cited by 20SourceScholar