← Search

Kai Yang

44 accepted papers

2026

CE-Nav: Flow-Guided Reinforcement Refinement for Cross-Embodiment Local Navigation

ICLR 2026poster

Generalizing local navigation policies across diverse robot morphologies is a critical challenge. Progress is often hindered by the need for costly and embodiment-specific data, the tight coupling of planning and control, and the "disastrous averaging" problem where deterministic models fail to capt…

Cited by 0SourcecodeScholar
2026

Debiased Model-based Representations for Sample-efficient Continuous Control

ICML 2026poster

Model-based representations recently stand out as a promising framework that embeds latent dynamics information into the representations for downstream off-policy actor-critic learning. It implicitly combines the advantages of both model-free and model-based approaches while avoiding the training co…

Cited by 0SourceScholar
2026

Physics-Inspired All-Pair Interaction Learning for 3D Dynamics Modeling

ICLR 2026poster

Modeling 3D dynamics is a fundamental problem in multi-body systems across scientific and engineering domains and has important practical implications in trajectory prediction and simulation. While recent GNN-based approaches have achieved strong performance by enforcing geometric symmetries, encodi…

Cited by 0SourcecodeScholar
2026

RoS-Guard: Robust and Scalable Online Change Detection with Delay-Optimal Guarantees

AAAI 2026technical

Online change detection (OCD) aims to rapidly identify change points in streaming data and is critical in applications such as power system monitoring, wireless network sensing, and financial anomaly detection. Existing OCD methods typically assume precise system knowledge, which is unrealistic due

Cited by 0SourcePDFScholar
2026

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models

ICML 2026poster

Reinforcement learning enhances the reasoning capabilities of large language models but often involves high computational costs due to rollout-intensive optimization. Online prompt selection presents a plausible solution by prioritizing informative prompts to improve training efficiency. However, cu…

Cited by 0SourceScholar
2026

Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners

ICLR 2026poster

Reinforcement Learning with Verifiable Reward (RLVR) effectively solves complex tasks but demands extremely long context lengths during training, leading to substantial computational costs. While multi-stage training can partially mitigate this, starting with overly short contexts often causes irrev…

Cited by 0SourcecodeScholar
2026

When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression

ICML 2026poster

Reasoning hallucinations in large language models (LLMs) often appear as fluent yet unsupported conclusions that violate either the given context or underlying factual knowledge. Although such failures are widely observed, the mechanisms by which decoder-only Transformers produce them remain poorly …

Cited by 0SourceScholar
2025

Contextures: Representations from Contexts

ICML 2025poster

Despite the empirical success of foundation models, we do not have a systematic characterization of the representations that these models learn. In this paper, we establish the contexture theory. It shows that a large class of representation learning methods can be characterized as learning from th…

Cited by 0SourcePDFScholar
2025

DTZO: Distributed Trilevel Zeroth Order Learning with Provable Non-Asymptotic Convergence

ICML 2025poster

Trilevel learning (TLL) with zeroth order constraints is a fundamental problem in machine learning, arising in scenarios where gradient information is inaccessible due to data privacy or model opacity, such as in federated learning, healthcare, and financial systems. These problems are notoriously d…

Cited by 0SourcePDFScholar
2025

From Sequence to Structure: Uncovering Substructure Reasoning in Transformers

NeurIPS 2025poster

Recent studies suggest that large language models (LLMs) possess the capability to solve graph reasoning tasks. Notably, even when graph structures are embedded within textual descriptions, LLMs can still effectively answer related questions. This raises a fundamental question: How can a decoder-onl…

Cited by 0SourceScholar
2025

How Numerical Precision Affects Arithmetical Reasoning Capabilities of LLMs

ACL 2025finding

Despite the remarkable success of transformer-based large language models (LLMs) across various domains, understanding and enhancing their mathematical capabilities remains a significant challenge. In this paper, we conduct a rigorous theoretical analysis of LLMs’ mathematical abilities, with a spec…

Cited by 0SourcePDFScholar
2025

Novelty-Guided Data Reuse for Efficient and Diversified Multi-Agent Reinforcement Learning

AAAI 2025technical

Recently, deep Multi-Agent Reinforcement Learning (MARL) has demonstrated its potential to tackle complex cooperative tasks, pushing the boundaries of AI in collaborative environments. However, the efficiency of these systems is often compromised by inadequate sample utilization and a lack of divers…

2024

BATON: Aligning Text-to-Audio Model Using Human Preference Feedback

IJCAI 2024poster

With the development of AI-Generated Content (AIGC), text-to-audio models are gaining widespread attention. However, it is challenging for these models to generate audio aligned with human preference due to the inherent information density of natural language and limited model understanding ability.…

2024

Design and Trajectory Tracking Control of CuRobot: A Cubic Reversible Robot

RA-L 2024

In field environments, numerous robots necessitate manual intervention for restoration of functionality post a turnover, resulting in diminished operational efficiency. This study presents an innovative design solution for a reversible omnidirectional mobile robot denoted as CuRobot, featuring a cub

Cited by 1SourceScholar
2024

Do Efficient Transformers Really Save Computation?

ICML 2024poster

As transformer-based language models are trained on increasingly large datasets and with vast numbers of parameters, finding more efficient alternatives to the standard Transformer has become very valuable. While many efficient Transformers and Transformer alternatives have been proposed, none provi…

Cited by 16SourcePDFScholar
2024

Exploration and Anti-Exploration with Distributional Random Network Distillation

ICML 2024poster

Exploration remains a critical issue in deep reinforcement learning for an agent to attain high returns in unknown environments. Although the prevailing exploration Random Network Distillation (RND) algorithm has been demonstrated to be effective in numerous environments, it often needs more discrim…

2024

Exploring Self-Explainable Street-Level IP Geolocation with Graph Information Bottleneck

ICASSP 2024accepted

Accurate IP geolocation is crucial for location-aware applications. While recent advances in router-centric IP graph methods have garnered attention, they face two persistent challenges: (1) the sparsity problem of IP graphs in rural areas and (2) the limited explainability of current IP geolocation…

Cited by 0SourceScholar
2024

Improving IP Geolocation With Target-Centric IP Graph (Student Abstract)

AAAI 2024technical

Accurate IP geolocation is indispensable for location-aware applications. While recent advances based on router-centric IP graphs are considered cutting-edge, one challenge remain: the prevalence of sparse IP graphs (14.24% with fewer than 10 nodes, 9.73% isolated) limits graph learning. To mitigate…

Cited by 0SourcePDFScholar
2024

Interpreting Temporal Knowledge Graph Reasoning (Student Abstract)

AAAI 2024technical

Temporal knowledge graph reasoning is an essential task that holds immense value in diverse real-world applications. Existing studies mainly focus on leveraging structural and sequential dependencies, excelling in tasks like entity and link prediction. However, they confront a notable interpretabili…

Cited by 2SourcePDFScholar
2024

Provably Convergent Federated Trilevel Learning

AAAI 2024technical

Trilevel learning, also called trilevel optimization (TLO), has been recognized as a powerful modelling tool for hierarchical decision process and widely applied in many machine learning applications, such as robust neural architecture search, hyperparameter optimization, and domain adaptation. Tack…

Cited by 6SourcePDFScholar
2024

Robust Beamforming for Downlink Multi-Cell Systems: A Bilevel Optimization Perspective

AAAI 2024technical

Utilization of inter-base station cooperation for information processing has shown great potential in enhancing the overall quality of communication services (QoS) in wireless communication networks. Nevertheless, such cooperations require the knowledge of channel state information (CSI) at base sta…

Cited by 3SourcePDFScholar
2024

SGM: A Dataset for 3D Garment Reconstruction from Single Hand-Drawn Sketch

ICASSP 2024accepted

High-fidelity garment reconstruction is essential for various applications such as garment design and virtual try-on. While image-based reconstruction methods have made significant progress with deep generative models, generating 3D models from hand-drawn sketches to meet design intentions remains c…

Cited by 0SourceScholar
2024

Tri-Level Navigator: LLM-Empowered Tri-Level Learning for Time Series OOD Generalization

NeurIPS 2024poster

Out-of-Distribution (OOD) generalization in machine learning is a burgeoning area of study. Its primary goal is to enhance the adaptability and resilience of machine learning models when faced with new, unseen, and potentially adversarial data that significantly diverges from their original training…

Cited by 4SourcePDFScholar
2024

Triadic-OCD: Asynchronous Online Change Detection with Provable Robustness, Optimality, and Convergence

ICML 2024poster

The primary goal of online change detection (OCD) is to promptly identify changes in the data stream. OCD problem find a wide variety of applications in diverse areas, e.g., security detection in smart grids and intrusion detection in communication networks. Prior research usually assumes precise kn…

Cited by 0SourcePDFScholar
2024

Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation

ICML 2024poster

In this work, we leverage the intrinsic segmentation of language sequences and design a new positional encoding method called Bilevel Positional Encoding (BiPE). For each position, our BiPE blends an intra-segment encoding and an inter-segment encoding. The intra-segment encoding identifies the loca…

2024

Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model

CVPR 2024poster

Using reinforcement learning with human feedback (RLHF) has shown significant promise in fine-tuning diffusion models. Previous methods start by training a reward model that aligns with human preferences then leverage RL techniques to fine-tune the underlying models. However crafting an efficient re…

2023

ASM: Adaptive Skinning Model for High-Quality 3D Face Modeling

ICCV 2023poster

The research fields of parametric face model and 3D face reconstruction have been extensively studied. However, a critical question remains unanswered: how to tailor the face model for specific reconstruction settings. We argue that reconstruction with multi-view uncalibrated images demands a new mo…

Cited by 6PDFScholar
2023

Asynchronous Distributed Bilevel Optimization

ICLR 2023poster

Bilevel optimization plays an essential role in many machine learning tasks, ranging from hyperparameter optimization to meta-learning. Existing studies on bilevel optimization, however, focus on either centralized or synchronous distributed setting. The centralized bilevel optimization approaches r…

2023

KerPrint: Local-Global Knowledge Graph Enhanced Diagnosis Prediction for Retrospective and Prospective Interpretations

AAAI 2023technical

While recent developments of deep learning models have led to record-breaking achievements in many areas, the lack of sufficient interpretation remains a problem for many specific applications, such as the diagnosis prediction task in healthcare. The previous knowledge graph(KG) enhanced approaches…

2023

VecoCare: Visit Sequences-Clinical Notes Joint Learning for Diagnosis Prediction in Healthcare Data

IJCAI 2023poster

Due to the insufficiency of electronic health records (EHR) data utilized in practical diagnosis prediction scenarios, most works are devoted to learning powerful patient representations either from structured EHR data (e.g., temporal medical events, lab test results, etc.) or unstructured data (e.g…

2022

A Unified Model for Multi-class Anomaly Detection

NeurIPS 2022accept

Despite the rapid advance of unsupervised anomaly detection, existing methods require to train separate models for different objects. In this work, we present UniAD that accomplishes anomaly detection for multiple classes with a unified framework. Under such a challenging setting, popular reconstruc…

2020

Learning Affordance Space in Physical World for Vision-based Robotic Object Manipulation

ICRA 2020poster

What is a proper representation for objects in manipulation? What would human try to perceive when manipulating a new object in a new environment? In fact, instead of focusing on the texture and illumination, human can infer the "affordance" [36] of the objects from vision. Here "affordance" describ…

Cited by 23SourceScholar
2020

Signal Sensing and Reconstruction Paradigms for a Novel Multi-Source Static Computed Tomography System

ICASSP 2020accepted

Conventional Computed Tomography (CT) systems use a single X-ray source and an arc of detectors mounted on a rotating gantry to acquire a set of projection data. Novel CT systems are now being pioneered in which a complete ring of distributed X-ray sources and detectors are electronically turned on…

Cited by 0SourceScholar
2019

A new robot skating on water surface intimating water striders based on flexible driving mechanism

ICRA 2019poster

The amazing ability of water striders on water surface has attracted many scholars. Especially the flexible driving mechanism enable the driving legs conform to the deformation of the water surface, which effectively improving water striders’ floating ability and stability. However, the current rese…

Cited by 8SourceScholar
2018

Deep Stock Representation Learning: From Candlestick Charts to Investment Decisions

ICASSP 2018accepted

We propose a novel investment decision strategy (IDS) based on deep learning. The performance of many IDSs is affected by stock similarity. Most existing stock similarity measurements have the problems: (a) The linear nature of many measurements cannot capture nonlinear stock dynamics; (b) The estim…

Cited by 0SourceScholar
2018

Fusing Object Context to Detect Functional Area for Cognitive Robots

ICRA 2018poster

A cognitive robot usually needs to perform multiple tasks in practice and needs to locate the desired area for each task. Since deep learning has achieved substantial progress in image recognition, to solve this area detection problem, it is straightforward to label a functional area (affordance) im…

Cited by 0SourceScholar