← Search

Kaiyan Zhao

14 accepted papers

2026

Covariance Volume Maximization for Embodied Latent Exploration in Deep Reinforcement Learning

ICML 2026poster

Efficient exploration remains a key challenge in deep reinforcement learning, especially for embodied agents operating in realistic environments with high-dimensional observations and complex dynamics. Recent latent exploration methods define bonuses in a learned latent space, but often struggle in …

Cited by 0SourceScholar
2026

DSAP: Enhancing Generalization in Goal-Conditioned Reinforcement Learning

AAAI 2026technical

Goal-conditioned Reinforcement Learning (RL) is a promising direction for training agents capable of tackling a variety of tasks. However, generalizing to new goals in different environments remains a central challenge for goal-conditioned RL agents. Existing methods often rely on state abstraction,

Cited by 0SourcePDFScholar
2026

Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training

ICML 2026poster

Determining an effective data mixture is a key factor in Large Language Model (LLM) pre-training, where models must balance general competence with proficiency on hard tasks such as math and code. However, identifying an optimal mixture remains an open challenge, as existing approaches either rely o…

Cited by 0SourceScholar
2026

E^2DT: Efficient and Effective Decision Transformer with Experience-Aware Sampling for Robotic Manipulation

ICRA 2026poster

In reinforcement learning (RL) for robotic manipulation, the Decision Transformer (DT) has emerged as an effective framework for addressing long-horizon tasks. However, DT’s performance depends heavily on the coverage of collected experiences. Without an active exploration mechanism, standard DT rel…

Cited by 0Scholar
2026

Explore to Learn: Latent Exploration Through Disentangled Synergy Patterns for Reinforcement Learning in Overactuated Control

AAAI 2026technical

Control in high-dimensional action spaces remains a fundamental challenge in reinforcement learning (RL), primarily due to inefficient exploration of the action space. While recent methods attempt to guide exploration, they often fall short of achieving the agility and coordination exhibited in biol

Cited by 0SourcePDFScholar
2026

Heuristic Self-Paced Learning for Domain Adaptive Semantic Segmentation under Adverse Conditions

CVPR 2026

The learning order of semantic classes significantly impacts unsupervised domain adaptation for semantic segmentation, especially under adverse weather conditions. Most existing curricula rely on handcrafted heuristics (e.g., fixed uncertainty metrics) and follow a static schedule, which fails to ad

Cited by 0SourceScholar
2026

Latent State-Predictive Exploration for Deep Reinforcement Learning

AAAI 2026technical

Reinforcement learning (RL) has achieved promising results in continuous control tasks, where efficient exploration of the state space is crucial for success. However, many recent RL approaches still struggle with sample inefficiency and insufficient exploration for long-horizon tasks, particularly

Cited by 0SourcePDFScholar
2026

RGMP: Recurrent Geometric-prior Multimodal Policy for Generalizable Humanoid Robot Manipulation

AAAI 2026technical

Humanoid robots exhibit significant potential in executing diverse human-level skills. However, current research predominantly relies on data-driven approaches that necessitate extensive training datasets to achieve robust multimodal decision-making capabilities and generalizable visuomotor control.

Cited by 0SourcePDFScholar
2025

BILE: An Effective Behavior-based Latent Exploration Scheme for Deep Reinforcement Learning

IJCAI 2025

Efficient exploration of state spaces is critical for the success of deep reinforcement learning (RL). While many methods leverage exploration bonuses to encourage exploration instead of relying solely on extrinsic rewards, these bonus-based approaches often face challenges with learning efficiency

Cited by 0SourcePDFScholar
2025

Efficient Diversity-based Experience Replay for Deep Reinforcement Learning

IJCAI 2025

Experience replay is widely used to improve learning efficiency in reinforcement learning by leveraging past experiences. However, existing experience replay methods, whether based on uniform or prioritized sampling, often suffer from low efficiency, particularly in real-world scenarios with high-di

Cited by 0SourcePDFScholar
2025

HeMoRa: Unsupervised Heuristic Consensus Sampling for Robust Point Cloud Registration

CVPR 2025poster

Heuristic information for consensus set sampling is essential for correspondence-based point cloud registration, but existing approaches typically rely on supervised learning or expert-driven parameter tuning. In this work, we propose HeMoRa, a new unsupervised framework that trains a Heuristic info…

2024

Enhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word Alignment

NAACL 2024findings

The field of cross-lingual sentence embeddings has recently experienced significant advancements, but research concerning low-resource languages has lagged due to the scarcity of parallel corpora. This paper shows that cross-lingual word representation in low-resource languages is notably under-alig…

2024

Rethinking Exploration in Reinforcement Learning with Effective Metric-Based Exploration Bonus

NeurIPS 2024spotlight

Enhancing exploration in reinforcement learning (RL) through the incorporation of intrinsic rewards, specifically by leveraging *state discrepancy* measures within various metric spaces as exploration bonuses, has emerged as a prevalent strategy to encourage agents to visit novel states. The critica…

Cited by 0SourcePDFScholar