← Search

Xiaoling Zhou

12 accepted papers

2026

ASKD: Reinforcement Learning-Style Knowledge Distillation with Quality-Adaptive Skewness

AAAI 2026technical

Knowledge distillation (KD) is a widely adopted technique for transferring the capabilities of large teacher models to smaller student models, thereby significantly reducing inference costs and memory consumption. However, existing KD methods are all constrained by an inherent greedy optimization ob

Cited by 0SourcePDFScholar
2026

Mind Dreamer: Untethering Imagination via Active Counterfactual Reasoning on Latent Manifolds

ICML 2026poster

Model-Based Reinforcement Learning (MBRL) leverages latent imagination for sample efficiency, yet remains constrained by **Historical Tethering**: imagination is typically initialized from observed states. This creates a learning asymmetry, where the world model’s manifold discovery outpaces the pol…

Cited by 0SourceScholar
2026

Tailoring the Training: Difficulty-Aware Learning Strategy Allocation for Large Language Models

ICML 2026poster

Although reinforcement learning (RL) enhances the reasoning capabilities of large language models (LLMs), it is primarily learned from the model's self-generated distribution, limiting its ability to acquire reasoning skills beyond its initial knowledge. To overcome this, we propose a Difficulty-Awa…

Cited by 0SourceScholar
2025

All-Optical Nonlinear Diffractive Deep Network for Ultrafast Image Denoising

CVPR 2025highlight

Image denoising poses a significant challenge in image processing, aiming to remove noise and artifacts from input images. However, current denoising algorithms implemented on electronic chips frequently encounter latency issues and demand substantial computational resources. In this paper, we intro…

Cited by 0SourcePDFScholar
2025

Boosting Resilience of Large Language Models through Causality-Driven Robust Optimization

NeurIPS 2025poster

Large language models (LLMs) have achieved remarkable achievements across diverse applications; however, they remain plagued by spurious correlations and the generation of hallucinated content. Despite extensive efforts to enhance the resilience of LLMs, existing approaches either rely on indiscrimi…

Cited by 0SourceScholar
2025

HaDeMiF: Hallucination Detection and Mitigation in Large Language Models

ICLR 2025poster

The phenomenon of knowledge hallucinations has raised substantial concerns about the security and reliability of deployed large language models (LLMs). Current methods for detecting hallucinations primarily depend on manually designed individual metrics, such as prediction uncertainty and consistenc…

Cited by 0SourcePDFScholar
2024

Boosting Model Resilience via Implicit Adversarial Data Augmentation

IJCAI 2024poster

Data augmentation plays a pivotal role in enhancing and diversifying training data. Nonetheless, consistently improving model performance in varied learning scenarios, especially those with inherent data biases, remains challenging. To address this, we propose to augment the deep features of samples…

Cited by 1SourcePDFScholar
2024

Enhancing In-Context Learning via Implicit Demonstration Augmentation

ACL 2024long

The emergence of in-context learning (ICL) enables large pre-trained language models (PLMs) to make predictions for unseen inputs without updating parameters. Despite its potential, ICL’s effectiveness heavily relies on the quality, quantity, and permutation of demonstrations, commonly leading to su…

Cited by 2SourcePDFScholar
2022

Combined Fast Control of Drifting State and Trajectory Tracking for Autonomous Vehicles Based on MPC Controller

ICRA 2022poster

Slipping may cause a vehicle out of control with serious accident potential. However, a kind of car slipping named “drifting” can be seen in professional contests. So, it is reasonable to apply drift maneuvers in autonomous driving. This article proposes a controller for the particular driving skill…

Cited by 16SourceScholar