← Search

Xiao Feng

8 accepted papers

2026

Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models

ICLR 2026poster

Although reinforcement learning with verifiable rewards (RLVR) shows promise in improving the reasoning ability of large language models (LLMs), the scaling up dilemma remains due to the reliance on human-annotated labels especially for complex tasks. Recent self-rewarding methods provide a label-fr…

Cited by 0SourcecodeScholar
2026

Landscape of Thoughts: Visualizing the Reasoning Process of Large Language Models

ICLR 2026poster

Numerous applications of large language models (LLMs) rely on their ability to perform step-by-step reasoning. However, the reasoning behavior of LLMs remains poorly understood, posing challenges to research, development, and safety. To address this gap, we introduce landscape of thoughts (LoT), the…

Cited by 0SourcecodeScholar
2026

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing

AAAI 2026technical

Large language models are extensively utilized in creative writing applications. Creative writing requires a balance between subjective writing quality (e.g., literariness and emotional expression) and objective constraint following (e.g., format requirements and word limits). Existing reinforcement

Cited by 0SourcePDFScholar
2025

From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?

ICML 2025poster

While existing benchmarks probe the reasoning abilities of large language models (LLMs) across diverse domains, they predominantly assess passive reasoning, providing models with all the information needed to reach a solution. By contrast, active reasoning—where an LLM must interact with external sy…

2025

PolarQuant: Leveraging Polar Transformation for Key Cache Quantization and Decoding Acceleration

NeurIPS 2025poster

The increasing demand for long-context generation has made the KV cache in large language models a bottleneck in memory consumption. Quantizing the cache to lower bit widths is an effective way to reduce memory costs; however, previous methods struggle with key cache quantization due to outliers, re…

Cited by 0SourcecodeScholar
2025

RecNet: Optimization for Dense Object Detection in Retail Scenarios Based on View Rectification

ICASSP 2025accepted

High-precision dense object detection in retail is crucial for automation, inventory management, and sales optimization. Our experiments revealed that detection models perform significantly better with frontal views than with oblique views, motivating the development of RecNet. RecNet utilizes a Rec…

Cited by 0SourceScholar