← Search

Pengcheng Wang

15 accepted papers

2026

DADP: Domain Adaptive Diffusion Policy

ICML 2026poster

Learning domain adaptive policies that can generalize to unseen transition dynamics, remains a fundamental challenge in learning-based control. Substantial progress has been made through domain representation learning to capture domain-specific information, thus enabling domain-aware decision making…

Cited by 0SourceScholar
2026

Mean Flow Policy with Instantaneous Velocity Constraint for One-step Action Generation

ICLR 2026oral

Learning expressive and efficient policy functions is a promising direction in reinforcement learning (RL). While flow-based policies have recently proven effective in modeling complex action distributions with a fast deterministic sampling process, they still face a trade-off between expressiveness…

Cited by 0SourceScholar
2026

Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL

ICML 2026poster

Maximum entropy has become a mainstream off-policy reinforcement learning (RL) framework for balancing exploitation and exploration. However, two bottlenecks still limit further performance gains: (1) non-stationary Q-value estimation stemming from the joint injection of entropy and the concurrent u…

Cited by 0SourceScholar
2026

REAR: Test-time Preference Realignment through Reward Decomposition

ICML 2026poster

Aligning large language models (LLMs) with diverse user preferences is a critical yet challenging task. While post-training methods can adapt models to specific needs, they often require costly data curation and additional training. Test-time scaling (TTS) presents an efficient, training-free altern…

Cited by 0SourceScholar
2026

ReCast: Reliability-aware Codebook-assisted Lightweight Time Series Forecasting

AAAI 2026technical

Time series forecasting is crucial for applications in various domains. Conventional methods often rely on global decomposition into trend, seasonal, and residual components, which become ineffective for real-world series dominated by local, complex, and highly dynamic patterns. Moreover, the high m

Cited by 0SourcePDFScholar
2026

WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving

ICML 2026poster

Deploying multiple models within shared GPU clusters is a key strategy to improve resource efficiency in large language model (LLM) serving. Existing multi-LLM serving systems improve GPU utilization at the cost of degraded inference performance, particularly time-to-first-token (TTFT). We attribute…

Cited by 0SourceScholar
2025

Residual-MPPI: Online Policy Customization for Continuous Control

ICLR 2025poster

Policies developed through Reinforcement Learning (RL) and Imitation Learning (IL) have shown great potential in continuous control tasks, but real-world applications often require adapting trained policies to unforeseen requirements. While fine-tuning can address such needs, it typically requires a…

Cited by 2SourcePDFScholar
2024

Active Prompting with Chain-of-Thought for Large Language Models

ACL 2024long

The increasing scale of large language models (LLMs) brings emergent abilities to various complex tasks requiring reasoning, such as arithmetic and commonsense reasoning. It is known that the effective design of task-specific prompts is critical for LLMs’ ability to produce high-quality answers. In…

Cited by 212SourcePDFScholar
2024

CK12: A Rounded K12 Knowledge Graph Based Benchmark for Chinese Holistic Cognition Evaluation

AAAI 2024technical

New NLP benchmarks are urgently needed to align with the rapid development of large language models (LLMs). We present a meticulously designed evaluation benchmark that leverages the knowledge graph. This evaluation comprises 584 level-1 knowledge points and 1,989 level-2 knowledge points, thereby e…

2024

OpenT2T: An Open-Source Toolkit for Table-to-Text Generation

EMNLP 2024system demonstrations

Table data is pervasive in various industries, and its comprehension and manipulation demand significant time and effort for users seeking to extract relevant information. Consequently, an increasing number of studies have been directed towards table-to-text generation tasks. However, most existing…

2024

Towards Robust Evidence-Aware Fake News Detection via Improving Semantic Perception

COLING 2024main

Evidence-aware fake news detection aims to determine the veracity of a given news (i.e., claim) with external evidences. We find that existing methods lack sufficient semantic perception and are easily blinded by textual expressions. For example, they still make the same prediction after we flip the…

2023

Two-Stage Video De-Raining with Spatio-Temporal Fusion and Illumination-Invariant Detail Preservation

ICASSP 2023accepted

Video de-raining is an important yet highly challenging task in the field of computer vision. Though numerous video de-raining methods are developed with encouraging performance, two major challenges for video de-raining are still unsatisfactorily solved and need to be further investigated as follow…

Cited by 0SourceScholar
2021

Gate Trimming: One-Shot Channel Pruning for Efficient Convolutional Neural Networks

ICASSP 2021accepted

Channel pruning is a promising technique of model compression and acceleration because it reduces the space and time complexity of convolutional neural networks (CNNs) while maintaining their performance. In existing methods, channel pruning is performed by iterative optimization or training with sp…

Cited by 0SourceScholar
2018

Adaptive Sparse Array Reconfiguration based on Machine Learning Algorithms

ICASSP 2018accepted

The sparse array design for adaptive beamforming has been recently formulated into combinatorial antenna selection problems, which belong to notorious NP-hard problems. As the commonly deployed convex relaxation algorithms are susceptible to local optima, several trials with different initial points…

Cited by 0SourceScholar