← Search

Quan Sheng

2 accepted papers

2026

Reward-Preserving Counterfactual State Editing for Offline Reinforcement Learning

ICML 2026poster

Transformer sequence models such as Decision Transformer can learn strong offline policies from logged trajectories, but they can suffer from causal confusion: reliance on spurious correlations that predict reward in the data but do not reflect the true causal mechanisms of the environment. We propo…

Cited by 0SourceScholar
2024

Automatic, Meta and Human Evaluation for Multimodal Summarization with Multimodal Output

NAACL 2024long

Multimodal summarization with multimodal output (MSMO) has attracted increasing research interests recently as multimodal summary could provide more comprehensive information compared to text-only summary, effectively improving the user experience and satisfaction. As one of the most fundamental com…