← Search

Peiyao Wang

7 accepted papers

2026

Reinforcing Structured Chain-of-Thought for Video Understanding

CVPR 2026

Multi-modal Large Language Models (MLLMs) show promise in video understanding. However, their reasoning often suffers from thinking drift and weak temporal comprehension, even when enhanced by Reinforcement Learning (RL) techniques like Group Relative Policy Optimization (GRPO). Moreover, existing R

Cited by 0SourceScholar
2026

VUDG: A Dataset for Video Understanding Domain Generalization

ICLR 2026poster

Video understanding has made remarkable progress in recent years, largely driven by advances in deep models and the availability of large-scale annotated datasets. However, the robustness of these models to domain shifts encountered in real-world video applications remains a critical yet underexplor…

Cited by 0SourceScholar
2026

W-EDIT: A Wavelet-Based Frequency-Aware Framework for Text-Driven Image Editing

ICLR 2026poster

While recent advances in Diffusion Transformers (DiTs) have significantly advanced text-to-image generation, text-driven image editing remains challenging. Existing approaches either struggle to balance structural preservation with flexible modifications or require costly fine-tuning of large models…

Cited by 0SourceScholar
2024

Efficient Temporal Action Segmentation via Boundary-aware Query Voting

NeurIPS 2024poster

Although the performance of Temporal Action Segmentation (TAS) has been improved in recent years, achieving promising results often comes with a high computational cost due to dense inputs, complex model structures, and resource-intensive post-processing requirements. To improve the efficiency while…

2022

DialAug: Mixing up Dialogue Contexts in Contrastive Learning for Robust Conversational Modeling

COLING 2022main

Retrieval-based conversational systems learn to rank response candidates for a given dialogue context by computing the similarity between their vector representations. However, training on a single textual form of the multi-turn context limits the ability of a model to learn representations that gen…

2022

The Royalflush System of Speech Recognition for M2met Challenge

ICASSP 2022accepted

This paper describes our RoyalFlush system for the track of multi-speaker automatic speech recognition (ASR) in the M2MeT challenge. We adopted the serialized output training (SOT) based multi-speakers ASR system with large-scale simulation data. Firstly, we investigated a set of front-end methods,…

Cited by 0SourceScholar
2020

SIRI: Spatial Relation Induced Network For Spatial Description Resolution

NeurIPS 2020poster

Spatial Description Resolution, as a language-guided localization task, is proposed for target location in a panoramic street view, given corresponding language descriptions. Explicitly characterizing an object-level relationship while distilling spatial relationships are currently absent but crucia…