← Search

Xiaorui Wang

24 accepted papers

2026

CombinationTS: A Modular Framework for Understanding Time-Series Forecasting Models

ICML 2026poster

Recent progress in time-series forecasting has led to rapidly increasing architectural complexity, yet many reported State-of-the-Art gains are statistically fragile or misattributed. We argue that progress requires a shift from model selection to modular attribution, identifying which components tr…

Cited by 0SourceScholar
2026

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

ICLR 2026poster

Deep Research Agents (DRAs) are emerging as one of the most practical classes of LLM-based agents. Given an open-ended research task, they find, analyze, and synthesize large numbers of online sources to produce a comprehensive report at the level of a research analyst. This can compress hours of ma…

Cited by 0SourcecodeScholar
2026

MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools

AAAI 2026technical

The Model Context Protocol (MCP) is rapidly emerging as a pivotal open standard, designed to enhance agent-tool integration and interoperability, and is positioned to unlock a new era of powerful, interconnected, and genuinely utilitarian agentic AI. However, despite MCP

Cited by 0SourcePDFScholar
2026

SDErasure: Concept-Specific Trajectory Shifting for Concept Erasure via Adaptive Diffusion Classifier

ICLR 2026poster

Concept erasure methods have proven effective in mitigating the potential for text‑to‑image diffusion models to produce harmful content. Nevertheless, prevailing methods based on post fine-tuning introduce substantial disruption to the original model’s parameter distribution and suffer from excessiv…

Cited by 0SourceScholar
2026

Test-Time Scaling with Reflective Generative Model

ICLR 2026poster

We introduce a new Reflective Generative Model (RGM), which obtains OpenAI o3-mini's performance via a novel Reflective Generative Form. This form focuses on high-quality reasoning trajectory selection and contains two novelties: 1) A unified interface for policy and process reward model: we share t…

Cited by 0SourcecodeScholar
2026

Thinker: Training LLMs in Hierarchical Thinking for Deep Search via Multi-Turn Interaction

AAAI 2026technical

Efficient retrieval of external knowledge bases and web pages is crucial for enhancing the reasoning abilities of LLMs. Previous works on training LLMs to leverage external retrievers for solving complex problems have predominantly employed end-to-end reinforcement learning. However, these approache

Cited by 0SourcePDFScholar
2025

A Portable Autonomous Underwater Vehicle With Multi-Thruster Propulsion: Design, Development, and Vision-Based Tracking Control

RA-L 2025

Autonomous underwater vehicles (AUVs) play a pivotal role in the exploration of marine resources. With the increasing complexity of underwater tasks, conventional torpedo-shaped AUVs exhibit significant limitations, particularly in complex and dynamic environments, due to restricted lateral translat

Cited by 6SourceScholar
2025

Adaptive Integral Sliding Mode Control for Attitude Tracking of Underwater Robots With Large Range Pitch Variations in Confined Spaces

RA-L 2025

Underwater robots play a crucial role in exploring aquatic environments. The ability to flexibly adjust their attitudes, especially the pitch, is essential for underwater robots to effectively accomplish tasks in confined spaces. However, the highly coupled six degrees of freedom dynamics resulting

Cited by 12SourceScholar
2025

Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach

EMNLP 2025

Creative writing is a key capability of Large Language Models (LLMs), with potential applications in literature, storytelling, and various creative domains. However, evaluating the creativity of machine-generated texts remains a significant challenge, as existing methods either rely on costly manual

Cited by 0SourcePDFScholar
2025

From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding

ACL 2025long

The pursuit of diverse, complex, and large-scale instruction data is crucial for automatically aligning large language models (LLMs). While there are methods capable of generating synthetic instructions at scale, they either suffer from limited grounding sources, leading to a narrow distribution, or…

2025

GRIP: A Graph-Based Reasoning Instruction Producer

NeurIPS 2025poster

Large-scale, high-quality data is essential for advancing the reasoning capabilities of large language models (LLMs). As publicly available Internet data becomes increasingly scarce, synthetic data has emerged as a crucial research direction. However, existing data synthesis methods often suffer fro…

Cited by 0SourceScholar
2025

IGD: Instructional Graphic Design with Multimodal Layer Generation

ICCV 2025poster

Graphic design visually conveys information and data by creating and combining text, images and graphics. Two-stage methods that rely primarily on layout generation lack creativity and intelligence, making graphic design still labor-intensive. Existing diffusion-based methods generate non-editable g…

2025

IterMeme: Expert-Guided Multimodal LLM for Interactive Meme Creation with Layout-Aware Generation

IJCAI 2025

Meme creation is a creative process that blends images and text. However, existing methods lack critical components, failing to support intent-driven caption-layout generation and personalized generation, making it difficult to generate high-quality memes. To address this limitation, we propose Iter

2025

MIRROR: Multi-agent Intra- and Inter-Reflection for Optimized Reasoning in Tool Learning

IJCAI 2025

Complex tasks involving tool integration pose significant challenges for Large Language Models (LLMs), leading to the emergence of multi-agent workflows as a promising solution. Reflection has emerged as an effective strategy for correcting erroneous trajectories in agentic workflows. However, exist

Cited by 0SourcePDFScholar
2025

SAGEPhos: Sage Bio-Coupled and Augmented Fusion for Phosphorylation Site Detection

ICLR 2025poster

Phosphorylation site prediction based on kinase-substrate interaction plays a vital role in understanding cellular signaling pathways and disease mechanisms. Computational methods for this task can be categorized into kinase-family-focused and individual kinase-targeted approaches. Individual kinase…

2023

Dynamic TF-TDNN: Dynamic Time Delay Neural Network Based on Temporal-Frequency Attention for Dialect Recognition

ICASSP 2023accepted

Dialect recognition aims to recognize dialect categories in utterances, which has been applied in many audio applications. Recently, various Time Delayed Neural Network (TDNN) based AI models are proposed to solve dialect recognition problems, such as D-TDNN, DMC-TDNN, and ECAPA-TDNN, however, most…

Cited by 0SourceScholar
2023

Filter Pruning Via Filters Similarity in Consecutive Layers

ICASSP 2023accepted

Filter pruning is widely adopted to compress and accelerate the Convolutional Neural Networks (CNNs), but most previous works ignore the relationship between filters and channels in different layers. Processing each layer independently fails to utilize the collaborative relationship across layers. I…

Cited by 0SourceScholar
2023

Improving Prosody for Cross-Speaker Style Transfer by Semi-Supervised Style Extractor and Hierarchical Modeling in Speech Synthesis

ICASSP 2023accepted

Cross-speaker style transfer in speech synthesis aims at transferring a style from source speaker to synthesized speech of a target speaker’s timbre. In most previous methods, the synthesized fine-grained prosody features often represent the source speaker’s average style, similar to the one-to-many…

Cited by 0SourceScholar
2023

NAS-DYMC: NAS-Based Dynamic Multi-Scale Convolutional Neural Network for Sound Event Detection

ICASSP 2023accepted

CNN+RNN models have become the mainstream approach for semi-supervised sound event detection, and the CNN part is mainly a stack of several 2D convolutional layers to capture the representations of the time-frequency features. However, conventional 2D convolution is of limited ability in capturing d…

Cited by 0SourceScholar
2022

EAD-Conformer: a Conformer-Based Encoder-Attention-Decoder-Network for Multi-Task Audio Source Separation

ICASSP 2022accepted

In this paper, we propose a Conformer-based network to improve the performance of multi-task audio source separation. This network, named EAD-Conformer, employs Conformer blocks to capture both local and global information, and an encoder-attention-decoder manner encourages the network to perform at…

Cited by 8SourceScholar
2022

K-Converter: An Unsupervised Singing Voice Conversion System

ICASSP 2022accepted

Singing voice conversion (SVC) converts a singer’s voice to another one’s voice while preserving the linguistic content. Recently, some SVC systems rely on supervised phonetic features extracted from pre-trained automatic speech recognition (ASR) models, increasing system complexity. Some end-toend…

Cited by 0SourceScholar
2022

Melons: Generating Melody With Long-Term Structure Using Transformers And Structure Graph

ICASSP 2022accepted

The creation of long melody sequences requires effective expression of coherent musical structure. However, there is no clear representation of musical structure. Recent works on music generation have suggested various approaches to deal with the structural information of music, but generating a ful…

Cited by 0SourceScholar
2021

One-Shot Voice Conversion Based on Speaker Aware Module

ICASSP 2021accepted

Voice conversion (VC) is a task to convert the voice of speech while preserving its linguistic content. Although several methods have been proposed to enable VC with non-parallel data, it is still difficult to model the voice without a great number of data or an adaptive process. In this paper, we p…

Cited by 0SourceScholar
2019

The Speechtransformer for Large-scale Mandarin Chinese Speech Recognition

ICASSP 2019accepted

Attention-based sequence-to-sequence architectures have made great progress in the speech recognition task. The SpeechTransformer, a no-recurrence encoder-decoder architecture, has shown promising results on small-scale speech recognition data sets in previous works. In this paper, we focus on a lar…

Cited by 0SourceScholar