← Search

Rui Xie

30 accepted papers

2026

Language Drift in Multilingual Retrieval-Augmented Generation: Characterization and Decoding-Time Mitigation

AAAI 2026technical

Multilingual Retrieval-Augmented Generation (RAG) enables large language models (LLMs) to perform knowledge-intensive tasks in multilingual settings by leveraging retrieved documents as external evidence. However, when the retrieved evidence differs in language from the user query and in-context exe

Cited by 0SourcePDFScholar
2026

MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs

ICLR 2026poster

Despite advancements in Multimodal Large Language Models (MLLMs), their proficiency in fine-grained video motion understanding remains critically limited. They often lack inter-frame differencing and tend to average or ignore subtle visual cues. Furthermore, while visual prompting has shown potentia…

Cited by 0SourceScholar
2026

TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents

AAAI 2026technical

Building Graphical User Interface (GUI) agents is a promising research direction, which simulates human interaction with computers or mobile phones to perform diverse GUI tasks. However, a major challenge in developing generalized GUI agents is the lack of sufficient trajectory data across various o

Cited by 0SourcePDFScholar
2025

All-Optical Nonlinear Diffractive Deep Network for Ultrafast Image Denoising

CVPR 2025highlight

Image denoising poses a significant challenge in image processing, aiming to remove noise and artifacts from input images. However, current denoising algorithms implemented on electronic chips frequently encounter latency issues and demand substantial computational resources. In this paper, we intro…

Cited by 0SourcePDFScholar
2025

DLFR-Gen: Diffusion-based Video Generation with Dynamic Latent Frame Rate

ICCV 2025poster

Diffusion Transformer (DiT)-based generation models have achieved remarkable success in video generation. However, their inherent computational demands pose significant efficiency challenges. In this paper, we exploit the inherent temporal non-uniformity of real-world videos, and observe that videos…

Cited by 0SourcePDFScholar
2025

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption

CVPR 2025poster

Text-to-video generation has evolved rapidly in recent years, delivering remarkable results. Training typically relies on video-caption paired data, which plays a crucial role in enhancing generation performance. However, current video captions often suffer from insufficient details, hallucinations…

2025

Learnable Infinite Taylor Gaussian for Dynamic View Rendering

CVPR 2025poster

Capturing the temporal evolution of Gaussian properties such as position, rotation, and scale is a challenging task due to the vast number of time-varying parameters and the limited photometric data available, which generally results in convergence issues, making it difficult to find an optimal solu…

Cited by 0SourcePDFScholar
2025

OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

ICLR 2025poster

Text-to-video (T2V) generation has recently garnered significant attention thanks to the large multi-modality model Sora. However, T2V generation still faces two important challenges: 1) Lacking a precise open sourced high-quality dataset. The previously popular video datasets, e.g.WebVid-10M and Pa…

Cited by 62SourcePDFScholar
2025

STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution

ICCV 2025poster

Image diffusion models have been adapted for real-world video super-resolution to tackle over-smoothing issues in GAN-based methods. However, these models struggle to maintain temporal consistency, as they are trained on static images, limiting their ability to capture temporal dynamics effectively.…

Cited by 0SourcePDFScholar
2025

Spatially-variant Blur Degradation Model Based on Depth Estimation

ICASSP 2025accepted

It is well known that the number of aligned images in the single image super-resolution (SISR) models training is limited. Synthesizing data is an effective way to address this issue. However, many degradation models only consider using spatially-invariant blur kernels to blur high-resolution (HR) i…

Cited by 0SourceScholar
2025

TerraFusion: Semi-Supervised Vision-Proprioception Fusion for Robust Terrain Classification

RA-L 2025

Terrain classification is essential for traversability estimation and planning of unmanned ground vehicles (UGVs) in complex environments. Most existing approaches utilize fully supervised learning to classify terrains based on either exteroceptive or proprioceptive sensor modalities. However, visio

Cited by 0SourceScholar
2024

Boosting Model Resilience via Implicit Adversarial Data Augmentation

IJCAI 2024poster

Data augmentation plays a pivotal role in enhancing and diversifying training data. Nonetheless, consistently improving model performance in varied learning scenarios, especially those with inherent data biases, remains challenging. To address this, we propose to augment the deep features of samples…

Cited by 1SourcePDFScholar
2024

CR-UTP: Certified Robustness against Universal Text Perturbations on Large Language Models

ACL 2024findings

It is imperative to ensure the stability of every prediction made by a language model; that is, a language’s prediction should remain consistent despite minor input variations, like word substitutions. In this paper, we investigate the problem of certifying a language model’s robustness against Univ…

2024

ELTA: An Enhancer against Long-Tail for Aesthetics-oriented Models

ICML 2024poster

Real-world datasets often exhibit long-tailed distributions, compromising the generalization and fairness of learning-based models. This issue is particularly pronounced in Image Aesthetics Assessment (IAA) tasks, where such imbalance is difficult to mitigate due to a severe distribution mismatch be…

Cited by 4SourcePDFScholar
2024

Enhancing In-Context Learning via Implicit Demonstration Augmentation

ACL 2024long

The emergence of in-context learning (ICL) enables large pre-trained language models (PLMs) to make predictions for unseen inputs without updating parameters. Despite its potential, ICL’s effectiveness heavily relies on the quality, quantity, and permutation of demonstrations, commonly leading to su…

Cited by 2SourcePDFScholar
2024

PURE: Aligning LLM via Pluggable Query Reformulation for Enhanced Helpfulness

EMNLP 2024finding

Aligning large language models (LLMs) with human values and preferences is a significant challenge. Training-based methods, such as reinforcement learning from human feedback (RLHF) and direct preference optimization (DPO), require substantial resources and are impractical for API-based LLMs. Post-p…

Cited by 3SourcePDFScholar
2024

PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

ICLR 2024poster

Instruction tuning large language models (LLMs) remains a challenging task, owing to the complexity of hyperparameter selection and the difficulty involved in evaluating the tuned models. To determine the optimal hyperparameters, an automatic, robust, and reliable evaluation benchmark is essential.…

2023

Causality-aware Concept Extraction based on Knowledge-guided Prompting

ACL 2023long

Concepts benefit natural language understanding but are far from complete in existing knowledge graphs (KGs). Recently, pre-trained language models (PLMs) have been widely used in text-based concept extraction (CE). However, PLMs tend to mine the co-occurrence associations from massive corpus as pre…

2023

Exploiting Pseudo Image Captions for Multimodal Summarization

ACL 2023findings

Multimodal summarization with multimodal output (MSMO) faces a challenging semantic gap between visual and textual modalities due to the lack of reference images for training. Our pilot investigation indicates that image captions, which naturally connect texts and images, can significantly benefit M…

2023

Guide and Select: A Transformer-Based Multimodal Fusion Method for Points of Interest Description Generation

ICASSP 2023accepted

The task of Points of Interest (POI) description generation aims to generate an objective and informative description for a given POI based on POI-related information. High-quality descriptions can better guide users and improve the performance of POI-related recommendation systems. A practical POI…

Cited by 0SourceScholar
2022

An Effective and Efficient Entity Alignment Decoding Algorithm via Third-Order Tensor Isomorphism

ACL 2022long

Entity alignment (EA) aims to discover the equivalent entity pairs between KGs, which is a crucial step for integrating multi-source KGs.For a long time, most researchers have regarded EA as a pure graph representation learning task and focused on improving graph encoders while paying little attenti…

2022

Can Pre-trained Language Models Interpret Similes as Smart as Human?

ACL 2022long

Simile interpretation is a crucial task in natural language processing. Nowadays, pre-trained language models (PLMs) have achieved state-of-the-art performance on many tasks. However, it remains under-explored whether PLMs can interpret similes or not. In this paper, we investigate the ability of PL…

2022

Making Parameter-efficient Tuning More Efficient: A Unified Framework for Classification Tasks

COLING 2022main

Large pre-trained language models (PLMs) have demonstrated superior performance in industrial applications. Recent studies have explored parameter-efficient PLM tuning, which only updates a small amount of task-specific parameters while achieving both high efficiency and comparable performance again…

2022

PlugAT: A Plug and Play Module to Defend against Textual Adversarial Attack

COLING 2022main

Adversarial training, which minimizes the loss of adversarially perturbed examples, has received considerable attention. However, these methods require modifying all model parameters and optimizing the model from scratch, which is parameter inefficient and unfriendly to the already deployed models.…

2022

Rethinking Image Aesthetics Assessment: Models, Datasets and Benchmarks

IJCAI 2022poster

Challenges in image aesthetics assessment (IAA) arise from that images of different themes correspond to different evaluation criteria, and learning aesthetics directly from images while ignoring the impact of theme variations on human visual perception inhibits the further development of IAA; howev…

2022

Retrieval Enhanced Segment Generation Neural Network for Task-Oriented Dialogue Systems

ICASSP 2022accepted

For task-oriented dialogue systems, Natural Language Generation (NLG) is the last and vital step which aims at generating an appropriate response according to the dialogue act (DA). While end-to-end neural networks have achieved promising performances on this task, the existing models still struggle…

Cited by 0SourceScholar
2020

Graph Enhanced Dual Attention Network for Document-Level Relation Extraction

COLING 2020main

Document-level relation extraction requires inter-sentence reasoning capabilities to capture local and global contextual information for multiple relational facts. To improve inter-sentence reasoning, we propose to characterize the complex interaction between sentences and potential relation instanc…

Cited by 83SourcePDFScholar
2019

Online Decentralized Leverage Score Sampling for Streaming Multidimensional Time Series

AISTATS 2019poster

Estimating the dependence structure of multidimensional time series data in real-time is challenging. With large volumes of streaming data, the problem becomes more difficult when the multidimensional data are collected asynchronously across distributed nodes, which motivates us to sample representa…

Cited by 25SourcePDFScholar