← Search

Rui Wu

13 accepted papers

2026

Bridging the Semantic Gap: Leveraging LLMs for Hierarchical Interest Evolution in Sequential Recommendation

IJCAI 2026

Accurate user behavior modeling is fundamental to the prediction of click-through rates (CTR) in industrial recommendation systems and online advertising. Traditional discriminative models, which rely on isolated ID features, struggle to capture the evolving nature of user intents across multiple ch

Cited by 0Scholar
2026

Dr.Occ: Depth- and Region-Guided 3D Occupancy from Surround-View Cameras for Autonomous Driving

CVPR 2026

3D semantic occupancy prediction is crucial for autonomous driving perception, offering comprehensive geometric scene understanding and semantic recognition. However, existing methods struggle with geometric misalignment in view transformation due to lack of pixel-level accurate depth estimation, an

Cited by 0SourceScholar
2026

ReFAct: Empowering Multimodal Web Agents with Visual and Context Focusing

CVPR 2026

Multimodal Web Search Agents demonstrate a practically valuable capability by fusing information from diverse modalities (e.g., text and vision), retrieved iteratively from the internet, to address complex user queries. However, the visual modality is prone to information overload, and the noise con

Cited by 0SourceScholar
2025

Adversarial Training for Graph Convolutional Networks: Stability and Generalization Analysis

IJCAI 2025

Recently, numerous methods have been proposed to enhance the robustness of the Graph Convolutional Networks (GCNs) for their vulnerability against adversarial attacks. Despite their empirical success, a significant gap remains in understanding GCNs' adversarial robustness from the theoretical perspe

Cited by 0SourcePDFScholar
2025

Generating Multimodal Driving Scenes via Next-Scene Prediction

CVPR 2025poster

Generative models in Autonomous Driving (AD) enable diverse scenario creation, yet existing methods fall short by only capturing a limited range of modalities, restricting the capability of generating controllable scenes for comprehensive evaluation of AD systems. In this paper, we introduce a multi…

2025

Human-Inspired Planning and Control of Shotcrete Robots based on Dynamical Systems Mapping

IROS 2025

Performing shotcrete operations at construction sites can be hazardous to humans and inefficient. Robots can offer a safer and more efficient alternative to assist in these tasks. We present a new planning strategy for shotcrete robots, including both the spraying and surface finishing phases, that

Cited by 1SourceScholar
2025

SUTBot: A Soft Umbrella-Like Tensegrity Robot With Elastic Struts for in-Pipe Locomotion

RA-L 2025

Compared with traditional in-pipe robots, tensegrity robots have exhibited many advantages such as light-weight, compliant, collapsible, low-cost, and rapidly manufacturable characteristics. However, published tensegrity in-pipe robots still have limited load capacity, because they rely on the stres

Cited by 6SourceScholar
2024

Biodegradable Gliding Paper Flyers Fabricated Through Inkjet Printing

IROS 2024poster

Seed-inspired minimalist microflyers are showing potential as dispersal platforms for sensor networks and seeds. Their function relies on the sheer number of low-cost flyers, inevitably raising concerns about the post-operation environmental impact. We propose a biodegradable paper glider platform f…

Cited by 0SourceScholar
2023

Generalization Bounds for Adversarial Metric Learning

IJCAI 2023poster

Recently, adversarial metric learning has been proposed to enhance the robustness of the learned distance metric against adversarial perturbations. Despite rapid progress in validating its effectiveness empirically, theoretical guarantees on adversarial robustness and generalization are far less und…

Cited by 1SourcePDFScholar
2021

You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection

NeurIPS 2021poster

Can Transformer perform $2\mathrm{D}$ object- and region-level recognition from a pure sequence-to-sequence perspective with minimal knowledge about the $2\mathrm{D}$ spatial structure? To answer this question, we present You Only Look at One Sequence (YOLOS), a series of object detection models bas…

2020

MCEN: Bridging Cross-Modal Gap between Cooking Recipes and Dish Images with Latent Variable Model

CVPR 2020poster

Nowadays, driven by the increasing concern on diet and health, food computing has attracted enormous attention from both industry and research community. One of the most popular research topics in this domain is Food Retrieval, due to its profound influence on health-oriented applications. In this p…

Cited by 71PDFScholar