← Search

Rui Yu

22 accepted papers

2026

BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining

ICML 2026poster

Effective data selection is essential for pretraining large language models (LLMs), enhancing efficiency and improving generalization to downstream tasks. However, existing approaches often require leveraging external pretrained models, making it difficult to disentangle the effects of data selectio…

Cited by 0SourceScholar
2026

Dynamic Momentum Recalibration in Online Gradient Learning

CVPR 2026

Stochastic Gradient Descent (SGD) and its momentum variants form the backbone of deep learning optimization, yet the underlying dynamics of their gradient behavior remain insufficiently understood. In this work, we reinterpret gradient updates through the lens of signal processing and reveal that fi

Cited by 0SourcecodeScholar
2026

From Detection to Association: Learning Discriminative Object Embeddings for Multi-Object Tracking

CVPR 2026

End-to-end multi-object tracking (MOT) methods have recently achieved remarkable progress by unifying detection and association within a single framework. Despite their strong detection performance, these methods suffer from relatively low association accuracy. Through detailed analysis, we observe

Cited by 0SourcecodeScholar
2025

Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval

ICCV 2025poster

Open-set 3D object retrieval (3DOR) is an emerging task aiming to retrieve 3D objects of unseen categories beyond the training set. Existing methods typically utilize all modalities (i.e., voxels, point clouds, multi-view images) and train specific backbones before fusion. However, they still strugg…

2025

DistillDrive: End-to-End Multi-Mode Autonomous Driving Distillation by Isomorphic Hetero-Source Planning Model

ICCV 2025poster

End-to-end autonomous driving has been recently seen rapid development, exerting a profound influence on both industry and academia. However, the existing work places excessive focus on ego-vehicle status as their sole learning objectives and lacks of planning-oriented understanding, which limits th…

2025

FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making

ICML 2025poster

Foundation Models (FMs) and World Models (WMs) offer complementary strengths in task generalization at different levels. In this work, we propose FOUNDER, a framework that integrates the generalizable knowledge embedded in FMs with the dynamic modeling capabilities of WMs to enable open-ended task s…

Cited by 0SourcePDFScholar
2025

JailBound: Jailbreaking Internal Safety Boundaries of Vision-Language Models

NeurIPS 2025poster

Vision-Language Models (VLMs) exhibit impressive performance, yet the integration of powerful vision encoders has significantly broadened their attack surface, rendering them increasingly susceptible to jailbreak attacks. However, lacking well-defined attack objectives, existing jailbreak methods of…

Cited by 0SourceScholar
2025

Reward Models in Deep Reinforcement Learning: A Survey

IJCAI 2025

In reinforcement learning (RL), agents continually interact with the environment and use the feedback to refine their behavior. To guide policy optimization, reward models are introduced as proxies of the desired objectives, such that when the agent maximizes the accumulated reward, it also fulfills

Cited by 0SourcePDFScholar
2025

Towards In-the-wild 3D Plane Reconstruction from a Single Image

CVPR 2025highlight

3D plane reconstruction from a single image is a crucial yet challenging topic in 3D computer vision. Previous state-of-the-art (SOTA) methods have focused on training their system on a single dataset from either indoor or outdoor domain, limiting their generalizability across diverse testing data.…

2023

Do We Need a New Foundation to Use Deep Learning to Monitor Weld Penetration?

RA-L 2023

Deep learning has been successfully used to automate the modeling process that trains a network/model from a given experimental dataset to calculate the output directly using high-dimensional complex raw data. However, the trained network is an inverse of the welding process (forward process) that p

Cited by 11SourceScholar
2022

Cascade Transformers for End-to-End Person Search

CVPR 2022poster

The goal of person search is to localize a target person from a gallery set of scene images, which is extremely challenging due to large scale variations, pose/viewpoint changes, and occlusions. In this paper, we propose the Cascade Occluded Attention Transformer (COAT) for end-to-end person search.…

Cited by 84PDFcodeScholar
2022

Design and Analysis of Truss Aerial Transportation System (TATS): The Lightweight Bar Spherical Joint Mechanism

IROS 2022poster

In aerial cooperative transportation missions, it has been recognized that for small-sized but heavy payloads, the cable-suspended framework is a preferred manner. However, to maintain proper safe flight distances, cables always stay inclined, which implies that horizontal force components have to b…

Cited by 3SourceScholar
2022

How to Accurately Monitor the Weld Penetration From Dynamic Weld Pool Serial Images Using CNN-LSTM Deep Learning Model?

RA-L 2022

This letter illustrates how a challenging problem be solved by assuring the adequacy of the raw information and using an appropriate deep learning network to extract the relevant information. The problem concerned is accurate monitoring of the penetration, in a fully penetrated weld pool, as quantif

Cited by 51SourceScholar
2022

Word Sense Disambiguation with Knowledge-Enhanced and Local Self-Attention-based Extractive Sense Comprehension

COLING 2022main

Word sense disambiguation (WSD), identifying the most suitable meaning of ambiguous words in the given contexts according to a predefined sense inventory, is one of the most classical and challenging tasks in natural language processing. Benefiting from the powerful ability of deep neural networks,…

2020

Data-driven Distributed State Estimation and Behavior Modeling in Sensor Networks

IROS 2020poster

Nowadays, the prevalence of sensor networks has enabled tracking of the states of dynamic objects for a wide spectrum of applications from autonomous driving to environmental monitoring and urban planning. However, tracking realworld objects often faces two key challenges: First, due to the limitati…

Cited by 7SourceScholar
2019

Modeling of Human Welders' Operations in Virtual Reality Human-Robot Interaction

RA-L 2019

This letter presents a virtual reality (VR) human–robot interaction welding system that allows human welders to manipulate a welding robot and undertake welding tasks naturally and intuitively via consumer-grade VR hardware (HTC Vive). In this system, human welders’ operations are captured by motion

Cited by 37SourceScholar
2018

Hard-Aware Point-to-Set Deep Metric for Person Re-identification

ECCV 2018poster

Person re-identification (re-ID) is a highly challenging task due to large variations of pose, viewpoint, illumination, and occlusion. Deep metric learning provides a satisfactory solution to person re-ID by training a deep network under supervision of metric loss, e.g., triplet loss. However, the p…

Cited by 180SourcePDFScholar
2015

Direct, Dense, and Deformable: Template-Based Non-Rigid 3D Reconstruction From RGB Video

ICCV 2015poster

In this paper we tackle the problem of capturing the dense, detailed 3D geometry of generic, complex non-rigid meshes using a single RGB-only commodity video camera and a direct approach. While robust and even real-time solutions exist to this problem if the observed scene is static, for non-rigid…

Cited by 115PDFScholar