← Search

Rui Dai

11 accepted papers

2026

Do Retrieval Augmented Language Models Know When They Don’t Know?

AAAI 2026technical

Existing large language models (LLMs) occasionally generate plausible yet factually incorrect responses, known as hallucinations. Two main approaches have been proposed to mitigate hallucinations: retrieval-augmented language models (RALMs) and refusal post-training. However, current research predom

Cited by 0SourcePDFScholar
2026

Explaining Data Mixing Scaling Laws

ICML 2026poster

Recent research has established empirical scaling laws to predict model performance on multi-domain data mixtures. However, a theoretical understanding of these model loss behaviors remains limited. In this work, we propose a unified framework to explain the underlying mechanics of data mixing. Our …

Cited by 0SourceScholar
2026

Urban Socio-Semantic Segmentation with Vision-Language Reasoning

ICLR 2026poster

As hubs of human activity, urban surfaces consist of a wealth of semantic entities. Segmenting these various entities from satellite imagery is crucial for a range of downstream applications. Current advanced segmentation models can reliably segment entities defined by physical attributes (e.g., bui…

Cited by 0SourcecodeScholar
2025

Leveraging Submodule Linearity Enhances Task Arithmetic Performance in LLMs

ICLR 2025poster

Task arithmetic is a straightforward yet highly effective strategy for model merging, enabling the resultant model to exhibit multi-task capabilities. Recent research indicates that models demonstrating linearity enhance the performance of task arithmetic. In contrast to existing methods that rely o…

2025

RoboNurse-VLA: Robotic Scrub Nurse System based on Vision-Language-Action Model

IROS 2025

In modern healthcare, the demand for autonomous robotic assistants has grown significantly, particularly in the operating room, where surgical tasks require precision and reliability. Robotic scrub nurses have emerged as a promising solution to improve efficiency and reduce human error during surger

Cited by 28SourcecodeScholar
2024

HYPERmotion: Learning Hybrid Behavior Planning for Autonomous Loco-manipulation

CoRL 2024poster

Enabling robots to autonomously perform hybrid motions in diverse environments can be beneficial for long-horizon tasks such as material handling, household chores, and work assistance. This requires extensive exploitation of intrinsic motion capabilities, extraction of affordances from rich environ…

Cited by 6SourcecodeScholar
2023

Moderately Distributional Exploration for Domain Generalization

ICML 2023poster

Domain generalization (DG) aims to tackle the distribution shift between training domains and unknown target domains. Generating new domains is one of the most effective approaches, yet its performance gain depends on the distribution discrepancy between the generated and target domains. Distributio…

2022

MS-TCT: Multi-Scale Temporal ConvTransformer for Action Detection

CVPR 2022poster

Action detection is an essential and challenging task, especially for densely labelled datasets of untrimmed videos. The temporal relation is complex in those datasets, including challenges like composite action, and co-occurring action. For detecting actions in those complex videos, efficiently cap…

Cited by 99PDFcodeScholar
2021

Learning an Augmented RGB Representation With Cross-Modal Knowledge Distillation for Action Detection

ICCV 2021poster

In video understanding, most cross-modal knowledge distillation (KD) methods are tailored for classification tasks, focusing on the discriminative representation of the trimmed videos. However, action detection requires not only categorizing actions, but also localizing them in untrimmed videos. The…

Cited by 49PDFScholar
2020

VPN: Learning Video-Pose Embedding for Activities of Daily Living

ECCV 2020poster

In this paper, we focus on the spatio-temporal aspect of recognizing Activities of Daily Living (ADL). ADL have two specific properties (i) subtle spatio-temporal patterns and (ii) similar visual patterns varying with time. Therefore, ADL may look very similar and often necessitate to look at their…

2019

Toyota Smarthome: Real-World Activities of Daily Living

ICCV 2019poster

The performance of deep neural networks is strongly influenced by the quantity and quality of annotated data. Most of the large activity recognition datasets consist of data sourced from the web, which does not reflect challenges that exist in activities of daily living. In this paper, we introduce…

Cited by 204PDFScholar