← Search

Rui Jiang

16 accepted papers

2026

Salus: Strategic Diagnostic Testing for Complex Diagnosis via Multi-Agent Reinforcement Learning

ICML 2026poster

Diagnosing complex diseases is inherently a sequential and iterative medical investigation process, in which a clinician strategically requests multiple rounds of diagnostic tests to differentiate among similar diseases until reaching a definitive diagnosis. Although large language models show great…

Cited by 0SourceScholar
2026

UniScene-MoTion: Unified Scene & Motion-aware Diffusion Transition Framework

AAAI 2026technical

Video transitions are critical for ensuring temporal coherence in edited media, yet existing methods often rely on handcrafted effects or relative-scale trajectories that fail to capture the physical structure of real-world scenes. In this work, we introduce a scale-aware video transition framework

Cited by 0SourcePDFScholar
2025

Design and Development of a Propulsion Induced Rolling Spherical Tensegrity Robot

IROS 2025

Spherical tensegrity structure has good dynamic stability, support strength and flexibility, and is widely used in the field of mobile robot research. Most of the tensegrity spherical robots deform themselves to make gravity work to realize the motion, but the deformation of both rods and ropes affe

Cited by 0SourceScholar
2025

Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection

ICCV 2025poster

The Mixture of Experts (MoE) architecture has excelled in Large Vision-Language Models (LVLMs), yet its potential in real-time open-vocabulary object detectors, which also leverage large-scale vision-language datasets but smaller models, remains unexplored. This work investigates this domain, reveal…

2025

Energy-Guided Optimization for Personalized Image Editing with Pretrained Text-to-Image Diffusion Models

AAAI 2025technical

The rapid advancement of pretrained text-driven diffusion models has significantly enriched applications in image generation and editing. However, as the demand for personalized content editing increases, new challenges emerge especially when dealing with arbitrary objects and complex scenes. Existi…

2025

Injecting Visual Features into Whisper for Parameter-Efficient Noise-Robust Audio-Visual Speech Recognition

ICASSP 2025accepted

Audio-visual speech recognition (AVSR) aims to enhance the robustness of an automatic speech recognition (ASR) systems by incorporating visual information from lip movements, especially in challenging noisy environments. Nevertheless, most current approaches either involve training from scratch or f…

Cited by 0SourceScholar
2025

ModRWKV: Transformer Multimodality in Linear Time

EMNLP 2025

Currently, most multimodal studies are based on large language models (LLMs) with quadratic-complexity Transformer architectures. While linear models like RNNs enjoy low inference costs, their application has been largely limited to the text-only modality. This work explores the capabilities of mode

2025

Open-Modality Latent Modality Interaction Maximization for Audio-Visual Learning

ICASSP 2025accepted

The utilization of multimodal cues enhances the effectiveness of specific cognitive tasks in audio-visual learning. However, on the one hand, designing a unified model for multimodal learning poses challenges due to the presence of information redundancy and modality noise. On the other hand, existi…

Cited by 0SourceScholar
2025

RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control

ICCV 2025poster

Recent advancements in camera-trajectory-guided image-to-video generation offer higher precision and better support for complex camera control compared to text-based approaches. However, they also introduce significant usability challenges, as users often struggle to provide precise camera parameter…

Cited by 0SourcePDFScholar
2024

EVS-assisted Joint Deblurring Rolling-Shutter Correction and Video Frame Interpolation through Sensor Inverse Modeling

CVPR 2024poster

Event-based Vision Sensors (EVS) gain popularity in enhancing CMOS Image Sensor (CIS) video capture. Nonidealities of EVS such as pixel or readout latency can significantly influence the quality of the enhanced images and warrant dedicated consideration in the design of fusion algorithms. A novel ap…

Cited by 2SourcePDFScholar
2023

Time-Aware Multiway Adaptive Fusion Network for Temporal Knowledge Graph Question Answering

ICASSP 2023accepted

Knowledge graphs (KGs) have received increasing attention due to its wide applications on natural language processing. However, its use case on temporal question answering (QA) has not been well-explored. Most of existing methods are developed based on pre-trained language models, which might not be…

Cited by 0SourceScholar
2022

ROSE: Robust Selective Fine-tuning for Pre-trained Language Models

EMNLP 2022main

Even though the large-scale language models have achieved excellent performances, they suffer from various adversarial attacks.A large body of defense methods has been proposed. However, they are still limited due to redundant attack search spaces and the inability to defend against various types of…

2019

Long-Reach Aerial Manipulation Employing Wire-Suspended Hand With Swing-Suppression Device

RA-L 2019

We herein propose a long-reach aerial manipulator to allow the manipulation of objects several meters away from an unmanned aerial vehicle (UAV). The system consists of a multirotor platform with winch mechanism and wire-suspended gripper module equipped with a camera and swing-suppression device. T

Cited by 25SourceScholar
2018

Airborne Docking for Multi-Rotor Aerial Manipulations

IROS 2018poster

We have proposed airborne docking using two multi-rotor aerial robots. This paper presents a transport multi-rotor UAV with winch mechanism and a small multi-rotor with onboard locolization and mobile manipulation system. The winch mechanism enables the UAV to lower and raise a bar to transport anot…

Cited by 44SourceScholar
2017

Geometric Map-Assisted Localization for Mobile Robots Based on Uniform-Gaussian Distribution

RA-L 2017

Drift and scale ambiguity are two main issues which reduce localization accuracy in monocular visual odometry (MVO). It is necessary to propose a unified model to represent these measurement uncertainties. In this paper, we present a geometric map-assisted localization approach for mobile robots equ

Cited by 18SourceScholar