← Search

Ruoyu Wang

44 accepted papers

2026

Few-Step Diffusion Sampling Through Instance-Aware Discretizations

CVPR 2026

Diffusion and flow matching models generate high-fidelity data by simulating paths defined by Ordinary or Stochastic Differential Equations (ODEs/SDEs), starting from a tractable prior distribution. The probability flow ODE formulation enables the use of advanced numerical solvers to accelerate samp

Cited by 0SourceScholar
2026

Improving Diffusion Generalization with Weak-to-Strong Segmented Guidance

CVPR 2026

Diffusion models generate synthetic images through an iterative refinement process. However, the misalignment between the simulation-free objective and the iterative process often causes accumulated gradient error along the sampling trajectory, which leads to unsatisfactory results and a failure to

Cited by 0SourcecodeScholar
2026

LongSplat: Online Generalizable 3D Gaussian Splatting from Long Sequence Images

AAAI 2026technical

3D Gaussian Splatting (3DGS) achieves high-fidelity novel view synthesis, but its application in online long-sequence scenarios is still restricted. Existing methods either rely on slow per-scene optimization or lack efficient frame-wise 3DGS updates, making them unsuitable for online long-sequence

Cited by 0SourcePDFScholar
2026

Restoring Initial Noise Sensitivity in Text-to-Image Distillation through Geometric Alignment

ICML 2026poster

Generative distillation significantly accelerates text-to-image (T2I) generation by compressing multi-step trajectories into few-step student models while preserving perceptual quality. However, existing distillation methods prioritize efficiency and output fidelity, often overlooking the preservati…

Cited by 0SourceScholar
2026

SPATIALLY AWARE SELF-SUPERVISED MODELS FOR MULTI-CHANNEL NEURAL SPEAKER DIARIZATION

ICASSP 2026poster

Self-supervised models such as WavLM have demonstrated strong performance for neural speaker diarization. However, these models are typically pre-trained on single-channel recordings, limiting their effectiveness in multi-channel scenarios. Existing diarization systems built on these models often re…

Cited by 0SourcePDFScholar
2026

Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs

AAAI 2026technical

Multi-view understanding, the ability to reconcile visual information across diverse viewpoints for effective navigation, manipulation, and 3D scene comprehension, is a fundamental challenge in Multi-Modal Large Language Models (MLLMs) to be used as embodied agents. While recent MLLMs have shown im

Cited by 0SourcePDFScholar
2025

Adaptive Stochastic Coefficients for Accelerating Diffusion Sampling

NeurIPS 2025poster

Diffusion-based generative processes, formulated as differential equation solving, frequently balance computational speed with sample quality. Our theoretical investigation of ODE- and SDE-based solvers reveals complementary weaknesses: ODE solvers accumulate irreducible gradient error along de…

Cited by 0SourcecodeScholar
2025

Distilling Parallel Gradients for Fast ODE Solvers of Diffusion Models

ICCV 2025poster

Diffusion models (DMs) have achieved state-of-the-art generative performance but suffer from high sampling latency due to their sequential denoising nature. Existing solver-based acceleration methods often face image quality degradation under a low-latency budget. In this paper, we propose the Ensem…

2025

Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings

ICASSP 2025accepted

Although fully end-to-end speaker diarization systems have made significant progress in recent years, modular systems often achieve superior results in real-world scenarios due to their greater adaptability and robustness. Historically, modular speaker diarization methods have seldom discussed how t…

Cited by 0SourceScholar
2025

L2COcc: Lightweight Camera-Centric Semantic Scene Completion via Distillation of LiDAR Model

IROS 2025

Semantic Scene Completion (SSC) constitutes a pivotal element in autonomous driving perception systems, tasked with inferring the 3D semantic occupancy of a scene from sensory data. To improve accuracy, prior research has implemented various computationally demanding and memory-intensive 3D operatio

Cited by 3SourcecodeScholar
2025

Latent Swap Joint Diffusion for 2D Long-Form Latent Generation

ICCV 2025poster

This paper introduces Swap Forward (SaFa), a modality-agnostic and efficient method to generate seamless and coherent long spectrum and panorama using a latent swap joint diffusion process across multi-views. We first investigate spectrum aliasing problem in spectrum-based audio generation caused by…

2025

Mitigating Visual Knowledge Forgetting in MLLM Instruction-tuning via Modality-decoupled Gradient Descent

EMNLP 2025

Recent MLLMs have demonstrated strong visual understanding and reasoning after large-scale multimodal pre-training. However, instruction-tuning is typically text-driven with limited visual supervision, leading to significant visual forgetting and degradation of pre-trained visual knowledge. Existing

Cited by 12SourcePDFScholar
2025

OCEAN: Offline Chain-of-thought Evaluation and Alignment in Large Language Models

ICLR 2025poster

Offline evaluation of LLMs is crucial in understanding their capacities, though current methods remain underexplored in existing research. In this work, we focus on the offline evaluation of the chain-of-thought capabilities and show how to optimize LLMs based on the proposed evaluation method. To e…

Cited by 0SourcePDFScholar
2025

QA-MDT: Quality-aware Masked Diffusion Transformer for Enhanced Music Generation

IJCAI 2025

Text-to-music (TTM) generation, which converts textual descriptions into audio, opens up innovative avenues for multimedia creation. Achieving high quality and diversity in this process demands extensive, high-quality data, which are often scarce in available datasets. Most open-source datasets freq

2025

RecNet: Optimization for Dense Object Detection in Retail Scenarios Based on View Rectification

ICASSP 2025accepted

High-precision dense object detection in retail is crucial for automation, inventory management, and sales optimization. Our experiments revealed that detection models perform significantly better with frontal views than with oblique views, motivating the development of RecNet. RecNet utilizes a Rec…

Cited by 0SourceScholar
2025

ScamNet: Toward Explainable Large Language Model-Based Fraudulent Shopping Website Detection

AAAI 2025technical

Fraudulent shopping websites pose a significant threat to online consumers and legitimate businesses: in 2023, victims of such scams reported $392 million in losses to the Federal Trade Commission. This alarming trend not only impacts individuals but also erodes societal trust in e-commerce, necessi…

2025

Self-Supervised Place Recognition by Refining Temporal and Featural Pseudo Labels From Panoramic Data

RA-L 2025

Visual place recognition (VPR) using deep networks has achieved state-of-the-art performance. However, most of them require a training set with ground truth sensor poses to obtain positive and negative samples of each observation's spatial neighborhood for supervised learning. When such information

Cited by 6SourcecodeScholar
2025

The Silent Assistant: NoiseQuery as Implicit Guidance for Goal-Driven Image Generation

ICCV 2025poster

In this work, we introduce NoiseQuery as a novel method for enhanced noise initialization in versatile goal-driven text-to-image (T2I) generation. Specifically, we propose to leverage an aligned Gaussian noise as implicit guidance to complement explicit user-defined inputs, such as text prompts, for…

2025

Weakly-supervised VLM-guided Partial Contrastive Learning for Visual Language Navigation

IROS 2025

Visual Language Navigation (VLN) is a fundamental task within the field of Embodied AI, focusing on the ability of agents to navigate complex environments based on natural language instructions. Despite the progress made by existing methods, these methods often present some common challenges. First,

Cited by 5SourceScholar
2024

A Spatial Long-Term Iterative Mask Estimation Approach for Multi-Channel Speaker Diarization and Speech Recognition

ICASSP 2024accepted

Deep learning (DL)-based speaker diarization methods have proven powerful performance comparing to traditional clustering-based methods for multi-talker speech diarization and recognition in farfield scenes. However, most DL-based approaches cannot utilize the spatial information well due to the poo…

Cited by 0SourceScholar
2024

A Study of Dropout-Induced Modality Bias on Robustness to Missing Video Frames for Audio-Visual Speech Recognition

CVPR 2024poster

Advanced Audio-Visual Speech Recognition (AVSR) systems have been observed to be sensitive to missing video frames performing even worse than single-modality models. While applying the common dropout techniques to the video modality enhances robustness to missing frames it simultaneously results in…

2024

Air Bumper: A Collision Detection and Reaction Framework for Autonomous MAV Navigation

ICRA 2024poster

Autonomous navigation in unknown environments with obstacles remains challenging for micro aerial vehicles (MAVs) due to their limited onboard computing and sensing resources. Although various collision avoidance methods have been developed, it is still possible for drones to collide with unobserved…

Cited by 4SourcecodeScholar
2024

Behind the Veil: Enhanced Indoor 3D Scene Reconstruction with Occluded Surfaces Completion

CVPR 2024poster

In this paper we present a novel indoor 3D reconstruction method with occluded surface completion given a sequence of depth readings. Prior state-of-the-art (SOTA) methods only focus on the reconstruction of the visible areas in a scene neglecting the invisible areas due to the occlusions e.g. the c…

Cited by 1SourcePDFScholar
2024

Diffusion in Diffusion: Cyclic One-Way Diffusion for Text-Vision-Conditioned Generation

ICLR 2024poster

Originating from the diffusion phenomenon in physics that describes particle movement, the diffusion generative models inherit the characteristics of stochastic random walk in the data space along the denoising trajectory. However, the intrinsic mutual interference among image regions contradicts th…

2024

Implicit Enhancement of Target Speaker in Speaker-Adaptive ASR through Efficient Joint Optimization

ICASSP 2024accepted

In multi-speaker scenarios, automatic speech recognition (ASR) models rely on pre-processed audio after speaker separation. However, when the target speaker is not accurately separated, ASR models face limitations in reaching their peak performance. To address this issue, we propose a speaker-adapti…

Cited by 0SourceScholar
2024

Multi-Modality Action Recognition Based on Dual Feature Shift in Vehicle Cabin Monitoring

ICASSP 2024accepted

Driver Action Recognition (DAR) is crucial in vehicle cabin monitoring systems. In real-world applications, it is common for vehicle cabins to be equipped with cameras featuring different modalities. However, multi-modality fusion strategies for the DAR task within car cabins have rarely been studie…

Cited by 0SourceScholar
2024

Neural Speaker Diarization Using Memory-Aware Multi-Speaker Embedding with Sequence-to-Sequence Architecture

ICASSP 2024accepted

We propose a novel neural speaker diarization system using memory-aware multi-speaker embedding with sequence-to-sequence architecture (NSD-MS2S), which integrates the strengths of memory-aware multi-speaker embedding (MA-MSE) and sequence-to-sequence (Seq2Seq) architecture, leading to improvement i…

Cited by 0SourceScholar
2024

SUP-NeRF: A Streamlined Unification of Pose Estimation and NeRF for Monocular 3D Object Reconstruction

ECCV 2024poster

"Monocular 3D reconstruction for categorical objects heavily relies on accurately perceiving each object’s pose. While gradient-based optimization in a NeRF framework updates the initial pose, this paper highlights that scale-depth ambiguity in monocular object reconstruction causes failures when th…

2024

TCLC-GS: Tightly Coupled LiDAR-Camera Gaussian Splatting for Autonomous Driving

ECCV 2024poster

"Most 3D Gaussian Splatting (3D-GS) based methods for urban scenes initialize 3D Gaussians directly with 3D LiDAR points, which not only underutilizes LiDAR data capabilities but also overlooks the potential advantages of fusing LiDAR with camera data. In this paper, we design a novel tightly couple…

2024

The Multimodal Information Based Speech Processing (MISP) 2023 Challenge: Audio-Visual Target Speaker Extraction

ICASSP 2024accepted

Previous Multimodal Information based Speech Processing (MISP) challenges mainly focused on audio-visual speech recognition (AVSR) with commendable success. However, the most advanced back-end recognition systems often hit performance limits due to the complex acoustic environments. This has prompte…

Cited by 0SourceScholar
2023

An Interactive System for Multiple-Task Linear Temporal Logic Path Planning

IROS 2023poster

Beyond programming robots to accomplish a single high-level task at a time, people also hope robots follow instructions and complete a series of tasks while meeting their requirements. This paper presents an interactive software system that consists of a multiple-task linear temporal logic (LTL) pat…

Cited by 0SourceScholar
2023

Breaking Correlation Shift via Conditional Invariant Regularizer

ICLR 2023poster

Recently, generalization on out-of-distribution (OOD) data with correlation shift has attracted great attentions. The correlation shift is caused by the spurious attributes that correlate to the class label, as the correlation between them may vary in training and test data. For such a problem, we s…

Cited by 8SourcePDFScholar
2023

Concavity-Induced Distance for Unoriented Point Cloud Decomposition

RA-L 2023

We propose Concavity-induced Distance (CID) as a novel way to measure the dissimilarity between a pair of points in an unoriented point cloud. CID indicates the likelihood of two points or two sets of points belonging to different convex parts of an underlying shape represented as a point cloud. Aft

Cited by 0SourcecodeScholar
2023

Quantum Transfer Learning Using the Large-Scale Unsupervised Pre-Trained Model Wavlm-Large for Synthetic Speech Detection

ICASSP 2023accepted

The development of quantum machine learning demonstrates its quantum advantages over traditional deep learning, which promises to discover new patterns on supervised classification datasets. This work proposes a classical-to-quantum transfer learning system based on the large-scale unsupervised pre-…

Cited by 0SourceScholar
2023

Sampling-based path planning under temporal logic constraints with real-time adaptation

ICRA 2023poster

Replanning in temporal logic tasks is extremely difficult during the online execution of robots. This study introduces an effective path planner that computes solutions for temporal logic goals and instantly adapts to non-static and partially unknown environments. Given prior knowledge and a task sp…

Cited by 3SourceScholar
2022

Characterization of Excess Risk for Locally Strongly Convex Population Risk

NeurIPS 2022accept

We establish upper bounds for the expected excess risk of models trained by proper iterative algorithms which approximate the local minima. Unlike the results built upon the strong globally strongly convexity or global growth conditions e.g., PL-inequality, we only require the population risk to be…

2022

Out-of-Distribution Generalization With Causal Invariant Transformations

CVPR 2022poster

In real-world applications, it is important and desirable to learn a model that performs well on out-of-distribution (OOD) data. Recently, causality has become a powerful tool to tackle the OOD generalization problem, with the idea resting on the causal mechanism that is invariant across domains of…

Cited by 80PDFScholar
2020

Real-Time Soft Body 3D Proprioception via Deep Vision-Based Sensing

RA-L 2020

Soft bodies made from flexible and deformable materials are popular in many robotics applications, but their proprioceptive sensing has been a long-standing challenge. In other words, there has hardly been a method to measure and model the high-dimensional 3D shapes of soft bodies with internal sens

Cited by 48SourcecodeScholar
2020

SPARE3D: A Dataset for SPAtial REasoning on Three-View Line Drawings

CVPR 2020poster

Spatial reasoning is an important component of human intelligence. We can imagine the shapes of 3D objects and reason about their spatial relations by merely looking at their three-view line drawings in 2D, with different levels of competence. Can deep networks be trained to perform spatial reasonin…

Cited by 20PDFcodeScholar