← Search

Rui Ding

13 accepted papers

2026

COPYLENS: Towards Copyrighted Characters Infringement Detection via Copyright-Aware Prompt Learning

CVPR 2026

Recent advances in text-to-image (T2I) generation can produce highly resembling images of copyrighted characters, often indistinguishable from official depictions, raising serious concerns about intellectual property infringement. Consequently, robust detection of copyright character infringement is

Cited by 0SourceScholar
2026

CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object Detection

CVPR 2026

Multi-camera 3D object detection (MC3D) has attracted increasing attention with the growing deployment of multi-sensor physical agents, such as robots and autonomous vehicles. However, MC3D models still struggle to generalize to unseen platforms with new multi-camera configurations. Current solution

Cited by 0SourceScholar
2026

Generalization of RLVR Using Causal Reasoning as a Testbed

ICLR 2026poster

Reinforcement learning with verifiable rewards (RLVR) has emerged as a promising paradigm for post-training large language models (LLMs) on complex reasoning tasks. Yet, the conditions under which RLVR yields robust generalization remain poorly understood. This paper provides an empirical study of R…

Cited by 0SourceScholar
2026

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models

CVPR 2026

Aligning text-to-video diffusion models with human preferences is crucial for generating high-quality videos. Existing Direct Preference Otimization (DPO) methods rely on multi-sample ranking and task-specific critic models, which is inefficient and often yields ambiguous global supervision. To addr

Cited by 0SourcecodeScholar
2026

RayD3D: Distilling Depth Knowledge Along the Ray for Robust Multi-View 3D Object Detection

AAAI 2026technical

Multi-view 3D detection with bird’s eye view (BEV) is crucial for autonomous driving and robotics, but its robustness in real-world is limited as it struggles to predict accurate depth values. A mainstream solution, cross-modal distillation, transfers depth information from LiDAR to camera models bu

Cited by 0SourcePDFScholar
2026

Test-Time Learning of Causal Structure from Interventional Data

ICML 2026poster

Supervised Causal Learning has shown promise in causal discovery, yet it often struggles with generalization across diverse interventional settings, particularly when intervention targets are unknown. To address this, we propose TICL (Test-time Interventional Causal Learning), a novel method that sy…

Cited by 0SourceScholar
2025

Consistent Feature Alignment for Cross-Modal Knowledge Distillation in Monocular 3D Object Detection

IROS 2025

Cross-modal knowledge distillation (CMKD) in monocular 3D object detection transfers LiDAR’s accurate depth information to compensate for the limitations of camera model. However, current methods directly align the intermediate features of the teacher and student networks, in which the modality gap

Cited by 0SourceScholar
2025

Learning Identifiable Structures Helps Avoid Bias in DNN-based Supervised Causal Learning

AISTATS 2025poster

Causal discovery is a structured prediction task that aims to predict causal relations among variables based on their data samples. Supervised Causal Learning (SCL) is an emerging paradigm in this field. Existing Deep Neural Network (DNN)-based methods commonly adopt the “Node-Edge approach”, in whi…

Cited by 0SourcecodeScholar
2025

QuartDepth: Post-Training Quantization for Real-Time Depth Estimation on the Edge

CVPR 2025poster

Monocular Depth Estimation (MDE) has emerged as a pivotal task in computer vision, supporting numerous real-world applications. However, deploying accurate depth estimation models on resource-limited edge devices, especially Application-Specific Integrated Circuits (ASICs), is challenging due to the…

2024

Text2Analysis: A Benchmark of Table Question Answering with Advanced Data Analysis and Unclear Queries

AAAI 2024technical

Tabular data analysis is crucial in various fields, and large language models show promise in this area. However, current research mostly focuses on rudimentary tasks like Text2SQL and TableQA, neglecting advanced analysis like forecasting and chart generation. To address this gap, we developed the…

2024

VeXKD: The Versatile Integration of Cross-Modal Fusion and Knowledge Distillation for 3D Perception

NeurIPS 2024poster

Recent advancements in 3D perception have led to a proliferation of network architectures, particularly those involving multi-modal fusion algorithms. While these fusion algorithms improve accuracy, their complexity often impedes real-time performance. This paper introduces VeXKD, an effective and V…

Cited by 0SourcePDFScholar
2023

Gradient-Based Graph Attention for Scene Text Image Super-resolution

AAAI 2023technical

Scene text image super-resolution (STISR) in the wild has been shown to be beneficial to support improved vision-based text recognition from low-resolution imagery. An intuitive way to enhance STISR performance is to explore the well-structured and repetitive layout characteristics of text and explo…

2022

ComGAN: Unsupervised Disentanglement and Segmentation via Image Composition

NeurIPS 2022accept

We propose ComGAN, a simple unsupervised generative model, which simultaneously generates realistic images and high semantic masks under an adversarial loss and a binary regularization. In this paper, we first investigate two kinds of trivial solutions in the compositional generation process, and de…

Cited by 12SourcePDFScholar