← Search

Rui Zhu

35 accepted papers

2026

Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use

ICML 2026poster

While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered by the dynamic nature of user queries, tool sets, and interaction dynamics. To address this generalization gap, we formalize **OpenAgent** (Tool-Use Agent in Ope…

Cited by 0SourceScholar
2026

Decomposition of Concept-Level Rules in Visual Scenes

ICLR 2026poster

Human cognition is compositional, and one can parse a visual scene into independent concepts and the corresponding concept-changing rules. By contrast, many vision-language systems process images holistically, with limited support for explicit decomposition. And previous methods of decomposing conce…

Cited by 0SourceScholar
2026

Envision, Attend, Then Respond: Counterfactual Hallucination Mitigation in Large Vision-Language Models

CVPR 2026

Large Vision-Language Models (LVLMs) often hallucinate when visual evidence conflicts with world knowledge, i.e., in counterfactual scenarios. We propose Envision-Attend-Respond (EnAR), a training-free framework that leverages visual priors to steer the model's attention toward counterfactual elemen

Cited by 0SourcecodeScholar
2026

Improving LLM-Based Recommenders with Conservative Generative Flow Networks

ICML 2026poster

Generative Flow Networks (GFlowNets) have recently been used to improve diversity and mitigate popularity bias in LLM-based recommender systems, yet most objectives are developed under online-style assumptions. In offline LLM-based recommendation, learning is constrained to a fixed logged dataset, y…

Cited by 0SourceScholar
2026

MS-CRL: Multi-Scale Global Path Planning with Progressive Curriculum Reinforcement Learning

ICRA 2026poster

Global path planning provides high-level guidance for autonomous navigation, supplying reference paths for downstream navigation and control modules. Deep Reinforcement Learning (DRL) has shown strong potential in this domain, but existing methods struggle with multi-scale map inputs. This limitatio…

Cited by 0Scholar
2025

Knowledge-Aware Co-Reasoning for Multidisciplinary Collaboration

EMNLP 2025

Large language models (LLMs) have shown significant potential to improve diagnostic performance for clinical professionals. Existing multi-agent paradigms rely mainly on prompt engineering, suffering from improper agent selection and insufficient knowledge integration. In this work, we propose a nov

Cited by 0SourcePDFScholar
2025

Oracle-MoE: Locality-preserving Routing in the Oracle Space for Memory-constrained Large Language Model Inference

ICML 2025poster

Mixture-of-Experts (MoE) is widely adopted to deploy Large Language Models (LLMs) on edge devices with limited memory budgets. Although MoE is, in theory, an inborn memory-friendly architecture requiring only a few activated experts to reside in the memory for inference, current MoE architectures ca…

Cited by 0SourcePDFScholar
2025

UMDATrack: Unified Multi-Domain Adaptive Tracking Under Adverse Weather Conditions

ICCV 2025poster

Visual object tracking has gained promising progress in past decades. Most of the existing approaches focus on learning target representation in well-conditioned daytime data, while for the unconstrained real-world scenarios with adverse weather conditions, e.g. nighttime or foggy environment, the t…

2025

UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?

NeurIPS 2025poster

With the rapid growth of video generative models (VGMs), it is essential to develop reliable and comprehensive automatic metrics for AI-generated videos (AIGVs). Existing methods either use off-the-shelf models optimized for other tasks or rely on human assessment data to train specialized evaluator…

Cited by 0SourcecodeScholar
2024

Improving Diffusion-Based Image Restoration with Error Contraction and Error Correction

AAAI 2024technical

Generative diffusion prior captured from the off-the-shelf denoising diffusion generative model has recently attained significant interest. However, several attempts have been made to adopt diffusion models to noisy inverse problems either fail to achieve satisfactory results or require a few thousa…

Cited by 4SourcePDFScholar
2024

Once Read is Enough: Domain-specific Pretraining-free Language Models with Cluster-guided Sparse Experts for Long-tail Domain Knowledge

NeurIPS 2024poster

Language models (LMs) only pretrained on a general and massive corpus usually cannot attain satisfying performance on domain-specific downstream tasks, and hence, applying domain-specific pretraining to LMs is a common and indispensable practice. However, domain-specific pretraining can be costly an…

Cited by 0SourcePDFScholar
2024

Research on Inverse Kinematics of Redundant Robotic Arms Based on Flexibility Index

RA-L 2024

This letter proposes an Improved Quantum Particle Swarm Optimization algorithm based on Directional Maneuverability constrained by Isotropic Indexes (DMII-IQPSO). It can solve the problems of large inverse kinematics solution errors and the inability to solve singular configurations of 7-Degree of F

Cited by 12SourceScholar
2024

SD-DiT: Unleashing the Power of Self-supervised Discrimination in Diffusion Transformer

CVPR 2024poster

Diffusion Transformer (DiT) has emerged as the new trend of generative diffusion models on image generation. In view of extremely slow convergence in typical DiT recent breakthroughs have been driven by mask strategy that significantly improves the training efficiency of DiT with additional intra-im…

Cited by 27SourcePDFScholar
2023

Factorized Inverse Path Tracing for Efficient and Accurate Material-Lighting Estimation

ICCV 2023oral

Inverse path tracing has recently been applied to joint material and lighting estimation, given geometry and multi-view HDR observations of an indoor scene. However, it has two major limitations: path tracing is expensive to compute, and ambiguities exist between reflection and emission. Our Facto…

Cited by 14PDFcodeScholar
2023

STINMatch: Semi-Supervised Semantic-Topological Iteration Network for Financial Risk Detection via News Label Diffusion

EMNLP 2023long main

Commercial news provide rich semantics and timely information for automated financial risk detection. However, unaffordable large-scale annotation as well as training data sparseness barrier the full exploitation of commercial news in risk detection. To address this problem, we propose a semi-superv…

Cited by 0SourceScholar
2022

A Theoretical Understanding of Gradient Bias in Meta-Reinforcement Learning

NeurIPS 2022accept

Gradient-based Meta-RL (GMRL) refers to methods that maintain two-level optimisation procedures wherein the outer-loop meta-learner guides the inner-loop gradient-based reinforcement learner to achieve fast adaptations. In this paper, we develop a unified framework that describes variations of GMRL…

2022

IRISformer: Dense Vision Transformers for Single-Image Inverse Rendering in Indoor Scenes

CVPR 2022oral

Indoor scenes exhibit significant appearance variations due to myriad interactions between arbitrarily diverse object shapes, spatially-changing materials, and complex lighting. Shadows, highlights, and inter-reflections caused by visible and invisible light sources require reasoning about long-rang…

Cited by 46PDFcodeScholar
2022

PhotoScene: Photorealistic Material and Lighting Transfer for Indoor Scenes

CVPR 2022poster

Most indoor 3D scene reconstruction methods focus on recovering 3D geometry and scene layout. In this work, we go beyond this to propose PhotoScene, a framework that takes input image(s) of a scene along with approximately aligned CAD geometry (either reconstructed automatically or manually specifie…

Cited by 31PDFcodeScholar
2022

Physically-Based Editing of Indoor Scene Lighting from a Single Image

ECCV 2022poster

"We present a method to edit complex indoor lighting from a single image with its predicted depth and light source segmentation masks. This is an extremely challenging problem that requires modeling complex light transport, and disentangling HDR lighting from material and geometry with only a partia…

Cited by 61SourcePDFScholar
2021

Improving Contrastive Learning by Visualizing Feature Transformation

ICCV 2021poster

Contrastive learning, which aims at minimizing the distance between positive pairs while maximizing that of negative ones, has been widely and successfully applied in unsupervised feature learning, where the design of positive and negative (pos/neg) pairs is one of its keys. In this paper, we attemp…

Cited by 104PDFcodeScholar
2021

OpenRooms: An Open Framework for Photorealistic Indoor Scene Datasets

CVPR 2021poster

We propose a novel framework for creating large-scale photorealistic datasets of indoor scenes, with ground truth geometry, material, lighting and semantics. Our goal is to make the dataset creation process widely accessible, allowing researchers to transform scans into datasets with highquality gro…

Cited by 93PDFScholar
2020

Deep Keypoint-Based Camera Pose Estimation with Geometric Constraints

IROS 2020poster

Estimating relative camera poses from consecutive frames is a fundamental problem in visual odometry (VO) and simultaneous localization and mapping (SLAM), where classic methods consisting of hand-crafted features and sampling-based outlier rejection have been a dominant choice for over a decade. Al…

Cited by 63SourcecodeScholar
2020

Generalization Bound of Gradient Descent for Non-Convex Metric Learning

NeurIPS 2020poster

Metric learning aims to learn a distance measure that can benefit distance-based methods such as the nearest neighbour (NN) classifier. While considerable efforts have been made to improve its empirical performance and analyze its generalization ability by focusing on the data structure and model co…

2020

Multi-Scale Representation Learning for Spatial Feature Distributions using Grid Cells

ICLR 2020spotlight

Unsupervised text encoding models have recently fueled substantial progress in NLP. The key idea is to use neural networks to convert words in texts to vector space representations (embeddings) based on word positions in a sentence and their contexts, which are suitable for end-to-end training of do…

Cited by 146SourcecodeScholar
2020

Single View Metrology in the Wild

ECCV 2020poster

Most 3D reconstruction methods may only recover scene properties up to a global scale ambiguity. We present a novel approach to single view metrology that can recover the absolute scale of a scene represented by 3D heights of objects or camera height above the ground as well as camera parameters of…

2019

ApolloCar3D: A Large 3D Car Instance Understanding Benchmark for Autonomous Driving

CVPR 2019poster

Autonomous driving has attracted remarkable attention from both industry and academia. An important task is to estimate 3D properties (e.g. translation, rotation and shape) of a moving or parked vehicle on the road. This task, while critical, is still under-researched in the computer vision communit…

Cited by 224PDFcodeScholar
2019

ScratchDet: Training Single-Shot Object Detectors From Scratch

CVPR 2019oral

Current state-of-the-art object objectors are fine-tuned from the off-the-shelf networks pretrained on large-scale classification dataset ImageNet, which incurs some additional problems: 1) The classification and detection have different degrees of sensitivity to translation, resulting in the learni…

Cited by 188PDFcodeScholar
2018

Learning Depth From Monocular Videos Using Direct Methods

CVPR 2018poster

The ability to predict depth from a single image - using recent advances in CNNs - is of increasing interest to the vision community. Unsupervised strategies to learning are particularly appealing as they can utilize much larger and varied monocular video datasets during learning without the need fo…

2017

Rethinking Reprojection: Closing the Loop for Pose-Aware Shape Reconstruction From a Single Image

ICCV 2017spotlight

An emerging problem in computer vision is the reconstruction of 3D shape and pose of an object from a single image. Hitherto, the problem has been addressed through the application of canonical deep learning methods to regress from the image directly to the 3D shape and pose labels. These approaches…

Cited by 121PDFScholar