← Search

Rui Tang

19 accepted papers

2026

Arcadia: Toward a Full-Lifecycle Framework for Embodied Lifelong Learning

CVPR 2026

We contend that embodied learning is fundamentally a lifecycle problem rather than a single-stage optimization. Systems that optimize only one link (data collection, simulation, learning, or deployment) rarely sustain improvement or generalize beyond narrow settings. We introduce Arcadia, a closed-l

Cited by 0SourceScholar
2026

Autonomous Partner Selection for Cooperative Multi-Agent Reinforcement Learning

AAAI 2026technical

In cooperative Multi-Agent Reinforcement Learning (MARL), the subgroup-wise learning is employed to assign sub-tasks to agents towards the enhancement of team collaboration. However, the present work is dependent on manually defined allocation criteria, which hinders its capacity to adapt to environ

Cited by 0SourcePDFScholar
2026

Beyond Monotonicity: Revisiting Factorization Principles in Multi-Agent Q-Learning

AAAI 2026technical

Value decomposition is a central approach in multi-agent reinforcement learning (MARL), enabling centralized training with decentralized execution by factorizing the global value function into local values. To ensure individual-global-max (IGM) consistency, existing methods either enforce monotonici

Cited by 0SourcePDFScholar
2026

CombinationTS: A Modular Framework for Understanding Time-Series Forecasting Models

ICML 2026poster

Recent progress in time-series forecasting has led to rapidly increasing architectural complexity, yet many reported State-of-the-Art gains are statistically fragile or misattributed. We argue that progress requires a shift from model selection to modular attribution, identifying which components tr…

Cited by 0SourceScholar
2026

Evolving Generalist Virtual Agents with Generative and Associative Memory

AAAI 2026technical

Generalist Virtual Agents (GVAs) powered by Multimodal Large Language Models (MLLMs) exhibit impressive capabilities. However, their long-term learning is hampered by a core limitation: a failure to evolve beyond existing trajectories. This stems from memory systems that treat experiences as isolate

Cited by 0SourcePDFScholar
2026

GeoLanG: Geometry-Aware Language-Guided Grasping with Unified RGB-D Multimodal Learning

ICRA 2026poster

Language-guided grasping has emerged as a promising paradigm for enabling robots to identify and manipulate target objects through natural language instructions, yet it remains highly challenging in cluttered or occluded scenes. Existing methods often rely on multi-stage pipelines that separate obje…

2026

Joint single-shot ToA and DoA estimation for VAA-based BLE ranging with phase ambiguity: A deep learning-based approach

ICASSP 2026poster

Conventional direction-of-arrival (DoA) estimation methods rely on multi-antenna arrays, which are costly to implement on size-constrained Bluetooth Low Energy (BLE) devices. Virtual antenna array (VAA) techniques enable DoA estimation with a single antenna, making angle estimation feasible on such…

Cited by 0SourcePDFScholar
2026

SpatiaLQA: A Benchmark for Evaluating Spatial Logical Reasoning in Vision-Language Models

CVPR 2026

Vision-Language Models (VLMs) have been increasingly applied in real-world scenarios due to their outstanding understanding and reasoning capabilities. Although VLMs have already demonstrated impressive capabilities in common visual question answering and logical reasoning, they still lack the abili

Cited by 0SourcecodeScholar
2026

Towards Physically Executable 3D Gaussian for Embodied Navigation

ICLR 2026poster

3D Gaussian Splatting (3DGS), a 3D representation method with photorealistic real-time rendering capabilities, is regarded as an effective tool for narrowing the sim-to-real gap. However, it lacks fine-grained semantics and physical executability for Visual-Language Navigation (VLN). To address this…

Cited by 0SourceScholar
2025

PrivDNFIS: Privacy-preserving and Efficient Deep Neuro-Fuzzy Inference System

AAAI 2025technical

Deep Neuro-Fuzzy Inference Systems (DNFIS) seamlessly fuse neural networks with the fuzzy inference system enabling intricate decision-making and knowledge representation, while upholding a commendable degree of adaptability and interpretability. However, the challenge of privacy-preserving inferenc…

2025

SpatialLM: Training Large Language Models for Structured Indoor Modeling

NeurIPS 2025poster

SpatialLM is a large language model designed to process 3D point cloud data and generate structured 3D scene understanding outputs. These outputs include architectural elements like walls, doors, windows, and oriented object boxes with their semantic categories. Unlike previous methods which exploit…

Cited by 0SourceScholar
2024

Low-Light Raw Image Enhancement on a Dataset Suffering Light Effects

ICASSP 2024accepted

Deep learning-based methods have achieved remarkable success in low-light image enhancement (LLIE). But most existing works are based on sRGB data and do not focus on the light effects in bright regions when enhancing low-light regions. This inevitably leads to excessive enhancement and saturation o…

Cited by 0SourceScholar
2023

ADU-Depth: Attention-based Distillation with Uncertainty Modeling for Depth Estimation

CoRL 2023poster

Monocular depth estimation is challenging due to its inherent ambiguity and ill-posed nature, yet it is quite important to many applications. While recent works achieve limited accuracy by designing increasingly complicated networks to extract features with limited spatial geometric cues from a sing…

Cited by 2SourceScholar
2023

I2-SDF: Intrinsic Indoor Scene Reconstruction and Editing via Raytracing in Neural SDFs

CVPR 2023poster

In this work, we present I^2-SDF, a new method for intrinsic indoor scene reconstruction and editing using differentiable Monte Carlo raytracing on neural signed distance fields (SDFs). Our holistic neural SDF-based framework jointly recovers the underlying shapes, incident radiance and materials fr…

2021

Layout-Guided Novel View Synthesis From a Single Indoor Panorama

CVPR 2021poster

Existing view synthesis methods mainly focus on the perspective images and have shown promising results. However, due to the limited field-of-view of the pinhole camera, the performance quickly degrades when large camera movements are adopted. In this paper, we make the first attempt to generate nov…

Cited by 28PDFcodeScholar
2020

Geometric Structure Based and Regularized Depth Estimation From 360 Indoor Imagery

CVPR 2020poster

Motivated by the correlation between the depth and the geometric structure of a 360 indoor image, we propose a novel learning-based depth estimation framework that leverages the geometric structure of a scene to conduct depth estimation. Specifically, we represent the geometric structure of an indoo…

Cited by 86PDFScholar
2020

Intelligent Home 3D: Automatic 3D-House Design From Linguistic Descriptions Only

CVPR 2020poster

Home design is a complex task that normally requires architects to finish with their professional skills and tools. It will be fascinating that if one can produce a house plan intuitively without knowing much knowledge about home design and experience of using complex designing tools, for example, v…

Cited by 47PDFcodeScholar
2020

Structured3D: A Large Photo-realistic Dataset for Structured 3D Modeling

ECCV 2020poster

Recently, there has been growing interest in developing learning-based methods to detect and utilize salient semi-global or global structures, such as junctions, lines, planes, cuboids, smooth surfaces, and all types of symmetries, for 3D scene modeling and understanding. However, the ground truth a…