← Search

Zhonglong Zheng

11 accepted papers

2026

AdaDepth: Exploiting Inherent Scene Information for Self-Supervised Depth Estimation in Dynamic Scenes

AAAI 2026technical

Self-supervised monocular depth estimation methods severely compromise accuracy in dynamic objects due to their static scene assumption. Existing approaches for dynamic scenes suffer from two critical shortcomings: 1) reliance on supervised segmentation models (requiring costly annotations) or comp

Cited by 1SourcePDFScholar
2026

Exploiting All Mamba Fusion for Efficient RGB-D Tracking

AAAI 2026technical

Despite the progress made through deep learning, existing Visual Object Tracking (VOT) frameworks struggle with real-world challenges. Recent approaches incorporate additional modalities like Depth, Thermal Infrared, and Language to enhance the robustness of VOT, particularly with the improvement of

Cited by 0SourcePDFScholar
2026

HyperGOOD: Towards Out-of-Distribution Detection in Hypergraphs

AAAI 2026technical

Out-of-distribution (OOD) detection plays a critical role in ensuring the robustness of machine learning models in open-world settings. While extensive efforts have been made in vision, language, and graph domains, the challenge of OOD detection in hypergraph-structured data remains unexplored. In t

Cited by 0SourcePDFScholar
2026

IGIANet: Illumination Guided Implicit Alignment Network for Infrared–Visible UAV Detection

AAAI 2026technical

Visible-Infrared (RGB-IR) Unmanned Aerial Vehicle (UAV) object detection integrates complementary cues from visible and infrared sensors, offering broad application potential. However, due to sensor parallax, it still faces the challenge of weak spatial misalignment, which significantly limits its p

Cited by 0SourcePDFScholar
2026

Neural Outline Cache for Real-time Anti-aliasing Font Rendering

AAAI 2026technical

Neural textures have emerged as pivotal assets in next-generation neural rendering pipelines. However, hardware limitations and programming interface constraints lead to suboptimal performance in multi-instance real-time rendering scenarios. This bottleneck becomes particularly acute for texture-int

Cited by 0SourcePDFScholar
2026

RoSAMDepth: Robust Self-supervised Depth Estimation Leveraging Segment Anything Model

CVPR 2026

Robust depth estimation aims to maintain high-quality depths across diverse conditions. However, most existing methods estimate depth without taking into account the object-level information. As a result, the predicted depth may easily deviate within objects and become blurred under adverse conditio

Cited by 0SourcecodeScholar
2026

Sparse Annotation, Dense Supervision: Unleashing Self-Training Power for Occupancy Prediction With 2D Labels

RA-L 2026

Serving as a fundamental task in robotic navigation and autonomous driving, occupancy prediction is gaining increasing attention for its fine-grained perception of the 3D environment. Most existing methods rely on dense 3D annotations, which are expensive, labor-intensive, and difficult to scale in

Cited by 1SourceScholar
2026

Test-Time Reinforcement Learning for Flow Matching

ICML 2026poster

Flow-matching has emerged as a leading framework for high-fidelity text-to-image generation. However, its alignment with human preferences through RL is often hindered by substantial computational overhead. In this paper, we introduce Flow-TTRL, the first test-time reinforcement learning framework t…

Cited by 0SourceScholar
2025

All Roads Lead to Rome: Exploring Edge Distribution Shifts for Heterophilic Graph Learning

IJCAI 2025

Heterophilic graph neural networks (GNNs) have gained prominence for their ability to learn effective representations in graphs with diverse, attribute-aware relationships. While existing methods leverage attribute inference during message passing to improve performance, they often struggle with cha

Cited by 0SourcePDFScholar
2021

Visual Tracking via Hierarchical Deep Reinforcement Learning

AAAI 2021technical

Visual tracking has achieved great progress due to numerous different algorithms. However, deep trackers based on classification or Siamese network still have their specific limitations. In this work, we show how to teach machines to track a generic object in videos like humans, who can use a few se…

Cited by 35SourcePDFScholar