← Search

Lijun Wang

26 accepted papers

2026

Entropy-Based Incremental Coverage Path Planning for Multi-UAV Persistent Monitoring

RA-L 2026

Oil spills continuously affect marine ecosystems and require rapid monitoring for effective emergency response. This letter tackles the problem of persistent monitoring for continuously changing and scattered oil spill regions through Entropy Based Incremental Coverage Path Planning (EICPP). By usin

Cited by 0SourceScholar
2026

Entropy-Based Incremental Coverage Path Planning for Multi-UAV Persistent Monitoring

ICRA 2026poster

Oil spills continuously affect marine ecosystems and require rapid monitoring for effective emergency response. This letter tackles the problem of persistent monitoring for continuously changing and scattered oil spill regions through Entropy-Based Incremental Coverage Path Planning (EICPP). By usin…

Cited by 0SourceScholar
2025

From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction

NeurIPS 2025poster

Despite remarkable progress in driving world models, their potential for autonomous systems remains largely untapped: the world models are mostly learned for world simulation and decoupled from trajectory planning. While recent efforts aim to unify world modeling and planning in a single framework,…

Cited by 0SourcecodeScholar
2025

Mono2Stereo: A Benchmark and Empirical Study for Stereo Conversion

CVPR 2025poster

With the rapid proliferation of 3D devices and the shortage of 3D content, stereo conversion is attracting increasing attention. Recent works introduce pretrained Diffusion Models (DMs) into this task. However, due to the scarcity of large-scale training data and comprehensive benchmarks, the optima…

Cited by 0SourcePDFScholar
2024

DME: Unveiling the Bias for Better Generalized Monocular Depth Estimation

AAAI 2024technical

This paper aims to design monocular depth estimation models with better generalization abilities. To this end, we have conducted quantitative analysis and discovered two important insights. First, the Simulation Correlation phenomenon, commonly seen in long-tailed classification problems, also exist…

2024

Large Occluded Human Image Completion via Image-Prior Cooperating

AAAI 2024technical

The completion of large occluded human body images poses a unique challenge for general image completion methods. The complex shape variations of human bodies make it difficult to establish a consistent understanding of their structures. Furthermore, as human vision is highly sensitive to human bodi…

2024

Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception

CVPR 2024highlight

Multimodal Large Language Model (MLLMs) leverages Large Language Models as a cognitive framework for diverse visual-language tasks. Recent efforts have been made to equip MLLMs with visual perceiving and grounding capabilities. However there still remains a gap in providing fine-grained pixel-level…

2024

PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety

ACL 2024long

Multi-agent systems, when enhanced with Large Language Models (LLMs), exhibit profound capabilities in collective intelligence. However, the potential misuse of this intelligence for malicious purposes presents significant risks. To date, comprehensive research on the safety issues associated with m…

2023

ARKitTrack: A New Diverse Dataset for Tracking Using Mobile RGB-D Data

CVPR 2023poster

Compared with traditional RGB-only visual tracking, few datasets have been constructed for RGB-D tracking. In this paper, we propose ARKitTrack, a new RGB-D tracking dataset for both static and dynamic scenes captured by consumer-grade LiDAR scanners equipped on Apple's iPhone and iPad. ARKitTrack c…

2023

Isomer: Isomerous Transformer for Zero-shot Video Object Segmentation

ICCV 2023poster

Recent leading zero-shot video object segmentation (ZVOS) works devote to integrating appearance and motion information by elaborately designing feature fusion modules and identically applying them in multiple feature stages. Our preliminary experiments show that with the strong long-range dependenc…

Cited by 15PDFcodeScholar
2023

Towards Deeply Unified Depth-aware Panoptic Segmentation with Bi-directional Guidance Learning

ICCV 2023oral

Depth-aware panoptic segmentation is an emerging topic in computer vision which combines semantic and geometric understanding for more robust scene interpretation. Recent works pursue unified frameworks to tackle this challenge but mostly still treat it as two individual learning tasks, which limits…

Cited by 12PDFcodeScholar
2022

Adaptive Co-Teaching for Unsupervised Monocular Depth Estimation

ECCV 2022poster

"Unsupervised depth estimation using photometric losses suffers from local minimum and training instability. We address this issue by proposing an adaptive co-teaching framework to distill the learned knowledge from unsupervised teacher networks to a student network. We design an ensemble architectu…

2022

MVSalNet:Multi-View Augmentation for RGB-D Salient Object Detection

ECCV 2022poster

"RGB-D salient object detection (SOD) enjoys significant advantages in understanding 3D geometry of the scene. However, the geometry information conveyed by depth maps are mostly under-explored in existing RGB-D SOD methods. In this paper, we propose a new framework to address this issue. We augment…

Cited by 44SourcePDFScholar
2022

Multi-Source Uncertainty Mining for Deep Unsupervised Saliency Detection

CVPR 2022poster

Deep learning-based image salient object detection (SOD) heavily relies on large-scale training data with pixel-wise labeling. High-quality labels involve intensive labor and are expensive to acquire. In this paper, we propose a novel multi-source uncertainty mining method to facilitate unsupervised…

Cited by 44PDFScholar
2022

You Only Infer Once: Cross-Modal Meta-Transfer for Referring Video Object Segmentation

AAAI 2022technical

We present YOFO (You Only inFer Once), a new paradigm for referring video object segmentation (RVOS) that operates in an one-stage manner. Our key insight is that the language descriptor should serve as target-specific guidance to identify the target object, while a direct feature fusion of image an…

Cited by 59SourcePDFScholar
2021

Can Scale-Consistent Monocular Depth Be Learned in a Self-Supervised Scale-Invariant Manner?

ICCV 2021poster

Geometric constraints are shown to enforce scale consistency and remedy the scale ambiguity issue in self-supervised monocular depth estimation. Meanwhile, scale-invariant losses focus on learning relative depth, leading to accurate relative depth prediction. To combine the best of both worlds, we l…

Cited by 49PDFScholar
2021

Video Annotation for Visual Tracking via Selection and Refinement

ICCV 2021poster

Deep learning based visual trackers entail offline pre-training on large volumes of video datasets with accurate bounding box annotations that are labor-expensive to achieve. We present a new framework to facilitate bounding box annotations for video sequences, which investigates a selection-and-ref…

Cited by 11PDFcodeScholar
2020

CLIFFNet for Monocular Depth Estimation with Hierarchical Embedding Loss

ECCV 2020poster

This paper proposes a hierarchical loss for monocular depth estimation, which measures the differences between the prediction and ground truth in hierarchical embedding spaces of depth maps. In order to find an appropriate embedding space, we design different architectures for hierarchical embedding…

2020

SDC-Depth: Semantic Divide-and-Conquer Network for Monocular Depth Estimation

CVPR 2020poster

Monocular depth estimation is an ill-posed problem, and as such critically relies on scene priors and semantics. Due to its complexity, we propose a deep neural network model based on a semantic divide-and-conquer approach. Our model decomposes a scene into semantic segments, such as object instance…

Cited by 160PDFScholar
2018

Structured Siamese Network for Real-Time Visual Tracking

ECCV 2018poster

Local structure of target objects are essential for robust tracking. However, existing methods based on deep neural networks mostly describe the target appearance from the global view, leading to high sensitivity to non-rigid appearance change and partial occlusion. In this paper, we circumvent this…

Cited by 325SourcePDFScholar
2017

Learning to Detect Salient Objects With Image-Level Supervision

CVPR 2017poster

Deep Neural Networks (DNNs) have substantially improved the state-of-the-art in salient object detection. However, training DNNs requires costly pixel-level annotations. In this paper, we leverage the observation that image-level tags provide important cues of foreground salient objects, and develop…

Cited by 1450PDFScholar
2016

STCT: Sequentially Training Convolutional Networks for Visual Tracking

CVPR 2016poster

Due to the limited amount of training samples, fine-tuning pre-trained deep models online is prone to over-fitting. In this paper, we propose a sequential training method for convolutional neural networks (CNNs) to effectively transfer pre-trained deep features for online applications. We regard a C…

Cited by 330PDFScholar
2015

Deep Networks for Saliency Detection via Local Estimation and Global Search

CVPR 2015poster

This paper presents a saliency detection algorithm by integrating both local estimation and global search. In the local estimation stage, we detect local saliency by using a deep neural network (DNN-L) which learns local patch features to determine the saliency value of each pixel. The estimated loc…

Cited by 829SourcePDFScholar