← Search

Jiandong Tian

21 accepted papers

2026

DuRP: Dual-Stage Physics-Embedded Learning for Joint Radiance and Polarization Restoration

ICML 2026poster

Polarization information is valuable for many computer vision applications. However, in hazy environments, polarization information is severely attenuated due to the degradation of captured polarized images. Existing dehazing methods struggle to effectively restore polarization information, as singl…

Cited by 0SourceScholar
2026

FocalPolicy: Frequency-Optimized Chunking and Locally Anchored Flow Matching for Coherent Visuomotor Policy

ICML 2026poster

Visuomotor policies aim to learn complex manipulation tasks from expert demonstrations. However, generating smooth and coherent trajectories remains challenging, as it requires balancing proximal precision with distal foresight. Existing approaches typically focus on optimizing intra-chunk action di…

Cited by 0SourceScholar
2025

D2ST-Adapter: Disentangled-and-Deformable Spatio-Temporal Adapter for Few-shot Action Recognition

ICCV 2025poster

Adapting pre-trained image models to video modality has proven to be an effective strategy for robust few-shot action recognition. In this work, we explore the potential of adapter tuning in image-to-video model adaptation and propose a novel video adapter tuning framework, called Disentangled-and-D…

2025

GLAM: Global-Local Variation Awareness in Mamba-based World Model

AAAI 2025technical

Mimicking the real interaction trajectory in the inference of the world model has been shown to improve the sample efficiency of model-based reinforcement learning (MBRL) algorithms. Many methods directly use known state sequences for reasoning. However, this approach fails to enhance the quality of…

2025

RIOcc: Efficient Cross-Modal Fusion Transformer with Collaborative Feature Refinement for 3D Semantic Occupancy Prediction

ICCV 2025poster

The multi-modal 3D semantic occupancy task provides a comprehensive understanding of the scene and has received considerable attention in the field of autonomous driving. However, existing methods mainly focus on processing large-scale voxels, which bring high computational costs and degrade details…

Cited by 0SourcePDFScholar
2024

Exploring Self- and Cross-Triplet Correlations for Human-Object Interaction Detection

AAAI 2024technical

Human-Object Interaction (HOI) detection plays a vital role in scene understanding, which aims to predict the HOI triplet in the form of . Existing methods mainly extract multi-modal features (e.g., appearance, object semantics, human pose) and then fuse them together to directly predict HOI triplet…

Cited by 5SourcePDFScholar
2024

GAFusion: Adaptive Fusing LiDAR and Camera with Multiple Guidance for 3D Object Detection

CVPR 2024poster

Recent years have witnessed the remarkable progress of 3D multi-modality object detection methods based on the Bird's-Eye-View (BEV) perspective. However most of them overlook the complementary interaction and guidance between LiDAR and camera. In this work we propose a novel multi-modality 3D objec…

Cited by 7SourcePDFScholar
2024

Integrating Scaling Strategy and Central Guided Voting for 3D Point Cloud Object Tracking

RA-L 2024

LiDAR-based 3D single object tracking has received remarkable attention due to its crucial role in robotics and autonomous driving. Most of them are based on hierarchical feature structures from PointNet++. However, existing based-stratified structure trackers ignore the fact that non-linearities in

Cited by 4SourceScholar
2024

Robust 3D Tracking with Quality-Aware Shape Completion

AAAI 2024technical

3D single object tracking remains a challenging problem due to the sparsity and incompleteness of the point clouds. Existing algorithms attempt to address the challenges in two strategies. The first strategy is to learn dense geometric features based on the captured sparse point cloud. Nevertheless,…

Cited by 6SourcePDFScholar
2024

SA²VP: Spatially Aligned-and-Adapted Visual Prompt

AAAI 2024technical

As a prominent parameter-efficient fine-tuning technique in NLP, prompt tuning is being explored its potential in computer vision. Typical methods for visual prompt tuning follow the sequential modeling paradigm stemming from NLP, which represents an input image as a flattened sequence of token embe…

2024

Unbiased Faster R-CNN for Single-source Domain Generalized Object Detection

CVPR 2024highlight

Single-source domain generalization (SDG) for object detection is a challenging yet essential task as the distribution bias of the unseen domain degrades the algorithm performance significantly. However existing methods attempt to extract domain-invariant features neglecting that the biased data lea…

Cited by 9SourcePDFScholar
2022

Few-Shot Object Detection by Knowledge Distillation Using Bag-of-Visual-Words Representations

ECCV 2022poster

"While fine-tuning based methods for few-shot object detection have achieved remarkable progress, a crucial challenge that has not been addressed well is the potential class-specific overfitting on base classes and sample-specific overfitting on novel classes. In this work we design a novel knowledg…

Cited by 18SourcePDFScholar
2022

Multi-faceted Distillation of Base-Novel Commonality for Few-Shot Object Detection

ECCV 2022poster

"Most of existing methods for few-shot object detection follow the fine-tuning paradigm, which potentially assumes that the class-agnostic generalizable knowledge can be learned and transferred implicitly from base classes with abundant samples to novel classes with limited samples via such a two-st…

2017

Depth and Image Restoration From Light Field in a Scattering Medium

ICCV 2017poster

Traditional imaging methods and computer vision algorithms are often ineffective when images are acquired in scattering media, such as underwater, fog, and biological tissue. Here, we explore the use of light field imaging and algorithms for image restoration and depth estimation that address the im…

Cited by 55PDFScholar
2017

DeshadowNet: A Multi-Context Embedding Deep Network for Shadow Removal

CVPR 2017spotlight

Shadow removal is a challenging task as it requires the detection/annotation of shadows as well as semantic understanding of the scene. In this paper, we propose an automatic and end-to-end deep neural network (DeshadowNet) to tackle these problems in a unified manner. DeshadowNet is designed with a…

Cited by 367PDFcodeScholar
2017

Video Desnowing and Deraining Based on Matrix Decomposition

CVPR 2017poster

The existing snow/rain removal methods often fail for heavy snow/rain and dynamic scene. One reason for the failure is due to the assumption that all the snowflakes/rain streaks are sparse in snow/rain scenes. The other is that the existing methods often can not differentiate moving objects and snow…

Cited by 199PDFScholar