← Search

Jian Ding

19 accepted papers

2026

High-Quality and Efficient Turbulence Mitigation with Events

CVPR 2026

Turbulence mitigation (TM) is highly ill-posed due to the stochastic nature of atmospheric turbulence. Most methods rely on multiple frames recorded by conventional cameras to capture stable patterns in natural scenarios. However, they inevitably suffer from a trade-off between accuracy and efficien

Cited by 0SourcecodeScholar
2025

Diffusion-Based Imaginative Coordination for Bimanual Manipulation

ICCV 2025poster

Bimanual manipulation is crucial in robotics, enabling complex tasks in industrial automation and household services. However, it poses significant challenges due to the high-dimensional action space and intricate coordination requirements. While video prediction has been recently studied for repres…

2025

InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows

EMNLP 2025

Understanding long-form videos, such as movies and TV episodes ranging from tens of minutes to two hours, remains a significant challenge for multi-modal models. Existing benchmarks often fail to test the full range of cognitive skills needed to process these temporally rich and narratively complex

Cited by 0SourcePDFScholar
2025

Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description

ICCV 2025poster

In this paper, we introduce Part-Aware Point Grounded Description (PaPGD), a challenging task aimed at advancing 3D multimodal learning for fine-grained, part-aware segmentation grounding and detailed explanation of 3D objects. Existing 3D datasets largely focus on either vision-only part segmentati…

Cited by 0SourcePDFScholar
2024

FreePoint: Unsupervised Point Cloud Instance Segmentation

CVPR 2024poster

Instance segmentation of point clouds is a crucial task in 3D field with numerous applications that involve localizing and segmenting objects in a scene. However achieving satisfactory results requires a large number of manual annotations which is time-consuming and expensive. To alleviate dependenc…

2024

Goldfish: Vision-Language Understanding of Arbitrarily Long Videos

ECCV 2024poster

"Most current LLM-based models for video understanding can process videos within minutes. However, they struggle with lengthy videos due to challenges such as “noise and redundancy”, as well as “memory and computation” constraints. In this paper, we present , a methodology tailored for comprehending…

Cited by 15SourcePDFScholar
2024

Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer

AAAI 2024technical

Never having seen an object and heard its sound simultaneously, can the model still accurately localize its visual position from the input audio? In this work, we concentrate on the Audio-Visual Localization and Segmentation tasks but under the demanding zero-shot and few-shot scenarios. To achieve…

2024

Uni3DL: A Unified Model for 3D Vision-Language Understanding

ECCV 2024poster

"We present Uni3DL, a unified model for 3D Vision-Language understanding. Distinct from existing unified 3D vision-language models that mostly rely on projected multi-view images and support limited tasks, Uni3DL operates directly on point clouds and significantly broadens the spectrum of tasks in t…

2024

VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding

NeurIPS 2024poster

We introduce a new benchmark designed to advance the development of general-purpose, large-scale vision-language models for remote sensing images. Although several vision-language datasets in remote sensing have been proposed to pursue this goal, existing datasets are typically tailored to single ta…

2023

Dynamic Coarse-To-Fine Learning for Oriented Tiny Object Detection

CVPR 2023poster

Detecting arbitrarily oriented tiny objects poses intense challenges to existing detectors, especially for label assignment. Despite the exploration of adaptive label assignment in recent oriented object detectors, the extreme geometry shape and limited feature of oriented tiny objects still induce…

2023

Few-Shot Object Detection via Variational Feature Aggregation

AAAI 2023technical

As few-shot object detectors are often trained with abundant base samples and fine-tuned on few-shot novel examples, the learned models are usually biased to base classes and sensitive to the variance of novel examples. To address this issue, we propose a meta-learning framework with two novel featu…

2023

HGFormer: Hierarchical Grouping Transformer for Domain Generalized Semantic Segmentation

CVPR 2023poster

Current semantic segmentation models have achieved great success under the independent and identically distributed (i.i.d.) condition. However, in real-world applications, test data might come from a different domain than training data. Therefore, it is important to improve model robustness against…

2022

Expanding Low-Density Latent Regions for Open-Set Object Detection

CVPR 2022poster

Modern object detectors have achieved impressive progress under the close-set setup. However, open-set object detection (OSOD) remains challenging since objects of unknown categories are often misclassified to existing known classes. In this work, we propose to identify unknown objects by separating…

Cited by 82PDFcodeScholar
2021

Correlation-Based Robust Linear Regression with Iterative Outlier Removal

ICASSP 2021accepted

Here we consider linear regression from the view of correlation and propose a robust regression algorithm. The main idea of this work is from the fact that the inliers lying in a low dimensional subspace are mostly correlated, and the presence of outliers leads to the decrease of correlation. We des…

Cited by 0SourceScholar
2021

DetCo: Unsupervised Contrastive Learning for Object Detection

ICCV 2021poster

We present DetCo, a simple yet effective self-supervised approach for object detection. Unsupervised pre-training methods have been recently designed for object detection, but they are usually deficient in image classification, or the opposite. Unlike them, DetCo transfers well on downstream instanc…

Cited by 408PDFcodeScholar
2019

Learning RoI Transformer for Oriented Object Detection in Aerial Images

CVPR 2019poster

Object detection in aerial images is an active yet challenging task in computer vision because of the bird's-eye view perspective, the highly complex backgrounds, and the variant appearances of objects. Especially when detecting densely packed objects in aerial images, methods relying on horizontal…

Cited by 1326PDFcodeScholar
2018

DOTA: A Large-Scale Dataset for Object Detection in Aerial Images

CVPR 2018poster

Object detection is an important and challenging problem in computer vision. Although the past decade has witnessed major advances in object detection in natural scenes, such successes have been slow to aerial imagery, not only because of the huge variation in the scale, orientation and shape of the…