← Search

Jun Cheng

48 accepted papers

2026

SURE: Semi-Dense Uncertainty-REfined Feature Matching

ICRA 2026poster

Establishing reliable image correspondences is essential for many robotic vision problems. However, existing methods often struggle in challenging scenarios with large viewpoint changes or textureless regions, where incorrect correspondences may still receive high similarity scores. This is mainly b…

2026

Safe and Efficient Control: A Subgraph-Augmented Hierarchical Reinforcement Learning Framework for Dynamically Reconfigurable Battery Systems

IJCAI 2026

Dynamically Reconfigurable Battery (DRB) systems employ power electronic switches to create dynamic topologies. They enable effective management of cell difference through real-time adjustment of cell connections. However, existing DRB control methods struggle to learn effective strategies due to sp

Cited by 0Scholar
2025

Debiased Distillation for Consistency Regularization

AAAI 2025technical

Knowledge distillation transfers "dark knowledge" from a large teacher model to a smaller student model, yielding a highly efficient network. To improve network's generalization ability, existing works use a larger temperature coefficient for knowledge distillation. Nevertheless, these methods may…

2025

EchoONE: Segmenting Multiple Echocardiography Planes in One Model

CVPR 2025poster

In clinical practice of echocardiography examinations, multiple planes containing the heart structures of different view are usually required in screening, diagnosis and treatment of cardiac disease. AI models for echocardiography have to be tailored for a specific plane due to the dramatic structur…

2025

Evidential Learning-based Certainty Estimation for Robust Dense Feature Matching

ICLR 2025poster

Dense feature matching methods aim to estimate a dense correspondence field between images. Inaccurate correspondence can occur due to the presence of unmatchable region, necessitating the need for certainty measurement. This is typically addressed by training a binary classifier to decide whether e…

Cited by 0SourcePDFScholar
2025

Human-Imperceptible, Machine-Recognizable Images

IJCAI 2025

Massive human-related data is collected to train neural networks for computer vision tasks. A major conflict is exposed relating to software engineers between better developing AI systems and distancing from the sensitive training data. To reconcile this conflict, the paper proposes an efficient pri

2025

LiftFeat: 3D Geometry-Aware Local Feature Matching

ICRA 2025

Robust and efficient local feature matching plays a crucial role in applications such as SLAM and visual localization for robotics. Despite great progress, it is still very challenging to extract robust and discriminative visual features in scenarios with drastic lighting changes, low texture areas,

Cited by 9SourcecodeScholar
2025

Occlusion-Aware 6D Pose Estimation with Visual Observation Guided Diffusion Model

IROS 2025

Category-level 6D pose estimation in cluttered and occluded environments is a challenging task. Most existing methods rely on deterministic point-based correspondences to estimate target poses, which cannot consider the uncertainty for occluded objects, and thus result in inferior performance. In th

Cited by 0SourceScholar
2025

Rectification-specific Supervision and Constrained Estimator for Online Stereo Rectification

CVPR 2025poster

Online stereo rectification is critical for autonomous vehicles and robots in dynamic environments, where factors such as vibration, temperature fluctuations, and mechanical stress can affect rectification accuracy and severely degrade downstream stereo depth estimation. Current dominant approaches…

Cited by 0SourcePDFScholar
2025

Task-Aware Clustering for Prompting Vision-Language Models

CVPR 2025poster

Prompt learning has attracted widespread attention in adapting vision-language models to downstream tasks. Existing methods largely rely on optimization strategies to ensure the task-awareness of learnable prompts. Due to the scarcity of task-specific data, overfitting is prone to occur. The resulti…

2025

YOLO-KED: A Novel Framework for Rotated Object Detection in Complex Environments

ICASSP 2025accepted

Rotated object detection aims to locate and classify objects with arbitrary orientations. In complex backgrounds, small rotated objects with limited salient features present challenges for standard backbones to extract high-quality, discriminative features. Additionally, traditional single-stage det…

Cited by 0SourceScholar
2024

Compensate Quantization Errors: Make Weights Hierarchical to Compensate Each Other

NAACL 2024findings

Emergent Large Language Models (LLMs) use their extraordinary performance and powerful deduction capacity to discern from traditional language models. However, the expenses of computational resources and storage for these LLMs are stunning, quantization then arises as a trending conversation. To add…

Cited by 2SourcePDFScholar
2024

Distill-then-prune: An Efficient Compression Framework for Real-time Stereo Matching Network on Edge Devices

ICRA 2024poster

In recent years, numerous real-time stereo matching methods have been introduced, but they often lack accuracy. These methods attempt to improve accuracy by introducing new modules or integrating traditional methods. However, the improvements are only modest. In this paper, we propose a novel strate…

Cited by 4SourceScholar
2024

Learning Intra-view and Cross-view Geometric Knowledge for Stereo Matching

CVPR 2024poster

Geometric knowledge has been shown to be beneficial for the stereo matching task. However prior attempts to integrate geometric insights into stereo matching algorithms have largely focused on geometric knowledge from single images while crucial cross-view factors such as occlusion and matching uniq…

2024

OphNet: A Large-Scale Video Benchmark for Ophthalmic Surgical Workflow Understanding

ECCV 2024poster

"Surgical scene perception via videos is critical for advancing robotic surgery, telesurgery, and AI-assisted surgery, particularly in ophthalmology. However, the scarcity of diverse and richly annotated video datasets has hindered the development of intelligent systems for surgical workflow analysi…

2024

Self-Supervised Monocular Depth Estimation with Effective Feature Fusion and Self Distillation

IROS 2024poster

Monocular depth estimation obtaining scene depth information from a single image is an important task in the field of computer vision. Constrained by the limitations of convolutional networks in conducting long-distance modeling and the underutilization of datasets, the generalization of existing mo…

Cited by 1SourceScholar
2024

Semantic-focused Patch Tokenizer with Multi-branch Mixer for Visual Place Recognition

ICRA 2024poster

Visual Place Recognition (VPR) is critical for navigation and loop closure in autonomous driving tasks, mitigating the impact of shift errors caused by dynamic changes in the environment. Due to the limited ability of backbone networks and extreme environmental changes, current methods fail to captu…

Cited by 0SourceScholar
2024

SuperJunction: Learning-Based Junction Detection for Retinal Image Registration

AAAI 2024technical

Keypoints-based approaches have shown to be promising for retinal image registration, which superimpose two or more images from different views based on keypoint detection and description. However, existing approaches suffer from ineffective keypoint detector and descriptor training. Meanwhile, the…

2023

Class-Aware Patch Embedding Adaptation for Few-Shot Image Classification

ICCV 2023poster

"A picture is worth a thousand words", significantly beyond mere a categorization. Accompanied by that, many patches of the image could have completely irrelevant meanings with the categorization if they were independently observed. This could significantly reduce the efficiency of a large family of…

Cited by 32PDFcodeScholar
2023

Efficiently Fusing Sparse Lidar for Enhanced Self-Supervised Monocular Depth Estimation

ICASSP 2023accepted

Monocular self-supervised depth estimation with a low-cost sensor is the mainstream solution to gathering dense depth maps for robots and autonomous driving. In this paper, based on the philosophy "less is more" (i.e., focusing only on valid pixels in sparse LiDAR), we propose a novel framework, Eff…

Cited by 0SourceScholar
2023

GSNet: Model Reconstruction Network for Category-level 6D Object Pose and Size Estimation

ICRA 2023poster

Category-level 6D pose and size estimation is to estimate the rotation, translation and size of the observed instance objects from an arbitrary angle in a cluttered scene. Compared with instance-level 6D pose estimation, there are two main challenges for category-level 6D pose estimation. One is tha…

Cited by 2SourceScholar
2023

Reject Decoding via Language-Vision Models for Text-to-Image Synthesis

AAAI 2023technical

Transformer-based text-to-image synthesis generates images from abstractive textual conditions and achieves prompt results. Since transformer-based models predict visual tokens step by step in testing, where the early error is hard to be corrected and would be propagated. To alleviate this issue, th…

2023

Score Priors Guided Deep Variational Inference for Unsupervised Real-World Single Image Denoising

ICCV 2023poster

Real-world single image denoising is crucial and practical in computer vision. Bayesian inversions combined with score priors now have proven effective for single image denoising but are limited to white Gaussian noise. Moreover, applying existing score-based methods for real-world denoising require…

Cited by 15PDFcodeScholar
2022

Text-to-Image Synthesis Based on Object-Guided Joint-Decoding Transformer

CVPR 2022poster

Object-guided text-to-image synthesis aims to generate images from natural language descriptions built by two-step frameworks, i.e., the model generates the layout and then synthesizes images from the layout and captions. However, such frameworks have two issues: 1) complex structure, since generati…

Cited by 17PDFScholar
2021

MFPN-6D : Real-time One-stage Pose Estimation of Objects on RGB Images

ICRA 2021poster

6D pose estimation of objects is an important part of robot grasping. The latest research trend on 6D pose estimation is to train a deep neural network to directly predict the 2D projection position of the 3D key points from the image, establish the corresponding relationship, and finally use Pespec…

Cited by 14SourceScholar
2021

Multi-Level Graph Encoding with Structural-Collaborative Relation Learning for Skeleton-Based Person Re-Identification

IJCAI 2021poster

Skeleton-based person re-identification (Re-ID) is an emerging open topic providing great value for safety-critical applications. Existing methods typically extract hand-crafted features or model skeleton dynamics from the trajectory of body joints, while they rarely explore valuable relation inform…

2020

Encoding Structure-Texture Relation with P-Net for Anomaly Detection in Retinal Images

ECCV 2020poster

Anomaly detection in retinal image refers to the identification of abnormality caused by various retinal diseases/lesions, by only leveraging normal images in training phase. Normal images from healthy subjects often have regular structures (e.g., the structured blood vessels in the fundus image, or…

2020

RiFeGAN: Rich Feature Generation for Text-to-Image Synthesis From Prior Knowledge

CVPR 2020poster

Text-to-image synthesis is a challenging task that generates realistic images from a textual sequence, which usually contains limited information compared with the corresponding image and so is ambiguous and abstractive. The limited textual information only describes a scene partly, which will compl…

Cited by 145PDFcodeScholar
2020

Self-Supervised Gait Encoding with Locality-Aware Attention for Person Re-Identification

IJCAI 2020poster

Gait-based person re-identification (Re-ID) is valuable for safety-critical applications, and using only 3D skeleton data to extract discriminative gait features for person Re-ID is an emerging open topic. Existing methods either adopt hand-crafted features or learn gait features by traditional supe…

2019

A Deep Step Pattern Representation for Multimodal Retinal Image Registration

ICCV 2019poster

This paper presents a novel feature-based method that is built upon a convolutional neural network (CNN) to learn the deep representation for multimodal retinal image registration. We coined the algorithm deep step patterns, in short DeepSPa. Most existing deep learning based methods require a set o…

Cited by 65PDFScholar
2019

Collect and Select: Semantic Alignment Metric Learning for Few-Shot Learning

ICCV 2019poster

Few-shot learning aims to learn latent patterns from few training examples and has shown promises in practice. However, directly calculating the distances between the query image and support image in existing methods may cause ambiguity because dominant objects can locate anywhere on images. To addr…

Cited by 183PDFcodeScholar
2019

Embedded Block Residual Network: A Recursive Restoration Model for Single-Image Super-Resolution

ICCV 2019poster

Single-image super-resolution restores the lost structures and textures from low-resolved images, which has achieved extensive attention from the research community. The top performers in this field include deep or wide convolutional neural networks, or recurrent neural networks. However, the method…

Cited by 158PDFScholar
2019

Topology Reconstruction of Tree-Like Structure in Images via Structural Similarity Measure and Dominant Set Clustering

CVPR 2019poster

The reconstruction and analysis of tree-like topological structures in the biomedical images is crucial for biologists and surgeons to understand biomedical conditions and plan surgical procedures. The underlying tree-structure topology reveals how different curvilinear components are anatomically…

Cited by 13PDFScholar
2015

A Low-Dimensional Step Pattern Analysis Algorithm With Application to Multimodal Retinal Image Registration

CVPR 2015poster

Existing feature descriptor-based methods on retinal image registration are mainly based on scale-invariant feature transform (SIFT) or partial intensity invariant feature descriptor (PIIFD). While these descriptors are often being exploited, they do not work very well upon unhealthy multimodal imag…

Cited by 45SourcePDFScholar