← Search

Jinhong Wang

9 accepted papers

2026

WDT-MD: Wavelet Diffusion Transformers for Microaneurysm Detection in Fundus Images

AAAI 2026technical

Microaneurysms (MAs), the earliest pathognomonic signs of Diabetic Retinopathy (DR), present as sub-60 μm lesions in fundus images with highly variable photometric and morphological characteristics, rendering manual screening not only labor-intensive but inherently error-prone. While diffusion-based

Cited by 2SourcePDFScholar
2025

All in One: Visual-Description-Guided Unified Point Cloud Segmentation

ICCV 2025poster

Unified segmentation of 3D point clouds is crucial for scene understanding, but is hindered by its sparse structure, limited annotations, and the challenge of distinguishing fine-grained object classes in complex environments. Existing methods often struggle to capture rich semantic and contextual i…

2025

Dual-level Fuzzy Learning with Patch Guidance for Image Ordinal Regression

IJCAI 2025

Ordinal regression bridges regression and classification by assigning objects to ordered classes. While human experts rely on discriminative patch-level features for decisions, current approaches are limited by the availability of only image-level ordinal labels, overlooking fine-grained patch-level

2025

OrderChain: Towards General Instruct-Tuning for Stimulating the Ordinal Understanding Ability of MLLM

ICCV 2025poster

Despite the remarkable progress of multimodal large language models (MLLMs), they continue to face challenges in achieving competitive performance on ordinal regression (OR; a.k.a. ordinal classification). To address this issue, this paper presents OrderChain, a novel and general prompting paradigm…

2025

Scalable Autoregressive Monocular Depth Estimation

CVPR 2025poster

This paper proposes a new autoregressive model as an effective and scalable monocular depth estimator. Our idea is simple: We tackle the monocular depth estimation (MDE) task with an autoregressive prediction paradigm, based on two core designs. First, our depth autoregressive model (DAR) treats the…

2024

Web2Code: A Large-scale Webpage-to-Code Dataset and Evaluation Framework for Multimodal LLMs

NeurIPS 2024poster

Multimodal large language models (MLLMs) have shown impressive success across modalities such as image, video, and audio in a variety of understanding and generation tasks. However, current MLLMs are surprisingly poor at understanding webpage screenshots and generating their corresponding HTML cod…

2023

Ord2Seq: Regarding Ordinal Regression as Label Sequence Prediction

ICCV 2023poster

Ordinal regression refers to classifying object instances into ordinal categories. It has been widely studied in many scenarios, such as medical disease grading and movie rating. Known methods focused only on learning inter-class ordinal relationships, but still incur limitations in distinguishing a…

Cited by 22PDFcodeScholar
2023

Robust Image Ordinal Regression with Controllable Image Generation

IJCAI 2023poster

Image ordinal regression has been mainly studied along the line of exploiting the order of categories. However, the issues of class imbalance and category overlap that are very common in ordinal regression were largely overlooked. As a result, the performance on minority categories is often unsatisf…

2021

An Underactuated Gripper based on Car Differentials for Self-Adaptive Grasping with Passive Disturbance Rejection

ICRA 2021poster

We introduce an underactuated differential-based robot gripper able to perform self-adaptive grasping with passive disturbance rejection. The gripper utilises three car differential systems to achieve self-adaptiveness with a single actuator: a base differential for distributing power from the motor…

Cited by 9SourceScholar