← Search

Jianfeng Ren

23 accepted papers

2026

Paper Folding Puzzles: Can Multimodal Large Language Models Perform Spatial Reasoning?

AAAI 2026technical

Multimodal Large Language Models (MLLMs) largely lag human-level performance on abstract visual reasoning (AVR), which requires models to infer latent rules from visual question sets and generalize them to novel scenarios. Most AVR benchmarks are constrained to narrow and repetitive 2D patterns, inv

Cited by 0SourcePDFScholar
2026

SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark

CVPR 2026

Fine-grained spatiotemporal reasoning on surgical videos is critical, yet the capabilities of Multi-modal Large Language Models (MLLMs) in this domain remain largely unexplored. To bridge this gap, we introduce SurgCoT, a unified benchmark for evaluating chain-of-thought (CoT) reasoning in MLLMs acr

Cited by 0SourcecodeScholar
2026

X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic Diagnosis

CVPR 2026

Despite significant progress in Multi-modal Large Language Models (MLLMs), their clinical reasoning capacity for multi-modal diagnosis remains largely unexamined. Current benchmarks, mostly single-modality data, can't evaluate progressive reasoning and cross-modal integration essential for clinical

Cited by 0SourcecodeScholar
2025

DARR: A Dual-Branch Arithmetic Regression Reasoning Framework for Solving Machine Number Reasoning

AAAI 2025technical

Abstract visual reasoning (AVR) is a critical ability of humans, and it has been widely studied, but arithmetic visual reasoning, a unique task in AVR to reason over number sense, is less studied in the literature. To facilitate this research, we construct a Machine Number Reasoning (MNR) dataset to…

2025

DBCR: Exploiting Both Intra-cluster and Extra-cluster Relations for Compositional Reasoning

ICASSP 2025accepted

Most existing models for abstract visual reasoning perform poorly in compositional visual reasoning (CVR), due to complex nature of compositional rules and difficulties in distinguishing tiny rule differences between outliers and normal images. To tackle the challenges, we propose a Dual-Branch Comp…

Cited by 0SourceScholar
2025

DSRF: A Dynamic and Scalable Reasoning Framework for Solving RPMs

NeurIPS 2025poster

Abstract Visual Reasoning (AVR) entails discerning latent patterns in visual data and inferring underlying rules. Existing solutions often lack scalability and adaptability, as deep architectures tend to overfit training data, and static neural networks fail to dynamically capture diverse rules. To…

Cited by 0SourcecodeScholar
2025

ERL-MPP: Evolutionary Reinforcement Learning with Multi-head Puzzle Perception for Solving Large-scale Jigsaw Puzzles of Eroded Gaps

AAAI 2025technical

Solving jigsaw puzzles has been extensively studied. While most existing models focus on solving either small-scale puzzles or puzzles with no gap between fragments, solving large-scale puzzles with gaps presents distinctive challenges in both image understanding and combinatorial optimization. To t…

Cited by 0SourcePDFScholar
2025

FineMotion: A Dataset and Benchmark with both Spatial and Temporal Annotation for Fine-grained Motion Generation and Editing

ICCV 2025poster

Generating realistic human motions from textual descriptions has undergone significant advancements. However, existing methods often overlook specific body part movements and their timing. In this paper, we address this issue by enriching the textual description with more details. Specifically, we p…

Cited by 0SourcePDFScholar
2025

Jointly Optimizing Data Discretization and Naive Bayes Classifier via Multi-Objective Optimization

ICASSP 2025accepted

Data discretization plays a critical role in enhancing the performance of the naive Bayes classifier. Traditional data discretization methods often utilize a two-stage framework, where data discretization and classification are optimized separately, leading to sub-optimal performance. To tackle the…

Cited by 0SourceScholar
2025

MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple Granularities

CVPR 2025poster

Recent motion-aware large language models have demonstrated promising potential in unifying motion comprehension and generation. However, existing approaches primarily focus on coarse-grained motion-text modeling, where text describes the overall semantics of an entire motion sequence in just a few…

2025

S³-Mamba: Small-Size-Sensitive Mamba for Lesion Segmentation

AAAI 2025technical

Small lesions play a critical role in early disease diagnosis and intervention of severe infections. Popular models often face challenges in segmenting small lesions, as it occupies only a minor portion of an image, while down-sampling operations may inevitably lose focus on local features of small…

Cited by 1SourcePDFScholar
2024

Regression Residual Reasoning with Pseudo-labeled Contrastive Learning for Uncovering Multiple Complex Compositional Relations

IJCAI 2024poster

Abstract Visual Reasoning (AVR) has been widely studied in literature. Our study reveals that AVR models tend to rely on appearance matching rather than a genuine understanding of underlying rules. We hence develop a challenging benchmark, Multiple Complex Compositional Reasoning (MC2R), composed of…

Cited by 4SourcePDFScholar
2024

Scale Optimization Using Evolutionary Reinforcement Learning for Object Detection on Drone Imagery

AAAI 2024technical

Object detection in aerial imagery presents a significant challenge due to large scale variations among objects. This paper proposes an evolutionary reinforcement learning agent, integrated within a coarse-to-fine object detection framework, to optimize the scale for more effective detection of obje…

2023

Confidence-Based Event-Centric Online Video Question Answering on a Newly Constructed ATBS Dataset

ICASSP 2023accepted

Deep neural networks facilitate video question answering (VideoQA), but the real-world applications on video streams such as CCTV and live cast place higher demands on the solver. To address the challenges of VideoQA on long videos of unknown length, we define a new set of problems called Online Ope…

Cited by 0SourceScholar
2023

Dual-Stream Siamese Vision Transformer With Mutual Attention For Radar Gait Verification

ICASSP 2023accepted

The inconspicuousness of human gait characteristic in radar signal makes it hard to differentiate different identities. In this work, a Dual-stream Siamese Vision Transformer with Mutual Attention is proposed to verify whether a pair of radar gait sequences originate from the same person or not. The…

Cited by 0SourceScholar
2023

Hierarchical ConViT with Attention-Based Relational Reasoner for Visual Analogical Reasoning

AAAI 2023technical

Raven’s Progressive Matrices (RPMs) have been widely used to evaluate the visual reasoning ability of humans. To tackle the challenges of visual perception and logic reasoning on RPMs, we propose a Hierarchical ConViT with Attention-based Relational Reasoner (HCV-ARR). Traditional solution methods o…

2023

Siamese-Discriminant Deep Reinforcement Learning for Solving Jigsaw Puzzles with Large Eroded Gaps

AAAI 2023technical

Jigsaw puzzle solving has recently become an emerging research area. The developed techniques have been widely used in applications beyond puzzle solving. This paper focuses on solving Jigsaw Puzzles with Large Eroded Gaps (JPwLEG). We formulate the puzzle reassembly as a combinatorial optimization…

Cited by 20SourcePDFScholar
2023

Solving Jigsaw Puzzle of Large Eroded Gaps Using Puzzlet Discriminant Network

ICASSP 2023accepted

Solving Jigsaw puzzles has recently become an emerging research topic. Traditionally, boundary similarities are utilized for puzzle reassembly. In this paper, we solve Jigsaw Puzzles of Large Eroded Gaps (JPLEG), where boundary similarities are weak and image semantics are the only feasible clues. I…

Cited by 0SourceScholar
2022

Attention-Based Dual-Stream Vision Transformer for Radar Gait Recognition

ICASSP 2022accepted

Radar gait recognition is robust to light variations and less infringement on privacy. Previous studies often utilize either spectrograms or cadence velocity diagrams. While the former shows the time-frequency patterns, the latter encodes the repetitive frequency patterns. In this work, a dual-strea…

Cited by 0SourceScholar
2022

Dynamic Texture Recognition Using PDV Hashing and Dictionary Learning on Multi-Scale Volume Local Binary Pattern

ICASSP 2022accepted

Spatial-temporal local binary pattern (STLBP) has been widely used in dynamic texture recognition. STLBP often encounters the high-dimension problem as its dimension increases exponentially, so that STLBP could only utilize a small neighborhood. To tackle this problem, we propose a method for dynami…

Cited by 0SourceScholar
2022

Spatial-Context-Aware Deep Neural Network for Multi-Class Image Classification

ICASSP 2022accepted

Multi-label image classification is a fundamental but challenging task in computer vision. Over the past few decades, solutions exploring relationships between semantic labels have made great progress. However, the underlying spatial-contextual information of labels is under-exploited. To tackle thi…

Cited by 0SourceScholar