← Search

Zhongyu Jiang

10 accepted papers

2026

On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs

ICML 2026poster

Reinforcement learning (RL) fine-tuning is now widely used to improve LLM reasoning, and recent work has begun extending it to vision-language models (VLMs). While RL-tuned VLMs can improve visual reasoning benchmark performance, they can still suffer from weak visual grounding, hallucinations, and …

Cited by 0SourceScholar
2025

Bayesian Optimization for Controlled Image Editing via LLMs

ACL 2025finding

In the rapidly evolving field of image generation, achieving precise control over generated content and maintaining semantic consistency remain significant limitations, particularly concerning grounding techniques and the necessity for model fine-tuning. To address these challenges, we propose Bayes…

Cited by 0SourcePDFScholar
2025

MambaMOT: State-Space Model as Motion Predictor for Multi-Object Tracking

ICASSP 2025accepted

In the field of multi-object tracking (MOT), traditional methods often rely on the Kalman filter for motion prediction, leveraging its strengths in linear motion scenarios. However, the inherent limitations of these methods become evident when confronted with complex, nonlinear motions and occlusion…

Cited by 0SourceScholar
2025

The Role of Deductive and Inductive Reasoning in Large Language Models

ACL 2025long

Large Language Models (LLMs) have demonstrated impressive capabilities in reasoning tasks, yet their reliance on static prompt structures and limited adaptability to complex scenarios remains a major challenge. In this paper, we propose the **Deductive and Inductive (DID)** method, a novel framework…

Cited by 0SourcePDFScholar
2024

2D Human Pose Estimation Calibration and Keypoint Visibility Classification

ICASSP 2024accepted

The confidence scores of 2D pose estimation are widely utilized in various fields, including multi-view 3D human pose estimation, skeleton-based human tracking, human action recognition, human re-identification, etc. Despite widespread use, confidence scores from 2D pose estimation methods are unrel…

Cited by 0SourceScholar
2024

A Density-Guided Temporal Attention Transformer for Indiscernible Object Counting in Underwater Videos

ICASSP 2024accepted

Dense object counting or crowd counting has come a long way thanks to the recent development in the vision community. However, indiscernible object counting, which aims to count the number of targets that are blended with respect to their surroundings, has been a challenge. Image-based object counti…

Cited by 0SourceScholar
2024

RT-Pose: A 4D Radar-Tensor based 3D Human Pose Estimation and Localization Benchmark

ECCV 2024poster

"Traditional methods for human localization and pose estimation (HPE), which mainly rely on RGB images as an input modality, confront substantial limitations in real-world applications due to privacy concerns. In contrast, radar-based HPE methods emerge as a promising alternative, characterized by d…

Cited by 5SourcePDFScholar
2023

Global Adaptation Meets Local Generalization: Unsupervised Domain Adaptation for 3D Human Pose Estimation

ICCV 2023poster

When applying a pre-trained 2D-to-3D Human Pose lifting model to a target unseen dataset, a large performance degradation is commonly encountered due to domain shift issues. We observe that the degradation is caused by two factors: 1) the large distribution gap over global positions of poses between…

Cited by 30PDFcodeScholar
2021

Hierarchical Pose Classification for Infant Action Analysis and Mental Development Assessment

ICASSP 2021accepted

Based on Alberta Infant Motor Scale (AIMS), a questionnaire that tracks an infant’s motor function, an infant’s mental development can be evaluated by recording poses a baby can achieve. Therefore, it is meaningful to propose a systematic image-based pose classifier to classify infant actions based…

Cited by 0SourceScholar