← Search

Jinrui Zhang

9 accepted papers

2026

Context Learning for Multi-Agent Discussion

ICLR 2026poster

Multi-Agent Discussion (MAD) has garnered increasing attention very recently, where multiple LLM instances collaboratively solve problems via structured discussion. However, we find that current MAD methods easily suffer from discussion inconsistency—LLMs fail to reach a coherent solution—due to the…

Cited by 0SourcecodeScholar
2026

LinProVSR: Linguistics-Knowledge Guided Progressive Disambiguation Network for Visual Speech Recognition

AAAI 2026technical

Visual Speech Recognition (VSR), commonly known as lipreading, enables the recognition of spoken text by analyzing lip visual features. Due to the subtlety of lip movements, its recognition is much harder than other motion recognition tasks. Existing VSR models face the challenge of viseme ambiguit

Cited by 0SourcePDFScholar
2026

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition

CVPR 2026

In controllable image generation, synthesizing coherent and consistent images from multiple reference inputs, i.e., **Multi-Image Composition** (MICo), remains a challenging problem, partly hindered by the lack of high-quality training data.To bridge this gap, we conduct a systematic study of MICo,

Cited by 0SourcecodeScholar
2026

Mitigating Error Accumulation in Continuous Navigation via Memory-Augmented Kalman Filtering

ICML 2026poster

Continuous prediction in complex environments is critical for Unmanned Aerial Vehicle (UAV). However, the existing Vision-Language Navigation (VLN) models follows the dead-reckoning, which iteratively predicts the next waypoint and updates its position, thereby constructing the complete trajectory. …

Cited by 0SourceScholar
2025

LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos

CVPR 2025poster

Despite impressive advancements in video understanding, most efforts remain limited to coarse-grained or visual-only video tasks. However, real-world videos encompass omni-modal information (vision, audio, and speech) with a series of events forming a cohesive storyline. The lack of multi-modal vide…

2025

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors

EMNLP 2025

Recent advancements in large video-language models have revolutionized video understanding tasks. However, their efficiency is significantly constrained by processing high volumes of visual tokens. Existing token compression strategies apply a fixed compression ratio, ignoring the variability in sem

2024

Multi-View Point Cloud Registration Based on Improved NDT Algorithm and ODM Optimization Method

RA-L 2024

The acquisition of targets' complete point cloud model is crucial for tasks such as 3D reconstruction and disordered grasping. Shooting targets from multiple perspectives and registering point clouds from different perspectives can obtain a relatively complete point cloud model. However, small scene

Cited by 8SourceScholar
2024

Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models

ECCV 2024poster

"Large vision-language models (LVLMs) have shown promising performance on a variety of vision-language tasks. However, they remain susceptible to hallucinations, generating outputs misaligned with visual content or instructions. While various mitigation strategies have been proposed, they often negl…

Cited by 7SourcePDFScholar
2023

Transferable Decoding with Visual Entities for Zero-Shot Image Captioning

ICCV 2023poster

Image-to-text generation aims to describe images using natural language. Recently, zero-shot image captioning based on pre-trained vision-language models (VLMs) and large language models (LLMs) has made significant progress. However, we have observed and empirically demonstrated that these methods a…

Cited by 53PDFcodeScholar