← Search

Jinsong Li

14 accepted papers

2026

Beyond Fixed: Training-Free Variable-Length Denoising for Diffusion Large Language Models

ICLR 2026poster

Diffusion Large Language Models (DLLMs) are emerging as a powerful alternative to the dominant Autoregressive Large Language Models, offering efficient parallel generation and capable global context modeling. However, the practical application of DLLMs is hindered by a critical architectural constra…

Cited by 0SourceScholar
2026

Game Ground Bench: Probing the Limits of LVLMs in Complex Semantic Grounding Across Game Universes

AAAI 2026technical

Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities, yet their ability to ground language in complex, interactive environments such as video games remains a critical frontier. Existing benchmarks are inadequate for this purpose: real-world datasets like RefCOCO introduce a

Cited by 0SourcePDFScholar
2026

ScaleCap: Scalable Image Captioning via Dual-Modality Debiasing

ICLR 2026poster

This paper presents ScaleCap, a scalable image captioning strategy that generates comprehensive and detailed image captions. The key challenges of high-quality image captioning lie in the inherent biases of LVLMs: multimodal bias resulting in imbalanced descriptive granularity, offering detailed acc…

Cited by 0SourcecodeScholar
2026

Visual Self-Refine: A Pixel-Guided Paradigm for Accurate Chart Parsing

ICLR 2026poster

While Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities for reasoning and self-correction at the textual level, these strengths provide minimal benefits for complex tasks centered on visual perception, such as Chart Parsing. Existing models often struggle with visually d…

Cited by 0SourcecodeScholar
2025

Dense Semantic Bird-Eye-View Map Generation from Sparse LiDAR Point Clouds via Distribution-aware Feature Fusion

IROS 2025

Semantic scene understanding in bird-eye view (BEV) plays a crucial role in autonomous driving. A common approach to generating BEV maps from LiDAR point-cloud data involves constructing a pillar-level representation by projecting 3D point clouds onto a 2D plane. This process partially discards spat

Cited by 0SourceScholar
2025

STC-Tracker: Spatiotemporal-Consistent Multi-Robot Collaboration Framework for Long-Term Dynamic Object Tracking

IROS 2025

Multi-robot cooperative tracking, as a vital sub-field of multi-robot collaboration, exhibits significant potential in areas such as military reconnaissance and emergency rescue. Conventional dynamic object tracking methods often face issues of incomplete target detection and even loss in complex sc

Cited by 0SourceScholar
2025

Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings

ACL 2025finding

Despite the strong performance of ColPali/ColQwen2 in Visualized Document Retrieval (VDR), its patch-level embedding approach leads to excessive memory usage. This empirical study investigates methods to reduce patch embeddings per page while minimizing performance degradation. We evaluate two token…

2024

Are We on the Right Way for Evaluating Large Vision-Language Models?

NeurIPS 2024poster

Large vision-language models (LVLMs) have recently achieved rapid progress, sparking numerous studies to evaluate their multi-modal capabilities. However, we dig into current evaluation works and identify two primary issues: 1) Visual content is unnecessary for many samples. The answers can be direc…

2024

Fast Temporal Logic Mission Planning of Multiple Robots: A Planning Decision Tree Approach

RA-L 2024

This work develops a fast mission planning framework named planning decision tree (PDT), that can handle large-scale multi-robot systems with temporal logic specifications in real time. Specifically, PDT builds a tree incrementally to represent the task progress. The system states are modeled by bot

Cited by 8SourceScholar
2024

SA²VP: Spatially Aligned-and-Adapted Visual Prompt

AAAI 2024technical

As a prominent parameter-efficient fine-tuning technique in NLP, prompt tuning is being explored its potential in computer vision. Typical methods for visual prompt tuning follow the sequential modeling paradigm stemming from NLP, which represents an input image as a flattened sequence of token embe…

2024

ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

ECCV 2024poster

"Modality alignment serves as the cornerstone for large multi-modal models (LMMs). However, the impact of different attributes (e.g., data type, quality, and scale) of training data on facilitating effective alignment is still under-explored. In this paper, we delve into the influence of training da…

2024

ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

NeurIPS 2024poster

We present the ShareGPT4Video series, aiming to facilitate the video understanding of large video-language models (LVLMs) and the video generation of text-to-video models (T2VMs) via dense and precise captions. The series comprises: 1) ShareGPT4Video, 40K GPT4V annotated dense captions of videos wit…

Cited by 156SourcePDFScholar
2023

R-LIOM: Reflectivity-Aware LiDAR-Inertial Odometry and Mapping

RA-L 2023

With the advent of solid-state LiDAR, a series of related studies have boosted the development of Simultaneous Localization and Mapping (SLAM). However, existing methods cannot work well in indoor environments. In the letter, the reflectivity measurement of the solid-state LiDAR is exploited to impr

Cited by 9SourceScholar