← Search

Jiaze Li

18 accepted papers

2026

BEST: Benchmarking Efficiency in Space and Time for LLM-Generated Code

ICML 2026poster

Large language models (LLMs) have revolutionized research in software engineering, and among various tasks, LLM-based code synthesis is promising. A recent line of benchmarks aims to evaluate LLM-generated codes in time efficiency, beyond their correctness. However, *space*, another vital aspect of …

Cited by 0SourceScholar
2026

ImageBindDC: Compressing Multi-modal Data with ImageBind-based Condensation

AAAI 2026technical

Data condensation techniques aim to synthesize a compact dataset from a larger one to enable efficient model training, yet while successful in unimodal settings, they often fail in multimodal scenarios where preserving intricate inter-modal dependencies is crucial. To address this, we introduce Imag

Cited by 0SourcePDFScholar
2026

InteractBench: Benchmarking LLMs on Competitive Programming under Unrevealed Information

ICML 2026poster

Competitive programming is increasingly being used to evaluate the algorithmic reasoning capabilities of large language models (LLMs). However, existing benchmarks primarily focus on full-information tasks where all problem inputs are provided upfront. This overlooks a critical dimension of algorith…

Cited by 0SourceScholar
2026

MSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video Understanding

CVPR 2026

Efficiently understanding long-form videos remains a fundamental challenge for Multimodal Large Language Models (MLLMs). In this paper, we present MLLM-Sampler Joint Evolution (MSJoE), a novel framework that jointly evolves the MLLM and a lightweight key-frame sampler for efficient long-form video u

Cited by 0SourceScholar
2026

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding

CVPR 2026

Self-reflection mechanisms that rely on purely text-based rethinking processes perform well in most multimodal tasks. However, when directly applied to long-form video understanding scenarios, they exhibit clear limitations. The fundamental reasons for this lie in two points: (1) long-form video und

Cited by 0SourceScholar
2026

TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding

CVPR 2026

We introduce TimeViper, a hybrid vision-language model designed to tackle challenges of long video understanding. Processing long videos demands both an efficient model architecture and an effective mechanism for handling extended temporal contexts. To this end, TimeViper adopts a hybrid Mamba-Trans

Cited by 0SourcecodeScholar
2026

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation

ICML 2026poster

Reinforcement learning has emerged as a principled post-training paradigm for Temporal Video Grounding (TVG) due to its on-policy optimization, yet existing GRPO-based methods remain fundamentally constrained by sparse reward signals and substantial computational overhead. We propose Video-OPD, an e…

Cited by 0SourceScholar
2026

Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension

ICML 2026poster

Existing LLM test-time scaling laws emphasize the emergence of self-reflective behaviors through extended reasoning length. Nevertheless, this vertical scaling strategy often encounters plateaus in exploration as the model becomes locked into specific thinking pattern. By shifting from depth to para…

Cited by 0SourceScholar
2025

MGSR: 2D/3D Mutual-boosted Gaussian Splatting for High-fidelity Surface Reconstruction under Various Light Conditions

ICCV 2025poster

Novel view synthesis (NVS) and surface reconstruction (SR) are essential tasks in 3D Gaussian Splatting (3DGS). Despite recent progress, these tasks are often addressed independently, with GS-based rendering methods struggling under diverse light conditions and failing to produce accurate surfaces,…

2025

Multilevel Semantic-Aware Model for AI-Generated Video Quality Assessment

ICASSP 2025accepted

The rapid development of diffusion models has greatly advanced AI-generated videos in terms of length and consistency recently, yet assessing AI-generated videos still remains challenging. Previous approaches have often focused on User-Generated Content(UGC), but few have targeted AI-Generated Video…

Cited by 0SourceScholar
2025

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes

CVPR 2025poster

Reference Audio-Visual Segmentation (Ref-AVS) aims to provide a pixel-wise scene understanding in Language-aided Audio-Visual Scenes (LAVS). This task requires the model to continuously segment objects referred to by text and audio from a video. Previous dual-modality methods always fail due to the…

Cited by 0SourcePDFScholar
2025

Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains

NeurIPS 2025poster

Large Language Models (LLMs) achieve superior performance through Chain-of-Thought (CoT) reasoning, but these token-level reasoning chains are computationally expensive and inefficient. In this paper, we introduce Compressed Latent Reasoning (CoLaR), a novel framework that dynamically compresses rea…

Cited by 0SourceScholar
2025

Towards Anytime Retrieval: A Benchmark for Anytime Person Re-Identification

IJCAI 2025

In real applications, person re-identification (ReID) expects to retrieve the target person at any time, including both daytime and nighttime, ranging from short-term to long-term. However, existing ReID tasks and datasets cannot meet this requirement, as they are constrained by available time and o

2025

UNIS: A Unified Framework for Achieving Unbiased Neural Implicit Surfaces in Volume Rendering

ICCV 2025poster

Reconstruction from multi-view images is a fundamental challenge in computer vision that has been extensively studied over the past decades. Recently, neural radiance fields have driven significant advancements, especially through methods using implicit functions and volume rendering, achieving high…

Cited by 0SourcePDFScholar
2025

iMOVE : Instance-Motion-Aware Video Understanding

ACL 2025finding

Enhancing the fine-grained instance spatiotemporal motion perception capabilities of Video Large Language Models is crucial for improving their temporal and general video understanding. However, current models struggle to perceive detailed and complex instance motions. To address these challenges, w…

Cited by 0SourcePDFScholar