← Search

Jing Bi

18 accepted papers

2026

Asynchronous Temporal Modeling with Two-Agent Framework for Streaming Dense Video Captioning

CVPR 2026

Streaming dense video captioning requires real-time processing of continuous visual input while determining precisely when and what to caption. Current approaches primarily focus on designing complex external memory mechanisms, failing to leverage Large Multimodal Models' (LMMs) inherent long-contex

Cited by 0SourceScholar
2026

Bridging Facial Understanding and Animation via Language Models

CVPR 2026

Text-guided human body animation has advanced rapidly, yet facial animation lags due to the scarcity of well-annotated, text-paired facial corpora. To close this gap, we leverage foundation generative models to synthesize a large, balanced corpus of facial behavior. We design prompts suite covering

Cited by 0SourceScholar
2026

Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting

AAAI 2026technical

In this work, we introduce CAT-V (Caption Anything in Video), a training-free framework for fine-grained object-centric video captioning of user-selected instances. CAT-V combines (i) a SAMURAI-based Segmenter for precise object masks across frames, (ii) a TRACE-Uni Temporal Analyzer for event bound

Cited by 0SourcePDFScholar
2026

Does Reasoning Improve Seeing? Understanding When Vision-Language Models Benefit from Thinking

ICML 2026poster

Vision–language models (VLMs) now support both direct Instruct and explicit-reasoning Thinking modes, but practitioners lack principled ways to decide when reasoning helps or how much computation to allocate at test time. We investigate whether VLMs encode meta-cognitive signals for adaptive inferen…

Cited by 0SourceScholar
2026

PGD-NO: A Neural Operator with Precomputed Geometry Decomposition for 3D Million-Scale physics simulations

ICML 2026poster

While neural PDE solvers have demonstrated significant potential for accelerating engineering simulations, existing architectures remain constrained by high memory consumption and the "single-node bottleneck," where the maximum processable mesh resolution is strictly limited by the VRAM of a single …

Cited by 0SourceScholar
2026

When to Think and When to Look: Uncertainty-Guided Lookback

CVPR 2026

Test-time "thinking" (i.e., generating explicit intermediate reasoning chains) is known to boost performance in large language models and has recently shown strong gains for large vision-language models (LVLMs). However, despite these promising results, there is still no systematic analysis of how t

Cited by 0SourcecodeScholar
2025

Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding

AAAI 2025technical

Large language models (LLMs) have demonstrated remarkable capabilities in natural language and multimodal domains. By fine-tuning multimodal LLMs with temporal annotations from well-annotated datasets, e.g., dense video captioning datasets, their temporal understanding capacity in video-language tas…

Cited by 5SourcePDFScholar
2025

Enhancing the Reasoning Capabilities of Small Language Models via Solution Guidance Fine-Tuning

COLING 2025main

Large language models (LLMs) have demonstrated remarkable performance across a wide range of tasks. Advances in prompt engineering and fine-tuning techniques have further enhanced their ability to address complex reasoning challenges. However, these advanced capabilities are often exclusive to model…

2025

MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness

NeurIPS 2025poster

Understanding perspective is fundamental to human visual perception, yet the extent to which multimodal large language models (MLLMs) internalize perspective geometry remains unclear. We introduce MMPerspective, the first benchmark specifically designed to systematically evaluate MLLMs' understandin…

Cited by 0SourcecodeScholar
2025

Unveiling Visual Perception in Language Models: An Attention Head Analysis Approach

CVPR 2025poster

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated remarkable progress in visual understanding. This impressive leap raises a compelling question: how can language models, initially trained solely on linguistic data, effectively interpret and process visual content? Th…

Cited by 4SourcePDFScholar
2025

VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?

CVPR 2025poster

The advancement of Multimodal Large Language Models (MLLMs) has enabled significant progress in multimodal understanding, expanding their capacity to analyze video content. However, existing evaluation benchmarks for MLLMs primarily focus on abstract video comprehension, lacking a detailed assessmen…

2025

ZeroSep: Separate Anything in Audio with Zero Training

NeurIPS 2025poster

Audio source separation is fundamental for machines to understand complex acoustic environments and underpins numerous audio applications. Current supervised deep learning approaches, while powerful, are limited by the need for extensive, task-specific labeled data and struggle to generalize to the…

Cited by 0SourceScholar
2024

OSCaR: Object State Captioning and State Change Representation

NAACL 2024findings

The capability of intelligent models to extrapolate and comprehend changes in object states is a crucial yet demanding aspect of AI research, particularly through the lens of human interaction in real-world settings. This task involves describing complex visual environments, identifying active objec…

2023

Multi-swarm Genetic Gray Wolf Optimizer with Embedded Autoencoders for High-dimensional Expensive Problems

ICRA 2023poster

High-dimensional expensive problems are often encountered in the design and optimization of complex robotic and automated systems and distributed computing systems, and they suffer from a time-consuming fitness evaluation process. It is extremely challenging and difficult to produce promising soluti…

Cited by 16SourceScholar
2023

Self-adaptive Teaching-learning-based Optimizer with Improved RBF and Sparse Autoencoder for Complex Optimization Problems

ICRA 2023poster

Evolutionary algorithms are commonly used to solve many complex optimization problems in such fields as robotics, industrial automation, and complex system design. Yet, their performance is limited when dealing with high-dimensional complex problems because they often require enormous computational…

Cited by 11SourceScholar
2022

Large-scale Network Traffic Prediction With LSTM and Temporal Convolutional Networks

ICRA 2022poster

Real-time and precise prediction for traffic of networks is critically important for allocating the optimal computing/network resources based on users' business requirements, analyzing the network performance, and realizing intelligent congestion control and high-accuracy anomaly detection. The dram…

Cited by 20SourceScholar
2021

Procedure Planning in Instructional Videos via Contextual Modeling and Model-Based Policy Learning

ICCV 2021poster

Learning new skills by observing humans' behaviors is an essential capability of AI. In this work, we leverage instructional videos to study humans' decision-making processes, focusing on learning a model to plan goal-directed actions in real-life videos. In contrast to conventional action recogniti…

Cited by 55PDFScholar