← Search

Ping Lu

5 accepted papers

2026

R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual Scenarios

AAAI 2026technical

Recently, rapid advancements have been made in multimodal large language models (MLLMs), especially in video understanding tasks. However, current research focuses on simple video scenarios, failing to reflect the complex and diverse nature of real-world audio-visual events in videos. To bridge this

Cited by 0SourcePDFScholar
2026

Unsat Core Prediction through Polarity-Aware Representation Learning over Clause-Literal Hypergraphs

ICML 2026poster

Graph neural networks have been widely used in Boolean satisfiability (SAT) tasks to learn structural information from SAT formulas. The goal of these studies is to solve SAT instances or to enhance SAT solvers, including tasks such as unsat-core prediction. However, most existing approaches model a…

Cited by 0SourceScholar
2025

Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection

ICCV 2025poster

The Mixture of Experts (MoE) architecture has excelled in Large Vision-Language Models (LVLMs), yet its potential in real-time open-vocabulary object detectors, which also leverage large-scale vision-language datasets but smaller models, remains unexplored. This work investigates this domain, reveal…

2024

Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models

ECCV 2024poster

"Large vision-language models (LVLMs) have shown promising performance on a variety of vision-language tasks. However, they remain susceptible to hallucinations, generating outputs misaligned with visual content or instructions. While various mitigation strategies have been proposed, they often negl…

Cited by 7SourcePDFScholar
2024

Trust Recognition in Human-Robot Cooperation Using EEG

ICRA 2024poster

Collaboration between humans and robots is becoming increasingly crucial in our daily life. In order to accomplish efficient cooperation, trust recognition is vital, empowering robots to predict human behaviors and make trust-aware decisions. Consequently, there is an urgent need for a generalized a…

Cited by 3SourcecodeScholar