← Search

Pengfei Gao

3 accepted papers

2026

UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models

CVPR 2026

With the advancement of multi-modal Large Language Models (LLMs), Video LLMs have been further developed to perform on holistic and specialized video understanding. However, existing works are limited to specialized video understanding tasks, failing to achieve a comprehensive and multi-grained vide

Cited by 0SourceScholar
2025

SoRFT: Issue Resolving with Subtask-oriented Reinforced Fine-Tuning

ACL 2025long

Mainstream issue-resolving frameworks predominantly rely on commercial models, leading to high costs and privacy concerns. Existing training approaches for issue resolving struggle with poor generalization and fail to fully leverage open-source development resources. We propose **S**ubtask-**o**rien…

2024

Cooperation Does Matter: Exploring Multi-Order Bilateral Relations for Audio-Visual Segmentation

CVPR 2024highlight

Recently an audio-visual segmentation (AVS) task has been introduced aiming to group pixels with sounding objects within a given video. This task necessitates a first-ever audio-driven pixel-level understanding of the scene posing significant challenges. In this paper we propose an innovative audio-…