← Search

Jiayi Tian

7 accepted papers

2026

OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs

ICLR 2026poster

Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities across audio and visual modalities, often neglecting either one of the modaliti…

Cited by 0SourcecodeScholar
2025

GeGS-PCR: Fast and Robust Color 3D Point Cloud Registration with Two-Stage Geometric-3DGS Fusion

NeurIPS 2025poster

We address the challenge of point cloud registration using color information, where traditional methods relying solely on geometric features often struggle in low-overlap and incomplete scenarios. To overcome these limitations, we propose GeGS-PCR, a novel two-stage method that combines geometric, c…

Cited by 0SourceScholar
2024

RoleAgent: Building, Interacting, and Benchmarking High-quality Role-Playing Agents from Scripts

NeurIPS 2024poster

Believable agents can empower interactive applications ranging from immersive environments to rehearsal spaces for interpersonal communication. Recently, generative agents have been proposed to simulate believable human behavior by using Large Language Models. However, the existing method heavily re…

Cited by 1SourcePDFScholar
2023

ICD-Face: Intra-class Compactness Distillation for Face Recognition

ICCV 2023poster

Knowledge distillation is an effective model compression method to improve the performance of a lightweight student model by transferring the knowledge of a well-performed teacher model, which has been widely adopted in many computer vision tasks, including face recognition (FR). The current FR dist…

Cited by 6PDFScholar
2022

Dual Regression for Efficient Hand Pose Estimation

ICRA 2022poster

Hand pose estimation constitutes prime attainment for human-machine interaction-based applications. Real-time operation is vital in such tasks. Thus, a reliable estimator should exhibit low computational complexity and high precision at the same time. Previous works have explored the regression tech…

Cited by 9SourceScholar
2022

SCMT: Self-Correction Mean Teacher for Semi-supervised Object Detection

IJCAI 2022poster

Semi-Supervised Object Detection (SSOD) aims to improve performance by leveraging a large amount of unlabeled data. Existing works usually adopt the teacher-student framework to enforce student to learn consistent predictions over the pseudo-labels generated by teacher. However, the performance of t…

Cited by 8SourcePDFScholar