← Search

Liuyue Xie

5 accepted papers

2026

MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX

AAAI 2026technical

We introduce MAVERIX (Multimodal Audio-Visual Evaluation and Recognition IndeX), a unified benchmark to probe video understanding in multimodal LLMs, encompassing video, audio, and text inputs with human performance baselines. Although recent advancements in audiovisual models have shown substantial

Cited by 0SourcePDFScholar
2026

Unified Spherical Frontend: Learning Rotation-Equivariant Representations of Spherical Images from Any Camera

CVPR 2026

Modern perception increasingly relies on fisheye, panoramic, and other wide field-of-view (FoV) cameras, yet most pipelines still apply planar CNNs designed for pinhole imagery on 2D grids, where pixel-space neighborhoods misrepresent physical adjacency and models are sensitive to global rotations.

Cited by 0SourceScholar
2025

AlignDiff: Learning Physically-Grounded Camera Alignment via Diffusion

ICCV 2025poster

Accurate camera calibration is a fundamental task for 3D perception, especially when dealing with real-world, in-the-wild environments where complex optical distortions are common. Existing methods often rely on pre-rectified images or calibration patterns, which limits their applicability and flexi…

Cited by 0SourcePDFScholar
2020

MuGNet: Multi-Resolution Graph Neural Network for Segmenting Large-Scale Pointclouds

CoRL 2020

In this paper, we propose a multi-resolution deep-learning architecture to segment dense large-scale pointclouds semantically. Dense pointcloud data require a computationally expensive feature encoding process before semantic segmentation. Previous work has used different approaches to drastically d

Cited by 0SourcePDFScholar