← Search

Xu Zhao

27 accepted papers

2026

LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding

AAAI 2026technical

The development of multimodal large language models (MLLMs) has advanced general video understanding. However, existing video evaluation benchmarks primarily focus on non-interactive videos, such as movies and recordings. To fill this gap, this paper proposes the first omnimodal benchmark for intera

Cited by 0SourcePDFScholar
2026

TritonGym: A Benchmark for Agentic LLM Workflows in Triton GPU Code Generation

ICML 2026poster

Large language models (LLMs) can already draft plausible Triton kernels, yet most existing evaluations still focus on single-shot generation and underplay tool use and feedback. We introduce *TritonGym*, a benchmark and orchestration framework for evaluating agentic workflows in GPU code generation.…

Cited by 0SourceScholar
2025

Bayesian Optimization for Controlled Image Editing via LLMs

ACL 2025finding

In the rapidly evolving field of image generation, achieving precise control over generated content and maintaining semantic consistency remain significant limitations, particularly concerning grounding techniques and the necessity for model fine-tuning. To address these challenges, we propose Bayes…

Cited by 0SourcePDFScholar
2025

Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking

AAAI 2025technical

Multimodal tracking has garnered widespread attention as a result of its ability to effectively address the inherent limitations of traditional RGB tracking. However, existing multimodal trackers mainly focus on the fusion and enhancement of spatial features or merely leverage the sparse temporal re…

2025

Extracting Sparse Specialist Models from Generalist Models

ICASSP 2025accepted

Recently, several generalist models such as Contrastive Language Image Pre-training (CLIP) have demonstrated their capabilities of performing diverse downstream tasks through zero-shot or few-shot guidance. When these generalist models are used for the specific downstream task where only a fraction…

Cited by 0SourceScholar
2025

The Role of Deductive and Inductive Reasoning in Large Language Models

ACL 2025long

Large Language Models (LLMs) have demonstrated impressive capabilities in reasoning tasks, yet their reliance on static prompt structures and limited adaptability to complex scenarios remains a major challenge. In this paper, we propose the **Deductive and Inductive (DID)** method, a novel framework…

Cited by 0SourcePDFScholar
2025

Weakly-Supervised Video Highlight Detection by Characteristic and Commonality Modeling

ICASSP 2025accepted

Video highlight detection is important for video understanding, as it localizes the attractive regions in the video automatically. Because the fully-supervised video highlight is expensive for its frame-level annotation, we present a novel network for weakly-supervised video highlight detection base…

Cited by 0SourceScholar
2024

Fluctuation-Based Adaptive Structured Pruning for Large Language Models

AAAI 2024technical

Network Pruning is a promising way to address the huge computing resource demands of the deployment and inference of Large Language Models (LLMs). Retraining-free is important for LLMs' pruning methods. However, almost all of the existing retraining-free pruning approaches for LLMs focus on unstruct…

2023

Movement Enhancement toward Multi-Scale Video Feature Representation for Temporal Action Detection

ICCV 2023poster

Boundary localization is a challenging problem in Temporal Action Detection (TAD), in which there are two main issues. First, the submergence of movement feature, i.e. the movement information in a snippet is covered by the scene information. Second, the scale of action, that is, the proportion of a…

Cited by 11PDFScholar
2023

Self-Evaluation Guided Beam Search for Reasoning

NeurIPS 2023poster

Breaking down a problem into intermediate steps has demonstrated impressive performance in Large Language Model (LLM) reasoning. However, the growth of the reasoning chain introduces uncertainty and error accumulation, making it challenging to elicit accurate final results. To tackle this challenge…

2023

ZBS: Zero-Shot Background Subtraction via Instance-Level Background Modeling and Foreground Selection

CVPR 2023poster

Background subtraction (BGS) aims to extract all moving objects in the video frames to obtain binary foreground segmentation masks. Deep learning has been widely used in this field. Compared with supervised-based BGS methods, unsupervised methods have better generalization. However, previous unsuper…

2022

Learning-Based Distortion Correction and Feature Detection for High Precision and Robust Camera Calibration

RA-L 2022

Camera calibration is a crucial technique which significantly influences the performance of many robotic systems. Robustness and high precision have always been the pursuit of diverse calibration methods. State-of-the-art calibration techniques, however, still suffer from inexact corner detection, r

Cited by 17SourceScholar
2022

Structural Triangulation: A Closed-Form Solution to Constrained 3D Human Pose Estimation

ECCV 2022poster

"We propose Structural Triangulation, a closed-form solution for optimal 3D human pose considering multi-view 2D pose estimations, calibrated camera parameters, and bone lengths. To start with, we focus on embedding structural constraints of human body in the process of 2D-to-3D inference using tria…

2018

BSN: Boundary Sensitive Network for Temporal Action Proposal Generation

ECCV 2018poster

Temporal action proposal generation is an important yet challenging problem, since temporal proposals with rich action content are indispensable for analysing real-world videos with long duration and high proportion irrelevant content. This problem requires methods not only generating proposals with…

2017

CoupleNet: Coupling Global Structure With Local Parts for Object Detection

ICCV 2017poster

The region-based Convolutional Neural Network (CNN) detectors such as Faster R-CNN or R-FCN have already shown promising results for object detection by combining the region proposal subnetwork and the classification subnetwork together. Although R-FCN has achieved higher detection speed while keepi…

Cited by 352PDFcodeScholar