← Search

Kaiwei Zhang

7 accepted papers

2026

Generalizable Video Quality Assessment via Weak-to-Strong Learning

CVPR 2026

Video quality assessment (VQA) seeks to predict the perceptual quality of a video in alignment with human visual perception, serving as a fundamental tool for quantifying quality degradation across video processing workflows. The dominant VQA paradigm relies on supervised training with human-labeled

Cited by 0SourcecodeScholar
2026

LifeEval: A Multimodal Benchmark for Assistive AI in Egocentric Daily Life Tasks

CVPR 2026

The rapid progress of Multimodal Large Language Models (MLLMs) marks a significant step toward artificial general intelligence, offering great potential for augmenting human capabilities. However, their ability to provide effective assistance in dynamic, real-world environments remains largely under

Cited by 0SourceScholar
2026

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond

ICML 2026poster

The success of large language models (LLMs) in scientific domains has heightened safety concerns, prompting numerous benchmarks to evaluate their scientific safety. Existing benchmarks often suffer from limited risk coverage and a reliance on subjective evaluation. To address thess problems, we intr…

Cited by 0SourceScholar
2026

SalDiff-DTM: A Novel Dual-Temporal Modulated Diffusion Model for Omnidirectional Images Scanpath Prediction

AAAI 2026technical

Scanpath prediction in omnidirectional images (ODIs) serves as a critical component for optimizing foveated rendering efficiency and enhancing interactive quality in virtual reality systems. However, existing scanpath prediction methods for ODIs still suffer from fundamental limitations: (1) inadequ

Cited by 0SourcePDFScholar
2026

VQAThinker: Exploring Generalizable and Explainable Video Quality Assessment via Reinforcement Learning

AAAI 2026technical

Video quality assessment (VQA) aims to objectively quantify perceptual quality degradation in alignment with human visual perception. Despite recent advances, existing VQA models still suffer from two critical limitations: poor generalization to out-of-distribution (OOD) videos and limited explainab

Cited by 0SourcePDFScholar
2025

Mesh Mamba: A Unified State Space Model for Saliency Prediction in Non-Textured and Textured Meshes

CVPR 2025poster

Mesh saliency enhances the adaptability of 3D vision by identifying and emphasizing regions that naturally attract visual attention. To investigate the interaction between geometric structure and texture in shaping visual attention, we establish a comprehensive mesh saliency dataset, which is the fi…

2025

Textured Mesh Saliency: Bridging Geometry and Texture for Human Perception in 3D Graphics

AAAI 2025technical

Textured meshes significantly enhance the realism and detail of objects by mapping intricate texture details onto the geometric structure of 3D models. This advancement is valuable across various applications, including entertainment, education, and industry. While traditional mesh saliency studies…