← Search

Dandan Zhu

12 accepted papers

2026

Adapter Shield: A Unified Framework with Built-in Authentication for Preventing Unauthorized Zero-Shot Image-to-Image Generation

CVPR 2026

With the rapid progress in diffusion models, image synthesis has advanced to the stage of zero-shot image-to-image generation, where high-fidelity replication of facial identities or artistic styles can be achieved using just one portrait or artwork, without modifying any model weights. Although the

Cited by 0SourceScholar
2026

CLIP2Pose: Frozen CLIP as Semantic Guide for Domain Adaptive Pose Estimation

AAAI 2026technical

Unsupervised domain adaptive pose estimation is a fundamental yet challenging task due to the need to transfer from labeled synthetic data to unlabeled real data. Nevertheless, the underlying pose semantics, which are governed by spatial structure, remain largely consistent across domains. This obse

Cited by 0SourcePDFScholar
2026

Generalizable Video Quality Assessment via Weak-to-Strong Learning

CVPR 2026

Video quality assessment (VQA) seeks to predict the perceptual quality of a video in alignment with human visual perception, serving as a fundamental tool for quantifying quality degradation across video processing workflows. The dominant VQA paradigm relies on supervised training with human-labeled

Cited by 0SourcecodeScholar
2026

LifeEval: A Multimodal Benchmark for Assistive AI in Egocentric Daily Life Tasks

CVPR 2026

The rapid progress of Multimodal Large Language Models (MLLMs) marks a significant step toward artificial general intelligence, offering great potential for augmenting human capabilities. However, their ability to provide effective assistance in dynamic, real-world environments remains largely under

Cited by 0SourceScholar
2026

RLSLM: A Hybrid Framework Combining Reinforcement Learning and a Rule-based Social Locomotion Model for Socially-aware Navigation

AAAI 2026technical

Navigating human-populated environments without causing discomfort is a critical capability for socially-aware agents. While rule-based approaches offer interpretability through predefined psychological principles, they often lack generalizability and flexibility. Conversely, data-driven methods can

Cited by 0SourcePDFScholar
2026

SalDiff-DTM: A Novel Dual-Temporal Modulated Diffusion Model for Omnidirectional Images Scanpath Prediction

AAAI 2026technical

Scanpath prediction in omnidirectional images (ODIs) serves as a critical component for optimizing foveated rendering efficiency and enhancing interactive quality in virtual reality systems. However, existing scanpath prediction methods for ODIs still suffer from fundamental limitations: (1) inadequ

Cited by 0SourcePDFScholar
2026

VQAThinker: Exploring Generalizable and Explainable Video Quality Assessment via Reinforcement Learning

AAAI 2026technical

Video quality assessment (VQA) aims to objectively quantify perceptual quality degradation in alignment with human visual perception. Despite recent advances, existing VQA models still suffer from two critical limitations: poor generalization to out-of-distribution (OOD) videos and limited explainab

Cited by 0SourcePDFScholar
2025

Mesh Mamba: A Unified State Space Model for Saliency Prediction in Non-Textured and Textured Meshes

CVPR 2025poster

Mesh saliency enhances the adaptability of 3D vision by identifying and emphasizing regions that naturally attract visual attention. To investigate the interaction between geometric structure and texture in shaping visual attention, we establish a comprehensive mesh saliency dataset, which is the fi…

2025

Pose as a Modality: A Psychology-Inspired Network for Personality Recognition with a New Multimodal Dataset

AAAI 2025technical

In recent years, predicting Big Five personality traits from multimodal data has received significant attention in artificial intelligence (AI). However, existing computational models often fail to achieve satisfactory performance. Psychological research has shown a strong correlation between pose a…

Cited by 0SourcePDFScholar
2025

Textured Mesh Saliency: Bridging Geometry and Texture for Human Perception in 3D Graphics

AAAI 2025technical

Textured meshes significantly enhance the realism and detail of objects by mapping intricate texture details onto the geometric structure of 3D models. This advancement is valuable across various applications, including entertainment, education, and industry. While traditional mesh saliency studies…

2022

Learning Invisible Markers for Hidden Codes in Offline-to-Online Photography

CVPR 2022poster

QR (quick response) codes are widely used as an offline-to-online channel to convey information (e.g., links) from publicity materials (e.g., display and print) to mobile devices. However, QR Codes are not favorable for taking up valuable space of publicity materials. Recent works propose invisible…

Cited by 37PDFScholar