← Search

Yaoyao Zhong

10 accepted papers

2026

Seeing Symbols, Missing Structure: A Real-World Handwritten Mathematical Expression Recognition Benchmark for Large Models

ICML 2026poster

Handwritten mathematical expression recognition (HMER) remains challenging in real-world educational scenarios, even with recent advances in large vision-language models. While these models often achieve high accuracy in local symbol transcription, their reliability in capturing two-dimensional math…

Cited by 0SourceScholar
2025

VIoTGPT: Learning to Schedule Vision Tools Towards Intelligent Video Internet of Things

AAAI 2025technical

Video Internet of Things (VIoT) has shown full potential in collecting an unprecedented volume of video data. How to schedule the domain-specific perceiving models and analyze the collected videos uniformly, efficiently, and especially intelligently to accomplish complicated tasks is challenging. To…

2024

Enhancing Generalization Of Invisible Facial Privacy Cloak Via Gradient Accumulation

ICASSP 2024accepted

The blooming of social media and face recognition (FR) systems has increased people’s concern about privacy and security. A new type of adversarial privacy cloak (class-universal) can be applied to all the images of regular users, to prevent malicious FR systems from acquiring their identity informa…

Cited by 0SourceScholar
2023

Enhancing Generalization of Universal Adversarial Perturbation through Gradient Aggregation

ICCV 2023poster

Deep neural networks are vulnerable to universal adversarial perturbation (UAP), an instance-agnostic perturbation capable of fooling the target model for most samples. Compared to instance-specific adversarial examples, UAP is more challenging as it needs to generalize across various samples and mo…

Cited by 30PDFcodeScholar
2022

Video Question Answering: Datasets, Algorithms and Challenges

EMNLP 2022main

This survey aims to sort out the recent advances in video question answering (VideoQA) and point towards future directions. We firstly categorize the datasets into 1) normal VideoQA, multi-modal VideoQA and knowledge-based VideoQA, according to the modalities invoked in the question-answer pairs, or…

2021

Adaptive Label Noise Cleaning With Meta-Supervision for Deep Face Recognition

ICCV 2021poster

The training of a deep face recognition system usually faces the interference of label noise in the training data. However, it is difficult to obtain a high-precision cleaning model to remove these noises. In this paper, we propose an adaptive label noise cleaning algorithm based on meta-learning fo…

Cited by 14PDFScholar
2020

Generate to Adapt: Resolution Adaption Network for Surveillance Face Recognition

ECCV 2020poster

Although deep learning techniques have largely improved face recognition, unconstrained surveillance face recognition (FR) is still an unsolved challenge, due to the limited training data and the gap of domain distribution. Previous methods mostly match low-resolution and high-resolution faces in di…

Cited by 27SourcePDFScholar
2019

Fair Loss: Margin-Aware Reinforcement Learning for Deep Face Recognition

ICCV 2019poster

Recently, large-margin softmax loss methods, such as angular softmax loss (SphereFace), large margin cosine loss (CosFace), and additive angular margin loss (ArcFace), have demonstrated impressive performance on deep face recognition. These methods incorporate a fixed additive margin to all the clas…

Cited by 111PDFScholar
2019

Unequal-Training for Deep Face Recognition With Long-Tailed Noisy Data

CVPR 2019poster

Large-scale face datasets usually exhibit a massive number of classes, a long-tailed distribution, and severe label noise, which undoubtedly aggravate the difficulty of training. In this paper, we propose a training strategy that treats the head data and the tail data in an unequal way, accompanying…

Cited by 151PDFScholar