← Search

Mengmi Zhang

19 accepted papers

2026

Learning to See Through a Baby's Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and Machines

CVPR 2026

Newborns perceive the world with low-acuity, color-degraded, and temporally continuous vision, which gradually sharpens as infants develop. To explore the ecological advantages of such staged "visual diets", we train self-supervised learning (SSL) models on object-centric videos under constraints th

Cited by 0SourcecodeScholar
2026

PEERING INTO THE UNKNOWN: ACTIVE VIEW SELECTION WITH NEURAL UNCERTAINTY MAPS FOR 3D RECONSTRUCTION

ICLR 2026poster

Imagine trying to understand the shape of a teapot by viewing it from the front—you might see the spout, but completely miss the handle. Some perspectives naturally provide more information than others. How can an AI system determine which viewpoint offers the most valuable insight for accurate and…

Cited by 0SourcecodeScholar
2025

Gazing at Rewards: Eye Movements as a Lens into Human and AI Decision-Making in Hybrid Visual Foraging

CVPR 2025poster

Imagine searching a collection of coins for quarters (0.25), dimes (0.10), nickels (0.05), and pennies (0.01)--a hybrid foraging task where observers search for multiple instances of multiple target types. In such tasks, how do target values and their prevalence influence foraging and eye movement b…

2025

Make Me Happier: Evoking Emotions Through Image Diffusion Models

ICCV 2025poster

Despite the rapid progress in image generation, emotional image editing remains under-explored. The semantics, context, and structure of an image can evoke emotional responses, making emotional image editing techniques valuable for various real-world applications, including treatment of psychologica…

Cited by 0SourcePDFScholar
2025

Seeing Sound, Hearing Sight: Uncovering Modality Bias and Conflict of AI models in Sound Localization

NeurIPS 2025spotlight

Imagine hearing a dog bark and instinctively turning toward the sound—only to find a parked car, while a silent dog sits nearby. Such moments of sensory conflict challenge perception, yet humans flexibly resolve these discrepancies, prioritizing auditory cues over misleading visuals to accurately lo…

Cited by 0SourcecodeScholar
2025

Unveiling AI's Blind Spots: An Oracle for In-Domain, Out-of-Domain, and Adversarial Errors

ICML 2025poster

AI models make mistakes when recognizing images—whether in-domain, out-of-domain, or adversarial. Predicting these errors is critical for improving system reliability, reducing costly mistakes, and enabling proactive corrections in real-world applications such as healthcare, finance, and autonomous…

2024

Adaptive Visual Scene Understanding: Incremental Scene Graph Generation

NeurIPS 2024poster

Scene graph generation (SGG) analyzes images to extract meaningful information about objects and their relationships. In the dynamic visual world, it is crucial for AI systems to continuously detect new objects and establish their relationships with existing ones. Recently, numerous studies have foc…

2024

Flow Snapshot Neurons in Action: Deep Neural Networks Generalize to Biological Motion Perception

NeurIPS 2024poster

Biological motion perception (BMP) refers to humans' ability to perceive and recognize the actions of living beings solely from their motion patterns, sometimes as minimal as those depicted on point-light displays. While humans excel at these tasks \textit{without any prior training}, current AI mod…

2023

Decoding the Enigma: Benchmarking Humans and AIs on the Many Facets of Working Memory

NeurIPS 2023poster

Working memory (WM), a fundamental cognitive process facilitating the temporary storage, integration, manipulation, and retrieval of information, plays a vital role in reasoning and decision-making tasks. Robust benchmark datasets that capture the multifaceted nature of WM are crucial for the effect…

2023

Label-Efficient Online Continual Object Detection in Streaming Video

ICCV 2023poster

Humans can watch a continuous video stream and effortlessly perform continual acquisition and transfer of new knowledge with minimal supervision yet retaining previously learnt experiences. In contrast, existing continual learning (CL) methods require fully annotated labels to effectively learn from…

Cited by 18PDFcodeScholar
2023

Learning to Learn: How to Continuously Teach Humans and Machines

ICCV 2023poster

Curriculum design is a fundamental component of education. For example, when we learn mathematics at school, we build upon our knowledge of addition to learn multiplication. These and other concepts must be mastered before our first algebra lesson, which also reinforces our addition and multiplicati…

Cited by 5PDFScholar
2023

Object-centric Learning with Cyclic Walks between Parts and Whole

NeurIPS 2023poster

Learning object-centric representations from complex natural environments enables both humans and machines with reasoning abilities from low-level perceptual features. To capture compositional entities of the scene, we proposed cyclic walks between perceptual features extracted from vision transform…

2023

Symbolic Replay: Scene Graph as Prompt for Continual Learning on VQA Task

AAAI 2023technical

VQA is an ambitious task aiming to answer any image-related question. However, in reality, it is hard to build such a system once for all since the needs of users are continuously updated, and the system has to implement new functions. Thus, Continual Learning (CL) ability is a must in developing ad…

2021

Visual Search Asymmetry: Deep Nets and Humans Share Similar Inherent Biases

NeurIPS 2021poster

Visual search is a ubiquitous and often challenging daily task, exemplified by looking for the car keys at home or a friend in a crowd. An intriguing property of some classical search tasks is an asymmetry such that finding a target A among distractors B can be easier than finding B among A. To eluc…

2021

When Pigs Fly: Contextual Reasoning in Synthetic and Natural Scenes

ICCV 2021poster

Context is of fundamental importance to both human and machine vision; e.g., an object in the air is more likely to be an airplane than a pig. The rich notion of context incorporates several aspects including physics rules, statistical co-occurrences, and relative object sizes, among others. While p…

Cited by 25PDFcodeScholar
2017

Deep Future Gaze: Gaze Anticipation on Egocentric Videos Using Adversarial Networks

CVPR 2017oral

We introduce a new problem of gaze anticipation on egocentric videos. This substantially extends the conventional gaze prediction problem to future frames by no longer confining it on the current frame. To solve this problem, we propose a new generative adversarial neural network based model, Deep F…

Cited by 131PDFcodeScholar