← Search

Zonglin Di

8 accepted papers

2026

CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation

CVPR 2026

Vision-Language Models (VLMs) have shown remarkable progress in Vision-Language Navigation (VLN), offering new possibilities for navigation decision-making that could benefit both robotic platforms and human users. However, real-world navigation is inherently conditioned by the agent's mobility cons

Cited by 0SourcecodeScholar
2026

Label Smoothing Improves Machine Unlearning

ICLR 2026poster

The objective of machine unlearning (MU) is to eliminate previously learned data from a model. However, it can be challenging to strike a balance between computation cost and performance when using existing MU techniques. Taking inspiration from the influence of label smoothing on model confidence a…

Cited by 0SourceScholar
2025

DiffTell: A High-Quality Dataset for Describing Image Manipulation Changes

ICCV 2025poster

The image difference captioning (IDC) task is to describe the distinctions between two images. However, existing datasets do not offer comprehensive coverage across all image-difference categories. In this work, we introduce a high-quality dataset, DiffTell with various types of image manipulations,…

Cited by 0SourcePDFScholar
2024

Navigation as Attackers Wish? Towards Building Robust Embodied Agents under Federated Learning

NAACL 2024long

Federated embodied agent learning protects the data privacy of individual visual environments by keeping data locally at each client (the individual environment) during training. However, since the local data is inaccessible to the server under federated learning, attackers may easily poison the tra…

Cited by 2SourcePDFScholar
2023

T2IAT: Measuring Valence and Stereotypical Biases in Text-to-Image Generation

ACL 2023findings

*Warning: This paper contains several contents that may be toxic, harmful, or offensive.*In the last few years, text-to-image generative models have gained remarkable success in generating images with unprecedented quality accompanied by a breakthrough of inference speed. Despite their rapid progres…

Cited by 32SourcePDFScholar
2021

Test-Time Personalization with a Transformer for Human Pose Estimation

NeurIPS 2021poster

We propose to personalize a 2D human pose estimator given a set of test images of a person without using any manual annotations. While there is a significant advancement in human pose estimation, it is still very challenging for a model to generalize to different unknown environments and unseen pers…

2019

Leveraging Crowdsourced GPS Data for Road Extraction From Aerial Imagery

CVPR 2019poster

Deep learning is revolutionizing the mapping industry. Under lightweight human curation, computer has generated almost half of the roads in Thailand on Open- StreetMap (OSM) using high resolution aerial imagery. Bing maps are displaying 125 million computer generated building polygons in the U.S. Wh…

Cited by 119PDFScholar