← Search

Zechen Li

8 accepted papers

2026

MetaStreet: Semi-Supervised Multimodal Learning for Street-Level Socioeconomic Prediction

ICML 2026poster

Predicting street-level socioeconomic indicators from street view imagery is fundamental to urban planning. Existing methods typically extract visual features via pretrained encoders and propagate information through graph-based learning, but they fail to fully exploit the structured, task-relevant,…

Cited by 0SourceScholar
2026

Seeing Clearly, Reasoning Confidently: Plug-and-Play Remedies for Vision Language Model Blindness

CVPR 2026

Vision language models (VLMs) have achieved remarkable success in broad visual understanding, yet they remain challenged by object-centric reasoning on rare objects due to the scarcity of such instances in pretraining data. While prior efforts alleviate this issue by retrieving additional data or in

Cited by 0SourcecodeScholar
2025

Cross-City Latent Space Alignment for Consistency Region Embedding

ICML 2025poster

Learning urban region embeddings has substantially advanced urban analysis, but their typical focus on individual cities leads to disparate embedding spaces, hindering cross-city knowledge transfer and the reuse of downstream task predictors. To tackle this issue, we present Consistent Region Embedd…

Cited by 0SourcePDFScholar
2025

PRISM: Pointcloud Reintegrated Inference via Segmentation and Cross-Attention for Manipulation

RA-L 2025

Robust imitation learning for robot manipulation requires comprehensive 3D perception, yet many existing methods struggle in cluttered environments. Fixed camera view approaches are vulnerable to perspective changes, and 3D point cloud techniques often limit themselves to keyframes predictions, redu

Cited by 2SourcecodeScholar
2025

SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition

EMNLP 2025

We introduce SensorLLM, a two-stage framework that enables Large Language Models (LLMs) to perform human activity recognition (HAR) from sensor time-series data. Despite their strong reasoning and generalization capabilities, LLMs remain underutilized for motion sensor data due to the lack of semant

2024

Urban Region Embedding via Multi-View Contrastive Prediction

AAAI 2024technical

Recently, learning urban region representations utilizing multi-modal data (information views) has become increasingly popular, for deep understanding of the distributions of various socioeconomic features in cities. However, previous methods usually blend multi-view information in a posteriors stag…

2022

QLEVR: A Diagnostic Dataset for Quantificational Language and Elementary Visual Reasoning

NAACL 2022findings

Synthetic datasets have successfully been used to probe visual question-answering datasets for their reasoning abilities. CLEVR (John- son et al., 2017), for example, tests a range of visual reasoning abilities. The questions in CLEVR focus on comparisons of shapes, colors, and sizes, numerical reas…

2021

On the Generation of Medical Dialogs for COVID-19

ACL 2021short

Under the pandemic of COVID-19, people experiencing COVID19-related symptoms have a pressing need to consult doctors. Because of the shortage of medical professionals, many people cannot receive online consultations timely. To address this problem, we aim to develop a medical dialog system that can…