← Search

Tianyuan Huang

4 accepted papers

2026

IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?

ICLR 2026poster

The webpage-to-code task requires models to understand visual representations of webpages and generate corresponding code. However, existing benchmarks primarily focus on static screenshot-to-code tasks, thereby overlooking the dynamic interactions fundamental to real-world web applications. To addr…

Cited by 0SourcecodeScholar
2025

BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks

EMNLP 2025

Braille plays a vital role in education and information accessibility for visually impaired individuals. However, Braille information processing faces challenges such as data scarcity and ambiguities in mixed-text contexts. We construct English and Chinese Braille Mixed Datasets (EBMD/CBMD) with mat

2024

CityPulse: Fine-Grained Assessment of Urban Change with Street View Time Series

AAAI 2024technical

Urban transformations have profound societal impact on both individuals and communities at large. Accurately assessing these shifts is essential for understanding their underlying causes and ensuring sustainable urban planning. Traditional measurements often encounter constraints in spatial and temp…

Cited by 4SourcePDFScholar
2024

SkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote Sensing

AAAI 2024technical

Remote sensing imagery, despite its broad applications in helping achieve Sustainable Development Goals and tackle climate change, has not yet benefited from the recent advancements of versatile, task-agnostic vision language models (VLMs). A key reason is that the large-scale, semantically diverse…