← Search

Jingqi Tian

3 accepted papers

2025

Ponder & Press: Advancing Visual GUI Agent towards General Computer Control

ACL 2025finding

Most existing GUI agents typically depend on non-vision inputs like HTML source code or accessibility trees, limiting flexibility across diverse software environments and platforms. Current multimodal large language models (MLLMs), though excel at using vision to ground real-world objects, often str…

Cited by 0SourcePDFScholar
2020

Social Data Assisted Multi-Modal Video Analysis For Saliency Detection

ICASSP 2020accepted

Video saliency should be taken into consideration to facilitate optimization of the end-to-end video production, delivery and consumption ecosystem to improve user experience at lowered cost. Although recent studies have significantly increased the accuracy of saliency prediction, the approaches are…

Cited by 0SourceScholar
2019

Information Entropy Based Feature Pooling for Convolutional Neural Networks

ICCV 2019poster

In convolutional neural networks (CNNs), we propose to estimate the importance of a feature vector at a spatial location in the feature maps by the network's uncertainty on its class prediction, which can be quantified using the information entropy. Based on this idea, we propose the entropy-based f…

Cited by 40PDFScholar