← Search

Weitai Kang

7 accepted papers

2026

GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement Learning

ICLR 2026poster

Graphical user interface visual grounding (GUI-VG)—a core capability for GUI agents—has primarily relied on supervised fine-tuning (SFT) of multimodal large language models (MLLMs), demanding extensive data curation and significant training costs. However, as MLLMs continue to advance and even cover…

Cited by 0SourceScholar
2026

VGent: Visual Grounding via Modular Design for Disentangling Reasoning and Prediction

CVPR 2026

Current visual grounding models are either based on a Multimodal Large Language Model (MLLM) that performs auto-regressive decoding, which is slow and risks hallucinations, or on re-aligning an LLM with vision features to learn new special or object tokens for grounding, which may undermine the LLM'

Cited by 0SourceScholar
2025

InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

NeurIPS 2025poster

This paper introduces \textsc{InfantAgent-Next}, a generalist agent capable of interacting with computers in a multimodal manner, encompassing text, images, audio, and video. Unlike existing approaches that either build intricate workflows around a single large model or only provide workflow modular…

Cited by 0SourcecodeScholar
2025

Intent3D: 3D Object Detection in RGB-D Scans Based on Human Intention

ICLR 2025poster

In real-life scenarios, humans seek out objects in the 3D world to fulfill their daily needs or intentions. This inspires us to introduce 3D intention grounding, a new task in 3D object detection employing RGB-D, based on human intention, such as "I want something to support my back." Closely rela…

Cited by 20SourcePDFScholar
2025

Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning

ICCV 2025poster

Recent advancements in 3D Large Language Models (3DLLMs) show their potential to build general-purpose agents in the 3D real world, yet challenges remain due to the lack of high-quality robust instruction-following data, leading to limited discriminative power and generalization of 3DLLMs. In this p…

2024

Token Transformation Matters: Towards Faithful Post-hoc Explanation for Vision Transformer

CVPR 2024poster

While Transformers have rapidly gained popularity in various computer vision applications post-hoc explanations of their internal mechanisms remain largely unexplored. Vision Transformers extract visual information by representing image regions as transformed tokens and integrating them via attentio…

Cited by 9SourcePDFScholar