← Search

Zitong Huang

6 accepted papers

2026

CGL: Advancing Continual GUI Learning via Reinforcement Fine-Tuning

CVPR 2026

Graphical User Interface (GUI) Agents, benefiting from recent advances in multimodal large language models (MLLM), have achieved significant development. However, due to the frequent updates of GUI applications, adapting to new tasks without forgetting old tasks in GUI continual learning remains an

Cited by 0SourceScholar
2026

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models

CVPR 2026

Aligning text-to-video diffusion models with human preferences is crucial for generating high-quality videos. Existing Direct Preference Otimization (DPO) methods rely on multi-sample ranking and task-specific critic models, which is inefficient and often yields ambiguous global supervision. To addr

Cited by 0SourcecodeScholar
2025

MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM

NeurIPS 2025poster

Multimodal hallucination in multimodal large language models (MLLMs) restricts the correctness of MLLMs. However, multimodal hallucinations are multi-sourced and arise from diverse causes. Existing benchmarks fail to adequately distinguish between perception-induced hallucinations and reasoning-indu…

Cited by 0SourceScholar
2023

ImaginaryNet: Learning Object Detectors without Real Images and Annotations

ICLR 2023poster

Without the demand of training in reality, humans are able of detecting a new category of object simply based on the language description on its visual characteristics. Empowering deep learning with this ability undoubtedly enables the neural network to handle complex vision tasks, e.g., object dete…

2022

W2N: Switching from Weak Supervision to Noisy Supervision for Object Detection

ECCV 2022poster

"Weakly-supervised object detection (WSOD) aims to train an object detector only requiring the image-level annotations. Recently, some works have managed to select the accurate boxes generated from a well-trained WSOD network to supervise a semi-supervised detection framework for better performance.…

2021

Boosting Weakly Supervised Object Detection via Learning Bounding Box Adjusters

ICCV 2021poster

Weakly-supervised object detection (WSOD) has emerged as an inspiring recent topic to avoid expensive instance-level object annotations. However, the bounding boxes of most existing WSOD methods are mainly determined by precomputed proposals, thereby being limited in precise object localization. In…

Cited by 60PDFcodeScholar