← Search

Qifan Yu

14 accepted papers

2026

Learning to Adapt: Self-Improving Web Agent via Cognitive-Aware Exploration

CVPR 2026

Recent advances in Multimodal Large Language Models (MLLMs) have led to promising progress in web agents. However, existing web agents often rely on handcrafted execution pipelines or expensive expert trajectories, limiting their adaptability to complex, dynamic environments. To address these challe

Cited by 0SourceScholar
2026

SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning

ICML 2026poster

Progressive Learning (PL) reduces pre-training computational overhead by gradually increasing model scale. While prior work has extensively explored depth expansion, width expansion remains significantly understudied, with the few existing methods limited to the early stages of training. However, ex…

Cited by 0SourceScholar
2026

WiseEdit: Benchmarking Cognition- and Creativity-Informed Image Editing

CVPR 2026

Recent image editing models boast next-level intelligent capabilities, facilitating cognition- and creativity-informed image editing. Yet, existing benchmarks provide too narrow a scope for evaluation, failing to holistically assess these advanced abilities. To address this, we introduce WiseEdit, a

Cited by 0SourcecodeScholar
2025

AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea

CVPR 2025poster

Instruction-based image editing aims to modify specific image elements with natural language instructions. However, current models in this domain often struggle to execute complex user instructions accurately, as they are trained on low-quality data with limited editing types. We present AnyEdit, a…

Cited by 21SourcePDFScholar
2025

Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark

ICML 2025poster

The development of Generalist Virtual Agents (GVAs) has shown significant promise in autonomous task execution. However, current training paradigms face critical limitations, including reliance on outcome supervision and labor-intensive human annotations. To address these challenges, we propose **Si…

2025

EvolvedGRPO: Unlocking Reasoning in LVLMs via Progressive Instruction Evolution

NeurIPS 2025poster

Recent advances in reinforcement learning (RL) methods such as Grouped Relative Policy Optimization (GRPO) have strengthened the reasoning capabilities of Large Vision-Language Models (LVLMs). However, due to the inherent entanglement between visual and textual modalities, applying GRPO to LVLMs oft…

Cited by 0SourcecodeScholar
2025

Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness

ICCV 2025poster

Instruction tuning fine-tunes pre-trained Multi-modal Large Language Models (MLLMs) to handle real-world tasks. However, the rapid expansion of visual instruction datasets introduces data redundancy, leading to excessive computational costs. We propose a collaborative framework, DataTailor, which le…

Cited by 0SourcePDFScholar
2025

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training

CVPR 2025poster

Video Large Language Models (Video-LLMs) have recently shown strong performance in basic video understanding tasks, such as captioning and coarse-grained question answering, but struggle with compositional reasoning that requires multi-step spatio-temporal inference across object relations, interact…

Cited by 4SourcePDFScholar
2025

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities

ICML 2025oral

As multimodal large language models (MLLMs) advance, MLLM-based virtual agents have demonstrated remarkable performance. However, existing benchmarks face significant limitations, including uncontrollable task complexity, extensive manual annotation, and a lack of multidimensional evaluation. In res…

2024

HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data

CVPR 2024poster

Multi-modal Large Language Models (MLLMs) tuned on machine-generated instruction-following data have demonstrated remarkable performance in various multimodal understanding and generation tasks. However the hallucinations inherent in machine-generated data which could lead to hallucinatory outputs i…

2024

Towards Unified Multimodal Editing with Enhanced Knowledge Collaboration

NeurIPS 2024spotlight

The swift advancement in Multimodal LLMs (MLLMs) also presents significant challenges for effective knowledge editing. Current methods, including intrinsic knowledge editing and external knowledge resorting, each possess strengths and weaknesses, struggling to balance the desired properties of relia…

2024

Unified Generative and Discriminative Training for Multi-modal Large Language Models

NeurIPS 2024poster

In recent times, Vision-Language Models (VLMs) have been trained under two predominant paradigms. Generative training has enabled Multimodal Large Language Models (MLLMs) to tackle various complex tasks, yet issues such as hallucinations and weak object discrimination persist. Discriminative trainin…

Cited by 3SourcePDFScholar
2023

Visually-Prompted Language Model for Fine-Grained Scene Graph Generation in an Open World

ICCV 2023poster

Scene Graph Generation (SGG) aims to extract <subject, predicate, object> relationships in images for vision understanding. Although recent works have made steady progress on SGG, they still suffer long-tail distribution that tail-predicates are more costly to train and hard to distinguish due to a…

Cited by 35PDFcodeScholar
2021

Flexoskeleton Fingers: 3D Printed Reconfigurable Ridges Enabling Multi-Functional and Low-Cost Underactuated Grasping

RA-L 2021

In this letter, we present a design and fabrication framework for soft, underactuated grippers that utilize reconfigurable laminate layers for finger stiffness modulation. The grippers consist of internal flexoskeleton layers, which are hybrid soft-rigid structures composed of a flexible thermoplast

Cited by 13SourceScholar