← Search

Yue Qiu

16 accepted papers

2026

VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models

AAAI 2026technical

Multimodal Large Language Models (MLLM) have enabled a wide range of advanced vision-language applications, including fine-grained object recognition and contextual understanding. When querying specific regions or objects in an image, human users naturally use "Visual Prompts" (VP) like bounding box

Cited by 0SourcePDFScholar
2025

Training Medical QA Models Based on Mixed Rewards from Multiple-Choice and Open-Ended Questions

EMNLP 2025

Reinforcement learning (RL) for large language models (LLMs) typically requires clear reward signals, which are often unavailable for open-ended (OE) questions where answer evaluation is ambiguous without scalable expert labeling. We investigate whether LLMs benefit from training on mixed data with

Cited by 0SourcePDFScholar
2025

VideoSetDiff: Identifying and Reasoning Similarities and Differences in Similar Videos

ICCV 2025poster

Recognizing subtle similarities and differences among sets of similar activities is central to many real-world applications, including skill acquisition, sports performance evaluation, and anomaly detection. Humans excel at such fine-grained analysis, which requires comprehensive video understanding…

Cited by 0SourcePDFScholar
2024

Conformal Prediction for Deep Classifier via Label Ranking

ICML 2024poster

Conformal prediction is a statistical framework that generates prediction sets containing ground-truth labels with a desired coverage guarantee. The predicted probabilities produced by machine learning models are generally miscalibrated, leading to large prediction sets in conformal prediction. To a…

2024

DailySTR: A Daily Human Activity Pattern Recognition Dataset for Spatio-temporal Reasoning

IROS 2024poster

Recognizing daily human activities is essential for domestic robots to assist humans effectively in indoor environments. These activities typically involve sequences of interactions between humans and objects across different locations and times within a household. Identifying these events and under…

Cited by 0SourceScholar
2024

Decoding Global Preferences: Temporal and Cooperative Dependency Modeling in Multi-Agent Preference-Based Reinforcement Learning

AAAI 2024technical

Designing accurate reward functions for reinforcement learning (RL) has long been challenging. Preference-based RL (PbRL) offers a promising approach by using human preferences to train agents, eliminating the need for manual reward design. While successful in single-agent tasks, extending PbRL to c…

2024

From Text to Trajectory: Exploring Complex Constraint Representation and Decomposition in Safe Reinforcement Learning

NeurIPS 2024poster

Safe reinforcement learning (RL) requires the agent to finish a given task while obeying specific constraints. Giving constraints in natural language form has great potential for practical scenarios due to its flexible transfer capability and accessibility. Previous safe RL methods with natural lang…

Cited by 0SourcePDFScholar
2024

Guided by the Way: The Role of On-the-route Objects and Scene Text in Enhancing Outdoor Navigation

ICRA 2024poster

In outdoor environments, Vision-and-Language Navigation (VLN) requires an agent to rely on multi-modal cues from real-world urban environments and natural language instructions. While existing outdoor VLN models predict actions using a combination of panorama and instruction features, this approach…

Cited by 1SourceScholar
2024

Indoor Scene Change Understanding (SCU): Segment, Describe, and Revert Any Change

IROS 2024poster

Understanding of scene changes is crucial for embodied AI applications, such as visual room rearrangement, where the agent must revert changes by restoring the objects to their original locations or states. Visual changes between two scenes, pre- and post-rearrangement, encompass two tasks: scene ch…

Cited by 2SourceScholar
2024

Subtle-Diff: A Dataset for Precise Recognition of Subtle Differences Among Visually Similar Objects

IROS 2024poster

Visual inspection robots used in factories and outdoor environments require the ability to accurately recognize visual differences between similar objects and further verbalize the recognition results to present the differences to humans. Despite the application of Large Language Models (LLMs) and m…

Cited by 0SourceScholar
2024

The STVchrono Dataset: Towards Continuous Change Recognition in Time

CVPR 2024poster

Recognizing continuous changes offers valuable insights into past historical events supports current trend analysis and facilitates future planning. This knowledge is crucial for a variety of fields such as meteorology and agriculture environmental science urban planning and construction tourism and…

Cited by 8SourcePDFScholar
2023

Graph Representation for Order-Aware Visual Transformation

CVPR 2023poster

This paper proposes a new visual reasoning formulation that aims at discovering changes between image pairs and their temporal orders. Recognizing scene dynamics and their chronological orders is a fundamental aspect of human cognition. The aforementioned abilities make it possible to follow step-by…

Cited by 4SourcePDFScholar
2023

Question Generation for Uncertainty Elimination in Referring Expressions in 3D Environments

ICRA 2023poster

We introduce a new task of question generation to eliminate the uncertainty of referring expressions in 3D indoor environments (3D-REQ). Referring to an object using natural language is one of the most common occurrences in daily human conversations; therefore, instructing robots to identify a certa…

Cited by 2SourceScholar
2023

Towards Long-delayed Sparsity: Learning a Better Transformer through Reward Redistribution

IJCAI 2023poster

Recently, Decision Transformer (DT) pioneered the offline RL into a contextual conditional sequence modeling paradigm, which leverages self-attended autoregression to learn from global target rewards, states, and actions. However, many applications have a severe delay of the above signals, such as t…

2021

Describing and Localizing Multiple Changes With Transformers

ICCV 2021poster

Existing change captioning studies have mainly focused on a single change. However, detecting and describing multiple changed parts in image pairs is essential for enhancing adaptability to complex scenarios. We solve the above issues from three aspects: (i) We propose a simulation-based multi-chang…

Cited by 65PDFScholar