← Search

Yi Cai

58 accepted papers

2026

Compositional Transformation Reasoning for Composed Video Retrieval

CVPR 2026

Composed Video Retrieval aims to retrieve a target video given a reference video and a textual modification describing the desired change. The core challenge lies in modeling compositional multimodal transformations, i.e., how entities, actions, and scenes evolve across video and language modalities

Cited by 0SourcecodeScholar
2026

EduDiag: A Benchmark for Educational Diagnostic Reasoning with Error Tracing and Correction on Large Multimodal Models

CVPR 2026

Large multimodal models (LMMs) have achieved impressive performance on multimodal reasoning, becoming crucial technology for the advancement of intelligent question-answering systems. In real-world educational scenarios, effective teaching extends far beyond providing answers. Experienced teachers a

Cited by 0SourceScholar
2026

Hybrid-DMKG: A Hybrid Reasoning Framework over Dynamic Multimodal Knowledge Graphs for Multimodal Multihop QA with Knowledge Editing

AAAI 2026technical

Multimodal Knowledge Editing (MKE) extends traditional knowledge editing to settings involving both textual and visual modalities. However, existing MKE benchmarks primarily assess final answer correctness, neglecting the quality of intermediate reasoning and robustness to visually rephrased inputs.

Cited by 0SourcePDFScholar
2026

Real-Time Millimeter-Accurate Underwater Pose Estimation Via Tightly-Coupled Fusion of Vision and Optical Tracking

ICRA 2026poster

Precise and high-frequency state estimation is required for advanced underwater robotic applications such as physical interaction and agile control, yet no single sensor can simultaneously provide both high accuracy and high update rates. Vision-based methods offer high-frequency updates but suffer …

Cited by 0SourceScholar
2026

Real-Time Millimeter-Accurate Underwater Pose Estimation via Tightly-Coupled Fusion of Vision and Optical Tracking

RA-L 2026

Precise and high-frequency state estimation is required for advanced underwater robotic applications such as physical interaction and agile control, yet no single sensor can simultaneously provide both high accuracy and high update rates. Vision-based methods offer high-frequency updates but suffer

Cited by 0SourceScholar
2026

Rethinking Explanation Evaluation Under the Retraining Scheme

AAAI 2026technical

Feature attribution has gained prominence as a tool for explaining model decisions, yet evaluating explanation quality remains challenging due to the absence of ground-truth explanations. To circumvent this, explanation-guided input manipulation has emerged as an indirect evaluation strategy, measur

Cited by 0SourcePDFScholar
2026

SRACG: A Code Generation Framework with Selective Retrieval Augmentation

AAAI 2026technical

Large Language Models (LLMs) have demonstrated remarkable performance in code generation, offering new possibilities for translating natural language into executable programs. To further enhance LLMs’ code generation capabilities, Retrieval-Augmented Generation (RAG) has emerged as a promising strat

Cited by 0SourcePDFScholar
2026

Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent Games

ICML 2026poster

While Large Language Models (LLMs) excel in certain reasoning tasks, they struggle in multi-agent games where the final outcome depends on the joint strategies of all agents. In multi-agent games, the non-stationarity of other agents brings significant challenges on the evaluation of the reasoning p…

Cited by 0SourceScholar
2025

Ad Hoc Teamwork via Offline Goal-Based Decision Transformers

ICML 2025poster

The ability of agents to collaborate with previously unknown teammates on the fly, known as ad hoc teamwork (AHT), is crucial in many real-world applications. Existing approaches to AHT require online interactions with the environment and some carefully designed teammates. However, these prerequisit…

Cited by 0SourcePDFScholar
2025

CADReview: Automatically Reviewing CAD Programs with Error Detection and Correction

ACL 2025long

Computer-aided design (CAD) is crucial in prototyping 3D objects through geometric instructions (i.e., CAD programs). In practical design workflows, designers often engage in time-consuming reviews and refinements of these prototypes by comparing them with reference images. To bridge this gap, we in…

Cited by 0SourcePDFScholar
2025

Classic4Children: Adapting Chinese Literary Classics for Children with Large Language Model

NAACL 2025findings

Chinese literary classics hold significant cultural and educational value, offering deep insights into morality, history, and human nature. These works often include classical Chinese and complex narratives, making them difficult for children to read. To bridge this gap, we introduce a child-friendl…

Cited by 0SourcePDFScholar
2025

Collaborative Multi-LoRA Experts with Achievement-based Multi-Tasks Loss for Unified Multimodal Information Extraction

IJCAI 2025

Multimodal Information Extraction (MIE) has gained attention for extracting structured information from multimedia sources. Traditional methods tackle MIE tasks separately, missing opportunities to share knowledge across tasks. Recent approaches unify these tasks into a generation problem using inst

2025

Content-free Logical Modification of Large Language Model by Disentangling and Modifying Logic Representation

AAAI 2025technical

Despite extensive training on diverse datasets and alignment with human values, large language models (LLMs) can still generate fallacious outputs. Additionally, the validity of LLM's outputs varies significantly depending on the content. It is crucial to ensure LLMs' logical consistency across diff…

2025

Explicitly Guided Difficulty-Controllable Visual Question Generation

AAAI 2025technical

Visual question generation (VQG) aims to generate questions from images automatically. While existing studies primarily focus on the quality of generated questions, such as fluency and relevance, the difficulty of the questions is also a crucial factor in assessing their quality. Question difficulty…

Cited by 0SourcePDFScholar
2025

Fine-Grained Features-based Code Search for Precise Query-Code Matching

COLING 2025main

Code search aims to quickly locate target code snippets from databases using natural language queries, which promotes code reusability. Existing methods can effectively obtain aligned token-level and query word-level features. However, these studies usually represent the semantics of code and query…

Cited by 1SourcePDFScholar
2025

GEFA: A General Feature Attribution Framework Using Proxy Gradient Estimation

ICML 2025poster

Feature attribution explains machine decisions by quantifying each feature's contribution. While numerous approaches rely on exact gradient measurements, recent work has adopted gradient estimation to derive explanatory information under query-level access, a restrictive yet more practical accessibi…

Cited by 0SourcePDFScholar
2025

Look Around Before Locating: Considering Content and Structure Information for Visual Grounding

AAAI 2025technical

As a long-term challenge and fundamental requirement in vision and language tasks, visual grounding aims to localize a target referred by a natural language query. The regional annotations form a superficial correlation between the subject of expression and some common visual entities, which hinder…

2025

MTRec: Learning to Align with User Preferences via Mental Reward Models

NeurIPS 2025poster

Recommendation models are predominantly trained using implicit user feedback, since explicit feedback is often costly to obtain. However, implicit feedback, such as clicks, does not always reflect users' real preferences. For example, a user might click on a news article because of its attractive he…

Cited by 0SourceScholar
2025

Motion Planning and Compensation Approaches for Autonomous Surface Manipulator Systems in Grasping Tasks on Water Surfaces

RA-L 2025

Autonomous surface manipulation systems (ASMSs) are novel robotics platforms composed of unmanned surface vehicles (USVs) and manipulators, and they can be used to recover floating objects on water surfaces. However, the improper positional relationship between the target object and the USV, and the

Cited by 1SourceScholar
2025

RMG: Real-Time Expressive Motion Generation with Self-collision Avoidance for 6-DOF Companion Robotic Arms

IROS 2025

The six-degree-of-freedom (6-DOF) robotic arm has gained widespread application in human-coexisting environments. While previous research has predominantly focused on functional motion generation, the critical aspect of expressive motion in human-robot interaction remains largely unexplored. This pa

Cited by 0SourceScholar
2025

RTADev: Intention Aligned Multi-Agent Framework for Software Development

ACL 2025finding

LLM-based Multi-agent frameworks have shown a great potential in solving real-world software development tasks, where the agents of different roles can communicate much more efficiently than humans. Despite their efficiency, LLM-based agents can hardly fully understand each other, which frequently c…

2025

Rethinking-based Code Summarization with Chain of Comments

COLING 2025main

Automatic code summarization aims to generate concise natural language descriptions (summary) for source code, which can free software developers from the heavy burden of manual commenting and software maintenance. Existing methods focus on learning a direct mapping from pure code to summaries, over…

2025

RuleEdit: Towards Rule-Level Knowledge Generalization to Mitigate Over-Editing in Large Language Models

ACL 2025finding

Knowledge editing emerges as a promising approach for updating target knowledge in Large Language Models (LLMs) in a timely manner, thereby preventing undesirable behaviors stemming from outdated, inaccurate, or incomplete knowledge. However, existing methods mainly focus on instance-level editing,…

Cited by 0SourcePDFScholar
2025

Sequence Structure Aware Retriever for Procedural Document Retrieval: A New Dataset and Baseline

EMNLP 2025

Execution failures are common in daily life when individuals perform procedural tasks, such as cooking or handicrafts making. Retrieving relevant procedural documents that align closely with both the content of steps and the overall execution sequence can help correct these failures with fewer modif

2025

Sound Bridge: Associating Egocentric and Exocentric Videos via Audio Cues

CVPR 2025poster

Understanding human behavior and the environmental information in the egocentric video is very challenging due to the invisibility of some actions (e.g., laughing and sneezing) and the local nature of the first-person view. Leveraging the corresponding exocentric video to provide global context has…

2025

Walk in Others’ Shoes with a Single Glance: Human-Centric Visual Grounding with Top-View Perspective Transformation

ACL 2025long

Visual perspective-taking, an ability to envision others’ perspectives from a single self-perspective, is vital in human-robot interactions. Thus, we introduce a human-centric visual grounding task and a dataset to evaluate this ability. Recent advances in vision-language models (VLMs) have shown po…

2024

A Logical Pattern Memory Pre-trained Model for Entailment Tree Generation

COLING 2024main

Generating coherent and credible explanations remains a significant challenge in the field of AI. In recent years, researchers have delved into the utilization of entailment trees to depict explanations, which exhibit a reasoning process of how a hypothesis is deduced from the supporting facts. Howe…

2024

Automated Defect Report Generation for Enhanced Industrial Quality Control

AAAI 2024technical

Defect detection is a pivotal aspect ensuring product quality and production efficiency in industrial manufacturing. Existing studies on defect detection predominantly focus on locating defects through bounding boxes and classifying defect types. However, their methods can only provide limited infor…

Cited by 4SourcePDFScholar
2024

Beyond Code: Evaluate Thought Steps for Complex Code Generation

COLING 2024main

Code generation aims to generate code in a general-purpose programming language, such as C++, based on natural language intents. Existing efforts primarily focus on relatively simple programming problems and fail to evaluate the thought process involved in complex programming scenarios. In this pape…

2024

Grounded Multimodal Procedural Entity Recognition for Procedural Documents: A New Dataset and Baseline

COLING 2024main

Much of commonsense knowledge in real world is the form of procudures or sequences of steps to achieve particular goals. In recent years, knowledge extraction on procedural documents has attracted considerable attention. However, they often focus on procedural text but ignore a common multimodal sce…

Cited by 1SourcePDFScholar
2024

Knowledge-Guided Cross-Topic Visual Question Generation

COLING 2024main

Visual question generation (VQG) task aims to generate high-quality questions based on the input image. Current methods primarily focus on generating questions containing specified content utilizing answers or question types as constraints. However, these constraints make it challenging to control t…

Cited by 2SourcePDFScholar
2024

On Gradient-like Explanation under a Black-box Setting: When Black-box Explanations Become as Good as White-box

ICML 2024poster

Attribution methods shed light on the explainability of data-driven approaches such as deep learning models by uncovering the most influential features in a to-be-explained decision. While determining feature attributions via gradients delivers promising results, the internal access required for acq…

2024

PoRank: A Practical Framework for Learning to Rank Policies

IJCAI 2024poster

In many real-world scenarios, we need to select from a set of candidate policies before online deployment. Although existing Off-policy evaluation (OPE) methods can be used to estimate the online performance, they suffer from high variance. Fortunately, we care only about the ranking of the candidat…

2024

Rethinking Two-Stage Referring Expression Comprehension: A Novel Grounding and Segmentation Method Modulated by Point

AAAI 2024technical

As a fundamental and challenging task in the vision and language domain, Referring Expression Comprehension (REC) has shown impressive improvements recently. However, for a complex task that couples the comprehension of abstract concepts and the localization of concrete instances, one-stage approach…

2023

CLEVR-Implicit: A Diagnostic Dataset for Implicit Reasoning in Referring Expression Comprehension

EMNLP 2023long main

Recently, pre-trained vision-language (VL) models have achieved remarkable success in various cross-modal tasks, including referring expression comprehension (REC). These models are pre-trained on the large-scale image-text pairs to learn the alignment between words in textual descriptions and objec…

Cited by 0SourceScholar
2023

Category-Guided Visual Question Generation (Student Abstract)

AAAI 2023technical

Visual question generation aims to generate high-quality questions related to images. Generating questions based only on images can better reduce labor costs and thus be easily applied. However, their methods tend to generate similar general questions that fail to ask questions about the specific co…

Cited by 1SourcePDFScholar
2023

Constructing Procedural Graphs with Multiple Dependency Relations: A New Dataset and Baseline

ACL 2023findings

Current structured and semi-structured knowledge bases mainly focus on representing descriptive knowledge but ignore another commonsense knowledge (Procedural Knowledge). To structure the procedural knowledge, existing methods are proposed to automatically generate flow graphs from procedural docume…

2023

Ensemble-in-One: Ensemble Learning within Random Gated Networks for Enhanced Adversarial Robustness

AAAI 2023technical

Adversarial attacks have threatened modern deep learning systems by crafting adversarial examples with small perturbations to fool the convolutional neural networks (CNNs). To alleviate that, ensemble training methods are proposed to facilitate better adversarial robustness by diversifying the vulne…

2023

Improving Named Entity Recognition via Bridge-based Domain Adaptation

ACL 2023findings

Recent studies have shown remarkable success in cross-domain named entity recognition (cross-domain NER). Despite the promising results, existing methods mainly utilize pre-training language models like BERT to represent words. As such, the original chaotic representations may challenge them to dist…

2023

Joint Multimodal Entity-Relation Extraction Based on Edge-Enhanced Graph Alignment Network and Word-Pair Relation Tagging

AAAI 2023technical

Multimodal named entity recognition (MNER) and multimodal relation extraction (MRE) are two fundamental subtasks in the multimodal knowledge graph construction task. However, the existing methods usually handle two tasks independently, which ignores the bidirectional interaction between them. This p…

2023

Linking People across Text and Images Based on Social Relation Reasoning

AAAI 2023technical

As a sub-task of visual grounding, linking people across text and images aims to localize target people in images with corresponding sentences. Existing approaches tend to capture superficial features of people (e.g., dress and location) that suffer from the incompleteness information across text an…

2023

Memory-Oriented Structural Pruning for Efficient Image Restoration

AAAI 2023technical

Deep learning (DL) based methods have significantly pushed forward the state-of-the-art for image restoration (IR) task. Nevertheless, DL-based IR models are highly computation- and memory-intensive. The surging demands for processing higher-resolution images and multi-task paralleling in practical…

Cited by 4SourcePDFScholar
2023

Rethinking Multimodal Entity and Relation Extraction from a Translation Point of View

ACL 2023long

We revisit the multimodal entity and relation extraction from a translation point of view. Special attention is paid on the misalignment issue in text-image datasets which may mislead the learning. We are motivated by the fact that the cross-modal misalignment is a similar problem of cross-lingual d…

2023

Scene Graph Enhanced Pseudo-Labeling for Referring Expression Comprehension

EMNLP 2023long findings

Referring Expression Comprehension (ReC) is a task that involves localizing objects in images based on natural language expressions. Most ReC methods typically approach the task as a supervised learning problem. However, the need for costly annotations, such as clear image-text pairs or region-text…

Cited by 0SourceScholar
2023

Segment-Level and Category-Oriented Network for Knowledge-Based Referring Expression Comprehension

ACL 2023findings

Knowledge-based referring expression comprehension (KB-REC) aims to identify visual objects referred to by expressions that incorporate knowledge. Existing methods employ sentence-level retrieval and fusion methods, which may lead to issues of similarity bias and interference from irrelevant informa…

2022

CLOSE: Curriculum Learning on the Sharing Extent towards Better One-Shot NAS

ECCV 2022poster

"One-shot Neural Architecture Search (NAS) has been widely used to discover architectures due to its efficiency. However, previous studies reveal that one-shot performance estimations of architectures might not be well correlated with their performances in stand-alone training because of the excessi…

2022

Towards Exploiting Sticker for Multimodal Sentiment Analysis in Social Media: A New Dataset and Baseline

COLING 2022main

Sentiment analysis in social media is challenging since posts are short of context. As a popular way to express emotion on social media, stickers related to these posts can supplement missing sentiments and help identify sentiments precisely. However, research about stickers has not been investigate…

2021

Entity Guided Question Generation with Contextual Structure and Sequence Information Capturing

AAAI 2021technical

Question generation is a challenging task and has attracted widespread attention in recent years. Although previous studies have made great progress, there are still two main shortcomings: First, previous work did not simultaneously capture the sequence information and structure information hidden i…

2021

Story Ending Generation with Multi-Level Graph Convolutional Networks over Dependency Trees

AAAI 2021technical

As an interesting and challenging task, story ending generation aims at generating a reasonable and coherent ending for a given story context. The key challenge of the task is to comprehend the context sufficiently and capture the hidden logic information effectively, which has not been well explore…

2020

A Two-phase Prototypical Network Model for Incremental Few-shot Relation Classification

COLING 2020main

Relation Classification (RC) plays an important role in natural language processing (NLP). Current conventional supervised and distantly supervised RC models always make a closed-world assumption which ignores the emergence of novel relations in open environment. To incrementally recognize the novel…

2020

Controllable Abstractive Sentence Summarization with Guiding Entities

COLING 2020main

Entities are the major proportion and build up the topic of text summaries. Although existing text summarization models can produce promising results of automatic metrics, for example, ROUGE, it is difficult to guarantee that an entity is contained in generated summaries. In this paper, we propose a…