← Search

Yihao Ding

10 accepted papers

2026

ToolTree: Efficient LLM Tool Planning via Dual-Feedback Monte Carlo Tree Search and Bidirectional Pruning

ICLR 2026poster

Large Language Model (LLM) agents are increasingly applied to complex, multi-step tasks that require interaction with diverse external tools across various domains. However, current LLM agent tool planning methods typically rely on greedy, reactive tool selection strategies that lack foresight and f…

Cited by 0SourceScholar
2025

Beyond Perception: Evaluating Abstract Visual Reasoning through Multi-Stage Task

ACL 2025finding

Current Multimodal Large Language Models (MLLMs) excel in general visual reasoning but remain underexplored in Abstract Visual Reasoning (AVR), which demands higher-order reasoning to identify abstract rules beyond simple perception. Existing AVR benchmarks focus on single-step reasoning, emphasizin…

2025

GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object Detector

CVPR 2025poster

We propose GO-N3RDet, a scene-geometry optimized multi-view 3D object detector enhanced by neural radiance fields. The key to accurate 3D object detection is in effective voxel representation. However, due to occlusion and lack of 3D information, constructing 3D features from multi-view 2D images is…

2025

Natural Language Processing in Support of Evidence-based Medicine: A Scoping Review

ACL 2025finding

Evidence-based medicine (EBM) is at the forefront of modern healthcare, emphasizing the use of the best available scientific evidence to guide clinical decisions. Due to the sheer volume and rapid growth of medical literature and the high cost of curation, there is a critical need to investigate Nat…

Cited by 0SourcePDFScholar
2025

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding

IJCAI 2025

Visually Rich Document Understanding (VRDU) has emerged as a critical field in document intelligence, enabling automated extraction of key information from complex documents across domains such as medical, financial, and educational applications. However, form-like documents pose unique challenges d

Cited by 0SourcePDFScholar
2024

3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding

ACL 2024findings

This paper presents a groundbreaking multimodal, multi-task, multi-teacher joint-grained knowledge distillation model for visually-rich form document understanding. The model is designed to leverage insights from both fine-grained and coarse-grained levels by facilitating a nuanced correlation betwe…

2024

MMVQA: A Comprehensive Dataset for Investigating Multipage Multimodal Information Retrieval in PDF-based Visual Question Answering

IJCAI 2024poster

Document Question Answering (QA) presents a challenge in understanding visually-rich documents (VRD), particularly with lengthy textual content. Existing studies primarily focus on real-world documents with sparse text, while challenges persist in comprehending the hierarchical semantic relations am…

2024

The Language Model Can Have the Personality: Joint Learning for Personality Enhanced Language Model (Student Abstract)

AAAI 2024technical

With the introduction of large language models, chatbots are becoming more conversational to communicate effectively and capable of handling increasingly complex tasks. To make a chatbot more relatable and engaging, we propose a new language model idea that maps the human-like personality. In this p…

Cited by 1SourcePDFScholar
2022

Doc-GCN: Heterogeneous Graph Convolutional Networks for Document Layout Analysis

COLING 2022main

Recognizing the layout of unstructured digital documents is crucial when parsing the documents into the structured, machine-readable format for downstream applications. Recent studies in Document Layout Analysis usually rely on visual cues to understand documents while ignoring other information, su…