← Search

Peng Wei

8 accepted papers

2026

DomainCQA: Crafting Knowledge-Intensive QA from Domain-Specific Charts

AAAI 2026technical

Chart Question Answering (CQA) evaluates Multimodal Large Language Models (MLLMs) on visual understanding and reasoning over chart data. However, existing benchmarks mostly test surface-level parsing, such as reading labels and legends, while overlooking deeper scientific reasoning. We propose Domai

Cited by 0SourcePDFScholar
2025

Efficient and Safe Trajectory Planning for Autonomous Agricultural Vehicle Headland Turning in Cluttered Orchard Environments

RA-L 2025

Autonomous agricultural vehicles (AAVs), including field robots and autonomous tractors, are becoming essential in modern farming by improving efficiency and reducing labor costs. A critical task in AAV operations is headland turning between crop rows. This task is challenging in orchards with limit

Cited by 7SourceScholar
2025

HIRAG: Hierarchical-Thought Instruction-Tuning Retrieval-Augmented Generation

EMNLP 2025

Retrieval-augmented generation (RAG) has become a fundamental paradigm for addressing the challenges faced by large language models in handling real-time information and domain-specific problems. Traditional RAG systems primarily rely on the in-context learning (ICL) capabilities of the large langua

Cited by 0SourcePDFScholar
2024

Learning to Plan for Retrieval-Augmented Large Language Models from Knowledge Graphs

EMNLP 2024finding

Improving the performance of large language models (LLMs) in complex question-answering (QA) scenarios has always been a research focal point. Recent studies have attempted to enhance LLMs’ performance by combining step-wise planning with external retrieval. While effective for advanced models like…

2023

Flowpose: Conditional Normalizing Flows for 3D Human Pose and Shape Estimation from Monocular Videos

ICASSP 2023accepted

Human motion modeling is essential for video-based 3D human pose and shape estimation. Most existing methods model human motion by learning a deterministic mapping from the input videos to the human body parameters, while the uncertainties such as occlusions and depth ambiguities are ignored. To add…

Cited by 0SourceScholar
2020

Spatio-Temporal and Geometry Constrained Network for Automobile Visual Odometry

ICASSP 2020accepted

Visual odometry (VO) is an essence of vision-based localization and mapping system where existing learning-based approaches utilize CNN and RNN to model camera motion and gain promising results. However, these methods lack full use of the relationship between spatial characteristics and temporal clu…

Cited by 0SourceScholar
2020

Unsupervised Monocular Visual-inertial Odometry Network

IJCAI 2020poster

Recently, unsupervised methods for monocular visual odometry (VO), with no need for quantities of expensive labeled ground truth, have attracted much attention. However, these methods are inadequate for long-term odometry task, due to the inherent limitation of only using monocular visual data and t…