← Search

Xuejing Liu

7 accepted papers

2026

Design of Bio-manta Based on Continuously Programmable Soft Pneumatic Actuators (CPSPAs)

RA-L 2026

Bionic autonomous underwater vehicles (AUVs) have been prevalent in underwater equipment for their excel features. Manta rays have been widely studied due to their highly efficient and maneuverable pectoral fin propulsion. Traditional single-hinge actuation methods have impeded bionic pectoral fin a

Cited by 0SourceScholar
2026

Revisiting Multimodal Positional Encoding in Vision–Language Models

ICLR 2026poster

Multimodal position encoding is essential for vision-language models, yet there has been little systematic investigation into multimodal position encoding. We conduct a comprehensive analysis of multimodal Rotary Positional Embedding (RoPE) by examining its two core components: position design and f…

Cited by 0SourcecodeScholar
2025

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

ICCV 2025poster

Large Multimodal Models (LMMs) have demonstrated impressive performance in recognizing document images with natural language instructions. However, it remains unclear to what extent capabilities in literacy with rich structure and fine-grained visual challenges. The current landscape lacks a compreh…

Cited by 0SourcePDFScholar
2024

PaDeLLM-NER: Parallel Decoding in Large Language Models for Named Entity Recognition

NeurIPS 2024poster

In this study, we aim to reduce generation latency for Named Entity Recognition (NER) with Large Language Models (LLMs). The main cause of high latency in LLMs is the sequential decoding process, which autoregressively generates all labels and mentions for NER, significantly increase the sequence le…

2023

Deeply Coupled Cross-Modal Prompt Learning

ACL 2023findings

Recent advancements in multimodal foundation models (e.g., CLIP) have excelled in zero-shot generalization. Prompt tuning involved in the knowledge transfer from foundation models to downstream tasks has gained significant attention recently. Existing prompt-tuning methods in cross-modal learning, h…

2020

Parsing-Based View-Aware Embedding Network for Vehicle Re-Identification

CVPR 2020poster

Vehicle Re-Identification is to find images of the same vehicle from various views in the cross-camera scenario. The main challenges of this task are the large intra-instance distance caused by different views and the subtle inter-instance discrepancy caused by similar vehicles. In this paper, we pr…

Cited by 255PDFcodeScholar
2019

Adaptive Reconstruction Network for Weakly Supervised Referring Expression Grounding

ICCV 2019poster

Weakly supervised referring expression grounding aims at localizing the referential object in an image according to the linguistic query, where the mapping between the referential object and query is unknown in the training stage. To address this problem, we propose a novel end-to-end adaptive recon…

Cited by 110PDFcodeScholar