← Search

Ziqi Zhao

12 accepted papers

2026

TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table Recognition

CVPR 2026

Table recognition (TR) aims to transform table images into semi-structured representations such as HTML or Markdown.As a core component of document parsing, TR has long relied on supervised learning, with recent efforts dominated by fine-tuning vision-language models (VLMs) using labeled data.While

Cited by 0SourcecodeScholar
2025

Geminio: Language-Guided Gradient Inversion Attacks in Federated Learning

ICCV 2025poster

Foundation models that bridge vision and language have made significant progress. While they have inspired many life-enriching applications, their potential for abuse in creating new threats remains largely unexplored. In this paper, we reveal that vision-language models (VLMs) can be weaponized to…

2025

ODD: Omni Differential Drive for Simultaneous Reconfiguration and Omnidirectional Mobility of Wheeled Robots

RA-L 2025

Wheeled robots are highly efficient in human living environments. However, conventional wheeled designs, limited by degrees of freedom, struggle to meet varying footprint needs and achieve omnidirectional mobility. This paper proposes a novel robot drive model inspired by human movements, termed as

Cited by 0SourceScholar
2023

Collaborative Trolley Transportation System with Autonomous Nonholonomic Robots

IROS 2023poster

Cooperative object transportation using multiple robots has been intensively studied in the control and robotics literature, but most approaches are either only applicable to omnidirectional robots or lack a complete navigation and decision-making framework that operates in real time. This paper pre…

Cited by 10SourceScholar
2023

Generalizing Few-Shot Named Entity Recognizers to Unseen Domains with Type-Related Features

EMNLP 2023long findings

Few-shot named entity recognition (NER) has shown remarkable progress in identifying entities in low-resource domains. However, few-shot NER methods still struggle with out-of-domain (OOD) examples due to their reliance on manual labeling for the target domain. To address this limitation, recent stu…

Cited by 0SourcecodeScholar
2022

Autonomous Magnetic Navigation Framework for Active Wireless Capsule Endoscopy Inspired by Conventional Colonoscopy Procedures

RA-L 2022

In recent years, simultaneous magnetic actuation and localization (SMAL) for active wireless capsule endoscopy (WCE) has been intensively studied to improve the efficiency and accuracy of gastrointestinal (GI) examination using WCE. In this letter, we draw inspiration from the conventional colonosco

Cited by 19SourceScholar
2022

External and Internal Sensor Fusion Based Localization Strategy for 6-DOF Pose Estimation of a Magnetic Capsule Robot

RA-L 2022

This paper introduces a novel localization approach for active capsule endoscopy that, for the first time, combines external magnetic field sensing and internal inertial sensing to realize 6-DOF pose estimation of a magnetic capsule robot. It utilizes an inertial measurement unit embedded in the cap

Cited by 21SourceScholar
2022

Robotic Autonomous Trolley Collection with Progressive Perception and Nonlinear Model Predictive Control

ICRA 2022poster

Autonomous mobile manipulation robots that can collect trolleys are widely used to liberate human resources and fight epidemics. Most prior robotic trolley collection solutions only detect trolleys with 2D poses or are merely based on spe-cific marks and lack the formal design of planning algorithms…

Cited by 20SourceScholar
2022

Robust Binary Models by Pruning Randomly-initialized Networks

NeurIPS 2022accept

Robustness to adversarial attacks was shown to require a larger model capacity, and thus a larger memory footprint. In this paper, we introduce an approach to obtain robust yet compact models by pruning randomly-initialized binary networks. Unlike adversarial training, which learns the model paramet…

2021

Hybrid Graph Convolutional Networks for Skeleton-Based and EEG-Based Jumping Action Recognition

IROS 2021poster

Kinematic information obtained directly from the skeletal model has been useful for jumping action recognition. Current research focuses on dynamic analysis based on the video stream. Although skeletal data can accurately capture the high-level information of human action, it ignores the brain’s pre…

Cited by 4SourceScholar
2021

Reciprocally Rotating Magnetic Actuation and Automatic Trajectory Following for Wireless Capsule Endoscopy

ICRA 2021poster

Active wireless capsule endoscopy (WCE) under magnetic actuation is a promising technology to reduce the inspection time and relieve the burden of physicians. In this paper, we propose a reciprocally rotating magnetic actuation method for trajectory following of a capsule and develop its dynamic mod…

Cited by 4SourceScholar
2020

Improved Multiple Objects Tracking based Autonomous Simultaneous Magnetic Actuation & Localization for WCE

ICRA 2020poster

Wireless Capsule Endoscopy (WCE) has the advantage of reducing the invasiveness and pain of gastrointestinal examinations. In this work, we propose a system aimed at autonomously accelerating and locating the WCE inside the intestine for clinical applications. A rotating magnet controlled by a robot…

Cited by 20SourceScholar