← Search

Shiqi Zhang

33 accepted papers

2026

Air-Know: Arbiter-Calibrated Knowledge-Internalizing Robust Network for Composed Image Retrieval

CVPR 2026

Composed Image Retrieval (CIR) has attracted significant attention due to its flexible multimodal query method, yet its development is severely constrained by the Noisy Triplet Correspondence (NTC) problem. Most existing robust learning methods rely on the "small loss hypothesis", but the unique sem

Cited by 0SourcecodeScholar
2026

CausalX: A Unified and Causally-Interpretable Plug-and-Play Model for Multi-modal Spatio-Temporal Forecasting

ICML 2026poster

Multi-modal spatio-temporal forecasting underpins many real-world applications but remains challenging due to the complex and evolving interactions across modalities and time steps. Moreover, the lack of interpretability in existing models limits their reliability in safety-critical scenarios. In th…

Cited by 0SourceScholar
2026

From Woofs to Words: Towards Intelligent Robotic Guide Dogs with Verbal Communication

AAAI 2026technical

Assistive robotics is an important subarea of robotics that focuses on the well-being of people with disabilities. A robotic guide dog is an assistive quadruped robot for assisting visually impaired people in obstacle avoidance and navigation. Enabling language capabilities on robotic guide dogs goe

Cited by 0SourcePDFScholar
2026

HABIT: Chrono-Synergia Robust Progressive Learning Framework for Composed Image Retrieval

AAAI 2026technical

Composed Image Retrieval (CIR) is a flexible image retrieval paradigm that enables users to accurately locate the target image through a multimodal query composed of a reference image and modification text. Although this task has demonstrated promising applications in personalized search and recomme

Cited by 0SourcePDFScholar
2026

M$^2$-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Mining

ICLR 2026poster

Graphical User Interface (GUI) agent is pivotal to advancing intelligent human-computer interaction paradigms. Constructing powerful GUI agents necessitates the large-scale annotation of high-quality user-behavior trajectory data (\textit{i.e.}, intent–trajectory pairs) for training. However, manual…

Cited by 0SourceScholar
2025

Data Quality Enhancement on the Basis of Diversity with Large Language Models for Text Classification: Uncovered, Difficult, and Noisy

COLING 2025main

In recent years, the use of large language models (LLMs) for text classification has attracted widespread attention. Despite this, the classification accuracy of LLMs has not yet universally surpassed that of smaller models. LLMs can enhance their performance in text classification through fine-tuni…

Cited by 1SourcePDFScholar
2025

GTA: Supervised-Guided Reinforcement Learning for Text Classification with Large Language Models

EMNLP 2025

In natural language processing (NLP) tasks, pure reinforcement learning fine-tuning methods often suffer from inefficient exploration and slow convergence; while supervised fine-tuning (SFT) methods, although efficient in training, have limited performance ceiling and less solid theoretical foundati

Cited by 0SourcePDFScholar
2025

IDOL: Meeting Diverse Distribution Shifts with Prior Physics for Tropical Cyclone Multi-Task Estimation

NeurIPS 2025poster

Tropical Cyclone (TC) estimation aims to accurately estimate various TC attributes in real time. However, distribution shifts arising from the complex and dynamic nature of TC environmental fields, such as varying geographical conditions and seasonal changes, present significant challenges to reliab…

Cited by 0SourceScholar
2025

Multi-Modal Entities Matter: Benchmarking Multi-Modal Entity Alignment

COLING 2025main

Multi-modal entity alignment (MMEA) is a long-standing task that aims to discover identical entities between different multi-modal knowledge graphs (MMKGs). However, most of the existing MMEA datasets consider the multi-modal data as the attributes of textual entities, while neglecting the correlati…

Cited by 0SourcePDFScholar
2025

OG-Gaussian: Occupancy Based Street Gaussians for Autonomous Driving

ICRA 2025

Accurate and realistic 3D scene reconstruction enables the lifelike creation of autonomous driving simulation environments. With advancements in 3D Gaussian Splatting (3DGS), previous studies have applied it to reconstruct complex dynamic driving scenes. These methods typically require expensive LiD

Cited by 5SourceScholar
2025

ORLA*: Mobile Manipulator-Based Object Rearrangement with Lazy A

ICRA 2025

Effectively performing object rearrangement is an essential skill for mobile manipulators, e.g., setting up a dinner table. A key challenge in such problems is deciding an appropriate ordering to effectively untangle object-object dependencies while considering the necessary motions for realizing ma

Cited by 10SourcecodeScholar
2025

TC-Diffuser: Bi-Condition Multi-Modal Diffusion for Tropical Cyclone Forecasting

AAAI 2025technical

Tropical cyclones (TCs) are complex weather systems with strong winds and heavy rainfall, causing substantial loss of life and property. Therefore, accurate TC forecasting is crucial for the effective prevention of disasters caused by TCs. TC forecasting can be regarded as a spatio-temporal predicti…

2024

OpenEQA: Embodied Question Answering in the Era of Foundation Models

CVPR 2024poster

We present a modern formulation of Embodied Question Answering (EQA) as the task of understanding an environment well enough to answer questions about it in natural language. An agent can achieve such an understanding by either drawing upon episodic memory exemplified by agents on smart glasses or b…

Cited by 118SourcePDFScholar
2024

Unrestricted Global Phase Bias-Aware Single-Channel Speech Enhancement with Conformer-Based Metric Gan

ICASSP 2024accepted

With the rapid development of neural networks in recent years, the ability of various networks to enhance the magnitude spectrum of noisy speech in the single-channel speech enhancement domain has become exceptionally outstanding. However, enhancing the phase spectrum using neural networks is often…

Cited by 0SourceScholar
2023

Learning to reason about contextual knowledge for planning under uncertainty

UAI 2023poster

Sequential decision-making (SDM) methods enable AI agents to compute an action policy toward achieving long-term goals under uncertainty. Existing research has shown that contextual knowledge in declarative forms can be used for improving the performance of SDM methods. However, the contextual knowl…

Cited by 0SourcePDFScholar
2023

Symbolic State Space Optimization for Long Horizon Mobile Manipulation Planning

IROS 2023poster

In existing task and motion planning (TAMP) research, it is a common assumption that experts manually specify the state space for task-level planning. A well-developed state space enables the desirable distribution of limited computational resources between task planning and motion planning. However…

Cited by 6SourceScholar
2023

Task and Motion Planning with Large Language Models for Object Rearrangement

IROS 2023poster

Multi-object rearrangement is a crucial skill for service robots, and commonsense reasoning is frequently needed in this process. However, achieving commonsense arrangements requires knowledge about objects, which is hard to transfer to robots. Large language models (LLMs) are one potential source o…

Cited by 192SourceScholar
2022

Efficient Dialog Policy Learning by Reasoning with Contextual Knowledge

AAAI 2022technical

Goal-oriented dialog policy learning algorithms aim to learn a dialog policy for selecting language actions based on the current dialog state. Deep reinforcement learning methods have been used for dialog policy learning. This work is motivated by the observation that, although dialog is a domain wi…

2022

Goal-oriented Vision-and-Dialog Navigation via Reinforcement Learning

EMNLP 2022finding

Vision-and-dialog navigation is a recent benchmark for evaluating the AI capabilities of perception, interaction, and decision making. While existing methods developed for this benchmark have demonstrated great successes, they mostly rely on large datasets, where data collection can be a challenge,…

Cited by 3SourcePDFScholar
2022

Learning Visualization Policies of Augmented Reality for Human-Robot Collaboration

CoRL 2022poster

In human-robot collaboration domains, augmented reality (AR) technologies have enabled people to visualize the state of robots. Current AR-based visualization policies are designed manually, which requires a lot of human efforts and domain knowledge. When too little information is visualized, human…

Cited by 7SourceScholar
2022

Visually Grounded Task and Motion Planning for Mobile Manipulation

ICRA 2022poster

Task and motion planning (TAMP) algorithms aim to help robots achieve task-level goals, while maintaining motion-level feasibility. This paper focuses on TAMP domains that involve robot behaviors that take extended periods of time (e.g., long-distance navigation). In this paper, we develop a visual…

Cited by 32SourceScholar
2021

ARROCH: Augmented Reality for Robots Collaborating with a Human

ICRA 2021poster

Human-robot collaboration frequently requires extensive communication, e.g., using natural language and gesture. Augmented reality (AR) has provided an alternative way of bridging the communication gap between robots and people. However, most current AR-based human-robot communication methods are un…

Cited by 45SourceScholar
2021

Label-Enhanced Hierarchical Contextualized Representation for Sequential Metaphor Identification

EMNLP 2021main

Recent metaphor identification approaches mainly consider the contextual text features within a sentence or introduce external linguistic features to the model. But they usually ignore the extra information that the data can provide, such as the contextual metaphor information and broader discourse…

Cited by 7SourcePDFScholar
2021

Learning to Guide Human Attention on Mobile Telepresence Robots with 360° Vision

IROS 2021poster

Mobile telepresence robots (MTRs) allow people to navigate and interact with a remote environment that is in a place other than the person’s true location. Thanks to the recent advances in 360° vision, many MTRs are now equipped with an all-degree visual perception capability. However, people’s visu…

Cited by 11SourceScholar
2021

Planning Multimodal Exploratory Actions for Online Robot Attribute Learning

RSS 2021poster

Robots frequently need to perceive object attributes; such as "red;" "heavy;" and "empty;" using multimodal exploratory actions; such as "look;" "lift;" and "shake." Robot attribute learning algorithms aim to learn an observation model for each perceivable attribute given an exploratory action. Once…

Cited by 4SourcePDFScholar
2019

Augmenting Knowledge through Statistical, Goal-oriented Human-Robot Dialog

IROS 2019poster

Some robots can interact with humans using natural language, and identify service requests through human-robot dialog. However, few robots are able to improve their language capabilities from this experience. In this paper, we develop a dialog agent for robots that is able to interpret user commands…

Cited by 30SourceScholar
2019

Task-Motion Planning with Reinforcement Learning for Adaptable Mobile Service Robots

IROS 2019poster

Task-motion planning (TMP) addresses the problem of efficiently generating executable and low-cost task plans in a discrete space such that the (initially unknown) action costs are determined by motion plans in a corresponding continuous space. A task-motion plan for a mobile service robot that beha…

Cited by 44SourceScholar
2017

Leveraging commonsense reasoning and multimodal perception for robot spoken dialog systems

IROS 2017poster

Probabilistic graphical models, such as partially observable Markov decision processes (POMDPs), have been used in stochastic spoken dialog systems to handle the inherent uncertainty in speech recognition and language understanding. Such dialog systems suffer from the fact that only a relatively sma…

Cited by 19SourceScholar