← Search

Zihang Wang

6 accepted papers

2026

LiteFT-PR: Lightweight and Fault-Tolerant LiDAR-Camera Fusion Network for Robust Place Recognition via Model Distillation

RA-L 2026

Place recognition (PR) is a key component of simultaneous localization and mapping (SLAM) in autonomous vehicles and robotics. By efficiently matching descriptors generated from the current scene with a prebuilt reference database, existing PR methods enable accurate vehicle re-localization. However

Cited by 0SourceScholar
2026

TAPO: Dynamic Teacher and Perturbed Answer Injection for Policy Optimization

AAAI 2026technical

Reinforcement learning (RL) has emerged as a powerful framework to improve the reasoning performance of large language models (LLMs), with approaches such as Group Relative Policy Optimization (GRPO) showing promising results. However, GRPO and its variants struggle with collapsed groups (i.e., all-

Cited by 0SourcePDFScholar
2025

A Continual Learning Approach for Embodied Question Answering with Generative Adversarial Imitation Learning

ICASSP 2025accepted

Embodied Question Answering (EQA) is a task in artificial intelligence where an intelligent agent is required to answer questions about its environment. For example, to answer a question such as "Is the TV on or off?", the agent must navigate to the room with the TV and answer with either "On." or "…

Cited by 0SourceScholar
2025

DAAC: Discrepancy-Aware Adaptive Contrastive Learning for Medical Time series

NeurIPS 2025poster

Medical time-series data play a vital role in disease diagnosis but suffer from limited labeled samples and single-center bias, which hinder model generalization and lead to overfitting. To address these challenges, we propose DAAC (Discrepancy-Aware Adaptive Contrastive learning), a learnable multi…

Cited by 0SourcecodeScholar
2025

KLFormer: Karhunen-Loève Transform for Robust 3D Human Pose Estimation

ICASSP 2025accepted

In the scope of 3D human pose estimation, the task encompasses estimating the 3D positions of key skeletal points (i.e., wrists, elbows, and knees) from a 2D image or video sequence. This technology demonstrates widespread applicability across diverse domains, encompassing domains such as kinematic…

Cited by 0SourceScholar
2020

Fast Intent Classification for Spoken Language Understanding Systems

ICASSP 2020accepted

Spoken Language Understanding (SLU) systems consist of several machine learning components operating together (e.g. intent classification, named entity recognition and resolution). Deep learning models have obtained state of the art results on several of these tasks, largely attributed to their bett…

Cited by 0SourceScholar