← Search

Jiahui Wang

26 accepted papers

2026

EPSegFZ: Efficient Point Cloud Semantic Segmentation for Few- and Zero-Shot Scenarios with Language Guidance

AAAI 2026technical

Recent approaches for few-shot 3D point cloud semantic segmentation typically require a two-stage learning process, i.e., a pre-training stage followed by a few-shot training stage. While effective, these methods face overreliance on pre-training, which hinders model flexibility and adaptability. So

Cited by 0SourcePDFScholar
2026

OmniMap: A General Mapping Framework Integrating Optics, Geometry, and Semantics

ICRA 2026poster

Robotic systems demand accurate and comprehensive 3D environment perception, requiring simultaneous capture of photo-realistic appearance (optical), precise layout shape (geometric), and open-vocabulary scene understanding (semantic). Existing methods typically achieve only partial fulfillment of th…

2026

PIPS: Planar Instance 3D Reconstruction Leveraging Planar Structural Priors

ICRA 2026poster

Planar structures, ubiquitous in man-made indoor environments, enable compact and accurate scene abstraction for various downstream tasks. Recent methods distill planar features into learning-based MVS geometries to obtain coherent 3D plane estimation from multi-view inputs. However, the lack of exp…

Cited by 0codeScholar
2026

Reliable LiDAR Loop Detection through Structural Descriptors and Semantic Graph Matching

ICRA 2026poster

Outdoor loop closure detection is essential for mitigating accumulated drift in SLAM and generating a global consistent map. Semantic graph matching methods utilize object-level topology for distinctive scene representation but rely on environments with rich and distinguishable objects. Moreover, ac…

Cited by 0codeScholar
2026

Video2Robo: 3DGS-based Synthetic Data from One Video Enables Scalable Robot Learning

CVPR 2026

Scalable robot learning is hindered by the high cost of acquiring diverse, high-quality embodied data. Existing data generation approaches partially mitigate this issue but typically depend on hard-to-access hardware and labor-intensive manual effort, with limited generalization to diverse scene con

Cited by 0SourceScholar
2025

Enhancing Multivariate Time-Series Domain Adaptation via Contrastive Frequency Graph Discovery and Language-Guided Adversary Alignment

AAAI 2025technical

Unsupervised domain adaptation (UDA) is a machine learning approach designed to minimize reliance on labeled data by aligning features between a labeled source domain and an unlabeled target domain, thereby reducing feature discrepancies, which is efficient for multivariate time series (MTS) predict…

Cited by 0SourcePDFScholar
2025

Focus-Then-Reuse: Fast Adaptation in Visual Perturbation Environments

NeurIPS 2025poster

Visual reinforcement learning has shown promise in various real-world applications. However, deploying policies in complex real-world environments with visual perturbations remains a significant challenge. We notice that humans tend to filter information at the object level prior to decision-making,…

Cited by 0SourcecodeScholar
2025

LGSDF: Continual Global Learning of Signed Distance Fields Aided by Local Updating

RA-L 2025

Implicit reconstruction of ESDF (Euclidean Signed Distance Field) involves training a neural network to regress the signed distance from any point to the nearest obstacle, which has the advantages of lightweight storage and continuous querying. However, existing algorithms usually rely on conflictin

Cited by 5SourcecodeScholar
2025

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples

COLING 2025main

Aligning Large Language Models (LLMs) with human feedback is crucial for their development. Existing preference optimization methods such as DPO and KTO, while improved based on Reinforcement Learning from Human Feedback (RLHF), are inherently derived from PPO, requiring a reference model that adds…

Cited by 0SourcePDFScholar
2025

Mix-Mask Augmentation and Self-Reconstruction for Cross-Domain Few-Shot Hyperspectral Image Classification

ICASSP 2025accepted

Recently, the metric-based prototypical methods achieves promising performance in few-shot learning (FSL) for hyperspectral image (HSI) classification. However, the existing models are easily affected by the noisy pixels of different categories around the center pixel of the patch, and tend to focus…

Cited by 0SourceScholar
2025

MultiPL-MoE: Multi-Programming-Lingual Extension of Large Language Models through Hybrid Mixture-of-Experts

EMNLP 2025

Despite LLMs’ excellent code creation capabilities, multilingual code generation remains extremely challenging. To address this, we intent to improve the multi-programming-lingual (MultiPL) performance of the base LLMs while retaining the most popular ones using restricted computational resources. W

2025

OpenMIGS: Multi-granularity Information-preserving Open-Vocabulary 3D Gaussian Splatting

IROS 2025

Open-vocabulary scene understanding is critical for robotics, yet existing 3D Gaussian Splatting (3DGS) methods rely on compressed feature embeddings, compromising semantic fidelity and fine-grained interpretation. Although utilizing uncompressed high-dimensional features offers a potential solution

Cited by 0SourcecodeScholar
2025

OpenMulti: Open-Vocabulary Instance-Level Multi-Agent Distributed Implicit Mapping

RA-L 2025

Multi-agent distributed collaborative mapping provides comprehensive and efficient representations for robots. However, existing approaches lack instance-level awareness and semantic understanding of environments, limiting their effectiveness for downstream applications. To address this issue, we pr

Cited by 3SourceScholar
2025

OpenObj: Open-Vocabulary Object-Level Neural Radiance Fields With Fine-Grained Understanding

RA-L 2025

In recent years, there has been a surge of interest in open-vocabulary 3D scene reconstruction facilitated by visual language models (VLMs), which showcase remarkable capabilities in open-set retrieval tasks. Although the semantic ambiguity of existing point-wise feature maps is alleviated by open-v

Cited by 12SourceScholar
2025

Reduced Effectiveness of Kolmogorov-Arnold Networks on Functions with Noise

ICASSP 2025accepted

It has been observed that even a small amount of noise introduced into the dataset can significantly degrade the performance of KAN. In this brief note, we aim to quantitatively evaluate the performance when noise is added to the dataset. We propose an oversampling technique combined with denoising…

Cited by 0SourceScholar
2025

SingRef6D: Monocular Novel Object Pose Estimation with a Single RGB Reference

NeurIPS 2025poster

Recent 6D pose estimation methods demonstrate notable performance but still face some practical limitations. For instance, many of them rely heavily on sensor depth, which may fail with challenging surface conditions, such as transparent or highly reflective materials. In the meantime, RGB-based sol…

Cited by 0SourceScholar
2025

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs

ICCV 2025poster

Multimodal Large Language Models (MLLMs) are commonly derived by extending pre-trained Large Language Models (LLMs) with visual capabilities. In this work, we investigate how MLLMs process visual inputs by analyzing their attention mechanisms. We reveal a surprising sparsity phenomenon: only a small…

2024

Design and Trajectory Tracking Control of CuRobot: A Cubic Reversible Robot

RA-L 2024

In field environments, numerous robots necessitate manual intervention for restoration of functionality post a turnover, resulting in diminished operational efficiency. This study presents an innovative design solution for a reversible omnidirectional mobile robot denoted as CuRobot, featuring a cub

Cited by 1SourceScholar
2024

Efficient Inference of Vision Instruction-Following Models with Elastic Cache

ECCV 2024poster

"In the field of instruction-following large vision-language models (LVLMs), the efficient deployment of these models faces challenges, notably due to the high memory demands of their key-value (KV) caches. Conventional cache management strategies for LLMs focus on cache eviction, which often fails…

2024

Feature Mixing-Based Active Learning for Multi-Label Text Classification

ICASSP 2024accepted

Active learning (AL) aims to reduce labeling costs by selecting the most valuable samples to annotate from a set of unlabeled data. However, recognizing these samples is particularly challenging in multi-label text classification tasks due to the high dimensionality but sparseness of label spaces. E…

Cited by 0SourceScholar
2024

LCP-Fusion: A Neural Implicit SLAM with Enhanced Local Constraints and Computable Prior

IROS 2024poster

Recently the dense Simultaneous Localization and Mapping (SLAM) based on neural implicit representation has shown impressive progress in hole filling and high-fidelity mapping. Nevertheless, existing methods either heavily rely on known scene bounds or suffer inconsistent reconstruction due to drift…

Cited by 0SourcecodeScholar
2024

OpenGraph: Open-Vocabulary Hierarchical 3D Graph Representation in Large-Scale Outdoor Environments

RA-L 2024

Environment representations endowed with sophisticated semantics are pivotal for facilitating seamless interaction between robots and humans, enabling them to effectively carry out various tasks. Open-vocabulary representation, powered by Visual-Language models (VLMs), possesses inherent advantages,

Cited by 38SourcecodeScholar
2024

SEC: More Accurate Clustering Algorithm via Structural Entropy

AAAI 2024technical

As one of the most popular machine learning tools in the field of unsupervised learning, clustering has been widely used in various practical applications. While numerous methods have been proposed for clustering, a commonly encountered issue is that the existing clustering methods rely heavily on l…

Cited by 0SourcePDFScholar
2023

Few-Shot Point Cloud Semantic Segmentation via Contrastive Self-Supervision and Multi-Resolution Attention

ICRA 2023poster

This paper presents an effective few-shot point cloud semantic segmentation approach for real-world applications. Existing few-shot segmentation methods on point cloud heavily rely on the fully-supervised pretrain with large annotated datasets, which causes the learned feature extraction bias to tho…

Cited by 16SourceScholar