← Search

Jie Wei

7 accepted papers

2026

EnergyAction: Unimanual to Bimanual Composition with Energy-Based Models

CVPR 2026

Recent advances in unimanual manipulation policies have achieved remarkable success across diverse robotic tasks through abundant training data and well-established model architectures. However, extending these capabilities to bimanual manipulation remains challenging due to the lack of bimanual dem

Cited by 0SourcecodeScholar
2026

EnsembleVLA: Ensemble Learning for Vision-Language Action Models

ICML 2026poster

Diverse Vision-language-action (VLA) models have been proposed and demonstrated remarkable capabilities in robotic manipulation. However, how to effectively ensemble VLAs to further enhance performance remains largely unexplored, as conventional ensemble techniques designed for discriminative tasks …

Cited by 0SourceScholar
2025

Enhancing Multimodal Named Entity Recognition through Adaptive Mixup Image Augmentation

COLING 2025main

Multimodal named entity recognition (MNER) extends traditional named entity recognition (NER) by integrating visual and textual information. However, current methods still face significant challenges due to the text-image mismatch problem. Recent advancements in text-to-image synthesis provide promi…

Cited by 0SourcePDFScholar
2025

Mix-Mask Augmentation and Self-Reconstruction for Cross-Domain Few-Shot Hyperspectral Image Classification

ICASSP 2025accepted

Recently, the metric-based prototypical methods achieves promising performance in few-shot learning (FSL) for hyperspectral image (HSI) classification. However, the existing models are easily affected by the noisy pixels of different categories around the center pixel of the patch, and tend to focus…

Cited by 0SourceScholar
2024

Efficient Temporal Action Segmentation via Boundary-aware Query Voting

NeurIPS 2024poster

Although the performance of Temporal Action Segmentation (TAS) has been improved in recent years, achieving promising results often comes with a high computational cost due to dense inputs, complex model structures, and resource-intensive post-processing requirements. To improve the efficiency while…

2023

Multi-Scale Receptive Field Graph Model for Emotion Recognition in Conversations

ICASSP 2023accepted

Emotion recognition in conversations (ERC) has gained more attention, where contextual information modeling and multimodal fusion have been the focus and challenges in recent years. In this paper, we proposed a Multi-Scale Receptive Field Graph model (MSRFG) to tackle the challenges of ERC. Specific…

Cited by 0SourceScholar