← Search

Jingjing Zhang

10 accepted papers

2026

Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment

CVPR 2026

Vision-Language-Action (VLA) models have emerged as a powerful framework that unifies perception, language, and control, enabling robots to perform diverse tasks through multimodal understanding. However, current VLA models typically contain massive parameters and rely heavily on large-scale robot d

Cited by 0SourcecodeScholar
2026

Self-guided Semantic Inspection for Zero-Shot Composed Image Retrieval

CVPR 2026

Zero-Shot Composed Image Retrieval (ZS-CIR) aims to retrieve target images using a composed query of a reference image and a textual modification, without relying on triplet-based supervision. As the two inputs describe related but semantically unaligned information, the key challenge lies in interp

Cited by 0SourcecodeScholar
2026

Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric Video

CVPR 2026

Humans develop visual intelligence through perceiving and interacting with their environment--a self-supervised learning process grounded in egocentric experience. Inspired by this, we ask how can artificial systems learn stable object representations from continuous, uncurated first-person videos w

Cited by 0SourceScholar
2025

Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction

ICASSP 2025accepted

The recent rapid development of auditory attention decoding (AAD) offers the possibility of using electroencephalography (EEG) as auxiliary information for target speaker extraction. However, effectively modeling long sequences of speech and resolving the identity of the target speaker from EEG sign…

Cited by 0SourceScholar
2025

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

IJCAI 2025

The brain-assisted target speaker extraction (TSE) aims to extract the attended speech from mixed speech by utilizing the brain neural activities, for example Electroencephalography (EEG). However, existing models overlook the issue of temporal misalignment between speech and EEG modalities, which h

2025

MHANet: Multi-scale Hybrid Attention Network for Auditory Attention Detection

IJCAI 2025

Auditory attention detection (AAD) aims to detect the target speaker in a multi-talker environment from brain signals, such as electroencephalography (EEG), which has made great progress. However, most AAD methods solely utilize attention mechanisms sequentially and overlook valuable multi-scale con

2025

SSM2Mel: State Space Model to Reconstruct Mel Spectrogram from the EEG

ICASSP 2025accepted

Decoding speech from brain signals is a challenging research problem that holds significant importance for studying speech processing in the brain. Although breakthroughs have been made in reconstructing the mel spectrograms of audio stimuli perceived by subjects at the word or letter level using no…

Cited by 0SourceScholar
2024

Model AI Assignments 2024

AAAI 2024technical

The Model AI Assignments session seeks to gather and dis- seminate the best assignment designs of the Artificial In- telligence (AI) Education community. Recognizing that as- signments form the core of student learning experience, we here present abstracts of five AI assignments from the 2024 sessi…

Cited by 0SourcePDFScholar
2019

Improved Latency-communication Trade-off for Map-shuffle-reduce Systems with Stragglers

ICASSP 2019accepted

In a distributed computing system operating according to the map-shuffle-reduce framework, coding data prior to storage can be useful both to reduce the latency caused by straggling servers and to decrease the inter-server communication load in the shuffle phase. In prior work, a concatenated coding…

Cited by 0SourceScholar