← Search

Haodong Jing

15 accepted papers

2026

Decoupling Primitive with Experts: Dynamic Feature Alignment for Compositional Zero-Shot Learning

ICLR 2026poster

Compositional Zero-Shot Learning (CZSL) investigates compositional generalization capacity to recognize unknown state-object pairs based on learned primitive concepts. Existing CZSL methods typically derive primitives features through a simple composition-prototype mapping, which is suboptimal for a…

Cited by 0SourceScholar
2026

EVOKE: Efficient and High-Fidelity EEG-to-Video Reconstruction via Decoupling Implicit Neural Representation

AAAI 2026technical

Visual neural decoding is an important research topic at the intersection of cognitive neuroscience and machine learning. While recent progress has been made in EEG-based neural decoding, reconstructing dynamic visual content remains challenging. In the field of EEG decoding, current models either u

Cited by 0SourcePDFScholar
2026

MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality

ICML 2026poster

Unified visual tokenization faces a fundamental trade-off: optimizing for high-fidelity pixel reconstruction (spatial equivariance) inherently conflicts with semantic abstraction (conceptual invariance). We identify the root cause as Manifold Misalignment, where naive joint optimization leads to con…

Cited by 0SourceScholar
2026

More Natural, More Real: Object-aware Gaussian Splatting for 3D Visual Decoding from Human Brain

CVPR 2026

Exploring human visual perception and understanding of the stereoscopic world represents a significant topic in computational neuroscience. Recent studies have provided rich Brain-3D datasets, conducted preliminary explorations into 3D visual reconstruction. However, existing research struggles to c

Cited by 0SourceScholar
2026

OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention

ICML 2026poster

Humans perceive the world through diverse modalities that operate synergistically to support a holistic understanding of their surroundings. However, existing omnimodal models still exhibit substantial performance degradation on visual tasks when the audio modality is incorporated. We identify this …

Cited by 0SourceScholar
2026

UniHOI: Unified Human-Object Interaction Understanding via Unified Token Space

AAAI 2026technical

In the field of human-object interaction (HOI), detection and generation are two dual tasks that have traditionally been addressed separately, hindering the development of comprehensive interaction understanding. To address this, we propose UniHOI, which jointly models HOI detection and generation v

Cited by 0SourcePDFScholar
2025

Beyond Brain Decoding: Visual-Semantic Reconstructions to Mental Creation Extension Based on fMRI

ICCV 2025poster

Decoding visual information from fMRI signals is an important pathway to understand how the brain represents the world, and is a cutting-edge field of artificial general intelligence. Decoding fMRI should not be limited to reconstructing visual stimuli, but also further transforming them into descri…

Cited by 0SourcePDFScholar
2025

Beyond Image Classification: A Video Benchmark and Dual-Branch Hybrid Discrimination Framework for Compositional Zero-Shot Learning

CVPR 2025poster

Human reasoning naturally combines concepts to identify unseen compositions, a capability that Compositional Zero-Shot Learning (CZSL) aims to replicate in machine learning models. However, we observe that focusing solely on typical image classification tasks in CZSL may limit models' compositional…

Cited by 0SourcePDFScholar
2025

NEED: Cross-Subject and Cross-Task Generalization for Video and Image Reconstruction from EEG Signals

NeurIPS 2025poster

Translating brain activity into meaningful visual content has long been recognized as a fundamental challenge in neuroscience and brain-computer interface research. Recent advances in EEG-based neural decoding have shown promise, yet two critical limitations remain in this area: poor generalization…

Cited by 0SourceScholar
2025

PlaneHEC: Efficient Hand-Eye Calibration for Multi-View Robotic Arm via Any Point Cloud Plane Detection

ICRA 2025

Hand-eye calibration is an important task in vision-guided robotic systems and is crucial for determining the transformation matrix between the camera coordinate system and the robot end-effector. Existing methods, for multi-view robotic systems, usually rely on accurate geometric models or manual a

Cited by 0SourceScholar
2025

Refiner: Fine-grained Cross-modal Concepts Refinement for Compositional Zero-Shot Learning

ICASSP 2025accepted

Recent Compositional Zero-Shot Learning (CZSL) methods increasingly adopt the pre-trained vision-language models to capture the contextual relations between image and text spaces. However, the single-class-token design from Transformer-based encoder inevitably captures contextual information from un…

Cited by 0SourceScholar
2025

SAMap: Semantic Alignment for HD Map Detection Domain Generalization Under Varying Weather and Lighting

IROS 2025

High-definition (HD) maps are crucial for autonomous driving systems. Despite recent advances in learning-based HD map prediction methods, these approaches experience significant performance degradation when encountering unseen weather or lighting conditions due to feature distribution discrepancies

Cited by 0SourceScholar
2025

See Through Their Minds: Learning Transferable Brain Decoding Models from Cross-Subject fMRI

AAAI 2025technical

Deciphering visual content from fMRI sheds light on the human vision system, but data scarcity and noise limit brain decoding model performance. Traditional approaches rely on subject-specific models, which are sensitive to training sample size. In this paper, we address data scarcity by proposing s…

2024

Complementing Onboard Sensors with Satellite Maps: A New Perspective for HD Map Construction

ICRA 2024poster

High-definition (HD) maps play a crucial role in autonomous driving systems. Recent methods have attempted to construct HD maps in real-time using vehicle onboard sensors. Due to the inherent limitations of onboard sensors, which include sensitivity to detection range and susceptibility to occlusion…

Cited by 18SourcecodeScholar
2024

MRSP: Learn Multi-Representations of Single Primitive for Compositional Zero-Shot Learning

ECCV 2024poster

"Compositional Zero-Shot Learning (CZSL) aims to classify unseen state-object compositions using seen primitives. Previous methods commonly map an identical primitive from different compositions to the same area within embedding space, aiming to establish primitive representation or assess decoding…

Cited by 0SourcePDFScholar