← Search

Chen Jiang

29 accepted papers

2026

Beyond Instance-Level Self-Supervision in 3D Multi-Modal Medical Imaging

ICML 2026poster

Self-supervised pre-training methods in medical imaging typically treat each individual as an isolated instance, learning representations through augmentation-based objectives or masked reconstruction. They often do not adequately capitalize on a key characteristic of physiological features: anatomi…

Cited by 0SourceScholar
2026

CARD: Towards Conditional Design of Multi-agent Topological Structures

ICLR 2026poster

Large language model (LLM)-based multi-agent systems have shown strong capabilities in tasks such as code generation and collaborative reasoning. However, the effectiveness and robustness of these systems critically depend on their communication topology, which is often fixed or statically learned,…

Cited by 0SourcecodeScholar
2026

Disco: Densely-overlapping Cell Instance Segmentation via Adjacency-aware Collaborative Coloring

ICLR 2026poster

Accurate cell instance segmentation is foundational for digital pathology analysis. Existing methods based on contour detection and distance mapping still face significant challenges in processing complex and dense cellular regions. Graph coloring-based methods provide a new paradigm for this task,…

Cited by 0SourcecodeScholar
2026

Tracing the Heart’s Pathways: ECG Representation Learning from a Cardiac Conduction Perspective

AAAI 2026technical

The multi-lead electrocardiogram (ECG) stands as a cornerstone of cardiac diagnosis. Recent strides in electrocardiogram self-supervised learning (eSSL) have brightened prospects for enhancing representation learning without relying on high-quality annotations. Yet earlier eSSL methods suffer a key

Cited by 0SourcePDFScholar
2026

Unsupervised Anomaly Detection in Dynamic Graphs via Compatibility Modeling and Boundary Learning

IJCAI 2026

Anomaly detection in dynamic graphs is essential for monitoring evolving systems such as transaction networks and online platforms. Yet existing methods remain limited in realistic edge-stream settings: snapshot-based approaches discretize continuous interactions and miss fine-grained temporal signa

Cited by 0Scholar
2025

BBScoreV2: Learning Time-Evolution and Latent Alignment from Stochastic Representation

EMNLP 2025

Autoregressive generative models play a key role in various language tasks, especially for modeling and evaluating long text sequences. While recent methods leverage stochastic representations to better capture sequence dynamics, encoding both temporal and structural dependencies and utilizing such

2025

Brain Bandit: A Biologically Grounded Neural Network for Efficient Control of Exploration

ICLR 2025oral

How to balance between exploration and exploitation in an uncertain environment is a central challenge in reinforcement learning. In contrast, humans and animals have demonstrated superior exploration efficiency in novel environments. To understand how the brain’s neural network controls exploration…

Cited by 0SourcePDFScholar
2025

ChromFound: Towards A Universal Foundation Model for Single-Cell Chromatin Accessibiltiy Data

NeurIPS 2025poster

The advent of single-cell Assay for Transposase-Accessible Chromatin using sequencing (scATAC-seq) offers an innovative perspective for deciphering regulatory mechanisms by assembling a vast repository of single-cell chromatin accessibility data. While foundation models have achieved significant suc…

Cited by 0SourcecodeScholar
2025

Interpreting Behaviors and Geometric Constraints as Knowledge Graphs for Robot Manipulation Control

IROS 2025

In this paper, we investigate the feasibility of using knowledge graphs to interpret actions and behaviors for robot manipulation control. Equipped with an uncalibrated visual servoing controller, we propose to use robot knowledge graphs to unify behavior trees and geometric constraints, conceptuali

Cited by 1SourceScholar
2025

Minimal Semantic Sufficiency Meets Unsupervised Domain Generalization

NeurIPS 2025poster

The generalization ability of deep learning has been extensively studied in supervised settings, yet it remains less explored in unsupervised scenarios. Recently, the Unsupervised Domain Generalization (UDG) task has been proposed to enhance the generalization of models trained with prevalent unsupe…

Cited by 0SourceScholar
2025

Point and Go: Intuitive Reference Frame Reallocation in Mode Switching for Assistive Robotics

ICRA 2025

Operating high degree of freedom robots can be difficult for users of wheelchair mounted robotic manipulators. Mode switching in Cartesian space has several drawbacks such as unintuitive control reference frames, separate translation and orientation control, and limited movement capabilities that hi

Cited by 0SourceScholar
2025

Robot Manipulation in Salient Vision Through Referring Image Segmentation and Geometric Constraints

ICRA 2025

In this paper, we perform robot manipulation activities in real-world environments with language contexts by integrating a compact referring image segmentation model into the robot's perception module. First, we propose CLIPU<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www

Cited by 4SourceScholar
2025

SegAnyPET: Universal Promptable Segmentation from Positron Emission Tomography Images

ICCV 2025poster

Positron Emission Tomography (PET) is a powerful molecular imaging tool that plays a crucial role in modern medical diagnostics by visualizing radio-tracer distribution to reveal physiological processes. Accurate organ segmentation from PET images is essential for comprehensive multi-systemic analys…

2025

Structure-aware Semantic Discrepancy and Consistency for 3D Medical Image Self-supervised Learning

ICCV 2025poster

3D medical image self-supervised learning (mSSL) holds great promise for medical analysis. Effectively supporting broader applications requires considering anatomical structure variations in location, scale, and morphology, which are crucial for capturing meaningful distinctions. However, previous m…

2025

Towards a Universal 3D Medical Multi-modality Generalization via Learning Personalized Invariant Representation

ICCV 2025poster

Variations in medical imaging modalities and individual anatomical differences pose challenges to cross-modality generalization in multi-modal tasks. Existing methods often concentrate exclusively on common anatomical patterns, thereby neglecting individual differences and consequently limiting thei…

2024

BBScore: A Brownian Bridge Based Metric for Assessing Text Coherence

AAAI 2024technical

Measuring the coherence of text is a vital aspect of evaluating the quality of written content. Recent advancements in neural coherence modeling have demonstrated their efficacy in capturing entity coreference and discourse relations, thereby enhancing coherence evaluation. However, many existing me…

2024

CLIPUNetr: Assisting Human-robot Interface for Uncalibrated Visual Servoing Control with CLIP-driven Referring Expression Segmentation

ICRA 2024poster

The classical human-robot interface in uncalibrated image-based visual servoing (UIBVS) relies on either human annotations or semantic segmentation with categorical labels. Both methods fail to match natural human communication and convey rich semantics in manipulation tasks as effectively as natura…

Cited by 1SourceScholar
2024

EVSMap: An Efficient Volumetric-Semantic Mapping Approach for Embedded Systems

IROS 2024poster

Despite significant progress in perception tasks such as 3D scene mapping and semantic information extraction using SLAM and deep learning, applying these techniques within computationally constrained embedded systems remains a challenge. In this work, we introduce a novel end-to-end framework for e…

Cited by 0SourceScholar
2023

Boundary-Aware Backward-Compatible Representation via Adversarial Learning in Image Retrieval

CVPR 2023poster

Image retrieval plays an important role in the Internet world. Usually, the core parts of mainstream visual retrieval systems include an online service of the embedding model and a large-scale vector database. For traditional model upgrades, the old model will not be replaced by the new one until th…

2023

InGVIO: A Consistent Invariant Filter for Fast and High-Accuracy GNSS-Visual-Inertial Odometry

RA-L 2023

Combining Global Navigation Satellite System (GNSS) with visual and inertial sensors can give smooth pose estimation without drifting. The fusion system gradually degrades to Visual-Inertial Odometry (VIO) with the number of satellites decreasing, which guarantees robust global navigation in GNSS un

Cited by 38SourcecodeScholar
2023

MDCS: More Diverse Experts with Consistency Self-distillation for Long-tailed Recognition

ICCV 2023poster

Recently, multi-expert methods have led to significant improvements in long-tail recognition (LTR). We summarize two aspects that need further enhancement to contribute to LTR boosting: (1) More diverse experts; (2) Lower model variance. However, the previous methods didn't handle them well. To this…

Cited by 16PDFcodeScholar
2023

TransVCL: Attention-Enhanced Video Copy Localization Network with Flexible Supervision

AAAI 2023technical

Video copy localization aims to precisely localize all the copied segments within a pair of untrimmed videos in video retrieval applications. Previous methods typically start from frame-to-frame similarity matrix generated by cosine similarity between frame-level features of the input video pair, an…

2022

A Large-Scale Comprehensive Dataset and Copy-Overlap Aware Evaluation Protocol for Segment-Level Video Copy Detection

CVPR 2022poster

In this paper, we introduce VCSL (Video Copy Segment Localization), a new comprehensive segment-level annotated video copy dataset. Compared with existing copy detection datasets restricted by either video-level annotation or small-scale, VCSL not only has two orders of magnitude more segment-level…

Cited by 18PDFcodeScholar
2020

Understanding Contexts Inside Robot and Human Manipulation Tasks through Vision-Language Model and Ontology System in Video Streams

IROS 2020poster

Manipulation tasks in daily life, such as pouring water, unfold through human intentions. Being able to process contextual knowledge from these Activities of Daily Living (ADLs) over time can help us understand manipulation intentions, which are essential for an intelligent robot to transition smoot…

Cited by 11SourceScholar
2019

Video Object Segmentation using Teacher-Student Adaptation in a Human Robot Interaction (HRI) Setting

ICRA 2019poster

Video object segmentation is an essential task in robot manipulation to facilitate grasping and learning affordances. Incremental learning is important for robotics in unstructured environments. Inspired by the children learning process, human robot interaction (HRI) can be utilized to teach robots…

Cited by 108SourcecodeScholar