← Search

JIE YANG

123 accepted papers

2026

AdaDepth: Exploiting Inherent Scene Information for Self-Supervised Depth Estimation in Dynamic Scenes

AAAI 2026technical

Self-supervised monocular depth estimation methods severely compromise accuracy in dynamic objects due to their static scene assumption. Existing approaches for dynamic scenes suffer from two critical shortcomings: 1) reliance on supervised segmentation models (requiring costly annotations) or comp

Cited by 1SourcePDFScholar
2026

Bridging Past and Future: Distribution-Aware Alignment for Time Series Forecasting

ICLR 2026poster

Although contrastive and other representation-learning methods have long been explored in vision and NLP, their adoption in modern time series forecasters remains limited. We believe they hold strong promise for this domain. To unlock this potential, we explicitly align past and future representatio…

Cited by 0SourcecodeScholar
2026

Bridging the Data Scarcity in Venous Thromboembolism Detection: A Deep Learning Framework for Large-scale Irregular Clinical Time Series

IJCAI 2026

Venous thromboembolism (VTE) is a common and life-threatening complication in cancer patients after treatment. Early risk assessment and detection of VTE primarily rely on clinical indicators, such as blood test results. However, existing studies are limited to static or snapshot-based models, faili

Cited by 0Scholar
2026

Bridging the Modality Reliability Gap in Drug-Target Interaction Prediction via a Confidence-aware Multimodal Fusion Framework

AAAI 2026technical

With the rapid advancement of deep learning, drug target interaction (DTI) prediction has seen substantial performance enhancements. However, existing methodologies face a critical, yet unaddressed challenge, i.e., the Modality Reliability Gap. Such a gap arises from the unpredictable variance in t

Cited by 0SourcePDFScholar
2026

Event-Fused Hybrid ANN-SNN Architecture for Low-Latency Object Detection in Automotive Vision

RA-L 2026

In advanced driver-assistance systems, current computer vision algorithms predominantly rely on frame-based RGB cameras, which suffer from high latency in high-speed or sudden-scenario applications due to fixed frame rates. In response to this challenge, event-based cameras have gained attention as

Cited by 0SourcecodeScholar
2026

Fair Graph Learning with Limited Sensitive Attribute Information

AAAI 2026technical

Graph neural networks (GNNs) excel at modeling graph-structured data but often inherit and amplify biases, leading to substantial efforts in developing fair GNNs. However, most existing approaches assume full access to sensitive attribute information, which is often impractical in real-world scenari

Cited by 0SourcePDFScholar
2026

Fore-Mamba3D: Mamba-based Foreground-Enhanced Encoding for 3D Object Detection

ICLR 2026poster

Linear modeling methods like Mamba have been merged as the effective backbone for the 3D object detection task. However, previous Mamba-based methods utilize the bidirectional encoding for the whole non-empty voxel sequence, which contains abundant useless background information in the scenes. Thoug…

Cited by 0SourcecodeScholar
2026

From Observations to States: Latent Time Series Forecasting

ICML 2026poster

Deep learning has achieved strong performance in Time Series Forecasting (TSF). However, we identify a critical representation paradox, termed Latent Chaos: models with accurate predictions often learn latent representations that are temporally disordered and lack continuity. We attribute this pheno…

Cited by 0SourceScholar
2026

MACRec: A Multi-View Subspace Alignment Framework for Contrastive Sampling Calibration in Recommendation

AAAI 2026technical

Graph Contrastive Learning (GCL) has proven effective in mitigating data sparsity and enhancing representation learning for recommendation. Yet, most GCL frameworks indiscriminately treat all non-anchor nodes as negatives during contrastive sampling, often leading to the false negative problem where

Cited by 0SourcePDFScholar
2026

Multi-View Clustering with Granularity-Aware Pseudo Supervision

AAAI 2026technical

Modern multi-view clustering (MVC) is dominated by two paradigms: multi-view fusion and pseudo-label-guided learning. Pseudo-labeling methods can suffer from confirmation bias; their reliance on a fixed-granularity supervision from an initial clustering can cause learned embeddings to drift from the

Cited by 0SourcePDFScholar
2026

No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection

CVPR 2026

The collection and detection of video anomaly data has long been a challenging problem due to its rare occurrence and spatio-temporal scarcity. Existing video anomaly detection (VAD) methods under perform in open-world scenarios. Key contributing factors include limited dataset diversity, and inadeq

Cited by 0SourcecodeScholar
2026

Object Fidelity Diffusion for Remote Sensing Image Generation

ICLR 2026poster

High-precision controllable remote sensing image generation is both meaningful and challenging. Existing diffusion models often produce low-fidelity objects due to their inability to adequately capture morphological details, which may affect the robustness and reliability of object detection models.…

Cited by 0SourcecodeScholar
2026

RECODE: A Benchmark for Research Code DEvelopment with Interactive Human Feedback

ICLR 2026poster

Large language models (LLMs) show the promise in supporting scientific research implementation, yet their ability to generate correct and executable code remains limited. Existing works largely adopt one-shot settings, ignoring the iterative and feedback-driven nature of realistic workflows of scien…

Cited by 0SourcecodeScholar
2026

RoSAMDepth: Robust Self-supervised Depth Estimation Leveraging Segment Anything Model

CVPR 2026

Robust depth estimation aims to maintain high-quality depths across diverse conditions. However, most existing methods estimate depth without taking into account the object-level information. As a result, the predicted depth may easily deviate within objects and become blurred under adverse conditio

Cited by 0SourcecodeScholar
2026

SMoFi: Step-wise Momentum Fusion for Split Federated Learning on Heterogeneous Data

AAAI 2026technical

Split Federated Learning is a system-efficient federated learning paradigm that leverages the rich computing resources at a central server to train model partitions. Data heterogeneity across silos, however, presents a major challenge undermining the convergence speed and accuracy of the global mode

Cited by 0SourcePDFScholar
2026

Sparse Annotation, Dense Supervision: Unleashing Self-Training Power for Occupancy Prediction With 2D Labels

RA-L 2026

Serving as a fundamental task in robotic navigation and autonomous driving, occupancy prediction is gaining increasing attention for its fine-grained perception of the 3D environment. Most existing methods rely on dense 3D annotations, which are expensive, labor-intensive, and difficult to scale in

Cited by 1SourceScholar
2026

Trustworthy Classification for Complex Social Surveys: A Memory-Enhanced Hierarchical Framework with Calibrated Uncertainty

AAAI 2026technical

Automated classification of complex social survey questionnaires is crucial for large-scale social science research but faces significant reliability challenges due to intricate hierarchical label structures, severe class imbalance, semantic ambiguity, and incomplete data coverage. Conventional clas

Cited by 0SourcePDFScholar
2026

V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning

ICML 2026poster

Multimodal large language models (MLLMs) have achieved remarkable success in general perception, yet complex multi-step visual reasoning remains a persistent challenge. Although recent agentic approaches incorporate tool use, they often neglect critical execution feedback. Consequently, they suffer …

Cited by 0SourceScholar
2025

Automatic MILP Model Construction for Multi-Robot Task Allocation and Scheduling Based on Large Language Models

IROS 2025

With the accelerated development of Industry 4.0, intelligent manufacturing systems increasingly require efficient task allocation and scheduling in multi-robot systems. However, existing methods rely on domain expertise and face challenges in adapting to dynamic production constraints. Additionally

Cited by 7SourceScholar
2025

BFS-Prover: Scalable Best-First Tree Search for LLM-based Automatic Theorem Proving

ACL 2025long

Recent advancements in large language models (LLMs) have spurred growing interest in automatic theorem proving using Lean4, where effective tree search methods are crucial for navigating the underlying large proof search spaces. While the existing approaches primarily rely on value functions and/or…

Cited by 0SourcePDFScholar
2025

CAD-GPT: Synthesising CAD Construction Sequence with Spatial Reasoning-Enhanced Multimodal LLMs

AAAI 2025technical

Computer-aided design (CAD) significantly enhances the efficiency, accuracy, and innovation of design processes by enabling precise 2D and 3D modeling, extensive analysis, and optimization. Existing methods for creating CAD models rely on latent vectors or point clouds, which are difficult to obtain…

Cited by 1SourcePDFScholar
2025

Collaborative Similarity Fusion and Consistency Recovery for Incomplete Multi-view Clustering

AAAI 2025technical

As partial samples are often absent in certain views, incomplete multi-view clustering has become a challenging task. To tackle data with missing views, current methods either utilize the data similarity relations to recover missing samples or primarily consider the available information of existing…

Cited by 0SourcePDFScholar
2025

Context-Informed Machine Translation of Manga using Multimodal Large Language Models

COLING 2025main

Due to the significant time and effort required for handcrafting translations, most manga never leave the domestic Japanese market. Automatic manga translation is a promising potential solution. However, it is a budding and underdeveloped field and presents complexities even greater than those found…

2025

Enhanced Denesity Peak Clustering for High-Dimensional Data

AAAI 2025technical

As a foundational clustering paradigm, Density Peak Clustering (DPC) partitions samples into clusters based on their density peaks, garnering widespread attention. However, traditional DPC methods usually focus on high-density regions, neglecting representative peaks in relatively low-density areas,…

2025

FB-Diff: Fourier Basis-guided Diffusion for Temporal Interpolation of 4D Medical Imaging

ICCV 2025poster

The temporal interpolation task for 4D medical imaging, plays a crucial role in clinical practice of respiratory motion modeling. Following the simplified linear-motion hypothesis, existing approaches adopt optical flow-based models to interpolate intermediate frames. However, realistic respiratory…

2025

Fix-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text

ICCV 2025poster

CLIP has shown promising performance across many short-text tasks in a zero-shot manner. However, limited by the input length of the text encoder, CLIP struggles on under-stream tasks with long-text inputs (>77 tokens). To remedy this issue, we propose FIX-CLIP, which includes three novel modules: (…

2025

Flick: Empowering Federated Learning with Commonsense Knowledge

NeurIPS 2025poster

Federated Learning (FL) has emerged as a privacy-preserving framework for training models on data generated at the edge. However, the heterogeneity of data silos (e.g., label skew and domain shift) often leads to inconsistent learning objectives and suboptimal model performance. Inspired by the data…

Cited by 0SourceScholar
2025

From GNNs to Trees: Multi-Granular Interpretability for Graph Neural Networks

ICLR 2025poster

Interpretable Graph Neural Networks (GNNs) aim to reveal the underlying reasoning behind model predictions, attributing their decisions to specific subgraphs that are informative. However, existing subgraph-based interpretable methods suffer from an overemphasis on local structure, potentially overl…

Cited by 0SourcePDFScholar
2025

GIM: A Million-scale Benchmark for Generative Image Manipulation Detection and Localization

AAAI 2025technical

The extraordinary ability of generative models emerges as a new trend in image editing and generating realistic images, posing a serious threat to the trustworthiness of multimedia data and driving the research of image manipulation detection and location (IMDL). However, the lack of a large-scale d…

2025

Glocal Information Bottleneck for Time Series Imputation

NeurIPS 2025poster

Time Series Imputation (TSI), which aims to recover missing values in temporal data, remains a fundamental challenge due to the complex and often high-rate missingness in real-world scenarios. Existing models typically optimize the point-wise reconstruction loss, focusing on recovering numerical val…

Cited by 0SourcecodeScholar
2025

Holistic Semantic Representation for Navigational Trajectory Generation

AAAI 2025technical

Trajectory generation has garnered significant attention from researchers in the field of spatio-temporal analysis, as it can generate substantial synthesized human mobility trajectories that enhance user privacy and alleviate data scarcity. However, existing trajectory generation methods often focu…

2025

Importance-Awareness Masking Network for Robust Document Retrieval

ICASSP 2025accepted

In this paper, we introduce the IMPortance-awaReness maskIng NeTwork (IMPRINT), a novel approach to enhance the robustness of document retrieval systems against query variations, particularly those containing misspellings. Unlike previous models that treat all query components (words/features) equal…

Cited by 0SourceScholar
2025

InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance Parsing

CVPR 2025highlight

Recent advances in 3D human-aware generation have made significant progress. However, existing methods still struggle with generating novel Human Object Interaction (HOI) from text, particularly for open-set objects. We identify three main challenges of this task: precise human-object relation reaso…

Cited by 0SourcePDFScholar
2025

Multi-view Clustering via Multi-granularity Ensemble

IJCAI 2025

Multi-view clustering aims to integrate complementary information from multiple views to improve clustering performance. However, existing ensemble-based methods suffer from information loss due to their reliance on single-granularity labels, limiting the discriminative capability of learned represe

Cited by 0SourcePDFScholar
2025

Revisiting Interpolation for Noisy Label Correction

AAAI 2025technical

Label correction methods are popular for their simple architecture in learning with noisy labels. However, they suffer severely from false label correction and achieve subpar performance compared with state-of-the-art methods. In this paper, we revisit the label correction methods through theoretica…

2025

Symmetric Bi-branch Modality-search Aggregation Network for Multi-modal Liver Segmentation

ICASSP 2025accepted

Medical image segmentation is crucial for diagnosis and surgical planning of liver diseases. The existing methods mainly focus on global or local features and neglect spatial dependencies among modalities and blurred boundaries. To tackle these challenges, we propose a symmetric bi-branch modality-s…

Cited by 0SourceScholar
2025

Towards Homogeneous Lexical Tone Decoding from Heterogeneous Intracranial Recordings

ICLR 2025poster

Recent advancements in brain-computer interfaces (BCIs) and deep learning have made decoding lexical tones from intracranial recordings possible, providing the potential to restore the communication ability of speech-impaired tonal language speakers. However, data heterogeneity induced by both physi…

Cited by 0SourcePDFScholar
2025

VIKI‑R: Coordinating Embodied Multi-Agent Cooperation via Reinforcement Learning

NeurIPS 2025poster

Coordinating multiple embodied agents in dynamic environments remains a core challenge in artificial intelligence, requiring both perception-driven reasoning and scalable cooperation strategies. While recent works have leveraged large language models (LLMs) for multi-agent planning, a few have begun…

Cited by 0SourceScholar
2025

VehicleWorld: A Highly Integrated Multi-Device Environment for Intelligent Vehicle Interaction

EMNLP 2025

Intelligent vehicle cockpits present unique challenges for API Agents, requiring coordination across tightly-coupled subsystems that exceed typical task environments’ complexity. Traditional Function Calling (FC) approaches operate statelessly, requiring multiple exploratory calls to build environme

2024

Bayesian Deep Predictive Coding for Snake-like Robotic Control in Unknown Terrains

IROS 2024

Effectively modeling the spatio-temporal interactions both internally and externally is a challenge in controlling multi-linked snake robots. This paper presents an effective method based on deep predictive coding: SnakeFormer, to address the aforementioned issue. The main contributions include: 1)

Cited by 0SourceScholar
2024

DMiT: Deformable Mipmapped Tri-Plane Representation for Dynamic Scenes

ECCV 2024poster

"Neural Radiance Fields (NeRF) have achieved remarkable progress on dynamic scenes with deformable objects. Nonetheless, most previous works required multi-view inputs or long training time (several hours), making it hard to apply them for real-world scenarios. Recent works dedicated to addressing b…

Cited by 0SourcePDFScholar
2024

Efficient Multi-view Unsupervised Feature Selection with Adaptive Structure Learning and Inference

IJCAI 2024poster

As data with diverse representations become high-dimensional, multi-view unsupervised feature selection has been an important learning paradigm. Generally, existing methods encounter the following challenges: (i) traditional solutions either concatenate different views or introduce extra parameters…

Cited by 11SourcePDFScholar
2024

Exploring Effective Stimulus Encoding via Vision System Modeling for Visual Prostheses

ICLR 2024poster

Visual prostheses are potential devices to restore vision for blind people, which highly depends on the quality of stimulation patterns of the implanted electrode array. However, existing processing frameworks prioritize the generation of stimulation while disregarding the potential impact of restor…

Cited by 0SourcePDFScholar
2024

F-HOI: Toward Fine-grained Semantic-Aligned 3D Human-Object Interactions

ECCV 2024poster

"Existing 3D human object interaction (HOI) datasets and models simply align global descriptions with the long HOI sequence, while lacking a detailed understanding of intermediate states and the transitions between states. In this paper, we argue that fine-grained semantic alignment, which utilizes…

Cited by 10SourcePDFScholar
2024

Fair Graph Learning Using Constraint-Aware Priority Adjustment and Graph Masking in River Networks

AAAI 2024technical

Accurate prediction of water quality and quantity is crucial for sustainable development and human well-being. However, existing data-driven methods often suffer from spatial biases in model performance due to heterogeneous data, limited observations, and noisy sensor data. To overcome these challen…

2024

FedTrans: Client-Transparent Utility Estimation for Robust Federated Learning

ICLR 2024poster

Federated Learning (FL) is an important privacy-preserving learning paradigm that plays an important role in the Intelligent Internet of Things. Training a global model in FL, however, is vulnerable to the noise in the heterogeneous data across the clients. In this paper, we introduce **FedTrans**,…

Cited by 0SourcePDFScholar
2024

Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

ECCV 2024poster

"In this paper, we develop an open-set object detector, called Grounding DINO, by marrying Transformer-based detector DINO with grounded pre-training, which can detect arbitrary objects with human inputs such as category names or referring expressions. The key solution of open-set object detection i…

2024

Guiding Clinical Reasoning with Large Language Models via Knowledge Seeds

IJCAI 2024poster

Clinical reasoning refers to the cognitive process that physicians employ in evaluating and managing patients. This process typically involves suggesting necessary examinations, diagnosing patients’ diseases, and selecting appropriate therapies, etc. Accurate clinical reasoning requires extensive me…

2024

Kernel PCA for Out-of-Distribution Detection

NeurIPS 2024poster

Out-of-Distribution (OoD) detection is vital for the reliability of Deep Neural Networks (DNNs). Existing works have shown the insufficiency of Principal Component Analysis (PCA) straightforwardly applied on the features of DNNs in detecting OoD data from In-Distribution (InD) data. The failure of P…

2024

KptLLM: Unveiling the Power of Large Language Model for Keypoint Comprehension

NeurIPS 2024poster

Recent advancements in Multimodal Large Language Models (MLLMs) have greatly improved their abilities in image understanding. However, these models often struggle with grasping pixel-level semantic details, e.g., the keypoints of an object. To bridge this gap, we introduce the novel challenge of Sem…

Cited by 1SourcePDFScholar
2024

MCM: Masked Cell Modeling for Anomaly Detection in Tabular Data

ICLR 2024poster

This paper addresses the problem of anomaly detection in tabular data, which is usually implemented in an one-class classification setting where the training set only contains normal samples. Inspired by the success of masked image/language modeling in vision and natural language domains, we extend…

Cited by 9SourcePDFScholar
2024

Masking the Unknown: Leveraging Masked Samples for Enhanced Data Augmentation

UAI 2024poster

Data Augmentation (DA) has become a widely adopted strategy for addressing data scarcity in numerous NLP tasks, especially in scenarios with limited resources or imbalanced classes. However, many existing augmentation techniques rely on randomness or additional resources, presenting challenges in bo…

Cited by 0SourcePDFScholar
2024

MedJourney: Benchmark and Evaluation of Large Language Models over Patient Clinical Journey

NeurIPS 2024poster

Large language models (LLMs) have demonstrated remarkable capabilities in language understanding and generation, leading to their widespread adoption across various fields. Among these, the medical field is particularly well-suited for LLM applications, as many medical tasks can be enhanced by LLMs.…

Cited by 1SourcePDFScholar
2024

OEE-CFC: A Dataset for Open Event Extraction from Chinese Financial Commentary

EMNLP 2024finding

To meet application needs, event extraction has shifted from simple entities to unconventional entities serving as event arguments. However, current corpora with unconventional entities as event arguments are limited in event types and lack rich multi-events and shared arguments. Financial commentar…

2024

OPUS: Occupancy Prediction Using a Sparse Set

NeurIPS 2024poster

Occupancy prediction, aiming at predicting the occupancy status within voxelized 3D environment, is quickly gaining momentum within the autonomous driving community. Mainstream occupancy prediction works first discretize the 3D environment into voxels, then perform classification on such dense grids…

2024

Open-World Human-Object Interaction Detection via Multi-modal Prompts

CVPR 2024poster

In this paper we develop MP-HOI a powerful Multi-modal Prompt-based HOI detector designed to leverage both textual descriptions for open-set generalization and visual exemplars for handling high ambiguity in descriptions realizing HOI detection in the open world. Specifically it integrates visual pr…

2024

Pre-trained Online Contrastive Learning for Insurance Fraud Detection

AAAI 2024technical

Medical insurance fraud has always been a crucial challenge in the field of healthcare industry. Existing fraud detection models mostly focus on offline learning scenes. However, fraud patterns are constantly evolving, making it difficult for models trained on past data to detect newly emerging frau…

2024

Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding

EMNLP 2024main

Graphical User Interfaces (GUIs) are central to our interaction with digital devices and growing efforts have been made to build models for various GUI understanding tasks. However, these efforts largely overlook an important GUI-referring task: screen reading based on user-indicated points, which w…

2024

Representation Learning across Feature and Topology Views with Output Correction for Graph Convolutional Networks

ICASSP 2024accepted

In Graph Convolutional Networks (GCNs), the aggregation of node features in graph convolutional learning is typically guided solely by the topology of the graphs. However, both network topology and node features provide unique and valuable information. Relying solely on topology cannot yield entirel…

Cited by 0SourceScholar
2024

Rethinking Fourier Transform from A Basis Functions Perspective for Long-term Time Series Forecasting

NeurIPS 2024poster

The interaction between Fourier transform and deep learning opens new avenues for long-term time series forecasting (LTSF). We propose a new perspective to reconsider the Fourier transform from a basis functions perspective. Specifically, the real and imaginary parts of the frequency components can…

2024

Revealing COVID-19’s Social Dynamics: Diachronic Semantic Analysis of Vaccine and Symptom Discourse on Twitter

EMNLP 2024finding

Social media is recognized as an important source for deriving insights into public opinion dynamics and social impacts due to the vast textual data generated daily and the ‘unconstrained’ behavior of people interacting on these platforms. However, such analyses prove challenging due to the semantic…

Cited by 0SourcePDFScholar
2024

SoftCLIP: Softer Cross-Modal Alignment Makes CLIP Stronger

AAAI 2024technical

During the preceding biennium, vision-language pre-training has achieved noteworthy success on several downstream tasks. Nevertheless, acquiring high-quality image-text pairs, where the pairs are entirely exclusive of each other, remains a challenging task, and noise exists in the commonly used data…

2024

Sparse Bayesian Synthetic Aperture Processing Based DOA Estimation with Deformed Towed Arrays

ICASSP 2024accepted

In this paper, we present a new synthetic aperture method for direction-of-arrival (DOA) estimation using a passive towed sonar array that is deformed during platform maneuver. With certain prior knowledge of source-array geometry, we propose to find the optimal maximum likelihood estimates of DOAs…

Cited by 0SourceScholar
2024

Tackling Vision Language Tasks through Learning Inner Monologues

AAAI 2024technical

Visual language tasks such as Visual Question Answering (VQA) or Visual Entailment (VE) require AI models to comprehend and reason with both visual and textual content. Driven by the power of Large Language Models (LLMs), two prominent methods have emerged: (1) the hybrid integration between LLMs an…

2024

Towards Better Data Exploitation in Self-Supervised Monocular Depth Estimation

RA-L 2024

Depth estimation plays an important role in robotic perception systems. The self-supervised monocular paradigm has gained significant attention since it can free training from the reliance on depth annotations. Despite recent advancements, existing self-supervised methods still underutilize the avai

Cited by 16SourcecodeScholar
2024

VSViG: Real-time Video-based Seizure Detection via Skeleton-based Spatiotemporal ViG

ECCV 2024poster

"An accurate and efficient epileptic seizure onset detection can significantly benefit patients. Traditional diagnostic methods, primarily relying on electroencephalograms (EEGs), often result in cumbersome and non-portable solutions, making continuous patient monitoring challenging. The video-based…

2023

CDFI: Cross Domain Feature Interaction for Robust Bronchi Lumen Detection

ICRA 2023poster

Endobronchial intervention is increasingly used as a minimally invasive means for the treatment of pulmonary diseases. In order to reduce the difficulty of manipulation in complex airway networks, robust lumen detection is essential for intraoperative guidance. However, these methods are sensitive t…

Cited by 1SourceScholar
2023

Design and Modeling of a Hybrid Soft Robotic Manipulator With Compliant Mechanism

RA-L 2023

The recent surge of interest in soft robotics has prompted many interesting soft robotic manipulator designs to be proposed. However, current soft robotic manipulators mainly focus on bending and rotation, while elongation has not received attention due to the limited extension/contraction range or

Cited by 18SourceScholar
2023

Explicit Box Detection Unifies End-to-End Multi-Person Pose Estimation

ICLR 2023poster

This paper presents a novel end-to-end framework with Explicit box Detection for multi-person Pose estimation, called ED-Pose, where it unifies the contextual learning between human-level (global) and keypoint-level (local) information. Different from previous one-stage methods, ED-Pose re-considers…

2023

FoPro: Few-Shot Guided Robust Webly-Supervised Prototypical Learning

AAAI 2023technical

Recently, webly supervised learning (WSL) has been studied to leverage numerous and accessible data from the Internet. Most existing methods focus on learning noise-robust models from web images while neglecting the performance drop caused by the differences between web domain and real-world domain.…

2023

GreenPLM: Cross-Lingual Transfer of Monolingual Pre-Trained Language Models at Almost No Cost

IJCAI 2023poster

Large pre-trained models have revolutionized natural language processing (NLP) research and applications, but high training costs and limited data resources have prevented their benefits from being shared equally amongst speakers of all the world's languages. To address issues of cross-linguistic ac…

2023

NeUDF: Leaning Neural Unsigned Distance Fields With Volume Rendering

CVPR 2023poster

Multi-view shape reconstruction has achieved impressive progresses thanks to the latest advances in neural implicit surface rendering. However, existing methods based on signed distance function (SDF) are limited to closed surfaces, failing to reconstruct a wide range of real-world objects that cont…

Cited by 58SourcePDFScholar
2023

NeuralSlice: Neural 3D Triangle Mesh Reconstruction via Slicing 4D Tetrahedral Meshes

ICML 2023poster

Learning-based high-fidelity reconstruction of 3D shapes with varying topology is a fundamental problem in computer vision and computer graphics. Recent advances in learning 3D shapes using explicit and implicit representations have achieved impressive results in 3D modeling. However, the template-b…

2023

Semantic Human Parsing via Scalable Semantic Transfer Over Multiple Label Domains

CVPR 2023poster

This paper presents Scalable Semantic Transfer (SST), a novel training paradigm, to explore how to leverage the mutual benefits of the data from different label domains (i.e. various levels of label granularity) to train a powerful human parsing network. In practice, two common application scenarios…

2023

USDNL: Uncertainty-Based Single Dropout in Noisy Label Learning

AAAI 2023technical

Deep Neural Networks (DNNs) possess powerful prediction capability thanks to their over-parameterization design, although the large model complexity makes it suffer from noisy supervision. Recent approaches seek to eliminate impacts from noisy labels by excluding data points with large loss values a…

2022

AMOS: A Large-Scale Abdominal Multi-Organ Benchmark for Versatile Medical Image Segmentation

NeurIPS 2022accept

Despite the considerable progress in automatic abdominal multi-organ segmentation from CT/MRI scans in recent years, a comprehensive evaluation of the models' capabilities is hampered by the lack of a large-scale benchmark from diverse clinical scenarios. Constraint by the high cost of collecting an…

2022

Answer Quality Aware Aggregation for Extractive QA Crowdsourcing

EMNLP 2022finding

Quality control is essential for creating extractive question answering (EQA) datasets via crowdsourcing. Aggregation across answers, i.e. word spans within passages annotated, by different crowd workers is one major focus for ensuring its quality. However, crowd workers cannot reach a consensus on…

2022

CROON: Automatic Multi-LiDAR Calibration and Refinement Method in Road Scene

IROS 2022poster

Sensor-based environmental perception is a crucial part of the autonomous driving system. In order to get an excellent perception of the surrounding environment, an intelligent system would configure multiple LiDARs (3D Light Detection and Ranging) to cover the distant and near space of the car. The…

Cited by 18SourcecodeScholar
2022

HSDF: Hybrid Sign and Distance Field for Modeling Surfaces with Arbitrary Topologies

NeurIPS 2022accept

Neural implicit function based on signed distance field (SDF) has achieved impressive progress in reconstructing 3D models with high fidelity. However, such approaches can only represent closed shapes. Recent works based on unsigned distance function (UDF) are proposed to handle both watertight and…

Cited by 21SourcePDFScholar
2022

IFRNet: Intermediate Feature Refine Network for Efficient Frame Interpolation

CVPR 2022poster

Prevailing video frame interpolation algorithms, that generate the intermediate frames from consecutive inputs, typically rely on complex model architectures with heavy parameters or large delay, hindering them from diverse real-time applications. In this work, we devise an efficient encoder-decoder…

Cited by 181PDFcodeScholar
2022

Laplacian Mesh Transformer: Dual Attention and Topology Aware Network for 3D Mesh Classification and Segmentation

ECCV 2022poster

"Deep learning-based approaches for shape understanding and processing tasks have attracted considerable attention. Despite the great progress that has been made, the existing approaches fail to efficiently capture sophisticated structure information and critical part features simultaneously, limiti…

Cited by 18SourcePDFScholar
2022

METS-CoV: A Dataset of Medical Entity and Targeted Sentiment on COVID-19 Related Tweets

NeurIPS 2022accept

The COVID-19 pandemic continues to bring up various topics discussed or debated on social media. In order to explore the impact of pandemics on people's lives, it is crucial to understand the public's concerns and attitudes towards pandemic-related entities (e.g., drugs, vaccines) on social media. H…

2022

Modeling Image Composition for Complex Scene Generation

CVPR 2022poster

We present a method that achieves state-of-the-art results on challenging (few-shot) layout-to-image generation tasks by accurately modeling textures, structures and relationships contained in a complex scene. After compressing RGB images into patch tokens, we propose the Transformer with Focal Atte…

Cited by 57PDFcodeScholar
2021

Discriminative Asymmetric Learning for Efficient Surgical Instrument Parsing

ICRA 2021poster

Semantic segmentation of surgical instruments provides essential priors for autonomous surgery. This task is however challenging since the fine-structure of surgical instruments requires the accurate segmentation of detailed regions in images. As the visual guidance for autonomous surgery, the algor…

Cited by 0SourceScholar
2021

MARTA: Leveraging Human Rationales for Explainable Text Classification

AAAI 2021technical

Explainability is a key requirement for text classification in many application domains ranging from sentiment analysis to medical diagnosis or legal reviews. Existing methods often rely on "attention" mechanisms for explaining classification results by estimating the relative importance of input un…

Cited by 56SourcePDFScholar
2021

OctField: Hierarchical Implicit Functions for 3D Modeling

NeurIPS 2021poster

Recent advances in localized implicit functions have enabled neural implicit representation to be scalable to large scenes. However, the regular subdivision of 3D space employed by these approaches fails to take into account the sparsity of the surface occupancy and the varying granularities of geom…

Cited by 38SourcePDFScholar
2021

Semi-Supervised Skin Lesion Segmentation with Learning Model Confidence

ICASSP 2021accepted

Segmentation of skin lesions is important for disease diagnoses and treatment planning. Over the years, semi-supervised methods using pseudo labels have boosted the segmentation performance with limited labeled data and abundant unlabeled data. However, the unreliable targets in pseudo labels might…

Cited by 0SourceScholar
2021

Single Image 3D Shape Retrieval via Cross-Modal Instance and Category Contrastive Learning

ICCV 2021poster

In this work, we tackle the problem of single image-based 3D shape retrieval (IBSR), where we seek to find the most matched shape of a given single 2D image from a shape repository. Most of the existing works learn to embed 2D images and 3D shapes into a common feature space and perform metric learn…

Cited by 40PDFcodeScholar
2021

Stable and Effective One-Step Method for Person Search

ICASSP 2021accepted

Person search, which requires both pedestrian detection and person re-identification, is a challenging computer vision task applied to real-world scenarios. The challenges faced by detection and re-identification, such as occlusion, poor illumination, confusing background, are still urgent for perso…

Cited by 0SourceScholar
2020

Adaptive Mixture Regression Network with Local Counting Map for Crowd Counting

ECCV 2020poster

The crowd counting task aims at estimating the number of people located in an image or a frame from videos. Existing methods widely adopt density maps as the training targets to optimize the point-to-point loss. While in testing phase, we only focus on the differences between the crowd numbers and t…

2020

Hierarchical Feature Embedding for Attribute Recognition

CVPR 2020poster

Attribute recognition is a crucial but challenging task due to viewpoint changes, illumination variations and appearance diversities, etc. Most of previous work only consider the attribute-level feature embedding, which might perform poorly in complicated heterogeneous conditions. To address this pr…

Cited by 61PDFScholar
2020

Structured Probabilistic End-to-End Learning from Crowds

IJCAI 2020poster

End-to-end learning from crowds has recently been introduced as an EM-free approach to training deep neural networks directly from noisy crowdsourced annotations. It models the relationship between true labels and annotations with a specific type of neural layer, termed as the crowd layer, which can…

Cited by 0SourcePDFScholar
2019

Inchworm-inspired soft climbing robot using microspine arrays

IROS 2019poster

Animals in nature, such as geckos, inchworms, and felines can climb on various surfaces using different mechanisms and serve as references for the study of bio-inspired robots. This paper presents an inchworm-inspired climbing robot that consists of soft body and feet. The soft robot is actuated by…

Cited by 16SourceScholar
2018

Cross-Scene Suture Thread Parsing for Robot Assisted Anastomosis based on Joint Feature Learning

IROS 2018poster

Task autonomy is an important consideration for the development of future surgical robots. For robot-assisted anastomosis, suture thread detection is a prerequisite for subsequent robot manipulation. Previous works on automatic thread detection are focused on the learning of the models with specific…

Cited by 11SourceScholar
2018

Multi-Stage Suture Detection for Robot Assisted Anastomosis Based on Deep Learning

ICRA 2018poster

The technique of robust suture detection is vital in many applications including trainee suturing skill evaluation, suture augmentation in robotic-assisted surgery and suture recognition for automatic suturing. Due to the complicated environment of surgery, the detection of a suture is challenged by…

Cited by 13SourceScholar
2018

Seeing Deeply and Bidirectionally: A Deep Learning Approach for Single Image Reflection Removal

ECCV 2018poster

Reflections often obstruct the desired scene when taking photos through glass panels. Removing unwanted reflection automatically from the photos is highly desirable. Traditional methods often impose certain priors or assumptions to target particular type(s) of reflection such as shifted double refle…

2018

Simultaneous Accurate Detection of Pulmonary Nodules and False Positive Reduction Using 3D CNNs

ICASSP 2018accepted

Accurate detection of nodules in CT images is vital for lung cancer diagnosis, which greatly influences the patient's chance for survival. Motivated by successful application of convolutional neural networks (CNNs) on natural images, we propose a computer-aided diagnosis (CAD) system for simultaneou…

Cited by 0SourceScholar
2017

Chromatic surface microstructures on bionic soft robots for non-contact deformation measurement

ICRA 2017poster

This paper presents a bionic soft robot with chromatic surface micro-structure (CSM), as a new approach for the measurement of body deformation of the soft robots. Firstly, the CSM films are fabricated by diffraction gratings mold using material polydimethylsiloxane (PDMS). Then, the CSM films are a…

Cited by 6SourceScholar
2017

From Motion Blur to Motion Flow: A Deep Learning Solution for Removing Heterogeneous Motion Blur

CVPR 2017poster

Removing pixel-wise heterogeneous motion blur is challenging due to the ill-posed nature of the problem. The predominant solution is to estimate the blur kernel by adding a prior, but extensive literature on the subject indicates the difficulty in identifying a prior which is suitably informative, a…

Cited by 504PDFScholar
2016

A MIL-based interactive approach for hotspot segmentation from bone scintigraphy

ICASSP 2016accepted

Bone scintigraphy is widely used to diagnose bone diseases. Accurate hotspot segmentation is a critical task for tumor metastasis diagnosis. In this paper, we propose an interactive approach to detect and extract hotspots in thoracic region based on a new multiple instance learning (MIL) method call…

Cited by 0SourceScholar
2016

Robust visual tracking via inverse nonnegative matrix factorization

ICASSP 2016accepted

The establishment of robust target appearance model over time is an overriding concern in visual tracking. In this paper, we propose an inverse nonnegative matrix factorization (NMF) method for robust appearance modeling. Rather than using a linear combination of nonnegative basis vectors for each t…

Cited by 0SourceScholar
2015

Modeling and closed-loop control of electromagnetic manipulation of a microparticle

ICRA 2015poster

Precise manipulation of microparticles has received considerable attention for its great potential applications to clinical medicine. Among the existing manipulation techniques, the method of magnetic force based manipulation exhibits great advantages for its minimally-invasive feature and insensiti…

Cited by 8SourceScholar
2015

Saliency Propagation From Simple to Difficult

CVPR 2015poster

Saliency propagation has been widely adopted for identifying the most attractive object in an image. The propagation sequence generated by existing saliency detection methods is governed by the spatial relationships of image regions, i.e., the saliency value is transmitted between two adjacent regio…

Cited by 179SourcePDFScholar
2015

Small target detection using an optimization-based filter

ICASSP 2015accepted

Small target detection is a critical problem in the Infrared Search And Track (IRST) system. Although it has been studied for years, there are some challenges remained, e.g. cloud edges and horizontal lines are likely to cause false alarms. This paper proposes a novel method using an optimization-ba…

Cited by 0SourceScholar