← Search

BIN YANG

138 accepted papers

2026

A Comprehensive Survey of Deep Learning for Multivariate Time Series Forecasting: A Channel Strategy Perspective

IJCAI 2026

Multivariate Time Series Forecasting (MTSF) plays a crucial role across diverse fields, ranging from economic, energy, to traffic. In recent years, deep learning has demonstrated outstanding performance in MTSF tasks. In MTSF, modeling the correlations among different channels is critical, as levera

Cited by 0Scholar
2026

API: Adaptive Prototype Imputation for Incomplete Multimodal Sentiment Analysis

ICML 2026poster

Multimodal sentiment analysis aims to infer human emotions by integrating signals from diverse modalities. However, missing modalities are common in real-world applications due to sensor failure, data corruption, or privacy concerns. Existing approaches typically follow two main paradigms: recovery-…

Cited by 0SourceScholar
2026

ARROW: An Adaptive Rollout and Routing Method for Global Weather Forecasting

ICLR 2026poster

Weather forecasting is a fundamental task in spatiotemporal data analysis, with broad applications across a wide range of domains. Existing data-driven forecasting methods typically model atmospheric dynamics over a fixed short time interval, e.g., 6 hours, and rely on naive autoregression-based rol…

Cited by 0SourcecodeScholar
2026

ASTGI: Adaptive Spatio-Temporal Graph Interactions for Irregular Multivariate Time Series Forecasting

ICLR 2026poster

Irregular multivariate time series (IMTS) are prevalent in critical domains like healthcare and finance, where accurate forecasting is vital for proactive decision-making. However, the asynchronous sampling and irregular intervals inherent to IMTS pose two core challenges for existing methods: (1) h…

Cited by 0SourcecodeScholar
2026

Accelerating Langevin Monte Carlo via Efficient Stochastic Runge-Kutta Methods beyond Log-Concavity

ICML 2026poster

Sampling from a high-dimensional probability distribution is a fundamental algorithmic task arising in wide-ranging applications across multiple disciplines, including scientific computing, computational statistics and machine learning. Langevin Monte Carlo (LMC) algorithms are among the most widely…

Cited by 0SourceScholar
2026

Aurora: Towards Universal Generative Multimodal Time Series Forecasting

ICLR 2026poster

Cross-domain generalization is very important in Time Series Forecasting because similar historical information may lead to distinct future trends due to the domain-specific characteristics. Recent works focus on building unimodal time series foundation models and end-to-end multimodal supervised mo…

Cited by 0SourcecodeScholar
2026

Automatic Unsupervised Ensemble Outlier Model Selection

ICML 2026poster

Unsupervised outlier detection is attractive because it eliminates the need for labeled data. Further, forming multi-model ensembles can improve detection robustness performance. However, composing an ensemble without labeled data is challenging. Naively composing ensembles can cause ensemble satura…

Cited by 0SourceScholar
2026

CoRA: Boosting Time Series Foundation Models for Multivariate Forecasting through Correlation-aware Adapter

ICLR 2026poster

Most existing Time Series Foundation Models (TSFMs) use channel independent modeling and focus on capturing and generalizing temporal dependencies, while neglecting the correlations among channels or overlook the different aspects of correlations. However, these correlations play a vital role in Mul…

Cited by 0SourcecodeScholar
2026

Cross-Modal Semantic Decoupling and Transfer for Text-to-Visible-Infrared Person Re-Identification

ICML 2026poster

Text-to-Image Person Re-Identification (TI-ReID) retrieves visible pedestrian images using text queries. Yet in low-light or nighttime settings, visible images lack sufficient identity details, while infrared images effectively capture pedestrian contours and textures. To enable all-day surveillance…

Cited by 0SourceScholar
2026

D-CORE: Incentivizing Task Decomposition in Large Reasoning Models for Complex Tool Use

ICML 2026poster

Effective tool use and reasoning are essential capabilities for large reasoning models (LRMs) to address complex real-world problems. Through empirical analysis, we identify a prevalent "Lazy Reasoning" phenomenon, where LRMs frequently engage in repetitive and meaningless reflective reasoning. This…

Cited by 0SourceScholar
2026

DAG: A Dual Correlation Network for Time Series Forecasting with Exogenous Variables

ICML 2026poster

Time series forecasting is essential in various domains. Compared to relying solely on endogenous variables (i.e., target variables), considering exogenous variables (i.e., covariates) provides additional predictive information and often leads to more accurate predictions. However, existing methods …

Cited by 0SourceScholar
2026

DGCPath: Distribution-Aware Generative Contrastive Framework for Self-supervised Path Representation Learning

IJCAI 2026

Due to proliferation of vehicle trajectory data from advanced sensing technologies, path representation learning has become a pivotal task in intelligent transportation systems. Although existing self-supervised approaches work to some extent, their dependence on deterministic contrastive learning p

Cited by 0Scholar
2026

DOS: Distilling Observable Softmaps of Zipfian Prototypes for Self-Supervised Point Representation

AAAI 2026technical

Recent advances in self-supervised learning (SSL) have shown tremendous potential for learning 3D point cloud representations without human annotations. However, SSL for 3D point clouds still faces critical challenges due to irregular geometry, shortcut-prone reconstruction, and unbalanced semantics

Cited by 0SourcePDFScholar
2026

Domain-Aware Suppression and Aggregation for Federated DG ReID

AAAI 2026technical

Federated domain generalization in person re-identification (FedDG-ReID) aims to learn a privacy-preserving server model from decentralized client source domains that generalizes to unseen domains. Existing approaches enhance the generalizability of the server model by increasing the diversity of c

Cited by 0SourcePDFScholar
2026

FedBPrompt: Federated Domain Generalization Person Re-Identification via Body Distribution Aware Visual Prompts

CVPR 2026

Federated Domain Generalization for Person Re-Identification (FedDG-ReID) aims to learn domain-invariant representations from decentralized data. Although Vision Transformers (ViTs) are widely adopted, their global attention often fails to distinguish pedestrians from high similarity backgrounds or

Cited by 0SourcecodeScholar
2026

GCGNet: Graph-Consistent Generative Network for Time Series Forecasting with Exogenous Variables

ICLR 2026poster

Exogenous variables offer valuable supplementary information for predicting future endogenous variables. Forecasting with exogenous variables needs to consider both past-to-future dependencies (i.e., temporal correlations) and the influence of exogenous variables on endogenous variables (i.e., chann…

Cited by 0SourcecodeScholar
2026

Interactive Person Retrieval via Multi-Turn Multimodal Conversation

ICML 2026poster

Traditional text-based person retrieval approaches typically rely on single-shot textual queries, which are generally incomplete or vague in real-world scenarios. Recently, chat-based person retrieval methods enable iterative query refinement via question-answering interactions between the system an…

Cited by 0SourceScholar
2026

Invariant Representation Learning for Source-Free Time Series Forecasting with LLM-Centric Proxy Denoising

ICML 2026poster

Effective time series forecasting enables various real-world applications, benefiting from the proliferation of mobile devices. However, the volume of time series data may vary significantly across domains due to low sampling rates and data regulations. To maximally create value from sparse data, th…

Cited by 0SourceScholar
2026

Learning-Based Observer for Coupled Disturbance

ICRA 2026poster

Achieving high-precision control for robotic systems is hindered by the low-fidelity dynamical model and external disturbances. Especially, the intricate coupling between internal uncertainties and external disturbances further exacerbates this challenge. This study introduces an effective and conve…

2026

Long Live The Balance: Information Bottleneck Driven Tree-based Policy Optimization

ICML 2026poster

Recent advances in online reinforcement learning (RL) for large language models (LLMs) have demonstrated promising performance in complex reasoning tasks. However, they often exhibit an imbalanced exploration–exploitation trade-off, resulting in unstable optimization and sub-optimal performance. We …

Cited by 0SourceScholar
2026

MdaIF: Robust One-Stop Multi-Degradation-Aware Image Fusion with Language-Driven Semantics

AAAI 2026technical

Infrared and visible image fusion aims to integrate complementary multi-modal information into a single fused result. However, existing methods 1) fail to account for the degradation visible images under adverse weather conditions, thereby compromising fusion performance; and 2) rely on fixed networ

Cited by 0SourcePDFScholar
2026

Multi-View Ensemble for Time Series Anomaly Detection via Coupling Flows

IJCAI 2026

Time series anomaly detection faces a critical challenge that different anomaly types require different detection mechanisms, yet single methods are inherently limited by their design biases. We propose FlowFuse, a multi-view ensemble framework with coupling flow-based score fusion for time series a

Cited by 0Scholar
2026

PATRA: Pattern-Aware Alignment and Balanced Reasoning for Time Series Question Answering

ICML 2026poster

Time series reasoning demands both the perception of complex dynamics and logical depth. However, existing LLM-based approaches exhibit two limitations: they often treat time series merely as text or images, failing to capture the patterns like trends and seasonalities needed to answer specific ques…

Cited by 0SourceScholar
2026

PCB-Bench: Benchmarking LLMs for Printed Circuit Board Placement and Routing

ICLR 2026poster

Recent advances in Large Language Models (LLMs) have enabled impressive capabilities across diverse reasoning and generation tasks. However, their ability to understand and operate on real-world engineering problems—such as Printed Circuit Board (PCB) placement and routing—remains underexplored due…

Cited by 0SourcecodeScholar
2026

Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents

ICML 2026spotlight

While GUI agents have advanced rapidly, they often lack the robustness to recover from their own errors, hindering real-world deployment. To bridge this gap at both the evaluation and data levels, we introduce GUI-RobustEval and propose Robustness-driven Trajectory Synthesis. GUI-RobustEval containi…

Cited by 0SourceScholar
2026

Rethinking Irregular Time Series Forecasting: A Simple Yet Effective Baseline

AAAI 2026technical

The forecasting of irregular multivariate time series (IMTS) is crucial in key areas such as healthcare, biomechanics, climate science, and astronomy. However, achieving accurate and practical predictions is challenging due to two main factors. First, the inherent irregularity and data missingness i

Cited by 0SourcePDFScholar
2026

SEER: Transformer-based Robust Time Series Forecasting via Automated Patch Enhancement and Replacement

ICML 2026poster

Time series forecasting is important in many fields that require accurate predictions for decision-making. Patching techniques, commonly used and effective in time series modeling, help capture temporal dependencies by dividing the data into patches. However, existing patch-based methods fail to dyn…

Cited by 0SourceScholar
2026

Similarity-Consistent Likelihood Diffusion enables Hidden Person Detection from Wall Reflections

CVPR 2026

Non-line-of-sight (NLOS) imaging seeks to recover hidden-scene information from indirect light transport beyond the direct line of sight. Existing NLOS methods can be broadly categorized into active and passive approaches. Active methods rely on controlled illumination and time-resolved sensors, but

Cited by 0SourceScholar
2026

SwiftTS: A Swift Selection Framework for Time Series Pre-trained Models via Multi-task Meta-Learning

ICLR 2026poster

Pre-trained models exhibit strong generalization to various downstream tasks. However, given the numerous models available in the model hub, identifying the most suitable one by individually fine-tuning is time-consuming. In this paper, we propose \textbf{SwiftTS}, a swift selection framework for ti…

Cited by 0SourcecodeScholar
2026

TeamWork: Multivariate Time Series Anomaly Detection via Asymmetric Role-aware Channel Modeling

ICML 2026poster

Multivariate time series anomaly detection remains challenging as it requires the joint modeling of variable relationships and temporal dependencies. Existing methods often struggle to balance channel relationship modeling and overlook the relative importance of different variables within multivaria…

Cited by 0SourceScholar
2026

Towards Cross-Modal Preservation, Consistency and Alignment for Privacy-Preserving Visible-Infrared Person Re-Identification

CVPR 2026

Privacy-preserving Person Re-Identification (PP-ReID) addresses the core privacy-utility trade-off in Re-ID by retrieving a person across multiple non-overlapping cameras while applying anonymization techniques to protect sensitive information. However, prior PP-ReID studies are confined to single-m

Cited by 0SourcecodeScholar
2026

Towards Foundation Models for 3D Scene Understanding: Instance-Aware Self-Supervised Learning for Point Clouds

CVPR 2026

Recent advances in self-supervised learning (SSL) for point clouds have substantially improved 3D scene understanding without human annotations. Existing approaches emphasize semantic awareness by enforcing feature consistency across augmented views or by masked scene modeling. However, the resultin

Cited by 0SourceScholar
2026

Towards Multimodal Time Series Anomaly Detection with Semantic Alignment and Condensed Interaction

ICLR 2026poster

Time series anomaly detection plays a critical role in many dynamic systems. However, previous approaches have primarily relied on unimodal numerical data, overlooking the importance of complementary information from other modalities. In this paper, we propose a novel multimodal time series anomaly…

Cited by 0SourcecodeScholar
2026

Towards Non-Stationary Time Series Forecasting with Temporal Stabilization and Frequency Differencing

AAAI 2026technical

Time series forecasting is critical for decision making across dynamic domains such as energy, finance, transportation, and cloud computing. However, real-world time series often exhibit non-stationarity, including temporal distribution shifts and spectral variability, which poses significant challe

Cited by 0SourcePDFScholar
2026

UniABG: Unified Adversarial View Bridging and Graph Correspondence for Unsupervised Cross-View Geo-Localization

AAAI 2026technical

Cross-view geo-localization (CVGL) matches query images (e.g., drone) to geographically corresponding opposite-view imagery (e.g., satellite). While supervised methods achieve strong performance, their reliance on extensive pairwise annotations limits scalability. Unsupervised alternatives avoid ann

Cited by 0SourcePDFScholar
2026

Unlocking the Value of Text: Event-Driven Reasoning and Multi-Level Alignment for Time Series Forecasting

ICLR 2026poster

Existing time series forecasting methods primarily rely on the numerical data itself. However, real-world time series exhibit complex patterns associated with multimodal information, making them difficult to predict with numerical data alone. While several multimodal time series forecasting methods…

Cited by 0SourcecodeScholar
2026

WHU-MARS: A Multispectral Aerial-Ground Benchmark Towards Any-Scenario Person Re-Identification

CVPR 2026

Recent person re-identification (ReID) leverages heterogeneous sensing with multiple modalities and viewpoints to improve robustness across diverse conditions. However, most approaches target predefined scenario pairs (e.g., visible-infrared or aerial-ground) and train separate task-specific models.

Cited by 0SourcecodeScholar
2025

$K^2$VAE: A Koopman-Kalman Enhanced Variational AutoEncoder for Probabilistic Time Series Forecasting

ICML 2025spotlight

Probabilistic Time Series Forecasting (PTSF) plays a crucial role in decision-making across various fields, including economics, energy, and transportation. Most existing methods excell at short-term forecasting, while overlooking the hurdles of Long-term Probabilistic Time Series Forecasting (LPTSF…

2025

A Deformable-Based Source-Free Unsupervised Domain Adaptation Method for Cervical Cell Detection

ICASSP 2025accepted

As the application of AI in cervical cancer cell detection expands, the demand for large-scale labeled pathological slide data has increased, resulting in a time-consuming and costly process. Furthermore, significant domain shifts caused by variations in sample collection, processing, and staining c…

Cited by 0SourceScholar
2025

Air Quality Prediction with Physics-Guided Dual Neural ODEs in Open Systems

ICLR 2025poster

Air pollution significantly threatens human health and ecosystems, necessitating effective air quality prediction to inform public policy. Traditional approaches are generally categorized into physics-based and data-driven models. Physics-based models usually struggle with high computational demands…

Cited by 3SourcePDFScholar
2025

An Empirical Study of Federated Prompt Learning for Vision Language Model

IJCAI 2025

The Vision Language Model (VLM) excels in aligning vision and language representations, and prompt learning has emerged as a key technique for adapting such models to downstream tasks. However, the application of prompt learning with VLM in federated learning (FL) scenarios remains underexplored. Th

2025

Assessing Pre-Trained Models for Transfer Learning Through Distribution of Spectral Components

AAAI 2025technical

Pre-trained model assessment for transfer learning aims to identify the optimal candidate for the downstream tasks from a model hub, without the need of time-consuming fine-tuning. Existing advanced works mainly focus on analyzing the intrinsic characteristics of the entire features extracted by eac…

Cited by 0SourcePDFScholar
2025

CATCH: Channel-Aware Multivariate Time Series Anomaly Detection via Frequency Patching

ICLR 2025poster

Anomaly detection in multivariate time series is challenging as heterogeneous subsequence anomalies may occur. Reconstruction-based methods, which focus on learning normal patterns in the frequency domain to detect diverse abnormal subsequences, achieve promising results, while still falling short o…

2025

Class-Aware PillarMix: Can Mixed Sample Data Augmentation Enhance 3D Object Detection with Radar Point Clouds?

IROS 2025

Due to the significant effort required for data collection and annotation in 3D perception tasks, mixed sample data augmentation (MSDA) has been widely studied to generate diverse training samples by mixing existing data. Among these methods, MixUp is a prominent approach that generates new samples

Cited by 1SourceScholar
2025

CrossAD: Time Series Anomaly Detection with Cross-scale Associations and Cross-window Modeling

NeurIPS 2025poster

Time series anomaly detection plays a crucial role in a wide range of real-world applications. Given that time series data can exhibit different patterns at different sampling granularities, multi-scale modeling has proven beneficial for uncovering latent anomaly patterns that may not be apparent at…

Cited by 0SourceScholar
2025

DBLoss: Decomposition-based Loss Function for Time Series Forecasting

NeurIPS 2025poster

Time series forecasting holds significant value in various domains such as economics, traffic, energy, and AIOps, as accurate predictions facilitate informed decision-making. However, the existing Mean Squared Error (MSE) loss function sometimes fails to accurately capture the seasonality or trend w…

Cited by 0SourceScholar
2025

Enhancing Diversity for Data-free Quantization

CVPR 2025poster

Model quantization is an effective way to compress deep neural networks and accelerate the inference time on edge devices. Existing quantization methods usually require original data for calibration during the compressing process, which may be inaccessible due to privacy issues. A common way is to g…

Cited by 1SourcePDFScholar
2025

Exploring the Distribution of Cell Subpopulations in Pancreatic Ductal Adenocarcinoma Slides by Joint Spatial Transcriptomics and Pathology Data

ICASSP 2025accepted

Current spatial transcriptomics (ST) technology can integrate stained pathological slides with RNA sequencing, providing precise information on gene expression and cell types. However, the unique handling of pathological slides required by ST technology can lead to image quality issues, impacting mo…

Cited by 0SourceScholar
2025

Federated Disentangled Tuning with Textual Prior Decoupling and Visual Dynamic Adaptation

ICML 2025poster

Federated Parameter-Efficient Fine-Tuning aims to adapt Vision-Language Models for downstream tasks in distributed environments. However, data heterogeneity across participants hinders collaborative effectiveness, necessitating personalized adaptation to cover distinct data distributions. Current pe…

2025

Learning Generalizable Skills from Offline Multi-Task Data for Multi-Agent Cooperation

ICLR 2025poster

Learning cooperative multi-agent policy from offline multi-task data that can generalize to unseen tasks with varying numbers of agents and targets is an attractive problem in many scenarios. Although aggregating general behavior patterns among multiple tasks as skills to improve policy transfer is…

2025

Learning to Factorize Spatio-Temporal Foundation Models

NeurIPS 2025spotlight

Spatio-Temporal Foundation Models (STFMs) promise zero/few-shot generalization across various datasets, yet joint spatio-temporal pretraining is computationally prohibitive and struggles with domain-specific spatial correlations. To this end, we introduce FactoST, a factorized STFM that decouples un…

Cited by 0SourceScholar
2025

LightGTS: A Lightweight General Time Series Forecasting Model

ICML 2025poster

Existing works on general time series forecasting build foundation models with heavy model parameters through large-scale multi-source pretraining. These models achieve superior generalization ability across various datasets at the cost of significant computational burdens and limitations in resourc…

Cited by 0SourcePDFScholar
2025

No More Sibling Rivalry: Debiasing Human-Object Interaction Detection

ICCV 2025poster

Detection transformers have been applied to human-object interaction (HOI) detection, enhancing the localization and recognition of human-action-object triplets in images. Despite remarkable progress, this study identifies a critical issue--"Toxic Siblings" bias--which hinders the interaction decode…

Cited by 0SourcePDFScholar
2025

Non-asymptotic Error Bounds in $\mathcal{W}_2$-Distance with Sqrt(d) Dimension Dependence and First Order Convergence for Langevin Monte Carlo beyond Log-Concavity

ICML 2025poster

Generating samples from a high dimensional probability distribution is a fundamental task with wide-ranging applications in the area of scientific computing, statistics and machine learning. This article revisits the popular Langevin Monte Carlo (LMC) sampling algorithms and provides a non-asymptoti…

Cited by 0SourcePDFScholar
2025

Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models

CVPR 2025award

Latent diffusion models with Transformer architectures excel at generating high-fidelity images. However, recent studies reveal an optimization dilemma in this two-stage design: while increasing the per-token feature dimension in visual tokenizers improves reconstruction quality, it requires substan…

2025

Rethinking Fair Federated Learning from Parameter and Client View

NeurIPS 2025poster

Federated Learning is a promising technique that enables collaborative machine learning while preserving participant privacy. With respect to multi-party collaboration, achieving performance fairness acts as a critical challenge in federated systems. Existing explorations mainly focus on considering…

Cited by 0SourcecodeScholar
2025

SPMC: Self-Purifying Federated Backdoor Defense via Margin Contribution

ICML 2025poster

Federated Learning (FL) enables collaborative training with privacy preservation but is vulnerable to backdoor attacks, where malicious clients degrade model performance on targeted inputs. These attacks exploit FL decentralized nature, while existing defenses, based on isolated behaviors and fixed…

2025

Splitting with Importance-aware Updating for Heterogeneous Federated Learning with Large Language Models

ICML 2025poster

Federated learning provides an efficient privacy-preserving distributed training framework for large language models, addressing the growing scarcity of publicly available training data while enabling the utilization of private datasets. While integrating large language model fine-tuning with federa…

2025

TokenMatcher: Diverse Tokens Matching for Unsupervised Visible-Infrared Person Re-Identification

AAAI 2025technical

Unsupervised visible-infrared person re-identification (US-VI-ReID) seeks to match infrared and visible images of the same individual without the use of annotations. Current methods typically derive cross-modal correspondences through a single global feature matching process for generating pseudo la…

2025

Towards a General Time Series Anomaly Detector with Adaptive Bottlenecks and Dual Adversarial Decoders

ICLR 2025poster

Time series anomaly detection plays a vital role in a wide range of applications. Existing methods require training one specific model for each dataset, which exhibits limited generalization capability across different target datasets, hindering anomaly detection performance in various scenarios wit…

Cited by 5SourcePDFScholar
2025

Towards a General Time Series Forecasting Model with Unified Representation and Adaptive Transfer

ICML 2025poster

With the growing availability of multi-domain time series data, there is an increasing demand for general forecasting models pre-trained on multi-source datasets to support diverse downstream prediction scenarios. Existing time series foundation models primarily focus on scaling up pre-training data…

Cited by 0SourcePDFScholar
2025

Unbiased Prototype Consistency Learning for Multi-Modal and Multi-Task Object Re-Identification

NeurIPS 2025spotlight

In object re-identification (ReID) task, both cross-modal and multi-modal retrieval methods have achieved notable progress. However, existing approaches are designed for specific modality and category (person or vehicle) retrieval task, lacking generalizability to others. Acquiring multiple task-spe…

Cited by 0SourcecodeScholar
2025

Unsupervised Visible-Infrared Person Re-identification under Unpaired Settings

ICCV 2025poster

Unsupervised visible-infrared person re-identification (USL-VI-ReID) aims to train a cross-modality retrieval model without labels, reducing the reliance on expensive cross-modality manual annotation. However, existing USL-VI-ReID methods rely on artificially cross-modality paired data as implicit s…

2024

Deep Regression for Biological Age Estimation in Multiple Organs: Investigations on 40, 000 Subjects of the UK Biobank

ICASSP 2024accepted

Age plays an important role in shaping medical decisions, but the biological changes associated with aging do not solely depend on the chronological age. Genetics, lifestyle, and environment cause variations in age-related characteristics, even within the same chronological age group. Biological age…

Cited by 0SourceScholar
2024

Dependency-aware Differentiable Neural Architecture Search

ECCV 2024poster

"UTF8gbsn Neural architecture search (NAS) reduces the burden of manual design by automatically building neural network architectures, among which differential NAS approaches such as DARTS, have gained popularity for the search efficiency. Despite achieving promising performance, the DARTS series me…

Cited by 2SourcePDFScholar
2024

Diversity-Aware Buffer for Coping with Temporally Correlated Data Streams in Online Test-Time Adaptation

ICASSP 2024accepted

Since distribution shifts are likely to occur after a model’s deployment and can drastically decrease the model’s performance, online test-time adaptation (TTA) continues to update the model during test-time, leveraging the current test data. In real-world scenarios, test data streams are not always…

Cited by 0SourceScholar
2024

Empowering Visible-Infrared Person Re-Identification with Large Foundation Models

NeurIPS 2024poster

Visible-Infrared Person Re-identification (VI-ReID) is a challenging cross-modal retrieval task due to significant modality differences, primarily resulting from the absence of color information in the infrared modality. The development of large foundation models like Large Language Models (LLMs) an…

Cited by 3SourcePDFScholar
2024

Part2Object: Hierarchical Unsupervised 3D Instance Segmentation

ECCV 2024poster

"Unsupervised 3D instance segmentation aims to segment objects from a 3D point cloud without any annotations. Existing methods face the challenge of either too loose or too tight clustering, leading to under-segmentation or over-segmentation. To address this issue, we propose Part2Object, hierarchic…

2024

Pathformer: Multi-scale Transformers with Adaptive Pathways for Time Series Forecasting

ICLR 2024poster

Transformers for time series forecasting mainly model time series from limited or fixed scales, making it challenging to capture different characteristics spanning various scales. We propose Pathformer, a multi-scale Transformer with adaptive pathways. It integrates both temporal resolution and temp…

2024

Position: What Can Large Language Models Tell Us about Time Series Analysis

ICML 2024poster

Time series analysis is essential for comprehending the complexities inherent in various real-world systems and applications. Although large language models (LLMs) have recently made significant strides, the development of artificial general intelligence (AGI) equipped with time series analysis capa…

Cited by 36SourcePDFScholar
2024

Shallow-Deep Collaborative Learning for Unsupervised Visible-Infrared Person Re-Identification

CVPR 2024poster

Unsupervised visible-infrared person re-identification (US-VI-ReID) centers on learning a cross-modality retrieval model without labels reducing the reliance on expensive cross-modality manual annotation. Previous US-VI-ReID works gravitate toward learning cross-modality information with the deep fe…

2024

TULIP: Transformer for Upsampling of LiDAR Point Clouds

CVPR 2024poster

LiDAR Upsampling is a challenging task for the perception systems of robots and autonomous vehicles due to the sparse and irregular structure of large-scale scene contexts. Recent works propose to solve this problem by converting LiDAR data from 3D Euclidean space into an image super-resolution prob…

2023

DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language Models

NeurIPS 2023poster

A long-standing goal of AI systems is to perform complex multimodal reasoning like humans. Recently, large language models (LLMs) have made remarkable strides in such multi-step reasoning on the language modality solely by leveraging the chain of thought (CoT) to mimic human thinking. However, the t…

Cited by 100SourcePDFScholar
2023

FLAG3D: A 3D Fitness Activity Dataset With Language Instruction

CVPR 2023poster

With the continuously thriving popularity around the world, fitness activity analytic has become an emerging research topic in computer vision. While a variety of new tasks and algorithms have been proposed recently, there are growing hunger for data resources involved in high-quality data, fine-gra…

2023

NIFF: Alleviating Forgetting in Generalized Few-Shot Object Detection via Neural Instance Feature Forging

CVPR 2023poster

Privacy and memory are two recurring themes in a broad conversation about the societal impact of AI. These concerns arise from the need for huge amounts of data to train deep neural networks. A promise of Generalized Few-shot Object Detection (G-FSOD), a learning paradigm in AI, is to alleviate the…

Cited by 16SourcePDFScholar
2023

Robust Mean Teacher for Continual and Gradual Test-Time Adaptation

CVPR 2023poster

Since experiencing domain shifts during test-time is inevitable in practice, test-time adaption (TTA) continues to adapt the model after deployment. Recently, the area of continual and gradual test-time adaptation (TTA) emerged. In contrast to standard TTA, continual TTA considers not only a single…

2023

Scale-Adaptive Tiny Object Detection Enhanced by Across-Scale and Shape-Preserved Semantic Location

ICASSP 2023accepted

In tiny object detection, the main challenges are tiny objects’ weak feature responses and possible semantic disappearance in deep networks. To address the problems, we proposed an Instance-level, Scale-adaptive, Shape-preserved, and Semantic-consistent Supervision (I4S) module for better locating t…

Cited by 0SourceScholar
2023

Top-K Visual Tokens Transformer: Selecting Tokens for Visible-Infrared Person Re-Identification

ICASSP 2023accepted

Visible modality and infrared modality person re-identification (VI-ReID) is an extremely important and challenging task. Existing works mainly focus on reducing the modality gap with Convolutional Neural Networks (CNN). However, the features extracted by CNN may contain useless identity-irrelevant…

Cited by 0SourceScholar
2023

Towards Grand Unified Representation Learning for Unsupervised Visible-Infrared Person Re-Identification

ICCV 2023poster

Unsupervised learning visible-infrared person re-identification (USL-VI-ReID) is an extremely important and challenging task, which can alleviate the issue of expensive cross-modality annotations. Existing works focus on handling the cross-modality discrepancy under unsupervised conditions. However,…

Cited by 52PDFcodeScholar
2023

Towards Unsupervised Object Detection From LiDAR Point Clouds

CVPR 2023poster

In this paper, we study the problem of unsupervised object detection from 3D point clouds in self-driving scenes. We present a simple yet effective method that exploits (i) point clustering in near-range areas where the point clouds are dense, (ii) temporal consistency to filter out noisy unsupervis…

Cited by 40SourcePDFScholar
2023

Visual Elements Mining as Prompts for Instruction Learning for Target-Oriented Multimodal Sentiment Classification

EMNLP 2023long findings

Target-oriented Multimodal Sentiment Classification (TMSC) aims to incorporate visual modality with text modality to identify the sentiment polarity towards a specific target within a sentence. To address this task, we propose a Visual Elements Mining as Prompts (VEMP) method, which describes the se…

Cited by 0SourceScholar
2022

An Unsupervised Domain Adaptive Approach for Multimodal 2D Object Detection in Adverse Weather Conditions

IROS 2022poster

Integrating different representations from complementary sensing modalities is crucial for robust scene interpretation in autonomous driving. While deep learning architectures that fuse vision and range data for 2D object detection have thrived in recent years, the corresponding modalities can degra…

Cited by 8SourceScholar
2022

BaLeNAS: Differentiable Architecture Search via the Bayesian Learning Rule

CVPR 2022poster

Differentiable Architecture Search (DARTS) has received massive attention in recent years, mainly because it significantly reduces the computational cost through weight sharing and continuous relaxation. However, more recent works find that existing differentiable NAS techniques struggle to outperfo…

Cited by 24PDFScholar
2022

Hyperverlet: A Symplectic Hypersolver for Hamiltonian Systems

AAAI 2022technical

Hamiltonian systems represent an important class of dynamical systems such as pendulums, molecular dynamics, and cosmic systems. The choice of solvers is significant to the accuracy when simulating Hamiltonian systems, where symplectic solvers show great significance. Recent advances in neural netwo…

2022

Interpreting Operation Selection in Differentiable Architecture Search: A Perspective from Influence-Directed Explanations

NeurIPS 2022accept

The Differentiable ARchiTecture Search (DARTS) has dominated the neural architecture search community due to its search efficiency and simplicity. DARTS leverages continuous relaxation to convert the intractable operation selection problem into a continuous magnitude optimization problem which can b…

Cited by 3SourcePDFScholar
2022

MT3: Meta Test-Time Training for Self-Supervised Test-Time Adaption

AISTATS 2022poster

An unresolved problem in Deep Learning is the ability of neural networks to cope with domain shifts during test-time, imposed by commonly fixing network parameters after training. Our proposed method Meta Test-Time Training (MT3), however, breaks this paradigm and enables adaption at test-time. We c…

2022

REMOTE: Reinforced Motion Transformation Network for Semi-supervised 2D Pose Estimation in Videos

AAAI 2022technical

Existing approaches for 2D pose estimation in videos often require a large number of dense annotations, which are costly and labor intensive to acquire. In this paper, we propose a semi-supervised REinforced MOtion Transformation nEtwork (REMOTE) to leverage a few labeled frames and temporal pose va…

Cited by 13SourcePDFScholar
2022

Simultaneous Depth Estimation and Localization for Cell Manipulation Based on Deep Learning

IROS 2022poster

Visual localization, which is a key technology to realize the automation of cell manipulation, has been widely studied. Since the depth of field of the microscope is narrow, the planar localization and depth estimation are usually coupled together. At present, most methods adopt the serial working m…

Cited by 3SourceScholar
2022

Triformer: Triangular, Variable-Specific Attentions for Long Sequence Multivariate Time Series Forecasting

IJCAI 2022poster

A variety of real-world applications rely on far future information to make decisions, thus calling for efficient and accurate long sequence multivariate time series forecasting. While recent attention-based forecasting models show strong abilities in capturing long-term dependencies, they still su…

2022

Wavelet-Based Unsupervised Label-to-Image Translation

ICASSP 2022accepted

Semantic Image Synthesis (SIS) is a subclass of image-to-image translation where a semantic layout is used to generate a photorealistic image. State-of-the-art conditional Generative Adversarial Networks (GANs) need a huge amount of paired data to accomplish this task while generic un-paired image-t…

Cited by 0SourceScholar
2022

Weighted Mutual Learning with Diversity-Driven Model Compression

NeurIPS 2022accept

Online distillation attracts attention from the community as it simplifies the traditional two-stage knowledge distillation process into a single stage. Online distillation collaboratively trains a group of peer models, which are treated as students, and all students gain extra knowledge from each o…

Cited by 10SourcePDFScholar
2021

A Model-Free Synchronous Control of Humanoid Robot Finger

ICRA 2021poster

For a multi-fingered robot hand, the individual control over single joints cannot guarantee their fine collaboration. For achieving a high-precision synchronization, a theory of synchronous control is introduced to multi-fingered robot hands. This paper introduced a new model-free and cross-coupling…

Cited by 2SourceScholar
2021

Automated Multi-Organ Segmentation in Pet Images Using Cascaded Training of a 3d U-Net and Convolutional Autoencoder

ICASSP 2021accepted

PET imaging is an important tool in clinical diagnostics, especially in oncology as it is able to visualize ongoing metabolic processes, e.g. caused by a tumor. Due to the low spatial resolution, a corresponding CT or MRI scan is normally necessary to gain knowledge about the physiological structure…

Cited by 0SourceScholar
2021

Diverse Complexity Measures for Dataset Curation in Self-Driving

IROS 2021poster

Modern self-driving systems heavily rely on deep learning. As a consequence, their performance is influenced significantly by the quality and richness of the training data. Data collection platforms can generate many hours of raw data on a daily basis, however, it is not feasible to label everything…

Cited by 16SourceScholar
2021

Multi-Class Uncertainty Calibration via Mutual Information Maximization-based Binning

ICLR 2021poster

Post-hoc multi-class calibration is a common approach for providing high-quality confidence estimates of deep neural network predictions. Recent work has shown that widely used scaling methods underestimate their calibration error, while alternative Histogram Binning (HB) methods often fail to prese…

2021

Perceive, Attend, and Drive: Learning Spatial Attention for Safe Self-Driving

ICRA 2021poster

In this paper, we propose an end-to-end self-driving network featuring a sparse attention module that learns to automatically attend to important regions of the input. The attention module specifically targets motion planning, whereas prior literature only applied attention in perception tasks. Lear…

Cited by 52SourceScholar
2021

Uncertainty-Based Biological Age Estimation of Brain MRI Scans

ICASSP 2021accepted

Age is an essential factor in modern diagnostic procedures. However, assessment of the true biological age (BA) remains a daunting task due to the lack of reference ground-truth labels. Current BA estimation approaches are either restricted to skeletal images or rely on non-imaging modalities that y…

Cited by 0SourceScholar
2021

Unsupervised Path Representation Learning with Curriculum Negative Sampling

IJCAI 2021poster

Path representations are critical in a variety of transportation applications, such as estimating path ranking in path recommendation systems and estimating path travel time in navigation systems. Existing studies often learn task-specific path representations in a supervised manner, which require a…

2020

DSDNet: Deep Structured self-Driving Network

ECCV 2020poster

In this paper, we propose the Deep Structured self-Driving Network (DSDNet), which performs object detection, motion prediction, and motion planning with a single neural network. Towards this goal, we develop a deep structured energy based model which considers the interactions between actors and pr…

Cited by 118SourcePDFScholar
2020

End-to-end Contextual Perception and Prediction with Interaction Transformer

IROS 2020poster

In this paper, we tackle the problem of detecting objects in 3D and forecasting their future motion in the context of self-driving. Towards this goal, we design a novel approach that explicitly takes into account the interactions between actors. To capture their spatial-temporal dependencies, we pro…

Cited by 148SourceScholar
2020

Learning Lane Graph Representations for Motion Forecasting

ECCV 2020poster

We propose a motion forecasting model that exploits a novel structured map representation as well as actor-map interactions. Instead of encoding vectorized maps as raster images, we construct a lane graph from raw map data to explicitly preserve the map structure. To capture the complex topology and…

2020

LiDARsim: Realistic LiDAR Simulation by Leveraging the Real World

CVPR 2020oral

We tackle the problem of producing realistic simulations of LiDAR point clouds, the sensor of preference for most self-driving vehicles. We argue that, by leveraging real data, we can simulate the complex world more realistically compared to employing virtual worlds built from CAD/procedural models.…

Cited by 265PDFScholar
2020

Physically Realizable Adversarial Examples for LiDAR Object Detection

CVPR 2020poster

Modern autonomous driving systems rely heavily on deep learning models to process point cloud sensory data; meanwhile, deep models have been shown to be susceptible to adversarial attacks with visually imperceptible perturbations. Despite the fact that this poses a security concern for the self-driv…

Cited by 287PDFScholar
2020

PnPNet: End-to-End Perception and Prediction With Tracking in the Loop

CVPR 2020poster

We tackle the problem of joint perception and motion forecasting in the context of self-driving vehicles. Towards this goal we propose PnPNet, an end-to-end model that takes as input sequential sensor data, and outputs at each time step object tracks and their future trajectories. The key component…

Cited by 221PDFScholar
2020

RadarNet: Exploiting Radar for Robust Perception of Dynamic Objects

ECCV 2020poster

We tackle the problem of exploiting Radar for perception in the context of self-driving as Radar provides complementary information to other sensors such as LiDAR or cameras in the form of Doppler velocity. The main challenges of using Radar are the noise and measurement ambiguities which have been…

Cited by 145SourcePDFScholar
2020

Recovering and Simulating Pedestrians in the Wild

CoRL 2020

Sensor simulation is a key component for testing the performance of self-driving vehicles and for data augmentation to better train perception systems. Typical approaches rely on artists to create both 3D assets and their animations to generate a new scenario. This, however, does not scale. In contr

Cited by 0SourcePDFScholar
2020

Supervised Canonical Correlation Analysis of Data on Symmetric Positive Definite Manifolds by Riemannian Dimensionality Reduction

ICASSP 2020accepted

Most computer vision problems entail data that reside on Riemannian manifolds. Canonical correlation analysis (CCA) is a powerful method that captures correlations between any two sets of matrices. In this paper, we propose a framework for a supervised CCA of manifold-based data. This framework aims…

Cited by 0SourceScholar
2020

Testing the Safety of Self-driving Vehicles by Simulating Perception and Prediction

ECCV 2020poster

We present a novel method for testing the safety of self-driving vehicles in simulation. We propose an alternative to sensor simulation, as sensor simulation is expensive and has large domain gaps. Instead, we directly simulate the outputs of the self-driving vehicle’s perception and prediction syst…

Cited by 29SourcePDFScholar
2020

V2VNet: Vehicle-to-Vehicle Communication for Joint Perception and Prediction

ECCV 2020poster

In this paper, we explore the use of vehicle-to-vehicle (V2V) communication to improve the perception and motion forecasting performance of self-driving vehicles. By intelligently aggregating the information received from multiple nearby vehicles, we can observe the same scene from different viewpoi…

2018

Automated Detection of High FDG Uptake Regions in CT Images

ICASSP 2018accepted

Combined PET-CT scan is an important diagnostic tool in modern medicine, e.g. for staging or treatment planning in the field of oncology. Especially in small structures, like a tumour, textural variations visible in a PET image are not visually recognizable within a CT scan from the same region. Thu…

Cited by 0SourceScholar
2018

Automatic Motion Artifact Detection for Whole-Body Magnetic Resonance Imaging

ICASSP 2018accepted

Magnetic resonance (MR) plays an important role in medical imaging. It can be flexibly tuned towards different applications for deriving a meaningful diagnosis. However, its long acquisition times and flexible parametrization make it on the other hand prone to artifacts which obscure the underlying…

Cited by 0SourceScholar
2018

Deep Continuous Fusion for Multi-Sensor 3D Object Detection

ECCV 2018poster

In this paper, we propose a novel 3D object detector that can exploit both LIDAR as well as cameras to perform very accurate localization. Towards this goal, we design an end-to-end learnable architecture that exploits continuous convolutions to fuse image and LIDAR feature maps at different levels…

Cited by 1168SourcePDFScholar
2018

Fast and Furious: Real Time End-to-End 3D Detection, Tracking and Motion Forecasting With a Single Convolutional Net

CVPR 2018poster

In this paper we propose a novel deep neural network that is able to jointly reason about 3D detection, tracking and motion forecasting given data captured by a 3D sensor. By jointly reasoning about these tasks, our holistic approach is more robust to occlusion as well as sparse data at range. O…

Cited by 834SourcePDFScholar
2018

Learning to Reweight Examples for Robust Deep Learning

ICML 2018oral

Deep neural networks have been shown to be very powerful modeling tools for many supervised learning tasks involving complex input patterns. However, they can also easily overfit to training set biases and label noises. In addition to various regularizers, example reweighting algorithms are popular…

2017

TorontoCity: Seeing the World With a Million Eyes

ICCV 2017spotlight

In this paper we introduce the TorontoCity benchmark, which covers the full greater Toronto area (GTA) with 712.5km2 of land, 8439km of road and around 400, 000 buildings. Our benchmark provides different perspectives of the world captured from airplanes, drones and cars driving around the city. Man…

Cited by 217PDFScholar
2017

Unsupervised image segmentation using convolutional autoencoder with total variation regularization as preprocessing

ICASSP 2017accepted

Conventional unsupervised image segmentation methods use color and geometric information and apply clustering algorithms over pixels. They preserve object boundaries well but often suffer from over-segmentation due to noise and artifacts in the images. In this paper, we contribute on a preprocessing…

Cited by 0SourceScholar
2016

Active learning for magnetic resonance image quality assessment

ICASSP 2016accepted

In medical imaging, the acquired images are usually analyzed by a human observer and rated with respect to a diagnostic question. However, this procedure is time-demanding and expensive. Further more, the lack of a reference image makes this task challenging. In order to support the human observer i…

Cited by 0SourceScholar
2015

Combining Compressed Sensing with motion correction in acquisition and reconstruction for PET/MR

ICASSP 2015accepted

In the field of oncology, simultaneous Positron-Emission-Tomography/Magnetic Resonance (PET/MR) scanners offer a great potential for improving diagnostic accuracy. However, to achieve a high Signal-to-Noise Ratio (SNR) for an accurate lesion detection and quantification in the PET/MR images, one has…

Cited by 0SourceScholar