← Search

Siqi Wang

26 accepted papers

2026

Alleviating Observation Bias via Causal-Invariant Meta-Learning for Unbalanced Incomplete Multi-view Clustering

ICML 2026poster

In incomplete multi-view clustering, unbalanced missingness is prevalent, where different views exhibit significantly varying missing rates, causing severe observation bias. This imbalance poses two core challenges: models develop serious learning biases by over-relying on low-missing-rate views whi…

Cited by 0SourceScholar
2026

Attention to Threat-Relevant Objects: Reasoning Detection in Autonomous Driving via Multimodal Large Language Models

AAAI 2026technical

Perceiving threats is an innate human instinct. During driving, humans naturally focus their attention on objects that pose real potential risks. Motivated by this observation, we shift the focus from traditional class-based detection to a novel task termed threat-oriented reasoning detection in aut

Cited by 0SourcePDFScholar
2026

CitySeeker: How Do VLMs Explore Embodied Urban Navigation with Implicit Human Needs?

ICLR 2026poster

Vision-Language Models (VLMs) have made significant progress in explicit instruction-based navigation; however, their ability to interpret implicit human needs (e.g., ''I am thirsty'') in dynamic urban environments remains underexplored. This paper introduces CitySeeker, a novel benchmark designed t…

Cited by 0SourcecodeScholar
2026

LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward

AAAI 2026technical

Navigation instruction generation for visually impaired (VI) individuals (NIG-VI) is critical yet relatively underexplored. This study focuses on generating precise, in-situ, step-by-step navigation instructions that are practically usable for VI users. Specifically, we propose LaF-GRPO (LLM-as-Foll

Cited by 0SourcePDFScholar
2026

MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement

AAAI 2026technical

Recent advances demonstrate that reinforcement learning with verifiable rewards (RLVR) significantly enhances the reasoning capabilities of large language models (LLMs). However, standard RLVR faces challenges with reward sparsity, where zero rewards from consistently incorrect candidate answers pro

Cited by 0SourcePDFScholar
2026

Noise-Aware Generalization: Robustness to In-Domain Noise and Out-of-Domain Generalization

ICLR 2026poster

Methods addressing Learning with Noisy Labels (LNL) and multi-source Domain Generalization (DG) use training techniques to improve downstream task performance in the presence of label noise or domain shifts, respectively. Prior work often explores these tasks in isolation, with only limited work t…

Cited by 0SourcecodeScholar
2026

RealRep: Generalized SDR-to-HDR Conversion via Attribute-Disentangled Representation Learning

AAAI 2026technical

High-Dynamic-Range Wide-Color-Gamut (HDR-WCG) technology is becoming increasingly widespread, driving a growing need for converting Standard Dynamic Range (SDR) content to HDR. Existing methods primarily rely on fixed tone mapping operators, which struggle to handle the diverse appearances and degra

Cited by 0SourcePDFScholar
2026

Scaling and Transferability of Annealing Strategies in Large Language Model Training

AAAI 2026technical

Learning rate scheduling is crucial for training large language models, yet understanding the optimal annealing strategies across different model configurations remains challenging. In this work, we investigate the transferability of annealing dynamics in large language model training and refine a g

Cited by 0SourcePDFScholar
2025

Achieving Speed-Accuracy Balance in Vision-based 3D Occupancy Prediction via Geometric-Semantic Disentanglement

AAAI 2025technical

Occupancy prediction plays a pivotal role in autonomous driving (AD) due to its capabilities of fine-grained 3D perception and general object recognition. However, existing methods often incur high computational costs, which conflict with AD's real-time demand. To this end, we redirect the focus fro…

2025

Adaptive Visual Servoing Control Barrier Function of Robotic Manipulators with Uncalibrated Camera

IROS 2025

This paper investigates the problem of safe visual servoing control of manipulators using an uncalibrated eye-in-hand camera based on control barrier functions (CBFs). Traditional CBFs are defined in the workspace, corresponding to the global coordinates of the base frame. However, when the camera’s

Cited by 0SourceScholar
2025

Exploring the Frontiers of Animation Video Generation in the Sora Era: Method, Dataset and Benchmark

IJCAI 2025

Animation has gained significant interest in the recent film and TV industry. Despite the success of advanced video generation models like Sora, Kling, and CogVideoX in generating natural videos, they lack the same effectiveness in handling animation videos. Evaluating animation video generation is

2025

MaxAuc: A Max-Plus-Based Auction Approach for Multi-Robot Allocations for Time-Ordered Temporal Logic Tasks

IROS 2025

In this paper, we investigate a multi-robot task allocation problem where a team of heterogeneous robots operates in a discrete workspace to achieve a set of tasks expressed by linear temporal logic formulas. In contrast to existing works, we further consider inter-task-time-order constraints, which

Cited by 0SourceScholar
2025

Revisiting Scaling Laws for Language Models: The Role of Data Quality and Training Strategies

ACL 2025long

Traditional scaling laws in natural language processing suggest that increasing model size and training data enhances performance. However, recent studies reveal deviations, particularly in large language models, where performance improvements decelerate—a phenomenon known as sub-scaling. This paper…

Cited by 0SourcePDFScholar
2024

CoSafe: Evaluating Large Language Model Safety in Multi-Turn Dialogue Coreference

EMNLP 2024main

As large language models (LLMs) constantly evolve, ensuring their safety remains a critical research issue. Previous red teaming approaches for LLM safety have primarily focused on single prompt attacks or goal hijacking. To the best of our knowledge, we are the first to study LLM safety in multi-tu…

2024

Diagnosis of Autism Spectrum Disorder Based on Contrastive Functional Connectivity Graph Learning Network

ICASSP 2024accepted

To reduce the dependence on tagged data, we proposed a Contrastive Functional Connectivity Graph Learning Network (CFCG-Net) for the diagnosis of autism spectrum disorder. CFCG-Net is mainly composed of three parts: construction of contrastive Functional Connection (FC) graphs, learning of contrasti…

Cited by 0SourceScholar
2024

Optimal Transport Guided Correlation Assignment for Multimodal Entity Linking

ACL 2024findings

Multimodal entity linking (MEL) aims to link ambiguous mentions in multimodal contexts to entities in a multimodal knowledge graph. A pivotal challenge is to fully leverage multi-element correlations between mentions and entities to bridge modality gap and enable fine-grained semantic matching. Exis…

2024

Scaling Laws Across Model Architectures: A Comparative Analysis of Dense and MoE Models in Large Language Models

EMNLP 2024main

The scaling of large language models (LLMs) is a critical research area for the efficiency and effectiveness of model training and deployment. Our work investigates the transferability and discrepancies of scaling laws between Dense Models and Mixture of Experts (MoE) models. Through a combination o…

Cited by 2SourcePDFScholar
2024

Synthesis of Temporally-Robust Policies for Signal Temporal Logic Tasks using Reinforcement Learning

ICRA 2024poster

This paper investigates the problem of designing control policies that satisfy high-level specifications described by signal temporal logic (STL) in unknown, stochastic environments. While many existing works concentrate on optimizing the spatial robustness of a system, our work takes a step further…

Cited by 5SourcecodeScholar
2023

CHAMMI: A benchmark for channel-adaptive models in microscopy imaging

NeurIPS 2023poster

Most neural networks assume that input images have a fixed number of channels (three for RGB images). However, there are many settings where the number of channels may vary, such as microscopy images where the number of channels changes depending on instruments and experimental goals. Yet, there has…

Cited by 11SourcePDFScholar
2022

Deep Anomaly Discovery From Unlabeled Videos via Normality Advantage and Self-Paced Refinement

CVPR 2022poster

While classic video anomaly detection (VAD) requires labeled normal videos for training, emerging unsupervised VAD (UVAD) aims to discover anomalies directly from fully unlabeled videos. However, existing UVAD methods still rely on shallow models to perform detection or initialization, and they are…

Cited by 51PDFcodeScholar
2022

LTMD: Learning Improvement of Spiking Neural Networks with Learnable Thresholding Neurons and Moderate Dropout

NeurIPS 2022accept

Spiking Neural Networks (SNNs) have shown substantial promise in processing spatio-temporal data, mimicking biological neuronal mechanisms, and saving computational power. However, most SNNs use fixed model regardless of their locations in the network. This limits SNNs’ capability of transmitting pr…

2022

Stgat-Mad : Spatial-Temporal Graph Attention Network For Multivariate Time Series Anomaly Detection

ICASSP 2022accepted

Anomaly detection in multivariate time series data is challenging due to complex temporal and feature correlations. This paper proposes a novel unsupervised multi-scale stacked spatial-temporal graph attention network for multivariate time series anomaly detection (STGAT-MAD). The core of our framew…

Cited by 0SourceScholar
2021

One-Pass Multi-View Clustering for Large-Scale Data

ICCV 2021poster

Existing non-negative matrix factorization based multi-view clustering algorithms compute multiple coefficient matrices respect to different data views, and learn a common consensus concurrently. The final partition is always obtained from the consensus with classical clustering techniques, such as…

Cited by 113PDFcodeScholar
2020

Self-Supervised Gait Encoding with Locality-Aware Attention for Person Re-Identification

IJCAI 2020poster

Gait-based person re-identification (Re-ID) is valuable for safety-critical applications, and using only 3D skeleton data to extract discriminative gait features for person Re-ID is an emerging open topic. Existing methods either adopt hand-crafted features or learn gait features by traditional supe…

2019

A bio-robotic remora disc with attachment and detachment capabilities for reversible underwater hitchhiking

ICRA 2019poster

Remoras employ their adhesive discs to rapidly attach to and detach from a wide range of marine surfaces. By analyzing high-speed images of remoras' (Echeneis naucrates) hitchhiking behavior, we describe the fish's detachment mechanism as a lip curling up to break the seal between the disc and subst…

Cited by 9SourceScholar
2019

Effective End-to-end Unsupervised Outlier Detection via Inlier Priority of Discriminative Network

NeurIPS 2019poster

Despite the wide success of deep neural networks (DNN), little progress has been made on end-to-end unsupervised outlier detection (UOD) from high dimensional data like raw images. In this paper, we propose a framework named E^3Outlier, which can perform UOD in a both effective and end-to-end manner…