← Search

Wenjing Yang

64 accepted papers

2026

Beyond Rational Illusion: Behaviorally Realistic Strategic Classification

ICML 2026poster

Strategic classification studies the interaction between decision models and agents who strategically manipulate their features for favorable outcomes. Existing SC frameworks typically rely on the idealized assumption that agents are strictly rational. However, evidence from behavioral economics and…

Cited by 0SourceScholar
2026

Causal Disentangled Anchor Learning for Scalable Fair Multi-view Clustering

ICML 2026poster

Existing fair multi-view clustering methods typically suffer from a severe trade-off between clustering utility and fairness, while incurring prohibitive quadratic complexity on large-scale datasets. To address these challenges, we propose Causal Disentangled Anchor Learning (CDAL), a novel framewor…

Cited by 0SourceScholar
2026

DR$^2$Seg: Decomposed Two-Stage Rollouts for Efficient Reasoning Segmentation in Multimodal Large Language Models

ICML 2026poster

Reasoning segmentation is an emerging vision-language task that requires reasoning over intricate text queries to precisely segment objects. However, existing methods typically suffer from overthinking, generating verbose reasoning chains that interfere with object localization in multimodal large l…

Cited by 0SourceScholar
2026

Detecting Unobserved Confounders: A Kernelized Regression Approach

AAAI 2026technical

Detecting unobserved confounders is crucial for reliable causal inference in observational studies. Existing methods require either linearity assumptions or multiple heterogeneous environments, limiting applicability to nonlinear single-environment settings. To bridge this gap, we propose Kernel Reg

Cited by 0SourcePDFScholar
2026

Do-Prompt: Causal Interventions Meet Variational Prompt Bottlenecks

ICML 2026poster

Multi-modal prompt learning is a parameter-efficient approach to adapt large vision--language models to downstream classification tasks. However, prompts can inadvertently evolve into a high-capacity pathway encoding environment-dependent spurious correlations that are only predictive in the source …

Cited by 0SourceScholar
2026

ECD: Evidence-guided Contrastive Decoding in Retrieval-Augmented Generation with Accurate Knowledge Reference Adjustment

AAAI 2026technical

Retrieval-Augmented Generation (RAG) enhances the quality of question answering by integrating external knowledge with internal knowledge. A robust RAG system needs to precisely regulate the dependence of the response on the two types of knowledge. The recently proposed context-aware contrastive dec

Cited by 0SourcePDFScholar
2026

Federated Multi-view Clustering for Remote Sensing Data

ICML 2026poster

The rapid expansion of remote sensing technology has generated massive amounts of unlabeled multi-view data distributed across different institutions. Analyzing this data presents significant challenges, as centralized processing incurs prohibitive communication costs and raises data privacy concern…

Cited by 0SourceScholar
2026

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning

CVPR 2026

When video reasoning requires external knowledge, many systems with large multimodal models (LMMs) adopt retrieval augmentation to supply the missing context. Appending textual or multi-clip evidence, however, forces heterogeneous signals into a single attention space. We observe diluted attention a

Cited by 0SourceScholar
2026

Hierarchical Anchor Graph Learning for Multi-View Clustering

ICML 2026poster

Multi-view clustering (MVC) is a fundamental task in heterogeneous data analysis, where anchor-based graph methods are widely adopted for their computational efficiency. However, existing approaches typically utilize static, single-layer anchors, failing to capture the multi-granularity nature of co…

Cited by 0SourceScholar
2026

Learning Kernelized Hypothesis for Hidden Confounder Detection

IJCAI 2026

Detecting hidden confounding is crucial for reliable causal analysis from observational data, directly determining which downstream causal inference method to be deployed. Inspired by the theory of higher-order regression, recent sample-efficient hypothesis testing strategies overcome the restrictiv

Cited by 0Scholar
2026

LogicSAGE: Neuro-Symbolic Reasoning with Socratic-Guided Enhancement

ICML 2026poster

Large Language Models (LLMs) often struggle with complex logical reasoning. Existing approaches typically rely on either purely neural reasoning in natural language or offloading to formal solvers via symbolic representations. However, both paradigms face significant limitations: while LLMs exhibit …

Cited by 0SourceScholar
2026

MLUBench: A Benchmark for Lifelong Unlearning Evaluation in MLLMs

ICML 2026poster

Multimodal large language models (MLLMs) are trained on massive multimodal data, making data unlearning increasingly important as data owners may request the removal of specific content. In practice, these requests often arrive sequentially over time, giving rise to the challenging problem of *MLLM …

Cited by 0SourceScholar
2026

Partial Fairness Awareness: Belief-Guided Strategic Mechanism for Strategic Agents

AAAI 2026technical

Strategic machine learning investigates scenarios where agents manipulate their features to receive favorable decisions from predictive models. To address fairness concerns intrinsic to strategic classification, recent work has introduced group-specific fairness constraints. However, current fairnes

Cited by 0SourcePDFScholar
2026

RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark

CVPR 2026

The integration of visual understanding and generation into unified multimodal models represents a significant stride toward general-purpose AI. However, a fundamental question remains unanswered by existing benchmarks: does this architectural unification actually enable synergetic interaction betwe

Cited by 0SourcecodeScholar
2026

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning

CVPR 2026

Video reasoning has advanced with large multimodal models (LMMs), yet their inference is often a single pass that returns an answer without verifying whether the reasoning is evidence-aligned. We introduce **Reinforce to Learn, Elect to Reason (RLER)**, a dual paradigm that decouples learning to pro

Cited by 0SourceScholar
2026

Transformers with Endogenous In-Context Learning: Bias Characterization and Mitigation

ICLR 2026poster

In-context learning (ICL) enables pre-trained transformers (TFs) to perform few-shot learning across diverse tasks, fostering growing research into its underlying mechanisms. However, existing studies typically assume a causally-sufficient regime, overlooking spurious correlations and prediction bia…

Cited by 0SourceScholar
2026

Unstitching the Chimera: Frame-Level Risk and Train-Free Mitigation for Video Hallucination

CVPR 2026

Hallucination limits the reliability of multimodal large language models (MLLMs), and it is particularly damaging in video where errors manifest as distorted narratives rather than single-frame mistakes. We introduce a frame-first study of **Chimera Hallucination**: model stitches visual segments th

Cited by 0SourceScholar
2026

Unveiling Prior-data Fitted Networks on Causal Effect Estimation: Pre-training or Finetuning?

ICML 2026poster

Amortized causal inference via Prior-data Fitted Networks (PFNs) has emerged as a promising paradigm, enabling zero-shot estimation of causal effects without the need for dataset-specific model tuning. However, the principled effectiveness of unified pre-training across general interventional regime…

Cited by 0SourceScholar
2026

When Tabular Foundation Models Meet Strategic Tabular Data: A Prior Alignment Approach

ICML 2026poster

Tabular foundation models via pretrained prior-data fitted networks (PFNs) achieve remarkable generalization performance on arbitrary testing tabular data, when sample distributions are independent of the deployed classifiers, i.e., a non-strategic regime. In a variety of real-world scenarios, howev…

Cited by 0SourceScholar
2025

Anchor-Prompt-based Segmentation and Embedding Model

ICASSP 2025accepted

Tackling multi-object tracking and segmentation (MOTS) can be attributed to a multi-task learning task, i.e., performing Segmentation and Identity Embedding jointly (SIEJ). Unfortunately, achieving optimal SIEJ is non-trivial, as it relies on different spatiotemporal features of objects. Besides, th…

Cited by 0SourceScholar
2025

Attribute-formed Class-specific Concept Space: Endowing Language Bottleneck Model with Better Interpretability and Scalability

CVPR 2025poster

Language Bottleneck Models (LBMs) are proposed to achieve interpretable image recognition by classifying images based on textual concept bottlenecks. However, current LBMs simply list all concepts together as the bottleneck layer, leading to the spurious cue inference problem and cannot generalized…

2025

Automated Exposure Mapping for Networked Interference

ICASSP 2025accepted

By characterizing interactions and influences across individuals, networked interference aims to estimate cross-individual treatment effects. For each individual, one of the central components of existing approaches is to manually design an exposure mapping from their neighboring covariates (includi…

Cited by 0SourceScholar
2025

BTPG: A Platform and Benchmark for Behavior Tree Planning in Everyday Service Robots

IJCAI 2025

Behavior Trees (BTs) are a widely used control architecture in robotics, renowned for their robustness and safety, which are especially crucial for everyday service robots. Recently, several methods have been proposed to automatically plan BTs to accomplish specific tasks. However, existing research

2025

Breaking the Gradient Barrier: Unveiling Large Language Models for Strategic Classification

NeurIPS 2025poster

Strategic classification (SC) explores how individuals or entities modify their features strategically to achieve favorable classification outcomes. However, existing SC methods, which are largely based on linear models or shallow neural networks, face significant limitations in terms of scalability…

Cited by 0SourceScholar
2025

ClusterUCB: Efficient Gradient-Based Data Selection for Targeted Fine-Tuning of LLMs

EMNLP 2025

Gradient-based data influence approximation has been leveraged to select useful data samples in the supervised fine-tuning of large language models. However, the computation of gradients throughout the fine-tuning process requires too many resources to be feasible in practice. In this paper, we prop

2025

Effective and Efficient Time-Varying Counterfactual Prediction with State-Space Models

ICLR 2025poster

Time-varying counterfactual prediction (TCP) from observational data supports the answer of when and how to assign multiple sequential treatments, yielding importance in various applications. Despite the progress achieved by recent advances, e.g., LSTM or Transformer based causal approaches, their c…

Cited by 0SourcePDFScholar
2025

Elastic Robust Unlearning of Specific Knowledge in Large Language Models

NeurIPS 2025poster

LLM unlearning aims to remove sensitive or harmful information within the model, thus reducing the potential risk of generating unexpected information. However, existing Preference Optimization (PO)-based unlearning methods suffer two limitations. First, their rigid reward setting limits the effect…

Cited by 0SourceScholar
2025

Enhancing Uncertainty Quantification in Large Language Models through Semantic Graph Density

UAI 2025

Large Language Models (LLMs) excel in language understanding but are susceptible to "confabulation," where they generate arbitrary, factually incorrect responses to uncertain questions. Detecting confabulation in question answering often relies on Uncertainty Quantification (UQ), which measures sema

Cited by 0SourcePDFScholar
2025

Environment Inference for Learning Generalizable Dynamical System

NeurIPS 2025spotlight

Data-driven methods offer efficient and robust solutions for analyzing complex dynamical systems but rely on the assumption of I.I.D. data, driving the development of generalization techniques for handling environmental differences. These techniques, however, are limited by their dependence on envir…

Cited by 0SourceScholar
2025

FutureNet-LoF: Joint Trajectory Prediction and Lane Occupancy Field Prediction with Future Context Encoding

ICRA 2025

Most prior motion prediction endeavors in autonomous driving have inadequately encoded future scenarios, leading to predictions that may fail to accurately capture the diverse movements of agents (e.g., vehicles or pedestrians). To address this, we propose FutureNet, which explicitly integrates init

Cited by 9SourceScholar
2025

GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution

NeurIPS 2025spotlight

Ultra-high-resolution (UHR) remote sensing (RS) imagery offers valuable data for Earth observation but pose challenges for existing multimodal foundation models due to two key bottlenecks: (1) limited availability of UHR training data, and (2) token explosion caused by the large image size. To addre…

Cited by 0SourcecodeScholar
2025

HBTP: Heuristic Behavior Tree Planning with Large Language Model Reasoning

ICRA 2025

Behavior Trees (BTs) are increasingly becoming a popular control structure in robotics due to their modularity, reactivity, and robustness. In terms of BT generation methods, BT planning shows promise for generating reliable BTs. However, the scalability of BT planning is often constrained by prolon

Cited by 6SourcecodeScholar
2025

Harnessing Massive Satellite Imagery with Efficient Masked Image Modeling

ICCV 2025poster

Masked Image Modeling (MIM) has become an essential method for building foundational visual models in remote sensing (RS). However, the limitations in size and diversity of existing RS datasets restrict the ability of MIM methods to learn generalizable representations. Additionally, conventional MIM…

2025

MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios

NeurIPS 2025poster

Multimodal Large Language Models (MLLMs) have achieved considerable accuracy in Optical Character Recognition (OCR) from static images. However, their efficacy in video OCR is significantly diminished due to factors such as motion blur, temporal variations, and visual effects inherent in video conte…

Cited by 0SourceScholar
2025

MRBTP: Efficient Multi-Robot Behavior Tree Planning and Collaboration

AAAI 2025technical

Multi-robot task planning and collaboration are critical challenges in robotics. While Behavior Trees (BTs) have been established as a popular control architecture and are plannable for a single robot, the development of effective multi-robot BT planning algorithms remains challenging due to the com…

2025

RoMA: Scaling up Mamba-based Foundation Models for Remote Sensing

NeurIPS 2025poster

Recent advances in self-supervised learning for Vision Transformers (ViTs) have fueled breakthroughs in remote sensing (RS) foundation models. However, the quadratic complexity of self-attention poses a significant barrier to scalability, particularly for large models and high-resolution images. Whi…

Cited by 0SourcecodeScholar
2025

Robust CLIP-Guided Deep Thinking: A Two-Stage Optimization Strategy for Enhancing Adversarial Robustness and Reliability in LVLMs

ICASSP 2025accepted

Large Vision-Language models (LVLMs) have demonstrated remarkable performance in a wide range of vision-language tasks as an efficient input/output system. However, the lack of adversarial robustness at the input side and the widespread hallucination phenomenon at the output side significantly under…

Cited by 0SourceScholar
2025

Robust Deterministic DOA Estimation Using α-divergence in Unknown Noise Fields with Sparse Sensor Arrays

ICASSP 2025accepted

In this paper, we address the problem of robust direction-of-arrival (DOA) estimation in unknown spatially cor-related noise fields using sensor arrays composed of subarrays in sparse configurations. In such arrays, the noise covariance matrix has a block-diagonal structure. The proposed robust DOA…

Cited by 0SourceScholar
2025

Transformer-Based Spatial-Temporal Counterfactual Outcomes Estimation

ICML 2025poster

The real world naturally has dimensions of time and space. Therefore, estimating the counterfactual outcomes with spatial-temporal attributes is a crucial problem. However, previous methods are based on classical statistical models, which still have limitations in performance and generalization. Thi…

2025

XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?

CVPR 2025highlight

The astonishing breakthrough of multimodal large language models (MLLMs) has necessitated new benchmarks to quantitatively assess their capabilities, reveal their limitations, and indicate future research directions. However, this is challenging in the context of remote sensing (RS), since the image…

2024

Coalition Formation Game Approach for Task Allocation in Heterogeneous Multi-Robot Systems under Resource Constraints

IROS 2024poster

This paper studies a case of the multi-robot task allocation (MRTA) problem, where each unmanned aerial vehicle (UAV) is endowed with multiple but limited resources. Completing each task necessitates UAVs to combine different resources through coalition formation, which will incur various costs incl…

Cited by 0SourceScholar
2024

Diversifying Cross-Domain Few-Shot Learning via Multimodal Image Editing

ICASSP 2024accepted

Standing out as one of the most widely used tools in Cross-Domain Few-Shot Learning (CDFSL), data augmentation forms the bedrock of numerous recent advancements. However, the current augmentations in CDFSL are limited in their ability to modify high-level semantic attributes, resulting in a lack of…

Cited by 0SourceScholar
2024

Integrating Intent Understanding and Optimal Behavior Planning for Behavior Tree Generation from Human Instructions

IJCAI 2024poster

Robots executing tasks following human instructions in domestic or industrial environments essentially require both adaptability and reliability. Behavior Tree (BT) emerges as an appropriate control architecture for these scenarios due to its modularity and reactivity. Existing BT generation methods…

2024

Scaling Few-Shot Learning for the Open World

AAAI 2024technical

Few-shot learning (FSL) aims to enable learning models with the ability to automatically adapt to novel (unseen) domains in open-world scenarios. Nonetheless, there exists a significant disparity between the vast number of new concepts encountered in the open world and the restricted available scale…

Cited by 4SourcePDFScholar
2024

Sequential Fusion Based Multi-Granularity Consistency for Space-Time Transformer Tracking

AAAI 2024technical

Regarded as a template-matching task for a long time, visual object tracking has witnessed significant progress in space-wise exploration. However, since tracking is performed on videos with substantial time-wise information, it is important to simultaneously mine the temporal contexts which have no…

Cited by 7SourcePDFScholar
2024

Task Allocation in Heterogeneous Multi-Robot Systems Based on Preference-Driven Hedonic Game

ICRA 2024poster

Multiple preferences between robots and tasks have been largely overlooked in previous research on Multi-Robot Task Allocation (MRTA) problems. In this paper, we propose a preference-driven approach based on hedonic game to address the task allocation problem of muti-robot systems in emergency rescu…

Cited by 1SourceScholar
2024

Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?

NeurIPS 2024poster

Causal reasoning capability is critical in advancing large language models (LLMs) towards artificial general intelligence (AGI). While versatile LLMs appear to have demonstrated capabilities in understanding contextual causality and providing responses that obey the laws of causality, it remains unc…

2023

Decomposition, Interaction, Reconstruction Meets Global Context Learning In Visual Tracking

ICASSP 2023accepted

Tensor decomposition and reconstruction attention is a promising global context learning approach because it can remain efficient while avoiding feature compression. To exploit its potential even further in visual tracking, we redesign a 3D tensor modeling paradigm, namely tensor Decomposition, Inte…

Cited by 0SourceScholar
2023

Domain Specified Optimization for Deployment Authorization

ICCV 2023poster

This paper explores Deployment Authorization (DPA) as a means of restricting the generalization capabilities of vision models on certain domains to protect intellectual property. Nevertheless, the current advancements in DPA are predominantly confined to fully supervised settings. Such settings requ…

Cited by 8PDFScholar
2023

Enhanced Dcf Tracker Regularized by Reliable Sample Construction

ICASSP 2023accepted

Discriminative correlation filter (DCF) is a highly efficient tracking technique using the circulant shifted samples of search images to update the template, so the reliability of input samples determines template quality. In this paper, we rethink the reliability problem of input samples in advance…

Cited by 0SourceScholar
2023

Evolving Physical Instinct for Morphology and Control Co-Adaption

IROS 2023poster

The capability of a robot to perform tasks depends not only on precise motion control, but also on a well-suited body morphology. Adapting both morphology and control of robots to improve their task performance has been a widely studied and long-standing issue. While the bio-inspired bi-level optimi…

Cited by 1SourceScholar
2023

GANet: Goal Area Network for Motion Forecasting

ICRA 2023poster

Predicting the future motion of road participants is crucial for autonomous driving but is extremely challenging due to staggering motion uncertainty. Recently, most motion forecasting methods resort to the goal-based strategy, i.e., predicting endpoints of motion trajectories as conditions to regre…

Cited by 89SourcecodeScholar
2023

MagicFusion: Boosting Text-to-Image Generation Performance by Fusing Diffusion Models

ICCV 2023poster

The advent of open-source AI communities has produced a cornucopia of powerful text-guided diffusion models that are trained on various datasets. While few explorations have been conducted on ensembling such models to combine their strengths. In this work, we propose a simple yet effective method ca…

Cited by 16PDFcodeScholar
2023

Memory-based Exploration-value Evaluation Model for Visual Navigation

ICRA 2023poster

We propose a hierarchical visual navigation solution, called Memory-based Exploration-value Evaluation Model (MEEM), to improve the agent's navigation performance. MEEM employs a hierarchical policy to tackle the challenge of sparse rewards, holds an episodic memory to store the historical informati…

Cited by 1SourceScholar
2023

Progressive Perception Learning for Distribution Modulation in Siamese Tracking

ICASSP 2023accepted

We explore an innovative view on distribution modulation to boost Siamese trackers. Specially, we observed two cases of possible distribution inconsistency in Siamese tracking: 1) Two branches with different sizes may be in different distribution ranges after a shared backbone (including BN layers).…

Cited by 0SourceScholar
2023

SODA: Robust Training of Test-Time Data Adaptors

NeurIPS 2023poster

Adapting models deployed to test distributions can mitigate the performance degradation caused by distribution shifts. However, privacy concerns may render model parameters inaccessible. One promising approach involves utilizing zeroth-order optimization (ZOO) to train a data adaptor to adapt the te…

2023

Task2Morph: Differentiable Task-Inspired Framework for Contact-Aware Robot Design

IROS 2023poster

Optimizing the morphologies and the controllers that adapt to various tasks is a critical issue in the field of robot design, aka. embodied intelligence. Previous works typically model it as a joint optimization problem and use search-based methods to find the optimal solution in the morphology spac…

Cited by 1SourceScholar
2022

Meta Discovery: Learning to Discover Novel Classes given Very Limited Data

ICLR 2022spotlight

In novel class discovery (NCD), we are given labeled data from seen classes and unlabeled data from unseen classes, and we train clustering models for the unseen classes. However, the implicit assumptions behind NCD are still unclear. In this paper, we demystify assumptions behind NCD and find that…

2021

BT Expansion: a Sound and Complete Algorithm for Behavior Planning of Intelligent Robots with Behavior Trees

AAAI 2021technical

Behavior Trees (BTs) have attracted much attention in the robotics field in recent years, which generalize existing control architectures and bring unique advantages for building robot systems. Automated synthesis of BTs can reduce human workload and build behavior models for complex tasks beyond th…

2021

Conditional Variational Capsule Network for Open Set Recognition

ICCV 2021poster

In open set recognition, a classifier has to detect unknown classes that are not known at training time. In order to recognize new categories, the classifier has to project the input samples of known classes in very compact and separated regions of the features space for discriminating samples of un…

Cited by 66PDFcodeScholar
2021

Dec-SGTS: Decentralized Sub-Goal Tree Search for Multi-Agent Coordination

AAAI 2021technical

Multi-agent coordination tends to benefit from efficient communication, where cooperation often happens based on exchanging information about what the agents intend to do, i.e. intention sharing. It becomes a key problem to model the intention by some proper abstraction. Currently, it is either too…

2021

TOHAN: A One-step Approach towards Few-shot Hypothesis Adaptation

NeurIPS 2021spotlight

In few-shot domain adaptation (FDA), classifiers for the target domain are trained with \emph{accessible} labeled data in the source domain (SD) and few labeled data in the target domain (TD). However, data usually contain private information in the current era, e.g., data distributed on personal ph…