← Search

Xin WANG

336 accepted papers

2026

A Causal Target for Learning to Defer Under Hidden Confounding

AAAI 2026technical

Learning decision policies from confounded observational data is a challenging task in causal inference, as unobserved confounders can lead to biased or suboptimal actions when relying solely on machine learning models. A synergistic approach is learning to defer, which decides when to act itself an

Cited by 0SourcePDFScholar
2026

Adaptive Mixture of Disentangled Experts for Dynamic Graphs under Distribution Shifts

ICLR 2026poster

Dynamic graph representation learning under distribution shifts has drawn an increasing amount of attention in the research community, given its wide applicability in real-world scenarios. Existing methods typically employ a fixed-architecture design to extract invariant patterns. However, there may…

Cited by 0SourceScholar
2026

AssetFormer: Modular 3D Assets Generation with Autoregressive Transformer

ICLR 2026poster

The digital industry demands high-quality, diverse modular 3D assets, especially for user-generated content (UGC). In this work, we introduce AssetFormer, an autoregressive Transformer-based model designed to generate modular 3D assets from textual descriptions. Our pilot study leverages real-world…

Cited by 0SourcecodeScholar
2026

Attention Hijacking: Backdooring Text Dataset Distillation via Semantic Anchors

ICML 2026poster

Dataset Distillation (DD) has emerged as a promising technique for compressing large-scale datasets into compact synthetic sets while preserving model performance. However, the security implications of this paradigm, particularly within the Transformer-based text classification domain, remain undere…

Cited by 0SourceScholar
2026

BA-GS: Bayesian Adaptive Gaussian Splatting for SFM-Free 3D Reconstruction

CVPR 2026

3D Gaussian Splatting (3DGS) has demonstrated exceptional performance in reconstruction and novel view synthesis tasks. However, its reliance on Structure-from-Motion preprocessing may lead to degraded performance under sparse-view scenarios. Recent works attempt to address this limitation by levera

Cited by 0SourceScholar
2026

Beyond Geometry: Artistic Disparity Synthesis for Immersive 2D-to-3D

CVPR 2026

Current 2D-to-3D conversion methods achieve geometric accuracy but are artistically deficient, failing to replicate the immersive and emotionally resonant experience of professional 3D cinema. This is because "geometric reconstruction" paradigms mistake deliberate artistic intent--such as strategic

Cited by 0SourceScholar
2026

Binary Message Passing for Generalizable Semi-Supervised Graph Anomaly Detection

AAAI 2026technical

Graph Neural Networks (GNNs) have achieved impressive performance in semi-supervised graph anomaly detection (GAD). While many GNN variants have been developed for this task, they largely focus on advanced message aggregation schemes, leaving the message routing aspect underexplored. We argue that t

Cited by 0SourcePDFScholar
2026

BuildingWorld: A Structured 3D Building Dataset for Urban Foundation Models

AAAI 2026technical

As digital twins become central to the transformation of modern cities, accurate and structured 3D building models emerge as a key enabler of high-fidelity, updatable urban representations. These models underpin diverse applications including energy modeling, urban planning, autonomous navigation, a

Cited by 0SourcePDFScholar
2026

CaliTex: Geometry-Calibrated Attention for View-Coherent 3D Texture Generation

CVPR 2026

Despite major advances brought by diffusion-based models, current 3D texture generation systems remain hindered by cross-view inconsistency -- textures that appear convincing from one viewpoint often fail to align across others. We find that this issue arises from attention ambiguity, where unstruct

Cited by 0SourceScholar
2026

CoSMo3D: Open-World Promptable 3D Semantic Segmentation through LLM-Guided Canonical Spatial Modeling

CVPR 2026

Open-world promptable 3D semantic segmentation remains brittle as semantics are inferred in the input sensor coordinates. Yet, humans, in contrast, interpret parts via functional roles in a canonical space -- wings extend laterally, handles protrude to the side, and legs support from below. Psychoph

Cited by 0SourcecodeScholar
2026

Cross-Scale Collaboration between LLMs and Lightweight Sequential Recommenders with Domain-Specific Latent Reasoning

AAAI 2026technical

Sequential recommendation aims to predict the next item based on historical interactions. To further enhance the reasoning capability in sequential recommendation, LLMs are employed to predict the next item or generate semantic IDs for item representation, given LLMs

Cited by 0SourcePDFScholar
2026

Demo2Tutorial: From Human Experience to Multimodal Software Tutorials

CVPR 2026

Human experience in digital environments offers a vast, underexplored resource of authentic, untrimmed interactions that contain rich procedural knowledge. We introduce Demo2Tutorial, a framework that transforms this experience captured via screen recordings and interaction logs into structured, mul

Cited by 0SourcecodeScholar
2026

Dual Mamba for Node-Specific Representation Learning: Tackling Over-Smoothing with Selective State Space Modeling

AAAI 2026technical

Over-smoothing remains a fundamental challenge in deep Graph Neural Networks (GNNs), where repeated message passing causes node representations to become indistinguishable. While existing solutions, such as residual connections and skip layers, alleviate this issue to some extent, they fail to expli

Cited by 0SourcePDFScholar
2026

Enabling Crab Driving for Rear Steering-Limited Vehicles via Coordinated Direct Yaw Moment Control and Steering Allocation

RA-L 2026

Crab driving has significant application potential in complex driving conditions such as maneuvering in confined spaces, high-speed lane changes, and emergency obstacle avoidance. However, its practical application is impeded by small rear-wheel steering angles in mass-production vehicles. To overco

Cited by 0SourceScholar
2026

FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning

ICLR 2026poster

Large language models (LLMs) increasingly rely on Chain-of-Thought (CoT) prompting to improve problem-solving and provide seemingly transparent explanations. However, growing evidence shows that CoT often fail to faithfully represent the underlying reasoning process, raising concerns about their rel…

Cited by 0SourcecodeScholar
2026

FakeWorld 1.0: An Omni modal Benchmark for Fake Media and Content

ICML 2026poster

The accelerating realism of AI-generated content has amplified the spread of deceptive information and eroded public trust. Prior works typically split the problem into two tracks, media authenticity, which concerns whether content is real or AI-generated, and content veracity, which concerns semant…

Cited by 0SourceScholar
2026

Fractal Camouflage: A Bio-Inspired Approach for Multi-Scale Adversarial Attacks in the Infrared Domain

CVPR 2026

Infrared pedestrian detection is crucial in safety-critical systems but remains vulnerable to adversarial attacks. Existing physical attacks often rely on fixed, static patterns. However, they often lack robustness across scales, as their hand-crafted or uniformly generated structures are fundamenta

Cited by 0SourceScholar
2026

Gaussian On-the-Fly Splatting: A Progressive Framework for Robust Near Real-Time 3DGS Optimization

RA-L 2026

3D Gaussian Splatting (3DGS) achieves high-fidelity rendering with real-time performance, but existing methods rely on offline training after full Structure-from-Motion (SfM) processing. In contrast, this work introduces Gaussian on-the-fly Splatting (abbreviated as On-the-Fly GS), a progressive fra

Cited by 4SourcecodeScholar
2026

HyperD: Hybrid Periodicity Decoupling Framework for Traffic Forecasting

AAAI 2026technical

Accurate traffic forecasting plays a vital role in intelligent transportation systems, enabling applications such as congestion control, route planning, and urban mobility optimization. However, traffic forecasting remains challenging due to two key factors: (1) complex spatial dependencies arising

Cited by 0SourcePDFScholar
2026

Inference Scaling Law for Retrieval Augmented Generation

AAAI 2026technical

Retrieval-augmented generation (RAG) has recently emerged as a powerful framework for knowledge-intensive natural language processing tasks, which leverages the strengths of both pre-trained language models and external knowledge. While significant progress has been made, the scaling behavior of the

Cited by 0SourcePDFScholar
2026

LUMIN: A Longitudinal Multi-modal Knowledge Decomposition Network for Predicting Breast Cancer Recurrence

AAAI 2026technical

Accurate prediction of breast cancer recurrence after treatment is essential for improving long-term outcomes. However, existing models are limited by three key challenges: (1) they typically rely on single-modal data, missing cross-modal interactions; (2) they analyze static snapshots, failing to c

Cited by 0SourcePDFScholar
2026

Learning Situated Awareness in the Real World

ICML 2026spotlight

A core aspect of human perception is *situated awareness*, the ability to relate ourselves to the surrounding physical environment and reason over possible actions in context. However, most existing benchmarks for multimodal foundation models (MFMs) emphasize **environment-centric** spatial relation…

Cited by 0SourceScholar
2026

LoRAGen: Structure-Aware Weight Space Learning for LoRA Generation

ICLR 2026poster

The widespread adoption of Low-Rank Adaptation (LoRA) for efficient fine-tuning of large language models has created demand for scalable parameter generation methods that can synthesize adaptation weights directly from task descriptions, avoiding costly task-specific training. We present LoRAGen, a…

Cited by 0SourcecodeScholar
2026

LumiTex: Towards High-Fidelity PBR Texture Generation with Illumination Context

ICLR 2026poster

Physically-based rendering (PBR) provides a principled standard for realistic material–lighting interactions in computer graphics. Despite recent advances in generating PBR textures, existing methods fail to address two fundamental challenges: 1) materials decomposition from image prompts under limi…

Cited by 0SourcecodeScholar
2026

Mechanistic Analysis of Cable Tension Effects on the Stiffness of Cable-Driven Serpentine Manipulators

RA-L 2026

This paper presents a mechanistic analysis of stiffness in cable-driven serpentine manipulators (CDSMs), incorporating both cable tension and cable stiffness. First, we derive an analytical stiffness model based on robot statics, identifying cable tension and stiffness as the dominant factors govern

Cited by 0SourceScholar
2026

Mechanistic Analysis of Cable Tension Effects on the Stiffness of Cable-Driven Serpentine Manipulators

ICRA 2026poster

This paper presents a mechanistic analysis of stiffness in cable-driven serpentine manipulators (CDSMs), incorporating both cable tension and cable stiffness. First, we derive an analytical stiffness model based on robot statics, identifying cable tension and stiffness as the dominant factors govern…

Cited by 0SourceScholar
2026

ModularAgent: A Task-Aware Modular Framework for Joint Optimization of Multimodal Large Language Models and World Models

CVPR 2026

Building generalist embodied agents requires a unified system that can interpret multimodal goals, model environment dynamics, and execute reliable actions across diverse real-world tasks. Multimodal large language models (MLLMs) offer strong semantic priors and cross-modal generalization, while wor

Cited by 0SourceScholar
2026

QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models

CVPR 2026

Vision-language-action (VLA) models unify perception, language, and control for embodied agents but face significant challenges in practical deployment due to rapidly increasing compute and memory demands, especially as models scale to longer horizons and larger backbones. To address these bottlenec

Cited by 0SourcecodeScholar
2026

RHYTHMBERT: A SELF-SUPERVISED LANGUAGE MODEL BASED ON LATENT REPRESENTATIONS OF ECG WAVEFORMS FOR HEART DISEASE DETECTION

ICASSP 2026oral

Electrocardiogram (ECG) analysis is crucial for diagnosing heart disease, but most self-supervised learning methods treat ECG as a generic time series, overlooking physiologic semantics and rhythm-level structure. Existing contrastive methods utilize augmentations that distort morphology, whereas ge…

Cited by 0SourcePDFScholar
2026

Reasoning Diffusion for Unpaired Test Time Out-of-distribution Text-Image to Video Generation

CVPR 2026

Text-image to video generation aims to synthesize a video conditioned on the given text-image inputs. Nevertheless, existing methods generally assume that the semantic information carried in the input text and image tends to be perfectly paired and temporally aligned, occurring simultaneously in the

Cited by 0SourceScholar
2026

RecEdit-Drive: 3D Reconstruction-Guided Spatiotemporal Video Editing for Autonomous Driving Scenes

CVPR 2026

High-quality video editing and processing are crucial in domains such as filmmaking and autonomous driving, where accurate visual refinement and data preparation are essential. However, it is challenging to achieve precise control over dynamic objects while maintaining spatiotemporal consistency. Cu

Cited by 0SourcecodeScholar
2026

SCOPE and SCION: Benchmark and Method for Ontology Induction and Fusion from Text

ICML 2026poster

Ontologies (schemas) are a key bottleneck for schema-grounded information extraction and knowledge graph construction, yet manual ontology engineering is expensive and schemas quickly fragment or drift across domains. We introduce SCOPE (Schema Construction and Ontology Induction Pipeline Evaluation…

Cited by 0SourceScholar
2026

SMART: A Surrogate Model for Predicting Application Runtime in Dragonfly Systems

AAAI 2026technical

The Dragonfly network, with its high-radix and low-diameter structure, is a leading interconnect in high-performance computing. A major challenge is workload interference on shared network links. Parallel discrete event simulation (PDES) is commonly used to analyze workload interference. However, h

Cited by 0SourcePDFScholar
2026

Scalable Semi-supervised Community Search via Graph Transformer on Attributed Heterogeneous Information Networks

AAAI 2026technical

Attributed heterogeneous information networks (AHINs) encode rich semantics through diverse node and edge types. Recent learning-based community search methods on AHINs have shown promising performance but face two major limitations: i) difficulty scaling to large graphs due to memory-intensive neig

Cited by 0SourcePDFScholar
2026

Selective Diffusion Distillation for Real-World High-Scale Image Super-Resolution

AAAI 2026technical

High-scale image super-resolution (SR) has become increasingly important with the rapid growth of mobile devices and high-resolution displays. However, current SR methods primarily focus on lower scales and generalize poorly to high-scale scenarios due to severe information loss and complex real-wor

Cited by 0SourcePDFScholar
2026

Structure Learning from Time-Series Data with Lag-Agnostic Structural Prior

ICLR 2026poster

Learning instantaneous and time-lagged causal relationships from time-series data is essential for uncovering fine-grained, temporally-aware interactions. Although this problem has been formulated as a continuous optimization task amenable to modern machine learning methods, existing approaches larg…

Cited by 0SourceScholar
2026

Target-Driven Policy Optimization for Sequential Counterfactual Outcome Control

ICML 2026poster

Identifying optimal intervention sequences from offline data to guide temporal systems toward target outcomes is a critical challenge with profound implications for fields like personalized medicine. While existing methods are mostly evaluated in offline settings, practical applications demand onlin…

Cited by 0SourceScholar
2026

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future

ICML 2026poster

Self-Rewarding Language Models propose an architecture in which the Large Language Models(LLMs) both generates responses and evaluates its own outputs via LLM-as-a-Judge prompting, dynamically improving its generative capabilities through iterative Direct Preference Optimization (DPO). However, our …

Cited by 0SourceScholar
2026

Temporal-aware Flow Matching for Video Generation with Temporally Coherent Motion

ICML 2026poster

Despite rapid advances in text-to-video generation, state-of-the-art generative models still suffer from producing temporally incoherent and unrealistic motion for videos. The key weakness of existing works is that they commonly treat videos as frame sequences and directly adopt Flow Matching object…

Cited by 0SourcecodeScholar
2026

Towards Context-Invariant Safety Alignment for Large Language Models

ICML 2026poster

Preference-based post-training aligns LLMs with human intent, yet safety behavior often remains brittle. A model may refuse a harmful request in a standard prompt but comply when the same intent is wrapped in adversarial wording. We suggest that robust safety requires context-invariant alignment, wh…

Cited by 0SourceScholar
2026

U2UData+: A Scalable Swarm UAVs Autonomous Flight Dataset for Embodied Long-horizon Tasks

AAAI 2026technical

Swarm UAV autonomous flight for Embodied Long-Horizon (ELH) tasks is crucial for advancing the low-altitude economy. However, existing methods focus only on specific basic tasks due to dataset limitations, failing in real-world deployment for ELH tasks. ELH tasks are not mere concatenations of basic

Cited by 0SourcePDFScholar
2026

Vulnerable Agent Identification in Large-Scale Multi-Agent Reinforcement Learning

ICML 2026poster

Partial agent failure becomes inevitable when systems scale up, making it crucial to identify the subset of agents whose failure causes worst-case system performance degradations. We study this Vulnerable Agent Identification (VAI) problem in large-scale multi-agent reinforcement learning (MARL). We…

Cited by 0SourceScholar
2026

rMMEA: Robust Multi-Modal Entity Alignment with Missing and Noise Visual Modality

AAAI 2026technical

Recently, multi-modal embedding methods have flourished in entity alignment. As state-of-the-art approaches evolve rapidly, visual modality (i.e., images) missing emerges as a critical challenge. While visual modality typically offers the most informative signals in multi-modal entity alignment (MME

Cited by 0SourcePDFScholar
2025

$\text{D}_{2}\text{O}$: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models

ICLR 2025poster

Efficient generative inference in Large Language Models (LLMs) is impeded by the growing memory demands of Key-Value (KV) cache, especially for longer sequences. Traditional KV Cache eviction strategies, which discard less critical KV-pairs based on attention scores, often degrade generation quality…

Cited by 0SourcePDFScholar
2025

3D-LMVIC: Learning-based Multi-View Image Compression with 3D Gaussian Geometric Priors

ICML 2025poster

Existing multi-view image compression methods often rely on 2D projection-based similarities between views to estimate disparities. While effective for small disparities, such as those in stereo images, these methods struggle with the more complex disparities encountered in wide-baseline multi-camer…

Cited by 0SourcePDFScholar
2025

A Diffusion Model over Directed Acyclic Graphs for Event Schema Generation

ICASSP 2025accepted

Event schema generation is crucial for understanding the structure and temporal relationships of complex events. In this paper, we introduce a novel Directed Acyclic Graph Diffusion Model (DAGDM) that integrates DAG characteristics within a diffusion framework to enhance the effectiveness of schema…

Cited by 0SourceScholar
2025

A Implies B: Circuit Analysis in LLMs for Propositional Logical Reasoning

NeurIPS 2025spotlight

Due to the size and complexity of modern large language models (LLMs), it has proven challenging to uncover the underlying mechanisms that models use to solve reasoning problems. For instance, is their reasoning for a specific problem localized to certain parts of the network? Do they break down the…

Cited by 0SourceScholar
2025

A Sequential Multi-Stage Approach for Code Vulnerability Detection via Confidence- and Collaboration-based Decision Making

EMNLP 2025

While large language models (LLMs) have shown strong capabilities across diverse domains, their application to code vulnerability detection holds great potential for identifying security flaws and improving software safety. In this paper, we propose a sequential multi-stage approach via confidence-

Cited by 0SourcePDFScholar
2025

AI-Face: A Million-Scale Demographically Annotated AI-Generated Face Dataset and Fairness Benchmark

CVPR 2025poster

AI-generated faces have enriched human life, such as entertainment, education, and art. However, they also pose misuse risks. Therefore, detecting AI-generated faces becomes crucial, yet current detectors show biased performance across different demographic groups. Mitigating biases can be done by d…

2025

Adaptive Dual Guidance Knowledge Distillation

AAAI 2025technical

Knowledge distillation (KD) aims to improve the performance of lightweight student networks under the guidance of pre-trained teachers. However, the large capacity gap between teachers and students limits the distillation gains. Previous methods addressing this problem have two weaknesses. First, mo…

Cited by 0SourcePDFScholar
2025

AutoGFM: Automated Graph Foundation Model with Adaptive Architecture Customization

ICML 2025oral

Graph foundation models (GFMs) aim to share graph knowledge across diverse domains and tasks to boost graph machine learning. However, existing GFMs rely on hand-designed and fixed graph neural network (GNN) architectures, failing to utilize optimal architectures *w.r.t.* specific domains and tasks…

Cited by 0SourcePDFScholar
2025

Behavior Importance-Aware Graph Neural Architecture Search for Cross-Domain Recommendation

AAAI 2025technical

Cross-domain recommendation (CDR) mitigates data sparsity and cold-start issues in recommendation systems. While recent CDR approaches using graph neural networks (GNNs) capture complex user-item interactions, they rely on manually designed architectures that are often suboptimal and labor-intensive…

2025

Beyond Single Images: Retrieval Self-Augmented Unsupervised Camouflaged Object Detection

ICCV 2025poster

At the core of Camouflaged Object Detection (COD) lies segmenting objects from their highly similar surroundings. Previous efforts navigate this challenge primarily through image-level modeling or annotation-based optimization. Despite advancing considerably, this commonplace practice hardly taps va…

2025

CMFNThinker: A Novel Cross-source Multi-modal Fake News Detection Model

ICASSP 2025accepted

The rapid development of social media platforms has accelerated the generation and spread of fake news. News on different platforms varies significantly in content and audience. It makes most existing fake news detection models, which rely on single-source datasets, struggle to perform well on news…

Cited by 0SourceScholar
2025

Complementary Advantages: Exploiting Cross-Field Frequency Correlation for NIR-Assisted Image Denoising

CVPR 2025poster

Existing single-image denoising algorithms often struggle to restore details when dealing with complex noisy images. The introduction of near-infrared (NIR) images offers new possibilities for RGB image denoising. However, due to the inconsistency between NIR and RGB images, the existing works still…

2025

Component-wise Self-Correction Network for Human Motion Prediction

ICASSP 2025accepted

Human motion prediction is a fundamental task in human-robot interaction and self-driving. Many existing human motion prediction methods use one encoder to embed the historical human poses and one decoder to predict future motion poses. We believe that it is possible to estimate the deviation of the…

Cited by 0SourceScholar
2025

ConDSeg: A General Medical Image Segmentation Framework via Contrast-Driven Feature Enhancement

AAAI 2025technical

Medical image segmentation plays an important role in clinical decision making, treatment planning, and disease tracking. However, it still faces two major challenges. On the one hand, there is often a "soft boundary" between foreground and background in medical images, with poor illumination and lo…

2025

Differentiable Structure Learning with Ancestral Constraints

ICML 2025poster

Differentiable structure learning of causal directed acyclic graphs (DAGs) is an emerging field in causal discovery, leveraging powerful neural learners. However, the incorporation of ancestral constraints, essential for representing abstract prior causal knowledge, remains an open research challeng…

Cited by 0SourcePDFScholar
2025

Disentangling Invariant Subgraph via Variance Contrastive Estimation under Distribution Shifts

ICML 2025poster

Graph neural networks (GNNs) have achieved remarkable success, yet most are developed under the in-distribution assumption and fail to generalize to out-of-distribution (OOD) environments. To tackle this problem, some graph invariant learning methods aim to learn invariant subgraph against distribut…

Cited by 0SourcePDFScholar
2025

Dyn-D^2P: Dynamic Differentially Private Decentralized Learning with Provable Utility Guarantee

IJCAI 2025

Most existing decentralized learning methods with differential privacy (DP) guarantee rely on constant gradient clipping bounds and fixed-level DP Gaussian noises for each node throughout the training process, leading to a significant accuracy degradation compared to non-private counterparts. In thi

Cited by 0SourcePDFScholar
2025

Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction Tuning

ICML 2025poster

Continual multimodal instruction tuning is crucial for adapting Multimodal Large Language Models (MLLMs) to evolving tasks. However, most existing methods adopt a fixed architecture, struggling with adapting to new tasks due to static model capacity. We propose to evolve the architecture under param…

Cited by 0SourcePDFScholar
2025

Enhancing Counterfactual Estimation: A Focus on Temporal Treatments

IJCAI 2025

In the medical field, treatment sequences significantly influence future outcomes through complex temporal interactions. Therefore, highlighting the role of temporal treatments within the model is crucial for accurate counterfactual estimation, which is often overlooked in current methods. To addres

2025

FIRE: Robust Detection of Diffusion-Generated Images via Frequency-Guided Reconstruction Error

CVPR 2025poster

The rapid advancement of diffusion models has significantly improved high-quality image generation, making generated content increasingly challenging to distinguish from real images and raising concerns about potential misuse. In this paper, we observe that diffusion models struggle to accurately re…

2025

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training

ICASSP 2025accepted

Large-scale Vision-Language Pre-training (VLP) has demonstrated remarkable success in the general domain. However, in the fashion domain, items are distinguished by fine-grained attributes such as texture and material, which are crucial for tasks such as retrieval. Existing models often fail to take…

Cited by 0SourceScholar
2025

FedSaaS: Class-Consistency Federated Semantic Segmentation via Global Prototype Supervision and Local Adversarial Harmonization

IJCAI 2025

Federated semantic segmentation enables pixel-level classification in images through collaborative learning while maintaining data privacy. However, existing research commonly overlooks the fine-grained class relationships within the semantic space when addressing heterogeneous problems, particularl

Cited by 0SourcePDFScholar
2025

From Abyssal Darkness to Blinding Glare: A Benchmark on Extreme Exposure Correction in Real World

ICCV 2025poster

Exposure correction aims to restore over/under-exposed images to well-exposed ones using a single network. However, existing methods mainly handle non-extreme exposure conditions and struggle with the severe luminance and texture loss caused by extreme exposure. Through a thorough investigation, we…

2025

Fusion meets Function: The Adaptive Selection-Generation Approach in Event Argument Extraction

COLING 2025main

Event Argument Extraction is a critical task of Event Extraction, focused on identifying event arguments within text. This paper presents a novel Fusion Selection-Generation-Based Approach, by combining the precision of selective methods with the semantic generation capability of generative methods…

2025

GBA-Net: A Method for 3D Brain Tumor Segmentation Based on Multi-scale Gaussian Boundary Attention

ICASSP 2025accepted

The complex nature of brain tumors, characterized by their individual shapes, sizes, and locations, as well as the presence of indistinct boundaries, presents a challenging task for precise automatic segmentation. While U-Net has been a top performer in medical image segmentation, it struggles with…

Cited by 1SourceScholar
2025

GTR-Loc: Geospatial Text Regularization Assisted Outdoor LiDAR Localization

NeurIPS 2025poster

Prevailing scene coordinate regression methods for LiDAR localization suffer from localization ambiguities, as distinct locations can exhibit similar geometric signatures — a challenge that current geometry-based regression approaches have yet to solve. Recent vision–language models show that textua…

Cited by 0SourcecodeScholar
2025

Generation-Augmented and Embedding Fusion in Document-Level Event Argument Extraction

COLING 2025main

Document-level event argument extraction is a crucial task that aims to extract arguments from the entire document, beyond sentence-level analysis. Prior classification-based models still fail to explicitly capture significant relationships and heavily relies on large-scale datasets. In this study,…

Cited by 0SourcePDFScholar
2025

GraphChain: Large Language Models for Large-scale Graph Analysis via Tool Chaining

NeurIPS 2025poster

Large Language Models (LLMs) face significant limitations when applied to large-scale graphs, struggling with context constraints and inflexible reasoning. We introduce GraphChain, a novel framework enabling LLMs to analyze large graphs by orchestrating dynamic sequences of specialized tools, mimick…

Cited by 0SourceScholar
2025

Implicit degree bias in the link prediction task

ICML 2025poster

Link prediction---a task of distinguishing actual hidden edges from random unconnected node pairs---is one of the quintessential tasks in graph machine learning. Despite being widely accepted as a universal benchmark and a downstream task for representation learning, the link prediction benchmark's…

2025

Improving Generalization for AI-Synthesized Voice Detection

AAAI 2025technical

AI-synthesized voice technology has the potential to create realistic human voices for beneficial applications, but it can also be misused for malicious purposes. While existing AI-synthesized voice detection models excel in intra-domain evaluation, they face challenges in generalizing across differ…

2025

JAQ: Joint Efficient Architecture Design and Low-Bit Quantization with Hardware-Software Co-Exploration

AAAI 2025technical

The co-design of neural network architectures, quantization precisions, and hardware accelerators offers a promising approach to achieving an optimal balance between performance and efficiency, particularly for model deployment on resource-constrained edge devices. In this work, we propose the JAQ F…

Cited by 0SourcePDFScholar
2025

Latte: Transfering LLMs' Latent-level Knowledge for Few-shot Tabular Learning

IJCAI 2025

Few-shot tabular learning, in which machine learning models are trained with a limited amount of labeled data, provides a cost-effective approach to addressing real-world challenges. The advent of Large Language Models (LLMs) has sparked interest in leveraging their pre-trained knowledge for few-sho

2025

MADiff: Text-Guided Fashion Image Editing with Mask Prediction and Attention-Enhanced Diffusion

ICASSP 2025accepted

Text-guided image editing model has achieved great success in general domain. However, directly applying these models to the fashion domain may encounter two issues: (1) Inaccurate localization of editing region; (2) Weak editing magnitude. To address these issues, the MADiff model is proposed. Spec…

Cited by 0SourceScholar
2025

MAR-3D: Progressive Masked Auto-regressor for High-Resolution 3D Generation

CVPR 2025highlight

Recent advances in auto-regressive transformers have revolutionized generative modeling across different domains, from language processing to visual generation, demonstrating remarkable capabilities. However, applying these advances to 3D generation presents three key challenges: the unordered natur…

Cited by 1SourcePDFScholar
2025

MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference

NAACL 2025long

Long-context Multimodal Large Language Models (MLLMs) that incorporate long text-image and text-video modalities, demand substantial computational resources as their multimodal Key-Value (KV) cache grows with increasing input lengths, challenging memory and time efficiency. For multimodal scenarios,…

2025

MEIT: Multimodal Electrocardiogram Instruction Tuning on Large Language Models for Report Generation

ACL 2025finding

Electrocardiogram (ECG) is the primary non-invasive diagnostic tool for monitoring cardiac conditions and is crucial in assisting clinicians. Recent studies have concentrated on classifying cardiac conditions using ECG data but have overlooked ECG report generation, which is time-consuming and requi…

2025

MERIT: Multi-Agent Collaboration for Unsupervised Time Series Representation Learning

ACL 2025finding

This paper studies the problem of unsupervised time series representation learning, which aims to map unlabeled time series data into a low-dimensional latent space for various downstream tasks. Previous works usually combine a range of augmentation strategies with contrastive learning to generate d…

Cited by 0SourcePDFScholar
2025

Mamba-Based Graph Convolutional Networks: Tackling Over-smoothing with Selective State Space

IJCAI 2025

Graph Neural Networks (GNNs) have shown great success in various graph-based learning tasks. However, it often faces the issue of over-smoothing as the model depth increases, which causes all node representations to converge to a single value and become indistinguishable. This issue stems from the i

2025

MoME: Mixture of Multi-Domain Experts for Multivariate Long-Term Series Forecasting

ICASSP 2025accepted

Time series forecasting is always important, with multivariate long-term series forecasting being its most challenging task. Here, the existing methods typically learn only in a single domain and focus on optimizing model structures, leading to incomplete information mining and imprecise predictions…

Cited by 0SourceScholar
2025

Modular-Cam: Modular Dynamic Camera-view Video Generation with LLM

AAAI 2025technical

Text-to-Video generation, which utilizes the provided text prompt to generate high-quality videos, has drawn increasing attention and achieved great success due to the development of diffusion models recently. Existing methods mainly rely on a pre-trained text encoder to capture the semantic informa…

2025

Modularized Self-Reflected Video Reasoner for Multimodal LLM with Application to Video Question Answering

ICML 2025poster

Multimodal Large Language Models (Multimodal LLMs) have shown their strength in Video Question Answering (VideoQA). However, due to the black-box nature of end-to-end training strategies, existing approaches based on Multimodal LLMs suffer from the lack of interpretability for VideoQA: they can neit…

Cited by 0SourcePDFScholar
2025

MuRL-DTI: A Multimodal Feature Fusion Reinforcement Learning Approach for Cold Start in Drug-Target Interactions

ICASSP 2025accepted

Drug Target Interaction (DTI) focuses on exploring the interactions between specific drug molecules and their biological targets to assess the efficacy and safety of drugs. Significant advancements have been made in integrating computational techniques compared to traditional approaches, including m…

Cited by 0SourceScholar
2025

NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls

EMNLP 2025

The resurgence of autonomous agents built using large language models (LLMs) to solve complex real-world tasks has brought increased focus on LLMs’ fundamental ability of tool or function calling. At the core of these agents, an LLM must plan, execute, and respond using external tools, APIs, and cus

2025

NoiseController: Towards Consistent Multi-view Video Generation via Noise Decomposition and Collaboration

ICCV 2025poster

High-quality video generation is crucial for many fields, including the film industry and autonomous driving. However, generating videos with spatiotemporal consistencies remains challenging. Current methods typically utilize attention mechanisms or modify noise to achieve consistent videos, neglect…

2025

Optimization Inspired Few-Shot Adaptation for Large Language Models

NeurIPS 2025spotlight

Large Language Models (LLMs) have demonstrated remarkable performance in real-world applications. However, adapting LLMs to novel tasks via fine-tuning often requires substantial training data and computational resources that are impractical in few-shot scenarios. Existing approaches, such as In-con…

Cited by 0SourceScholar
2025

Out-of-Distribution Generalized Graph Anomaly Detection with Homophily-aware Environment Mixup

NeurIPS 2025poster

Graph anomaly detection (GAD) is widely prevalent in scenarios such as financial fraud detection, anti-money laundering, and social bot detection. However, structural distribution shifts are commonly observed in real-world GAD data due to selection bias, resulting in reduced homophily. Existing GAD…

Cited by 0SourceScholar
2025

Pattern-Guided Adaptive Prior for Structure Learning

NeurIPS 2025poster

Learning the causality between variables, known as DAG structure learning, is critical yet challenging due to issues such as insufficient data and noise. While prior knowledge can improve the learning process and refine the DAG structure, incorporating prior knowledge is not without pitfalls. In par…

Cited by 0SourceScholar
2025

Preserving AUC Fairness in Learning with Noisy Protected Groups

ICML 2025poster

The Area Under the ROC Curve (AUC) is a key metric for classification, especially under class imbalance, with growing research focus on optimizing AUC over accuracy in applications like medical image analysis and deepfake detection. This leads to fairness in AUC optimization becoming crucial as bias…

2025

Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers

CVPR 2025poster

Recent advancements in diffusion models, particularly the architectural transformation from UNet-based models to Diffusion Transformers (DiTs), significantly improve the quality and scalability of image and video generation. However, despite their impressive capabilities, the substantial computation…

2025

RLMiniStyler: Light-weight RL Style Agent for Arbitrary Sequential Neural Style Generation

IJCAI 2025

Arbitrary style transfer aims to apply the style of any given artistic image to another content image. Still, existing deep learning-based methods often require significant computational costs to generate diverse stylized results. Motivated by this, we propose a novel reinforcement learning-based fr

2025

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

ICML 2025spotlight

Post-training quantization (PTQ) of large language models (LLMs) holds the promise in reducing the prohibitive computational cost at inference time. Quantization of all weight, activation and key-value (KV) cache tensors to 4-bit without significantly degrading generalizability is challenging, due t…

2025

Restricted Global-Aware Graph Filters Bridging GNNs and Transformer for Node Classification

NeurIPS 2025poster

Transformers have been widely regarded as a promising direction for breaking through the performance bottlenecks of Graph Neural Networks (GNNs), primarily due to their global receptive fields. However, a recent empirical study suggests that tuned classical GNNs can match or even outperform state-of…

Cited by 0SourceScholar
2025

SCALM: Detecting Bad Practices in Smart Contracts Through LLMs

AAAI 2025technical

As the Ethereum platform continues to mature and gain widespread usage, it is crucial to maintain high standards of smart contract writing practices. While bad practices in smart contracts may not directly lead to security issues, they do elevate the risk of encountering problems. Therefore, to unde…

2025

SVD-LLM V2: Optimizing Singular Value Truncation for Large Language Model Compression

NAACL 2025long

Despite significant advancements, the practical deployment of Large Language Models (LLMs) is often hampered by their immense sizes, highlighting the need for effective compression techniques. Singular Value Decomposition (SVD) emerges as a promising method for compressing LLMs. However, existing SV…

2025

SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression

ICLR 2025poster

The advancements in Large Language Models (LLMs) have been hindered by their substantial sizes, which necessitates LLM compression methods for practical deployment. Singular Value Decomposition (SVD) offers a promising solution for LLM compression. However, state-of-the-art SVD-based LLM compression…

2025

SafeVid: Toward Safety Aligned Video Large Multimodal Models

NeurIPS 2025poster

As Video Large Multimodal Models (VLMMs) rapidly advance, their inherent complexity introduces significant safety challenges, particularly the issue of mismatched generalization where static safety alignments fail to transfer to dynamic video contexts. We introduce SafeVid, a framework designed to…

Cited by 0SourceScholar
2025

Self-supervised Masked Graph Autoencoder via Structure-aware Curriculum

ICML 2025spotlight

Self-supervised learning (SSL) on graph-structured data has attracted considerable attention recently. Masked graph autoencoder, as one promising generative graph SSL approach that aims to recover masked parts of the input graph data, has shown great success in various downstream graph tasks. Howeve…

Cited by 0SourcePDFScholar
2025

Shifting Spotlight for Co-supervision: A Simple yet Efficient Single-branch Network to See Through Camouflage

ICASSP 2025accepted

Camouflaged object detection (COD) remains a challenging task in computer vision. Existing methods often resort to additional branches for edge supervision, incurring substantial computational costs. To address this, we propose the Co-Supervised Spotlight Shifting Network (CS<sup xmlns:mml="http://w…

Cited by 0SourceScholar
2025

TAPT: Test-Time Adversarial Prompt Tuning for Robust Inference in Vision-Language Models

CVPR 2025poster

Large pre-trained Vision-Language Models (VLMs) such as CLIP have demonstrated excellent zero-shot generalizability across various downstream tasks. However, recent studies have shown that the inference performance of CLIP can be greatly degraded by small adversarial perturbations, especially its vi…

2025

Towards An Integrated Approach for Expressive Piano Performance Synthesis from Music Scores

ICASSP 2025accepted

This paper presents an integrated system that transforms symbolic music scores into expressive piano performance audio. By combining a Transformer-based Expressive Performance Rendering (EPR) model with a fine-tuned neural MIDI synthesiser, our approach directly generates expressive audio performanc…

Cited by 0SourceScholar
2025

Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards

CVPR 2025poster

Diffusion models have achieved remarkable success in text-to-image generation. However, their practical applications are hindered by the misalignment between generated images and corresponding text prompts. To tackle this issue, reinforcement learning (RL) has been considered for diffusion model fin…

2025

Trajectory-LLM: A Language-based Data Generator for Trajectory Prediction in Autonomous Driving

ICLR 2025poster

Vehicle trajectory prediction is a crucial aspect of autonomous driving, which requires extensive trajectory data to train prediction models to understand the complex, varied, and unpredictable patterns of vehicular interactions. However, acquiring real-world data is expensive, so we advocate using…

2025

Understanding the Information Propagation Effects of Communication Topologies in LLM-based Multi-Agent Systems

EMNLP 2025

The communication topology in large language model-based multi-agent systems fundamentally governs inter-agent collaboration patterns, critically shaping both the efficiency and effectiveness of collective decision-making. While recent studies for communication topology automated design tend to cons

2025

Unifying Unsupervised Graph-Level Anomaly Detection and Out-of-Distribution Detection: A Benchmark

ICLR 2025poster

To build safe and reliable graph machine learning systems, unsupervised graph-level anomaly detection (GLAD) and unsupervised graph-level out-of-distribution (OOD) detection (GLOD) have received significant attention in recent years. Though these two lines of research share the same objective, they…

2025

Universal Online Temporal Calibration for Optimization-Based Visual-Inertial Navigation Systems

ICRA 2025

6-Degree of Freedom (6DoF) motion estimation with a combination of visual and inertial sensors is a growing area with numerous real-world applications. However, precise calibration of the time offset between these two sensor types is a prerequisite for accurate and robust tracking. To address this,

Cited by 0SourcecodeScholar
2025

Variational Counterfactual Intervention Planning to Achieve Target Outcomes

ICML 2025poster

A key challenge in personalized healthcare is identifying optimal intervention sequences to guide temporal systems toward target outcomes, a novel problem we formalize as counterfactual target achievement. In addressing this problem, directly adopting counterfactual estimation methods face compoundi…

Cited by 0SourcePDFScholar
2024

A Dual-module Framework for Counterfactual Estimation over Time

ICML 2024poster

Efficiently and effectively estimating counterfactuals over time is crucial for optimizing treatment strategies. We present the Adversarial Counterfactual Temporal Inference Network (ACTIN), a novel framework with dual modules to enhance counterfactual estimation. The balancing module employs a dist…

Cited by 2SourcePDFScholar
2024

A Novel Friction Measuring Method and Its Application to Improve the Static Modeling Accuracy of Cable-Driven Continuum Manipulators

RA-L 2024

Cable-driven continuum manipulators exhibit high flexibility and dexterity, leading to their increased popularity in recent years. Friction analysis is a crucial problem for these manipulators. Previous research has introduced friction models that are applicable to dynamic states where the direction

Cited by 6SourceScholar
2024

Adaptive Spatial-Temporal Hypergraph Fusion Learning for Next POI Recommendation

ICASSP 2024accepted

Next point-of-interest (POI) recommendation has been a trending task to provide next POI suggestions. Most existing sequential-based and graph-based methods have endeavored to model user visiting behaviors and achieved considerable performances. However, they have either modeled user interests at a…

Cited by 0SourceScholar
2024

Adversarial Prompt Tuning for Vision-Language Models

ECCV 2024poster

"With the rapid advancement of multimodal learning, pre-trained Vision-Language Models (VLMs) such as CLIP have demonstrated remarkable capacities in bridging the gap between visual and language modalities. However, these models remain vulnerable to adversarial attacks, particularly in the image mod…

2024

Can Large-Scale Vocoded Spoofed Data Improve Speech Spoofing Countermeasure with a Self-Supervised Front End?

ICASSP 2024accepted

A speech spoofing countermeasure (CM) that discriminates between unseen spoofed and bona fide data requires diverse training data. While many datasets use spoofed data generated by speech synthesis systems, it was recently found that data vocoded by neural vocoders were also effective as the spoofed…

Cited by 0SourceScholar
2024

Causal language modeling can elicit search and reasoning capabilities on logic puzzles

NeurIPS 2024poster

Causal language modeling using the Transformer architecture has yielded remarkable capabilities in Large Language Models (LLMs) over the last few years. However, the extent to which fundamental search and reasoning capabilities emerged within LLMs remains a topic of ongoing debate. In this work, we…

2024

ComCLIP: Training-Free Compositional Image and Text Matching

NAACL 2024long

Contrastive Language-Image Pretraining (CLIP) has demonstrated great zero-shot performance for matching images and text. However, it is still challenging to adapt vision-language pretrained models like CLIP to compositional image and text matching — a more challenging image and text matching task re…

2024

Continuous Optical Zooming: A Benchmark for Arbitrary-Scale Image Super-Resolution in Real World

CVPR 2024poster

Most current arbitrary-scale image super-resolution (SR) methods has commonly relied on simulated data generated by simple synthetic degradation models (e.g. bicubic downsampling) at continuous various scales thereby falling short in capturing the complex degradation of real-world images. This limit…

2024

CurBench: Curriculum Learning Benchmark

ICML 2024poster

Curriculum learning is a training paradigm where machine learning models are trained in a meaningful order, inspired by the way humans learn curricula. Due to its capability to improve model generalization and convergence, curriculum learning has gained considerable attention and has been widely app…

2024

Data-Augmented Curriculum Graph Neural Architecture Search under Distribution Shifts

AAAI 2024technical

Graph neural architecture search (NAS) has achieved great success in designing architectures for graph data processing.However, distribution shifts pose great challenges for graph NAS, since the optimal searched architectures for the training graph data may fail to generalize to the unseen test grap…

Cited by 9SourcePDFScholar
2024

Data-Centric Explainable Debiasing for Improving Fairness in Pre-trained Language Models

ACL 2024findings

Human-like social bias of pre-trained language models (PLMs) on downstream tasks have attracted increasing attention. The potential flaws in the training data are the main factor that causes unfairness in PLMs. Existing data-centric debiasing strategies mainly leverage explicit bias words (defined a…

2024

Differentiable Structure Learning with Partial Orders

NeurIPS 2024poster

Differentiable structure learning is a novel line of causal discovery research that transforms the combinatorial optimization of structural models into a continuous optimization problem. However, the field has lacked feasible methods to integrate partial order constraints, a critical prior informati…

Cited by 1SourcePDFScholar
2024

DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation

ICLR 2024poster

Subject-driven text-to-image generation aims to generate customized images of the given subject based on the text descriptions, which has drawn increasing attention. Existing methods mainly resort to finetuning a pretrained generative model, where the identity-relevant information (e.g., the boy) an…

2024

Disentangled Continual Graph Neural Architecture Search with Invariant Modular Supernet

ICML 2024poster

The existing graph neural architecture search (GNAS) methods assume that the graph tasks are static during the search process, ignoring the ubiquitous scenarios where sequential graph tasks come in a continual fashion. Moreover, existing GNAS works resort to entangled graph factors during the archit…

Cited by 10SourcePDFScholar
2024

Disentangled Graph Self-supervised Learning for Out-of-Distribution Generalization

ICML 2024poster

Graph out-of-distribution (OOD) generalization, aiming to generalize graph neural networks (GNNs) under distribution shifts between training and testing environments, has attracted ever-increasing attention recently. However, existing literature heavily relies on sufficient task-dependent graph labe…

Cited by 11SourcePDFScholar
2024

Efficient Sharpness-Aware Minimization for Molecular Graph Transformer Models

ICLR 2024poster

Sharpness-aware minimization (SAM) has received increasing attention in computer vision since it can effectively eliminate the sharp local minima from the training trajectory and mitigate generalization degradation. However, SAM requires two sequential gradient computations during the optimization o…

2024

Enhancing Video Super-Resolution via Implicit Resampling-based Alignment

CVPR 2024highlight

In video super-resolution it is common to use a frame-wise alignment to support the propagation of information over time. The role of alignment is well-studied for low-level enhancement in video but existing works overlook a critical step -- resampling. We show through extensive experiments that for…

Cited by 15SourcePDFScholar
2024

Exponential Hardness of Optimization from the Locality in Quantum Neural Networks

AAAI 2024technical

Quantum neural networks (QNNs) have become a leading paradigm for establishing near-term quantum applications in recent years. The trainability issue of QNNs has garnered extensive attention, spurring demand for a comprehensive analysis of QNNs in order to identify viable solutions. In this work, we…

Cited by 4SourcePDFScholar
2024

FUG: Feature-Universal Graph Contrastive Pre-training for Graphs with Diverse Node Features

NeurIPS 2024poster

Graph Neural Networks (GNNs), known for their effective graph encoding, are extensively used across various fields. Graph self-supervised pre-training, which trains GNN encoders without manual labels to generate high-quality graph representations, has garnered widespread attention. However, due to t…

2024

From Text Segmentation to Enhanced Representation Learning: A Novel Approach to Multi-Label Classification for Long Texts

EMNLP 2024finding

Multi-label text classification (MLTC) is an important task in the field of natural language processing. Most existing models rely on high-quality text representations provided by pre-trained language models (PLMs). They hence face the challenge of input length limitation caused by PLMs, when dealin…

2024

Gorilla: Large Language Model Connected with Massive APIs

NeurIPS 2024poster

Large Language Models (LLMs) have seen an impressive wave of advances, with models now excelling in a variety of tasks, such as mathematical reasoning and program synthesis. However, their potential to effectively use tools via API calls remains unfulfilled. This is a challenging task even for today…

2024

Heuristic-Driven, Type-Specific Embedding in Parallel Spaces for Enhancing Knowledge Graph Reasoning

ICASSP 2024accepted

Knowledge Graph Reasoning aims to derive new insights from existing Knowledge Graphs (KGs) and address any missing or incomplete data. Existing models primarily rely on explicit information while neglecting the implicit constraints imposed by entity types on relations types. For example, when the en…

Cited by 0SourceScholar
2024

In2SET: Intra-Inter Similarity Exploiting Transformer for Dual-Camera Compressive Hyperspectral Imaging

CVPR 2024poster

Dual-camera compressive hyperspectral imaging (DCCHI) offers the capability to reconstruct 3D hyperspectral image (HSI) by fusing compressive and panchromatic (PAN) image which has shown great potential for snapshot hyperspectral imaging in practice. In this paper we introduce a novel DCCHI reconstr…

2024

Investigation on the multi-solution problem of the kinetostatics of cable-driven continuum manipulators

ICRA 2024poster

Cable-driven continuum manipulators have gained considerable attention due to their high dexterity and inherent structural compliance, making them a popular research topic. However, previous studies have overlooked the kinetostatics of these manipulators, which can result in a multi-solution problem…

Cited by 0SourceScholar
2024

Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models

NeurIPS 2024poster

Large language models (LLMs) and vision-language models (VLMs) have demonstrated remarkable performance across a wide range of tasks and domains. Despite this promise, spatial understanding and reasoning—a fundamental component of human cognition—remains under-explored. We propose SpatialEval, a nov…

2024

MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations

ICLR 2024poster

Multimodal intent recognition poses significant challenges, requiring the incorporation of non-verbal modalities from real-world contexts to enhance the comprehension of human intentions. However, most existing multimodal intent benchmark datasets are limited in scale and suffer from difficulties in…

2024

Mitigate Extrinsic Social Bias in Pre-trained Language Models via Continuous Prompts Adjustment

EMNLP 2024main

Although pre-trained language models (PLMs) have been widely used in natural language understandings (NLU), they are still exposed to fairness issues. Most existing extrinsic debiasing methods rely on manually curated word lists for each sensitive groups to modify training data or to add regular con…

Cited by 2SourcePDFScholar
2024

Molecular Data Programming: Towards Molecule Pseudo-labeling with Systematic Weak Supervision

CVPR 2024poster

The premise for the great advancement of molecular machine learning is dependent on a considerable amount of labeled data. In many real-world scenarios the labeled molecules are limited in quantity or laborious to derive. Recent pseudo-labeling methods are usually designed based on a single domain k…

Cited by 1SourcePDFScholar
2024

Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQA

ACL 2024long

Multipanel images, commonly seen as web screenshots, posters, etc., pervade our daily lives. These images, characterized by their composition of multiple subfigures in distinct layouts, effectively convey information to people. Toward building advanced multimodal AI applications, such as agents that…

Cited by 20SourcePDFScholar
2024

Multimodal Graph Neural Architecture Search under Distribution Shifts

AAAI 2024technical

Multimodal graph neural architecture search (MGNAS) has shown great success for automatically designing the optimal multimodal graph neural network (MGNN) architecture by leveraging multimodal representation, crossmodal information and graph structure in one unified framework. However, existing MGNA…

Cited by 6SourcePDFScholar
2024

Navigation as Attackers Wish? Towards Building Robust Embodied Agents under Federated Learning

NAACL 2024long

Federated embodied agent learning protects the data privacy of individual visual environments by keeping data locally at each client (the individual environment) during training. However, since the local data is inaccessible to the server under federated learning, attackers may easily poison the tra…

Cited by 2SourcePDFScholar
2024

Non-asymptotic Approximation Error Bounds of Parameterized Quantum Circuits

NeurIPS 2024spotlight

Understanding the power of parameterized quantum circuits (PQCs) in accomplishing machine learning tasks is one of the most important questions in quantum machine learning. In this paper, we focus on the PQC expressivity for general multivariate function classes. Previously established Universal App…

Cited by 3SourcePDFScholar
2024

Object Correlation Matrix for Two-Stage Object Detection Network

ICASSP 2024accepted

The relationship between various objects in real life is very important and universal. However, existing object detection models, especially Two-stage models, mostly rely solely on instance learning of individual objects, which use limited global information to extract regions of interest and neglec…

Cited by 0SourceScholar
2024

Pioneering Reliable Assessment in Text-to-Image Knowledge Editing: Leveraging a Fine-Grained Dataset and an Innovative Criterion

EMNLP 2024finding

During pre-training, the Text-to-Image (T2I) diffusion models encode factual knowledge into their parameters. These parameterized facts enable realistic image generation, but they may become obsolete over time, thereby misrepresenting the current state of the world. Knowledge editing techniques aim…

2024

PokeMQA: Programmable knowledge editing for Multi-hop Question Answering

ACL 2024long

Multi-hop question answering (MQA) is one of the challenging tasks to evaluate machine’s comprehension and reasoning abilities, where large language models (LLMs) have widely achieved the human-comparable performance. Due to the dynamics of knowledge facts in real world, knowledge editing has been e…

2024

Preserving Fairness Generalization in Deepfake Detection

CVPR 2024poster

Although effective deepfake detection models have been developed in recent years recent studies have revealed that these models can result in unfair performance disparities among demographic groups such as race and gender. This can lead to particular groups facing unfair targeting or exclusion from…

2024

PrivSGP-VR: Differentially Private Variance-Reduced Stochastic Gradient Push with Tight Utility Bounds

IJCAI 2024poster

In this paper, we propose a differentially private decentralized learning method (termed PrivSGP-VR) which employs stochastic gradient push with variance reduction and guarantees (epsilon, delta)-differential privacy (DP) for each node. Our theoretical analysis shows that, under DP Gaussian noise wi…

Cited by 0SourcePDFScholar
2024

Rethinking Independent Cross-Entropy Loss For Graph-Structured Data

ICML 2024poster

Graph neural networks (GNNs) have exhibited prominent performance in learning graph-structured data. Considering node classification task, based on the i.i.d assumption among node labels, the traditional supervised learning simply sums up cross-entropy losses of the independent training nodes and ap…

2024

Rethinking Propagation for Unsupervised Graph Domain Adaptation

AAAI 2024technical

Unsupervised Graph Domain Adaptation (UGDA) aims to transfer knowledge from a labelled source graph to an unlabelled target graph in order to address the distribution shifts between graph domains. Previous works have primarily focused on aligning data from the source and target graph in the represen…

2024

Self-Supervised Learning for Enhancing Spatial Awareness in Free-Hand Sketches

IJCAI 2024poster

Free-hand sketch, as a versatile medium of communication, can be viewed as a collection of strokes arranged in a spatial layout to convey a concept. Due to the abstract nature of the sketches, changes in stroke position may make them difficult to recognize. Recently, Graphic sketch representations a…

2024

Spoofing Attack Augmentation: Can Differently-Trained Attack Models Improve Generalisation?

ICASSP 2024accepted

A reliable deepfake detector or spoofing countermeasure (CM) should be robust in the face of unpredictable spoofing attacks. To encourage the learning of more generaliseable artefacts, rather than those specific only to known attacks, CMs are usually exposed to a broad variety of different attacks d…

Cited by 0SourceScholar
2024

Synvox2: Towards A Privacy-Friendly Voxceleb2 Dataset

ICASSP 2024accepted

The success of deep learning in speaker recognition relies heavily on the use of large datasets. However, the data-hungry nature of deep learning methods has already being questioned on account the ethical, privacy, and legal concerns that arise when using large-scale datasets of natural speech coll…

Cited by 0SourceScholar
2024

Two-Stage Video Shadow Detection via Temporal-Spatial Adaption

ECCV 2024poster

"Video Shadow Detection (VSD) is an important computer vision task focusing on detecting and segmenting shadows throughout the entire video sequence. Despite their remarkable performance, existing VSD methods and datasets mainly focus on the dominant and isolated shadows. Consequently, VSD under com…

2024

Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal Utterances

ACL 2024long

Discovering the semantics of multimodal utterances is essential for understanding human language and enhancing human-machine interactions. Existing methods manifest limitations in leveraging nonverbal information for discerning complex semantics in unsupervised scenarios. This paper introduces a nov…

2024

VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding

NeurIPS 2024poster

Existing Video Corpus Moment Retrieval (VCMR) is limited to coarse-grained understanding that hinders precise video moment localization when given fine-grained queries. In this paper, we propose a more challenging fine-grained VCMR benchmark requiring methods to localize the best-matched moment from…

2024

ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models

ACL 2024findings

In our work, we explore the synergistic capabilities of pre-trained vision-and-language models (VLMs) and large language models (LLMs) on visual commonsense reasoning (VCR) problems. We find that VLMs and LLMs-based decision pipelines are good at different kinds of VCR problems. Pre-trained VLMs exh…

Cited by 9SourcePDFScholar
2024

WelQrate: Defining the Gold Standard in Small Molecule Drug Discovery Benchmarking

NeurIPS 2024poster

While deep learning has revolutionized computer-aided drug discovery, the AI community has predominantly focused on model innovation and placed less emphasis on establishing best benchmarking practices. We posit that without a sound model evaluation framework, the AI community's efforts cannot reac…

Cited by 0SourcePDFScholar
2023

Adversarially Robust Neural Architecture Search for Graph Neural Networks

CVPR 2023poster

Graph Neural Networks (GNNs) obtain tremendous success in modeling relational data. Still, they are prone to adversarial attacks, which are massive threats to applying GNNs to risk-sensitive domains. Existing defensive methods neither guarantee performance facing new data/tasks or adversarial attack…

Cited by 25SourcePDFScholar
2023

Aerial Vision-and-Dialog Navigation

ACL 2023findings

The ability to converse with humans and follow natural language commands is crucial for intelligent unmanned aerial vehicles (a.k.a. drones). It can relieve people’s burden of holding a controller all the time, allow multitasking, and make drone control more accessible for people with disabilities o…

2023

Alternating Updates for Efficient Transformers

NeurIPS 2023spotlight

It has been well established that increasing scale in deep transformer networks leads to improved quality and performance. However, this increase in scale often comes with prohibitive increases in compute cost and inference latency. We introduce Alternating Updates (AltUp), a simple-to-implement met…

Cited by 6SourcePDFScholar
2023

AutoGT: Automated Graph Transformer Architecture Search

ICLR 2023top-5%

Although Transformer architectures have been successfully applied to graph data with the advent of Graph Transformer, current design of Graph Transformer still heavily relies on human labor and expertise knowledge to decide proper neural architectures and suitable graph encoding strategies at each T…

Cited by 27SourcePDFScholar
2023

CIMI4D: A Large Multimodal Climbing Motion Dataset Under Human-Scene Interactions

CVPR 2023poster

Motion capture is a long-standing research problem. Although it has been studied for decades, the majority of research focus on ground-based movements such as walking, sitting, dancing, etc. Off-grounded actions such as climbing are largely overlooked. As an important type of action in sports and fi…

Cited by 30SourcePDFScholar
2023

Can Knowledge of End-to-End Text-to-Speech Models Improve Neural Midi-to-Audio Synthesis Systems?

ICASSP 2023accepted

With the similarity between music and speech synthesis from symbolic input and the rapid development of text-to-speech (TTS) techniques, it is worthwhile to explore ways to improve the MIDI-to-audio performance by borrowing from TTS techniques. In this study, we analyze the shortcomings of a TTS-bas…

Cited by 0SourceScholar
2023

Clothes Grasping and Unfolding Based on RGB-D Semantic Segmentation

ICRA 2023poster

Clothes grasping and unfolding is a core step in robotic-assisted dressing. Most existing works leverage depth images of clothes to train a deep learning-based model to recognize suitable grasping points. These methods often utilize physics engines to synthesize depth images to reduce the cost of re…

Cited by 5SourceScholar
2023

Controlling Neural Style Transfer with Deep Reinforcement Learning

IJCAI 2023poster

Controlling the degree of stylization in the Neural Style Transfer (NST) is a little tricky since it usually needs hand-engineering on hyper-parameters. In this paper, we propose the first deep Reinforcement Learning (RL) based architecture that splits one-step style transfer into a step-wise proces…

Cited by 1SourcePDFScholar
2023

Curriculum Co-disentangled Representation Learning across Multiple Environments for Social Recommendation

ICML 2023poster

There exist complex patterns behind the decision-making processes of different individuals across different environments. For instance, in a social recommender system, various user behaviors are driven by highly entangled latent factors from two environments, i.e., consuming environment where users…

Cited by 24SourcePDFScholar
2023

Curriculum Multi-Negative Augmentation for Debiased Video Grounding

AAAI 2023technical

Video Grounding (VG) aims to locate the desired segment from a video given a sentence query. Recent studies have found that current VG models are prone to over-rely the groundtruth moment annotation distribution biases in the training set. To discourage the standard VG model's behavior of exploiting…

2023

Decouple Before Interact: Multi-Modal Prompt Learning for Continual Visual Question Answering

ICCV 2023poster

In the real world, a desirable Visual Question Answering model is expected to provide correct answers to new questions and images in a continual setting (recognized as CL-VQA). However, existing works formulate CLVQA from a vision-only or language-only perspective, and straightforwardly apply the un…

Cited by 24PDFScholar
2023

Detection of Real-Time Deepfakes in Video Conferencing with Active Probing and Corneal Reflection

ICASSP 2023accepted

The COVID pandemic has led to the wide adoption of online video calls in recent years. However, the increasing reliance on video calls provides opportunities for new impersonation attacks by fraudsters using the advanced real-time DeepFakes. Real-time DeepFakes pose new challenges to detection metho…

Cited by 0SourceScholar
2023

Doubly Right Object Recognition: A Why Prompt for Visual Rationales

CVPR 2023poster

Many visual recognition models are evaluated only on their classification accuracy, a metric for which they obtain strong performance. In this paper, we investigate whether computer vision models can also provide correct rationales for their predictions. We propose a "doubly right" object recognitio…

2023

Dynamic Heterogeneous Graph Attention Neural Architecture Search

AAAI 2023technical

Dynamic heterogeneous graph neural networks (DHGNNs) have been shown to be effective in handling the ubiquitous dynamic heterogeneous graphs. However, the existing DHGNNs are hand-designed, requiring extensive human efforts and failing to adapt to diverse dynamic heterogeneous graph scenarios. In th…

2023

Hiding Speaker's Sex in Speech Using Zero-Evidence Speaker Representation in an Analysis/Synthesis Pipeline

ICASSP 2023accepted

The use of modern vocoders in an analysis/synthesis pipeline allows us to investigate high-quality voice conversion that can be used for privacy purposes. Here, we propose to transform the speaker embedding and the pitch in order to hide the sex of the speaker. ECAPA-TDNN-based speaker representatio…

Cited by 0SourceScholar
2023

HoloAssist: an Egocentric Human Interaction Dataset for Interactive AI Assistants in the Real World

ICCV 2023poster

Building an interactive AI assistant that can perceive, reason, and collaborate with humans in the real world has been a long-standing pursuit in the AI community. This work is part of a broader research effort to develop intelligent agents that can interactively guide humans through performing task…

Cited by 55PDFcodeScholar
2023

JR2Net: Joint Monocular 3D Face Reconstruction and Reenactment

AAAI 2023technical

Face reenactment and reconstruction benefit various applications in self-media, VR, etc. Recent face reenactment methods use 2D facial landmarks to implicitly retarget facial expressions and poses from driving videos to source images, while they suffer from pose and expression preservation issues fo…

Cited by 3SourcePDFScholar
2023

Joint Data-Task Generation for Auxiliary Learning

NeurIPS 2023poster

Current auxiliary learning methods mainly adopt the methodology of reweighing losses for the manually collected auxiliary data and tasks. However, these methods heavily rely on domain knowledge during data collection, which may be hardly available in reality. Therefore, current methods will become l…

Cited by 3SourcePDFScholar
2023

Large Language Models with Controllable Working Memory

ACL 2023findings

Large language models (LLMs) have led to a series of breakthroughs in natural language processing (NLP), partly owing to the massive amounts of world knowledge they memorize during pretraining. While many downstream applications provide the model with an informational context to aid its underlying t…

Cited by 148SourcePDFScholar
2023

Multi-task Graph Neural Architecture Search with Task-aware Collaboration and Curriculum

NeurIPS 2023poster

Graph neural architecture search (GraphNAS) has shown great potential for automatically designing graph neural architectures for graph related tasks. However, multi-task GraphNAS capable of handling multiple tasks simultaneously has been largely unexplored in literature, posing great challenges to c…

Cited by 13SourcePDFScholar
2023

On the Benefits of Learning to Route in Mixture-of-Experts Models

EMNLP 2023long main

Mixture-of-Expert (MoE) Transformer models, such as the Switch Transformer, allow us to successfully scale up model sizes while keeping the amount of compute time fixed. Prior work has established the computational efficiency benefits of using these models. A core component of these models is a rout…

Cited by 0SourceScholar
2023

Prompt Tuning Pushes Farther, Contrastive Learning Pulls Closer: A Two-Stage Approach to Mitigate Social Biases

ACL 2023long

As the representation capability of Pre-trained Language Models (PLMs) improve, there is growing concern that they will inherit social biases from unprocessed corpora. Most previous debiasing techniques used Counterfactual Data Augmentation (CDA) to balance the training corpus. However, CDA slightly…

Cited by 12SourcePDFScholar
2023

Public Opinion Field Effect Fusion in Representation Learning for Trending Topics Diffusion

NeurIPS 2023poster

Trending topic diffusion and prediction analysis is an important problem and has been well studied in social networks. Representation learning is an effective way to extract node embeddings, which can help for topic propagation analysis by completing downstream tasks such as link prediction and node…

2023

RMBench: Benchmarking Deep Reinforcement Learning for Robotic Manipulator Control

IROS 2023poster

Reinforcement learning is used to tackle complex tasks with high-dimensional sensory inputs. Over the past decade, a wide range of reinforcement learning algorithms have been developed, with recent progress benefiting from deep learning for raw sensory signal representation. This raises a natural qu…

Cited by 4SourcecodeScholar