← Search

Yang Yang

314 accepted papers

2026

$G^2$-Reader: Dual Evolving Graphs for Multimodal Document QA

ICML 2026poster

Retrieval-augmented generation is a practical paradigm for question answering over long documents, but it remains brittle for multimodal reading where text, tables, and figures are interleaved across many pages. First, flat chunking breaks document-native structure and cross-modal alignment, yieldin…

Cited by 1SourceScholar
2026

AdaS: Adaptive Gradient Descent for Spiking Transformers

ICML 2026poster

Transformer-based Spiking Neural Networks (SNNs) combine Transformer performance with SNN energy efficiency through an event-driven self-attention mechanism. However, Spiking Transformers still lag behind their Artificial Neural Network (ANN) counterparts. Most existing studies address this issue th…

Cited by 0SourceScholar
2026

AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition

CVPR 2026

Vision-Language Models (VLMs) have achieved remarkable success in visual question answering tasks, but their reliance on large numbers of visual tokens introduces significant computational overhead. While existing efficient VLM approaches reduce visual tokens through fixed-ratio compression, they op

Cited by 0SourcecodeScholar
2026

Adaptive Debiasing Tsallis Entropy for Test-Time Adaptation

ICLR 2026poster

Mainstream Test-Time Adaptation (TTA) methods for adapting vision-language models, e.g., CLIP, typically rely on Shannon Entropy (SE) at test time to measure prediction uncertainty and inconsistency. However, since CLIP has a built-in bias from pretraining on highly imbalanced web-crawled data, SE i…

Cited by 0SourcecodeScholar
2026

Adaptive Multiscale Binary Expansion Tests for Independence

ICML 2026poster

This paper introduces a new family of adaptive, distribution-free independence tests for multivariate random vectors based on binary expansion coefficients, supported by a tractable asymptotic theory. Our first key contribution establishes a general equivalence between independence testing and testi…

Cited by 0SourceScholar
2026

Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion Model

ICLR 2026poster

Diffusion models have recently emerged as powerful tools for camera simulation, enabling both geometric transformations and realistic optical effects. Among these, image-based bokeh rendering has shown promising results, but diffusion for video bokeh remains unexplored. Existing image-based methods…

Cited by 0SourcecodeScholar
2026

Assembling the Mind's Mosaic: Towards EEG Semantic Intent Decoding

ICLR 2026poster

Enabling natural communication through brain–computer interfaces (BCIs) remains one of the most profound challenges in neuroscience and neurotechnology. While existing frameworks offer partial solutions, they are constrained by oversimplified semantic representations and a lack of interpretability.…

Cited by 0SourceScholar
2026

Calibrating Uncertainty for Zero-Shot Adversarial CLIP

ICML 2026poster

CLIP delivers strong zero-shot classification but remains highly vulnerable to adversarial attacks. Prior adversarial fine-tuning work largely focuses on matching the predicted logits between clean and adversarial examples, which overlooks uncertainty calibration and may degrade the zero-shot genera…

Cited by 0SourceScholar
2026

DWTSG: Parameter-Efficient Fine-Tuning of Large Pre-trained Models via Discrete Wavelet Transform and Subband Guidance

AAAI 2026technical

Fully fine-tuning large pre-trained models for each downstream task is impractical due to prohibitive memory, computation, and storage costs. Although parameter-efficient fine-tuning (PEFT) methods address this issue, leading methods like LoRA still exhibit linear scaling of trainable parameters wit

Cited by 0SourcePDFScholar
2026

DiffWind: Physics-Informed Differentiable Modeling of Wind-Driven Object Dynamics

ICLR 2026poster

Modeling wind-driven object dynamics from video observations is highly challenging due to the invisibility and spatio–temporal variability of wind, as well as the complex deformations of objects. We present DiffWind, a physics-informed differentiable framework that unifies wind–object interaction mo…

Cited by 0SourcecodeScholar
2026

Distilling Future Temporal Knowledge with Masked Feature Reconstruction for 3D Object Detection

AAAI 2026technical

Camera-based temporal 3D object detection has shown impressive results in autonomous driving, with offline models improving accuracy by using future frames. Knowledge distillation (KD) can be an appealing framework for transferring rich information from offline models to online models. However, exis

Cited by 0SourcePDFScholar
2026

Domain Adaptive Object Detection via Dynamic Causal Refinement

ICML 2026poster

Domain Adaptive Object Detection (DAOD) addresses the challenge of transferring object detectors from labeled source domains to unlabeled target domains. Existing domain adaptation methods primarily rely on feature distribution alignment, which enhances domain-invariant features (statistical invaria…

Cited by 0SourceScholar
2026

Efficient Autoregressive Inference for Transformer Probabilistic Models

ICLR 2026poster

Transformer-based models for amortized probabilistic inference, such as neural processes, prior-fitted networks, and tabular foundation models, excel at single-pass *marginal* prediction. However, many real-world applications require coherent *joint distributions* that capture dependencies between p…

Cited by 0SourceScholar
2026

EpiCoCo: De Novo Epitope Generation via MHC-Context Co-Modeling and Contrastive Affinity Guidance

ICML 2026poster

The *de novo* generation of high-affinity epitopes tailored to specific major histocompatibility complex (MHC) proteins is a pivotal challenge in computational immunotherapy. However, current methods struggle to effectively integrate the MHC context into the generation process, and often fail to gua…

Cited by 0SourceScholar
2026

Experience Transfer for Multimodal LLM Agents in Minecraft Game

CVPR 2026

Multimodal LLM agents operating in complex game environments must continually reuse past experience to solve new tasks efficiently. In this work, we propose Echo, a transfer-oriented memory framework that enables agents to derive actionable knowledge from prior interactions rather than treating memo

Cited by 0SourceScholar
2026

From Dialogue to Destination: Geography-Aware Large Language Models with Multimodal Fusion for Conversational Recommendation

AAAI 2026technical

Conversational Recommender Systems (CRS) aim to provide personalized recommendations by interacting with users through natural language dialogue. However, in scenarios requiring deep geospatial awareness, existing methods, including those based on Large Language Models (LLMs), still face significant

Cited by 0SourcePDFScholar
2026

From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection

ICML 2026poster

With rapid advances in audio-visual generative models, reliable forgery detection becomes increasingly critical. Existing methods for audio-visual deepfake detection typically rely on cross-modal inconsistencies. In singing, rhythmic vocalization weakens this coupling and introduces a nontrivial dom…

Cited by 0SourceScholar
2026

Generating-Filtering-Ranking: A Three-Stage MultiModal Data Augmentation Framework Under Partial Modality Missing

AAAI 2026technical

Multimodal data significantly improves the performance of pretrained models, but its practical application is often limited by missing or incomplete data across modalities. There are two key challenges that existing methods of synthesizing missing data face: (1) semantic inaccuracies due to model ha

Cited by 0SourcePDFScholar
2026

GeoPurify: A Data-Efficient Geometric Distillation Framework for Open-Vocabulary 3D Segmentation

ICLR 2026poster

Recent attempts to transfer features from 2D Vision–Language Models (VLMs) to 3D semantic segmentation expose a persistent trade-off. Directly projecting 2D features into 3D yields noisy and fragmented predictions, whereas enforcing geometric coherence necessitates costly training pipelines and larg…

Cited by 0SourcecodeScholar
2026

History to Future: Evolving Agent with Experience and Thought for Zero-shot Vision-and-Language Navigation

CVPR 2026

Vision-and-Language Navigation in Continuous Environment (VLN-CE) requires an agent to follow language instructions to navigate the target destination. With the advancement of large language models (LLMs), recent efforts have explored adapting them for zero-shot VLN-CE, offering a promising solution

Cited by 0SourceScholar
2026

Hyperspectral Image Fusion with Spectral-Band and Fusion-Scale Agnosticism

ICML 2026poster

Current deep learning models for Multispectral and Hyperspectral Image Fusion (MS/HS fusion) are typically designed for fixed spectral bands and spatial scales, which limits their transferability across diverse sensors. To address this, we propose SSA, a universal framework for MS/HS fusion with spe…

Cited by 0SourceScholar
2026

Kronecker Generative Networks: A General Neural Architecture for Parameter-Efficient Learning Across Classification Tasks

ICML 2026poster

Modern neural networks derive much of their effectiveness from rich connectivity patterns. Yet, existing architectures often fix the topology at either the sparse or dense extremes, thereby limiting structural flexibility and analysis. We propose Kronecker Generative Networks (KGNs), an algebraic fr…

Cited by 0SourceScholar
2026

LATO: 3D Mesh Flow Matching with Structured TOpology Preserving LAtents

ICML 2026poster

In this paper, we introduce LATO, a novel topology-preserving latent representation that enables scalable, flow matching-based synthesis of explicit 3D meshes. LATO represents a mesh as a Vertex Displacement Field (VDF) anchored on surface, incorporating a sparse voxel Variational Autoencoder (VAE) …

Cited by 0SourceScholar
2026

LC3: Long Cross-Language Code Clone Detection Enhanced by Opcode Sequences and Affinity Aggregation

AAAI 2026technical

Cross-language code clone detection, which identifies functionally similar code across programming languages, is critical for ensuring synchronized evolution and reducing maintenance costs in multi-platform software development. While zero-shot approaches have emerged as a practical solution to data

Cited by 0SourcePDFScholar
2026

LLaVA-FA: Learning Fourier Approximation for Compressing Large Multimodal Models

ICLR 2026poster

Large multimodal models (LMMs) have achieved impressive performance on various vision-language tasks, but their substantial computational and memory costs hinder their practical deployment. Existing compression methods often decouple low-rank decomposition and quantization, leading to compounded rec…

Cited by 0SourceScholar
2026

Learning Global Hypothesis Space for Enhancing Synergistic Reasoning Chain

ICLR 2026poster

Chain-of-Thought (CoT) has emerged as an effective paradigm to enhance the reasoning ability of large language models (LLMs) in complex tasks. However, existing approaches still face two major challenges: (1) the lack of a global mechanism to integrate and interact across diverse reasoning hypothese…

Cited by 0SourceScholar
2026

Learning Molecular Chirality via Chiral Determinant Kernels

ICLR 2026poster

Chirality is a fundamental molecular property that governs stereospecific behavior in chemistry and biology. Capturing chirality in machine learning models remains challenging due to the geometric complexity of stereochemical relationships and the limitations of traditional molecular representations…

Cited by 0SourcecodeScholar
2026

LinearSR: Unlocking Linear Attention for Stable and Efficient Image Super-Resolution

ICLR 2026poster

Generative models for Image Super-Resolution (SR) are increasingly powerful, yet their reliance on self-attention's quadratic complexity ($O(N^2)$) creates a major computational bottleneck. Linear Attention offers an $O(N)$ solution, but its promise for photorealistic SR has remained largely untappe…

Cited by 0SourcecodeScholar
2026

MaskGuide: Efficient Distillation for Deployable Lightweight Segmentation in Marine Environments

RA-L 2026

The growing demand for efficient image segmentation in marine ecological studies is currently constrained by two key factors: the high computational requirements of models such as the Segment Anything Model (SAM) and the degraded accuracy of lightweight models in underwater environments. To overcome

Cited by 0SourceScholar
2026

MeshRipple: Structured Autoregressive Generation of Artist-Meshes

CVPR 2026

Meshes serve as a primary representation for 3D assets. Autoregressive mesh generators serialize faces into sequences and train on truncated segments with sliding-window inference to cope with memory limits. However, this mismatch breaks long-range geometric dependencies, producing holes and fragmen

Cited by 0SourceScholar
2026

NOTAM-Evolve: A Knowledge-Guided Self-Evolving Optimization Framework with LLMs for NOTAM Interpretation

AAAI 2026technical

Accurate interpretation of Notices To Airmen (NOTAMs) is critical for aviation safety, yet their condensed and cryptic language poses significant challenges to both manual and automated processing. Existing automated systems are typically limited to "Shallow Parsing," failing to extract the actionab

Cited by 0SourcePDFScholar
2026

Neural Dynamics Self-Attention for Spiking Transformers

ICLR 2026poster

Integrating Spiking Neural Networks (SNNs) with Transformer architectures offers a promising pathway to balance energy efficiency and performance, particularly for edge vision applications. However, existing Spiking Transformers face two critical challenges: i) a substantial performance gap relative…

Cited by 0SourceScholar
2026

On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs

ICML 2026poster

Reinforcement learning (RL) fine-tuning is now widely used to improve LLM reasoning, and recent work has begun extending it to vision-language models (VLMs). While RL-tuned VLMs can improve visual reasoning benchmark performance, they can still suffer from weak visual grounding, hallucinations, and …

Cited by 0SourceScholar
2026

On the Feasibility of Using MultiModal LLMs to Execute AR Social Engineering Attacks

AAAI 2026technical

Augmented Reality (AR) and Multimodal Large Language Models (LLMs) are rapidly evolving, providing unprecedented capabilities for human-computer interaction. However, their integration introduces a new attack surface for Social Engineering (SE). In this paper, we systematically investigate the feasi

Cited by 0SourcePDFScholar
2026

Positional Encoding for Spiking Transformers

ICML 2026poster

Spiking Neural Networks (SNNs) demonstrate superior energy efficiency over conventional Artificial Neural Networks (ANNs). Recent advances in Transformer-based SNNs have shown encouraging performance by seamlessly integrating spike-driven computation with Transformer architectures. Positional inform…

Cited by 0SourceScholar
2026

PriorGuide: Test-Time Prior Adaptation for Simulation-Based Inference

ICLR 2026poster

Amortized simulator-based inference offers a powerful framework for tackling Bayesian inference in computational fields such as engineering or neuroscience, increasingly leveraging modern generative methods like diffusion models to map observed data to model parameters or future predictions. These a…

Cited by 0SourceScholar
2026

RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer

ICASSP 2026oral

Audio-driven portrait animation aims to synthesize realistic and natural talking head videos from an input audio signal and a single reference image. While existing methods achieve high-quality results by leveraging high-dimensional intermediate representations and explicitly modeling motion dynamic…

Cited by 0SourcePDFScholar
2026

RLIE: Rule Generation with Logistic Regression, Iterative Refinement, and Evaluation for Large Language Models

ICML 2026poster

Large Language Models (LLMs) can propose natural-language rules, circumventing the reliance on a predefined predicate space in traditional rule learning. However, existing LLM-based methods often neglect the global interactions among rules, and the potential of using fine-grained rule importance sco…

Cited by 0SourceScholar
2026

Rethinking Time-Series Imputation as Conditional Inference along Temporal Evolution

ICML 2026poster

Real-world time-series data often suffer from missing observations, hindering long-range temporal modeling. However, most existing imputation methods formulate imputation as conditional reconstruction over limited context, which restricts temporal information propagation and fails to explicitly mode…

Cited by 0SourceScholar
2026

Robust Spiking Neural Networks Against Adversarial Attacks

ICLR 2026poster

Spiking Neural Networks (SNNs) represent a promising paradigm for energy-efficient neuromorphic computing due to their bio-plausible and spike-driven characteristics. However, the robustness of SNNs in complex adversarial environments remains significantly constrained. In this study, we theoretical…

Cited by 0SourceScholar
2026

STARK: Strategic Team of Agents for Refining Kernels

ICLR 2026poster

The efficiency of GPU kernels is central to the progress of modern AI, yet optimizing them remains a difficult and labor-intensive task due to complex interactions between memory hierarchies, thread scheduling, and hardware-specific characteristics. While recent advances in large language models (LL…

Cited by 0SourceScholar
2026

Seeing Beyond Illusion: Generalized and Efficient Mirror Detection

AAAI 2026technical

Reflective imaging enables the mirror imagings and physical entities to possess identical attributes, e.g., color and shape. Current mirror detection (MD) methods primarily rely on designing functional components to establish the correlation and disparities between the imagings and entities, thereby

Cited by 0SourcePDFScholar
2026

SmoothSpike: Spiking Transformer with Learnable Hadamard Transformation

ICML 2026spotlight

Spiking Neural Networks (SNNs) that leverage sparse binary spikes and temporal dynamics have emerged as energy-efficient alternatives to Artificial Neural Networks (ANNs). However, SNNs suffer from limited representational capacity due to the discrete nature of spikes. Existing solutions extending s…

Cited by 0SourceScholar
2026

SpikingLM: Towards Fully Spiking Language Model

ICML 2026poster

Spiking Neural Networks (SNNs) offer a promising avenue toward energy-efficient language modeling by replacing multiply-accumulate operations with sparse, event-driven computation. However, constructing fully spiking language models reveals two fundamental challenges: (1) gradient degradation from d…

Cited by 0SourceScholar
2026

TC-Pade: Trajectory-Consistent Pade Approximation for Diffusion Acceleration

CVPR 2026

Despite achieving state-of-the-art generation quality, diffusion models are hindered by the substantial computational burden of their iterative sampling process. While feature caching techniques achieve effective acceleration at higher step counts (e.g., 50 steps), they exhibit critical limitations

Cited by 0SourceScholar
2026

TCS Jumper: A Bio-Inspired Jumping Robot Featuring High Energy Density via Synergistic Deformation

RA-L 2026

Maximizing the energy density of springs is a key consideration in designing spring-driven jumping robots. Inspired by the synergistic deformation mechanism of springtails, which involves the coupled action of rotational and bending deformations in their furcula, we created the TCS Jumper (Torsion s

Cited by 0SourceScholar
2026

TP-Spikformer: Token Pruned Spiking Transformer

ICLR 2026poster

Spiking neural networks (SNNs) offer an energy-efficient alternative to traditional neural networks due to their event-driven computing paradigm. However, recent advancements in spiking transformers have focused on improving accuracy with large-scale architectures, which require significant computat…

Cited by 0SourceScholar
2026

Temporal Interaction in Spiking Transformers with Multi-Delay Mixer

CVPR 2026

Spiking Neural Networks (SNNs) have gained significant attention due to their event-driven computational paradigm, making them promising for neuromorphic computing. In recent years, the integration of SNNs and Transformer architectures has made remarkable progress in various tasks. However, existing

Cited by 0SourceScholar
2026

Text summarization via global structure awareness

ICLR 2026poster

Text summarization is a core task in natural language processing (NLP). With the rapid growth of information, handling long documents has become increasingly demanding, making summarization essential. Existing research mainly focuses on model improvements and sentence-level pruning, but often overlo…

Cited by 0SourceScholar
2026

Thermally Activated Dual-Modal Adversarial Clothing against AI Surveillance Systems

CVPR 2026

Adversarial patches have emerged as a popular privacy-preserving approach for resisting AI-driven surveillance systems. However, their conspicuous appearance makes them difficult to deploy in real-world scenarios. In this paper, we propose a thermally activated adversarial wearable designed to ensur

Cited by 0SourceScholar
2026

Towards Generalizable AI-Generated Image Detection via Image-Adaptive Prompt Learning

CVPR 2026

In AI-generated image detection, current cutting-edge methods typically adapt pre-trained foundation models through partial-parameter fine-tuning. However, these approaches often struggle to generalize to forgeries from unseen generators, as the fine-tuned models capture only limited patterns from t

Cited by 0SourceScholar
2026

Trimming the Fat: Redundancy-Aware Acceleration Framework for DGNNs

AAAI 2026technical

Temporal graphs are essential for modeling complex real-world systems, such as social interactions, financial transactions, and recommendation systems, but the high computational cost and model complexity of dynamic graph neural networks (DGNNs) pose significant challenges for practical deployment.

Cited by 0SourcePDFScholar
2026

UniFLoW: Universal Multi-Modal Federated LoRA Fine-Tuning Framework with Analytical Aggregation

ICML 2026poster

As Multimodal Large Language Models (MLLMs) continue to be trained, the availability of public data diminishes, limiting the possibility for further training and adaptation. However, private data remains an underutilized yet valuable resource. Federated Learning (FL) enables decentralized training o…

Cited by 0SourceScholar
2026

Views Attention Fusion of Granular-ball Fuzzy Representations Split for Improved Multi-view Clustering

AAAI 2026technical

Multi-View Clustering (MVC) is a pivotal multi-view learning paradigm widely adopted across various fields. Despite recent advances, existing methods primarily focus on enhancing the performance of fused multi-view representation, often neglecting the issue of Representation Degradation (RD) arising

Cited by 0SourcePDFScholar
2025

A Parameter-Efficient Tuning Framework for Language-Guided Object Grounding and Robot Grasping

ICRA 2025

The language-guided robot grasping task requires a robot agent to integrate multimodal information from both visual and linguistic inputs to predict actions for target-driven grasping. While recent approaches utilizing Multimodal Large Language Models (MLLMs) have shown promising results, their exte

Cited by 7SourceScholar
2025

A Soft Active Surface Gripper for Safe In Hand Manipulation of Fragile Objects

IROS 2025

This paper introduces a soft active surface gripper designed to manipulate fragile objects safely. This gripper consists of two fingers, each equipped with two compliant pneumatic actuators and a soft active surface. The gripper utilizes the elastic belt as its soft active surface, which is driven b

Cited by 0SourceScholar
2025

APIMig: A Project-Level Cross-Multi-Version API Migration Framework Based on Evolution Knowledge Graph

IJCAI 2025

API migration is essential for software maintenance due to the rapid evolution of third-party libraries where API elements may change continuously through updates. There are two main challenges for API migration at the project level, especially across multiple versions: 1) lack of specific library e

Cited by 0SourcePDFScholar
2025

Agri-CM3: A Chinese Massive Multi-modal, Multi-level Benchmark for Agricultural Understanding and Reasoning

ACL 2025long

Multi-modal Large Language Models (MLLMs) integrating images, text, and speech can provide farmers with accurate diagnoses and treatment of pests and diseases, enhancing agricultural efficiency and sustainability. However, existing benchmarks lack comprehensive evaluations, particularly in multi-lev…

2025

An Effective and Secure Federated Multi-View Clustering Method with Information-Theoretic Perspective

ICML 2025poster

Recently, federated multi-view clustering (FedMVC) has gained attention for its ability to mine complementary clustering structures from multiple clients without exposing private data. Existing methods mainly focus on addressing the feature heterogeneity problem brought by views on different clients…

Cited by 0SourcePDFScholar
2025

BRIGHT-VO: Brightness-Guided Hybrid Transformer for Visual Odometry with Multi-modality Refinement Module

IJCAI 2025

Visual odometry (VO) plays a crucial role in autonomous driving, robotic navigation, and other related tasks by estimating the position and orientation of a camera based on visual input. Significant progress has been made in data-driven VO methods, particularly those leveraging deep learning techniq

2025

BSO: Binary Spiking Online Optimization Algorithm

ICML 2025poster

Binary Spiking Neural Networks (BSNNs) offer promising efficiency advantages for resource-constrained computing. However, their training algorithms often require substantial memory overhead due to latent weights storage and temporal processing requirements. To address this issue, we propose Binary S…

2025

BiTSpoke: A Leg-Wheel Robot With Single-Motor Driven Actively-Transformable Spoke Wheels

RA-L 2025

To address the issue that current transformable spoke-wheeled leg-wheel robots cannot simultaneously achieve simple structure, stable locomotion and open step climbing ability, a novel robot design method is proposed. The robot called BiTSpoke is driven by two transformable spoke wheels, with a pass

Cited by 4SourceScholar
2025

Binary Event-Driven Spiking Transformer

IJCAI 2025

Transformer-based Spiking Neural Networks (SNNs) introduce a novel event-driven self-attention paradigm that combines the high performance of Transformers with the energy efficiency of SNNs. However, the larger model size and increased computational demands of the Transformer structure limit their p

2025

Bipolar Self-attention for Spiking Transformers

NeurIPS 2025spotlight

Harnessing the event-driven characteristic, Spiking Neural Networks (SNNs) present a promising avenue toward energy-efficient Transformer architectures. However, existing Spiking Transformers still suffer significant performance gaps compared to their Artificial Neural Network counterparts. Through…

Cited by 0SourceScholar
2025

CDTR: Semantic Alignment for Video Moment Retrieval Using Concept Decomposition Transformer

AAAI 2025technical

Video Moment Retrieval (VMR) involves locating specific moments within a video based on natural language queries. However, existing VMR methods that employ various strategies for cross-modal alignment still face challenges such as limited understanding of fine-grained semantics, semantic overlap, an…

Cited by 0SourcePDFScholar
2025

Chain-of-Focus Prompting: Leveraging Sequential Visual Cues to Prompt Large Autoregressive Vision Models

ICLR 2025poster

In-context learning (ICL) has revolutionized natural language processing by enabling models to adapt to diverse tasks with only a few illustrative examples. However, the exploration of ICL within the field of computer vision remains limited. Inspired by Chain-of-Thought (CoT) prompting in the langua…

Cited by 0SourcePDFScholar
2025

Cypher-RI: Reinforcement Learning for Integrating Schema Selection into Cypher Generation

NeurIPS 2025poster

The increasing utilization of graph databases across various fields stems from their capacity to represent intricate interconnections. Nonetheless, exploiting the full capabilities of graph databases continues to be a significant hurdle, largely because of the inherent difficulty in translating natu…

Cited by 0SourceScholar
2025

Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence Modeling

NeurIPS 2025poster

The explosive growth in sequence length has intensified the demand for effective and efficient long sequence modeling. Benefiting from intrinsic oscillatory membrane dynamics, Resonate-and-Fire (RF) neurons can efficiently extract frequency components from input signals and encode them into spatiote…

Cited by 0SourceScholar
2025

Efficient Adaptation of Pre-trained Vision Transformer underpinned by Approximately Orthogonal Fine-Tuning Strategy

ICCV 2025poster

A prevalent approach in Parameter-Efficient Fine-Tuning (PEFT) of pre-trained Vision Transformers (ViT) involves freezing the majority of the backbone parameters and solely learning low-rank adaptation weight matrices to accommodate downstream tasks. These low-rank matrices are commonly derived thro…

2025

Evolving Minds: Logic-Informed Inference from Temporal Action Patterns

ICML 2025poster

Understanding human mental states—such as intentions and desires—is crucial for natural AI-human collaboration. However, this is challenging because human actions occur irregularly over time, and the underlying mental states that drive these actions are unobserved. To tackle this, we propose a novel…

Cited by 0SourcePDFScholar
2025

Exploring Stiffness Gradient Effects in Magnetically Induced Metamorphic Materials via Continuum Simulation and Validation

IROS 2025

Magnetic soft continuum robots are capable of bending with remote control in confined space environments, and they have been applied in various bioengineering contexts. As one type of ferromagnetic soft continuums, the Magnetically Induced Metamorphic Materials (MIMMs)-based continuum (MC) exhibits

Cited by 0SourceScholar
2025

FFBGNet: Full-Flow Bidirectional Feature Fusion Grasp Detection Network Based on Hybrid Architecture

RA-L 2025

Effectively integrating the complementary information from RGB-D images presents a significant challenge in robotic grasping. In this letter, we propose a full-flow bidirectional feature fusion grasp detection network (FFBGNet) based on a hybrid architecture to generate accurate grasp poses from RGB

Cited by 3SourceScholar
2025

Financial Language Model Evaluation (FLaME)

ACL 2025finding

Language Models (LMs) have demonstrated impressive capabilities with core Natural Language Processing (NLP) tasks. The effectiveness of LMs for highly specialized knowledge-intensive tasks in finance remains difficult to assess due to major gaps in the methodologies of existing evaluation frameworks…

2025

Generating Full-field Evolution of Physical Dynamics from Irregular Sparse Observations

NeurIPS 2025poster

Modeling and reconstructing multidimensional physical dynamics from sparse and off-grid observations presents a fundamental challenge in scientific research. Recently, diffusion-based generative modeling shows promising potential for physical simulation. However, current approaches typically operate…

Cited by 0SourceScholar
2025

Generating Is Believing: Membership Inference Attacks against Retrieval-Augmented Generation

ICASSP 2025accepted

Retrieval-Augmented Generation (RAG) is a state-of-the-art technique that mitigates issues such as hallucinations and knowledge staleness in Large Language Models (LLMs) by retrieving relevant knowledge from an external database to assist in content generation. Existing research has demonstrated pot…

Cited by 0SourceScholar
2025

Implicit Counterfactual Learning for Audio-Visual Segmentation

ICCV 2025poster

Audio-visual segmentation (AVS) aims to segment objects in videos based on audio cues. Existing AVS methods are primarily designed to enhance interaction efficiency but pay limited attention to modality representation discrepancies and imbalances. To overcome this, we propose the implicit counterfac…

Cited by 0SourcePDFScholar
2025

KAA: Kolmogorov-Arnold Attention for Enhancing Attentive Graph Neural Networks

ICLR 2025poster

Graph neural networks (GNNs) with attention mechanisms, often referred to as attentive GNNs, have emerged as a prominent paradigm in advanced GNN models in recent years. However, our understanding of the critical process of scoring neighbor nodes remains limited, leading to the underperformance of m…

2025

KGE Calibrator: An Efficient Probability Calibration Method of Knowledge Graph Embedding Models for Trustworthy Link Prediction

EMNLP 2025

Knowledge graph embedding (KGE) models are designed for the task of link prediction, which aims to infer missing triples by learning representations for entities and relations. While KGE models excel at ranking-based link prediction, the critical issue of probability calibration has been largely ove

2025

LASeR: Towards Diversified and Generalizable Robot Design with Large Language Models

ICLR 2025poster

Recent advances in Large Language Models (LLMs) have stimulated a significant paradigm shift in evolutionary optimization, where hand-crafted search heuristics are gradually replaced with LLMs serving as intelligent search operators. However, these studies still bear some notable limitations, includ…

2025

Layer-Animate for Transparent Video Generation

ICASSP 2025accepted

Transparent videos with alpha channels play a crucial role in film production, advertising, and augmented reality fields. However, there is currently no available method for producing transparent videos. Traditional methods are time-consuming and labor-intensive, and employing alternative approaches…

Cited by 0SourceScholar
2025

Leveraging Asynchronous Spiking Neural Networks for Ultra Efficient Event-Based Visual Processing

AAAI 2025technical

Event cameras encode visual information by generating asynchronous and sparse event streams, which hold great potential for low latency and low power consumption. Despite many successful implementations of event camera-based applications, most of them accumulate the events into frames and then utili…

Cited by 0SourcePDFScholar
2025

Local Conditional Controlling for Text-to-Image Diffusion Models

AAAI 2025technical

Diffusion models have exhibited impressive prowess in the text-to-image task. Recent methods add image-level structure controls, e.g., edge and depth maps, to manipulate the generation process together with text prompts to obtain desired images. This controlling process is globally operated on the e…

2025

MSE-Adapter: A Lightweight Plugin Endowing LLMs with the Capability to Perform Multimodal Sentiment Analysis and Emotion Recognition

AAAI 2025technical

Current multimodal sentiment analysis (MSA) and emotion recognition in conversations (ERC) methods based on pre-trained language models exhibit two primary limitations: 1) Once trained for MSA and ERC tasks, these pre-trained language models lose their original generalized capabilities. 2) They dema…

2025

MTGIB-UNet: A Multi-Task Graph Information Bottleneck and Uncertainty Weighted Network for ADMET Prediction

IJCAI 2025

Accurate prediction of ADMET (Absorption, Distribution, Metabolism, Excretion, and Toxicity) properties is crucial in drug development, as these properties directly impact a drug's efficacy and safety. However, existing multi-task learning models often face challenges related to noise interference a

Cited by 0SourcePDFScholar
2025

MTTM: Memory-Augmented with Mamba for 3D Medical Images Analysis

ICASSP 2025accepted

The rapid advancement of artificial intelligence has propelled the healthcare industry into a new era of diagnostic precision. A pivotal component of this evolution is the accurate classification of 3D medical images, which necessitates extracting robust feature representations capable of effectivel…

Cited by 0SourceScholar
2025

MedSegFactory: Text-Guided Generation of Medical Image-Mask Pairs

ICCV 2025poster

This paper presents **MedSegFactory**, a versatile medical synthesis framework that generates high-quality paired medical images and segmentation masks across modalities and tasks. It aims to serve as an unlimited data repository, supplying image-mask pairs to enhance existing segmentation tools. Th…

Cited by 0SourcePDFScholar
2025

Memory-Free and Parallel Computation for Quantized Spiking Neural Networks

ICASSP 2025accepted

Quantized Spiking Neural Networks (QSNNs) offer superior energy efficiency and are well-suited for deployment on resource-limited edge devices. However, limited bit-width weight and membrane potential result in a notable performance decline. In this study, we first identify a new underlying cause fo…

Cited by 0SourceScholar
2025

Microtitre Plate Image Augmentation with Generative Adversarial Networks

ICASSP 2025accepted

Antibiotic Susceptibility Testing (AST) based on microorganism culturing is the gold-standard technique to determine whether a pathogen is susceptible or resistant to available antibiotics. While broth microdilution offers a potential high-throughput method for AST, reading and interpreting microtit…

Cited by 0SourceScholar
2025

Mono3DVLT: Monocular-Video-Based 3D Visual Language Tracking

CVPR 2025poster

Visual-Language Tracking (VLT) is emerging as a promising paradigm to bridge the human-machine performance gap. For single objects, VLT broadens the problem scope to text-driven video comprehension. Yet, this direction is still confined to 2D spatial extents, currently lacking the ability to deal wi…

2025

Multi-Label Node Classification with Label Influence Propagation

ICLR 2025poster

Graphs are a complex and versatile data structure used across various domains, with possibly multi-label nodes playing a particularly crucial role. Examples include proteins in PPI networks with multiple functions and users in social or e-commerce networks exhibiting diverse interests. Tackling mu…

Cited by 0SourcePDFScholar
2025

Multi-Label Test-Time Adaptation with Bound Entropy Minimization

ICLR 2025poster

Mainstream test-time adaptation (TTA) techniques endeavor to mitigate distribution shifts via entropy minimization for multi-class classification, inherently increasing the probability of the most confident class. However, when encountering multi-label instances, the primary challenge stems from the…

2025

Multi-View Graph Clustering via Node-Guided Contrastive Encoding

ICML 2025poster

Multi-view clustering has gained significant attention for integrating multi-view information in multimedia applications. With the growing complexity of graph data, multi-view graph clustering (MVGC) has become increasingly important. Existing methods primarily use Graph Neural Networks (GNNs) to en…

Cited by 0SourcePDFScholar
2025

Online Fraud Detection via Test-Time Retrieval-Based Representation Enrichment

AAAI 2025technical

Anti-fraud machine learning systems are perpetually confronted with the significant challenge of concept drift, driven by the continuous and intense evolution of fraudulent techniques. That is, outdated models trained on historical fraudulent behaviors often fall short in addressing the evolving tac…

Cited by 0SourcePDFScholar
2025

PanTS: The Pancreatic Tumor Segmentation Dataset

NeurIPS 2025poster

PanTS is a large-scale, multi-institutional dataset curated to advance research in pancreatic CT analysis. It contains 36,390 CT scans from 145 medical centers, with expert-validated, voxel-wise annotations of over 993,000 anatomical structures, covering pancreatic tumors, pancreas head, body, and t…

Cited by 0SourceScholar
2025

QP-SNN: Quantized and Pruned Spiking Neural Networks

ICLR 2025poster

Brain-inspired Spiking Neural Networks (SNNs) leverage sparse spikes to encode information and operate in an asynchronous event-driven manner, offering a highly energy-efficient paradigm for machine intelligence. However, the current SNN community focuses primarily on performance improvement by deve…

Cited by 0SourcePDFScholar
2025

Quantized Spike-driven Transformer

ICLR 2025poster

Spiking neural networks (SNNs) are emerging as a promising energy-efficient alternative to traditional artificial neural networks (ANNs) due to their spike-driven paradigm. However, recent research in the SNN domain has mainly focused on enhancing accuracy by designing large-scale Transformer struct…

2025

RCTrans: Radar-Camera Transformer via Radar Densifier and Sequential Decoder for 3D Object Detection

AAAI 2025technical

In radar-camera 3D object detection, the radar point clouds are sparse and noisy, which causes difficulties in fusing camera and radar modalities. To solve this, we introduce a novel query-based detection method named Radar-Camera Transformer (RCTrans). Specifically, we first design a Radar Dense En…

2025

RLKGF: Reinforcement Learning from Knowledge Graph Feedback Without Human Annotations

ACL 2025finding

Reinforcement Learning from Human Feedback (RLHF) has been shown to effectively align large language models (LLMs) with human knowledge. However, the lack of human preference labels remains a significant bottleneck when applying RLHF to a downstream domain. Humans in RLHF play a critical role in inj…

2025

RadGPT: Constructing 3D Image-Text Tumor Datasets

ICCV 2025poster

Cancers identified in CT scans are usually accompanied by detailed radiology reports, but publicly available CT datasets often lack these essential reports. This absence limits their usefulness for developing accurate report generation AI. To address this gap, we present AbdomenAtlas 3.0, the first…

2025

Reaction Prediction via Interaction Modeling of Symmetric Difference Shingle Sets

NeurIPS 2025poster

Chemical reaction prediction remains a fundamental challenge in organic chemistry, where existing machine learning models face two critical limitations: sensitivity to input permutations (molecule/atom orderings) and inadequate modeling of substructural interactions governing reactivity. These short…

Cited by 0SourceScholar
2025

RealisHuman: A Two-Stage Approach for Refining Malformed Human Parts in Generated Images

AAAI 2025technical

In recent years, diffusion models have revolutionized visual generation, outperforming traditional frameworks like Generative Adversarial Networks (GANs). However, generating images of humans with realistic semantic parts, such as hands and faces, remains a significant challenge due to their intric…

2025

Rethinking Multimodal Learning from the Perspective of Mitigating Classification Ability Disproportion

NeurIPS 2025oral

Multimodal learning (MML) is significantly constrained by modality imbalance, leading to suboptimal performance in practice. While existing approaches primarily focus on balancing the learning of different modalities to address this issue, they fundamentally overlook the inherent disproportion in mo…

Cited by 0SourcecodeScholar
2025

S$^2$NN: Sub-bit Spiking Neural Networks

NeurIPS 2025poster

Spiking Neural Networks (SNNs) offer an energy-efficient paradigm for machine intelligence, but their continued scaling poses challenges for resource-limited deployment. Despite recent advances in binary SNNs, the storage and computational demands remain substantial for large-scale networks. To furt…

Cited by 0SourceScholar
2025

SCoder: Progressive Self-Distillation for Bootstrapping Small-Scale Data Synthesizers to Empower Code LLMs

EMNLP 2025

Existing code large language models (LLMs) often rely on large-scale instruction data distilled from proprietary LLMs for fine-tuning, which typically incurs high costs. In this paper, we explore the potential of small-scale open-source LLMs (e.g., 7B) as synthesizers for high-quality code instructi

2025

SFma-Unet: A Mamba-Based Spatial-Frequency Fusion Network for Medical Image Segmentation

ICASSP 2025accepted

Recently, Mamba-based methods have gained popularity in medical image segmentation due to their ability to model long-range dependencies with linear computational complexity. However, current segmentation methods often face challenges such as low contrast, blurred boundaries, and unclear backgrounds…

Cited by 0SourceScholar
2025

SKE-MSA: Enhancing Representation Learning with VAD Lexicon for Multimodal Sentiment Analysis

ICASSP 2025accepted

Most existing Multimodal Sentiment Analysis (MSA) models that are enhanced with external sentiment knowledge primarily focus on integrating this knowledge during the multimodal fusion stage, while overlooking its potential benefits during the multimodal representation learning process. In this study…

Cited by 0SourceScholar
2025

Spiking Vision Transformer with Saccadic Attention

ICLR 2025poster

The combination of Spiking Neural Networks (SNNs) and Vision Transformers (ViTs) holds potential for achieving both energy efficiency and high performance, particularly suitable for edge vision applications. However, a significant performance gap still exists between SNN-based ViTs and their ANN cou…

Cited by 1SourcePDFScholar
2025

Strengthen Out-of-Distribution Detection Capability with Progressive Self-Knowledge Distillation

ICML 2025poster

Out-of-distribution (OOD) detection aims to ensure AI system reliability by rejecting inputs outside the training distribution. Recent work shows that memorizing atypical samples during later stages of training can hurt OOD detection, while strategies for forgetting them show promising improvements.…

Cited by 0SourcePDFScholar
2025

SyncGaussian: Stable 3D Gaussian-Based Talking Head Generation with Enhanced Lip Sync via Discriminative Speech Features

IJCAI 2025

Generating high-fidelity talking heads that maintain stable head poses and achieve robust lip sync remains a significant challenge. Although methods based on 3D Gaussian Splatting (3DGS) offer a promising solution via point-based deformation, they suffer from inconsistent head dynamics and mismatche

Cited by 0SourcePDFScholar
2025

Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language Models

NeurIPS 2025poster

In this work, we address the task of table image to LaTeX code generation, with the goal of automating the reconstruction of high-quality, publication-ready tables from visual inputs. A central challenge of this task lies in accurately handling complex tables—those with large sizes, deeply nested st…

Cited by 0SourceScholar
2025

Tactile sensing soft fingertip with dual air bag structure for an anthropomorphic robotic hand

IROS 2025

Tactile sensing plays a crucial role to empower robotic hands with improved grasping and manipulating abilities. In this paper, we propose an anthropomorphic robotic hand design with dual air bag sensors integrated soft fingertips to achieve tactile sensing. The air bag sensor is low-cost, easy-to-b

Cited by 0SourceScholar
2025

Testing Conditional Independence with Deep Neural Network Based Binary Expansion Testing (DeepBET)

AISTATS 2025poster

This paper focuses on testing conditional independence between two random variables ($X$ and $Y$) given a set of high-dimensional confounding variables ($Z$). The high dimensionality of these confounding variables presents a challenge, often resulting in inflated type-I errors or insufficient power…

Cited by 0SourceScholar
2025

Three-Dimension Tip Force Perception and Axial Contact Location Identification for Flexible Endoscopy Using Tissue-Compliant Soft Distal Attachment Cap Sensors

ICRA 2025

In endoluminal surgeries, inserting a flexible endo-scope is one of the fundamental procedures. During this process, vision remains the primary feedback, while the perception of tactile magnitude and location is insufficient. This limitation can hinder the clinician's efficiency when navigating the

Cited by 0SourceScholar
2025

Towards Accurate Binary Spiking Neural Networks: Learning with Adaptive Gradient Modulation Mechanism

AAAI 2025technical

Binary Spiking Neural Networks (BSNNs) inherit the event-driven paradigm of SNNs, while also adopting the reduced storage burden of binarization techniques. These distinct advantages grant BSNNs lightweight and energy-efficient characteristics, rendering them ideal for deployment on resource-constra…

2025

Towards Equilibrium: An Instantaneous Probe-and-Rebalance Multimodal Learning Approach

IJCAI 2025

The multimodal imbalance problem has been extensively studied to prevent the undesirable scenario where multimodal performance falls below that of unimodal models. However, existing methods typically assess the strength of modalities and perform learning simultaneously under the imbalanced status. T

2025

UltraFastCrackSeg: A Lightweight Real-Time Crack Segmentation Model with Task-Oriented Pretraining

ICRA 2025

Crack segmentation is pivotal for structural health monitoring, enabling the timely maintenance of critical infrastructure such as bridges and roads. However, existing deep learning models are often too computationally intensive for deployment on resource-constrained devices. To address this limitat

Cited by 0SourcecodeScholar
2025

UnCo: Uncertainty-Driven Collaborative Framework of Large and Small Models for Grounded Multimodal NER

EMNLP 2025

Grounded Multimodal Named Entity Recognition (GMNER) is a new information extraction task. It requires models to extract named entities and ground them to real-world visual objects. Previous methods, relying on domain-specific fine-tuning, struggle with unseen multimodal entities due to limited know

2025

Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media Manipulation

CVPR 2025poster

To tackle the threat of fake news, the task of detecting and grounding multi-modal media manipulation (DGM4) has received increasing attention. However, most state-of-the-art methods fail to explore the fine-grained consistency within local content, usually resulting in an inadequate perception of d…

2025

Unveiling the Spatial-temporal Effective Receptive Fields of Spiking Neural Networks

NeurIPS 2025poster

Spiking Neural Networks (SNNs) demonstrate significant potential for energy-efficient neuromorphic computing through an event-driven paradigm. While training methods and computational models have greatly advanced, SNNs struggle to achieve competitive performance in visual long-sequence modeling task…

Cited by 0SourcecodeScholar
2024

3S-TSE: Efficient Three-Stage Target Speaker Extraction for Real-Time and Low-Resource Applications

ICASSP 2024accepted

Target speaker extraction (TSE) aims to isolate a specific voice from multiple mixed speakers relying on a registerd sample. Since voiceprint features usually vary greatly, current end-to-end neural networks require large model parameters which are computational intensive and impractical for real-ti…

Cited by 0SourceScholar
2024

A Rigid-Flexible Coupling Oscillator for Pneumatic Autonomous Robots

RA-L 2024

The traditional oscillators are limited by their material characteristics and structural configuration, exhibiting a low oscillation frequency and low output power under pressure input. This results in a slow locomotion speed for the robot. To address these limitations, this work presents a compact,

Cited by 1SourceScholar
2024

Adaptive Uncertainty-Based Learning for Text-Based Person Retrieval

AAAI 2024technical

Text-based person retrieval aims at retrieving a specific pedestrian image from a gallery based on textual descriptions. The primary challenge is how to overcome the inherent heterogeneous modality gap in the situation of significant intra-class variation and minimal inter-class variation. Existing…

2024

An Efficient Prototype-Based Clustering Approach for Edge Pruning in Graph Neural Networks to Battle Over-Smoothing

IJCAI 2024poster

Topology augmentation is a popular strategy to address the issue of over-smoothing in graph neural networks (GNNs). To prevent potential distortion of node representations, an essential principle is to enhance the separability between embeddings of nodes from different classes while preserving smoot…

2024

An Expert is Worth One Token: Synergizing Multiple Expert LLMs as Generalist via Expert Token Routing

ACL 2024long

We present Expert-Token-Routing, a unified generalist framework that facilitates seamless integration of multiple expert LLMs. Our framework represents expert LLMs as special expert tokens within the vocabulary of a meta LLM. The meta LLM can route to an expert LLM like generating new tokens. Expert…

2024

Binaural Angular Separation Network

ICASSP 2024accepted

We propose a neural network model that can separate target speech sources from interfering sources at different angular regions using two microphones. The model is trained with simulated room impulse responses (RIRs) using omnidirectional microphones without needing to collect real RIRs. By relying…

Cited by 0SourceScholar
2024

Bounding Box Stability against Feature Dropout Reflects Detector Generalization across Environments

ICLR 2024spotlight

Bounding boxes uniquely characterize object detection, where a good detector gives accurate bounding boxes of categories of interest. However, in the real-world where test ground truths are not provided, it is non-trivial to find out whether bounding boxes are accurate, thus preventing us from asses…

2024

Bridging Gaps: Federated Multi-View Clustering in Heterogeneous Hybrid Views

NeurIPS 2024poster

Recently, federated multi-view clustering (FedMVC) has emerged to explore cluster structures in multi-view data distributed on multiple clients. Many existing approaches tend to assume that clients are isomorphic and all of them belong to either single-view clients or multi-view clients. While these…

2024

CDPNet: Cross-Modal Dual Phases Network for Point Cloud Completion

AAAI 2024technical

Point cloud completion aims at completing shapes from their partial. Most existing methods utilized shape’s priors information for point cloud completion, such as inputting the partial and getting the complete one through an encoder-decoder deep learning structure. However, it is very often to easi…

Cited by 7SourcePDFScholar
2024

CIFAR-10-Warehouse: Broad and More Realistic Testbeds in Model Generalization Analysis

ICLR 2024poster

Analyzing model performance in various unseen environments is a critical research problem in the machine learning community. To study this problem, it is important to construct a testbed with out-of-distribution test sets that have broad coverage of environmental discrepancies. However, existing tes…

Cited by 7SourcePDFScholar
2024

CLGSI: A Multimodal Sentiment Analysis Framework based on Contrastive Learning Guided by Sentiment Intensity

NAACL 2024findings

Recently, contrastive learning has begun to gain popularity in multimodal sentiment analysis (MSA). However, most of existing MSA methods based on contrastive learning lacks more detailed learning of the distribution of sample pairs with different sentiment intensity differences in the contrastive l…

2024

Can Graph Neural Networks Expose Training Data Properties? An Efficient Risk Assessment Approach

NeurIPS 2024poster

Graph neural networks (GNNs) have attracted considerable attention due to their diverse applications. However, the scarcity and quality limitations of graph data present challenges to their training process in practical settings. To facilitate the development of effective GNNs, companies and researc…

2024

Con4m: Context-aware Consistency Learning Framework for Segmented Time Series Classification

NeurIPS 2024poster

Time Series Classification (TSC) encompasses two settings: classifying entire sequences or classifying segmented subsequences. The raw time series for segmented TSC usually contain Multiple classes with Varying Duration of each class (MVD). Therefore, the characteristics of MVD pose unique challenge…

2024

DMNet: Self-comparison Driven Model for Subject-independent Seizure Detection

NeurIPS 2024poster

Automated seizure detection (ASD) using intracranial electroencephalography (iEEG) is critical for effective epilepsy treatment. However, the significant domain shift of iEEG signals across subjects poses a major challenge, limiting their applicability in real-world clinical scenarios. In this paper…

Cited by 1SourcePDFScholar
2024

DWLR: Domain Adaptation under Label Shift for Wearable Sensor

IJCAI 2024poster

Wearable sensors play a crucial role in real-world scenarios, such as human activity recognition, sleep monitoring and electrocardiogram monitoring. However, deploying classifiers on them is challenged by distribution shifts across users and devices. Unsupervised domain adaptation (UDA) is proposed…

Cited by 0SourcePDFScholar
2024

Design and Kinematic Modeling of a Pneumatic Soft Bellow-Type Wrist

RA-L 2024

In this letter, we aim to develop an efficient and accurate kinematic model of a soft wrist that consists of pneumatic bellows configured in parallel, toward dexterous manipulation in confined space. The challenge arises from the distributed nature and deformation-dependency of the generated pneumat

Cited by 8SourceScholar
2024

Dialogues Are Not Just Text: Modeling Cognition for Dialogue Coherence Evaluation

AAAI 2024technical

The generation of logically coherent dialogues by humans relies on underlying cognitive abilities. Based on this, we redefine the dialogue coherence evaluation process, combining cognitive judgment with the basic text to achieve a more human-like evaluation. We propose a novel dialogue evaluation fr…

2024

Diffusion Models as Optimizers for Efficient Planning in Offline RL

ECCV 2024poster

"Diffusion models have shown strong competitiveness in offline reinforcement learning tasks by formulating decision-making as sequential generation. However, the practicality of these methods is limited due to the lengthy inference processes they require. In this paper, we address this problem by de…

2024

Digital Task-Oriented Communication with Hardware-Limited Task-Based Quantization

ICASSP 2024accepted

Task-oriented communication exploits the task to improve communication efficiency. Most existing works on task-oriented communication transmit analog signals without quantization, which limits its application in digital communication systems. This paper studies digital task-oriented communication sy…

Cited by 0SourceScholar
2024

Disentangling Domain and General Representations for Time Series Classification

IJCAI 2024poster

Modeling time series data has become a very at tractive research topic due to its wide application, such as human activity recognition, financial forecasting and sensor-based automatic system monitoring. Recently deep learning models have shown great advances in modeling the time series data but the…

2024

Efficient Adaptation of Pre-trained Vision Transformer via Householder Transformation

NeurIPS 2024poster

A common strategy for Parameter-Efficient Fine-Tuning (PEFT) of pre-trained Vision Transformers (ViTs) involves adapting the model to downstream tasks by learning a low-rank adaptation matrix. This matrix is decomposed into a product of down-projection and up-projection matrices, with the bottleneck…

Cited by 1SourcePDFScholar
2024

Enhancing Learning-Based Binary Code Similarity Detection Model through Adversarial Training with Multiple Function Variants

EMNLP 2024finding

Compared to identifying binary versions of the same function under different compilation options, existing Learning-Based Binary Code Similarity Detection (LB-BCSD) methods exhibit lower accuracy in recognizing functions with the same functionality but different implementations. To address this issu…

Cited by 0SourcePDFScholar
2024

Ensemble Diversity Facilitates Adversarial Transferability

CVPR 2024poster

With the advent of ensemble-based attacks the transferability of generated adversarial examples is elevated by a noticeable margin despite many methods only employing superficial integration yet ignoring the diversity between ensemble models. However most of them compromise the latent value of the d…

2024

Exploring Correlations of Self-Supervised Tasks for Graphs

ICML 2024poster

Graph self-supervised learning has sparked a research surge in training informative representations without accessing any labeled data. However, our understanding of graph self-supervised learning remains limited, and the inherent relationships between various self-supervised tasks are still unexplo…

2024

Extracting Training Data from Molecular Pre-trained Models

NeurIPS 2024poster

Graph Neural Networks (GNNs) have significantly advanced the field of drug discovery, enhancing the speed and efficiency of molecular identification. However, training these GNNs demands vast amounts of molecular data, which has spurred the emergence of collaborative model-sharing initiatives. These…

2024

Facilitating Multimodal Classification via Dynamically Learning Modality Gap

NeurIPS 2024poster

Multimodal learning falls into the trap of the optimization dilemma due to the modality imbalance phenomenon, leading to unsatisfactory performance in real applications. A core reason for modality imbalance is that the models of each modality converge at different rates. Many attempts naturally focu…

2024

Fast Updating Truncated SVD for Representation Learning with Sparse Matrices

ICLR 2024poster

Updating truncated Singular Value Decomposition (SVD) has extensive applications in representation learning. The continuous evolution of massive-scaled data matrices in practical scenarios highlights the importance of aligning SVD-based models with fast-paced updates. Recent methods for updating tru…

Cited by 2SourcePDFScholar
2024

Fine-Tuning Graph Neural Networks by Preserving Graph Generative Patterns

AAAI 2024technical

Recently, the paradigm of pre-training and fine-tuning graph neural networks has been intensively studied and applied in a wide range of graph mining tasks. Its success is generally attributed to the structural consistency between pre-training and downstream datasets, which, however, does not hold…

2024

Generative Enzyme Design Guided by Functionally Important Sites and Small-Molecule Substrates

ICML 2024poster

Enzymes are genetically encoded biocatalysts capable of accelerating chemical reactions. How can we automatically design functional enzymes? In this paper, we propose EnzyGen, an approach to learn a unified model to design enzymes across all functional families. Our key idea is to generate an enzyme…

2024

Goal-Reaching Policy Learning from Non-Expert Observations via Effective Subgoal Guidance

CoRL 2024poster

In this work, we address the challenging problem of long-horizon goal-reaching policy learning from non-expert, action-free observation data. Unlike fully labeled expert data, our data is more accessible and avoids the costly process of action labeling. Additionally, compared to online learning, whi…

Cited by 1SourcecodeScholar
2024

Hardware-Based Time Synchronization for a Multi-Sensor System

IROS 2024poster

Accurate time synchronization is crucial for multisensor fusion, which is widely used in mobile robotics, autonomous driving, and virtual reality. Despite many advancements, precise multi-sensor synchronization is still challenging due to the sensors’ internal characteristics, data filtering, disjoi…

Cited by 0SourceScholar
2024

Hierarchical Speaker Representation for Target Speaker Extraction

ICASSP 2024accepted

Target speaker extraction aims to isolate a specific speaker’s voice from a composite of multiple sound sources, guided by an enrollment utterance or called anchor. Current methods predominantly derive speaker embeddings from the anchor and integrate them into the separation network to separate the…

Cited by 0SourceScholar
2024

InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks

ICML 2024poster

In this paper, we introduce InfiAgent-DABench, the first benchmark specifically designed to evaluate LLM-based agents on data analysis tasks. Agents need to solve these tasks end-to-end by interacting with an execution environment. This benchmark contains DAEval, a dataset consisting of 603 data ana…

2024

Low-Rank Rescaled Vision Transformer Fine-Tuning: A Residual Design Approach

CVPR 2024poster

Parameter-efficient fine-tuning for pre-trained Vision Transformers aims to adeptly tailor a model to downstream tasks by learning a minimal set of new adaptation parameters while preserving the frozen majority of pre-trained parameters. Striking a balance between retaining the generalizable represe…

2024

MA-Stereo: Real-Time Stereo Matching via Multi-Scale Attention Fusion and Spatial Error-Aware Refinement

RA-L 2024

Stereo matching is a fundamental task in computer vision. Real-time stereo matching has recently shown great potential in robotics and autonomous driving applications. However, the existing cost aggregation in real-time stereo matching suffers from accuracy limitations in ill-posed regions. Furtherm

Cited by 4SourceScholar
2024

Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

EMNLP 2024finding

Large language models (LLMs) have demonstrated impressive reasoning capabilities, particularly in textual mathematical problem-solving. However, existing open-source image instruction fine-tuning datasets, containing limited question-answer pairs per image, do not fully exploit visual information to…

2024

Measuring Bargaining Abilities of LLMs: A Benchmark and A Buyer-Enhancement Method

ACL 2024findings

Bargaining is an important and unique part of negotiation between humans. As LLM-driven agents learn to negotiate and act like real humans, how to evaluate agents’ bargaining abilities remains an open problem.For the first time, we formally described the Bargaining task as an asymmetric incomplete i…

2024

Measuring Task Similarity and Its Implication in Fine-Tuning Graph Neural Networks

AAAI 2024technical

The paradigm of pre-training and fine-tuning graph neural networks has attracted wide research attention. In previous studies, the pre-trained models are viewed as universally versatile, and applied for a diverse range of downstream tasks. In many situations, however, this practice results in limite…

2024

MorphVAE: Advancing Morphological Design of Voxel-Based Soft Robots with Variational Autoencoders

AAAI 2024technical

Soft robot design is an intricate field with unique challenges due to its complex and vast search space. In the past literature, evolutionary computation algorithms, including novel probabilistic generative models (PGMs), have shown potential in this realm. However, these methods are sample ineffici…

2024

Noise-Aware Image Captioning with Progressively Exploring Mismatched Words

AAAI 2024technical

Image captioning aims to automatically generate captions for images by learning a cross-modal generator from vision to language. The large amount of image-text pairs required for training is usually sourced from the internet due to the manual cost, which brings the noise with mismatched relevance th…

2024

Origami Actuator with Tunable Limiting Layer for Reconfigurable Soft Robotic Grasping

IROS 2024poster

This paper presents a soft actuator inspired by origami and a tunable strain limiting layer, which is proposed for reconfigurable soft robotic grasping. Main structure of the actuator is based on Miura origami which generates extension under pressurized air while a limiting layer with tunable length…

Cited by 0SourceScholar
2024

PowerPM: Foundation Model for Power Systems

NeurIPS 2024poster

The proliferation of abundant electricity time series (ETS) data presents numerous opportunities for various applications within power systems, including demand-side management, grid stability, and consumer behavior analysis. Deep learning models have advanced ETS modeling by effectively capturing s…

2024

Privacy-Aware Joint Source-Channel Coding For Image Transmission Based On Disentangled Information Bottleneck

ICASSP 2024accepted

Current privacy-aware joint source-channel coding (JSCC) works aim at avoiding private information transmission by adversarially training the JSCC encoder and decoder under specific signal-to-noise ratios (SNRs) of eavesdroppers. However, these approaches incur additional computational and storage r…

Cited by 0SourceScholar
2024

Region-aware Distribution Contrast: A Novel Approach to Multi-Task Partially Supervised Learning

ECCV 2024poster

"In this study, we address the intricate challenge of multi-task dense prediction, encompassing tasks such as semantic segmentation, depth estimation, and surface normal estimation, particularly when dealing with partially annotated data (MTPSL). The complexity arises from the absence of complete ta…

2024

STREAMVC: Real-Time Low-Latency Voice Conversion

ICASSP 2024accepted

We present StreamVC, a streaming voice conversion solution that preserves the content and prosody of any source speech while matching the voice timbre from any target speech. Unlike previous approaches, StreamVC produces the resulting waveform at low latency from the input signal even on a mobile pl…

Cited by 0SourceScholar
2024

ScanERU: Interactive 3D Visual Grounding Based on Embodied Reference Understanding

AAAI 2024technical

Aiming to link natural language descriptions to specific regions in a 3D scene represented as 3D point clouds, 3D visual grounding is a very fundamental task for human-robot interaction. The recognition errors can significantly impact the overall accuracy and then degrade the operation of AI systems…

2024

Self-Sensing Origami-Inspired Soft Twisting Actuators and Its Application in Soft Robots

RA-L 2024

The good compliance of soft robots provides a reliable safety environment for human-robot interaction; however, it also creates challenges for adding sensors to soft robots. In this letter, we propose a self-sensing origami-inspired soft twisting actuator. The actuator is designed based on the struc

Cited by 13SourceScholar
2024

Spike-based Neuromorphic Model for Sound Source Localization

NeurIPS 2024poster

Biological systems possess remarkable sound source localization (SSL) capabilities that are critical for survival in complex environments. This ability arises from the collaboration between the auditory periphery, which encodes sound as precisely timed spikes, and the auditory cortex, which performs…

Cited by 6SourcePDFScholar
2024

TAI++: Text as Image for Multi-Label Image Classification by Co-Learning Transferable Prompt

IJCAI 2024poster

The recent introduction of prompt tuning based on pre-trained vision-language models has dramatically improved the performance of multi-label image classification. However, some existing strategies that have been explored still have drawbacks, i.e., either exploiting massive labeled visual data at a…

2024

Tabular Insights, Visual Impacts: Transferring Expertise from Tables to Images

ICML 2024spotlight

Transferring knowledge across diverse data modalities is receiving increasing attention in machine learning. This paper tackles the task of leveraging expert-derived, yet expensive, tabular data to enhance image-based predictions when tabular data is unavailable during inference. The primary challen…

Cited by 2SourcePDFScholar
2024

Task-agnostic Distillation of Encoder-Decoder Language Models

COLING 2024main

Finetuning pretrained language models (LMs) have enabled appealing performance on a diverse array of tasks. The intriguing task-agnostic property has driven a shifted focus from task-specific to task-agnostic distillation of LMs. While task-agnostic, compute-efficient, performance-preserved LMs can…

Cited by 2SourcePDFScholar
2024

Towards Fair Graph Federated Learning via Incentive Mechanisms

AAAI 2024technical

Graph federated learning (FL) has emerged as a pivotal paradigm enabling multiple agents to collaboratively train a graph model while preserving local data privacy. Yet, current efforts overlook a key issue: agents are self-interested and would hesitant to share data without fair and satisfactory i…

2024

United We Stand, Divided We Fall: Fingerprinting Deep Neural Networks via Adversarial Trajectories

NeurIPS 2024poster

In recent years, deep neural networks (DNNs) have witnessed extensive applications, and protecting their intellectual property (IP) is thus crucial. As a non-invasive way for model IP protection, model fingerprinting has become popular. However, existing single-point based fingerprinting methods are…

Cited by 0SourcePDFScholar
2024

Unsupervised Domain Adaptative Temporal Sentence Localization with Mutual Information Maximization

AAAI 2024technical

Temporal sentence localization (TSL) aims to localize a target segment in a video according to a given sentence query. Though respectable works have made decent achievements in this task, they severely rely on abundant yet expensive manual annotations for training. Moreover, these trained data-depen…

Cited by 7SourcePDFScholar
2024

Untethered Soft Rolling Robot Based on Pneumatic-Tendon Coupled Actuation

RA-L 2024

The Soft Rolling Robot (SRR) excels in adaptability across diverse natural terrains, demonstrating significant flexibility, environmental interaction capabilities, and impact resistance. However, the development of Untethered Soft Rolling Robots (USRR) faces challenges in locomotion speed and contro

Cited by 4SourceScholar
2024

Unveiling Latent Causal Rules: A Temporal Point Process Approach for Abnormal Event Explanation

AISTATS 2024poster

In high-stakes systems such as healthcare, it is critical to understand the causal reasons behind unusual events, such as sudden changes in patient’s health. Unveiling the causal reasons helps with quick diagnoses and precise treatment planning. In this paper, we propose an automated method for unco…

2024

Visual Servo Control of a Conceptual Magnetically Anchored and Guided Flexible Endoscope

IROS 2024poster

This paper presents a conceptual magnetically anchored and guided flexible endoscope for minimally invasive surgery (MIS). Leveraging both the magnetic coupling between the external and internal permanent magnets and the bending of a flexible joint, the endoscope offers improved maneuver-ability and…

Cited by 0SourceScholar
2024

Visual-Tactile Perception Based Control Strategy for Complex Robot Peg-in-Hole Process via Topological and Geometric Reasoning

RA-L 2024

Peg-hole-insertion processes of diverse shapes are typical contact-rich tasks, which need the accurate representation of object's shape, pose, and peg-hole contact states. The visual-tactile sensor can perceive the relative moving trend between the gripper and the grasped object, which could be appl

Cited by 10SourceScholar
2024

Weakly-Supervised Mirror Detection via Scribble Annotations

AAAI 2024technical

Mirror detection is of great significance for avoiding false recognition of reflected objects in computer vision tasks. Existing mirror detection frameworks usually follow a supervised setting, which relies heavily on high quality labels and suffers from poor generalization. To resolve this, we inst…

2023

A Novel Obstacle-Avoidance Solution With Non-Iterative Neural Controller for Joint-Constrained Redundant Manipulators

IROS 2023poster

Obstacle avoidance (OA) and joint-limit avoidance (JLA) are essential for redundant manipulators to ensure safe and reliable robotic operations. One solution to OA and JLA is to incorporate the involved constraints into a quadratic programming (QP), by solving which OA and JLA can be achieved. There…

Cited by 1SourceScholar
2023

Adversarial Object Rearrangement in Constrained Environments with Heterogeneous Graph Neural Networks

IROS 2023poster

Adversarial object rearrangement in the real world (e.g., previously unseen or oversized items in kitchens and stores) could benefit from understanding task scenes, which inherently entail heterogeneous components such as current objects, goal objects, and environmental constraints. The semantic rel…

Cited by 3SourceScholar
2023

Better with Less: A Data-Active Perspective on Pre-Training Graph Neural Networks

NeurIPS 2023poster

Pre-training on graph neural networks (GNNs) aims to learn transferable knowledge for downstream tasks with unlabeled data, and it has recently become an active research area. The success of graph pre-training models is often attributed to the massive amount of input data. In this paper, however, we…