← Search

Wei Wang.

463 accepted papers

2026

$\alpha$-DPO: Robust Preference Alignment for Diffusion Models via $\alpha$ Divergence

ICLR 2026poster

Diffusion models have demonstrated remarkable success in high-fidelity image generation, yet aligning them with human preferences remains challenging. Direct Preference Optimization (DPO) offers a promising framework, but its effectiveness is critically hindered by noisy data arising from mislabeled…

Cited by 0SourceScholar
2026

ARLArena: Demystifying Policy Gradient Stability in Agentic Reinforcement Learning

ICML 2026poster

Agentic reinforcement learning (ARL) has rapidly gained attention as a promising paradigm for training agents to solve complex, multi-step interactive tasks. In this paper, we first propose $\textbf{ARLArena}$, a fair and systematic analysis framework that encompasses a broad spectrum of ARL algorit…

Cited by 0SourceScholar
2026

Accessible, Realistic, and Fair Evaluation of Positive-Unlabeled Learning Algorithms

ICLR 2026poster

Positive-unlabeled (PU) learning is a weakly supervised binary classification problem, in which the goal is to learn a binary classifier from only positive and unlabeled data, without access to negative data. In recent years, many PU learning algorithms have been developed to improve model performan…

Cited by 0SourceScholar
2026

Adaptive Hallucination Alleviation in Multimodal Large Language Models: From Strategic Data Selection to Severity-Guided Training

AAAI 2026technical

Multimodal Large Language Models (MLLMs) have recently achieved strong performance across a variety of multimodal tasks. However, they still suffer from various forms of hallucination, which hinder their practical deployment. Prior approaches often struggle to efficiently construct high-quality hall

Cited by 0SourcePDFScholar
2026

Agentic Proposing: Enhancing Large language Model Reasoning via Compositional Skill Synthesis

ICML 2026poster

Advancing complex reasoning in large language models relies on high-quality, verifiable datasets, yet human annotation remains cost-prohibitive and difficult to scale. Current synthesis paradigms often face a recurring trade-off: maintaining structural validity typically restricts problem complexity…

Cited by 0SourceScholar
2026

All-in-One Slider for Attribute Manipulation in Diffusion Models

CVPR 2026

Text-to-image (T2I) diffusion models have made significant strides in generating high-quality images. However, progressively manipulating certain attributes of generated images to meet the desired user expectations remains challenging, particularly for content with rich details, such as human faces.

Cited by 0SourcecodeScholar
2026

Ambiguity-aware Truncated Flow Matching for Ambiguous Medical Image Segmentation

AAAI 2026technical

A simultaneous enhancement of accuracy and diversity of predictions remains a challenge in ambiguous medical image segmentation (AMIS) due to the inherent trade-offs. While truncated diffusion probabilistic models (TDPMs) hold strong potential with a paradigm optimization, existing TDPMs suffer from

Cited by 0SourcePDFScholar
2026

BRIDGING THE GAP: TRANSFORMING NATURAL LANGUAGE QUESTIONS INTO SQL QUERIES VIA ABSTRACT QUERY PATTERN AND CONTEXTUAL SCHEMA MARKUP

ICASSP 2026poster

Large language models have demonstrated excellent performance in many tasks, including Text-to-SQL, due to their powerful in-context learning capabilities. They are becoming the mainstream approach for Text-to-SQL. However, these methods still have a significant gap compared to human performance, es…

Cited by 0SourcePDFScholar
2026

Beyond Magic Words: Sharpness-Aware Prompt Evolving for Robust Large Language Models with TARE

ICLR 2026poster

The performance of Large Language Models (LLMs) hinges on carefully engineered prompts. However, prevailing prompt optimization methods, ranging from heuristic edits and reinforcement learning to evolutionary search, primarily target point-wise accuracy. They seldom enforce paraphrase invariance or…

Cited by 0SourceScholar
2026

Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training

ICLR 2026poster

Reinforcement fine-tuning (RFT) often suffers from reward over-optimization, where a policy model hacks the reward signals to achieve high scores while producing low-quality outputs. Our theoretical analysis shows that the key lies in reward misspecification at the high-reward tail: the inability to…

Cited by 0SourcecodeScholar
2026

Concurrent-Learning Based Relative Localization in Shape Formation of Robot Swarms (I)

ICRA 2026poster

In this article, we address the shape formation problem for massive robot swarms in environments where external localization systems are unavailable.Achieving this task effectively with solely onboard measurements is still scarcely explored and faces some practical challenges.To solve this challengi…

Cited by 0Scholar
2026

ConstructAI: From Real-Time Safety Insight to Skill Growth in Deployed Construction AI Systems

AAAI 2026technical

Ensuring safety in power grid construction remains a critical yet challenging task, as existing monitoring approaches often lack scalability, timeliness, and adaptability to diverse on-site conditions. To address these limitations, we present ConstructAI, a deployed AI-driven safety management syste

Cited by 0SourcePDFScholar
2026

CoordAR: One-Reference 6D Pose Estimation of Novel Objects via Autoregressive Coordinate Map Generation

AAAI 2026technical

Object 6D pose estimation, a crucial task for robotics and augmented reality applications, becomes particularly challenging when dealing with novel objects whose 3D models are not readily available. To reduce dependency on 3D models, recent studies have explored one-reference-based pose estimation,

Cited by 0SourcePDFScholar
2026

DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing Model

ICML 2026poster

Significant progress has been made in the field of Instruction-based Image Editing Models (IIEMs). However, while these models demonstrate plausible adherence to instructions and strong reasoning ability on current benchmarks, their ability to edit small objects remains underexplored, despite its im…

Cited by 0SourceScholar
2026

Decision-Driven Orthogonal Learning with Complementary Feature Mining for Robust Synthetic Image Detection

AAAI 2026technical

The widespread and inconsistent compression applied by Online Social Networks severely degrades the performance of synthetic image detectors. We attribute this degradation to two main issues: 1) the model confuses forgery artifacts with compression artifacts, and 2) compression erodes crucial discri

Cited by 0SourcePDFScholar
2026

Decoupled Residual Denoising Diffusion Models for Unified and Data Efficient Image-to-Image Translation

CVPR 2026

We propose Decoupled Residual Denoising Diffusion models (DRDD) for unified and data-efficient image-to-image (I2I) translation. While diffusion models have advanced I2I translation in terms of quality and diversity, we uncover a previously under-explored property in diffusion models. Crucially, bey

Cited by 0SourcecodeScholar
2026

Design and Control of Centimeter-Scale Reconfigurable Aquatic Modular Robots

RA-L 2026

This letter presents the Centimeter-scale Autonomous Reconfigurable Platform (CARP), a low-cost, self-assembling Autonomous Surface Vehicle (ASV) designed for modular and cooperative maritime operations. CARP features a 90 mm square top section with four independent active latching mechanisms, enabl

Cited by 0SourceScholar
2026

DiGraphHal-Bench: Evaluating Multimodal Large Language Models on Complex Directed Graphs

CVPR 2026

While prior research on Multimodal Large Language Model (MLLM) hallucinations has primarily examined cross-modal inconsistencies in natural images, hallucination over complex graph structures remains underexplored.Concurrently, there is a lack of robust evaluation for fine-grained reasoning integrat

Cited by 0SourcecodeScholar
2026

Diffusion Reconstruction-based Data Likelihood Estimation for Core-Set Selection

AAAI 2026technical

Existing core-set selection methods predominantly rely on heuristic scoring signals such as training dynamics or model uncertainty, lacking explicit modeling of data likelihood. This omission may hinder the constructed subset from capturing subtle yet critical distributional structures that underpin

Cited by 0SourcePDFScholar
2026

Elucidating the Design Space of Arbitrary-Noise-Based Diffusion Models

CVPR 2026

Although EDM aims to unify the design space of diffusion models, its reliance on fixed Gaussian noise prevents it from explaining emerging flow-based methods that diffuse arbitrary noise. Moreover, our study reveals that EDM's forcible injection of Gaussian noise has adverse effects on image restora

Cited by 0SourcecodeScholar
2026

Emotion and Intention Guided Multi-Modal Learning for Sticker Response Selection

AAAI 2026technical

Stickers are widely used in online communication to convey emotions and implicit intentions. The Sticker Response Selection (SRS) task aims to select the most contextually appropriate sticker based on the dialogue. However, existing methods typically rely on semantic matching and model emotional and

Cited by 0SourcePDFScholar
2026

FedCure: Mitigating Participation Bias in Semi-Asynchronous Federated Learning with Non-IID Data

AAAI 2026technical

While semi-asynchronous federated learning (SAFL) combines the efficiency of synchronous training with the flexibility of asynchronous updates, it inherently suffers from participation bias, which is further exacerbated by non-IID data distributions. More importantly, hierarchical architecture shift

Cited by 0SourcePDFScholar
2026

Frequency-Aligned Cross-Modal Learning with Top-K Wavelet Fusion and Dynamic Expert Routing for Enhanced Retinal Disease Diagnosis

AAAI 2026technical

Multimodal fusion of color fundus photography (CFP) and optical coherence tomography (OCT) B-scan images has demonstrated superior diagnostic potential for retinal diseases compared to single-modality approaches. However, existing fusion paradigms - whether through naive concatenation or attention m

Cited by 0SourcePDFScholar
2026

From Basis to Basis: Gaussian Particle Representation for Interpretable PDE Operators

ICML 2026poster

Learning PDE dynamics for fluids increasingly relies on neural operators and Transformer-based models, yet these approaches often lack interpretability and struggle with localized, high-frequency structures while incurring quadratic cost in spatial samples. We propose to represent fields with a \emp…

Cited by 0SourceScholar
2026

From Detection to Association: Learning Discriminative Object Embeddings for Multi-Object Tracking

CVPR 2026

End-to-end multi-object tracking (MOT) methods have recently achieved remarkable progress by unifying detection and association within a single framework. Despite their strong detection performance, these methods suffer from relatively low association accuracy. Through detailed analysis, we observe

Cited by 0SourcecodeScholar
2026

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling

AAAI 2026technical

Recent advancements in improving the reasoning capabilities of Large Language Models have underscored the efficacy of Process Reward Models (PRMs) in addressing intermediate errors through structured feedback mechanisms. This study analyzes PRMs from multiple perspectives, including training methodo

Cited by 0SourcePDFScholar
2026

GDP: Enhancing End-To-End Autonomous Driving with Goal-Driven Planner

ICRA 2026poster

End-to-end (E2E) autonomous driving has emerged as a promising paradigm with the pervasive power of model architectures and the availability of large-scale driving datasets. Despite tremendous efforts in recent research, most E2E driving frameworks rely on rather general driving commands, such as "G…

Cited by 0Scholar
2026

GradPower: Powering Gradients for Faster Language Model Pre-Training

ICML 2026poster

We propose **GradPower**, a lightweight gradient-transformation technique for accelerating language model pre-training. Given a gradient vector $\boldsymbol{g}=(g\_{i})\_{i}$, GradPower first applies the elementwise `sign-power` transformation: $ \varphi_p(\boldsymbol{g}) = \left({\rm sign}(g\_i)|g\…

Cited by 0SourceScholar
2026

HEARTS: Benchmarking LLM Reasoning on Health Time Series

ICML 2026poster

The rise of large language models (LLMs) has shifted time series analysis from narrow analytics to general-purpose reasoning. Yet, existing benchmarks cover only a small set of health time series modalities and tasks, failing to reflect the diverse domains and extensive temporal dependencies inheren…

Cited by 0SourceScholar
2026

Harnessing Spectrum Video for Subject-Level Few-Shot and Cross-Montage EEG Generalization

ICML 2026poster

Existing EEG models are limited by electrode heterogeneity and rigid "channel-first" architectures that treat sensors as independent features. We propose Brain Signal Rendering (BSR), which reinterprets EEG as a physical projection of neural activity and transforms raw signals into geometry-aware Sp…

Cited by 0SourceScholar
2026

ICLR: Inter-Chrominance and Luminance Interaction for Natural Color Restoration in Low-Light Image Enhancement

AAAI 2026technical

Low-Light Image Enhancement (LLIE) task aims at improving contrast while restoring details and textures for images captured in low-light conditions. HVI color space has made significant progress in this task by enabling precise decoupling of chrominance and luminance. However, for the interaction of

Cited by 0SourcePDFScholar
2026

KnowLCP: Knowledge Augmented Lane Change Prediction for Autonomous Driving

AAAI 2026technical

Lane change prediction, encompassing both intention recognition and trajectory forecasting, is essential for the safe operation of autonomous vehicles in mixed-traffic environments. Existing models predominantly follow a data-driven paradigm, learning directly from historical vehicle states through

Cited by 0SourcePDFScholar
2026

Learning from Label Proportions via Proportional Value Classification

ICLR 2026poster

Learning from Label Proportions~(LLP) aims to use bags of instances associated with the proportions of each label within the bag to learn an instance-level classifier. Proportion matching is a widely used strategy that aligns the average model outputs of all instances in a bag with the label proport…

Cited by 0SourcecodeScholar
2026

M2I2: Learning Efficient Multi-Agent Communication via Masked State Modeling and Intention Inference

AAAI 2026technical

Communication is essential in coordinating the behaviors of multiple agents. However, existing methods primarily emphasize content, timing, and partners for information sharing, often neglecting the critical aspect of integrating shared information. This gap can significantly impact agents

Cited by 0SourcePDFScholar
2026

MEASURING PROSODY DIVERSITY IN ZERO-SHOT TTS: A NEW METRIC, BENCHMARK, AND EXPLORATION

ICASSP 2026poster

Prosody diversity is essential for achieving naturalness and expressiveness in zero-shot text-to-speech (TTS). However, frequently used acoustic metrics capture only partial views of prosodic variation and correlate poorly with human perception, leaving the problem of reliably quantifying prosody di…

Cited by 10SourcePDFScholar
2026

MPA: Multimodal Prototype Augmentation for Few-Shot Learning

AAAI 2026technical

Recently, Few-shot Learning (FSL) has become a popular task that aims to recognize new classes from only a few labeled examples and has been widely applied in fields such as natural science, remote sensing, and medical images. However, most existing methods focus only on the visual modality and comp

Cited by 0SourcePDFScholar
2026

MambaOVSR: Multiscale Fusion with Global Motion Modeling for Chinese Opera Video Super-Resolution

AAAI 2026technical

Chinese opera is celebrated for preserving classical art. However, early filming equipment limitations have degraded videos of last-century performances by renowned artists (e.g., low frame rates and resolution), hindering archival efforts. Although space-time video super-resolution (STVSR) has adva

Cited by 0SourcePDFScholar
2026

Mind the Gap: The Divergence Between Human and LLM-Generated Tasks

AAAI 2026technical

Humans constantly generate a diverse range of tasks guided by internal motivations. While generative agents powered by large language models (LLMs) aim to simulate this complex behavior, it remains uncertain whether they operate on similar cognitive principles. To address this, we conducted a task-g

Cited by 0SourcePDFScholar
2026

Motion-Aware Animatable Gaussian Avatars Deblurring

CVPR 2026

The creation of 3D human avatars from multi-view videos is a significant yet challenging task in computer vision. However, existing techniques rely on high-quality, sharp images as input, which are often impractical to obtain in real-world scenarios due to variations in human motion speed and intens

Cited by 0SourcecodeScholar
2026

Noisy-Pair Robust Representation Alignment for Positive-Unlabeled Learning

ICLR 2026poster

Positive-Unlabeled (PU) learning aims to train a binary classifier (positive vs. negative) where only limited positive data and abundant unlabeled data are available. While widely applicable, state-of-the-art PU learning methods substantially underperform their supervised counterparts on complex dat…

Cited by 0SourcecodeScholar
2026

PU-BENCH: A UNIFIED BENCHMARK FOR RIGOROUS AND REPRODUCIBLE PU LEARNING

ICLR 2026poster

Positive-Unlabeled (PU) learning, a challenging paradigm for training binary classifiers from only positive and unlabeled samples, is fundamental to many applications. While numerous PU learning methods have been proposed, the research is systematically hindered by the lack of a standardized and com…

Cited by 0SourcecodeScholar
2026

Physics-Consistent Diffusion for Efficient Fluid Super-Resolution via Multiscale Residual Correction

CVPR 2026

Existing image SR and generic diffusion models transfer poorly to fluid SR: they are sampling-intensive, ignore physical constraints, and often yield spectral mismatch and spurious divergence. We address fluid super-resolution (SR) with **ReMD** (**Re**sidual-**M**ultigrid **D**iffusion), a physics-

Cited by 0SourcecodeScholar
2026

Position: Beyond Prediction: Toward Verifiable Physiological Waveform Reasoning with Foundation Models and Agentic LLMs

ICML 2026poster

Physiological waveforms (e.g., ECG, PPG, EEG) encode clinically meaningful information in fine-grained morphology, precise timing, and cross-channel dynamics, yet most machine learning systems still treat them as generic time series and optimize end-to-end prediction. In this position paper, **we ar…

Cited by 0SourceScholar
2026

Positive-Unlabeled Learning with Extreme Scarcity of Labeled Positives

ICML 2026poster

Positive-Unlabeled (PU) learning is a weakly-supervised paradigm that trains a binary classifier from labeled positive and unlabeled instances. In PU risk estimation, the empirical risk consists of an unlabeled term and a positive term. In this paper, we observe that when labeled positives are scarc…

Cited by 0SourceScholar
2026

Preference Leakage: A Contamination Problem in LLM-as-a-judge

ICLR 2026poster

Large Language Models (LLMs) as judges and LLM-based data synthesis have emerged as two fundamental LLM-driven data annotation methods in model development. While their combination significantly enhances the efficiency of model training and evaluation, little attention has been given to the potentia…

Cited by 0SourcecodeScholar
2026

R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?

ICLR 2026poster

Recent trends in test-time scaling for reasoning models (e.g., OpenAI o1, DeepSeek-R1) have led to remarkable improvements through long Chain-of-Thought (CoT). However, existing benchmarks mainly focus on immediate, single-horizon tasks, failing to adequately evaluate models’ ability to understand a…

Cited by 0SourcecodeScholar
2026

ROI-GSurFisher: Next Best View Selection for Active Gaussian Splatting Via Fisher Information of ROI-Selected Gaussian Surfels

ICRA 2026poster

Next Best View (NBV) selection is critical for achieving high-quality 3D reconstruction in unknown environments. This paper presents an active NBV selection approach tailored for Gaussian Splatting (GS), a widely adopted 3D reconstruction technique that has recently gained significant attention and …

Cited by 0codeScholar
2026

RaLD: Generating High-Resolution 3D Radar Point Clouds with Latent Diffusion

AAAI 2026technical

Millimeter-wave radar offers a promising sensing modality for autonomous systems thanks to its robustness in adverse conditions and low cost. However, its utility is significantly limited by the sparsity and low resolution of radar point clouds, which poses challenges for tasks requiring dense and a

Cited by 0SourcePDFScholar
2026

RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis

AAAI 2026technical

We introduce RacketVision, a novel dataset and benchmark for advancing computer vision in sports analytics, covering table tennis, tennis, and badminton. The dataset is the first to provide large-scale, fine-grained annotations for racket pose alongside traditional ball positions, enabling research

Cited by 0SourcePDFScholar
2026

RepetitionCurse: Measuring and Understanding Router Imbalance in Mixture-of-Experts LLMs under DoS Stress

ICML 2026poster

Mixture-of-Experts architectures have become the standard for efficient LLM scaling, typically employing expert parallelism to distribute experts across devices. However, the absence of explicit load balancing constraints during inference allows adversarial inputs to trigger severe routing concentra…

Cited by 0SourceScholar
2026

Rethinking Consistent Multi-Label Classification under Inexact Supervision

ICLR 2026poster

Partial multi-label learning and complementary multi-label learning are two popular weakly supervised multi-label classification paradigms that aim to alleviate the high annotation costs of collecting precisely annotated multi-label data. In partial multi-label learning, each instance is annotated w…

Cited by 0SourceScholar
2026

Rethinking Token Reduction for Large Vision-Language Models

CVPR 2026

Large Vision-Language Models (LVLMs) excel in visual understanding and reasoning, but the excessive visual tokens lead to high inference costs. Although recent token reduction methods mitigate this issue, they mainly target single-turn Visual Question Answering (VQA), leaving the more practical mult

Cited by 0SourcecodeScholar
2026

Rethinking the Sampling Criteria in Reinforcement Learning for LLM Reasoning: A Competence-Difficulty Alignment Perspective

AAAI 2026technical

The low sampling efficiency during the rollout phase poses a significant challenge to scaling reinforcement learning for large language model reasoning. Existing methods attempt to improve efficiency by scheduling problems based on problem difficulties. However, these approaches suffer from unstabl

Cited by 0SourcePDFScholar
2026

SE-Diff: Simulator and Experience Enhanced Diffusion Model for Comprehensive ECG Generation

ICLR 2026poster

Cardiovascular disease (CVD) is a leading cause of mortality worldwide. Electrocardiograms (ECGs) are the most widely used non-invasive tool for cardiac assessment, yet large, well-annotated ECG corpora are scarce due to cost, privacy, and workflow constraints. Generating ECGs can aid mechanistic un…

Cited by 0SourcecodeScholar
2026

SLCFormer: Spectral-Local Context Transformer with Physics-Grounded Flare Synthesis for Nighttime Flare Removal

AAAI 2026technical

Lens flare is a common nighttime artifact caused by strong light sources scattering within camera lenses, leading to hazy streaks, halos, and glare that degrade visual quality. However, existing methods usually fail to effectively address nonuniform scattered flares, which severely reduces their app

Cited by 0SourcePDFScholar
2026

SLIM: Secure and Efficient Inference for Large Language Models on Untrusted Devices via TEEs

ICML 2026poster

Deploying large language models (LLMs) on untrusted hardware entails a risk of weight extraction, which can lead to unauthorized replication and misuse of the model. A practical approach is to leverage Trusted Execution Environments (TEEs) and protect model security by obfuscating model weights. How…

Cited by 0SourceScholar
2026

SSTODE: Ocean-Atmosphere Physics-Informed Neural ODEs for Sea Surface Temperature Prediction

AAAI 2026technical

Sea Surface Temperature (SST) is crucial for understanding upper-ocean thermal dynamics and ocean-atmosphere interactions, which have profound economic and social impacts. While data-driven models show promise in SST prediction, their black-box nature often limits interpretability and overlooks key

Cited by 0SourcePDFScholar
2026

Semi-LAR: Semi-supervised Contrastive Learning with Linear Attention for Removal of Nighttime Flares

ICML 2026poster

Lens flare removal is challenging due to the large spatial extent of flare artifacts and their entangle-ment with scene structures, while existing meth-ods heavily rely on large-scale paired data. We propose a semi-supervised flare removal frame-work that enables stable learning from unlabeled image…

Cited by 0SourceScholar
2026

Sequential Information Bottleneck Fusion: Towards Robust and Generalizable Multi-Modal Brain Tumor Segmentation

ICLR 2026poster

Brain tumor segmentation in multi-modal MRIs poses significant challenges when one or more modalities are missing. Recent approaches commonly employ parallel fusion strategies; however, these methods often risk losing crucial shared information across modalities, which can degrade segmentation perfo…

Cited by 0SourceScholar
2026

Socratic-Geo: Synthetic Data Generation and Cross-Modal Geometric Reasoning via Multi-Agent Interaction

CVPR 2026

Multimodal Large Language Models (MLLMs) have significantly advanced vision-language understanding. However, even state-of-the-art models struggle with geometric reasoning, revealing a critical bottleneck: the extreme scarcity of high-quality image-text pairs. Human annotation is prohibitively expen

Cited by 0SourceScholar
2026

Spatial Short-Term Fourier Transform Based Single-Channel Single-Fiber 3D Shape Sensing

RA-L 2026

Accurate three-dimensional (3D) shape sensing is vital for continuum robots in minimally invasive surgery. Conventional optical fiber methods depend on multi-fiber or multicore configurations, increasing integration complexity and associated costs. Single-fiber approaches support miniaturization, bu

Cited by 0SourceScholar
2026

Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning

ICLR 2026poster

Large Language Models (LLMs) often struggle with challenging, multi-step reasoning problems due to a fundamental learning gap -- Reinforcement Learning with Verifiable Rewards (RLVR) suffers from sparse rewards when correct solutions are rarely sampled, while Supervised Fine-Tuning (SFT) tends to ov…

Cited by 0SourceScholar
2026

Towards Understanding the Dynamics of Low-Rank Adaptation

ICML 2026poster

Low-Rank Adaptation (LoRA) is a widely used parameter-efficient fine-tuning technique, and previous works have studied the update dynamics of LoRA, showing that updating via the low-rank matrix $\mathbf{A}$ can be viewed as a process within the compressed subspace defined by $\mathbf{A}^{\top} \math…

Cited by 0SourceScholar
2026

Towards an Early Warning System for Ocean Heat Extremes Through AI-Ocean Dynamics Synergy

IJCAI 2026

Ocean heat extremes, including marine heatwaves and the El Ni\~no–Southern Oscillation (ENSO), exert profound impacts on marine ecosystems and socio-economic stability. Establishing robust early warning systems is critical for proactive risk management; however, conventional predictive models often

Cited by 0Scholar
2026

Trustworthy Classification for Complex Social Surveys: A Memory-Enhanced Hierarchical Framework with Calibrated Uncertainty

AAAI 2026technical

Automated classification of complex social survey questionnaires is crucial for large-scale social science research but faces significant reliability challenges due to intricate hierarchical label structures, severe class imbalance, semantic ambiguity, and incomplete data coverage. Conventional clas

Cited by 0SourcePDFScholar
2026

U2E: Uncertainty-Aware Modeling and Uncertainty-Guided Exploration with Deep Ensemble for Quadrupedal Robot

ICRA 2026poster

Reinforcement learning has facilitated agile locomotion in quadrupedal robots. However, most works remain highly dependent on the accuracy of simulation models in describing real-world robot dynamics. Consequently, policy transfer from simulation to hardware is still hindered by the well-known sim-t…

Cited by 0Scholar
2026

Uncertainty-Aware Clarification in LLM Agents with Information Gain

ICML 2026poster

Large Language Model (LLM) agents often operate under underspecified user instructions, where latent uncertainty over user intent leads to erroneous tool actions. To address this challenge, we propose a goal-oriented clarification framework that aligns clarification behavior with ambiguity resolutio…

Cited by 0SourceScholar
2026

What Do Agents Learn from Trajectory-SFT: Semantics or Interfaces?

ICML 2026spotlight

Large language models are increasingly evaluated as interactive agents, yet standard agent benchmarks conflate two qualitatively distinct sources of success: semantic tool-use and interface-specific interaction pattern memorization. Because both mechanisms can yield identical task success on the ori…

Cited by 0SourceScholar
2025

A Data-Efficient Progressive Learning Framework for Robot Scooping Task

ICRA 2025

Robot scooping is a challenging and important task in robotic tool manipulation research due to the complex relationship between the robot, the tool, and target objects/environment. Taking into account different tools, different target objects and varying environments, the required scooping manipula

Cited by 0SourceScholar
2025

A Kinematic Constrained Batch Informed Trees Algorithm With Varied Density Sampling for Mobile Robot Path Planning

RA-L 2025

we proposed a novel Kinematic Batch Informed Trees algorithm (K-BIT*) to solve problems of the low efficiency, poor geometric smoothness and local optimum when conducting path planning for mobile robots. A variable density sampling strategy is designed which can automatically adjust the searching ra

Cited by 3SourceScholar
2025

A Lightweight Sparse Interaction Network for Time Series Forecasting

AAAI 2025technical

Recent work shows that linear models can outperform several transformer models in long-term time-series forecasting (TSF). However, instead of explicitly performing temporal interaction through self-attention, linear models implicitly perform it based on stacked MLP structures, which may be insuffic…

Cited by 0SourcePDFScholar
2025

A Multi-annotated and Multi-modal Dataset for Wide-angle Video Quality Assessment

ICASSP 2025accepted

Wide-angle video is favored for its wide viewing angle and ability to capture a large area of scenery, making it an ideal choice for sports and adventure recording. However, wide-angle video is prone to deformation, exposure and other distortions, resulting in poor video quality and affecting the pe…

Cited by 0SourceScholar
2025

A Survey of Large Language Models in Psychotherapy: Current Landscape and Future Directions

ACL 2025finding

Mental health is increasingly critical in contemporary healthcare, with psychotherapy demanding dynamic, context-sensitive interactions that traditional NLP methods struggle to capture. Large Language Models (LLMs) offer significant potential for addressing this gap due to their ability to handle ex…

Cited by 0SourcePDFScholar
2025

A Trusted Lesion-assessment Network for Interpretable Diagnosis of Coronary Artery Disease in Coronary CT Angiography

AAAI 2025technical

Coronary Artery Disease (CAD) poses a significant threat to cardiovascular patients worldwide, underscoring the critical importance of automated CAD diagnostic technologies in clinical practice. Previous technologies for lesion assessment in Coronary CT Angiography (CCTA) images have been insufficie…

2025

A novel event-based structured light system for high-precision and high-speed depth sensing

IROS 2025

This paper presents a novel event-based depth sensing system with line laser scan. Our main contribution involves both hardware and software improvements to previous state-of-the-art works. The polygon mirror scanner is designed to steer line laser with a constant velocity, which minimizes non-linea

Cited by 0SourceScholar
2025

AI-Enhanced Automatic Design of Efficient Underwater Gliders

ICRA 2025

The development of novel autonomous underwater gliders has been hindered by limited shape diversity, primarily due to the reliance on traditional design tools that depend heavily on manual trial and error. Building an automated design framework is challenging due to the complexities of representing

Cited by 0SourceScholar
2025

Accelerate Parallelizable Reasoning via Parallel Decoding within One Sequence

EMNLP 2025

Recent advances in reasoning models have demonstrated significant improvements in accuracy by employing detailed and comprehensive reasoning processes. However, generating these lengthy reasoning sequences is computationally expensive and time-consuming. To address this inefficiency, we leverage the

2025

AdaDHP: Fine-Grained Fine-Tuning via Dual Hadamard Product and Adaptive Parameter Selection

ACL 2025long

With the continuously expanding parameters, efficiently adapting large language models to downstream tasks is crucial in resource-limited conditions. Many parameter-efficient fine-tuning methods have emerged to address this challenge. However, they lack flexibility, like LoRA requires manually selec…

Cited by 0SourcePDFScholar
2025

Adaptive Median Smoothing: Adversarial Defense for Unlearned Text-to-Image Diffusion Models at Inference Time

ICML 2025poster

Text-to-image (T2I) diffusion models have raised concerns about generating inappropriate content, such as "*nudity*". Despite efforts to erase undesirable concepts through unlearning techniques, these unlearned models remain vulnerable to adversarial inputs that can potentially regenerate such conte…

Cited by 0SourcePDFScholar
2025

Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models

EMNLP 2025

Multi-modal large language models (MLLMs) have achieved remarkable success in fine-grained visual understanding across a range of tasks. However, they often encounter significant challenges due to inadequate alignment for fine-grained knowledge, which restricts their ability to accurately capture lo

Cited by 0SourcePDFScholar
2025

AgentRefine: Enhancing Agent Generalization through Refinement Tuning

ICLR 2025poster

Large Language Model (LLM) based agents have proved their ability to perform complex tasks like humans. However, there is still a large gap between open-sourced LLMs and commercial models like the GPT series. In this paper, we focus on improving the agent generalization capabilities of LLMs via inst…

Cited by 5SourcePDFScholar
2025

An Inflatable Deployable Origami Grasper for Adaptive and High-Load Grasping

IROS 2025

Robotic graspers are essential for enhancing the efficiency and versatility of robots in grasping tasks. In this paper, we propose a novel inflatable deployable origami grasper with a rigid-flexible coupling structure. The proposed grasper can achieve multiple deployment configurations under a singl

Cited by 0SourceScholar
2025

Arbitrary Reading Order Scene Text Spotter with Local Semantics Guidance

AAAI 2025technical

Scene text spotting has attracted the enthusiasm of relative researchers in recent years. Most existing scene text spotters follow the detection-then-recognition paradigm, where the vanilla detection module hardly determines the reading order and leads to failure recognition. After rethinking the au…

Cited by 2SourcePDFScholar
2025

Bayesian Active Learning for Bivariate Causal Discovery

ICML 2025poster

Determining the direction of relationships between variables is fundamental for understanding complex systems across scientific domains. While observational data can uncover relationships between variables, it cannot distinguish between cause and effect without experimental interventions. To effecti…

Cited by 0SourcePDFScholar
2025

BearLLM: A Prior Knowledge-Enhanced Bearing Health Management Framework with Unified Vibration Signal Representation

AAAI 2025technical

We propose a bearing health management framework leveraging large language models (BearLLM), a novel multimodal model that unifies multiple bearing-related tasks by processing user prompts and vibration signals. Specifically, we introduce a prior knowledge-enhanced unified vibration signal represent…

2025

Bilevel Learning for Low-Light Image Enhancement and Detection

ICASSP 2025accepted

Object detection in low-light scenes is a challenging but widely discussed topic in computer vision. A common approach in low-light object detection involves employing cascaded architectures to connect enhancement and detection networks. This strategy aims to bridge the gap between low-light and nor…

Cited by 0SourceScholar
2025

Bridging and Modeling Correlations in Pairwise Data for Direct Preference Optimization

ICLR 2025poster

Direct preference optimization (DPO), a widely adopted offline preference optimization algorithm, aims to align large language models (LLMs) with human-desired behaviors using pairwise preference data. However, the generation of the winning response and the losing response within pairwise data are t…

2025

CSD: Weather forecasting with graph neural network based on cross-scale diffusivity

ICASSP 2025accepted

Automated weather stations play a pivotal role in fine-grained weather forecasting, due to their cost-effectiveness and global deployment potential. Data-driven methods, particularly deep learning techniques, have emerged as potent tools for precise weather forecasting. However, capturing meaningful…

Cited by 0SourceScholar
2025

CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories

NAACL 2025long

The increasing complexity of computer science research projects demands more effective tools for deploying code repositories. Large Language Models (LLMs), such as Anthropic Claude and Meta Llama, have demonstrated significant advancements across various fields of computer science research, includin…

2025

Can GRPO Boost Complex Multimodal Table Understanding?

EMNLP 2025

Existing table understanding methods face challenges due to complex table structures and intricate logical reasoning. While supervised finetuning (SFT) dominates existing research, reinforcement learning (RL), such as Group Relative Policy Optimization (GRPO), has shown promise but struggled with lo

2025

Can Large Language Models Master Complex Card Games?

NeurIPS 2025poster

Complex games have long been an important benchmark for testing the progress of artificial intelligence algorithms. AlphaGo, AlphaZero, and MuZero have defeated top human players in Go and Chess, garnering widespread societal attention towards artificial intelligence. Concurrently, large language mo…

Cited by 0SourcecodeScholar
2025

ChatCAD: An MLLM-Guided Framework for Zero-shot CAD Drawing Restoration

ICASSP 2025accepted

CAD drawing restoration is one of the most urgent needs in industrial manufacturing. The existing research focuses on the digitization of CAD drawings, However, there are actually many problems in digitized CAD drawings due to the upgrading of engineering drafting software, and it is difficult to re…

Cited by 0SourceScholar
2025

CoSER: Coordinating LLM-Based Persona Simulation of Established Roles

ICML 2025poster

Role-playing language agents (RPLAs) have emerged as promising applications of large language models (LLMs). However, simulating established characters presents a challenging task for RPLAs, due to the lack of authentic character datasets and nuanced evaluation methods using such data. In this paper…

2025

CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding

NeurIPS 2025oral

Coral reefs are vital yet vulnerable ecosystems that require continuous monitoring to support conservation. While coral reef images provide essential information in coral monitoring, interpreting such images remains challenging due to the need for domain expertise. Visual Question Answering (VQA), p…

Cited by 0SourceScholar
2025

Critical Forgetting-Based Multi-Scale Disentanglement for Deepfake Detection

AAAI 2025technical

Recent face forgery detection methods based on disentangled representation learning utilize paired images for cross-reconstruction, aiming to extract forgery-relevant attributes and forgery-irrelevant content. However, there still exist the following issues that may comprise the detector performance…

Cited by 0SourcePDFScholar
2025

DCI: Dual-Conditional Inversion for Boosting Diffusion-Based Image Editing

NeurIPS 2025poster

Diffusion models have achieved remarkable success in image generation and editing tasks. Inversion within these models aims to recover the latent noise representation for a real or generated image, enabling reconstruction, editing, and other downstream tasks. However, to date, most inversion approac…

Cited by 0SourcecodeScholar
2025

DERI: Cross-Modal ECG Representation Learning with Deep ECG-Report Interaction

IJCAI 2025

Electrocardiogram (ECG) is widely used to diagnose cardiac conditions via deep learning methods. Although existing self-supervised learning (SSL) methods have achieved great performance in learning representation for ECG-based cardiac conditions classification, the clinical semantics can not be effe

2025

DatawiseAgent: A Notebook-Centric LLM Agent Framework for Adaptive and Robust Data Science Automation

EMNLP 2025

Existing large language model (LLM) agents for automating data science show promise, but they remain constrained by narrow task scopes, limited generalization across tasks and models, and over-reliance on state-of-the-art (SOTA) LLMs. We introduce DatawiseAgent, a notebook-centric LLM agent framewor

Cited by 0SourcePDFScholar
2025

Deep Hierarchies and Invariant Disease-Indicative Feature Learning for Computer Aided Diagnosis of Multiple Fundus Diseases

AAAI 2025technical

With the advancement of computer vision, numerous models have been proposed for screening of fundus diseases. However, the recognition of multiple fundus diseases is often hampered by the simultaneous presence of multiple disease types and the confluence of lesion types in fundus images. This paper…

Cited by 0SourcePDFScholar
2025

Deep Opinion-Unaware Blind Image Quality Assessment by Learning and Adapting from Multiple Annotators

IJCAI 2025

Existing deep neural network (DNN)-based blind image quality assessment (BIQA) methods primarily rely on human-rated datasets for training. However, collecting human labels is extremely time-consuming and labor-intensive, posing a significant bottleneck for practical applications. To address this ch

2025

Detecting Conversational Mental Manipulation with Intent-Aware Prompting

COLING 2025main

Mental manipulation severely undermines mental wellness by covertly and negatively distorting decision-making. While there is an increasing interest in mental health care within the natural language processing community, progress in tackling manipulation remains limited due to the complexity of dete…

2025

Detoxifying Large Language Models via the Diversity of Toxic Samples

EMNLP 2025

Eliminating toxicity from Large Language Models (LLMs) is crucial for ensuring user safety. However, current methods have limitations in the analysis and utilization of toxic samples, failing to fully harness their potential. Through comparative analysis of toxic and safe samples, we discover that t

2025

Dialect-SQL: An Adaptive Framework for Bridging the Dialect Gap in Text-to-SQL

EMNLP 2025

Text-to-SQL is the task of translating natural language questions into SQL queries based on relational databases. Different databases implement their own SQL dialects, leading to variations in syntax. As a result, SQL queries designed for one database may not execute properly in another, creating a

2025

Disentangle Nighttime Lens Flares: Self-supervised Generation-based Lens Flare Removal

AAAI 2025technical

Lens flares arise from light reflection and refraction within sensor arrays, whose diverse types include glow, veiling glare, reflective flare and so on. Existing methods are specialized for one specific type only, and overlook the simultaneous occurrence of multiple typed lens flares, which is comm…

Cited by 1SourcePDFScholar
2025

Don’t Forget the Enjoin: FocalLoRA for Instruction Hierarchical Alignment in Large Language Models

NeurIPS 2025poster

Recent studies reveal that large language models (LLMs) often struggle to resolve conflicting instructions embedded within hierarchical prompts, resulting in decreased compliance with system-level directives and compromising the reliability of safety-critical applications. While earlier approaches a…

Cited by 0SourceScholar
2025

Dual-View Interaction-Aware Lane Change Prediction for Autonomous Driving

AAAI 2025technical

As artificial intelligence techniques evolve, we are approaching a critical moment for the widespread deployment of autonomous vehicles. Subsequently, the emergence of mixed-autonomy traffic environments presents formidable challenges to autonomous vehicles, especially for the accurate prediction of…

Cited by 0SourcePDFScholar
2025

ECG2TOK: ECG Pre-Training with Self-Distillation Semantic Tokenizers

IJCAI 2025

Self-supervised learning (SSL) has garnered increasing attention in electrocardiogram (ECG) analysis for its effectiveness in resource-limited settings. Existing state-of-the-art SSL methods rely on time-frequency detail reconstruction, but due to the inherent redundancy of ECG signals and individua

2025

Enhancing Vision-Language Models with Morphological and Taxonomic Knowledge: Towards Coral Recognition for Ocean Health

AAAI 2025technical

Coral reefs play a crucial role in marine ecosystems, offering a nutrient-rich environment and safe shelter for numerous marine species. Automated coral image recognition aids in monitoring ocean health at a scale without experts' manual effort. Recently, large vision-language models like CLIP have…

Cited by 0SourcePDFScholar
2025

Enhancing the Performance of Global Model by Improving the Adaptability of Local Models in Federated Learning

IJCAI 2025

Federated learning enables the clients to collaboratively train a global model, which is aggregated from local models. Due to the heterogeneous data distributions over clients and data privacy in federated learning, it is difficult to train local models to achieve a well-performed global model. In t

Cited by 0SourcePDFScholar
2025

Ex-VAD: Explainable Fine-grained Video Anomaly Detection Based on Visual-Language Models

ICML 2025poster

With advancements in visual language models (VLMs) and large language models (LLMs), video anomaly detection (VAD) has progressed beyond binary classification to fine-grained categorization and multidimensional analysis. However, existing methods focus mainly on coarse-grained detection, lacking ano…

Cited by 0SourcePDFScholar
2025

Federated Weakly Supervised Video Anomaly Detection with Multimodal Prompt

AAAI 2025technical

Video anomaly detection (VAD) aims at locating the abnormal events in videos. Recently, the Weakly Supervised VAD has made great progress, which only requires video-level annotations when training. In practical applications, different institutions may have different types of abnormal videos. However…

2025

Finding Local Diffusion Schrodinger Bridge using Kolmogorov-Arnold Network

CVPR 2025poster

In image generation, Schrodinger Bridge (SB)-based methods theoretically enhance the efficiency and quality compared to the diffusion models by finding the least costly path between two distributions. However, they are computationally expensive and time-consuming when applied to complex image data.…

2025

Flow Field Reconstruction with Sensor Placement Policy Learning

NeurIPS 2025poster

Flow‐field reconstruction from sparse sensor measurements remains a central challenge in modern fluid dynamics, as the need for high‐fidelity data often conflicts with practical limits on sensor deployment. Existing deep learning–based methods have demonstrated promising results, but they typically…

Cited by 0SourceScholar
2025

GBA-Net: A Method for 3D Brain Tumor Segmentation Based on Multi-scale Gaussian Boundary Attention

ICASSP 2025accepted

The complex nature of brain tumors, characterized by their individual shapes, sizes, and locations, as well as the presence of indistinct boundaries, presents a challenging task for precise automatic segmentation. While U-Net has been a top performer in medical image segmentation, it struggles with…

Cited by 1SourceScholar
2025

GIVE: Structured Reasoning of Large Language Models with Knowledge Graph Inspired Veracity Extrapolation

ICML 2025poster

Existing approaches based on context prompting or reinforcement learning (RL) to improve the reasoning capacities of large language models (LLMs) depend on the LLMs' internal knowledge to produce reliable Chain-Of-Thought (CoT). However, no matter the size of LLMs, certain problems cannot be resolve…

2025

Gen-SQL: Efficient Text-to-SQL By Bridging Natural Language Question And Database Schema With Pseudo-Schema

COLING 2025main

With the prevalence of Large Language Models (LLMs), recent studies have shifted paradigms and leveraged LLMs to tackle the challenging task of Text-to-SQL. Because of the complexity of real world databases, previous works adopt the retrieve-then-generate framework to retrieve relevant database sche…

2025

GradOT: Training-free Gradient-preserving Offsite-tuning for Large Language Models

ACL 2025long

The rapid growth of large language models (LLMs) with traditional centralized fine-tuning emerges as a key technique for adapting these models to domain-specific challenges, yielding privacy risks for both model and data owners. One promising solution, called offsite-tuning (OT), is proposed to addr…

2025

Hierarchical Frequency Tagging Probe (HFTP): A Unified Approach to Investigate Syntactic Structure Representations in Large Language Models and the Human Brain

NeurIPS 2025poster

Large Language Models (LLMs) demonstrate human-level or even superior language abilities, effectively modeling syntactic structures, yet the specific computational units responsible remain unclear. A key question is whether LLM behavioral capabilities stem from mechanisms akin to those in the human…

Cited by 0SourcecodeScholar
2025

How Do Large Language Models Perform on PDE Discovery: A Coarse-to-fine Perspective

EMNLP 2025

This paper studies the problem of how to use large language models (LLMs) to identify the underlying partial differential equations (PDEs) out of very limited observations of a physical system. Previous methods usually utilize physical-informed neural networks (PINNs) to learn the PDE solver and coe

Cited by 0SourcePDFScholar
2025

HyperTrans: Efficient Hypergraph-Driven Cross-Domain Pattern Transfer in Image Anomaly Detection

IJCAI 2025

Anomaly detection plays a pivotal role in industrial quality assurance processes, with cross-domain problems, exemplified by the model upgrade from RGB to 3D, being prevalent in real-world scenarios yet remaining systematically underexplored. To address the severe challenges posed by the extreme lac

2025

IDEA-Bench: How Far are Generative Models from Professional Designing?

CVPR 2025poster

Recent advancements in image generation models enable the creation of high-quality images and targeted modifications based on textual instructions. Some models even support multimodal complex guidance and demonstrate robust task generalization capabilities. However, they still fall short of meeting…

2025

IR-MFGL: Image-Represented Magnetic Field Global Localization in Repetitive Environments

RA-L 2025

Global localization is an essential ingredient for autonomous mobile robots. However, existing global localization systems primarily rely on Global Navigation Satellite System (GNSS), infrastructures, or visual/LiDAR-based place recognition, which suffer from enclosed/semi-enclosed GNSS-denied envir

Cited by 0SourceScholar
2025

Inferring from Logits: Exploring Best Practices for Decoding-Free Generative Candidate Selection

ACL 2025long

Generative Language Models rely on autoregressive decoding to produce the output sequence token by token. Many tasks such as preference optimization, require the model to produce task-level output consisting of multiple tokens directly by selecting candidates from a pool as predictions. Determining…

Cited by 0SourcePDFScholar
2025

Influence-Guided Diffusion for Dataset Distillation

ICLR 2025poster

Dataset distillation aims to streamline the training process by creating a compact yet effective dataset for a much larger original dataset. However, existing methods often struggle with distilling large, high-resolution datasets due to prohibitive resource costs and limited performance, primarily s…

2025

Instruction-Tuning Data Synthesis from Scratch via Web Reconstruction

ACL 2025finding

The improvement of LLMs’ instruction-following capabilities depends critically on the availability of high-quality instruction-response pairs. While existing automatic data synthetic methods alleviate the burden of manual curation, they often rely heavily on either the quality of seed data or strong…

2025

IntelliCockpitBench: A Comprehensive Benchmark to Evaluate VLMs for Intelligent Cockpit

ACL 2025finding

The integration of sophisticated Vision-Language Models (VLMs) in vehicular systems is revolutionizing vehicle interaction and safety, performing tasks such as Visual Question Answering (VQA). However, a critical gap persists due to the lack of a comprehensive benchmark for multimodal VQA models in…

2025

Intra and Inter Parser-Prompted Transformers for Effective Image Restoration

AAAI 2025technical

We propose Intra and Inter Parser-Prompted Transformers (PPTformer) that explore useful features from visual foundation models for image restoration. Specifically, PPTformer contains two parts: an Image Restoration Network (IRNet) for restoring images from degraded observations and a Parser-Prompted…

2025

InverseCoder: Self-improving Instruction-Tuned Code LLMs with Inverse-Instruct

AAAI 2025technical

Recent advancements in open-source code large language models (LLMs) have been driven by fine-tuning on the data generated from powerful closed-source LLMs, which are expensive to obtain. This paper explores whether it is possible to use a fine-tuned open-source model to generate additional data to…

2025

LawDNet: Enhanced Audio-Driven Lip Synthesis via Local Affine Warping Deformation

ICASSP 2025accepted

In the domain of photorealistic talking head generation, the fidelity of audio-driven lip motion synthesis is essential for realistic virtual interactions. Existing methods face two key challenges: a lack of vivacity due to limited diversity in generated lip poses and noticeable anamorphose motions…

Cited by 0SourceScholar
2025

Learning Robust Image Watermarking with Lossless Cover Recovery

ICCV 2025poster

Watermarking as a traceable authentication technology has been widely applied in image copyright protection. However, most existing watermarking methods embed watermarks by adding irremovable perturbations to the cover image, causing permanent distortion. To address this issue, we propose a novel wa…

2025

Localized Data Shapley: Accelerating Valuation for Nearest Neighbor Algorithms

NeurIPS 2025poster

Data Shapley values provide a principled approach for quantifying the contribution of individual training examples to machine learning models. However, computing these values often requires computational complexity that is exponential in the data size, and this has led researchers to pursue efficien…

Cited by 0SourceScholar
2025

Long-context Language Models Fail in Basic Retrieval Tasks Without Sufficient Reasoning Steps

EMNLP 2025

Long-context language models (LCLMs), characterized by their extensive context window, are becoming popular. However, despite the fact that they are nearly perfect at standard long-context retrieval tasks, our evaluations demonstrate they fail in some basic cases. Later, we find they can be well add

2025

MMGIA: Gradient Inversion Attack Against Multimodal Federated Learning via Intermodal Correlation

IJCAI 2025

Multimodal federated learning (MMFL) enables collaborative model training across multiple modalities, such as images and text, without requiring direct data sharing. However, the inherent correlations between modalities introduce new privacy vulnerabilities, making MMFL more susceptible to gradient

Cited by 0SourcePDFScholar
2025

MedFact: A Large-scale Chinese Dataset for Evidence-based Medical Fact-checking of LLM Responses

EMNLP 2025

Medical fact-checking has become increasingly critical as more individuals seek medical information online. However, existing datasets predominantly focus on human-generated content, leaving the verification of content generated by large language models (LLMs) relatively unexplored. To address this

2025

Memorize and Rank: Elevating Large Language Models for Clinical Diagnosis Prediction

AAAI 2025technical

Clinical diagnosis prediction models, when provided with a patient's medical history, aim to detect potential diseases early, facilitating timely intervention and improving prognostic outcomes. However, the inherent scarcity of patient data and large disease candidate space often pose challenges in…

Cited by 3SourcePDFScholar
2025

Memory-augmented Query Reconstruction for LLM-based Knowledge Graph Reasoning

ACL 2025finding

Large language models (LLMs) have achieved remarkable performance on knowledge graph question answering (KGQA) tasks by planning and interacting with knowledge graphs. However, existing methods often confuse tool utilization with knowledge reasoning, harming readability of model outputs and giving r…

2025

MetaScientist: A Human-AI Synergistic Framework for Automated Mechanical Metamaterial Design

NAACL 2025system demonstrations

The discovery of novel mechanical metamaterials, whose properties are dominated by their engineered structures rather than chemical composition, is a knowledge-intensive and resource-demanding process. To accelerate the design of novel metamaterials, we present MetaScientist, a human-in-the-loop sys…

2025

MetricEmbedding: Accelerate Metric Nearness by Tropical Inner Product

ICML 2025poster

The Metric Nearness Problem involves restoring a non-metric matrix to its closest metric-compliant form, addressing issues such as noise, missing values, and data inconsistencies. Ensuring metric properties, particularly the $O(N^3)$ triangle inequality constraints, presents significant computation…

Cited by 0SourcePDFScholar
2025

Mixture of In-Context Prompters for Tabular PFNs

ICLR 2025poster

Recent benchmarks find In-Context Learning (ICL) outperforms both deep learning and tree-based algorithms on small tabular datasets. However, on larger datasets, ICL for tabular learning suffers in both efficiency and effectiveness. In terms of efficiency, transformers incur linear space and quadrat…

Cited by 10SourcePDFScholar
2025

Multi-Label Few-Shot Image Classification via Pairwise Feature Augmentation and Flexible Prompt Learning

AAAI 2025technical

Multi-label few-shot image classification is a crucial and challenging task due to limited annotated data and elusive category specificity. However, research on this topic is still in the rudimentary stage and few methods are available. Existing methods either leverage data augmentation to alleviate…

Cited by 0SourcePDFScholar
2025

Multi-view Evidential Learning-based Medical Image Segmentation

AAAI 2025technical

Medical image segmentation provides useful information about the shape and size of organs, which is beneficial for improving diagnosis, analysis, and treatment. Despite traditional deep learning-based models can extract domain-specific knowledge, they face a generalization bottleneck due to the limi…

Cited by 0SourcePDFScholar
2025

MutationGuard: A Graph and Temporal-Spatial Neural Method for Detecting Mutation Telecommunication Fraud

IJCAI 2025

Telecommunication fraud refers to deceptive activities in the field of communication services. This research focuses on a category of fraud identified as ''mutation telecommunication fraud". There is currently a lack of research on mutation telecommunication fraud detection, allowing this type of fr

2025

NUMINA: A Natural Understanding Benchmark for Multi-dimensional Intelligence and Numerical Reasoning Abilities

EMNLP 2025

Recent advancements in 2D multimodal large language models (MLLMs) have significantly improved performance in vision-language tasks. However, extending these capabilities to 3D environments remains a distinct challenge due to the complexity of spatial reasoning. Nevertheless, existing 3D benchmarks

2025

OVG-HQ: Online Video Grounding with Hybrid-modal Queries

ICCV 2025poster

Video grounding (VG) task focuses on locating specific moments in a video based on a query, usually in text form. However, traditional VG struggles with some scenarios like streaming video or queries using visual cues. To fill this gap, we present a new task named Online Video Grounding with Hybrid-…

Cited by 0SourcePDFScholar
2025

Object-Level Backdoor Attacks in RGB-T Semantic Segmentation with Cross-Modality Trigger Optimization

IJCAI 2025

The escalating threat of backdoor risks in deep vision models is a pressing concern. Existing research on backdoor attacks is often confined to a single modality, neglecting the challenges posed by multi-modality scene perception. This work is a pioneer of backdoor attacks in RGB-Thermal (RGB-T) sem

Cited by 0SourcePDFScholar
2025

Online Residual Model Learning for Model Predictive Control of Autonomous Surface Vehicles in Real-World Environments

IROS 2025

Model predictive control (MPC) relies on an accurate dynamics model to achieve precise and safe robot operation. In complex and dynamic aquatic environments, developing an accurate model that captures hydrodynamic details and accounts for environmental disturbances like waves, currents, and winds is

Cited by 0SourceScholar
2025

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

NeurIPS 2025poster

We introduce *OpenVLThinker*, one of the first open-source large vision–language models (LVLMs) to exhibit sophisticated chain-of-thought reasoning, achieving notable performance gains on challenging visual reasoning tasks. While text-based reasoning models (e.g., Deepseek R1) show promising results…

Cited by 0SourceScholar
2025

PUER: Boosting Few-shot Positive-Unlabeled Entity Resolution with Reinforcement Learning

EMNLP 2025

Entity resolution is a fundamental problem in data management that aims to identify all duplicate entries within collections of multi-attribute tuples. Most existing works focus on supervised learning, relying on large amounts of high-quality labeled data, including both positive and negative tuple

2025

Partial Reconstruction Error for Deepfake Detection

ICASSP 2025accepted

The rapid development of deepfake technology poses a formidable challenge to personal privacy and security, underscoring the urgent need for deepfake detection. Recently, the methods based on the reconstruction error, such as DIRE and RECCE, achieve impressive performance in forgery detection. Howev…

Cited by 0SourceScholar
2025

PatchST: A Patch Spatial-Temporal Network for Large-Scale Traffic Forecasting

ICASSP 2025accepted

Traffic prediction plays a critical role in mitigating congestion, optimizing traffic flow, and enhancing urban mobility. Accurate predictions contribute directly to improved safety, reduced travel times, and a more efficient transportation system. For example, when a surge in traffic is anticipated…

Cited by 0SourceScholar
2025

Promptable Representation Distribution Learning and Data Augmentation for Gigapixel Histopathology WSI Analysis

AAAI 2025technical

Gigapixel image analysis, particularly for whole slide images (WSIs), often relies on multiple instance learning (MIL). Under the paradigm of MIL, patch image representations are extracted and then fixed during the training of the MIL classifiers for efficiency consideration. However, the invariance…

2025

Protein Large Language Models: A Comprehensive Survey

EMNLP 2025

Protein-specific large language models (ProteinLLMs) are revolutionizing protein science by enabling more efficient protein structure prediction, function annotation, and design. While existing surveys focus on specific aspects or applications, this work provides the first comprehensive overview of

2025

Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content

CVPR 2025poster

Evaluating text-to-vision content hinges on two crucial aspects: **visual quality** and **alignment**. While significant progress has been made in developing objective models to assess these dimensions, the performance of such models heavily relies on the scale and quality of human annotations. Acco…

2025

QCRD: Quality-guided Contrastive Rationale Distillation for Large Language Models

EMNLP 2025

The deployment of large language models (LLMs) faces considerable challenges concerning resource constraints and inference efficiency. Recent research has increasingly focused on smaller, task-specific models enhanced by distilling knowledge from LLMs. However, prior studies have often overlooked th

Cited by 0SourcePDFScholar
2025

RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering

NeurIPS 2025poster

In real-world scenarios, providing user queries with visually enhanced responses can considerably benefit understanding and memory, underscoring the great value of interleaved image-text generation. Despite recent progress, like the visual autoregressive model that unifies text and image processing…

Cited by 0SourcecodeScholar
2025

Realistic Evaluation of Deep Partial-Label Learning Algorithms

ICLR 2025spotlight

Partial-label learning (PLL) is a weakly supervised learning problem in which each example is associated with multiple candidate labels and only one is the true label. In recent years, many deep PLL algorithms have been developed to improve model performance. However, we find that some early develop…

Cited by 1SourcePDFScholar
2025

Relative Localization of Asynchronous Agents Based on Hybrid Active-Passive Two-Way Ranging

ICASSP 2025accepted

To position agents equipped with ultrawideband (UWB) devices without requiring continuous clock synchronization, the primary measurement currently used is the time of flight between agents, estimated through active two-way ranging (TWR). However, the cumbersome signal exchange mechanism in active TW…

Cited by 0SourceScholar
2025

Rethink GraphODE Generalization within Coupled Dynamical System

ICML 2025spotlight

Coupled dynamical systems govern essential phenomena across physics, biology, and engineering, where components interact through complex dependencies. While Graph Ordinary Differential Equations (GraphODE) offer a powerful framework to model these systems, their **generalization** capabilities degra…

Cited by 0SourcePDFScholar
2025

Rethinking Joint Maximum Mean Discrepancy for Visual Domain Adaptation

NeurIPS 2025oral

In domain adaption (DA), joint maximum mean discrepancy (JMMD), as a famous distribution-distance metric, aims to measure joint probability distribution difference between the source domain and target domain, while it is still not fully explored and especially hard to be applied into a subspace-lear…

Cited by 0SourceScholar
2025

S3E: Self-Supervised State Estimation for Radar-Inertial System

ICCV 2025poster

Millimeter-wave radar for state estimation is gaining significant attention for its affordability and reliability in harsh conditions. Existing localization solutions typically rely on post-processed radar point clouds as landmark points. Nonetheless, the inherent sparsity of radar point clouds, gho…

2025

Safe Motion Planning and Control Using Predictive and Adaptive Barrier Methods for Autonomous Surface Vessels

IROS 2025

Safe motion planning is essential for autonomous vessel operations, especially in challenging spaces such as narrow inland waterways. However, conventional motion planning approaches are often computationally intensive or overly conservative. This paper proposes a safe motion planning strategy combi

Cited by 1SourceScholar
2025

SampleMix: A Sample-wise Pre-training Data Mixing Strategy by Coordinating Data Quality and Diversity

EMNLP 2025

Existing pretraining data mixing methods for large language models (LLMs) typically follow a domain-wise methodology, a top-down process that first determines domain weights and then performs uniform data sampling across each domain. However, these approaches neglect significant inter-domain overlap

Cited by 0SourcePDFScholar
2025

ScaleOT: Privacy-utility-scalable Offsite-tuning with Dynamic LayerReplace and Selective Rank Compression

AAAI 2025technical

Offsite-tuning is a privacy-preserving method for tuning large language models (LLMs) by sharing a lossy compressed emulator from the LLM owners with data owners for downstream task tuning. This approach protects the privacy of both the model and data owners. However, current offsite tuning methods…

2025

SdalsNet: Self-Distilled Attention Localization and Shift Network for Unsupervised Camouflaged Object Detection

AAAI 2025technical

Unsupervised camouflaged object detection (UCOD) poses significant challenges, primarily attributed to the absence of human labels. Existing UCOD methodologies, leveraging attention mechanisms, often struggle to achieve precise localization of camouflaged objects. To overcome this limitation, we int…

2025

Symmetry-Preserving Conformer Ensemble Networks for Molecular Representation Learning

NeurIPS 2025poster

Molecular representation learning has emerged as a promising approach for modeling molecules with deep learning in chemistry and beyond. While 3D geometric models effectively capture molecular structure, they typically process single static conformers, overlooking the inherent flexibility and dynami…

Cited by 0SourceScholar
2025

TIDE-Net: A Physics-Based Graph Model for Predicting Tropical Cyclone Impacts on Estuarine Systems

ICASSP 2025accepted

Estuarine systems, located at the interface of land, ocean, and atmosphere, are vital to global ecosystems and economies due to their rich exchanges among multiple environments. Tropical cyclones cause significant fluctuations in salinity and other environmental parameters within estuaries, impactin…

Cited by 0SourceScholar
2025

Template-Driven LLM-Paraphrased Framework for Tabular Math Word Problem Generation

AAAI 2025technical

Solving tabular math word problems (TMWPs) has become a critical role in evaluating the mathematical reasoning ability of large language models (LLMs), where large-scale TMWP samples are commonly required for fine-tuning. Since the collection of high-quality TMWP datasets is costly and time-consumin…

2025

Time-IMM: A Dataset and Benchmark for Irregular Multimodal Multivariate Time Series

NeurIPS 2025poster

Time series data in real-world applications such as healthcare, climate modeling, and finance are often irregular, multimodal, and messy, with varying sampling rates, asynchronous modalities, and pervasive missingness. However, existing benchmarks typically assume clean, regularly sampled, unimodal…

Cited by 0SourcecodeScholar
2025

Toward Comprehensive Semantic Prompt for Region Contrastive Learning Underwater Image Enhancement

ICASSP 2025accepted

Underwater image enhancement (UIE) focuses on mitigating image quality degradation due to light absorption and scattering. However, most existing methods enhance images via a global and uniform manner, neglecting the inherent semantic information in different regions, which may cause the network to…

Cited by 0SourceScholar
2025

Towards Better Robustness Against Natural Corruptions in Document Tampering Localization

AAAI 2025technical

Marvelous advances have been exhibited in recent document tampering localization (DTL) systems. However, confronted with corrupted tampered document images, their vulnerability is fatal in real-world scenarios. While robustness against adversarial attack has been extensively studied by adversarial t…

2025

Uni-Zipper: A Multi-modal Perception Framework of Deformable Objects with Unpaired Data

IROS 2025

Multi-modal perception plays a crucial role in preventing deformation and damage during the robotic manipulation of deformable objects. However, integrating new heterogeneous modalities into existing robotic perception frameworks remains a significant challenge, primarily due to the need for massive

Cited by 0SourceScholar
2025

UnifiedMLLM: Enabling Unified Representation for Multi-modal Multi-tasks With Large Language Model

NAACL 2025findings

Significant advancements has recently been achieved in the field of multi-modal large language models (MLLMs), demonstrating their remarkable capabilities in understanding and reasoning across diverse tasks. However, these models are often trained for specific tasks and rely on task-specific input-o…

2025

Unlocking the Potential of Lightweight Quantized Models for Deepfake Detection

IJCAI 2025

Deepfake detection is increasingly crucial due to the rapid rise of AI-generated content. Existing methods achieve high performance relying on computationally intensive large models, making real-time detection on resource-constrained edge devices challenging. Given that deepfake detection is a binar

2025

Unsupervised Region-Based Image Editing of Denoising Diffusion Models

AAAI 2025technical

Although diffusion models have achieved remarkable success in the field of image generation, their latent space remains under-explored. Current methods for identifying semantics within latent space often rely on external supervision, such as textual information and segmentation masks. In this paper,…

Cited by 0SourcePDFScholar
2025

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought

NeurIPS 2025poster

Recent advancements in reasoning capability of Multimodal Large Language Models (MLLMs) demonstrate its effectiveness in tackling complex visual tasks. However, existing MLLM-based Video Anomaly Detection (VAD) methods remain limited to shallow anomaly descriptions without deep reasoning. In this pa…

Cited by 0SourcecodeScholar
2025

When Evolution Strategy Meets Language Models Tuning

COLING 2025main

Supervised Fine-tuning has been pivotal in training autoregressive language models, yet it introduces exposure bias. To mitigate this, Post Fine-tuning, including on-policy and off-policy methods, has emerged as a solution to enhance models further. However, each has its limitations regarding perfor…

2025

When Inverse Data Outperforms: Exploring the Pitfalls of Mixed Data in Multi-Stage Fine-Tuning

EMNLP 2025

Existing work has shown that o1-level performance can be achieved with limited data distillation, but most existing methods focus on unidirectional supervised fine-tuning (SFT), overlooking the intricate interplay between diverse reasoning patterns. In this paper, we construct r1k, a high-quality re

2025

XCotton: Advancing AI-Enabled Hardware/Software Integrated System for Foreign Fiber Cleaning

AAAI 2025technical

Cotton is a critical agricultural product and industrial raw material, playing a key role in the national economies and people's living conditions, particularly in developing countries. However, cotton picking and processing often result in the contamination with various foreign fibers, such as hair…

Cited by 0SourcePDFScholar