← Search

Yang Liu

766 accepted papers

2026

A Causal Marriage between VLM and IRM from Understanding to Reasoning

CVPR 2026

Vision-Language Models (VLMs) like CLIP exhibit extraordinary out-of-distribution (OOD) generalization, while the theoretical foundations underlying this robustness remain largely unexplored. This work establishes a connection between CLIP and Invariant Risk Minimization (IRM), the principled paradi

Cited by 0SourcecodeScholar
2026

AlignFlow: Improving Flow-based Generative Models with Semi-Discrete Optimal Transport

ICLR 2026poster

Flow-based Generative Models (FGMs) effectively transform noise into a data distribution, and coupling the noise and data in the training of FGM by Optimal Transport (OT) improves the straightness of the flow paths. However, existing OT- based couplings are difficult to combine with modern models an…

Cited by 0SourcecodeScholar
2026

AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language Models

ICLR 2026poster

The rapid development and widespread adoption of Audio Large Language Models (ALLMs) require a rigorous assessment of their trustworthiness. However, existing evaluation frameworks, primarily designed for text, are not equipped to handle the unique vulnerabilities introduced by audio’s acoustic prop…

Cited by 0SourcecodeScholar
2026

Autoregressive Image Generation with Masked Bit Modeling

ICML 2026poster

This paper challenges the dominance of continuous pipelines in visual generation. We systematically investigate the performance gap between discrete and continuous methods. Contrary to the belief that discrete tokenizers are intrinsically inferior, we demonstrate that the disparity arises primarily …

Cited by 0SourceScholar
2026

BabyVision: Visual Reasoning Beyond Language

ICML 2026poster

While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a crucial fact: state-of-the-art MLLMs consistently fail on basic visual tasks that …

Cited by 0SourceScholar
2026

Beyond Structure: Invariant Crystal Property Prediction with Pseudo-Particle Ray Diffraction

ICLR 2026poster

Crystal property prediction, governed by quantum mechanical principles, is computationally prohibitive to solve exactly for large many-body systems using traditional density functional theory. While machine learning models have emerged as efficient approximations for large-scale applications, their…

Cited by 0SourcecodeScholar
2026

BotVA: Combating Social Bots via Variational Feature Augmentation and Adversarial Graph Learning

IJCAI 2026

Social bots threaten online platforms by spreading disinformation and manipulating public discourse. Graph neural networks have emerged as effective tools for bot detection by modeling user interactions, yet two fundamental challenges limit their practical deployment: severe class imbalance where bo

Cited by 0Scholar
2026

Bridging Perception and Planning: Towards End-To-End Planning for Signal Temporal Logic Tasks

ICRA 2026poster

We investigate the task and motion planning problem for Signal Temporal Logic (STL) specifications in robotics. Existing STL methods rely on pre-defined maps or mobility representations, which are ineffective in unstruc- tured real-world environments. We propose the Structured- MoE STL Planner (S-MS…

2026

CARD: Coarse-to-fine Autoregressive Modeling with Radix-based Decomposition for Transferable Free Energy Estimation

ICML 2026poster

Estimating free energy differences quantifies thermodynamic preferences in molecular interactions, which is central to chemistry and drug discovery. Despite fruitful progress, existing methods still face key limitations: classical computational approaches remain prohibitively expensive due to their …

Cited by 0SourceScholar
2026

CloDS: Visual-Only Unsupervised Cloth Dynamics Learning in Unknown Conditions

ICLR 2026poster

Deep learning has demonstrated remarkable capabilities in simulating complex dynamic systems. However, existing methods require known physical properties as supervision or inputs, limiting their applicability under unknown conditions. To explore this challenge, we introduce Cloth Dynamics Grounding…

Cited by 0SourcecodeScholar
2026

Controllable Financial Market Generation with Diffusion Guided Meta Agent

AAAI 2026technical

Generative modeling has transformed many fields, such as language and visual modeling, while its application in financial markets remains under-explored. As the minimal unit within a financial market is an order, order-flow modeling represents a fundamental generative financial task. However, curren

Cited by 0SourcePDFScholar
2026

CrossEarth-Gate: Fisher-Guided Adaptive Tuning Engine for Efficient Adaptation of Cross-Domain Remote Sensing Semantic Segmentation

CVPR 2026

In Remote Sensing (RS), Parameter-Efficient Fine-Tuning (PEFT) has emerged as a key approach to activate the generalizable representation ability of foundation models for downstream tasks. However, existing specialized PEFT methods often fail when applied to large-scale Earth observation tasks, as t

Cited by 0SourceScholar
2026

Cubemap-Based LiDAR-Inertial Odometry with Intensity Assistance

ICRA 2026poster

We present CUBE-LIO, a LiDAR-inertial odometry framework that leverages direct photometric constraints from LiDAR intensity to improve robustness in geometrically degenerate environments. At its core is an efficient cubemap projection that maps LiDAR intensity onto six cube faces, eliminating pole s…

Cited by 0Scholar
2026

DDP-WM: Disentangled Dynamics Prediction for Efficient World Models

ICML 2026poster

World models are essential for autonomous robotic planning. However, the substantial computational overhead of existing dense Transformer-based models significantly hinders real-time deployment. To address this efficiency-performance bottleneck, we introduce DDP-WM, a novel world model centered on t…

Cited by 0SourceScholar
2026

DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning

ICLR 2026poster

Unlearning in Large Language Models (LLMs) is crucial for protecting private data and removing harmful knowledge. Most existing approaches rely on fine-tuning to balance unlearning efficiency with general language capabilities. However, these methods typically require training or access to retain da…

Cited by 0SourcecodeScholar
2026

Deep Reinforcement Learning Based Autonomous Drift System for Abrupt Obstacle Avoidance

RA-L 2026

Autonomous vehicles face significant challenges in executing emergency obstacle avoidance maneuvers beyond conventional driving limits. Previous approaches, relying on vehicle dynamics modeling or simplified learning methods, often struggle with generalization to diverse scenarios. This paper presen

Cited by 0SourcecodeScholar
2026

DiffRP: Diffusion-Driven Promising Region Prediction for Sampling-Based Path Planning

ICRA 2026poster

Utilizing neural networks to predict potential regions containing optimal paths in advance and subsequently biasing the sampling probability towards these promising regions has been proven to effectively enhance the path planning efficiency of sampling-based algorithms. %In complex scenarios, unifor…

Cited by 0SourceScholar
2026

DirectFisheye-GS: Enabling Native Fisheye Input in Gaussian Splatting with Cross-View Joint Optimization

CVPR 2026

3D Gaussian Splatting (3DGS) has enabled efficient 3D scene reconstruction from everyday images with real-time, high-fidelity rendering, greatly advancing VR/AR applications. Fisheye cameras, with their wider field of view (FOV), promise high-quality reconstructions from fewer inputs and have recent

Cited by 0SourceScholar
2026

Doctor-R1: Mastering Clinical Inquiry with Experiential Agentic Reinforcement Learning

ICLR 2026poster

The professionalism of a human doctor in outpatient service depends on two core abilities: the ability to make accurate medical decisions and the medical consultation skill to conduct strategic, empathetic patient inquiry. Existing Large Language Models (LLMs) have achieved remarkable accuracy on me…

Cited by 0SourcecodeScholar
2026

DrugTrail: Explainable Drug Discovery via Structured Reasoning and Druggability‑Tailored Preference Optimization

ICLR 2026poster

Machine learning promises to revolutionize drug discovery, but its "black-box" nature and narrow focus limit adoption by experts. While Large Language Models (LLMs) offer a path forward with their broad knowledge and interactivity, existing methods remain data-intensive and lack transparent reasonin…

Cited by 0SourceScholar
2026

DynamicsBoost: Dynamic Plausible Video Generation via Annotation-Free Continuation Preference Optimization

CVPR 2026

Despite significant progress in text-to-video generation, current models still suffer from unrealistic dynamics, temporal inconsistency, and unstable semantic alignment. Existing preference alignment approaches rely on costly and often ambiguous human or VLM-based video preference annotation, which

Cited by 0SourceScholar
2026

EE-RL: Vision Language Guided Reinforcement Learning with Explorer and Expert model for End-to-End Autonomous Driving

CVPR 2026

End-to-end driving frameworks, which directly map raw sensor data to vehicle control commands, have shown remarkable potential. However, their performance often deteriorates in sparse-critical scenarios, where rare but safety-sensitive events occur. To address this problem, we propose Explorer-Exper

Cited by 0SourcecodeScholar
2026

Enhancing Safety and Manipulability of Redundant Manipulators: Accelerated Motion Generation in Dynamic Environments

ICRA 2026poster

Motion generation in dynamic environments is crucial for human-machine interaction with redundant manipulators. In this context, we propose the Enhancing Safety and Manipulability (ESM) scheme, which integrates geometry-based dynamic obstacle avoidance, manipulability optimization,trajectory trackin…

Cited by 0SourceScholar
2026

Eva-Tracker: ESDF-Update-Free, Visibility-Aware Planning with Target Reacquisition for Robust Aerial Tracking

ICRA 2026poster

The Euclidean Signed Distance Field (ESDF) is widely used in visibility evaluation to prevent occlusions and collisions during tracking. However, frequent ESDF updates introduce considerable computational overhead. To address this issue, we propose Eva-Tracker, a visibility-aware trajectory planning…

2026

Exploring Spatial Intelligence from a Generative Perspective

CVPR 2026

Spatial intelligence is essential for multimodal large language models, yet current benchmarks largely assess it only from an understanding perspective. We ask whether modern generative or unified multimodal models also possess generative spatial intelligence (GSI)--the ability to respect and manipu

Cited by 0SourcecodeScholar
2026

FedHarmony: Harmonizing Heterogeneous Label Correlations in Federated Multi-Label Learning

CVPR 2026

Federated Multi-Label Learning is a distributed paradigm where multiple clients possess heterogeneous multi-label data and perform collaborative learning under privacy constraints without sharing raw data. However, modeling label correlations under heterogeneous distributions remains challenging. Du

Cited by 0SourceScholar
2026

From Prompts to Responses: Dual-Sided Data Leakage and Defense in Split Large Language Models

ICML 2026poster

Large language models (LLMs) are increasingly deployed in privacy-sensitive domains, where users must balance the risk of data exposure through external APIs against the high computational cost of local deployment. Split learning has therefore emerged as a promising paradigm for LLM fine-tuning and …

Cited by 0SourceScholar
2026

From Static to Dynamic: Exploring Self-supervised Image-to-Video Representation Transfer Learning

CVPR 2026

Recent studies have made notable progress in video representation learning by transferring image-pretrained models to video tasks, typically with complex temporal modules and video fine-tuning. However, fine-tuning heavy modules may compromise inter-video semantic separability, i.e., the essential a

Cited by 0SourcecodeScholar
2026

GIPO: Gaussian Importance Sampling Policy Optimization

ICML 2026poster

Post-training with reinforcement learning (RL) has recently shown strong promise for advancing multimodal agents beyond supervised imitation. However, RL remains limited by poor data efficiency, particularly in settings where interaction data are scarce and quickly become outdated. To address this c…

Cited by 0SourceScholar
2026

Granulon: Awakening Pixel-Level Visual Encoders with Adaptive Multi-Granularity Semantics for MLLM

CVPR 2026

Recent advances in multimodal large language models largely rely on CLIP-based visual encoders, which emphasize global semantic alignment but struggle with fine-grained visual understanding. In contrast, DINOv3 provides strong pixel-level perception yet lacks coarse-grained semantic abstraction, lea

Cited by 0SourcecodeScholar
2026

Graph VQ-Transformer (GVT): Fast and Accurate Molecular Generation via High-Fidelity Discrete Latents

AAAI 2026technical

The de novo generation of molecules with desirable properties is a critical challenge, where diffusion models are computationally intensive and autoregressive models struggle with error propagation. In this work, we introduce the Graph VQ-Transformer (GVT), a two-stage generative framework that achi

Cited by 0SourcePDFScholar
2026

HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models

CVPR 2026

In multimodal large language models (MLLMs), the surge of visual tokens significantly increases the inference time and computational overhead, making them impractical for real-time or resource-constrained applications.Visual token pruning is a promising strategy for reducing the cost of MLLM inferen

Cited by 0SourcecodeScholar
2026

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models

CVPR 2026

Vision-Language-Action (VLA) models have recently enabled robotic manipulation by grounding visual and linguistic cues into actions. However, most VLAs assume the Markov property, relying only on the current observation and thus suffering from temporal myopia that degrades long-horizon coherence. In

Cited by 0SourcecodeScholar
2026

Hidden in the Noise: Unveiling Backdoors in Audio LLMs Alignment Through Latent Acoustic Pattern Triggers

AAAI 2026technical

As Audio Large Language Models (ALLMs) emerge as powerful tools for speech processing, their safety implications demand urgent attention. While considerable research has explored textual and vision safety, audio’s distinct characteristics present significant challenges. This paper first investigates

Cited by 0SourcePDFScholar
2026

Human-AI Curation Synergy: Scaling Preference Data Curation via Human-Guided AI Feedback

ICLR 2026poster

Despite the critical role of reward models (RMs) in reinforcement learning from human feedback (RLHF), current state-of-the-art open RMs perform poorly on most existing evaluation benchmarks, failing to capture the spectrum of nuanced and sophisticated human preferences. Even approaches incorporatin…

Cited by 0SourcecodeScholar
2026

I2E: Real-Time Image-to-Event Conversion for High-Performance Spiking Neural Networks

AAAI 2026technical

Spiking neural networks (SNNs) promise highly energy-efficient computing, but their adoption is hindered by a critical scarcity of event-stream data. This work introduces I2E, an algorithmic framework that resolves this bottleneck by converting static images into high-fidelity event streams. By simu

Cited by 0SourcePDFScholar
2026

ICM-Fusion: In-Context Meta-Optimized LoRA Fusion for Multi-Task Adaptation

AAAI 2026technical

Enabling multi-task adaptation in pre-trained Low-Rank Adaptation (LoRA) models is crucial for enhancing their generalization capabilities. Most existing pre-trained LoRA fusion methods decompose weight matrices, sharing similar parameters, while fusion divergent ones. However, this paradigm inevit

Cited by 0SourcePDFScholar
2026

Image-Text Knowledge Modeling for Unsupervised Multi-Scenario Person Re-Identification

AAAI 2026technical

We propose unsupervised multi-scenario (UMS) person re-identification (ReID) as a new task that expands ReID across diverse scenarios (cross-resolution, clothing change, etc.) within a single coherent framework. To tackle UMS-ReID, we introduce image-text knowledge modeling (ITKM) -- a three-stage f

Cited by 0SourcePDFScholar
2026

Improving Sustainability of Adversarial Examples in Class-Incremental Learning

AAAI 2026technical

Current adversarial examples (AEs) are typically designed for static models. However, with the wide application of Class-Incremental Learning (CIL), models are no longer static and need to be updated with new data distributed and labeled differently from the old ones. As a result, existing AEs often

Cited by 0SourcePDFScholar
2026

InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language Models

ICLR 2026poster

Advanced reasoning in large language models has achieved remarkable performance on challenging tasks, but the prevailing long-context reasoning paradigm faces critical limitations: quadratic computational scaling with sequence length, reasoning constrained by maximum context boundaries, and performa…

Cited by 0SourcecodeScholar
2026

InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation

ICML 2026poster

While Large Language Models (LLMs) hold promise for automating science and education, generating interactive scientific demonstrations demands a complex synthesis of deep domain knowledge and precise reactive coding. Current benchmarks fail to capture this synergy, largely bifurcating into static co…

Cited by 0SourceScholar
2026

Intra-Modal Neighbors Never Lie: Rectifying Inter-Modal Noisy Correspondence via Graph-Based Intra-Modal Reasoning

ICML 2026poster

Large-scale web-harvested datasets have fueled the progress of cross-modal retrieval but inevitably suffer from \textit{noisy correspondence}, which severely degrades model generalization. Existing methods primarily address this by filtering out noise or seeking a substitute label, yet they predomin…

Cited by 0SourceScholar
2026

Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment

ICLR 2026poster

Alignment is vital for safely deploying large language models (LLMs). Existing techniques are either reward-based--train a reward model on preference pairs and optimize with reinforcement learning (RL)--or reward-free--directly fine-tune on ranked outputs. Recent research show that well-tuned reward…

Cited by 0SourceScholar
2026

JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence

ICLR 2026poster

The scope of neural code intelligence is rapidly expanding beyond text-based source code to encompass the rich visual outputs that programs generate. This visual dimension is critical for advanced applications like flexible content generation and precise, program-driven editing of visualizations. Ho…

Cited by 0SourcecodeScholar
2026

Kimi-Dev: Agentless Training as Skill Prior for SWE-agents

ICLR 2026poster

Large Language Models (LLMs) are increasingly applied to software engineering (SWE), with SWE-bench as a key benchmark. Solutions are split into SWE-Agent frameworks with multi-turn interactions and workflow-based Agentless methods with single-turn verifiable steps. We argue these paradigms are not…

Cited by 0SourcecodeScholar
2026

Label Smoothing Improves Machine Unlearning

ICLR 2026poster

The objective of machine unlearning (MU) is to eliminate previously learned data from a model. However, it can be challenging to strike a balance between computation cost and performance when using existing MU techniques. Taking inspiration from the influence of label smoothing on model confidence a…

Cited by 0SourceScholar
2026

Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation

CVPR 2026

Recent vision-language-action (VLA) models for multi-task robot manipulation often rely on fixed camera setups and shared visual encoders, which limit their performance under occlusions and during cross-task transfer. To address these challenges, we propose Task-aware Virtual View Exploration (TVVE)

Cited by 0SourcecodeScholar
2026

LookasideVLN: Direction-Aware Aerial Vision-and-Language Navigation

CVPR 2026

Aerial Vision-and-Language Navigation (Aerial VLN) enables unmanned aerial vehicles (UAVs) to follow natural language instructions and navigate complex urban environments.While recent advances have achieved progress through large-scale memory graphs and lookahead path planning, they remain limited b

Cited by 0SourceScholar
2026

MAGIC: Mastering Physical Adversarial Generation in Context Through Collaborative LLM Agents

AAAI 2026technical

Physical adversarial attacks in driving scenarios can expose critical vulnerabilities in visual perception models. However, developing such attacks remains non-trivial due to diverse real-world environmental influences. Existing approaches either struggle to generalize to dynamic environments or fai

Cited by 0SourcePDFScholar
2026

MAS$^2$: Self-Generative, Self-Configuring, Self-Rectifying Multi-Agent Systems

ICLR 2026poster

The past two years have witnessed the meteoric rise of Large Language Model (LLM)-powered multi-agent systems (MAS), which harness collective intelligence and exhibit a remarkable trajectory toward self-evolution. This paradigm has rapidly progressed from manually engineered systems that require bes…

Cited by 0SourcecodeScholar
2026

MIDAS: Multi-Image Dispersion and Semantic Reconstruction for Jailbreaking MLLMs

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have achieved remarkable performance but remain vulnerable to jailbreak attacks that can induce harmful content and undermine their secure deployment. Previous studies have shown that introducing additional inference steps, which disrupt security attention, c…

Cited by 0SourcecodeScholar
2026

MathFimer: Enhancing Mathematical Reasoning by Expanding Reasoning Steps through Fill-in-the-Middle Task

ICLR 2026poster

Mathematical reasoning represents a critical frontier in advancing large language models (LLMs). While step-by-step approaches have emerged as the dominant paradigm for mathematical problem-solving in LLMs, the quality of reasoning steps in training data fundamentally constrains model performance. R…

Cited by 0SourceScholar
2026

Memoria-Bench: A Comprehensive Benchmark for Evaluating Memory in Long-Horizon Autonomous Agents

ICML 2026poster

Memory is a core capability of autonomous agents, yet existing benchmarks evaluate it primarily in constrained settings such as short dialogues or synthetic tasks, failing to reflect realistic agent deployments. We present \textbf{Memoria-Bench}, a benchmark for evaluating agent memory grounded in c…

Cited by 0SourceScholar
2026

Message Tuning Outshines Graph Prompt Tuning: A Prismatic Space Perspective

ICML 2026poster

Graph Foundation Models (GFMs), built upon the *Pre-training and Adaptation* paradigm, have emerged as a research hotspot in graph learning. For GNN-based GFMs, graph prompt tuning has become the prevailing adaptation method for downstream tasks. Although recent methods explain why graph prompt tuni…

Cited by 0SourceScholar
2026

Modeling Trend Dynamics with Variational Neural ODEs for Information Popularity Prediction

AAAI 2026technical

Predicting the future popularity of information in online social networks is a crucial yet challenging task, due to the complex spatiotemporal dynamics underlying information diffusion. Existing methods typically use structural or sequential patterns within the observation window as direct inputs fo

Cited by 0SourcePDFScholar
2026

NPRIP: Nucleus-to-Periphery Retrieval-Iterative Prompting for Improved Abstractive Summarization in Low-Resource Mongolian

IJCAI 2026

Large language models often face challenges in low-resource agglutinative language text summarization tasks due to poorly designed prompts, leading to core information dilution, reduced fidelity, and critical information loss caused by the complex grammatical structures of agglutinative languages. F

Cited by 0Scholar
2026

Now You See That: Learning End-to-End Humanoid Locomotion from Raw Pixels

RSS 2026poster

Achieving robust vision-based humanoid locomotion remains challenging due to two fundamental issues: the sim-toreal gap introduces significant perception noise that degrades performance on fine-grained tasks, and training a unified policy across diverse terrains is hindered by conflicting learning o…

Cited by 0SourceScholar
2026

Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search

ICLR 2026poster

As Large Language Models (LLMs) are increasingly used, their security risks have drawn increasing attention. Existing research reveals that LLMs are highly susceptible to jailbreak attacks, with effectiveness varying across language contexts. This paper investigates the role of classical Chinese in…

Cited by 0SourcecodeScholar
2026

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding

CVPR 2026

Video Temporal Grounding (VTG), the task of localizing video segments from text queries, struggles in open-world settings due to limited dataset scale and semantic diversity, causing performance gaps between common and rare concepts. To overcome these limitations, we introduce OmniVTG, a new large-s

Cited by 0SourcecodeScholar
2026

PCSR: Pseudo-label Consistency-Guided Sample Refinement for Noisy Correspondence Learning

AAAI 2026technical

Cross-modal retrieval aims to align different modalities via semantic similarity. However, existing methods often assume that image-text pairs are perfectly aligned, overlooking Noisy Correspondences in real data. These misaligned pairs misguide similarity learning and degrade retrieval performance.

Cited by 0SourcePDFScholar
2026

PE-SGD: Differentially Private Deep Learning via Evolution of Gradient Subspace for Text

ICLR 2026poster

Differentially Private Stochastic Gradient Descent (DP-SGD) and its variants like DP-Adam ensure data privacy by injecting noise into per-sample gradients. Although effective with large private datasets, their performance degrades significantly when private training data is limited. Recent works lev…

Cited by 0SourcecodeScholar
2026

PhyScene3D: Physically Consistent 3D Interactive Tabletop Scene Generation

ICML 2026poster

Generating physically consistent 3D tabletop scenes is a fundamental yet underexplored problem for interactive and generalist robotic learning. The challenge stems from dense object hierarchies and irregular affordances. Existing methods, ranging from decoupled symbolic solvers to end-to-end regress…

Cited by 0SourceScholar
2026

PhysPatch: A Physically Realizable and Transferable Adversarial Patch Attack for Multimodal Large Language Models-based Autonomous Driving Systems

AAAI 2026technical

Multimodal Large Language Models (MLLMs) are becoming integral to autonomous driving (AD) systems due to their strong vision-language reasoning capabilities. However, MLLMs are vulnerable to adversarial attacks—particularly adversarial patch attacks—which can pose serious threats in real-world scen

Cited by 0SourcePDFScholar
2026

Probing How Scalable Table Data Enhances General Long-Context Reasoning

ICML 2026poster

As real-world tasks grow increasingly complex, long-context reasoning has become a core capability for Large Language Models (LLMs). However, few studies explore which data types are effective for long-context reasoning and why. We find that structured table data with periodic structures shows stron…

Cited by 0SourceScholar
2026

ProtSAE: Disentangling and Interpreting Protein Language Models via Semantically-Guided Sparse Autoencoders

AAAI 2026technical

Sparse Autoencoder (SAE) has emerged as a powerful tool for mechanistic interpretability of large language models. Recent works apply SAE to protein language models (PLMs), aiming to extract and analyze biologically meaningful features from their latent spaces. However, SAE suffers from semantic en

Cited by 0SourcePDFScholar
2026

RABot: Reinforcement-Guided Graph Augmentation for Imbalanced and Noisy Social Bot Detection

AAAI 2026technical

Social bot detection is pivotal for safeguarding the integrity of online information ecosystems. Although recent graph neural network (GNN) solutions achieve strong results, they remain hindered by two practical challenges: (i) severe class imbalance arising from the high cost of generating bots, an

Cited by 0SourcePDFScholar
2026

ROBUST MULTIMODAL REPRESENTATION LEARNING IN HEALTHCARE

ICASSP 2026poster

Medical multimodal representation learning aims to integrate heterogeneous data into unified patient representations to support clinical outcome prediction. However, real-world medical datasets commonly contain systematic biases from multiple sources, which poses significant challenges for medical m…

Cited by 0SourcePDFScholar
2026

ROBUST MULTIMODAL REPRESENTATION LEARNING IN HEALTHCARE

ICASSP 2026poster

Medical multimodal representation learning aims to integrate heterogeneous data into unified patient representations to support clinical outcome prediction. However, real-world medical datasets commonly contain systematic biases from multiple sources, which poses significant challenges for medical m…

Cited by 0SourcePDFScholar
2026

ReFocusEraser: Refocusing for Small Object Removal with Robust Context-Shadow Repair

ICLR 2026poster

Existing diffusion-based object removal and inpainting methods often fail to recover the fine structural and textural details of small objects. This is primarily due to the VAE encoder’s downsampling, which inevitably compresses small masked regions and causes significant detail loss, while the deco…

Cited by 0SourcecodeScholar
2026

Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs

ICML 2026poster

Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in reasoning and generation tasks and are increasingly deployed in real-world applications. However, their explicit chain-of-thought (CoT) mechanism introduces new security risks, making them particularly vulnerable to jailbreak…

Cited by 0SourceScholar
2026

Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs

ICML 2026poster

Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in reasoning and generation tasks and are increasingly deployed in real-world applications. However, their explicit chain-of-thought (CoT) mechanism introduces new security risks, making them particularly vulnerable to jailbreak…

Cited by 0SourceScholar
2026

Recurrent Reasoning with Vision-Language Models for Estimating Long-Horizon Embodied Task Progress

CVPR 2026

Accurately estimating task progress is critical for embodied agents to plan and execute long-horizon, multi-step tasks. Despite promising advances, existing Vision-Language Models (VLMs) based methods primarily leverage their video understanding capabilities, while neglecting their complex reasoning

Cited by 0SourceScholar
2026

RepIt: Steering Language Models with Concept-Specific Refusal Vectors

ICLR 2026poster

Current safety evaluations of language models rely on benchmark-based assessments that may miss targeted vulnerabilities. We present RepIt, a simple and data-efficient framework for isolating concept-specific representations in LM activations. While existing steering methods already achieve high att…

Cited by 0SourcecodeScholar
2026

Robust Learning from Noisily Labeled Long-Tailed Data via Fairness Regularizer

AAAI 2026technical

Both long-tailed and noisily labeled data frequently appear in real-world applications and impose significant challenges for learning. Most prior works treat either problem in an isolated way and do not explicitly consider the coupling effects of the two. Our empirical observation reveals that such

Cited by 0SourcePDFScholar
2026

Rounded or Streamlined Head? Bridging Concept Bottleneck Models and Attribute-Described Object Parts

CVPR 2026

A faithful decision-making process requires models to ground human-understandable concepts both spatially (where they appear in the image) and causally (how they influence the prediction). Recent advances in Vision-Language Models (VLMs) enable concept-level alignment and have inspired Concept Bottl

Cited by 0SourceScholar
2026

Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Models

CVPR 2026

Spatial reasoning focuses on locating target objects based on spatial relations in 3D scenes, which plays a crucial role in developing intelligent embodied agents. Due to the limited availability of 3D scene-language paired data, it is challenging to train models with strong reasoning ability from s

Cited by 0SourcecodeScholar
2026

SceneTransporter: Optimal Transport-Guided Compositional Latent Diffusion for Single-Image Structured 3D Scene Generation

ICLR 2026poster

We introduce SceneTransporter, an end-to-end framework for structured 3D scene generation from a single image. While existing methods generate part-level 3D objects, they often fail to organize these parts into distinct instances in open-world scenes. Through a debiased clustering probe, we reveal a…

Cited by 0SourcecodeScholar
2026

Shedding Light on VLN Robustness: A Black-box Framework for Indoor Lighting-based Adversarial Attack

CVPR 2026

Vision-and-Language Navigation (VLN) agents have made remarkable progress, but their robustness remains insufficiently studied. Existing adversarial evaluations often rely on perturbations that manifest as unusual textures rarely encountered in everyday indoor environments. Errors under such contriv

Cited by 0SourcecodeScholar
2026

Stabilizing Self-Consuming Diffusion Models with Latent Space Filtering

AAAI 2026technical

As synthetic data proliferates across the Internet, it is often reused to train successive generations of generative models. This creates a "self-consuming loop" that can lead to training instability or *model collapse*. Common strategies to address the issue---such as accumulating historical traini

Cited by 0SourcePDFScholar
2026

SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety Perspectives

ICML 2026poster

We introduce STEERINGSAFETY, a benchmark for evaluating representation steering methods across nine safety perspectives spanning 18 datasets. While prior work highlights general capabilities of representation steering, we focus on safety perspectives including bias, harmfulness, hallucination, socia…

Cited by 0SourceScholar
2026

THEMIS: Towards Holistic Evaluation of MLLMs for Scientific Paper Fraud Forensics

ICLR 2026poster

We present **THEMIS**, a novel multi-task benchmark designed to comprehensively evaluate Multimodal Large Language Models (MLLMs) on visual fraud reasoning within real-world academic scenarios. Compared to existing benchmarks, THEMIS introduces three major advancements. (1) **Real-world Scenarios &…

Cited by 0SourcecodeScholar
2026

TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding

ICML 2026poster

Chain-of-thought (CoT) reasoning has proven effective for enhancing problem-solving in large language models. However, when applied to multimodal LLMs (MLLMs), existing CoT approaches suffer from a fundamental limitation: \textit{they perform reasoning entirely in text without accessing visual featu…

Cited by 0SourceScholar
2026

TaRO: Temporal-Aware Reasoning Optimization for Video Temporal Grounding

ICML 2026poster

Multi-modal Large Language Models (MLLMs) have achieved remarkable progress in video temporal grounding (VTG) with the introduction of reinforcement learning (RL) for generating reasoning paths. However, existing models often produce superficial reasoning, such as providing generic video description…

Cited by 0SourceScholar
2026

Taming I2V models for Image HOI Editing: A Cognitive Benchmark and Agentic Self-Correcting Framework

ICML 2026poster

Current image editing methods excels at static attributes but fails at complex Human-Object Interactions (HOI), a critical challenge unaddressed by existing benchmarks that conflate HOI with static attributes, relying on global metrics incapable of simultaneously assessing dynamic interaction validi…

Cited by 0SourceScholar
2026

Task-Adaptive Parameter-Efficient Fine-Tuning for Weather Foundation Models

ICLR 2026poster

While recent advances in machine learning have equipped Weather Foundation Models (WFMs) with substantial generalization capabilities across diverse downstream tasks, the escalating computational requirements associated with their expanding scale increasingly hinder practical deployment. Current Par…

Cited by 0SourceScholar
2026

Temporal Interaction in Spiking Transformers with Multi-Delay Mixer

CVPR 2026

Spiking Neural Networks (SNNs) have gained significant attention due to their event-driven computational paradigm, making them promising for neuromorphic computing. In recent years, the integration of SNNs and Transformer architectures has made remarkable progress in various tasks. However, existing

Cited by 0SourceScholar
2026

Temporal-Consistent Video Restoration with Pre-trained Diffusion Models

AAAI 2026technical

Video restoration (VR) aims to recover high-quality videos from degraded ones. Although recent zero-shot VR methods using pre-trained diffusion models (DMs) show good promise, they suffer from approximation errors during reverse diffusion and insufficient temporal consistency. Moreover, dealing with

Cited by 0SourcePDFScholar
2026

TianQuan-S2S: A Subseasonal-to-Seasonal Global Weather Model via Incorporate Climatology State

ICLR 2026poster

Accurate Subseasonal-to-Seasonal (S2S) forecasting is vital for decision-making in agriculture, energy production, and emergency management. However, it remains a challenging and underexplored problem due to the chaotic nature of the weather system. Recent data-driven studies have shown promising re…

Cited by 0SourcecodeScholar
2026

Towards Robust Multi-Modal Semantic Segmentation with Teacher-Student Framework and Hybrid Prototype Distillation

CVPR 2026

Multimodal semantic segmentation (MMSS) faces significant challenges in real-world applications due to incomplete, degraded, or missing sensor data. To address this, we propose RobustSeg, an efficient teacher-student framework that enhances model robustness under missing-modality conditions while ma

Cited by 0SourceScholar
2026

Trustworthy Federated Label Distribution Learning under Annotation Quality Disparity

ICML 2026poster

Label Distribution Learning (LDL) models supervision as an instance-wise probability distribution, enabling fine-grained learning under inherent ambiguity, but its success relies on high-fidelity label distributions that are costly to obtain and thus often noisy. Motivated by privacy-sensitive appli…

Cited by 0SourceScholar
2026

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models

ICML 2026spotlight

Uniform Discrete Diffusion (UDM) has recently emerged as a promising paradigm for discrete generative modeling; however, its integration with reinforcement learning remains largely unexplored. We observe that naively adapting GRPO to UDM leads to unstable training and marginal performance. To addres…

Cited by 0SourceScholar
2026

Unified Sequence Modeling for Remote Sensing: A Parameter-Efficient Foundation Model via Prompt-Driven Granularity Alignment

IJCAI 2026

Current remote sensing (RS) perception systems suffer from task heterogeneity, necessitating distinct architectures for classification, localization, and reasoning. While vision--language models (VLMs) offer a route toward unification, their computational cost can hinder deployment. In this work, we

Cited by 0Scholar
2026

Uniform Discrete Diffusion with Metric Path for Video Generation

ICLR 2026poster

Continuous-space video generation has advanced rapidly, while discrete approaches lag behind due to error accumulation and long-context inconsistency. In this work, we revisit discrete generative modeling and present Uniform discRete diffuSion with metric pAth (URSA), a simple yet powerful framework…

Cited by 0SourcecodeScholar
2026

Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs

ICML 2026poster

Unlearning in large language models (LLMs) aims to remove specified data, but its efficacy is typically assessed with task-level metrics like accuracy and perplexity. We demonstrate that these metrics are often misleading, as models can appear to forget while their original behavior is easily restor…

Cited by 0SourcecodeScholar
2026

VDFE: Difference-Aware 3D Scene Editing with Non-Intrusive Video Diffusion Priors for Multi-View Consistency and Efficiency

CVPR 2026

Text-driven 3D editing, enabled by advancements in 3D reconstruction techniques such as NeRF and 3D Gaussian Splatting, aims to provide intuitive scene customization. However, existing methods frequently exhibit limitations in controllability and consistency. To address these shortcomings, we propos

Cited by 0SourceScholar
2026

VIRUS: Injecting Persistent Cognitive Pathogens into Stateful Zero-Shot Object Navigation Agents

ICML 2026poster

Zero-Shot Object Navigation (ZSON) agents rely on continuously updated internal states to support long-horizon planning and decision-making. However, existing methods heavily depend on the observational outputs of vision-language models (VLMs) during state updates and lack explicit validation of per…

Cited by 0SourceScholar
2026

VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models

ICLR 2026poster

Large reasoning models such as OpenAI o1 and DeepSeek-R1 have demonstrated remarkable performance in complex reasoning tasks. A critical component of their training is the incorporation of reference-based reward systems within reinforcement learning (RL), where model outputs are evaluated against gr…

Cited by 0SourcecodeScholar
2026

VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning

CVPR 2026

Most of the multi-agent video understanding frameworks adopt static and non-learnable tool invocation mechanisms, which limit the discovery of diverse clues essential for robust perception and reasoning regarding temporally or spatially complex videos. To address this challenge, we propose a novel M

Cited by 0SourceScholar
2026

Visual-Friendly Concept Protection via Selective Adversarial Perturbations

AAAI 2026technical

Personalized concept generation by tuning diffusion models with a few images raises potential legal and ethical concerns regarding privacy and intellectual property rights. Researchers attempt to prevent malicious personalization using adversarial perturbations. However, previous efforts have mainly

Cited by 0SourcePDFScholar
2026

WeatherSyn: An Instruction Tuning MLLM For Weather Forecasting Report Generation

ICML 2026poster

Accurate weather forecast reporting enables individuals and communities to better plan daily activities, agricultural operations, and transportation. However, the current reporting process primarily relies on manual analysis of multi-source data, which often leads to information overload and reduced…

Cited by 0SourceScholar
2026

Weight Decay may matter more than µP for Learning Rate Transfer in Practice

ICLR 2026poster

Transferring the optimal learning rate from small to large neural networks can enable efficient training at scales where hyperparameter tuning is otherwise prohibitively expensive. To this end, the Maximal Update Parameterization (µP) proposes a learning rate scaling designed to keep the update dyna…

Cited by 0SourcecodeScholar
2025

3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment

ICRA 2025

The 3D weakly-supervised visual grounding task aims to localize oriented 3D boxes in point clouds based on natural language descriptions without requiring annotations to guide model learning. This setting presents two primary challenges: category-level ambiguity and instance-level complexity. Catego

Cited by 2SourceScholar
2025

ACC-Collab: An Actor-Critic Approach to Multi-Agent LLM Collaboration

ICLR 2025poster

Large language models (LLMs) have demonstrated a remarkable ability to serve as general-purpose tools for various language-based tasks. Recent works have demonstrated that the efficacy of such models can be improved through iterative dialog between multiple models. While these paradigms show…

Cited by 0SourcePDFScholar
2025

AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning

ICCV 2025poster

Visual Robot Manipulation (VRM) aims to enable a robot to follow natural language instructions based on robot states and visual observations, and therefore requires costly multi- modal data. To compensate for the deficiency of robot data, existing approaches have employed vision-language pre- traini…

2025

ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models

ACL 2025long

Active perception, a crucial human capability, involves setting a goal based on the current understanding of the environment and performing actions to achieve that goal. Despite significant efforts in evaluating Multimodal Large Language Models (MLLMs), active perception has been largely overlooked.…

2025

AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization

CVPR 2025poster

Recently, model merging methods have demonstrated powerful strengths in combining abilities on various tasks from multiple Large Language Models (LLMs). While previous model merging methods mainly focus on merging homogeneous models with identical architecture, they meet challenges when dealing with…

2025

AdsQA: Towards Advertisement Video Understanding

ICCV 2025poster

Large language models (LLMs) have taken a great step towards AGI. Meanwhile, an increasing number of domain-specific problems such as math and programming boost these general-purpose models to continuously evolve via learning deeper expertise. Now is thus the time further to extend the diversity of…

2025

Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment

NeurIPS 2025poster

Multimodal large language models (MLLMs) remain vulnerable to transferable adversarial examples. While existing methods typically achieve targeted attacks by aligning global features—such as CLIP’s [CLS] token—between adversarial and target samples, they often overlook the rich local information enc…

Cited by 0SourcecodeScholar
2025

Adversarial Robust Memory-Based Continual Learner

ICCV 2025poster

Despite the remarkable advances that have been made in continual learning, the adversarial vulnerability of such methods has not been fully discussed. We delve into the adversarial robustness of memory-based continual learning algorithms and observe limited robustness improvement by directly applyin…

2025

Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text Matching

ICCV 2025poster

Enabling Visual Semantic Models to effectively handle multi-view description matching has been a longstanding challenge. Existing methods typically learn a set of embeddings to find the optimal match for each view's text and compute similarity. However, the visual and text embeddings learned through…

2025

Apply Hierarchical-Chain-of-Generation to Complex Attributes Text-to-3D Generation

CVPR 2025poster

Recent text-to-3D generation models have demonstrated remarkable abilities in producing high-quality 3D assets. Despite their great advancements, current models struggle to generate satisfying 3D objects with complex attributes. The difficulty for such complex attributes 3D generation arises from tw…

2025

Asymmetric Visual Semantic Embedding Framework for Efficient Vision-Language Alignment

AAAI 2025technical

Learning visual semantic similarity is a critical challenge in bridging the gap between images and texts. However, there exist inherent variations between vision and language data, such as information density, i.e., images can contain textual information from multiple different views, which makes it…

2025

Balancing Preservation and Modification: A Region and Semantic Aware Metric for Instruction-Based Image Editing

ICML 2025poster

Instruction-based image editing, which aims to modify the image faithfully towards instruction while preserving irrelevant content unchanged, has made advanced progresses. However, there still lacks a comprehensive metric for assessing the editing quality. Existing metrics either require high costs…

2025

Beyond the Destination: A Novel Benchmark for Exploration-Aware Embodied Question Answering

ICCV 2025poster

Embodied Question Answering (EQA) is a challenging task in embodied intelligence that requires agents to dynamically explore 3D environments, actively gather visual information, and perform multi-step reasoning to answer questions. However, current EQA approaches suffer from critical limitations in…

2025

Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal Representations

NeurIPS 2025poster

The growing scale of evaluation tasks has led to the widespread adoption of automated evaluation using LLMs, a paradigm known as “LLM-as-a-judge”. However, improving its alignment with human preferences without complex prompts or fine-tuning remains challenging. Previous studies mainly optimize base…

Cited by 0SourceScholar
2025

COSMIC: Generalized Refusal Direction Identification in LLM Activations

ACL 2025finding

Large Language Models encode behaviors like refusal within their activation space, but identifying these behaviors remains challenging. Existing methods depend on predefined refusal templates detectable in output tokens or manual review. We introduce **COSMIC** (Cosine Similarity Metrics for Inversi…

2025

CPSea: Large-scale cyclic peptide-protein complex dataset for machine learning in cyclic peptide design

NeurIPS 2025poster

Cyclic peptides exhibit better binding affinity and proteolytic stability compared to their linear counterparts. However, the development of cyclic peptide design models is hindered by the scarcity of data. To address this, we introduce **CPSea**(**C**yclic **P**eptide **Sea**), a dataset of 2.71 mi…

Cited by 0SourcecodeScholar
2025

CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation

ACL 2025finding

Large language models (LLMs) have shown great potential in natural language processing tasks, but their application to machine translation (MT) remains challenging due to pretraining on English-centric data and the complexity of reinforcement learning from human feedback (RLHF). Direct Preference Op…

Cited by 0SourcePDFScholar
2025

Can Large Language Models Derive High-Level Cognition from Low-Level and Fragmented Foundational Information?

AAAI 2025technical

As one of the key technologies leading to Artificial General Intelligence (AGI), Large Language Models (LLMs) have achieved remarkable accomplishments. Exploring the capabilities of LLMs is crucial for scientific research, and many studies propose new challenges from various aspects to explore the b…

2025

CirT: Global Subseasonal-to-Seasonal Forecasting with Geometry-inspired Transformer

ICLR 2025poster

Accurate Subseasonal-to-Seasonal (S2S) climate forecasting is pivotal for decision-making including agriculture planning and disaster preparedness but is known to be challenging due to its chaotic nature. Although recent data-driven models have shown promising results, their performance is limited b…

2025

CityGaussianV2: Efficient and Geometrically Accurate Reconstruction for Large-Scale Scenes

ICLR 2025poster

Recently, 3D Gaussian Splatting (3DGS) has revolutionized radiance field reconstruction, manifesting efficient and high-fidelity novel view synthesis. However, accurately representing surfaces, especially in large and complex scenarios, remains a significant challenge due to the unstructured nature…

Cited by 5SourcePDFScholar
2025

CoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language Models

CVPR 2025poster

Vision-Language Models (VLMs) have recently witnessed significant progress in visual comprehension. As the permitting length of image context grows, VLMs can now comprehend a broader range of views and spaces. Current benchmarks provide insightful analysis of VLMs in tasks involving complex visual i…

2025

CodexGraph: Bridging Large Language Models and Code Repositories via Code Graph Databases

NAACL 2025long

Large Language Models (LLMs) excel in stand-alone code tasks like HumanEval and MBPP, but struggle with handling entire code repositories. This challenge has prompted research on enhancing LLM-codebase interaction at a repository scale. Current solutions rely on similarity-based retrieval or manual…

2025

Cognitive Load Monitoring via Earable Acoustic Sensing

ICASSP 2025accepted

The rapid adoption of ear-worn devices (earables) has shown significant potential for continuous health monitoring. Despite their close proximity to the human brain and diverse sensing capabilities, the exploration of earable sensing in relation to cognitive function remains underexplored. Building…

Cited by 0SourceScholar
2025

Compiler-R1: Towards Agentic Compiler Auto-tuning with Reinforcement Learning

NeurIPS 2025poster

Compiler auto-tuning optimizes pass sequences to improve performance metrics such as Intermediate Representation (IR) instruction count. Although recent advances leveraging Large Language Models (LLMs) have shown promise in automating compiler tuning, two significant challenges still remain: the abs…

Cited by 0SourcecodeScholar
2025

ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion Transfer

CVPR 2025poster

The development of Text-to-Video (T2V) generation has made motion transfer possible, enabling the control of video motion based on existing footage. However, current methods have two limitations: 1) struggle to handle multi-subjects videos, failing to transfer specific subject motion; 2) struggle to…

2025

Conformal Anomaly Detection in Event Sequences

ICML 2025poster

Anomaly detection in continuous-time event sequences is a crucial task in safety-critical applications. While existing methods primarily focus on developing a superior test statistic, they fail to provide guarantees regarding the false positive rate (FPR), which undermines their reliability in pract…

Cited by 0SourcePDFScholar
2025

Contrastive Private Data Synthesis via Weighted Multi-PLM Fusion

ICML 2025poster

Substantial quantity and high quality are the golden rules of making a good training dataset with sample privacy protection equally important. Generating synthetic samples that resemble high-quality private data while ensuring Differential Privacy (DP), a formal privacy guarantee, promises scalabili…

2025

Crabs: Consuming Resource via Auto-generation for LLM-DoS Attack under Black-box Settings

ACL 2025finding

Large Language Models (LLMs) have demonstrated remarkable performance across diverse tasks yet still are vulnerable to external threats, particularly LLM Denial-of-Service (LLM-DoS) attacks. Specifically, LLM-DoS attacks aim to exhaust computational resources and block services. However, existing st…

2025

Cross-modal Causal Relation Alignment for Video Question Grounding

CVPR 2025highlight

Video question grounding (VideoQG) requires models to answer the questions and simultaneously infer the relevant video segments to support the answers. However, existing VideoQG methods usually suffer from spurious cross-modal correlations, leading to a failure to identify the dominant visual scenes…

2025

DAPO : Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage-Based Policy Optimization

NeurIPS 2025spotlight

The role of reinforcement learning (RL) in enhancing the reasoning of large language models (LLMs) is becoming increasingly significant. Despite the success of RL in many scenarios, there are still many challenges in improving the reasoning of LLMs. One key challenge is the sparse reward, which intr…

Cited by 0SourceScholar
2025

DRDM: A Disentangled Representations Diffusion Model for Synthesizing Realistic Person Images

ICASSP 2025accepted

Person image synthesis with controllable body poses and appearances is an essential task owing to the practical needs in the context of virtual try-on, image editing and video production. However, existing methods face significant challenges with details missing, limbs distortion and the garment sty…

Cited by 0SourceScholar
2025

DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering

CVPR 2025poster

3D Question Answering (3D QA) requires the model to comprehensively understand its situated 3D scene described by the text, then reason about its surrounding environment and answer a question under that situation. However, existing methods usually rely on global scene perception from pure 3D point c…

2025

DanmakuTPPBench: A Multi-modal Benchmark for Temporal Point Process Modeling and Understanding

NeurIPS 2025poster

We introduce DanmakuTPPBench, a comprehensive benchmark designed to advance multi-modal Temporal Point Process (TPP) modeling in the era of Large Language Models (LLMs). While TPPs have been widely studied for modeling temporal event sequences, existing datasets are predominantly unimodal, hinderin…

Cited by 0SourcecodeScholar
2025

DeMo: Decoupled Feature-Based Mixture of Experts for Multi-Modal Object Re-Identification

AAAI 2025technical

Multi-modal object Re-IDentification (ReID) aims to retrieve specific objects by combining complementary information from multiple modalities. Existing multi-modal object ReID methods primarily focus on the fusion of heterogeneous features. However, they often overlook the dynamic quality changes in…

2025

Defending LVLMs Against Vision Attacks Through Partial-Perception Supervision

ICML 2025poster

Recent studies have raised significant concerns regarding the vulnerability of Large Vision Language Models (LVLMs) to maliciously injected or perturbed input images, which can mislead their responses. Existing defense methods show that such vision attacks are sensitive to image modifications especi…

Cited by 0SourcePDFScholar
2025

Depth-PC: Sim-to-Real Transfer for Zero-Shot Visual Servoing via Cross-Modal Fusion

RA-L 2025

Visual servoing techniques guide robotic motion using visual information to accomplish manipulation tasks, requiring high precision and robustness against noise. Traditional methods often require prior knowledge and are susceptible to external disturbances. Learning-driven alternatives, while promis

Cited by 0SourcecodeScholar
2025

DepthVanish: Optimizing Adversarial Interval Structures for Stereo-Depth-Invisible Patches

NeurIPS 2025poster

Stereo depth estimation is a critical task in autonomous driving and robotics, where inaccuracies (such as misidentifying nearby objects as distant) can lead to dangerous situations. Adversarial attacks against stereo depth estimation can help revealing vulnerabilities before deployment. Previous wo…

Cited by 0SourcecodeScholar
2025

DexMGNet: Multi-Mode Dexterous Grasping in Cluttered Scenes With Generative Models

RA-L 2025

Dexterous grasping is a crucial technique in humanoid robot manipulation. However, existing methods still fall short in effectively detecting dexterous grasps in cluttered environments. In this work, we propose DexMGNet, a novel multi-mode dexterous grasping framework designed for such challenging s

Cited by 1SourceScholar
2025

DiffTell: A High-Quality Dataset for Describing Image Manipulation Changes

ICCV 2025poster

The image difference captioning (IDC) task is to describe the distinctions between two images. However, existing datasets do not offer comprehensive coverage across all image-difference categories. In this work, we introduce a high-quality dataset, DiffTell with various types of image manipulations,…

Cited by 0SourcePDFScholar
2025

DisTime: Distribution-based Time Representation for Video Large Language Models

ICCV 2025poster

Despite advances in general video understanding, Video Large Language Models (Video-LLMs) face challenges in precise temporal localization due to discrete time representations and limited temporally aware datasets. Existing methods for temporal expression either conflate time with text-based numeric…

2025

Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs

ICLR 2025oral

Large Language Models (LLMs) are increasingly deployed as chatbots, yet their ability to personalize responses to user preferences remains limited. We introduce PrefEval, a benchmark for evaluating LLMs' ability to infer, memorize and adhere to user preferences in long-context conversational setting…

2025

Do Large Language Models excel in Complex Logical Reasoning with Formal Language?

EMNLP 2025

Large Language Models (LLMs) have been shown to achieve breakthrough performances on complex logical reasoning tasks. Nevertheless, most existing research focuses on employing formal language to guide LLMs for deriving reliable reasoning paths, with systematic evaluations of these capabilities still

2025

DoGA: Enhancing Grounded Object Detection via Grouped Pre-Training with Attributes

AAAI 2025technical

Recent advances in vision-language pre-training have significantly enhanced the model capabilities on grounded object detection. However, these studies often pre-train with coarse-grained text prompts, such as plain category names and brief grounded phrases. This limitation curtails the model's capa…

2025

Domain-aware Node Representation Learning for Graph Out-of-Distribution Generalization

ICASSP 2025accepted

Graph Neural Networks (GNNs) have demonstrated impressive success across diverse fields when data satisfies in-distribution (ID) assumption. Nevertheless, GNN performance significantly declines in cases of distribution shifts between training and testing graph data. This degradation primarily stems…

Cited by 0SourceScholar
2025

DongbaMIE: A Multimodal Information Extraction Dataset for Evaluating Semantic Understanding of Dongba Pictograms

EMNLP 2025

Dongba pictographic is the only pictographic script still in use in the world. Its pictorial ideographic features carry rich cultural and contextual information. However, due to the lack of relevant datasets, research on semantic understanding of Dongba hieroglyphs has progressed slowly. To this end

2025

Dynamic Graph Learning with Static Relations for Credit Risk Assessment

AAAI 2025technical

Credit risk assessment has increasingly become a prominent research field due to the dramatically increased incidents of financial default. Traditional graph-based methods have been developed to detect defaulters within user-merchant commercial payment networks. However, these methods face challenge…

Cited by 0SourcePDFScholar
2025

Efficient Dynamic Clustering-Based Document Compression for Retrieval-Augmented-Generation

EMNLP 2025

Retrieval-Augmented Generation (RAG) has emerged as a widely adopted approach for knowledge injection during large language model (LLM) inference in recent years. However, due to their limited ability to exploit fine-grained inter-document relationships, current RAG implementations face challenges i

2025

Efficient Universal Goal Hijacking with Semantics-guided Prompt Organization

ACL 2025long

Universal goal hijacking is a kind of prompt injection attack that forces LLMs to return a target malicious response for arbitrary normal user prompts. The previous methods achieve high attack performance while being too cumbersome and time-consuming. Also, they have concentrated solely on optimizat…

2025

Efficiently Editing Mixture-of-Experts Models with Compressed Experts

EMNLP 2025

Mixture-of-Experts (MoE) models have become a key approach for scaling large language models efficiently by activating only a subset of experts during training and inference. Typically, the number of activated experts presents a trade-off: fewer experts reduce computational costs, while more experts

Cited by 0SourcePDFScholar
2025

End-to-End Driving with Online Trajectory Evaluation via BEV World Model

ICCV 2025poster

End-to-end autonomous driving has achieved remarkable progress by integrating perception, prediction, and planning into a fully differentiable framework. Yet, to fully realize its potential, an effective online trajectory evaluation is indispensable to ensure safety. By forecasting the future outcom…

2025

Enhancing Safety and Manipulability of Redundant Manipulators: Accelerated Motion Generation in Dynamic Environments

RA-L 2025

Motion generation in dynamic environments is crucial for human-machine interaction with redundant manipulators. In this context, we propose the Enhancing Safety and Manipulability (ESM) scheme, which integrates geometry-based dynamic obstacle avoidance, manipulability optimization, trajectory tracki

Cited by 1SourceScholar
2025

Equilibrium Policy Generalization: A Reinforcement Learning Framework for Cross-Graph Zero-Shot Generalization in Pursuit-Evasion Games

NeurIPS 2025poster

Equilibrium learning in adversarial games is an important topic widely examined in the fields of game theory and reinforcement learning (RL). Pursuit-evasion game (PEG), as an important class of real-world games from the fields of robotics and security, requires exponential time to be accurately sol…

Cited by 0SourceScholar
2025

Erasing Concept Combination from Text-to-Image Diffusion Model

ICLR 2025poster

Advancements in the text-to-image diffusion model have raised security concerns due to their potential to generate images with inappropriate themes such as societal biases and copyright infringements. Current studies have made notable progress in preventing the model from generating images containin…

Cited by 1SourcePDFScholar
2025

Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia

NeurIPS 2025poster

Large Language Model (LLM) agents have demonstrated impressive capabilities for social interaction and are increasingly being deployed in situations where they might engage with both human and artificial agents. These interactions represent a critical frontier for LLM-based agents, yet existing eval…

Cited by 0SourceScholar
2025

Exploring Enhanced Contextual Information for Video-Level Object Tracking

AAAI 2025technical

Contextual information at the video level has become increasingly crucial for visual object tracking. However, existing methods typically use only a few tokens to convey this information, which can lead to information loss and limit their ability to fully capture the context. To address this issue,…

2025

Exploring Structural Degradation in Dense Representations for Self-supervised Learning

NeurIPS 2025poster

In this work, we observe a counterintuitive phenomenon in self-supervised learning (SSL): longer training may impair the performance of dense prediction tasks (e.g., semantic segmentation). We refer to this phenomenon as Self-supervised Dense Degradation (SDD) and demonstrate its consistent presence…

Cited by 0SourcecodeScholar
2025

FaVe: Factored and Verified Search Rationale for Long-form Answer

ACL 2025finding

Targeting long-form question-answering, chain-of-query (CoQ) has been studied, integrating chain-of-thought (CoT) with retrieval-augmented generation. CoQ answers the complex question step-by-step, through simpler subquestions (SQs) from which relevant knowledge is retrieved. By doing so, CoQ aims t…

Cited by 0SourcePDFScholar
2025

FastPERT: Towards Fast Microservice Application Latency Prediction via Structural Inductive Bias over PERT Networks

AAAI 2025technical

The recent surge in popularity of cloud-native applications using microservice architectures has led to a focus on accurate end-to-end latency prediction for proactive resource allocation. Existing models leverage Graph Transformers to Microservice Call Graphs or the Program Evaluation and Review Te…

Cited by 0SourcePDFScholar
2025

FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging

ICCV 2025poster

We present FinMMR, a novel bilingual multimodal benchmark tailored to evaluate the reasoning capabilities of multimodal large language models (MLLMs) in financial numerical reasoning tasks. Compared to existing benchmarks, our work introduces three significant advancements. (1) Multimodality: We met…

Cited by 0SourcePDFScholar
2025

FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and Challenging

ACL 2025long

We introduce **FinanceReasoning**, a novel benchmark designed to evaluate the reasoning capabilities of large reasoning models (LRMs) in financial numerical reasoning problems. Compared to existing benchmarks, our work provides three key advancements. (1) **Credibility**: We update 15.6% of the ques…

2025

FlashAudio: Rectified Flow for Fast and High-Fidelity Text-to-Audio Generation

ACL 2025long

Recent advancements in latent diffusion models (LDMs) have markedly enhanced text-to-audio generation, yet their iterative sampling processes impose substantial computational demands, limiting practical deployment. While recent methods utilizing consistency-based distillation aim to achieve few-step…

2025

From Coarse to Fine: A Matching and Alignment Framework for Unsupervised Cross-View Geo-Localization

AAAI 2025technical

Cross-view geo-localization aims at determining the geographic location of a query image by matching the reference images. The matching pairs can be captured from diverse perspectives, such as those from satellites and drones. Most existing methods are supervised that require input of location-label…

Cited by 0SourcePDFScholar
2025

G2: Guided Generation for Enhanced Output Diversity in LLMs

EMNLP 2025

Large Language Models (LLMs) have demonstrated exceptional performance across diverse natural language processing tasks. However, these models exhibit a critical limitation in output diversity, often generating highly similar content across multiple attempts. This limitation significantly affects ta

2025

GFM-Planner: Perception-Aware Trajectory Planning with Geometric Feature Metric

IROS 2025

Like humans who rely on landmarks for orientation, autonomous robots depend on feature-rich environments for accurate localization. In this paper, we propose the GFM-Planner, a perception-aware trajectory planning framework based on the geometric feature metric, which enhances LiDAR localization acc

Cited by 1SourceScholar
2025

GFPack++: Attention-Driven Gradient Fields for Optimizing 2D Irregular Packing

ICCV 2025poster

2D irregular packing is a classic combinatorial optimization problem with various applications, such as material utilization and texture atlas generation. Due to its NP-hard nature, conventional numerical approaches typically encounter slow convergence and high computational costs. Previous research…

2025

Generating Full-field Evolution of Physical Dynamics from Irregular Sparse Observations

NeurIPS 2025poster

Modeling and reconstructing multidimensional physical dynamics from sparse and off-grid observations presents a fundamental challenge in scientific research. Recently, diffusion-based generative modeling shows promising potential for physical simulation. However, current approaches typically operate…

Cited by 0SourceScholar
2025

Generative Video Diffusion for Unseen Novel Semantic Video Moment Retrieval

AAAI 2025technical

Video moment retrieval (VMR) aims to locate the most likely video moment(s) corresponding to a text query in untrimmed videos. Training of existing methods is limited by the lack of diverse and generalisable VMR datasets, hindering their ability to generalise moment-text associations to queries cont…

Cited by 0SourcePDFScholar
2025

GeoDANO: Geometric VLM with Domain Agnostic Vision Encoder

EMNLP 2025

We introduce GeoDANO, a geometric vision-language model (VLM) with a domain-agnostic vision encoder, for solving plane geometry problems. Although VLMs have been employed for solving geometry problems, their ability to recognize geometric features remains insufficiently analyzed. To address this gap

2025

HSI: A Holistic Style Injector for Arbitrary Style Transfer

CVPR 2025poster

Attention-based arbitrary style transfer methods have gained significant attention recently due to their impressive ability to synthesize style details. However, the point-wise matching within the attention mechanism may overly focus on local patterns such that neglect the remarkable global features…

Cited by 0SourcePDFScholar
2025

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding

ICCV 2025poster

In this paper, we tackle the task of online video temporal grounding (OnVTG), which requires the model to locate events related to a given text query within a video stream. Unlike regular video temporal grounding, OnVTG requires the model to make predictions without observing future frames. As onlin…

2025

How Do Multimodal Large Language Models Handle Complex Multimodal Reasoning? Placing Them in An Extensible Escape Game

ICCV 2025poster

The rapid advancing of Multimodal Large Language Models (MLLMs) has spurred interest in complex multimodal reasoning tasks in the real-world and virtual environment, which require coordinating multiple abilities, including visual perception, visual reasoning, spatial awareness, and target deduction.…

2025

How Sememic Components Can Benefit Link Prediction for Lexico-Semantic Knowledge Graphs?

EMNLP 2025

Link Prediction (LP) aims to predict missing triple information within a Knowledge Graph (KG). Existing LP methods have sought to improve the performance by integrating structural and textual information. However, for lexico-semantic KGs designed to document fine-grained sense distinctions, these ty

2025

Human and AI Perceptual Differences in Image Classification Errors

AAAI 2025technical

Artificial intelligence (AI) models for computer vision trained with supervised machine learning are assumed to solve classification tasks by imitating human behavior learned from training labels. Most efforts in recent vision research focus on measuring the model task performance using standardized…

Cited by 0SourcePDFScholar
2025

HyperCRS: Hypergraph-Aware Multi-Grained Preference Learning to Burst Filter Bubbles in Conversational Recommendation System

ACL 2025finding

The filter bubble is a notorious issue in Recommender Systems (RSs), characterized by users being confined to a limited corpus of information or content that strengthens and amplifies their pre-established preferences and beliefs. Most existing methods primarily aim to analyze filter bubbles in the…

2025

INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning

NeurIPS 2025poster

Large Multimodal Models (LMMs) have made significant breakthroughs with the advancement of instruction tuning. However, while existing models can understand images and videos at a holistic level, they still struggle with instance-level understanding that requires a more fine-grained comprehension an…

Cited by 0SourceScholar
2025

Improved Techniques for Optimization-Based Jailbreaking on Large Language Models

ICLR 2025poster

Large language models (LLMs) are being rapidly developed, and a key component of their widespread deployment is their safety-related alignment. Many red-teaming efforts aim to jailbreak LLMs, where among these efforts, the Greedy Coordinate Gradient (GCG) attack's success has led to a growing intere…

2025

Improving Data Efficiency via Curating LLM-Driven Rating Systems

ICLR 2025poster

Instruction tuning is critical for adapting large language models (LLMs) to downstream tasks, and recent studies have demonstrated that small amounts of human-curated data can outperform larger datasets, challenging traditional data scaling laws. While LLM-based data quality rating systems offer a c…

Cited by 3SourcePDFScholar
2025

Incentivizing LLMs to Self-Verify Their Answers

NeurIPS 2025poster

Large Language Models (LLMs) have demonstrated remarkable progress in complex reasoning tasks through both post-training and test-time scaling laws. While prevalent test-time scaling approaches are often realized by using external reward models to guide the model generation process, we find that onl…

Cited by 0SourcecodeScholar
2025

InversionGNN: A Dual Path Network for Multi-Property Molecular Optimization

ICLR 2025poster

Exploring chemical space to find novel molecules that simultaneously satisfy multiple properties is crucial in drug discovery. However, existing methods often struggle with trading off multiple properties due to the conflicting or correlated nature of chemical properties. To tackle this issue, we i…

2025

Knowledge Image Matters: Improving Knowledge-Based Visual Reasoning with Multi-Image Large Language Models

ACL 2025long

We revisit knowledge-based visual reasoning (KB-VR) in light of modern advances in multimodal large language models (MLLMs), and make the following contributions: (i) We propose Visual Knowledge Card (VKC) – a novel image that incorporates not only internal visual knowledge (e.g., scene-aware inform…

Cited by 0SourcePDFScholar