← Search

Wei Li

345 accepted papers

2026

Artemis: Structured Visual Reasoning for Perception Policy Learning

ICML 2026poster

Recent reinforcement-learning frameworks for visual perception policy usually incorporate intermediate reasoning chains expressed in natural language. Empirical observations indicate that such purely linguistic intermediate reasoning often reduces performance on perception tasks. We argue that the c…

Cited by 0SourceScholar
2026

Azimuth-LIO: Robust LiDAR-Inertial Odometry via Azimuth-Aware Voxelization and Probabilistic Fusion

RA-L 2026

Voxel-based LiDAR–inertial odometry (LIO) is accurate and efficient but can suffer from geometric inconsistencies when single-Gaussian voxel models indiscriminately merge observations from conflicting viewpoints. To address this limitation, we propose Azimuth-LIO, a robust voxel-based LIO framework

Cited by 0SourceScholar
2026

BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation

RSS 2026poster

Equipping embodied agents with the ability to reason about tasks, foresee physical outcomes, and generate precise actions is essential for general-purpose manipulation. While recent Vision-Language-Action (VLA) models have leveraged pre-trained foundation models, they typically focus on either lingu…

Cited by 0SourceScholar
2026

Beyond Sequential Tools: A Unified VLM Agent System for Photographic Post-Processing via Dynamic Multi-Expert Fusion

CVPR 2026

Real-world image restoration is challenged by complex, coupled degradations. Existing "all-in-one" models often lack generalization, while agentic systems suffer from inefficient sequential tool invocation. We propose a VLM-guided one-shot framework for universal photographic post-processing. Our sy

Cited by 0SourcecodeScholar
2026

Branch, or Layer? Zeroth-Order Optimization for Continual Learning of Vision-Language Models

AAAI 2026technical

Vision-Language Continual Learning (VLCL) has attracted significant research attention for its robust capabilities, and the adoption of Parameter-Efficient Fine-Tuning (PEFT) strategies is enabling these models to achieve competitive performance with substantially reduced resource consumption. Howev

Cited by 0SourcePDFScholar
2026

CSD: Content-aware Speculative Decoding for Efficient Image Generation

ICML 2026poster

Speculative decoding (SD) has emerged as a key solution to accelerate the inference of autoregressive models. However, in the field of image generation, it faces the challenge of low acceptance rates, and directly relaxing its criteria leads to degradation in image quality. In this paper, we propose…

Cited by 0SourceScholar
2026

CliCARE: Grounding Large Language Models in Clinical Guidelines for Decision Support over Longitudinal Cancer Electronic Health Records

AAAI 2026technical

Large Language Models (LLMs) hold significant promise for improving clinical decision support and reducing physician burnout by synthesizing complex, longitudinal cancer Electronic Health Records (EHRs). However, their implementation in this critical field faces three primary challenges: the inabili

Cited by 0SourcePDFScholar
2026

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation

CVPR 2026

Current Vision-Language-Action (VLA) models primarily focus on mapping 2D observations to actions but exhibit notable limitations in spatiotemporal perception and reasoning: 1) spatial representations often rely on additional sensors, introducing substantial computational overhead; 2) visual reasoni

Cited by 0SourcecodeScholar
2026

DarkFarseer: Robust Spatio-Temporal Kriging Under Graph Sparsity and Noise

AAAI 2026technical

The rapid expansion of the Internet of Things (IoT) has created a growing demand for large-scale sensor deployment. However, the high cost of physical sensors limits the scalability and coverage of sensor networks, making fine-grained sensing difficult. Inductive Spatio-Temporal Kriging (ISK) addres

Cited by 0SourcePDFScholar
2026

Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing

ICLR 2026poster

In recent years, integrating multimodal understanding and generation into a single unified model has emerged as a promising paradigm. While this approach achieves strong results in text-to-image (T2I) generation, it still struggles with precise image editing. We attribute this limitation to an imbal…

Cited by 0SourceScholar
2026

Entropy-aware Span-Constrained Optimal Transport for Robust Cross-Tokenizer Knowledge Distillation

ICML 2026poster

Existing Cross-Tokenizer Knowledge Distillation (CTKD) methods fail to outperform simple supervised fine-tuning when vocabulary overlap is low due to severe alignment noise. We identify this phenomenon as the **``Low-Overlap negative transfer regime,''** To overcome this, we propose **Entropy-aware …

Cited by 0SourceScholar
2026

Escaping Low-Rank Traps: Interpretable Visual Concept Learning via Implicit Vector Quantization

ICLR 2026poster

Concept Bottleneck Models (CBMs) achieve interpretability by interposing a human-understandable concept layer between perception and label prediction. The foundation of CBMs lies in the many-to-many mapping that translates high-dimensional visual features to a set of discrete concepts. However, we…

Cited by 0SourceScholar
2026

Failure Detection With Zero-Shot Error Correction in Robotic Manipulation

RA-L 2026

Diffusion Policy (DP) is effective for imitation learning in robotic manipulation, yet likelihood-based replanning methods lack self-correction and often fail once execution deviates. Vision-Language Models (VLMs) offer strong spatial reasoning for failure handling but incur prohibitive latency for

Cited by 0SourceScholar
2026

FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance Customization

AAAI 2026technical

Garment-centric fashion image generation aims to synthesize realistic and controllable human models dressing a given garment, which has attracted growing interest due to its practical applications in e-commerce. The key challenges of the task lie in two aspects: (1) faithfully preserving the garment

Cited by 0SourcePDFScholar
2026

From Semantics to Spectrum: A New Lens on Graph Augmentation Strategy

AAAI 2026technical

Graph augmentation is a cornerstone of effective graph contrastive learning, yet existing methods often rely on random designed perturbations, which may distort latent semantics and impair representation quality. In this work, we argue that semantic consistency can be effectively approximated by low

Cited by 0SourcePDFScholar
2026

From Time Series Analysis to Question Answering: A Survey in the LLM Era

IJCAI 2026

Recently, Large Language Models (LLMs) have introduced a novel paradigm in Time Series Analysis (TSA), leveraging strong language capabilities to support tasks such as forecasting and anomaly detection. However, these analysis tasks cannot adequately cover temporal language tasks, such as interpreta

Cited by 0Scholar
2026

GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and a Comprehensive Multimodal Dataset Towards General Medical AI

AAAI 2026technical

Despite significant advancements in general AI, its effectiveness in the medical domain is limited by the lack of specialized medical knowledge. To address this, we formulate GMAI-VL-5.5M, a multimodal medical dataset created by converting hundreds of specialized medical datasets with various annot

Cited by 0SourcePDFScholar
2026

GRIP: Latent Field-Guided Graph Policy for Budget-Constrained Multi-Agent Routing

AAAI 2026technical

Subset selection under budget constraints is critical in applications like multi-robot patrolling, crime deterrence, and targeted marketing, where multiple agents must jointly select targets and plan feasible routes. We formalize this challenge as Multi-Subset Selection with Budget-Constrained Routi

Cited by 0SourcePDFScholar
2026

Hydra-Nav: Object Navigation via Adaptive Dual-Process Reasoning

ICML 2026poster

While large vision-language models (VLMs) show promise for object goal navigation, current methods still struggle with low success rates and inefficient localization of unseen objects—failures primarily attributed to weak temporal-spatial reasoning. Meanwhile, recent attempts to inject reasoning int…

Cited by 0SourceScholar
2026

I2CD: An Invertible Causal Framework for Compositional Zero-Shot Learning via Disentangle-Compose-Disentangle

AAAI 2026technical

Compositional Zero-Shot Learning (CZSL) addresses the challenge of recognizing unseen attribute-object compositions in images, representing a fundamental challenge in artificial intelligence. Current approaches, which primarily focus on semantic alignment or distribution independence of primitives,

Cited by 0SourcePDFScholar
2026

Image-to-Point Cloud Feature Back-Projection for Multimodal Training of 3D Semantic Segmentation

CVPR 2026

The effective integration and utilization of multimodal data acquired from image cameras and LiDAR is of paramount importance for perception systems. This paper proposes **I**mage-to-**P**oint Cloud **F**eature Back-**P**rojection (**IPFP**), a novel method for training multimodal fusion networks th

Cited by 0SourceScholar
2026

In-Context Generation with Regional Constraints for Instructional Video Editing

ICML 2026poster

The In-context generation paradigm has demonstrated strong power in instructional image editing for better synthesis quality. Nevertheless, shaping such in-context learning for instructional video editing is not trivial. Without specifying editing regions, the results can suffer from the issue of in…

Cited by 0SourceScholar
2026

InfoCom: Kilobyte-Scale Communication-Efficient Collaborative Perception with Information Bottleneck

AAAI 2026technical

Precise environmental perception is critical for the reliability of autonomous driving systems. While collaborative perception mitigates the limitations of single-agent perception through information sharing, it encounters a fundamental communication-performance trade-off. Existing communication-eff

Cited by 0SourcePDFScholar
2026

JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization

ICLR 2026poster

This paper introduces JavisDiT, a novel Joint Audio-Video Diffusion Trans- former designed for synchronized audio-video generation (JAVG). Based on the powerful Diffusion Transformer (DiT) architecture, JavisDiT simultaneously generates high-quality audio and video content from open-ended user promp…

Cited by 0SourcecodeScholar
2026

LAOF: Robust Latent Action Learning with Optical Flow Constraints

CVPR 2026

Learning latent actions from large-scale videos is crucial for the pre-training of scalable embodied foundation models, yet existing methods often struggle with action-irrelevant distractors. Although incorporating action supervision can alleviate these distractions, its effectiveness is restricted

Cited by 0SourcecodeScholar
2026

LearnIR: Learnable Posterior Sampling for Real-World Image Restoration

ICLR 2026poster

Image restoration in real-world conditions is highly challenging due to heterogeneous degradations such as haze, noise, shadows, and blur. Existing diffusion-based methods remain limited: conditional generation struggles to balance fidelity and realism, inversion-based approaches accumulate errors,…

Cited by 0SourcecodeScholar
2026

Light-X: Generative 4D Video Rendering with Camera and Illumination Control

ICLR 2026poster

Recent advances in illumination control extend image-based methods to video, yet still facing a trade-off between lighting fidelity and temporal consistency. Moving beyond relighting, a key step toward generative modeling of real-world scenes is the joint control of camera trajectory and illuminatio…

Cited by 12SourcecodeScholar
2026

Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generation

ICML 2026poster

Reverse Chain-of-Thought Generation (RCG) synthesizes reasoning traces from query-answer pairs, but runs the risk of producing post-hoc rationalizations: when models can see the answer during generation, the answer serves as a cognitive anchor that shapes the entire explanation. We formalize this ph…

Cited by 0SourceScholar
2026

Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search

ICLR 2026poster

Recent advances in large multimodal models have leveraged image-based tools with reinforcement learning to tackle visual problems. However, existing open-source approaches often exhibit monotonous reasoning patterns and allow only a limited number of interaction turns, making them inadequate for dif…

Cited by 0SourcecodeScholar
2026

Monocular Vehicle Pose and Shape Reconstruction via Dynamic Context Adaptation and Progressive Geometry Refinement

AAAI 2026technical

Accurate reconstruction of 3D vehicle pose and shape from monocular images is challenging, particularly for distant objects in autonomous driving. Existing methods often suffer from geometric ambiguity in depth estimation and structural hollowness in shape recovery, primarily due to inadequate multi

Cited by 0SourcePDFScholar
2026

MvP-ECR: Multi-Perspective Emotion-Cause Reasoning for Empathetic Dialogue

AAAI 2026technical

The empathetic dialogue systems aim to recognize user emotions and generate appropriate empathetic responses. However, existing approaches predominantly rely on dialogue history, contextual descriptions, and emotion category labels, failing to model the causal relationship between emotions and their

Cited by 0SourcePDFScholar
2026

Octopus: History-Free Gradient Orthogonalization for Continual Learning in Multimodal Large Language Models

CVPR 2026

Continual learning in multimodal large language models (MLLMs) aims to sequentially acquire knowledge while mitigating catastrophic forgetting, yet existing methods face inherent limitations: architecture-based approaches incur additional computational overhead and often generalize poorly to new tas

Cited by 0SourceScholar
2026

Outlier Matters: Efficient Long-to-Short Reasoning via Outlier-Guided Model Merging

AAAI 2026technical

Large Reasoning Language Models (LRMs) have recently shown remarkable performance in complex reasoning tasks, but their extensive reasoning chains incur substantial computational overhead. To address this challenge, we propose Outlier-aware Reasoning Conciseness Adaptive Merge (ORCA), a novel plug-a

Cited by 0SourcePDFScholar
2026

Prefill-Time Intervention for Mitigating Hallucination in Large Vision-Language Models

CVPR 2026

Large Vision-Language Models (LVLMs) have achieved remarkable progress in visual-textual understanding, yet their reliability is critically undermined by hallucinations, i.e., the generation of factually incorrect or inconsistent responses.While recent studies using steering vectors demonstrated pro

Cited by 0SourceScholar
2026

ReactID: Synchronizing Realistic Actions and Identity in Personalized Video Generation

ICLR 2026poster

Personalized video generation faces a fundamental trade-off between identity consistency and action realism: overly rigid identity preservation often leads to unnatural motion, while emphasis on action dynamics can compromise subject fidelity. This tension stems from three interrelated challenges: i…

Cited by 0SourceScholar
2026

Real Garment Benchmark (RGBench): A Comprehensive Benchmark for Robotic Garment Manipulation Featuring a High-Fidelity Scalable Simulator

AAAI 2026technical

While there has been significant progress to use simulated data to learn robotic manipulation of rigid objects, applying its success to deformable objects has been hindered by the lack of both deformable object models and realistic non-rigid body simulators. In this paper, we present Real Garment Be

Cited by 0SourcePDFScholar
2026

SEMANTIC REFORMULATION ENTROPY FOR ROBUST HALLUCINATION DETECTION IN QA TASKS

ICASSP 2026poster

Reliable question answering with large language models (LLMs) is challenged by hallucinations, fluent but factually incorrect outputs arising from epistemic uncertainty. Existing entropy-based semantic-level uncertainty estimation methods are limited by sampling noise and unstable clustering of vari…

Cited by 0SourcePDFScholar
2026

SPSC: Sparse and Scalable Multi-Modal 3D Occupancy Prediction for Autonomous Driving

AAAI 2026technical

3D semantic occupancy prediction offers a nuanced representation of the surrounding environment, which is crucial for ensuring the safety of autonomous driving. However, fine-grained scene representations inevitably result in cubic growth in data scale, which imposes substantial demands on model arc

Cited by 0SourcePDFScholar
2026

Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory

ICLR 2026poster

We introduce M3-Agent, a novel multimodal agent framework equipped with long-term memory. Like humans, M3-Agent can process real-time visual and auditory inputs to build and update episodic and semantic memories, gradually accumulating world knowledge. Its memory is organized in an entity-centric, m…

Cited by 0SourcecodeScholar
2026

SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation

AAAI 2026technical

Vision-Language-Action (VLA) models have advanced in robotic manipulation, yet practical deployment remains hindered by two key limitations: **1) perceptual redundancy**, where irrelevant visual inputs are processed inefficiently, and **2) superficial instruction-vision alignment**, which hampers se

Cited by 0SourcePDFScholar
2026

Solving Spatial-Spectral Fusion with Latent Spectral Operators

ICML 2026poster

Existing deep spatial–spectral fusion (SSF) methods typically learn the fusion mapping in the coordinate domain using convolutions and attentions, making it hard to scale across varying spatial resolutions and offering limited control over the frequency content of the reconstructions, which may furt…

Cited by 0SourceScholar
2026

Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging

AAAI 2026technical

Mixture of Experts (MoE) LLMs face significant obstacles due to their massive parameter scale, which imposes memory, storage, and deployment challenges. Although recent expert merging methods aim to achieve greater efficiency by consolidating several experts, they are fundamentally hindered by param

Cited by 0SourcePDFScholar
2026

SynthRGB-T: Language-Vision Guided Image Translation for Diversity Synthesis

CVPR 2026

Bridging the modality gap between infrared and visible imagery is critical for cross-modal understanding and for enriching multimodal benchmarks. However, existing approaches remain confined to one-to-one mappings and are typically evaluated on unidirectional or closed-set scenarios. To address this

Cited by 0SourceScholar
2026

Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation

ICLR 2026poster

Camera-centric understanding and generation are two cornerstones of spatial intelligence, yet they are typically studied in isolation. We present Puffin, a unified camera-centric multimodal model that extends spatial awareness along the camera dimension. Puffin integrates language regression and dif…

Cited by 0SourcecodeScholar
2026

Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought Compression

ICML 2026poster

Chain-of-Thought (CoT) reasoning successfully enhances the reasoning capabilities of Large Language Models (LLMs), yet it incurs substantial computational overhead for inference. Existing CoT compression methods often suffer from a critical loss of logical fidelity at high compression ratios, result…

Cited by 0SourceScholar
2026

UniMedVL: Unifying Medical Multimodal Understanding and Generation through Observation-Knowledge-Analysis

ICML 2026poster

Medical diagnosis demands models that can process multimodal medical inputs, such as medical images and patient histories, and generate diverse outputs including textual reports and visual content, such as annotations or segmentation masks. Despite this need, existing medical AI models disrupt this …

Cited by 0SourceScholar
2026

Unifying Language-Action Understanding and Generation for Autonomous Driving

CVPR 2026

Vision-Language-Action (VLA) models are emerging as a promising paradigm for end-to-end autonomous driving, valued for their potential to leverage world knowledge and reason about complex driving scenes. However, existing methods suffer from two critical limitations: a persistent misalignment betwee

Cited by 0SourcecodeScholar
2026

VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset

CVPR 2026

Directly editing ultra-high-resolution (UHR) images is valuable but underexplored, primarily due to the lack of high-quality data and the challenge in modeling high-frequency texture details. We introduce VINS-120K, the first large-scale dataset for instruction-based UHR image editing, comprising 12

Cited by 0SourceScholar
2026

VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language tasks yet remain limited in long video understanding due to the limited context window. Consequently, prevailing approaches tend to rely on uniform frame sampling or static pre-selection, which might overlo…

Cited by 0SourceScholar
2026

Why Does RL Generalize Better Than SFT? A Data-Centric Perspective on VLM Post-Training

CVPR 2026

Post-training of large-scale Vision-Language Models (VLMs) reveals a pronounced generalization gap: models fine-tuned with Reinforcement Learning (RL) consistently achieve superior out-of-distribution (OOD) performance compared to those trained with Supervised Fine-Tuning (SFT). This paper posits a

Cited by 0SourcecodeScholar
2026

video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM

ICML 2026poster

Long-duration streaming video understanding is fundamental for future AI agents, yet remains limited by ineffective long-term memory. We introduce video-SALMONN S, a memory-enhanced streaming audio-visual large language model that processes over 3-hour videos at $1$ FPS and $360$p resolution, outper…

Cited by 0SourceScholar
2025

A Mamba-based Network for Semi-supervised Singing Melody Extraction Using Confidence Binary Regularization

ICASSP 2025accepted

Singing melody extraction (SME) is a key task in the field of music information retrieval. However, existing methods are facing several limitations: firstly, prior models use transformers to capture the contextual dependencies, which requires quadratic computation resulting in low efficiency in the…

Cited by 0SourceScholar
2025

AIRA: Activation-Informed Low-Rank Adaptation for Large Models

ICCV 2025poster

Low-Rank Adaptation (LoRA) is a widely used method for efficiently fine-tuning large models by introducing low-rank matrices into weight updates. However, existing LoRA techniques fail to account for activation information, such as outliers, which significantly impact model performance. This omissio…

2025

AdaCo: Overcoming Visual Foundation Model Noise in 3D Semantic Segmentation via Adaptive Label Correction

AAAI 2025technical

Recently, Visual Foundation Models (VFMs) have shown a remarkable generalization performance in 3D perception tasks. However, their effectiveness in large-scale outdoor datasets remains constrained by the scarcity of accurate supervision signals, the extensive noise caused by variable outdoor cond…

2025

AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents

ICLR 2025poster

Autonomous agents that execute human tasks by controlling computers can enhance human productivity and application accessibility. However, progress in this field will be driven by realistic and reproducible benchmarks. We present AndroidWorld, a fully functional Android environment that provides rew…

2025

Audio-centric Video Understanding Benchmark without Text Shortcut

EMNLP 2025

Audio often serves as an auxiliary modality in video understanding tasks of audio-visual large language models (LLMs), merely assisting in the comprehension of visual information. However, a thorough understanding of videos significantly depends on auditory information, as audio offers critical cont

2025

AugKD: Ingenious Augmentations Empower Knowledge Distillation for Image Super-Resolution

ICLR 2025poster

Knowledge distillation (KD) compresses deep neural networks by transferring task-related knowledge from cumbersome pre-trained teacher models to more compact student models. However, vanilla KD for image super-resolution (SR) networks yields only limited improvements due to the inherent nature of SR…

Cited by 0SourcePDFScholar
2025

BayesKD: Bayesian Knowledge Distillation for Compact LLMs in Constrained Fine-tuning Scenarios

ACL 2025finding

Large language models (LLMs) have revolutionized various domains with their remarkable capabilities, but their massive parameter sizes pose significant challenges for fine-tuning and inference, especially in resource-constrained environments. Conventional compression methods often result in substant…

Cited by 0SourcePDFScholar
2025

BeatKAN: An Efficient and Drum-Attuned Beat Tracking Method Using Kolmogorov-Arnold Networks

ICASSP 2025accepted

In this paper, we propose an efficient and drum-attuned beat tracking method based on Kolmogorov-Arnold networks (KAN). Traditional MLP-based frameworks struggle with complex musical signals due to limited capacity in modeling intricate patterns. Inspired by KAN’s efficient ability to capture comple…

Cited by 0SourceScholar
2025

Breaking Information Isolation: Accelerating MRI via Inter-sequence Mapping and Progressive Masking

AAAI 2025technical

Deep unfolding network (DUN) has shed new light on multi-sequence MRI reconstruction, providing both high interpretability and acceptable performance. However, current approaches still suffer from the plight of information isolation, i.e., learning features of multi-suquences individually and leavin…

Cited by 0SourcePDFScholar
2025

CA2Point: Learning Keypoint Detection and Description with Context Aggregation and Cross Augmentation

IROS 2025

Keypoint detection and description are fundamental tasks for a variety of computer vision applications. Due to the limited receptive field of convolutional neural networks, most existing methods based on deep learning mainly focus on the local features, instead of taking into account the global cont

Cited by 0SourcecodeScholar
2025

CBQ: Cross-Block Quantization for Large Language Models

ICLR 2025spotlight

Post-training quantization (PTQ) has played a pivotal role in compressing large language models (LLMs) at ultra-low costs. Although current PTQ methods have achieved promising results by addressing outliers and employing layer- or block-wise loss optimization techniques, they still suffer from signi…

Cited by 13SourcePDFScholar
2025

ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area

AAAI 2025technical

Large Language Models (LLMs) have achieved remarkable success and have been applied across various scientific fields, including chemistry. However, many chemical tasks require the processing of visual information, which cannot be successfully handled by existing chemical LLMs. This brings a growing…

2025

CoPEFT: Fast Adaptation Framework for Multi-Agent Collaborative Perception with Parameter-Efficient Fine-Tuning

AAAI 2025technical

Multi-agent collaborative perception is expected to significantly improve perception performance by overcoming the limitations of single-agent perception through exchanging complementary information. However, training a robust collaborative perception model requires collecting sufficient training da…

2025

CogVLA: Cognition-Aligned Vision-Language-Action Models via Instruction-Driven Routing & Sparsification

NeurIPS 2025poster

Recent Vision-Language-Action (VLA) models built on pre-trained Vision-Language Models (VLMs) require extensive post-training, resulting in high computational overhead that limits scalability and deployment. Existing sparsification strategies—such as Mixture-of-Depths, layer skipping, and early exit…

Cited by 0SourcecodeScholar
2025

Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank Adaptation

ICML 2025poster

Low-rank adaptation (LoRA) has emerged as a leading parameter-efficient fine-tuning technique for adapting large foundation models, yet it often locks adapters into suboptimal minima near their initialization. This hampers model generalization and limits downstream operators such as adapter merging…

2025

Competitive Fair Scheduling with Predictions

ICLR 2025poster

Beyond the worst-case analysis of algorithms, the learning-augmented framework considers that an algorithm can leverage possibly imperfect predictions about the unknown variables to have guarantees tied to the prediction quality. We consider online non-clairvoyant scheduling to minimize the max-stre…

Cited by 0SourcePDFScholar
2025

Controllable Human-centric Keyframe Interpolation with Generative Prior

NeurIPS 2025poster

Existing interpolation methods use pre‑trained video diffusion priors to generate intermediate frames between sparsely sampled keyframes. In the absence of 3D geometric guidance, these methods struggle to produce plausible results for complex, articulated human motions and offer limited control over…

Cited by 0SourceScholar
2025

Cross-Modality Fusion Mamba for All-in-One Extreme Weather-Degraded Image Restoration

ICASSP 2025accepted

A major obstacle for high-level tasks is the unpredictable image degradation. While several architectures proposed to address this, they fail under extreme degradation. Therefore, we introduce a novel cross-modality pipeline called AIRMamba, designed to holistically and robustly restore images degra…

Cited by 0SourceScholar
2025

Delta Decompression for MoE-based LLMs Compression

ICML 2025poster

Mixture-of-Experts (MoE) architectures in large language models (LLMs) achieve exceptional performance, but face prohibitive storage and memory requirements. To address these challenges, we present $D^2$-MoE, a new delta decompression compressor for reducing the parameters of MoE LLMs. Based on obse…

2025

Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent

ICCV 2025poster

Despite the progress in text-to-image generation, semantic image editing remains a challenge. Inversion-based algorithms unavoidably introduce reconstruction errors, while instruction-based models mainly suffer from limited dataset quality and scale. To address these problems, we propose a descripti…

Cited by 0SourcePDFScholar
2025

DoF-Gaussian: Controllable Depth-of-Field for 3D Gaussian Splatting

CVPR 2025poster

Recent advances in 3D Gaussian Splatting (3D-GS) have shown remarkable success in representing 3D scenes and generating high-quality, novel views in real-time. However, 3D-GS and its variants assume that input images are captured based on pinhole imaging and are fully in focus. This assumption limit…

2025

Dynamic Contrastive Knowledge Distillation for Efficient Image Restoration

AAAI 2025technical

Knowledge distillation (KD) is a valuable yet challenging approach that enhances a compact student network by learning from a high-performance but cumbersome teacher model. However, previous KD methods for image restoration overlook the state of the student during the distillation, adopting a fixed…

2025

Dynamic Graph Convolutional Networks with Spatiotemporal Missing Pattern Awareness

ICASSP 2025accepted

Missing data is ubiquitous phenomenon in the time series community, significantly challenging forecasting due to incomplete ground truth and sparse data. Most previous Multi-variate Time Series Forecasting with Missing Values (MTSFMV) approaches usually assume static missing patterns, neglecting the…

Cited by 0SourceScholar
2025

Dynamic Object Queries for Transformer-based Incremental Object Detection

ICASSP 2025accepted

Incremental object detection (IOD) aims to sequentially learn new classes, while maintaining the capability to locate and identify old ones. Prior methodologies mainly tackle catastrophic forgetting through knowledge distillation and exemplar replay, ignoring the conflict between limited model capac…

Cited by 0SourceScholar
2025

ECC: An Emotion-Cause Conversation Dataset for Empathy Response

EMNLP 2025

The empathy dialogue system requires understanding emotions and their underlying causes. However, existing datasets mainly focus on emotion labels, while cause annotations are added post hoc through costly and subjective manual processes. This leads to three limitations: subjective bias in cause lab

2025

Efficient Fine-Tuning of Large Models via Nested Low-Rank Adaptation

ICCV 2025poster

Low-Rank Adaptation (LoRA) has become a popular paradigm for fine-tuning large models, but it still necessitates a substantial number of training parameters. To address this issue, we first conduct comprehensive empirical studies on parameter-efficient LoRA structure. Then, we establish design guide…

2025

Efficient Spiking Point Mamba for Point Cloud Analysis

ICCV 2025poster

Bio-inspired Spiking Neural Networks (SNNs) provide an energy-efficient way to extract 3D spatio-temporal features. However, existing 3D SNNs have struggled with long-range dependencies until the recent emergence of Mamba, which offers superior computational efficiency and sequence modeling capabili…

2025

Encode Errors: Representational Retrieval of In-Context Demonstrations for Multilingual Grammatical Error Correction

ACL 2025finding

Grammatical Error Correction (GEC) involves detecting and correcting the wrong usage of grammar. While large language models (LLMs) with in-context learning (ICL) capabilities have shown significant progress on various natural language processing (NLP) tasks, their few-shot performance on GEC remain…

Cited by 0SourcePDFScholar
2025

Expectation Confirmation Preference Optimization for Multi-Turn Conversational Recommendation Agent

ACL 2025finding

Recent advancements in Large Language Models (LLMs) have significantly propelled the development of Conversational Recommendation Agents (CRAs). However, these agents often generate short-sighted responses that fail to sustain user guidance and meet expectations. Although preference optimization has…

2025

Explanation based In-Context Demonstrations Retrieval for Multilingual Grammatical Error Correction

NAACL 2025long

Grammatical error correction (GEC) aims to correct grammatical, spelling, and semantic errors in natural language text. With the growing of large language models (LLMs), direct text generation has gradually become the focus of the GEC methods, and few-shot in-context learning presents a cost-effecti…

2025

Exploring Graph-aware Reasoning and Bidirectional Selection for Vision-Language Navigation

ICASSP 2025accepted

Inspired by structured state space models and graph neural network modeling, we proposes a novel graph-aware reasoning (GAR) model to effectively solve the problem between memory utilization efficiency and reasoning navigation. First, we introduce graph networks into the navigation framework to enha…

Cited by 0SourceScholar
2025

F-LMM: Grounding Frozen Large Multimodal Models

CVPR 2025poster

Endowing Large Multimodal Models (LMMs) with visual grounding capability can significantly enhance AIs' understanding of the visual world and their interaction with humans. However, existing methods typically fine-tune the parameters of LMMs to learn additional segmentation tokens and overfit ground…

2025

FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation

ACL 2025long

Implementing new features in repository-level codebases is a crucial application of code generation models. However, current benchmarks lack a dedicated evaluation framework for this capability. To fill this gap, we introduce FEA-Bench, a benchmark designed to assess the ability of large language mo…

2025

Feature Refinement Decomposition and Relation Preference Enhancement for Remote Sensing Change Detection

ICASSP 2025accepted

Remote Sensing Change Detection (RSCD) is essential for identifying alterations within geographical landscapes based on dual-temporal imagery. Current methods often enhance global modeling through the use of Transformers or integrate global and local features in a coarse-grained manner. The former t…

Cited by 0SourceScholar
2025

Free4D: Tuning-free 4D Scene Generation with Spatial-Temporal Consistency

ICCV 2025poster

We present Free4D, a novel tuning-free framework for 4D scene generation from a single image. Existing methods either focus on object-level generation, making scene-level generation infeasible, or rely on large-scale multi-view video datasets for expensive training, with limited generalization abili…

2025

Fuel-Optimal Operational Speed Planning for Autonomous Trucking on Highways

ICRA 2025

The rapid advancement of autonomous driving technology, particularly in autonomous trucking on highways, shows great value for enhancing efficiency and reducing costs in the logistics industry. In this work, we define the full-trip speed planning problem for autonomous trucks under delivery time and

Cited by 0SourceScholar
2025

GEMD-UNet: Graph Structure Enhanced Multi-dimensional Learning Unet for Cloud Detection

ICASSP 2025accepted

Cloud detection (CD) in remote sensing images is commonly used in satellite imaging and laser communication. UNet-based methods with multi-level feature caching and interaction learning, are popular for superior CD performance. However, most current CD methods focus on spatial feature enhancement th…

Cited by 0SourceScholar
2025

GIM: A Million-scale Benchmark for Generative Image Manipulation Detection and Localization

AAAI 2025technical

The extraordinary ability of generative models emerges as a new trend in image editing and generating realistic images, posing a serious threat to the trustworthiness of multimedia data and driving the research of image manipulation detection and location (IMDL). However, the lack of a large-scale d…

2025

GoHD: Gaze-oriented and Highly Disentangled Portrait Animation with Rhythmic Poses and Realistic Expressions

AAAI 2025technical

Audio-driven talking head generation necessitates seamless integration of audio and visual data amidst the challenges posed by diverse input portraits and intricate correlations between audio and facial motions. In response, we propose a robust framework GoHD designed to produce highly realistic, ex…

2025

HOMO-Feature: Cross-Arbitrary-Modal Image Matching with Homomorphism of Organized Major Orientation

ICCV 2025poster

An exploration of cross-arbitrary-modal image invariant feature extraction and matching is made, with a purely handcrafted full-chain algorithm, Homomorphism of Organized Major Orientation (HOMO), being proposed. Instead of using deep models to conduct data-driven black-box learning, we introduce a…

2025

Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

ICCV 2025poster

Unifying visual understanding and generation within a single multimodal framework remains a significant challenge, as the two inherently heterogeneous tasks require representations at different levels of granularity. Current approaches that utilize vector quantization (VQ) or variational autoencoder…

2025

HeR-DRL:Heterogeneous Relational Deep Reinforcement Learning for Single-Robot and Multi-Robot Crowd Navigation

RA-L 2025

Crowd navigation has garnered significant research attention in recent years, particularly with the advent of DRL-based methods. Current DRL-based methods have extensively explored interaction relationships in single-robot scenarios. However, the heterogeneity of multiple interaction relationships i

Cited by 4SourceScholar
2025

Hierarchical Visual Policy Learning for Long-Horizon Robot Manipulation in Densely Cluttered Scenes

ICRA 2025

In this work, we focus on addressing the long-horizon packing tasks in densely cluttered scenes. Such tasks require policies to effectively manage severe occlusions among objects and continually produce precise actions based on visual observations. We propose a vision-based Hierarchical policy for C

Cited by 0SourceScholar
2025

Hybrid Boundary Physics-Informed Neural Networks for Solving Navier-Stokes Equations with Complex Boundary

NeurIPS 2025poster

Physics-informed neural networks (PINN) have achieved notable success in solving partial differential equations (PDE), yet solving the Navier-Stokes equations (NSE) with complex boundary conditions remains a challenging task. In this paper, we introduce a novel Hybrid Boundary PINN (HB-PINN) method…

Cited by 0SourceScholar
2025

Hyper-Modality Enhancement for Multimodal Sentiment Analysis with Missing Modalities

NeurIPS 2025poster

Multimodal Sentiment Analysis (MSA) aims to infer human emotions by integrating complementary signals from diverse modalities. However, in real-world scenarios, missing modalities are common due to data corruption, sensor failure, or privacy concerns, which can significantly degrade model performanc…

Cited by 0SourceScholar
2025

ISPDiffuser: Learning RAW-to-sRGB Mappings with Texture-Aware Diffusion Models and Histogram-Guided Color Consistency

AAAI 2025technical

RAW-to-sRGB mapping, or the simulation of the traditional camera image signal processor (ISP), aims to generate DSLR-quality sRGB images from raw data captured by smartphone sensors. Despite achieving comparable results to sophisticated handcrafted camera ISP solutions, existing learning-based metho…

2025

Improving 5G Positioning Through Signal-to-Noise Ratio Recognition Training

ICASSP 2025accepted

Fifth-generation communication technology enables advanced indoor positioning with its high bandwidth and frequency capabilities. However, indoor environment variability causes signal propagation fluctuations, making existing models inadequate for accurate location estimation. In this paper, we demo…

Cited by 0SourceScholar
2025

Improving LLM Video Understanding with 16 Frames Per Second

ICML 2025poster

Human vision is dynamic and continuous. However, in video understanding with multimodal large language models (LLMs), existing methods primarily rely on static features extracted from images sampled at a fixed low frame rate of frame-per-second (FPS) $\leqslant$2, leading to critical visual informat…

2025

Knowledge Distillation with Multi-granularity Mixture of Priors for Image Super-Resolution

ICLR 2025spotlight

Knowledge distillation (KD) is a promising yet challenging model compression approach that transmits rich learning representations from robust but resource-demanding teacher models to efficient student models. Previous methods for image super-resolution (SR) are often tailored to specific teacher-st…

Cited by 4SourcePDFScholar
2025

LION-FS: Fast & Slow Video-Language Thinker as Online Video Assistant

CVPR 2025poster

First-person video assistants are highly anticipated to enhance our daily life through online video dialogue. However, existing online video assistants often sacrifice assistant efficacy for real-time efficiency by processing low-frame-rate videos with coarse-grained visual features. To overcome the…

2025

LLaVA-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

ICLR 2025spotlight

Visual instruction tuning has made considerable strides in enhancing the capabilities of Large Multimodal Models (LMMs). However, existing open LMMs largely focus on single-image tasks, their applications to multi-image scenarios remains less explored. Additionally, prior LMM research separately tac…

Cited by 0SourcePDFScholar
2025

Learn to Memorize: Scalable Continual Learning in Semiparametric Models with Mixture-of-Neighbors Induction Memory

ACL 2025long

Semiparametric language models (LMs) have shown promise in various Natural Language Processing (NLP) tasks. However, they utilize non-parametric memory as static storage, which lacks learning capability and remains disconnected from the internal information flow of the parametric models, limiting sc…

Cited by 0SourcePDFScholar
2025

Learning to Control Free-Form Soft Swimmers

NeurIPS 2025poster

Swimming in nature achieves remarkable performance through diverse morphological adaptations and intricate solid-fluid interaction, yet exploring this capability in artificial soft swimmers remains challenging due to the high-dimensional control complexity and the computational cost of resolving hyd…

Cited by 0SourcecodeScholar
2025

Leveraging SD Map to Augment HD Map-based Trajectory Prediction

CVPR 2025poster

Latest trajectory prediction models in real-world autonomous driving systems often rely on online High-Definition (HD) maps to understand the road environment.However, online HD maps suffer from perception errors and feature redundancy, which hinder the performance of HD map-based trajectory predict…

Cited by 0SourcePDFScholar
2025

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

CVPR 2025poster

Recent video large language models (Video LLMs) often depend on costly human annotations or proprietary APIs (e.g., GPT-4o) to produce training data, which limits their training at scale. In this paper, we explore large-scale training for Video LLM with cheap automatic speech recognition (ASR) trans…

2025

MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search

NeurIPS 2025poster

Large language models (LLMs) have shown promise in automating scientific hypothesis generation, yet existing approaches primarily yield coarse-grained hypotheses lacking critical methodological and experimental details. We introduce and formally define the new task of fine-grained scientific hypothe…

Cited by 0SourceScholar
2025

MoE-SVD: Structured Mixture-of-Experts LLMs Compression via Singular Value Decomposition

ICML 2025poster

Mixture of Experts (MoE) architecture improves Large Language Models (LLMs) with better scaling, but its higher parameter counts and memory demands create challenges for deployment. In this paper, we present MoE-SVD, a new decomposition-based compression framework tailored for MoE LLMs without any e…

2025

Monte Carlo Tree Search Based Prompt Autogeneration for Jailbreak Attacks against LLMs

COLING 2025main

Jailbreak attacks craft specific prompts or append adversarial suffixes to prompts, thereby inducing language models to generate harmful or unethical content and bypassing the model’s safety guardrails. With the recent blossom of large language models (LLMs), there’s a growing focus on jailbreak att…

2025

NaFV-Net: An Adversarial Four-view Network for Mammogram Classification

AAAI 2025technical

Breast cancer remains a leading cause of mortality among women, with millions of new cases diagnosed annually. Early detection through screening is crucial. Using neural networks to improve the accuracy of breast cancer screening has become increasingly important. In accordance with radiologists' pr…

2025

Odysseus Navigates the Sirens’ Song: Dynamic Focus Decoding for Factual and Diverse Open-Ended Text Generation

ACL 2025long

Large Language Models (LLMs) are increasingly required to generate text that is both factually accurate and diverse across various open-ended applications. However, current stochastic decoding methods struggle to balance such objectives. We introduce Dynamic Focus Decoding (DFD), a novel plug-and-pl…

Cited by 0SourcePDFScholar
2025

OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

ICLR 2025spotlight

Image-text interleaved data, consisting of multiple images and texts arranged in a natural document format, aligns with the presentation paradigm of internet data and closely resembles human reading habits. Recent studies have shown that such data aids multimodal in-context learning and maintains th…

2025

OpenHuEval: Evaluating Large Language Model on Hungarian Specifics

ACL 2025finding

We introduce OpenHuEval, the first benchmark for LLMs focusing on the Hungarian language and specifics. OpenHuEval is constructed from a vast collection of Hungarian-specific materials sourced from multiple origins. In the construction, we incorporated the latest design principles for evaluating LLM…

2025

OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining

ICCV 2025poster

Vision-language pretraining (VLP) enables open-world generalization beyond predefined labels, a critical capability in surgery due to the diversity of procedures, instruments, and patient anatomies. However, applying VLP to ophthalmic surgery presents unique challenges, including limited vision-lang…

2025

Point Clean-label Backdoor Attack for Specific Classes via Feature Entanglement

ICASSP 2025accepted

Point cloud classifiers have been recently demonstrated to be vulnerable to backdoor attacks. The infected model functions normally on clean data, yet its predictions are errors when triggers are encountered. Currently, the point clean-label backdoor attack (PointCBA) method utilizes feature disenta…

Cited by 0SourceScholar
2025

RACQC: Advanced Retrieval-Augmented Generation for Chinese Query Correction

EMNLP 2025

In web search scenarios, erroneous queries frequently degrade users’ experience through irrelevant results, underscoring the pivotal role of Chinese Spelling Check (CSC) systems. Although large language models (LLMs) exhibit remarkable capabilities across many tasks, they face critical challenges in

2025

SRefiner: Soft-Braid Attention for Multi-Agent Trajectory Refinement

ICCV 2025poster

Accurate prediction of multi-agent future trajectories is crucial for autonomous driving systems to make safe and efficient decisions. Trajectory refinement has emerged as a key strategy to enhance prediction accuracy. However, existing refinement methods often overlook the topological relationships…

2025

SSFSL: Self-Supervised and Few-Shot Learning for Cross-Domain Hyperspectral Image Classification

ICASSP 2025accepted

Few-shot learning (FSL) has gained increasing attention in hyperspectral image (HSI) classification due to its ability to perform cross-domain classification with minimal labeled samples. However, existing FSL methods overlook the continuity of HSI spectral sequences and fail to utilize the large am…

Cited by 0SourceScholar
2025

Self-Supervised Uncertainty-Guided Refinement for Robust Joint Optical Flow and Depth Estimation

ICASSP 2025accepted

Jointly estimating the optical flow and depth tasks in real-world scenes presents considerable hurdles due to some phenomena, such as occlusion, ambiguous textures, and illumination variation. The lack of guidance from the labeled data makes these challenges harder to overcome. This paper presents a…

Cited by 0SourceScholar
2025

Solving Partial Differential Equations via Radon Neural Operator

NeurIPS 2025poster

Neural operator is considered a popular data-driven alternative to traditional partial differential equation (PDE) solvers. However, most current solutions, whether fulfilling computations in frequency, Laplacian, and wavelet domains, all deviate far from the intrinsic PDE space. While with meticulo…

Cited by 0SourcecodeScholar
2025

Spy Inside: Scalable Verification of Dependable Transformers for Event Time Series Systems

ICASSP 2025accepted

Event time series appear in many software scenarios and are a necessary data type in data analytics systems. Transformers are the preferred type of sequential neural network for advanced analytics on event time series, particularly due to their significant contributions to the recent surge of large…

Cited by 0SourceScholar
2025

TermDiffuSum: A Term-guided Diffusion Model for Extractive Summarization of Legal Documents

COLING 2025main

Extractive summarization for legal documents aims to automatically extract key sentences from legal texts to form concise summaries. Recent studies have explored diffusion models for extractive summarization task, showcasing their remarkable capabilities. Despite these advancements, these models oft…

2025

The Role of Video Generation in Enhancing Data-Limited Action Understanding

IJCAI 2025

Video action understanding tasks in real-world scenarios often suffer from data limitations. In this paper, we address the data-limited action understanding problem by bridging data scarcity. We propose a novel method that leverages a text-to-video diffusion transformer to generate annotated data fo

Cited by 0SourcePDFScholar
2025

TimE: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios

NeurIPS 2025spotlight

Temporal reasoning is pivotal for Large Language Models (LLMs) to comprehend the real world. However, existing works neglect the real-world challenges for temporal reasoning: (1) intensive temporal information, (2) fast-changing event dynamics, and (3) complex temporal dependencies in social interac…

Cited by 0SourcecodeScholar
2025

Ultra Lightweight Singing Melody Extraction via Combination of Convolution and MLP

ICASSP 2025accepted

Singing melody extraction serves as an important foundation in the realm of music information retrieval (MIR). Although fully convolutional neural networks (CNNs) are commonly employed for singing melody extraction, they are constrained by inductive biases and face challenges in establishing long ra…

Cited by 0SourceScholar
2025

Unsupervised Image-to-Image Style Transfer via Dual-Condition Diffusion Models*

ICASSP 2025accepted

Style transfer is an artistic research topic within a series of generative tasks, and it is quite challenging to stably and reliably guide the generation of images with the expected content and style. Early methods followed a predefined style, defined by a reference image, and applied it to another…

Cited by 0SourceScholar
2025

V-Fusion: 2D Detection-enhanced Multimodal 3D BEV Object Detection

ICASSP 2025accepted

Integrating information from multiple sensors enhances the performance of autonomous vehicle perception systems. However, current multimodal 3D object detection methods focus on unifying modalities into a bird’s-eye view (BEV) representation, which overlooks the inherent characteristics of camera pe…

Cited by 0SourceScholar
2025

Vibration-Aware Trajectory Optimization for Mobile Robots in Wild Environments via Physics-Informed Neural Network

IROS 2025

The suspension system, through effective damping of vibrations and shocks, can enhance the stability of wheeled robots traversing challenging terrain. Because the suspension system decouples the rigid correspondence between terrain changes and robot vibrations, considering suspension modeling in tra

Cited by 0SourceScholar
2025

VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning

NeurIPS 2025poster

Recent advancements in vision-language models (VLMs) have improved performance by increasing the number of visual tokens, which are often significantly longer than text tokens. However, we observe that most real-world scenarios do not require such an extensive number of visual tokens. While the perf…

Cited by 0SourceScholar
2025

Visual Evidence Prompting Mitigates Hallucinations in Large Vision-Language Models

ACL 2025long

Large Vision-Language Models (LVLMs) have shown impressive progress by integrating visual perception with linguistic understanding to produce contextually grounded outputs. Despite these advancements achieved, LVLMs still suffer from the hallucination problem, e.g., they tend to produce content that…

Cited by 0SourcePDFScholar
2025

Wasserstein Heterogeneous Graph Neural Networks for Uncertainty-Aware Anomaly Detection

ICASSP 2025accepted

Graph anomaly detection, a critical topic in graph mining, has garnered significant research interest and found applications across diverse domains such as attack event detection, spam review identification, and financial fraud prevention. Graph Neural Networks (GNNs) have emerged as the dominant ap…

Cited by 0SourceScholar
2025

Weakly Supervised Semantic Segmentation via Progressive Confidence Region Expansion

CVPR 2025poster

Weakly supervised semantic segmentation (WSSS) has garnered considerable attention due to its effective reduction of annotation costs. Most approaches utilize Class Activation Maps (CAM) to produce pseudo-labels, thereby localizing target regions using only image-level annotations. However, the prev…

2025

WildAvatar: Learning In-the-wild 3D Avatars from the Web

CVPR 2025poster

Existing research on avatar creation is typically limited to laboratory datasets, which require high costs against scalability and exhibit insufficient representation of the real world. On the other hand, the web abounds with off-the-shelf real-world human videos, but these videos vary in quality an…

Cited by 2SourcePDFScholar
2025

ZeroFlow: Overcoming Catastrophic Forgetting is Easier than You Think

ICML 2025poster

Backpropagation provides a generalized configuration for overcoming catastrophic forgetting. Optimizers such as SGD and Adam are commonly used for weight updates in continual learning and continual pre-training. However, access to gradient information is not always feasible in practice due to black-…

Cited by 1SourcePDFScholar
2025

video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model

ICML 2025poster

While recent advancements in reasoning optimization have significantly enhanced the capabilities of large language models (LLMs), existing efforts to improve reasoning have been limited to solving mathematical problems and focusing on visual graphical inputs, neglecting broader applications in gener…

2024

A Safe and Efficient Timed-Elastic-Band Planner for Unstructured Environments

IROS 2024poster

In unstructured environments with complex obstacles and obscure road boundaries, the local planner faces more severe challenges in terms of safety and real-time performance. In order to fulfill these emerging requirements, we propose a novel Timed-Elastic-Band approach for unstructured environments,…

Cited by 2SourceScholar
2024

Adaptive Layer Sparsity for Large Language Models via Activation Correlation Assessment

NeurIPS 2024poster

Large Language Models (LLMs) have revolutionized the field of natural language processing with their impressive capabilities. However, their enormous size presents challenges for deploying them in real-world applications. Traditional compression techniques, like pruning, often lead to suboptimal per…

2024

Adaptive Motion Scaling for Robot-Assisted Microsurgery Based on Hybrid Offline Reinforcement Learning and Damping Control

ICRA 2024poster

Motion scaling is essential to empower users to conduct precise manipulation during teleoperation for robot-assisted microsurgery (RAMS). A constant, small motion scaling ratio can enhance the precision of teleoperation but hinder the operator from quickly reaching distant targets. The concept of se…

Cited by 1SourceScholar
2024

An Unsupervised Framework for Adaptive Context-aware Simplified-Traditional Chinese Character Conversion

COLING 2024main

Traditional Chinese character is an important carrier of Chinese culture, and is still actively used in many areas. Automatic conversion between traditional and simplified Chinese characters can help modern people understand traditional culture and facilitate communication among different regions. P…

2024

AutoOS: Make Your OS More Powerful by Exploiting Large Language Models

ICML 2024poster

With the rapid development of Artificial Intelligence of Things (AIoT), customizing and optimizing operating system (OS) kernel configurations for various AIoT application scenarios is crucial for maximizing system performance. However, existing approaches falter due to the overwhelming problem comp…

Cited by 3SourcePDFScholar
2024

Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to Reasoning-Memorization Correlations

ACL 2024long

We introduce CHARM, the first benchmark for comprehensively and in-depth evaluating the commonsense reasoning ability of large language models (LLMs) in Chinese, which covers both globally known and Chinese-specific commonsense. We evaluated 7 English and 12 Chinese-oriented LLMs on CHARM, employing…

2024

CGS-Mask: Making Time Series Predictions Intuitive for All

AAAI 2024technical

Artificial intelligence (AI) has immense potential in time series prediction, but most explainable tools have limited capabilities in providing a systematic understanding of important features over time. These tools typically rely on evaluating a single time point, overlook the time ordering of inpu…

Cited by 1SourcePDFScholar
2024

Connecting Speech Encoder and Large Language Model for ASR

ICASSP 2024accepted

The impressive capability and versatility of large language models (LLMs) have aroused increasing attention in automatic speech recognition (ASR), with several pioneering studies attempting to build integrated ASR models by connecting a speech encoder with an LLM. This paper presents a comparative s…

Cited by 0SourceScholar
2024

DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object Detection

AAAI 2024technical

Vehicle-to-Everything (V2X) collaborative perception has recently gained significant attention due to its capability to enhance scene understanding by integrating information from various agents, e.g., vehicles, and infrastructure. However, current works often treat the information from each agent e…

2024

Detection-Correction Structure via General Language Model for Grammatical Error Correction

ACL 2024long

Grammatical error correction (GEC) is a task dedicated to rectifying texts with minimal edits, which can be decoupled into two components: detection and correction. However, previous works have predominantly focused on direct correction, with no prior efforts to integrate both into a single model. M…

2024

ESP: Extro-Spective Prediction for Long-term Behavior Reasoning in Emergency Scenarios

ICRA 2024poster

Emergent-scene safety is the key milestone for fully autonomous driving, and reliable on-time prediction is essential to maintain safety in emergency scenarios. However, these emergency scenarios are long-tailed and hard to collect, which restricts the system from getting reliable predictions. In th…

Cited by 1SourcecodeScholar
2024

Enhancing Discourse Dependency Parsing with Sentence Dependency Parsing: A Unified Generative Method Based on Code Representation

EMNLP 2024finding

Due to the high complexity of Discourse Dependency Parsing (DDP) tasks, their existing annotation resources are relatively scarce compared to other NLP tasks, and different DDP tasks also have significant differences in annotation schema. These issues have led to the dilemma of low resources for DDP…

Cited by 0SourcePDFScholar
2024

Exploring Structured Semantic Priors Underlying Diffusion Score for Test-time Adaptation

NeurIPS 2024poster

Capitalizing on the complementary advantages of generative and discriminative models has always been a compelling vision in machine learning, backed by a growing body of research. This work discloses the hidden semantic structure within score-based generative models, unveiling their potential as eff…

2024

Extending Large Language Models for Speech and Audio Captioning

ICASSP 2024accepted

Multimodal large language models (LLMs) have shown promising visual perception abilities by connecting with image encoders, but their performance on auditory tasks has not yet been widely investigated. Meanwhile, automatic speech recognition (ASR) and automatic audio captioning (AAC) are often achie…

Cited by 0SourceScholar
2024

Extending Multilingual ASR to New Languages Using Supplementary Encoder and Decoder Components

ICASSP 2024accepted

Extending multilingual automatic speech recognition (mASR) systems to new languages poses challenges, particularly when training data for existing languages is limited or unavailable. To tackle this issue, we suggest utilizing supplementary encoder and decoder components. Specifically, we propose ap…

Cited by 0SourceScholar
2024

Face Adapter for Pre-Trained Diffusion Models with Fine-Grained ID and Attribute Control

ECCV 2024poster

"Current face reenactment and swapping methods mainly rely on GAN frameworks, but recent focus has shifted to pre-trained diffusion models for their superior generation capabilities. However, training these models is resource-intensive, and the results have not yet achieved satisfactory performance…

Cited by 27SourcePDFScholar
2024

GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI

NeurIPS 2024poster

Large Vision-Language Models (LVLMs) are capable of handling diverse data types such as imaging, text, and physiological signals, and can be applied in various fields. In the medical field, LVLMs have a high potential to offer substantial assistance for diagnosis and treatment. Before that, it is cr…

2024

HeterGCL: Graph Contrastive Learning Framework on Heterophilic Graph

IJCAI 2024poster

Graph Contrastive Learning (GCL) has attracted significant research attention due to its self-supervised ability to learn robust node representations. Unfortunately, most methods primarily focus on homophilic graphs, rendering them less effective for heterophilic graphs. In addition, the complexity…

2024

HybridBooth: Hybrid Prompt Inversion for Efficient Subject-Driven Generation

ECCV 2024poster

"Recent advancements in text-to-image diffusion models have shown remarkable creative capabilities with textual prompts, but generating personalized instances based on specific subjects, known as subject-driven generation, remains challenging. To tackle this issue, we present a new hybrid framework…

Cited by 4SourcePDFScholar
2024

Hypergraph-Based Session Modeling: A Multi-Collaborative Self-Supervised Approach for Enhanced Recommender Systems

COLING 2024main

Session-based recommendation (SBR) is a challenging task that involves predicting a user’s next item click based on their recent session history. Presently, many state-of-the-art methodologies employ graph neural networks to model item transitions. Notwithstanding their impressive performance, graph…

Cited by 4SourcePDFScholar
2024

IS-Fusion: Instance-Scene Collaborative Fusion for Multimodal 3D Object Detection

CVPR 2024highlight

Bird's eye view (BEV) representation has emerged as a dominant solution for describing 3D space in autonomous driving scenarios. However objects in the BEV representation typically exhibit small sizes and the associated point cloud context is inherently sparse which leads to great challenges for rel…

2024

Identity-Consistent Diffusion Network for Grading Knee Osteoarthritis Progression in Radiographic Imaging

ECCV 2024poster

"Knee osteoarthritis (KOA), a common form of arthritis that causes physical disability, has become increasingly prevalent in society. Employing computer-aided techniques to automatically assess the severity and progression of KOA can greatly benefit KOA treatment and disease management. Particularly…

Cited by 1SourcePDFScholar
2024

Improving Context Understanding in Multimodal Large Language Models via Multimodal Composition Learning

ICML 2024poster

Previous efforts using frozen Large Language Models (LLMs) for visual understanding, via image captioning or image-text retrieval tasks, face challenges when dealing with complex multimodal scenarios. In order to enhance the capabilities of Multimodal Large Language Models (MLLM) in comprehending th…

2024

Improving Robustness of GNN-based Anomaly Detection by Graph Adversarial Training

COLING 2024main

Graph neural networks (GNNs) play a fundamental role in anomaly detection, excelling at the identification of node anomalies by aggregating information from neighboring nodes. Nonetheless, they exhibit vulnerability to attacks, with even minor alterations in the graph structure or node attributes re…

Cited by 7SourcePDFScholar
2024

InstructEval: Instruction-Tuned Text Evaluator from Human Preference

ACL 2024findings

This paper explores to construct a general text evaluator based on open-source Large Language Models (LLMs), a domain predominantly occupied by commercial counterparts such as GPT-4. Recognizing the limitations of open-source models like Llama in evaluative tasks, we introduce InstructEval, a genera…

2024

InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

NeurIPS 2024poster

The Large Vision-Language Model (LVLM) field has seen significant advancements, yet its progression has been hindered by challenges in comprehending fine-grained visual content due to limited resolution. Recent efforts have aimed to enhance the high-resolution understanding capabilities of LVLMs, ye…

2024

Interpretable Composition Attribution Enhancement for Visio-linguistic Compositional Understanding

EMNLP 2024main

Contrastively trained vision-language models such as CLIP have achieved remarkable progress in vision and language representation learning. Despite the promising progress, their proficiency in compositional reasoning over attributes and relations (e.g., distinguishing between “the car is underneath…

Cited by 0SourcePDFScholar
2024

MM-ChatAlign: A Novel Multimodal Reasoning Framework based on Large Language Models for Entity Alignment

EMNLP 2024finding

Multimodal entity alignment (MMEA) integrates multi-source and cross-modal knowledge graphs, a crucial yet challenging task for data-centric applications.Traditional MMEA methods derive the visual embeddings of entities and combine them with other modal data for alignment by embedding similarity com…

2024

MVSGaussian: Fast Generalizable Gaussian Splatting Reconstruction from Multi-View Stereo

ECCV 2024poster

"We present MVSGaussian, a new generalizable 3D Gaussian representation approach derived from Multi-View Stereo (MVS) that can efficiently reconstruct unseen scenes. Specifically, 1) we leverage MVS to encode geometry-aware Gaussian representations and decode them into Gaussian parameters. 2) To fur…

2024

Make Continual Learning Stronger via C-Flat

NeurIPS 2024poster

How to balance the learning ’sensitivity-stability’ upon new task training and memory preserving is critical in CL to resolve catastrophic forgetting. Improving model generalization ability within each learning phase is one solution to help CL learning overcome the gap in the joint knowledge space.…

2024

Mertech: Instrument Playing Technique Detection Using Self-Supervised Pretrained Model with Multi-Task Finetuning

ICASSP 2024accepted

Instrument playing techniques (IPTs) constitute a pivotal component of musical expression. However, the development of automatic IPT detection methods suffers from limited labeled data and inherent class imbalance issues. In this paper, we propose to apply a self-supervised learning model pre-traine…

Cited by 0SourceScholar
2024

Microrobotic Flight Enabled by Ultralight Ion Thrusters with High Thrust-to-Weight Ratio and Low Fabrication Cost

ICRA 2024poster

Flying microrobots have garnered growing research interest owing to their technological intricacies and suitability for various applications leveraging miniaturized size. Electrohydrodynamic (EHD) thrust offers advantages by generating propulsion without moving parts, but real-world use is limited b…

Cited by 2SourceScholar
2024

Multi-Modal Disordered Representation Learning Network for Description-Based Person Search

AAAI 2024technical

Description-based person search aims to retrieve images of the target identity via textual descriptions. One of the challenges for this task is to extract discriminative representation from images and descriptions. Most existing methods apply the part-based split method or external models to explore…

Cited by 4SourcePDFScholar
2024

NPC: Neural Predictive Control for Fuel-Efficient Autonomous Trucks

ICRA 2024poster

Fuel efficiency is a crucial aspect of long-distance cargo transportation by oil-powered trucks that economize on costs and decrease carbon emissions. Current predictive control methods depend on an accurate model of vehicle dynamics and engine, including weight, drag coefficient, and the Brake-spec…

Cited by 0SourceScholar
2024

NeuMA: Neural Material Adaptor for Visual Grounding of Intrinsic Dynamics

NeurIPS 2024poster

While humans effortlessly discern intrinsic dynamics and adapt to new scenarios, modern AI systems often struggle. Current methods for visual grounding of dynamics either use pure neural-network-based simulators (black box), which may violate physical laws, or traditional physical simulators (white…

2024

OMG-Seg: Is One Model Good Enough For All Segmentation?

CVPR 2024poster

In this work we address various segmentation tasks each traditionally tackled by distinct or partially unified models. We propose OMG-Seg One Model that is Good enough to efficiently and effectively handle all the segmentation tasks including image semantic instance and panoptic segmentation as well…

2024

On the Effects of Data Scale on UI Control Agents

NeurIPS 2024spotlight

Autonomous agents that control user interfaces to accomplish human tasks are emerging. Leveraging LLMs to power such agents has been of special interest, but unless fine-tuned on human-collected task demonstrations, performance is still relatively low. In this work we study whether fine-tuning alone…

Cited by 14SourcePDFScholar
2024

PEACE: A Dataset of Pharmaceutical Care for Cancer Pain Analgesia Evaluation and Medication Decision

NeurIPS 2024poster

Over half of cancer patients experience long-term pain management challenges. Recently, interest has grown in systems for cancer pain treatment effectiveness assessment (TEA) and medication recommendation (MR) to optimize pharmacological care. These systems aim to improve treatment effectiveness by…

2024

PQ-SAM: Post-training Quantization for Segment Anything Model

ECCV 2024poster

"Segment anything model (SAM) is a promising prompt-guided vision foundation model to segment objects of interest. However, the extensive computational requirements of SAM have limited its applicability in resource-constraint edge devices. Post-training quantization (PTQ) is an effective potential f…

Cited by 5SourcePDFScholar
2024

Physics-Constrained Comprehensive Optical Neural Networks

NeurIPS 2024poster

With the advantages of low latency, low power consumption, and high parallelism, optical neural networks (ONN) offer a promising solution for time-sensitive and resource-limited artificial intelligence applications. However, the performance of the ONN model is often diminished by the gap between the…

Cited by 1SourcePDFScholar
2024

Provable Acceleration of Nesterov’s Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks

IJCAI 2024poster

Due to its simplicity and efficiency, the first-order gradient method has been extensively employed in training neural networks. Although the optimization problem of the neural network is non-convex, recent research has proved that the first-order method is capable of attaining a global minimum duri…

Cited by 0SourcePDFScholar