← Search

Xiang Li

376 accepted papers

2026

A Semi-Active Occupational Shoulder Exoskeleton for Overhead Work With Free Mode and Personalized Assistive Torque

RA-L 2026

Current passive or semi-active shoulder exoskeletons for overhead work provide fixed assistive torque for all participants and tasks, which lacks adaptability. In addition, due to the need to store energy at low elevation angles, they may increase physical demand on the user when assistance is not r

Cited by 0SourceScholar
2026

A Semi-Active Occupational Shoulder Exoskeleton for Overhead Work with Free Mode and Personalized Assistive Torque

ICRA 2026poster

Current passive or semi-active shoulder exoskeletons for overhead work provide fixed assistive torque for all participants and tasks, which lacks adaptability. In addition, due to the need to store energy at low elevation angles, they may increase physical demand on the user when assistance is not r…

Cited by 0SourceScholar
2026

ALSO: Adversarial Online Strategy Optimization for Social Agents

ICML 2026poster

Social simulation provides a compelling testbed for studying social intelligence, where agents interact through multi-turn dialogues under evolving contexts and strategically adapting opponents. Such environments are inherently non-stationary, requiring agents to dynamically adjust their strategies …

Cited by 0SourceScholar
2026

AirSim360: A Panoramic Simulation Platform within Drone View

CVPR 2026

The field of 360-degree omnidirectional understanding has been receiving increasing attention for advancing spatial intelligence. However, the lack of large-scale and diverse data remains a major limitation. In this work, we propose AirSim360, a simulation platform for omnidirectional data from aeri

Cited by 0SourcecodeScholar
2026

Analyze–Compose–Execute: A Dynamic Dialogue Framework for Multi-Agent Debate

AAAI 2026technical

Multi-Agent Debate (MAD) is an emerging paradigm that leverages the reasoning abilities of Large Language Models (LLMs) by encouraging them to collaboratively solve problems through human-like discussions. However, current MAD methods typically constrain agents to follow fixed discussion pipelines,

Cited by 0SourcePDFScholar
2026

Arm-Aware Guided Dexterous Grasp Generation With Arm-Agnostic Grasp Models

RA-L 2026

Dexterous grasp generation that considers armrelated constraints is crucial in real-world scenarios involving armenvironment collision avoidance, workspace boundary grasps, and consecutive grasping. Existing hand-centric grasp models, which primarily focus on the floating hand's pose, are insufficie

Cited by 0SourcecodeScholar
2026

Beyond Adapter Retrieval: Latent Geometry-Preserving Composition via Sparse Task Projection

AAAI 2026technical

Recent advances in parameter-efficient transfer learning have demonstrated the utility of composing LoRA adapters from libraries of pretrained modules. However, most existing approaches rely on simple retrieval heuristics or uniform averaging, which overlook the latent structure of task relationshi

Cited by 0SourcePDFScholar
2026

Break the Trade-off Between Watermark Strength and Speculative Sampling Efficiency for Language Models

ICLR 2026poster

Watermarking is a principled approach for tracing the provenance of large language model (LLM) outputs, but its deployment in practice is hindered by inference inefficiency. Speculative sampling accelerates inference, with efficiency improving as the acceptance rate between draft and target models i…

Cited by 0SourceScholar
2026

Confusion-Aware Spectral Regularizer for Long-Tailed Recognition

CVPR 2026

Long-tailed image classification remains a long-standing challenge, as real-world data typically follow highly imbalanced distributions where a few head classes dominate and many tail classes contain only limited samples. This imbalance biases feature learning toward head categories and leads to sig

Cited by 0SourcecodeScholar
2026

CoorGrasp: Coordinated Contact Control for Adaptive Dexterous Grasping under Uncertainty

ICRA 2026poster

While recent research has focused heavily on dexterous grasp pose generation, less attention has been devoted to the execution of planned grasps. Under shape and position uncertainty, open-loop execution often yields uncoordinated contacts, causing undesired in-hand object motion and even grasp fail…

Cited by 0codeScholar
2026

DanceTogether: Generating Interactive Multi-Person Video without Identity Drifting

ICLR 2026poster

Controllable video generation (CVG) has advanced rapidly, yet current systems falter when more than one actor must move, interact, and exchange positions under noisy control signals. We address this gap with DanceTogether, the first end-to-end diffusion framework that turns a single reference image…

Cited by 0SourceScholar
2026

DenoDet V2: Phase-Amplitude Cross Denoising for SAR Object Detection

AAAI 2026technical

One of the primary challenges in Synthetic Aperture Radar (SAR) object detection lies in the pervasive influence of coherent noise. As a common practice, most existing methods, whether handcrafted approaches or deep learning-based methods, employ the analysis or enhancement of object spatial-domain

Cited by 0SourcePDFScholar
2026

Denoising Neural Reranker for Recommender Systems

ICLR 2026poster

For multi-stage recommenders in industry, a user request would first trigger a simple and efficient retriever module that selects and ranks a list of relevant items, then the recommender calls a slower but more sophisticated reranking model that refines the item list exposure to the user. To consist…

Cited by 0SourcecodeScholar
2026

Disentangled Hypergraph Network with Implicit Structure Learning for Mobility Social Relationship Inference

IJCAI 2026

Inferring social relationships from users' mobile data holds significant value for personalized recommendations. Most methods model user interactions based on co-occurrence records, achieving impressive success in capturing social signals. However, despite these advancements, current techniques stil

Cited by 0Scholar
2026

Dual-Channel Hybrid Graph Neural Network for Mobility Social Relationship Inference

IJCAI 2026

Inferring latent social ties from large-scale spatiotemporal mobility traces is a foundational AI task with broad applicability. Existing hypergraph-based methods often model higher-order relations by treating hyperedges as static snapshots, thus failing to capture the temporal dynamics and co-evolu

Cited by 0Scholar
2026

Dual-Kernel Adapter: Expanding Spatial Horizons for Data-Constrained Medical Image Analysis

ICLR 2026poster

Adapters have become a widely adopted strategy for efficient fine-tuning of foundation models, particularly in resource-constrained settings. However, their performance under extreme data scarcity—common in medical imaging due to high annotation costs, privacy regulations, and fragmented datasets—re…

Cited by 0SourceScholar
2026

Dynamic Legged Ball Manipulation on Rugged Terrains With Hierarchical Reinforcement Learning

RA-L 2026

Achieving reliable object manipulation while traversing complex terrains is the missing link between agile quadruped locomotion and practical autonomy. Specifically, using traditional end-to-end reinforcement learning (RL) for dynamic ball manipulation in rugged environments presents two key challen

Cited by 1SourceScholar
2026

Efficient Transcoder Adaptation for Fine-Tuned Models: Revealing Medical Reasoning Mechanisms in Large Language Models

AAAI 2026technical

Large language models (LLMs) suffer from a lack of decision-making transparency, limiting their deployment in high-stakes domains such as healthcare. We propose a mechanistic interpretability framework that introduces two novel paradigms: Medical Fine-Tuning with Frozen Attention Layers (FTFA) and P

Cited by 0SourcePDFScholar
2026

Ego-1K - A Large-Scale Multiview Video Dataset for Egocentric Vision

CVPR 2026

We present Ego-1K, a large-scale, time-synchronized collection of egocentric multiview videos designed to advance neural 3D video synthesis, dynamic scene understanding, and embodied perception. The dataset contains nearly 1,000 short egocentric videos taken with a custom rig with 12 synchronous cam

Cited by 0SourceScholar
2026

Ego-PMOVE: Prompt-aware Mixture of View Experts Network for Egocentric Gaze Prediction

AAAI 2026technical

Egocentric gaze prediction serves as a critical indicator for decoding human visual attention and cognitive processes, but its inherently limited field of view creates prediction challenges. Although exo-view data provides supplementary contextual information, it exhibits significant spatial and sem

Cited by 0SourcePDFScholar
2026

Embodied Task Planning via Graph-Informed Action Generation with Large Lanaguage Model

ICML 2026poster

While Large Language Models (LLMs) have demonstrated strong zero-shot reasoning capabilities, their deployment as embodied agents still faces fundamental challenges in long-horizon planning. Unlike open-ended text generation, embodied agents must decompose high-level intent into actionable sub-goals…

Cited by 1SourceScholar
2026

Evaluating the Representation Space of Diffusion Models via Self-Supervised Principles

ICML 2026poster

Diffusion models are effective generative frameworks with strong representation learning capabilities, yet the intrinsic properties that govern their semantic structure and generalization remain poorly understood. Drawing inspiration from self-supervised representation learning (SSL), we introduce a…

Cited by 0SourceScholar
2026

Flexible Trajectory Planning for Autonomous Vehicles Via Environmental Assessment in Extreme Scenarios

ICRA 2026poster

Trajectory planning is a core task in autonomous driving. However, in diverse extreme scenarios characterized by unstructured obstacles, there is a lack of solutions that provide efficient computation, safety, and scene generalization capabilities. To address this issue, we propose a two-stage spati…

Cited by 0Scholar
2026

From Swept Contact to Pose: Probe-Aware Registration Via Complementary-Shape Docking

ICRA 2026poster

Accurate registration between a prior model and the real scene is essential for high-precision robotic manipulation, yet optical methods suffer from long calibration chains, line-of-sight constraints, and fabrication errors. We propose a calibration-free alternative that reformulates contact registr…

2026

GAF: Gaussian Action Field As a 4D Representation for Dynamic World Modeling in Robotic Manipulation

ICRA 2026poster

Accurate scene perception is critical for vision-based robotic manipulation. Existing approaches typically follow either a Vision-to-Action V-A paradigm, predicting actions directly from visual inputs, or a Vision-to-3D-to-Action V-3D-A paradigm, leveraging intermediate 3D representations. However, …

2026

GS-UVCE: Gaussian Splatting-Driven Unsupervised Visual Consistency Enhancement for Underwater 3D Scene Reconstruction

ICRA 2026poster

Underwater 3D scene reconstruction is critical for the operation of underwater robotics, yet remains highly challenging due to the semi-transparent water medium, which introduces optical distortions, light scattering, and severe visibility degradation. Therefore, effective underwater image enhanceme…

Cited by 0Scholar
2026

GarmentGPT: Compositional Garment Pattern Generation via Discrete Latent Tokenization

ICLR 2026poster

Apparel is a fundamental component of human appearance, making garment digitalization critical for digital human creation. However, sewing pattern creation traditionally relies on the intuition and extensive experience of skilled artisans. This manual bottleneck significantly hinders the scalability…

Cited by 0SourcecodeScholar
2026

Generalization of Diffusion Models Arises with a Balanced Representation Space

ICLR 2026poster

Diffusion models generate high-quality, diverse images with great generalizability, yet when overfit to the training objective, they may memorize training samples. We analyze memorization and generalization of diffusion models through the lens of representation learning. Using a two-layer ReLU denoi…

Cited by 0SourcecodeScholar
2026

GeoBayes: Probabilistic Image Geo-Localization Inference via Sequential Bayesian Updating

AAAI 2026technical

Image geo-localization aims to determine the geographic location of a query image. While Multimodal Large Language Models (MLLMs) show potential for this task due to their rich world knowledge and explainable abilities, they often struggle with confirmation bias, i.e., committing to early, potential

Cited by 0SourcePDFScholar
2026

GigaMoE: Sparsity-Guided Mixture of Experts for Efficient Gigapixel Object Detection

AAAI 2026technical

Object detection in High-Resolution Wide (HRW) shots, or gigapixel images, presents unique challenges due to extreme object sparsity and vast scale variations. State-of-the-art methods like SparseFormer have pioneered sparse processing by selectively focusing on important regions, yet they apply a u

Cited by 0SourcePDFScholar
2026

GoalRank: Group-Relative Optimization for a Large Ranking Model

ICLR 2026poster

Mainstream ranking approaches typically follow a Generator–Evaluator two-stage paradigm, where a generator produces candidate lists and an evaluator selects the best one. Recent work has attempted to enhance performance by expanding the number of candidate lists, for example, through multi-generator…

Cited by 0SourcecodeScholar
2026

GradientStabilizer: Fix the Norm, Not the Gradient

ICML 2026poster

Training instability in modern deep learning systems is frequently triggered by rare but extreme gradient-norm spikes, which can induce oversized parameter updates, corrupt optimizer state, and lead to slow recovery or divergence. Widely used safeguards such as gradient clipping mitigate these failu…

Cited by 0SourceScholar
2026

HSI-GPT2: A Dual-Granularity Large Motion Reasoning Model with Diffusion Refinement for Human-Scene Interaction

CVPR 2026

Unified interpreting and synthesizing human behaviors within 3D environments is vital for advancing spatial intelligence and humanoid robotics. Despite recent advancements (e.g., HSI-GPT), two fundamental capabilities expected of a unified model--understanding and generation--still lag behind specia

Cited by 0SourceScholar
2026

Hierarchical Semantic-Acoustic Modeling via Semi-Discrete Residual Representations for Expressive End-to-End Speech Synthesis

ICLR 2026poster

Generative models for speech synthesis face a fundamental trade-off: discrete tokens ensure stability but sacrifice expressivity, while continuous signals retain acoustic richness but suffer from error accumulation due to task entanglement. This challenge has driven the field towards multi-stage pip…

Cited by 0SourcecodeScholar
2026

HippoTune: A Hippocampal Associative Loop–Inspired Fine-Tuning Method for Continual Learning

ICLR 2026poster

Studies have shown that catastrophic forgetting primarily stems from the difficulty of reactivating old memories; although parameter-efficient fine-tuning can mitigate forgetting while keeping most model parameters frozen, it still falls short in fully reawakening knowledge of prior tasks. In contra…

Cited by 0SourcecodeScholar
2026

Human Cognition Inspired RAG with Knowledge Graph for Complex Problem Solving

AAAI 2026technical

Large Language Models (LLMs) have demonstrated significant potential across various domains. However, they often struggle with integrating external knowledge and performing complex reasoning, leading to hallucinations and unreliable outputs. Retrieval Augmented Generation (RAG) has emerged as a prom

Cited by 0SourcePDFScholar
2026

Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction

ICLR 2026poster

The pursuit of human-like conversational agents has long been guided by the Turing test. For modern speech-to-speech (S2S) systems, a critical yet unanswered question is whether they can converse like humans. To tackle this, we conduct the first Turing test for S2S systems, collecting 2,968 human ju…

Cited by 0SourceScholar
2026

Initialization is Half the Battle: Generating Diverse Images from a Guidance Potential Posterior

ICML 2026spotlight

Despite the remarkable fidelity of generative models, they frequently suffer from mode collapse. Existing strategies for enhancing diversity predominantly focus on intervening during the generation trajectory. We identify a critical oversight that the standard Gaussian initialization often causes tr…

Cited by 0SourceScholar
2026

Kinematics-Aware Diffusion Policy With Consistent 3D Observation and Action Space for Whole-Arm Robotic Manipulation

RA-L 2026

Full-configuration control of robotic manipulators with awareness of whole-arm kinematics is crucial for many manipulation scenarios involving body collision avoidance or body-object interactions, making it insufficient to consider only the end-effector poses in policy learning. The typical approach

Cited by 1SourceScholar
2026

LLM-MatLogic: Executable Exchange Contracts for Knowledge-Graph Query Answering with Scoped Negation

ICML 2026poster

LLM-to-KG systems frequently fail on exclusion-rich questions because natural-language negation is both scope-sensitive and evidence-dependent: it may constrain only one subgoal/branch and only certain supporting paths, yet such attachment is rarely explicit in text. We propose the Executable Exchan…

Cited by 0SourceScholar
2026

LPPG-RL: Lexicographically Projected Policy Gradient Reinforcement Learning with Subproblem Exploration

AAAI 2026technical

Lexicographic multi-objective problems, which consist of multiple conflicting subtasks with explicit priorities, are common in real-world applications. Despite the advantages of Reinforcement Learning (RL) in single tasks, extending conventional RL methods to prioritized multiple objectives remains

Cited by 0SourcePDFScholar
2026

MedREK: Retrieval-Based Editing for Medical LLMs with Key-Aware Prompts

ICML 2026poster

LLMs hold great promise for healthcare applications, but fast-changing medical knowledge can quickly make their outputs outdated or inaccurate, limiting use in high-stakes settings. Model editing can update LLMs without full retraining, but parameter-based methods often break locality and are risky …

Cited by 0SourceScholar
2026

MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs

ICLR 2026poster

Despite advancements in Multimodal Large Language Models (MLLMs), their proficiency in fine-grained video motion understanding remains critically limited. They often lack inter-frame differencing and tend to average or ignore subtle visual cues. Furthermore, while visual prompting has shown potentia…

Cited by 0SourceScholar
2026

Multimodal Policy Internalization for Conversational Agents

ICLR 2026poster

Modern conversational agents such as ChatGPT and Alexa+ have become indispensable in everyday life. To handle diverse business requirements and enable agentic capabilities, these LLM-based systems often rely on predefined policies, which specify instructions such as model metadata, response styles,…

Cited by 0SourceScholar
2026

Multiplex Heterogeneous Graph Neural Networks with Euclidean-Riemannian Mutual Space Synergy

AAAI 2026technical

Multiplex heterogeneous networks are common in real-world scenarios, where entities interact through diverse types of relations across multiple semantic layers. Recent advances in multiplex heterogeneous graph neural networks have achieved remarkable results by incorporating node and relation types

Cited by 0SourcePDFScholar
2026

NAIPv2: Debiased Pairwise Learning for Efficient Paper Quality Estimation

ICLR 2026poster

The ability to estimate the quality of scientific papers is central to how both humans and AI systems will advance scientific knowledge in the future. However, existing LLM-based estimation methods suffer from high inference cost, whereas the faster direct score regression approach is limited by sca…

Cited by 0SourcecodeScholar
2026

Pixel Motion Diffusion is What We Need for Robot Control

CVPR 2026

We present DAWN (Diffusion is All We Need for robot control), a unified diffusion-based framework for language-conditioned robotic manipulation that bridges high-level motion intent and low-level robot action via structured pixel motion representation. In DAWN, both the high-level and low-level cont

Cited by 0SourcecodeScholar
2026

Point2RBox-v3: Self-Bootstrapping from Point Annotations via Integrated Pseudo-Label Refinement and Utilization

ICLR 2026poster

Driven by the growing need for Oriented Object Detection (OOD), learning from point annotations under a weakly-supervised framework has emerged as a promising alternative to costly and laborious manual labeling. In this paper, we discuss two deficiencies in existing point-supervised methods: ineffic…

Cited by 0SourcecodeScholar
2026

Reconstructing KV Caches with Cross-Layer Fusion for Enhanced Transformers

ICLR 2026poster

Transformer decoders have achieved strong results across tasks, but the memory required for the KV cache becomes prohibitive at long sequence lengths. Although Cross-layer KV Cache sharing (e.g., YOCO, CLA) offers a path to mitigate KV Cache bottleneck, it typically underperforms within-layer method…

Cited by 0SourcecodeScholar
2026

RoadGIE: Towards A Global-Scale Aerial Benchmark for Generalizable Interactive Road Extraction

CVPR 2026

Accurate road segmentation from aerial imagery is fundamental to many geospatial applications. However, existing datasets often suffer from limited scene diversity, low semantic granularity, and poor structural continuity, restricting their generalization across environments. To address these challe

Cited by 0SourcecodeScholar
2026

SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View Fusion

CVPR 2026

Error detection is crucial in industrial training, healthcare, and assembly quality control. Most existing work assumes a single-view setting and cannot handle the practical case where a third-person (exo) demonstration is used to assess a first-person (ego) imitation. We formalize Ego->Exo Imitatio

Cited by 0SourceScholar
2026

SLIP-RS: Structured-Attribute Language-Image Pre-Training for Remote Sensing Object Detection

ICML 2026poster

Existing language-image pre-training for remote sensing object detection is constrained by Monolithic Label Learning, which relies on exhaustively enumerating open-set categories via black-box data to acquire fine-grained representations, creating a dependency incompatible with the domain's inherent…

Cited by 0SourceScholar
2026

SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection

AAAI 2026technical

With the rapid advancement of remote sensing technology, high-resolution multi-modal imagery is now more widely accessible. Conventional object detection models are trained on a single dataset, often restricted to a specific imaging modality and annotation format. However, such an approach overlooks

Cited by 0SourcePDFScholar
2026

ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows

ICLR 2026poster

Large Language Models (LLMs) have extended their impact beyond Natural Language Processing, substantially fostering the development of interdisciplinary research. Recently, various LLM-based agents have been developed to assist scientific discovery progress across multiple aspects and domains. Among…

Cited by 0SourcecodeScholar
2026

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment

ICML 2026poster

Neural speech codecs based on Vector-Quantized VAEs (VQ-VAEs) are core audio tokenizers for speech LLMs, yet their reconstruction fidelity is bottlenecked by quantization error. Instead of modifying the quantizer or increasing model capacity—common approaches that complicate downstream language mode…

Cited by 0SourceScholar
2026

Sentiment-aware Rating-based Recommendation via Semantic-enhanced Item Alignment

IJCAI 2026

Leveraging review texts to mine deep user preferences is vital for recommendation. However, existing methods neglect the positive-negative counteraction and rely on noisy hard sentiment thresholds. Furthermore, the feature density asymmetry causes dense semantic features to overwhelm sparse collabor

Cited by 0Scholar
2026

SpatioTemporal Difference Network for Video Depth Super-Resolution

AAAI 2026technical

Depth super-resolution has achieved impressive performance, and the incorporation of multi-frame information further enhances reconstruction quality. Nevertheless, statistical analyses reveal that video depth super-resolution remains affected by pronounced long-tailed distributions, with the long-ta

Cited by 0SourcePDFScholar
2026

Strip R-CNN: Large Strip Convolution for Remote Sensing Object Detection

AAAI 2026technical

In this paper, we show that current approaches using large square kernels or transformer-based global modeling aggregate contextual information uniformly across spatial dimensions, leading to feature dilution and localization errors for elongated targets. To mitigate this issue, we propose Strip R-C

Cited by 0SourcePDFScholar
2026

T-POP: Test-Time Personalization with Online Preference Feedback

ICML 2026poster

Personalizing large language models (LLMs) to individual user preferences is a critical step beyond generating generically helpful responses. However, current personalization methods are ill-suited for new users, as they typically require either slow, resource-intensive fine-tuning or a substantial …

Cited by 0SourceScholar
2026

TCoT: Trajectory Chain-of-Thoughts for Robotic Manipulation with Failure Recovery in Vision-Language-Action Model

AAAI 2026technical

Recent advances in vision-language-action (VLA) models have demonstrated impressive generalization for robotic manipulation. However, these models often operate by directly mapping visual and linguistic inputs to subsequent actions, lacking intermediate task planning, along with failure detection an

Cited by 0SourcePDFScholar
2026

TRM-VLA: Temporal-Aware Chain-of-Thought Reasoning and Memorization for Vision-Language-Action Models

CVPR 2026

Vision-Language-Action (VLA) models have emerged as a powerful paradigm for general robotic manipulation. However, existing approaches typically omit intermediate reasoning steps and directly regress actions, limiting reasoning interpretability and performance in long-horizon or compositional tasks.

Cited by 0SourceScholar
2026

TrustGen: A Platform of Dynamic Benchmarking on the Trustworthiness of Generative Foundation Models

ICLR 2026poster

Generative foundation models (GenFMs), such as large language models and text-to-image systems, have demonstrated remarkable capabilities in various downstream applications. As they are increasingly deployed in high-stakes applications, assessing their trustworthiness has become both a critical nece…

Cited by 0SourceScholar
2026

UltraHiT: A Hierarchical Transformer Architecture for Generalizable Internal Carotid Artery Robotic Ultrasonography

ICRA 2026poster

Carotid ultrasound is crucial for the assessment of cerebrovascular health, particularly the internal carotid artery (ICA). While previous research has explored automating carotid ultrasound, none has tackled the challenging ICA. This is primarily due to its deep location, tortuous course, and signi…

2026

UniChange: Unifying Change Detection with Multimodal Large Language Model

CVPR 2026

Change detection (CD) is a fundamental task for monitoring and analysing land cover dynamics. While recent high performance models and high quality datasets have significantly advanced the field, a critical limitation persists. Current models typically acquire limited knowledge from single-type anno

Cited by 0SourcecodeScholar
2026

Unified Episodic and Semantic Memory via Modulating Transformer FeedForward Layers

ICML 2026poster

It is widely recognized that, after generative pre-training, Transformer FeedForward layers implicitly function as semantic memory, encoding linguistic and factual knowledge, while the contexts in key–value (KV) cache contain raw events, serving as the source of models' episodic memory. In this work…

Cited by 0SourceScholar
2026

Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted Pretraining

ICML 2026poster

Heterogeneous multi-modal remote sensing object detection aims to accurately detect objects from diverse sensors (e.g., RGB, SAR, Infrared). Existing approaches largely adopt a late alignment paradigm, in which modality alignment and task-specific optimization are entangled during downstream fine-tu…

Cited by 0SourceScholar
2026

Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models

AAAI 2026technical

Recent advancements in Large Video Language Models (LVLMs) have highlighted their potential for multi-modal understanding, yet evaluating their factual grounding in videos remains a critical unsolved challenge. To address this gap, we introduce Video SimpleQA, the first comprehensive benchmark tailo

Cited by 0SourcePDFScholar
2026

VistaBot: View-Robust Robot Manipulation Via Spatiotemporal-Aware View Synthesis

ICRA 2026poster

Recently, end-to-end robotic manipulation models have gained significant attention for their generalizability and scalability. However, they often suffer from limited robustness to camera viewpoint changes when training with a fixed camera. In this paper, we propose VistaBot, a novel framework that …

2026

WOW-Seg: A Word-free Open World Segmentation Model

ICLR 2026poster

Open world image segmentation aims to achieve precise segmentation and semantic understanding of targets within images by addressing the infinitely open set of object categories encountered in the real world. However, traditional closed-set segmentation approaches struggle to adapt to complex open…

Cited by 0SourceScholar
2026

When Scores Learn Geometry: Rate Separations under the Manifold Hypothesis

ICLR 2026poster

Score-based methods, such as diffusion models and Bayesian inverse problems, are often interpreted as learning the data distribution in the low-noise limit ($\sigma \to 0$). In this work, we propose an alternative perspective: their success arises from implicitly learning the data manifold rather th…

Cited by 0SourceScholar
2025

A Survey of LLM-based Agents in Medicine: How far are we from Baymax?

ACL 2025finding

Large Language Models (LLMs) are transforming healthcare through LLM-based agents that can understand and assist with medical tasks. This survey examines the architectures, applications, and challenges of LLM-based agents in medicine. We analyze key components including system profiles, clinical pla…

Cited by 0SourcePDFScholar
2025

ASD-iLLM:An Intervention Large Language Model for Autistic Children based on Real Clinical Dialogue Intervention Dataset

EMNLP 2025

Currently, leveraging large language models (LLMs) for autism intervention is a significant yet challenging task, particularly when directly employing LLMs as an intervention doctor. Researchers have mainly focused on using prompt engineering for role play as an intervention doctor and integrating a

2025

Advancing Textual Prompt Learning with Anchored Attributes

ICCV 2025poster

Textual-based prompt learning methods primarily employ multiple learnable soft prompts and hard class tokens in a cascading manner as text inputs, aiming to align image and text (category) spaces for downstream tasks. However, current training is restricted to aligning images with predefined known c…

2025

BookBot: A Robotic Manipulation Benchmark for Voice-Driven Book Recognition and Grasping in Cluttered Environments

IROS 2025

Books, as enduring repositories of cultural heritage as well as knowledge, play a fundamental role in human development. Although advances in embodied AI and robotics revolutionize automation in domains, e.g., manufacturing and logistics, robotic book manipulation remains an underexplored frontier.

Cited by 0SourcecodeScholar
2025

CAARMA: Class Augmentation with Adversarial Mixup Regularization

EMNLP 2025

Speaker verification is a typical zero-shot learning task, where inference of unseen classes is performed by comparing embeddings of test instances to known examples. The models performing inference must hence naturally generate embeddings that cluster same-class instances compactly, while maintaini

2025

Can Large Language Models Act as Ensembler for Multi-GNNs?

EMNLP 2025

Graph Neural Networks (GNNs) have emerged as powerful models for learning from graph-structured data. However, GNNs lack the inherent semantic understanding capability of rich textual node attributes, limiting their effectiveness in applications. On the other hand, we empirically observe that for ex

2025

ChatCAD: An MLLM-Guided Framework for Zero-shot CAD Drawing Restoration

ICASSP 2025accepted

CAD drawing restoration is one of the most urgent needs in industrial manufacturing. The existing research focuses on the digitization of CAD drawings, However, there are actually many problems in digitized CAD drawings due to the upgrading of engineering drafting software, and it is difficult to re…

Cited by 0SourceScholar
2025

Corruption-Robust Variance-aware Algorithms for Generalized Linear Bandits under Heavy-tailed Rewards

UAI 2025

Stochastic linear bandits have recently received significant attention in sequential decision-making. However, real-world challenges such as heavy-tailed noise, reward corruption, and nonlinear reward functions remain difficult to address. To tackle these difficulties, we propose GAdaOFUL, a novel a

2025

Coupling-based Convergence Diagnostic and Stepsize Scheme for Stochastic Gradient Descent

AAAI 2025technical

The convergence behavior of Stochastic Gradient Descent (SGD) crucially depends on the stepsize configuration. When using a constant stepsize, the SGD iterates form a Markov chain, enjoying fast convergence during the initial transient phase. However, when reaching stationarity, the iterates oscilla…

2025

Cue3D: Quantifying the Role of Image Cues in Single-Image 3D Generation

NeurIPS 2025spotlight

Humans and traditional computer vision methods rely on a diverse set of monocular cues to infer 3D structure from a single image, such as shading, texture, silhouette, etc. While recent deep generative models have dramatically advanced single-image 3D generation, it remains unclear which image cues…

Cited by 0SourceScholar
2025

DISTA-Net: Dynamic Closely-Spaced Infrared Small Target Unmixing

ICCV 2025poster

Resolving closely-spaced small targets in dense clusters presents a significant challenge in infrared imaging, as the overlapping signals hinder precise determination of their quantity, sub-pixel positions, and radiation intensities. While deep learning has advanced the field of infrared small targe…

2025

DRARL: Disengagement-Reason-Augmented Reinforcement Learning for Efficient Improvement of Autonomous Driving Policy

IROS 2025

With the increasing presence of automated vehicles on open roads under driver supervision, disengagement cases are becoming more prevalent. While some data-driven planning systems attempt to directly utilize these disengagement cases for policy improvement, the inherent scarcity of disengagement dat

Cited by 3SourceScholar
2025

DVS-Aware Visual Perception for Pose Estimation of Mobile Robots with Neuromorphic Implementation

ICRA 2025

The Dynamic Vision Sensor (DVS) is a distinctive visual sensor that exclusively responds to alterations in pixel brightness, enabling the real-time capture of swift and subtle movements with reduced power consumption and data bandwidth requirements. This paper proposes a DVS-aware visual perception

Cited by 1SourceScholar
2025

DeMAC: Enhancing Multi-Agent Coordination with Dynamic DAG and Manager-Player Feedback

EMNLP 2025

Multi-agent systems (MAS) powered by large language models (LLMs) have shown potential in tackling multifaceted problems through advanced understanding and reasoning. However, they struggle to adapt to evolving task dependencies and to handle uncertainties, such as shifting priorities or unpredictab

Cited by 0SourcePDFScholar
2025

Deep Height Decoupling for Precise Vision-Based 3D Occupancy Prediction

ICRA 2025

The task of vision-based 3D occupancy prediction aims to reconstruct 3D geometry and estimate its semantic classes from 2D color images, where the 2D-to-3D view transformation is an indispensable step. Most previous methods conduct forward projection, such as BEVPooling and VoxelPooling, both of whi

Cited by 17SourcecodeScholar
2025

Demystifying Small Language Models for Edge Deployment

ACL 2025long

Small language models (SLMs) have emerged as a promising solution for deploying resource-constrained devices, such as smartphones and Web of Things. This work presents the first comprehensive study of over 60 SLMs such as Microsoft Phi and Google Gemma that are publicly accessible. Our findings show…

Cited by 0SourcePDFScholar
2025

Distribution-aware Fairness Learning in Medical Image Segmentation From A Control-Theoretic Perspective

ICML 2025spotlight

Ensuring fairness in medical image segmentation is critical due to biases in imbalanced clinical data acquisition caused by demographic attributes (e.g., age, sex, race) and clinical factors (e.g., disease severity). To address these challenges, we introduce Distribution-aware Mixture of Experts (dM…

2025

ECHOPulse: ECG Controlled Echocardio-gram Video Generation

ICLR 2025poster

Echocardiography (ECHO) is essential for cardiac assessments, but its video quality and interpretation heavily relies on manual expertise, leading to inconsistent results from clinical and portable devices. ECHO video generation offers a solution by improving automated monitoring through synthetic d…

Cited by 5SourcePDFScholar
2025

Enhancing Cognition and Explainability of Multimodal Foundation Models with Self-Synthesized Data

ICLR 2025poster

Large Multimodal Models (LMMs), or Vision-Language Models (VLMs), have shown impressive capabilities in a wide range of visual tasks. However, they often struggle with fine-grained visual reasoning, failing to identify domain-specific objectives and provide justifiable explanations for their predict…

2025

Enhancing Indoor Occupancy Prediction via Sparse Query-Based Multi-Level Consistent Knowledge Distillation

RA-L 2025

Occupancy prediction provides critical geometric and semantic understanding for robotics but faces efficiency-accuracy trade-offs. Current dense methods suffer computational waste on empty voxels, while sparse query-based approaches lack robustness in diverse and complex indoor scenes. In this paper

Cited by 1SourceScholar
2025

Enhancing LLM-based Hatred and Toxicity Detection with Meta-Toxic Knowledge Graph

ACL 2025finding

The rapid growth of social media platforms has raised significant concerns regarding online content toxicity. When Large Language Models (LLMs) are used for toxicity detection, two key challenges emerge: 1) the absence of domain-specific toxicity knowledge leads to false negatives; 2) the excessive…

2025

Explain-Analyze-Generate: A Sequential Multi-Agent Collaboration Method for Complex Reasoning

COLING 2025main

Exploring effective collaboration among multiple large language models (LLMs) represents an active research direction, with multiagent debate (MAD) emerging as a popular approach. MAD involves LLMs independently generating responses and refining their own responses by incorporating feedback from oth…

Cited by 14SourcePDFScholar
2025

Fading to Grow: Growing Preference Ratios via Preference Fading Discrete Diffusion for Recommendation

NeurIPS 2025poster

Recommenders aim to rank items from a discrete item corpus in line with user interests, yet suffer from extremely sparse user preference data. Recent advances in diffusion models have inspired diffusion-based recommenders, which alleviate sparsity by injecting noise during a forward process to preve…

Cited by 0SourceScholar
2025

Flow-Aware Navigation of Magnetic Micro-Robots in Complex Fluids via PINN-Based Prediction

IROS 2025

While magnetic micro-robots have demonstrated significant potential across various applications, including drug delivery and microsurgery, the open issue of precise navigation and control in complex fluid environments is crucial for in vivo implementation. This paper introduces a novel flow-aware na

Cited by 1SourceScholar
2025

From Words to Worth: Newborn Article Impact Prediction with LLM

AAAI 2025technical

Predicting the future impact of newly published articles is pivotal for advancing scientific discovery in an era of unprecedented scholarly expansion. This paper introduces a promising approach, leveraging the capabilities of LLMs to predict the future impact of newborn articles solely based on titl…

Cited by 2SourcePDFScholar
2025

HSI-GPT: A General-Purpose Large Scene-Motion-Language Model for Human Scene Interaction

CVPR 2025highlight

While flourishing developments have been witnessed in text-to-motion generation, synthesizing physically realistic, controllable, language-conditioned Human Scene Interactions (HSI) remains a relatively underexplored landscape. Current HSI methods naively rely on conditional Variational AutoEncoder…

Cited by 0SourcePDFScholar
2025

Hierarchically Controlled Deformable 3D Gaussians for Talking Head Synthesis

AAAI 2025technical

Audio-driven talking head synthesis is a critical task in digital human modeling. While recent advances using diffusion models and Neural Radiance Fields (NeRF) have improved visual quality, they often require substantial computational resources, limiting practical deployment. We present a novel fra…

Cited by 1SourcePDFScholar
2025

Idea23D: Collaborative LMM Agents Enable 3D Model Generation from Interleaved Multimodal Inputs

COLING 2025main

With the success of 2D diffusion models, 2D AIGC content has already transformed our lives. Recently, this success has been extended to 3D AIGC, with state-of-the-art methods generating textured 3D models from single images or text. However, we argue that current 3D AIGC methods still don’t fully un…

2025

ImageFolder: Autoregressive Image Generation with Folded Tokens

ICLR 2025poster

Image tokenizers are crucial for visual generative models, \eg, diffusion models (DMs) and autoregressive (AR) models, as they construct the latent representation for modeling. Increasing token length is a common approach to improve image reconstruction quality. However, tokenizers with longer token…

2025

In-Pipe Navigation Development Environment and a Smooth Path Planning Method on Pipeline Surface

ICRA 2025

Autonomous in-pipe inspection robots can automatically navigate through complex pipeline networks and detect potential risks from corrosion and defects, demonstrating great potential for replacing costly manual inspections. However, there is no publicly available simulation environment where researc

Cited by 6SourceScholar
2025

Initializing and Retrofitting Key-Value Adaptors for Traceable Model Editing

ACL 2025finding

As the insight of knowledge storage in language models deepens, the ability to perform CRUD (Create, Read, Update, Delete) operations on language models becomes increasingly indispensable for satisfying the demands of managing rapidly updating knowledge. Considering the high cost of fine-tuning lang…

2025

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption

CVPR 2025poster

Text-to-video generation has evolved rapidly in recent years, delivering remarkable results. Training typically relies on video-caption paired data, which plays a crucial role in enhancing generation performance. However, current video captions often suffer from insufficient details, hallucinations…

2025

LLMsPark: A Benchmark for Evaluating Large Language Models in Strategic Gaming Contexts

EMNLP 2025

As large language models (LLMs) advance across diverse tasks, the need for comprehensive evaluation beyond single metrics becomes increasingly important.To fully assess LLM intelligence, it is crucial to examine their interactive dynamics and strategic behaviors.We present LLMsPark, a game theory–ba

2025

LLaRA: Supercharging Robot Learning Data for Vision-Language Policy

ICLR 2025poster

Vision Language Models (VLMs) have recently been leveraged to generate robotic actions, forming Vision-Language-Action (VLA) models. However, directly adapting a pretrained VLM for robotic control remains challenging, particularly when constrained by a limited number of robot demonstrations. In this…

2025

Learning to Plan Like the Human Brain via Visuospatial Perception and Semantic-Episodic Synergistic Decision-Making

NeurIPS 2025poster

Motion planning in high-dimensional continuous spaces remains challenging due to complex environments and computational constraints. Although learning-based planners, especially graph neural network (GNN)-based, have significantly improved planning performance, they still struggle with inaccurate gr…

Cited by 0SourceScholar
2025

Let Your Features Tell The Differences: Understanding Graph Convolution By Feature Splitting

ICLR 2025poster

Graph Neural Networks (GNNs) have demonstrated strong capabilities in processing structured data. While traditional GNNs typically treat each feature dimension equally important during graph convolution, we raise an important question: **Is the graph convolution operation equally beneficial for each…

2025

Let’s Be Self-generated via Step by Step: A Curriculum Learning Approach to Automated Reasoning with Large Language Models

ACL 2025finding

While Chain of Thought (CoT) prompting approaches have significantly consolidated the reasoning capabilities of large language models (LLMs), they still face limitations that require extensive human effort or have performance needs to be improved. Existing endeavors have focused on bridging these ga…

Cited by 0SourcePDFScholar
2025

Leveraging Large Language Models for Node Generation in Few-Shot Learning on Text-Attributed Graphs

AAAI 2025technical

Text-attributed graphs have recently garnered significant attention due to their wide range of applications in web domains. Existing methodologies employ word embedding models for acquiring text representations as node features, which are subsequently fed into Graph Neural Networks (GNNs) for traini…

2025

LogiGraph: Logical Reasoning with Contrastive Learning and Lightweight Graph Networks

COLING 2025main

Logical reasoning is a crucial factor in machine reading comprehension tasks (MRC). Existing methods suffer from the balance between semantic and explicit logical relation representations, in which some emphasize contextual semantics, while others pay more attention to explicit logical features. Add…

Cited by 0SourcePDFScholar
2025

MaskDGNN: Self-Supervised Dynamic Graph Neural Networks with Activeness-aware Temporal Masking

IJCAI 2025

Integrating dynamics into graph neural networks (GNNs) provides deeper insights into the evolution of dynamic graphs, thereby enhancing the temporal representation in real-world dynamic network problems. Existing methods extracting critical information from dynamic graphs face two key challenges, ei

2025

Masked Autoencoders Are Effective Tokenizers for Diffusion Models

ICML 2025spotlight

Recent advances in latent diffusion models have demonstrated their effectiveness for high-resolution image synthesis. However, the properties of the latent space from tokenizer for better learning and generation of diffusion models remain under-explored. Theoretically and empirically, we find that i…

Cited by 8SourcePDFScholar
2025

Mitigating Spurious Correlations via Counterfactual Contrastive Learning

EMNLP 2025

Identifying causal relationships rather than spurious correlations between words and class labels plays a crucial role in building robust text classifiers. Previous studies proposed using causal effects to distinguish words that are causally related to the sentiment, and then building robust text cl

Cited by 0SourcePDFScholar
2025

Mitigating the Privacy–Utility Trade-off in Decentralized Federated Learning via f-Differential Privacy

NeurIPS 2025spotlight

Differentially private (DP) decentralized Federated Learning (FL) allows local users to collaborate without sharing their data with a central server. However, accurately quantifying the privacy budget of private FL algorithms is challenging due to the co-existence of complex algorithmic components s…

Cited by 0SourceScholar
2025

Multi-Sensor Object Anomaly Detection: Unifying Appearance, Geometry, and Internal Properties

CVPR 2025poster

Object anomaly detection is essential for industrial quality inspection, yet traditional single-sensor methods face critical limitations. They fail to capture the wide range of anomaly types, as single sensors are often constrained to either external appearance, geometric structure, or internal prop…

2025

Multi-clue Consistency Learning to Bridge Gaps Between General and Oriented Object in Semi-supervised Detection

AAAI 2025technical

While existing semi-supervised object detection (SSOD) methods perform well in general scenes, they encounter challenges in handling oriented objects in aerial images. We experimentally find three gaps between general and oriented object detection in semi-supervised learning: 1) Sampling inconsist…

2025

Multi-level Relevance Document Identifier Learning for Generative Retrieval

ACL 2025long

Generative Retrieval (GR) introduces a new information retrieval paradigm that directly generates unique document identifiers (DocIDs). The key challenge of GR lies in creating effective yet discrete DocIDs that preserve semantic relevance for similar documents while differentiating dissimilar ones.…

2025

Multimodal Document-level Triple Extraction via Dynamic Graph Enhancement and Relation-Aware Reflection

EMNLP 2025

Multimodal documents, which are among the most prevalent data formats, combine a large amount of textual and visual content. Extracting structured triples knowledge from these documents is a highly valuable task, aimed at helping users efficiently acquire key entities and their relationships. Howeve

2025

Non-Contact Dexterous Micromanipulation With Multiple Optoelectronic Robots

RA-L 2025

Micromanipulation systems leverage automation and robotic technologies to improve the precision, repeatability, and efficiency of various tasks at the microscale. However, current approaches are typically limited to specific objects or tasks, which necessitates the use of custom tools and specialize

Cited by 5SourceScholar
2025

Not All Layers of LLMs Are Necessary During Inference

IJCAI 2025

Due to the large number of parameters, the inference phase of Large Language Models (LLMs) is resource-intensive. However, not all requests posed to LLMs are equally difficult to handle. Through analysis, we show that for some tasks, LLMs can achieve results comparable to the final output at some in

Cited by 0SourcePDFScholar
2025

On the Empirical Power of Goodness-of-Fit Tests in Watermark Detection

NeurIPS 2025spotlight

Large language models (LLMs) raise concerns about content authenticity and integrity because they can generate human-like text at scale. Text watermarks, which embed detectable statistical signals into generated text, offer a provable way to verify content origin. Many detection methods rely on pivo…

Cited by 0SourceScholar
2025

OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

ICLR 2025poster

Text-to-video (T2V) generation has recently garnered significant attention thanks to the large multi-modality model Sora. However, T2V generation still faces two important challenges: 1) Lacking a precise open sourced high-quality dataset. The previously popular video datasets, e.g.WebVid-10M and Pa…

Cited by 62SourcePDFScholar
2025

PA-RAG: RAG Alignment via Multi-Perspective Preference Optimization

NAACL 2025long

The emergence of Retrieval-augmented generation (RAG) has alleviated the issues of outdated and hallucinatory content in the generation of large language models (LLMs), yet it still reveals numerous limitations. When a general-purpose LLM serves as the RAG generator, it often suffers from inadequate…

2025

PRDetect: Perturbation-Robust LLM-generated Text Detection Based on Syntax Tree

NAACL 2025findings

As LLM-generated text becomes increasingly prevalent on the internet, often containing hallucinations or biases, detecting such content has emerged as a critical area of research.Recent methods have demonstrated impressive performance in detecting text generated entirely by LLMs.However, in real-wor…

2025

Permitted Knowledge Boundary: Evaluating the Knowledge-Constrained Responsiveness of Large Language Models

EMNLP 2025

With the advancement of large language models (LLMs), recent research has raised concerns about their controllability.. In this paper, we argue for the importance of Knowledge-Constrained Responsiveness (KCR), ensuring that LLMs comply with human-defined constraints. However, KCR is an implicit and

2025

Preference Adaptive and Sequential Text-to-Image Generation

ICML 2025poster

We address the problem of interactive text-to-image (T2I) generation, designing a reinforcement learning (RL) agent which iteratively improves a set of generated images for a user through a sequence of prompt expansions. Using human raters, we create a novel dataset of sequential preferences, which…

Cited by 0SourcePDFScholar
2025

REOBench: Benchmarking Robustness of Earth Observation Foundation Models

NeurIPS 2025poster

Earth observation foundation models have shown strong generalization across multiple Earth observation tasks, but their robustness under real-world perturbations remains underexplored. To bridge this gap, we introduce REOBench, the first comprehensive benchmark for evaluating the robustness of Earth…

Cited by 0SourcecodeScholar
2025

RSAR: Restricted State Angle Resolver and Rotated SAR Benchmark

CVPR 2025poster

Rotated object detection has made significant progress in the optical remote sensing. However, advancements in the Synthetic Aperture Radar (SAR) field are laggard behind, primarily due to the absence of a large-scale dataset. Annotating such a dataset is inefficient and costly. A promising solution…

2025

Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think

NeurIPS 2025oral

REPA and its variants effectively mitigate training challenges in diffusion models by incorporating external visual representations from pretrained models, through alignment between the noisy hidden projections of denoising networks and foundational clean image representations. We argue that the ext…

Cited by 0SourcecodeScholar
2025

Rethinking Point Cloud Data Augmentation: Topologically Consistent Deformation

ICML 2025poster

Data augmentation has been widely used in machine learning. Its main goal is to transform and expand the original data using various techniques, creating a more diverse and enriched training dataset. However, due to the disorder and irregularity of point clouds, existing methods struggle to enrich g…

2025

Revisiting End-to-End Learning with Slide-level Supervision in Computational Pathology

NeurIPS 2025poster

Pre-trained encoders for offline feature extraction followed by multiple instance learning (MIL) aggregators have become the dominant paradigm in computational pathology (CPath), benefiting cancer diagnosis and prognosis. However, performance limitations arise from the absence of encoder fine-tuning…

Cited by 0SourcecodeScholar
2025

Robotic In-Hand Manipulation for Large-Range Precise Object Movement: The RGMC Champion Solution

RA-L 2025

In-hand manipulation using multiple dexterous fingers is a critical robotic skill that can reduce the reliance on large arm motions, thereby saving space and energy. This letter focuses on in-grasp object movement, which refers to manipulating an object to a desired pose through only finger motions

Cited by 9SourceScholar
2025

SGCD: Subtask-Guided Causal-Debiasing Framework for Robust Cross-Utterance Sentiment Quadruple Extraction in Dialogues

EMNLP 2025

The rise of digital social media has generated a vast amount of conversational data on platforms like Twitter and Reddit, allowing users to express sentiments through multi-turn dialogues. Dialogue-level aspect-based sentiment quadruple analysis (DiaASQ) seeks to extract structured information in th

2025

SKIntern: Internalizing Symbolic Knowledge for Distilling Better CoT Capabilities into Small Language Models

COLING 2025main

Small Language Models (SLMs) are attracting attention due to the high computational demands and privacy concerns of Large Language Models (LLMs). Some studies fine-tune SLMs using Chains of Thought (CoT) data distilled from LLMs, aiming to enhance their reasoning ability. Furthermore, Some CoT disti…

2025

Scalable Benchmarking and Robust Learning for Noise-Free Ego-Motion and 3D Reconstruction from Noisy Video

ICLR 2025poster

We aim to redefine robust ego-motion estimation and photorealistic 3D reconstruction by addressing a critical limitation: the reliance on noise-free data in existing models. While such sanitized conditions simplify evaluation, they fail to capture the unpredictable, noisy complexities of real-world…

2025

See the World, Discover Knowledge: A Chinese Factuality Evaluation for Large Vision Language Models

ACL 2025finding

The evaluation of factual accuracy in large vision language models (LVLMs) has lagged behind their rapid development, making it challenging to fully reflect these models’ knowledge capacity and reliability. In this paper, we introduce the first factuality-based visual question-answering benchmark in…

Cited by 0SourcePDFScholar
2025

See through the Dark: Learning Illumination-affined Representations for Nighttime Occupancy Prediction

NeurIPS 2025poster

Occupancy prediction aims to estimate the 3D spatial distribution of occupied regions along with their corresponding semantic labels. Existing vision-based methods perform well on daytime benchmarks but struggle in nighttime scenarios due to limited visibility and challenging lighting conditions. To…

Cited by 0SourcecodeScholar
2025

Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous Driving

ICLR 2025poster

Understanding world dynamics is crucial for planning in autonomous driving. Recent methods attempt to achieve this by learning a 3D occupancy world model that forecasts future surrounding scenes based on current observation. However, 3D occupancy labels are still required to produce promising result…

2025

SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models

ICCV 2025poster

The increasing application of multi-modal large language models (MLLMs) across various sectors has spotlighted the essence of their output reliability and accuracy, particularly their ability to produce content grounded in factual information (e.g. common and domain-specific knowledge). In this work…

Cited by 0SourcePDFScholar
2025

SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer

CVPR 2025poster

Efficient image tokenization with high compression ratios remains a critical challenge for training generative models.We present SoftVQ-VAE, a continuous image tokenizer that leverages soft categorical posteriors to aggregate multiple codewords into each latent token, substantially increasing the re…

2025

Symmetry Strikes Back: From Single-Image Symmetry Detection to 3D Generation

CVPR 2025highlight

Symmetry is a ubiquitous and fundamental property in the visual world, serving as a critical cue for perception and structure interpretation. This paper investigates the detection of 3D reflection symmetry from a single RGB image, and reveals its significant benefit on single-image 3D generation. We…

Cited by 0SourcePDFScholar
2025

TF-Mamba: Text-enhanced Fusion Mamba with Missing Modalities for Robust Multimodal Sentiment Analysis

EMNLP 2025

Multimodal Sentiment Analysis (MSA) with missing modalities has attracted increasing attention recently. While current Transformer-based methods leverage dense text information to maintain model robustness, their quadratic complexity hinders efficient long-range modeling and multimodal fusion. To th

2025

Text Detoxification: Data Efficiency, Semantic Preservation and Model Generalization

EMNLP 2025

The widespread dissemination of toxic content on social media poses a serious threat to both online environments and public discourse, highlighting the urgent need for detoxification methods that effectively remove toxicity while preserving the original semantics.However, existing approaches often s

2025

TreeEval: Benchmark-Free Evaluation of Large Language Models through Tree Planning

AAAI 2025technical

Recently, numerous new benchmarks have been established to evaluate the performance of large language models (LLMs) via either computing a holistic score or employing another LLM as a judge. However, these approaches suffer from data leakage due to the open access of the benchmark and inflexible ev…

2025

UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts

ICASSP 2025accepted

Emotional Text-to-Speech (E-TTS) synthesis has garnered significant attention in recent years due to its potential to revolutionize human-computer interaction. However, current E-TTS approaches often struggle to capture the intricacies of human emotions, primarily relying on oversimplified emotional…

Cited by 0SourceScholar
2025

UltraDP: Generalizable Carotid Ultrasound Scanning with Force-Aware Diffusion Policy

IROS 2025

Ultrasound scanning is a critical imaging technique for real-time, non-invasive diagnostics. However, variations in patient anatomy and complex human-in-the-loop interactions pose significant challenges for autonomous robotic scanning. Existing ultrasound scanning robots are commonly limited to rela

Cited by 3SourceScholar
2025

Understanding Long Videos with Multimodal Language Models

ICLR 2025poster

Large Language Models (LLMs) have allowed recent LLM-based approaches to achieve excellent performance on long-video understanding benchmarks. We investigate how extensive world knowledge and strong reasoning skills of underlying LLMs influence this strong performance. Surprisingly, we discover that…

2025

Understanding Representation Dynamics of Diffusion Models via Low-Dimensional Modeling

NeurIPS 2025poster

Diffusion models, though originally designed for generative tasks, have demonstrated impressive self-supervised representation learning capabilities. A particularly intriguing phenomenon in these models is the emergence of unimodal representation dynamics, where the quality of learned features peaks…

Cited by 0SourceScholar
2025

Who You Are Matters: Bridging Interests and Social Roles via LLM-Enhanced Logic Recommendation

NeurIPS 2025poster

Recommender systems filter contents/items valuable to users by inferring preferences from user features and historical behaviors. Mainstream approaches follow the learning-to-rank paradigm, which focus on discovering and modeling item topics (e.g., categories), and capturing user preferences on the…

Cited by 0SourcecodeScholar
2025

XFormParser: A Simple and Effective Multimodal Multilingual Semi-structured Form Parser

COLING 2025main

In the domain of Document AI, parsing semi-structured image form is a crucial Key Information Extraction (KIE) task. The advent of pre-trained multimodal models significantly empowers Document AI frameworks to extract key information from form documents in different formats such as PDF, Word, and im…

2024

3DCoMPaT200: Language Grounded Large-Scale 3D Vision Dataset for Compositional Recognition

NeurIPS 2024poster

Understanding objects in 3D at the part level is essential for humans and robots to navigate and interact with the environment. Current datasets for part-level 3D object understanding encompass a limited range of categories. For instance, the ShapeNet-Part and PartNet datasets only include 16, and 2…

2024

A General Framework for Learning from Weak Supervision

ICML 2024poster

Weakly supervised learning generally faces challenges in applicability to various scenarios with diverse weak supervision and in scalability due to the complexity of existing algorithms, thereby hindering the practical deployment. This paper introduces a general framework for learning from weak supe…

2024

A Supervised Information Enhanced Multi-Granularity Contrastive Learning Framework for EEG Based Emotion Recognition

ICASSP 2024accepted

This study introduces a novel Supervised Info-enhanced Contrastive Learning framework for EEG based Emotion Recognition (SI-CLEER). SI-CLEER employs multi-granularity contrastive learning to create robust EEG contextual representations, potentially improving emotion recognition effectiveness. Unlike…

Cited by 0SourceScholar
2024

A Unified Interaction Control Framework for Safe Robotic Ultrasound Scanning with Human-Intention-Aware Compliance

IROS 2024

The ultrasound scanning robot operates in environments where frequent human-robot interactions occur. Most existing control methods for ultrasound scanning address only one specific interaction situation or implement hard switches between controllers for different situations, which compromises both

Cited by 6SourceScholar
2024

Achieving Near-Optimal Convergence for Distributed Minimax Optimization with Adaptive Stepsizes

NeurIPS 2024poster

In this paper, we show that applying adaptive methods directly to distributed minimax problems can result in non-convergence due to inconsistency in locally computed adaptive stepsizes. To address this challenge, we propose D-AdaST, a Distributed Adaptive minimax method with Stepsize Tracking. The k…

Cited by 0SourcePDFScholar
2024

AlphaFin: Benchmarking Financial Analysis with Retrieval-Augmented Stock-Chain Framework

COLING 2024main

The task of financial analysis primarily encompasses two key areas: stock trend prediction and the corresponding financial question answering. Currently, machine learning and deep learning algorithms (ML&DL) have been widely applied for stock trend predictions, leading to significant progress. Howev…

2024

AltNeRF: Learning Robust Neural Radiance Field via Alternating Depth-Pose Optimization

AAAI 2024technical

Neural Radiance Fields (NeRF) have shown promise in generating realistic novel views from sparse scene images. However, existing NeRF approaches often encounter challenges due to the lack of explicit 3D supervision and imprecise camera poses, resulting in suboptimal outcomes. To tackle these issues,…

Cited by 2SourcePDFScholar
2024

AutoPRM: Automating Procedural Supervision for Multi-Step Reasoning via Controllable Question Decomposition

NAACL 2024long

Recent advancements in large language models (LLMs) have shown promise in multi-step reasoning tasks, yet their reliance on extensive manual labeling to provide procedural feedback remains a significant impediment. To address this challenge, in this paper, we propose a novel self-supervised framewor…

Cited by 25SourcePDFScholar
2024

AutoPrep: An Automatic Preprocessing Framework for In-The-Wild Speech Data

ICASSP 2024accepted

Recently, the utilization of extensive open-sourced text data has significantly advanced the performance of text-based large language models (LLMs). However, the use of in-the-wild large-scale speech data in the speech technology community remains constrained. One reason for this limitation is that…

Cited by 0SourceScholar
2024

Automated Peer Reviewing in Paper SEA: Standardization, Evaluation, and Analysis

EMNLP 2024finding

In recent years, the rapid increase in scientific papers has overwhelmed traditional review mechanisms, resulting in varying quality of publications. Although existing methods have explored the capabilities of Large Language Models (LLMs) for automated scientific reviewing, their generated contents…

2024

Beyond Read-Only: Crafting a Comprehensive Chinese Text-to-SQL Dataset for Database Manipulation and Query

NAACL 2024findings

Text-to-SQL aims to convert natural language into structured query language, which is a challenging task. Current research focuses mainly on read operations and ignores other aspects of database operations such as create, update, and delete operations. The benchmark datasets as well as models that h…

2024

Biomedical Visual Instruction Tuning with Clinician Preference Alignment

NeurIPS 2024poster

Recent advancements in multimodal foundation models have showcased impressive capabilities in understanding and reasoning with visual and textual information. Adapting these foundation models trained for general usage to specialized domains like biomedicine requires large-scale domain-specific instr…

2024

Boosting Language Models Reasoning with Chain-of-Knowledge Prompting

ACL 2024long

Recently, Chain-of-Thought (CoT) prompting has delivered success on complex reasoning tasks, which aims at designing a simple prompt like “Let’s think step by step” or multiple in-context exemplars with well-designed rationales to elicit Large Language Models (LLMs) to generate intermediate reasonin…

2024

CM-TTS: Enhancing Real Time Text-to-Speech Synthesis Efficiency through Weighted Samplers and Consistency Models

NAACL 2024findings

Neural Text-to-Speech (TTS) systems find broad applications in voice assistants, e-learning, and audiobook creation. The pursuit of modern models, like Diffusion Models (DMs), holds promise for achieving high-fidelity, real-time speech synthesis. Yet, the efficiency of multi-step sampling in Diffusi…

2024

Completing Visual Objects via Bridging Generation and Segmentation

ICML 2024poster

This paper presents a novel approach to object completion, with the primary goal of reconstructing a complete object from its partially visible components. Our method, named MaskComp, delineates the completion process through iterative stages of generation and segmentation. In each iteration, the ob…

Cited by 6SourcePDFScholar
2024

Conjoin after Decompose: Improving Few-Shot Performance of Named Entity Recognition

COLING 2024main

Prompt-based methods have been widely used in few-shot named entity recognition (NER). In this paper, we first conduct a preliminary experiment and observe that the key to affecting the performance of prompt-based NER models is the capability to detect entity boundaries. However, most existing model…

2024

Contact-Implicit Model Predictive Control for Dexterous In-hand Manipulation: A Long-Horizon and Robust Approach

IROS 2024poster

Dexterous in-hand manipulation is an essential skill of production and life. However, the highly stiff and mutable nature of contacts limits real-time contact detection and inference, degrading the performance of model-based methods. Inspired by recent advances in contact-rich locomotion and manipul…

Cited by 5SourceScholar
2024

Cross-model Control: Improving Multiple Large Language Models in One-time Training

NeurIPS 2024poster

The number of large language models (LLMs) with varying parameter scales and vocabularies is increasing. While they deliver powerful performance, they also face a set of common optimization needs to meet specific requirements or standards, such as instruction following or avoiding the output of sens…

2024

CrossKD: Cross-Head Knowledge Distillation for Object Detection

CVPR 2024poster

Knowledge Distillation (KD) has been validated as an effective model compression technique for learning compact object detectors. Existing state-of-the-art KD methods for object detection are mostly based on feature imitation. In this paper we present a general and effective prediction mimicking dis…

2024

Crossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised Learning

ICRA 2024poster

Diffusion models have been adopted for behavioral cloning in a sequence modeling fashion, benefiting from their exceptional capabilities in modeling complex data distributions. The standard diffusion-based policy iteratively denoises action sequences from random noise conditioned on the input states…

Cited by 29SourcecodeScholar
2024

DCDepth: Progressive Monocular Depth Estimation in Discrete Cosine Domain

NeurIPS 2024poster

In this paper, we introduce DCDepth, a novel framework for the long-standing monocular depth estimation task. Moving beyond conventional pixel-wise depth estimation in the spatial domain, our approach estimates the frequency coefficients of depth patches after transforming them into the discrete cos…

2024

DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object Detection

AAAI 2024technical

Vehicle-to-Everything (V2X) collaborative perception has recently gained significant attention due to its capability to enhance scene understanding by integrating information from various agents, e.g., vehicles, and infrastructure. However, current works often treat the information from each agent e…

2024

Decoding Natural Images from EEG for Object Recognition

ICLR 2024poster

Electroencephalography (EEG) signals, known for convenient non-invasive acquisition but low signal-to-noise ratio, have recently gained substantial attention due to the potential to decode natural images. This paper presents a self-supervised framework to demonstrate the feasibility of learning imag…

2024

Distilling Knowledge from Large-Scale Image Models for Object Detection

ECCV 2024poster

"Large-scale image models have made great progress in recent years, pushing the boundaries of many vision tasks, , object detection. Considering that deploying large models is impractical in many scenes due to expensive computation overhead, this paper presents a new knowledge distillation method, w…

Cited by 0SourcePDFScholar
2024

Efficient LLM Jailbreak via Adaptive Dense-to-sparse Constrained Optimization

NeurIPS 2024poster

Recent research indicates that large language models (LLMs) are susceptible to jailbreaking attacks that can generate harmful content. This paper introduces a novel token-level attack method, Adaptive Dense-to-Sparse Constrained Optimization (ADC), which has been shown to successfully jailbreak mult…

2024

Efficient Model Learning and Adaptive Tracking Control of Magnetic Micro-Robots for Non-Contact Manipulation

ICRA 2024poster

Magnetic microrobots can be navigated by an external magnetic field to autonomously move within living organisms with complex and unstructured environments. Potential applications include drug delivery, diagnostics, and therapeutic interventions. Existing techniques commonly impart magnetic properti…

Cited by 3SourceScholar
2024

Every Answer Matters: Evaluating Commonsense with Probabilistic Measures

ACL 2024long

Large language models have demonstrated impressive performance on commonsense tasks; however, these tasks are often posed as multiple-choice questions, allowing models to exploit systematic biases. Commonsense is also inherently probabilistic with multiple correct answers. The purpose of “boiling wa…

2024

Eye-gaze Guided Multi-modal Alignment for Medical Representation Learning

NeurIPS 2024poster

In the medical multi-modal frameworks, the alignment of cross-modality features presents a significant challenge. However, existing works have learned features that are implicitly aligned from the data, without considering the explicit relationships in the medical context. This data-reliance may lea…

2024

Federated Learning via Consensus Mechanism on Heterogeneous Data: A New Perspective on Convergence

ICASSP 2024accepted

Federated learning (FL) on heterogeneous data (non-IID data) has recently received great attention. Most existing methods focus on studying the convergence guarantees for the global objective. While these methods can guarantee the decrease of the global objective in each communication round, they fa…

Cited by 0SourceScholar
2024

Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report Generation

ACL 2024long

Fine-grained vision-language models (VLM) have been widely used for inter-modality local alignment between the predefined fixed patches and textual words. However, in medical analysis, lesions exhibit varying sizes and positions, and using fixed patches may cause incomplete representations of lesion…

Cited by 13SourcePDFScholar
2024

Generalized Hole-Filling Strategy for Overlapping Hole-Existing Coprime Arrays for DOA Estimation

ICASSP 2024accepted

The holes in difference coarrays (DCA) of coprime arrays (CA) limit the extension of the aperture thus causing the waste of resources. In this paper, we propose a generalized hole-filling strategy for hole-existing CAs with overlapping subarrays, which allows to extend the aperture completely and ac…

Cited by 0SourceScholar
2024

Imprecise Label Learning: A Unified Framework for Learning with Various Imprecise Label Configurations

NeurIPS 2024poster

Learning with reduced labeling standards, such as noisy label, partial label, and supplementary unlabeled data, which we generically refer to as imprecise label, is a commonplace challenge in machine learning tasks. Previous methods tend to propose specific designs for every emerging imprecise label…

2024

Improving Continual Learning of Acoustic Scene Classification via Mutual Information Optimization

ICASSP 2024accepted

Continual learning, which aims to incrementally accumulate knowledge, has been an increasingly significant but challenging research topic for deep models that are prone to catastrophic forgetting. In this paper, we propose a novel replay-based continual learning approach in the context of class-incr…

Cited by 0SourceScholar
2024

In-Hand 3D Object Reconstruction from a Monocular RGB Video

AAAI 2024technical

Our work aims to reconstruct a 3D object that is held and rotated by a hand in front of a static RGB camera. Previous methods that use implicit neural representations to recover the geometry of a generic hand-held object from multi-view images achieved compelling results in the visible part of the o…

2024

In-Hand Following of Deformable Linear Objects Using Dexterous Fingers with Tactile Sensing

IROS 2024

Most research on deformable linear object (DLO) manipulation assumes rigid grasping. However, beyond rigid grasping and re-grasping, in-hand following is also an essential skill that humans use to dexterously manipulate DLOs, which requires continuously changing the grasp point by in-hand sliding wh

Cited by 13SourceScholar
2024

MMAD:Multi-modal Movie Audio Description

COLING 2024main

Audio Description (AD) aims to generate narrations of information that is not accessible through unimodal hearing in movies to aid the visually impaired in following film narratives. Current solutions rely heavily on manual work, resulting in high costs and limited scalability. While automatic metho…