← Search

Yu Zhang

368 accepted papers

2026

A Brain Graph Foundation Model: Pre-Training and Prompt-Tuning across Broad Atlases and Disorders

ICLR 2026poster

As large language models (LLMs) continue to revolutionize AI research, there is a growing interest in building large-scale brain foundation models to advance neuroscience. While most existing brain foundation models are pre-trained on time-series signals or connectome features, we propose a novel gr…

Cited by 0SourcecodeScholar
2026

ATEX-CF: Attack-Informed Counterfactual Explanations for Graph Neural Networks

ICLR 2026poster

Counterfactual explanations offer an intuitive way to interpret graph neural networks (GNNs) by identifying minimal changes that alter a model’s prediction, thereby answering “what must differ for a different outcome?”. In this work, we propose a novel framework, ATEX-CF that unifies adversarial att…

Cited by 0SourcecodeScholar
2026

Adaptive Visual Autoregressive Acceleration via Dual-Linkage Entropy Analysis

ICML 2026poster

Visual AutoRegressive modeling (VAR) suffers from substantial computational cost due to the massive token count involved. Failing to account for the continuous evolution of modeling dynamics, existing VAR token reduction methods face three key limitations: heuristic stage partition, non-adaptive sch…

Cited by 0SourceScholar
2026

An Open-Ended Benchmark and Formal Framework for Adjuvant Research with MLLM

ICLR 2026poster

Adjuvants play a critical role in modulating immune responses and are central to the development of vaccines and immunotherapies. Yet progress in this field is constrained by data scarcity and incomplete understanding of mechanisms of action, which limit the transition from experience-based design t…

Cited by 0SourceScholar
2026

Cross-Distill: Multi-Manifold and Viewpoint-Decoupled Distillation for Cross-View Geo-Localization

ICRA 2026poster

Abstract— Cross-View Geo-Localization (CVGL) localizes a query image via retrieval from georeferenced satellite imagery,yet severe viewpoint variation remains a central challenge. Recent advances often rely on heavy backbones or add-on modules that achieve high accuracy but are impractical on resour…

Cited by 0Scholar
2026

Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning

ICLR 2026poster

We aim to improve the reasoning capabilities of language models via reinforcement learning with verifiable rewards (RLVR). Recent RLVR post-trained models like DeepSeek-R1 have demonstrated reasoning abilities on mathematical and coding tasks. However, prior studies suggest that using RLVR alone to…

Cited by 0SourcecodeScholar
2026

DialogueVPR: Towards Conversational Visual Place Recognition

CVPR 2026

Inspired by how humans communicate spatial information, language-guided geo-localization has gained significant traction for its intuitive and practical value. Despite this progress, most methods still rely on a static, one-shot retrieval paradigm, which fails to handle the ambiguity and incompleten

Cited by 0SourcecodeScholar
2026

Discrete Latent Features Ablate Adversarial Attack: A Robust Prompt Tuning Framework for VLMs

ICLR 2026poster

While adversarial fine-tuning can enhance the robustness of vision-language models (VLMs), such approaches are computationally expensive. Adversarial prompt tuning has emerged as a practical alternative. However, existing methods are limited by their reliance on vulnerable continuous image features.…

Cited by 0SourceScholar
2026

Dynamic Momentum Recalibration in Online Gradient Learning

CVPR 2026

Stochastic Gradient Descent (SGD) and its momentum variants form the backbone of deep learning optimization, yet the underlying dynamics of their gradient behavior remain insufficiently understood. In this work, we reinterpret gradient updates through the lens of signal processing and reveal that fi

Cited by 0SourcecodeScholar
2026

DynamicGTR: Leveraging Graph Topology Representation Preferences to Boost VLM Capabilities on Graph QAs

CVPR 2026

Vision-Language Models (VLMs) have emerged as versatile solutions for zero-shot question answering (QA) across various domains. However, enabling VLMs to effectively comprehend structured graphs and perform accurate, efficient QA remains challenging. Existing approaches typically rely on a single ty

Cited by 5SourceScholar
2026

Efficient and Effective In-context Demonstration Selection with Coreset

AAAI 2026technical

In-context learning (ICL) has emerged as a powerful paradigm for Large Visual Language Models (LVLMs), enabling them to leverage a few examples directly from input contexts. However, the effectiveness of this approach is heavily reliant on the selection of demonstrations, a process that is NP-hard.

Cited by 0SourcePDFScholar
2026

Embracing Positional Bias in Multiple-Choice Question Answering via Permutation Equivariant Neural Networks

AAAI 2026technical

Several studies have demonstrated that large language models (LLMs) exhibit positional bias when answering multiple-choice questions (MCQs). Previous methods have identified such bias to be detrimental, leading to the development of techniques to mitigate it. However, we observe that certain permuta

Cited by 0SourcePDFScholar
2026

Evaluating and Steering Modality Preferences in Multi-modal LLMs

ICML 2026poster

Multi-modal large language models (MLLMs) have achieved remarkable success on complex multi-modal tasks. However, it remains insufficiently explored whether they exhibit \textit{modality preference}, a tendency to favor one modality over another when processing multi-modal contexts. To study this qu…

Cited by 0SourceScholar
2026

FPI-DET: A FACE–PHONE INTERACTION DATASET FOR PHONE-USE DETECTION AND UNDERSTANDING

ICASSP 2026poster

The widespread use of mobile devices has created new challenges for vision systems in safety monitoring, workplace productivity assessment, and attention management. Detecting whether a person is using a phone requires not only object recognition but also an understanding of behavioral context, whic…

Cited by 0SourcePDFScholar
2026

FairJudge : An Adaptive, Debiased, and Consistent LLM-as-a-Judge

ICML 2026poster

Existing LLM-as-a-Judge systems suffer from three fundamental limitations: \textbf{limited adaptivity} to task and domain-specific evaluation criteria, \textbf{systematic biases} driven by non-semantic cues such as position, length, format, and model provenance, and \textbf{evaluation inconsistency}…

Cited by 0SourceScholar
2026

Foresight Diffusion: Improving Sampling Consistency in Predictive Diffusion Models

ICLR 2026poster

Diffusion and flow-based models have enabled significant progress in generation tasks across various modalities and have recently found applications in predictive learning. However, unlike typical generation tasks that encourage sample diversity, predictive learning entails different sources of stoc…

Cited by 0SourceScholar
2026

GLAD: Bidirectional Structure-Attribute Alignment via Latent Graph Diffusion Models

ICML 2026poster

Learning on graphs with missing node attributes is a prevalent yet challenging problem in real-world scenarios, as graph neural networks (GNNs) typically rely on complete attribute information. Existing solutions often employ adversarial learning in a shared latent space to align graph structure and…

Cited by 0SourceScholar
2026

Gastric-X: A Multimodal Multi-Phase Benchmark Dataset for Advancing Vision-Language Models in Gastric Cancer Analysis

CVPR 2026

Recent vision-language models (VLMs) have shown strong generalization and multimodal reasoning abilities in natural domains. However, their application to medical diagnosis remains limited by the lack of comprehensive and structured datasets that capture real clinical workflows. To advance the devel

Cited by 0SourceScholar
2026

GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation

CVPR 2026

Driving World Models (DWMs) have been developing rapidly with the advances of generative models. However, existing DWMs lack 3D scene understanding capabilities and can only generate content conditioned on input data, without the ability to interpret or reason about the driving environment. Moreover

Cited by 0SourcecodeScholar
2026

Graph2Video: Leveraging Video Models to Model Dynamic Graph Evolution

AAAI 2026technical

Dynamic graphs are common in real‑world systems such as social media, recommender systems, and traffic networks. Existing dynamic graph models for link prediction often fall short in capturing the full complexity of temporal evolution. They tend to overlook fine‑grained variations in interaction or

Cited by 0SourcePDFScholar
2026

HAD: Heterogeneity-Aware Distillation for Lifelong Heterogeneous Learning

CVPR 2026

Lifelong learning aims to preserve knowledge acquired from previous tasks while incorporating knowledge from a sequence of new tasks. However, most prior work explores only streams of homogeneous tasks (*e.g.*, only classification tasks) and neglects the scenario of learning across heterogeneous tas

Cited by 0SourcecodeScholar
2026

In-The-Flow Agentic System Optimization for Effective Planning and Tool Use

ICLR 2026oral

Outcome-driven reinforcement learning has advanced reasoning in large language models (LLMs), but prevailing tool-augmented approaches train a single, monolithic policy that interleaves thoughts and tool calls under full context; this scales poorly with long horizons and diverse tools and generalize…

Cited by 0SourcecodeScholar
2026

Learning Coupled Continuous-Time Latent Dynamics from Irregular Events

ICML 2026spotlight

Modeling dynamic dependencies from irregularly sampled event sequences is a fundamental challenge in modern machine learning. In many real-world systems, individual-level states evolve continuously over time while being simultaneously influenced by population-level distributional dynamics. However, …

Cited by 0SourceScholar
2026

MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling

ICLR 2026poster

Training large language models with FP8 formats offers significant efficiency gains. However, the reduced numerical precision of FP8 poses challenges for stable and accurate training. Current frameworks preserve training performance using mixed-granularity quantization, i.e., applying per-group quan…

Cited by 0SourceScholar
2026

Markovian Scale Prediction: A New Era of Visual Autoregressive Generation

CVPR 2026

Visual AutoRegressive modeling (VAR) based on next-scale prediction has revitalized autoregressive visual generation. Although its full-context dependency, i.e., modeling all previous scales for next-scale prediction, facilitates more stable and comprehensive representation learning by leveraging co

Cited by 0SourceScholar
2026

NaRA: Noise-Aware LoRA for Parameter-Efficient Fine-Tuning of Diffusion LLMs

ICML 2026poster

Diffusion Large Language Models (dLLMs) have emerged as a promising non-autoregressive generative paradigm. Given the prohibitive computational cost of full fine-tuning, Parameter-Efficient Fine-Tuning (PEFT) has become the standard approach. However, existing PEFT methods (e.g., LoRA), originally t…

Cited by 0SourceScholar
2026

PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video Generation

CVPR 2026

Hand-object interaction (HOI) reconstruction and synthesis are becoming central to embodied AI and AR/VR. Yet, despite rapid progress, existing HOI generation research remains fragmented across three disjoint tracks: (1) pose-only synthesis that predicts MANO trajectories without producing pixels; (

Cited by 0SourcecodeScholar
2026

Post-Hoc Refinement for Multitask Symbolic Regression via Consensus-Accelerated Shapley Analysis

AAAI 2026technical

Multitask genetic programming (MTGP) is one of the primary methods for solving multitask symbolic regression (MTSR), the problem of discovering mathematical expressions for multiple interconnected tasks simultaneously. However, conventional MTGP approaches discard a wealth of valuable knowledge from

Cited by 0SourcePDFScholar
2026

Real-Time Coverage Path Planning for Bone-Aware Robotic Ultrasound Scanning

RA-L 2026

Robotic ultrasound systems offer significant potential for automated organ coverage scanning, yet acoustic shadows from bone structures can cause incomplete scans and missed diagnoses. This work presents a real-time coverage path planning method for bone-aware robotic ultrasound scanning. A voxel ba

Cited by 0SourceScholar
2026

Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning

ICLR 2026poster

Recent breakthroughs in reasoning language models have significantly advanced text-based reasoning. On the other hand, Multi-modal Large Language Models (MLLMs) still lag behind, hindered by their outdated internal LLMs. Upgrading these is often prohibitively expensive, as it requires complete visio…

Cited by 0SourcecodeScholar
2026

Revisiting Contrastive Learning in Collaborative Filtering via Parallel Graph Filters

AAAI 2026technical

Graph Contrastive Learning (GCL) has recently emerged as a powerful paradigm for modeling user–item interactions and learning high-quality representations in recommender systems. While existing GCL-based methods benefit from data augmentation and sampling strategies, they often overlook the inherent

Cited by 0SourcePDFScholar
2026

Robust Integrative Analysis of Multi-omics Datasets via Nuclear-norm Maximization

AAAI 2026technical

Spatially multimodal omics technologies provide unprecedented opportunities to address cellular heterogeneity within tissue contexts. However, learning robust and informative latent representations from such complex data remains a significant challenge. Existing graph-based methods often rely on sta

Cited by 0SourcePDFScholar
2026

S3Audio: Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer

ICML 2026poster

Real-time and accurate spatial audio generation is pivotal for delivering an immersive experience. However, existing spatial audio synthesis technologies are often encumbered by a tradeoff between generation quality and high inference latency, as well as difficulty in capturing precise spatial infor…

Cited by 0SourcecodeScholar
2026

SAFL-Geo: Structure-Aware Feature Learning with Fusion Loss for Infrared-Visible Geo-Localization

ICRA 2026poster

Cross-modal Visual Geo-localization often aims to retrieve a satellite visible-light image of the same geographic lo cation from a large-scale database using an infrared image cap tured by an unmanned aerial vehicle (UAV), thereby achieving precise localization. This capability is crucial for autono…

Cited by 0Scholar
2026

SAME: Spatial-Aware Multimodal Egocentric Human Pose Estimation

AAAI 2026technical

Egocentric human pose estimation (HPE) plays a crucial role in immersive applications such as virtual and augmented reality. However, existing methods relying on either visual or sparse inertial data alone often suffer from occlusion or ill-posed problems. In this work, we propose SAME, a novel spat

Cited by 0SourcePDFScholar
2026

SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance–Diversity Data Selection

ICML 2026poster

Fine-tuning large language models often undermines their safety alignment, a problem further amplified by harmful fine-tuning attacks in which adversarial data removes safeguards and induces unsafe behaviors. We propose SPARD, a defense framework that integrates Safety-Projected Alternating optimiza…

Cited by 0SourceScholar
2026

SafeNLIDB: A Privacy-Preserving Safety Alignment Framework for LLM-based Natural Language Database Interfaces

AAAI 2026technical

The rapid advancement of Large Language Models (LLMs) has driven significant progress in Natural Language Interface to Database (NLIDB). However, the widespread adoption of LLMs has raised critical privacy and security concerns. During interactions, LLMs may unintentionally expose confidential datab

Cited by 0SourcePDFScholar
2026

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark

CVPR 2026

Long video understanding (LVU) remains a core challenge in multimodal learning. Although recent vision-language models (VLMs) have made notable progress, existing benchmarks mainly focus on either fine-grained perception or coarse summarization, offering limited insight into temporal understanding o

Cited by 0SourceScholar
2026

Semi-Supervised Noise Adaptation: Transferring Knowledge from Noise Domain

ICML 2026poster

Transfer learning aims to facilitate the learning of a target domain by transferring knowledge from a source domain. The source domain typically contains semantically meaningful samples (*e.g.*, images) to facilitate effective knowledge transfer. However, a recent study observes that the noise domai…

Cited by 0SourceScholar
2026

Sketch2CAD: Generative Adversarial Network for Automated Conversion of Hand-Drawn Sketches to Parametric CAD Models

ICRA 2026poster

This paper addresses the labor-intensive process of converting imprecise hand-drawn sketches into precise, parametric CAD sketches. We present Sketch2CAD, a novel deep learning framework that leverages generative adversarial networks (GANs) to automate this conversion. Our approach consists of two m…

Cited by 0Scholar
2026

SocialNav: Training Human-Inspired Foundation Model for Socially-Aware Embodied Navigation

CVPR 2026

Embodied navigation that adheres to social norms remains an open research challenge. Our SocialNav is a foundational model for socially-aware navigation with a hierarchical "brain-action" architecture, capable of understanding high-level social norms and generating low-level, socially compliant traj

Cited by 0SourcecodeScholar
2026

StormInsight: Hierarchical Environmental Forcing and Vertical Coupling for Weather System Evolution

ICML 2026poster

Nowcasting forms the first line of defense against rapidly evolving weather hazards, where even minutes of delay can lead to severe societal impacts. However, existing systems predominantly extrapolate 2D radar reflectivity, which struggles under rapid intensification regimes. We introduce \N, a mul…

Cited by 0SourceScholar
2026

Swordsman: Entropy-Driven Adaptive Block Partition for Efficient Diffusion Language Models

ICML 2026poster

Block-wise decoding effectively improves the inference speed and quality in diffusion language models (DLMs) by combining inter-block sequential denoising and intra-block parallel unmasking. However, existing block-wise decoding methods typically partition blocks in a rigid and fixed manner, which i…

Cited by 0SourceScholar
2026

Toward Robust Multilingual Adaptation of LLMs for Low-Resource Languages

ICML 2026poster

Large language models (LLMs) continue to struggle with low-resource languages, primarily due to limited training data, translation noise, and unstable cross-lingual alignment. To address these challenges, we propose LiRA (Linguistic Robust Anchoring for LLMs)—a plug-and-play framework that requires …

Cited by 0SourceScholar
2026

When Transformers Meet Mamba: A Hybrid Transformer-Mamba Network for Video Object Detection

CVPR 2026

Video object detection has gained notable progress with the advent of transformers. While transformers excel at modeling long-range contextual dependencies, the quadratic complexity limits their efficiency in long-sequence processing. In contrast, Mamba offers greater efficiency in modeling long seq

Cited by 0SourceScholar
2026

XTransfer: Modality-Agnostic Few-Shot Model Transfer for Human Sensing at the Edge

ICML 2026poster

Deep learning for human sensing on edge systems presents significant potential for smart applications. However, its training and development are hindered by the limited availability of sensor data and resource constraints of edge systems. While transferring pre-trained models to different sensing ap…

Cited by 0SourceScholar
2025

$\text{D}_{2}\text{O}$: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models

ICLR 2025poster

Efficient generative inference in Large Language Models (LLMs) is impeded by the growing memory demands of Key-Value (KV) cache, especially for longer sequences. Traditional KV Cache eviction strategies, which discard less critical KV-pairs based on attention scores, often degrade generation quality…

Cited by 0SourcePDFScholar
2025

A Unified Taxonomy-Guided Instruction Tuning Framework for Entity Set Expansion and Taxonomy Expansion

ACL 2025finding

Entity set expansion, taxonomy expansion, and seed-guided taxonomy construction are three representative tasks that can be applied to automatically populate an existing taxonomy with emerging concepts. Previous studies view them as three separate tasks. Therefore, their proposed techniques usually w…

2025

Adaptive Wavelet-Positional Encoding for High-Frequency Information Learning in Implicit Neural Representation

AAAI 2025technical

Implicit Neural Representation (INR) has shown great potential in constructing the complex nature signal as a continuous implicit function. However, the representation results are incomplete since different components of the signal correspond to different frequencies and neural network inherently te…

Cited by 0SourcePDFScholar
2025

AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity

EMNLP 2025

Large Language Models (LLMs) with extended context lengths face significant computational challenges during the pre-filling phase, primarily due to the quadratic complexity of self-attention. Existing methods typically employ dynamic pattern matching and block-sparse low-level implementations. Howev

2025

BANER: Boundary-Aware LLMs for Few-Shot Named Entity Recognition

COLING 2025main

Despite the recent success of two-stage prototypical networks in few-shot named entity recognition (NER), challenges such as over/under-detected false spans in the span detection stage and unaligned entity prototypes in the type classification stage persist. Additionally, LLMs have not proven to be…

2025

CAFE-AD: Cross-Scenario Adaptive Feature Enhancement for Trajectory Planning in Autonomous Driving

ICRA 2025

Imitation learning based planning tasks on the nuPlan dataset have gained great interest due to their potential to generate human-like driving behaviors. However, open-loop training on the nuPlan dataset tends to cause causal confusion during closed-loop testing, and the dataset also presents a long

Cited by 2SourcecodeScholar
2025

CaDA: Cross-Problem Routing Solver with Constraint-Aware Dual-Attention

ICML 2025poster

Vehicle routing problems (VRPs) are significant combinatorial optimization problems (COPs) holding substantial practical importance. Recently, neural combinatorial optimization (NCO), which involves training deep learning models on extensive data to learn vehicle routing heuristics, has emerged as a…

2025

ChemActor: Enhancing Automated Extraction of Chemical Synthesis Actions with LLM-Generated Data

ACL 2025long

With the increasing interest in robotic synthesis in the context of organic chemistry, the automated extraction of chemical procedures from literature is critical. However, this task remains challenging due to the inherent ambiguity of chemical language and the high cost of human annotation required…

2025

ChromFound: Towards A Universal Foundation Model for Single-Cell Chromatin Accessibiltiy Data

NeurIPS 2025poster

The advent of single-cell Assay for Transposase-Accessible Chromatin using sequencing (scATAC-seq) offers an innovative perspective for deciphering regulatory mechanisms by assembling a vast repository of single-cell chromatin accessibility data. While foundation models have achieved significant suc…

Cited by 0SourcecodeScholar
2025

Cloud-Native Fog Robotics: Model-Based Deployment and Evaluation of Real-Time Applications

RA-L 2025

As the field of robotics evolves, robots become increasingly multi-functional and complex. Currently, there is a need for solutions that enhance flexibility and computational power without compromising real-time performance. The emergence of fog computing and cloud-native approaches addresses these

Cited by 7SourceScholar
2025

Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank Adaptation

ICML 2025poster

Low-rank adaptation (LoRA) has emerged as a leading parameter-efficient fine-tuning technique for adapting large foundation models, yet it often locks adapters into suboptimal minima near their initialization. This hampers model generalization and limits downstream operators such as adapter merging…

2025

Corrupted but Not Broken: Understanding and Mitigating the Negative Impacts of Corrupted Data in Visual Instruction Tuning

EMNLP 2025

Visual Instruction Tuning (VIT) aims to enhance Multimodal Large Language Models (MLLMs), yet its effectiveness is often compromised by corrupted datasets with issues such as hallucinated content, incorrect responses, and poor OCR quality. Previous approaches to address these challenges have focused

Cited by 0SourcePDFScholar
2025

CrossQG: Improving Difficulty-Controllable Question Generation through Consistency Enhancement

EMNLP 2025

Automatically generating questions with controlled difficulty has great application value, especially in the field of education. Although large language models are capable of generating questions of various difficulty levels, the generated questions often fail to align with the given target difficul

2025

DGO-VINS: A Visual-Inertial SLAM for Dynamic Environments With Geometric Constraint and Adaptive State Optimization

RA-L 2025

Traditional SLAM performs well in static environments, but experiences degeneration of localization accuracy and stability in dynamic settings. To enhance performance in dynamic environments, this letter presents DGO-VINS, a real-time dynamic visual-inertial SLAM system based on geometric constraint

Cited by 3SourceScholar
2025

DSDIR: A Two-Stage Method for Addressing Noisy Long-Tailed Problems in Malicious Traffic Detection

ICASSP 2025accepted

In recent years, deep learning based malicious traffic detection (MTD) systems have demonstrated remarkable success. However, their effectiveness tend to decrease because most malicious traffic datasets are suffered from noisy-labeled and long-tailed problems. While numerous approaches have been dev…

Cited by 0SourceScholar
2025

Design and Control of SeparaTrek: A Hybrid Aerial-Ground Robot with Separable and Combinative Locomotion Parts

IROS 2025

The hybrid aerial-ground robots combine ground mobility and aerial flight capability, which are often designed for executing multi-terrain tasks. However, most of the existing hybrid aerial-ground robots integrate the ground and aerial functionality into one single whole system. This leads to functi

Cited by 0SourceScholar
2025

Direction Informed Trees (DIT*): Optimal Path Planning via Direction Filter and Direction Cost Heuristic

ICRA 2025

Optimal path planning requires finding a series of feasible states from the starting point to the goal to optimize objectives. Popular path planning algorithms, such as Effort Informed Trees (EIT*), employ effort heuristics to guide the search. Effective heuristics are accurate and computationally e

Cited by 1SourceScholar
2025

Discarding the Crutches: Adaptive Parameter-Efficient Expert Meta-Learning for Continual Semantic Parsing

COLING 2025main

Continual Semantic Parsing (CSP) enables parsers to generate SQL from natural language questions in task streams, using minimal annotated data to handle dynamically evolving databases in real-world scenarios. Previous works often rely on replaying historical data, which poses privacy concerns. Recen…

Cited by 1SourcePDFScholar
2025

DocAssistant: Integrating Key-region Reading and Step-wise Reasoning for Robust Document Visual Question Answering

EMNLP 2025

Understanding the multimodal documents is essential for accurately extracting relevant evidence and using it for reasoning. Existing document understanding models struggle to focus on key information and tend to generate answers straightforwardly, ignoring evidence from source documents and lacking

Cited by 0SourcePDFScholar
2025

Dynamical Diffusion: Learning Temporal Dynamics with Diffusion Models

ICLR 2025poster

Diffusion models have emerged as powerful generative frameworks by progressively adding noise to data through a forward process and then reversing this process to generate realistic samples. While these models have achieved strong performance across various tasks and modalities, their application to…

2025

EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

CVPR 2025poster

GPT-4o, an omni-modal model that enables vocal conversations with diverse emotions and tones, marks a milestone for omni-modal foundation models. However, empowering Large Language Models to perceive and generate images, texts, and speeches end-to-end with publicly available data remains challenging…

Cited by 23SourcePDFScholar
2025

Enhancing Text-to-Image Diffusion Transformer via Split-Text Conditioning

NeurIPS 2025poster

Current text-to-image diffusion generation typically employs complete-text conditioning. Due to the intricate syntax, diffusion transformers (DiTs) inherently suffer from a comprehension defect of complete-text captions. One-fly complete-text input either overlooks critical semantic details or cause…

Cited by 0SourceScholar
2025

EnvPoser: Environment-aware Realistic Human Motion Estimation from Sparse Observations with Uncertainty Modeling

CVPR 2025poster

Estimating full-body motion using the tracking signals of head and hands from VR devices holds great potential for various applications. However, the sparsity and unique distribution of observations present a significant challenge, resulting in an ill-posed problem with multiple feasible solutions (…

2025

ExVideo: Extending Video Diffusion Models via Parameter-Efficient Post-Tuning

IJCAI 2025

Recently, advancements in video synthesis have attracted significant attention. Video synthesis models have demonstrated the practical applicability of diffusion models in creating dynamic visual content. Despite these advancements, the extension of video lengths remains constrained by computational

2025

Filling Memory Gaps: Enhancing Continual Semantic Parsing via SQL Syntax Variance-Guided LLMs Without Real Data Replay

AAAI 2025technical

Continual Semantic Parsing (CSP) aims to train parsers to convert natural language questions into SQL across tasks with limited annotated examples, adapting to dynamically updated databases in real-world scenarios. Previous studies mitigate this challenge by replaying historical data or employing pa…

2025

Focus on Local: Finding Reliable Discriminative Regions for Visual Place Recognition

AAAI 2025technical

Visual Place Recognition (VPR) is aimed at predicting the location of a query image by referencing a database of geotagged images. For VPR task, often fewer discriminative local regions in an image produce important effects while mundane background regions do not contribute or even cause perceptual…

2025

ForestCast: Open-Ended Event Forecasting with Semantic News Forest

EMNLP 2025

Open-ended event forecasting (OEEF) seeks to predict future events from a given context without being restricted to a predefined scope or format. It plays a crucial role in domains such as risk management and financial decision making. Although large language models show potential for OEEF, existing

2025

From Experts to a Generalist: Toward General Whole-Body Control for Humanoid Robots

NeurIPS 2025spotlight

Achieving general agile whole-body control on humanoid robots remains a major challenge due to diverse motion demands and data conflicts. While existing frameworks excel in training single motion-specific policies, they struggle to generalize across highly varied behaviors due to conflicting control…

Cited by 0SourceScholar
2025

GARD: A Geometry-Informed and Uncertainty-Aware Baseline Method for Zero-Shot Roadside Monocular Object Detection

RA-L 2025

Roadside camera-based perception methods are in high demand for developing efficient vehicle-infrastructure collaborative perception systems. By focusing on object-level depth prediction, we explore the potential benefits of integrating environmental priors into such systems and propose a geometry-b

Cited by 1SourceScholar
2025

Gassidy: Gaussian Splatting SLAM in Dynamic Environments

ICRA 2025

3D Gaussian Splatting (3DGS) allows flexible adjustments to scene representation, enabling continuous optimization of scene quality during dense visual simultaneous localization and mapping (SLAM) in static environments. However, 3DGS faces challenges in handling environmental disturbances from dyna

Cited by 14SourceScholar
2025

Gaussian Mixture Model for Graph Domain Adaptation

IJCAI 2025

Unsupervised domain adaptation (UDA) has been widely studied with the goal of transferring knowledge from a label-rich source domain to a related but unlabeled target domain. Most UDA techniques achieve this by reducing the feature discrepancies between the two domains to learn domain-invariant feat

Cited by 0SourcePDFScholar
2025

HeadMap: Locating and Enhancing Knowledge Circuits in LLMs

ICLR 2025poster

Large language models (LLMs), through pretraining on extensive corpora, encompass rich semantic knowledge and exhibit the potential for efficient adaptation to diverse downstream tasks. However, the intrinsic mechanisms underlying LLMs remain unexplored, limiting the efficacy of applying these model…

2025

HiRA: Parameter-Efficient Hadamard High-Rank Adaptation for Large Language Models

ICLR 2025oral

We propose Hadamard High-Rank Adaptation (HiRA), a parameter-efficient fine-tuning (PEFT) method that enhances the adaptability of Large Language Models (LLMs). While Low-rank Adaptation (LoRA) is widely used to reduce resource demands, its low-rank updates may limit its expressiveness for new tasks…

2025

HomoMatcher: Achieving Dense Feature Matching with Semi-Dense Efficiency by Homography Estimation

AAAI 2025technical

Feature matching between image pairs is a fundamental problem in computer vision that drives many applications, such as SLAM. Recently, semi-dense matching approaches have achieved substantial performance enhancements and established a widely-accepted coarse-to-fine paradigm. However, the majority…

Cited by 0SourcePDFScholar
2025

Image Watermarks are Removable using Controllable Regeneration from Clean Noise

ICLR 2025poster

Image watermark techniques provide an effective way to assert ownership, deter misuse, and trace content sources, which has become increasingly essential in the era of large generative models. A critical attribute of watermark techniques is their robustness against various manipulations. In this pap…

2025

Improving Efficiency of Answer Set Planning with Rough Solutions from Large Language Models for Robotic Task Planning

IJCAI 2025

Answer Set Programming (ASP) planning can be used to refine the rough solutions generated by Large Language Models (LLMs) to handle specific restrictions of actions, i.e., reconstruct the rough solutions to be executable, for robotic task planning. However, it is still challenging to efficiently sol

2025

InImageTrans: Multimodal LLM-based Text Image Machine Translation

ACL 2025finding

Multimodal large language models (MLLMs) have shown remarkable capabilities across various downstream tasks. However, when MLLMs are transferred to the text image machine translation (TiMT) task, preliminary experiments reveal that MLLMs suffer from serious repetition and omission hallucinations. To…

2025

Instantaneous Contact Localization on A Magnetically Transduced Tapered Whisker

IROS 2025

The whisker-inspired tactile sensor is advantageous for enhancing robotic perception in proximate range and darkness via non-intrusive contacts. However, localizing contact along the whisker shaft is challenging due to the non-injective mapping between tangential contacts and the resulting bending m

Cited by 0SourceScholar
2025

Inter-sentence Context Modeling and Structure-aware Representation Enhancement for Conversational Sentiment Quadruple Extraction

EMNLP 2025

Conversational aspect-based sentiment quadruple analysis (DiaASQ) is a newly-emergent task aiming to extract quadruples of target-aspect-opinion-sentiment from a conversation text. Existing studies struggle to capture complete dialogue semantics, largely due to inadequate inter-utterance modeling an

Cited by 0SourcePDFScholar
2025

LHMM: A Tightly-Coupled LiDAR-Inertial Hybrid-Map Matching Approach for Robust and Efficient Global Localization

IROS 2025

LiDAR map matching (LMM) faces two key challenges: the enormous number of point clouds imposes constraints on storage and computation, and traditional two-stage frameworks suffer from initial guess errors during degeneration. This paper presents LHMM, a hybrid-map framework that first compresses the

Cited by 0SourceScholar
2025

LI-GS: Gaussian Splatting With LiDAR Incorporated for Accurate Large-Scale Reconstruction

RA-L 2025

Large-scale 3D reconstruction is critical in the field of robotics, and the potential of 3D Gaussian Splatting (3DGS) for achieving accurate object-level reconstruction has been demonstrated. However, ensuring geometric accuracy in outdoor and unbounded scenes remains a significant challenge. This s

Cited by 32SourceScholar
2025

Learning Visuotactile Skills With Two Multifingered Hands

ICRA 2025

Aiming to replicate human-like dexterity, perceptual experiences, and motion patterns, we explore learning from human demonstrations using a bimanual system with multifingered hands and visuotactile data. Two significant challenges exist: the lack of an affordable and accessible teleoperation system

Cited by 119SourcecodeScholar
2025

MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations

NeurIPS 2025poster

Humans rely on multisensory integration to perceive spatial environments, where auditory cues enable sound source localization in three-dimensional space. Despite the critical role of spatial audio in immersive technologies such as VR/AR, most existing multimodal datasets provide only monaural audi…

Cited by 0SourcecodeScholar
2025

Mixture of insighTful Experts (MoTE): The Synergy of Reasoning Chains and Expert Mixtures in Self-Alignment

ACL 2025long

As the capabilities of large language models (LLMs) continue to expand, aligning these models with human values remains a significant challenge. Recent studies show that reasoning abilities contribute significantly to model safety, while integrating Mixture-of-Experts (MoE) architectures can further…

Cited by 0SourcePDFScholar
2025

MoPFormer: Motion-Primitive Transformer for Wearable-Sensor Activity Recognition

NeurIPS 2025poster

Human Activity Recognition (HAR) with wearable sensors is challenged by limited interpretability, which significantly impacts cross-dataset generalization. To address this challenge, we propose Motion-Primitive Transformer (MoPFormer), a novel self-supervised framework that enhances interpretability…

Cited by 0SourceScholar
2025

Multi-Label Ranking Loss Minimization for Matrix Completion

AAAI 2025technical

The common matrix completion methods minimize the rank of the matrix to be completed in addition to the Hamming loss between the incomplete and completed matrices. The rank of matrix measures the linear relation among the vectors of matrix, which may introduce ambiguity for data recovery. To cope wi…

2025

NaFV-Net: An Adversarial Four-view Network for Mammogram Classification

AAAI 2025technical

Breast cancer remains a leading cause of mortality among women, with millions of new cases diagnosed annually. Early detection through screening is crucial. Using neural networks to improve the accuracy of breast cancer screening has become increasingly important. In accordance with radiologists' pr…

2025

Neural-Lyapunov Fusion: Stable Dynamical System Learning for Robotic Motion Generation

IROS 2025

Point-to-point and periodic motions are ubiquitous in the world of robotics. To master these motions, Autonomous Dynamic System (ADS) based algorithms are fundamental in the domain of Learning from Demonstration (LfD). However, these algorithms face the significant challenge of balancing precision i

Cited by 0SourceScholar
2025

Object-level Correlation for Few-Shot Segmentation

ICCV 2025poster

Few-shot semantic segmentation (FSS) aims to segment objects of novel categories in the query images given only a few annotated support samples. Existing methods primarily build the image-level correlation between the support target object and the entire query image. However, this correlation contai…

Cited by 0SourcePDFScholar
2025

Open Your Eyes: Vision Enhances Message Passing Neural Networks in Link Prediction

ICML 2025poster

Message-passing graph neural networks (MPNNs) and structural features (SFs) are cornerstones for the link prediction task. However, as a common and intuitive mode of understanding, the potential of visual perception has been overlooked in the MPNN community. For the first time, we equip MPNNs with v…

Cited by 0SourcePDFScholar
2025

Pre-Training Graph Contrastive Masked Autoencoders are Strong Distillers for EEG

ICML 2025poster

Effectively utilizing extensive unlabeled high-density EEG data to improve performance in scenarios with limited labeled low-density EEG data presents a significant challenge. In this paper, we address this challenge by formulating it as a graph transfer learning and knowledge distillation problem.…

Cited by 0SourcePDFScholar
2025

Protein Large Language Models: A Comprehensive Survey

EMNLP 2025

Protein-specific large language models (ProteinLLMs) are revolutionizing protein science by enabling more efficient protein structure prediction, function annotation, and design. While existing surveys focus on specific aspects or applications, this work provides the first comprehensive overview of

2025

RISED: Accurate and Efficient RGB-Colorized Mapping Using Image Selection and Point Cloud Densification

ICRA 2025

Recent advances in robotics have underscored the critical role of colorized point clouds in enhancing environmental perception accuracy. However, conventional multisensor fusion Simultaneous Localization and Mapping (SLAM) systems typically employ all available images indiscriminately for point clou

Cited by 1SourceScholar
2025

SPE Attention: Making Attention Equivariant to Semantic-Preserving Permutation for Code Processing

EMNLP 2025

Codes serve as the fundamental language for human to communicate with machines, and various Transformer-based models are trained to process codes in recent advancements. A unique symmetry of code is its semantic-preserving permutation, which allows certain lines to be rearranged without altering the

Cited by 0SourcePDFScholar
2025

SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning

NeurIPS 2025poster

Multimodal large language models (MLLMs) have shown promising capabilities in reasoning tasks, yet still struggle significantly with complex problems requiring explicit self-reflection and self-correction, especially compared to their unimodal text-based counterparts. Existing reflection methods are…

Cited by 0SourceScholar
2025

STARS: A Unified Framework for Singing Transcription, Alignment, and Refined Style Annotation

ACL 2025finding

Recent breakthroughs in singing voice synthesis (SVS) have heightened the demand for high-quality annotated datasets, yet manual annotation remains prohibitively labor-intensive and resource-intensive. Existing automatic singing annotation (ASA) methods, however, primarily tackle isolated aspects of…

2025

Safety-Critical Control with Saliency Detection for Mobile Robots in Dynamic Multi-Obstacle Environments

ICRA 2025

This paper proposes a novel dual-filter architecture utilizing RGB-D camera data and dynamic control barrier functions (D-CBFs) for real-time obstacle avoidance in unstructured environments. The proposed method efficiently handles static, suddenly appearing, and dynamic obstacles, maintaining consis

Cited by 1SourceScholar
2025

Sharpness-Aware Black-Box Optimization

ICLR 2025poster

Black-box optimization algorithms have been widely used in various machine learning problems, including reinforcement learning and prompt fine-tuning. However, directly optimizing the training loss value, as commonly done in existing black-box optimization methods, could lead to suboptimal model qua…

Cited by 0SourcePDFScholar
2025

SpatialSplat: Efficient Semantic 3D from Sparse Unposed Images

ICCV 2025poster

A major breakthrough in 3D reconstruction is the feedforward paradigm to generate pixel-wise 3D points or Gaussian primitives from sparse, unposed images. To further incorporate semantics while avoiding the significant memory and storage costs of high-dimensional semantic features, existing methods…

Cited by 0SourcePDFScholar
2025

Strategic A/B testing via Maximum Probability-driven Two-armed Bandit

ICML 2025poster

Detecting a minor average treatment effect is a major challenge in large-scale applications, where even minimal improvements can have a significant economic impact. Traditional methods, reliant on normal distribution-based or expanded statistics, often fail to identify such minor effects because of…

Cited by 0SourcePDFScholar
2025

TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis

ACL 2025finding

Customizable multilingual zero-shot singing voice synthesis (SVS) has various potential applications in music composition and short video dubbing. However, existing SVS models overly depend on phoneme and note boundary annotations, limiting their robustness in zero-shot scenarios and producing poor…

2025

TechSinger: Technique Controllable Multilingual Singing Voice Synthesis via Flow Matching

AAAI 2025technical

Singing voice synthesis has made remarkable progress in generating natural and high-quality voices. However, existing methods rarely provide precise control over vocal techniques such as intensity, mixed voice, falsetto, bubble, and breathy tones, thus limiting the expressive potential of synthetic…

2025

Text-guided Multimodal Fusion for the Multimodal Emotion and Intent Joint Understanding

ICASSP 2025accepted

Emotion and Intent Joint Understanding in Multi-modal Conversation is a challenging task in the field of affective computing, aiming to decode the semantic information manifested in the multimodal conversational while simultaneously inferring the emotions and intents of the utterance. To address thi…

Cited by 0SourceScholar
2025

Versatile Framework for Song Generation with Prompt-based Control

EMNLP 2025

Song generation focuses on producing controllable high-quality songs based on various prompts. However, existing methods struggle to generate vocals and accompaniments with prompt-based control and proper alignment. Additionally, they fall short in supporting various tasks. To address these challeng

2024

A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery

EMNLP 2024main

In many scientific fields, large language models (LLMs) have revolutionized the way text and other modalities of data (e.g., molecules and proteins) are handled, achieving superior performance in various applications and augmenting the scientific discovery process. Nevertheless, previous surveys on…

2024

Adaptive Stochastic Gradient Algorithm for Black-box Multi-Objective Learning

ICLR 2024poster

Multi-objective optimization (MOO) has become an influential framework for various machine learning problems, including reinforcement learning and multi-task learning. In this paper, we study the black-box multi-objective optimization problem, where we aim to optimize multiple potentially conflictin…

Cited by 5SourcePDFScholar
2024

AutoFusion: Autonomous Visual Geolocation and Online Dense Reconstruction for UAV Cluster

ICRA 2024poster

Real-time dense reconstruction using Unmanned Aerial Vehicle (UAV) is becoming increasingly popular in large-scale rescue and environmental monitoring tasks. However, due to the energy constraints of a single UAV, the efficiency can be greatly improved through the collaboration of multi-UAVs. Nevert…

Cited by 0SourceScholar
2024

ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language Model

NeurIPS 2024poster

Visual object tracking aims to locate a targeted object in a video sequence based on an initial bounding box. Recently, Vision-Language~(VL) trackers have proposed to utilize additional natural language descriptions to enhance versatility in various applications. However, VL trackers are still infer…

Cited by 6SourcePDFScholar
2024

CogDPM: Diffusion Probabilistic Models via Cognitive Predictive Coding

ICML 2024poster

Predictive Coding (PC) is a theoretical framework in cognitive science suggesting that the human brain processes cognition through spatiotemporal prediction of visual world. Existing studies have developed spatiotemporal prediction neural networks based on the PC theroy, emulating its two core mecha…

Cited by 1SourcePDFScholar
2024

Design, Modeling and Adaptive Robust Control of a Spatial Symmetric Omni-Directional Aerial Vehicle

RA-L 2024

The increasing popularity of Unmanned Aerial Vehicles has led to heightened demand for their ability to interact with the environment along with the large flight envelopes. In this letter, we propose a new spatial symmetric fully actuated multirotor (FAM) that is capable of achieving fast, smooth at

Cited by 6SourceScholar
2024

Dissect Black Box: Interpreting for Rule-Based Explanations in Unsupervised Anomaly Detection

NeurIPS 2024poster

In high-stakes sectors such as network security, IoT security, accurately distinguishing between normal and anomalous data is critical due to the significant implications for operational success and safety in decision-making. The complexity is exacerbated by the presence of unlabeled data and the op…

Cited by 0SourcePDFScholar
2024

Dynamic Inertial Poser (DynaIP): Part-Based Motion Dynamics Learning for Enhanced Human Pose Estimation with Sparse Inertial Sensors

CVPR 2024poster

This paper introduces a novel human pose estimation approach using sparse inertial sensors addressing the shortcomings of previous methods reliant on synthetic data. It leverages a diverse array of real inertial motion capture data from different skeleton formats to improve motion diversity and mode…

2024

ERASOR++: Height Coding Plus Egocentric Ratio Based Dynamic Object Removal for Static Point Cloud Mapping

ICRA 2024poster

Mapping plays a crucial role in location and navigation within automatic systems. However, the presence of dynamic objects in 3D point cloud maps generated from scan sensors can introduce map distortion and long traces, thereby posing challenges for accurate mapping and navigation. To address this i…

Cited by 4SourceScholar
2024

Elliptical K-Nearest Neighbors - Path Optimization via Coulomb’s Law and Invalid Vertices in C-space Obstacles

IROS 2024poster

Path planning has long been an important and active research area in robotics. To address challenges in high-dimensional motion planning, this study introduces the Force Direction Informed Trees (FDIT*), a sampling-based planner designed to enhance speed and cost-effectiveness in pathfinding. FDIT*…

Cited by 1SourceScholar
2024

FARFusion: A Practical Roadside Radar-Camera Fusion System for Far-Range Perception

RA-L 2024

Far-range perception through roadside sensors is crucial to the effectiveness of intelligent transportation systems. The main challenge of far-range perception is due to the difficulty of performing accurate object detection and tracking under far distances <italic xmlns:mml="http://www.w3.org/1998/

Cited by 24SourceScholar
2024

Flexible Informed Trees (FIT*): Adaptive Batch-Size Approach in Informed Sampling-Based Path Planning

IROS 2024poster

In path planning, anytime almost-surely asymptotically optimal planners dominate the benchmark of sampling-based planners. A notable example is Batch Informed Trees (BIT*), where planners iteratively determine paths to batches of vertices within the exploration area. However, utilizing a consistent…

Cited by 6SourceScholar
2024

Forward-Backward Reasoning in Large Language Models for Mathematical Verification

ACL 2024findings

Self-Consistency samples diverse reasoning chains with answers and chooses the final answer by majority voting. It is based on forward reasoning and cannot further improve performance by sampling more reasoning chains when saturated. To further boost performance, we introduce backward reasoning to v…

Cited by 23SourcePDFScholar
2024

GITA: Graph to Visual and Textual Integration for Vision-Language Graph Reasoning

NeurIPS 2024poster

Large Language Models (LLMs) are increasingly used for various tasks with graph structures. Though LLMs can process graph information in a textual format, they overlook the rich vision modality, which is an intuitive way for humans to comprehend structural information and conduct general graph reaso…

2024

GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks

NeurIPS 2024spotlight

The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited diversity of languages and singers, absence of multi-technique information and real…

2024

Gated Slot Attention for Efficient Linear-Time Sequence Modeling

NeurIPS 2024poster

Linear attention Transformers and their gated variants, celebrated for enabling parallel training and efficient recurrent inference, still fall short in recall-intensive tasks compared to traditional Transformers and demand significant resources for training from scratch. This paper introduces Gated…

2024

Graph Chain-of-Thought: Augmenting Large Language Models by Reasoning on Graphs

ACL 2024findings

Large language models (LLMs), while exhibiting exceptional performance, suffer from hallucinations, especially on knowledge-intensive tasks. Existing works propose to augment LLMs with individual text units retrieved from external knowledge corpora to alleviate the issue. However, in many domains, t…

2024

Harmonic Retrieval for Non-Circular Coherent Signals via Double Decoupled Atomic Norm Minimization

ICASSP 2024accepted

This paper studies super-resolution harmonic retrieval for strictly non-circular coherent signals. We develop gridless sparse representations of both their covariance and pseudo-covariance matrices over a common matrix-form atom set. This enables the decoupled atomic norm minimization (D-ANM) techni…

Cited by 0SourceScholar
2024

HiPose: Hierarchical Binary Surface Encoding and Correspondence Pruning for RGB-D 6DoF Object Pose Estimation

CVPR 2024poster

In this work we present a novel dense-correspondence method for 6DoF object pose estimation from a single RGB-D image. While many existing data-driven methods achieve impressive performance they tend to be time-consuming due to their reliance on rendering-based refinement approaches. To circumvent t…

2024

IDGen: Item Discrimination Induced Prompt Generation for LLM Evaluation

NeurIPS 2024poster

As Large Language Models (LLMs) become more capable of handling increasingly complex tasks, the evaluation set must keep pace with these advancements to ensure it remains sufficiently discriminative. Item Discrimination (ID) theory, which is widely used in educational assessment, measures the abilit…

2024

IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers

ECCV 2024poster

"Generative training has been demonstrated to be powerful for building visual-language models. However, on zero-shot discriminative benchmarks, there is still a performance gap between models trained with generative and discriminative objectives. In this paper, we aim to narrow this gap by improving…

Cited by 3SourcePDFScholar
2024

Knowledge-aware Attention Network for Medication Effectiveness Prediction

COLING 2024main

The first 24 hours’ medication plan is critical to patients with serious or life-threatening illnesses and injuries. An appropriate medication can result in a lower mortality, a shorter length stay and a higher APACHE score. However, in clinical practice, the medication plan is often error-prone, es…

Cited by 0SourcePDFScholar
2024

LEFormer: A Hybrid CNN-Transformer Architecture for Accurate Lake Extraction from Remote Sensing Imagery

ICASSP 2024accepted

Lake extraction from remote sensing images is challenging due to the complex lake shapes and inherent data noises. Existing methods suffer from blurred segmentation boundaries and poor foreground modeling. This paper proposes a hybrid CNN-Transformer architecture, called LEFormer, for accurate lake…

Cited by 0SourceScholar
2024

Learning a Stable Dynamic System with a Lyapunov Energy Function for Demonstratives Using Neural Networks

ICRA 2024poster

Autonomous Dynamic System (DS)-based algorithms hold a pivotal and foundational role in the field of Learning from Demonstration (LfD). Nevertheless, they confront the formidable challenge of striking a delicate balance between achieving precision in learning and ensuring the overall stability of th…

Cited by 2SourceScholar
2024

MLIP: Efficient Multi-Perspective Language-Image Pretraining with Exhaustive Data Utilization

ICML 2024poster

Contrastive Language-Image Pretraining (CLIP) has achieved remarkable success, leading to rapid advancements in multimodal studies. However, CLIP faces a notable challenge in terms of *inefficient data utilization*. It relies on a single contrastive supervision for each image-text pair during repres…

Cited by 3SourcePDFScholar
2024

MTMamba: Enhancing Multi-Task Dense Scene Understanding by Mamba-Based Decoders

ECCV 2024poster

"Multi-task dense scene understanding, which learns a model for multiple dense prediction tasks, has a wide range of application scenarios. Modeling long-range dependency and enhancing cross-task interactions are crucial to multi-task dense prediction. In this paper, we propose MTMamba, a novel Mamb…

2024

MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

ICLR 2024spotlight

Large language models (LLMs) have pushed the limits of natural language understanding and exhibited excellent problem-solving ability. Despite the great success, most existing open-source LLMs (\eg, LLaMA-2) are still far away from satisfactory for solving mathematical problems due to the complex re…

2024

Multi-Task Interactive Robot Fleet Learning with Visual World Models

CoRL 2024poster

Recent advancements in large-scale multi-task robot learning offer the potential for deploying robot fleets in household and industrial settings, enabling them to perform diverse tasks across various environments. However, AI-enabled robots often face challenges with generalization and robustness wh…

Cited by 4SourcecodeScholar
2024

Multilingual and Fully Non-Autoregressive ASR with Large Language Model Fusion: A Comprehensive Study

ICASSP 2024accepted

In the era of large models, the autoregressive nature of decoding often results in latency serving as a significant bottleneck. We propose a non-autoregressive LM-fused ASR system that effectively leverages the parallelization capabilities of accelerator hardware. Our approach combines the Universal…

Cited by 19SourceScholar
2024

Multimodal Modeling for Spoken Language Identification

ICASSP 2024accepted

Spoken language identification refers to the task of automatically predicting the spoken language in a given utterance. Conventionally, it is modeled as a speech-based language identification task. Prior techniques have been constrained to a single modality; however in the case of video data there i…

Cited by 0SourceScholar
2024

NC-SDF: Enhancing Indoor Scene Reconstruction Using Neural SDFs with View-Dependent Normal Compensation

CVPR 2024poster

State-of-the-art neural implicit surface representations have achieved impressive results in indoor scene reconstruction by incorporating monocular geometric priors as additional supervision. However we have observed that multi-view inconsistency between such priors poses a challenge for high-qualit…

Cited by 2SourcePDFScholar
2024

Nemesis: Normalizing the Soft-prompt Vectors of Vision-Language Models

ICLR 2024spotlight

With the prevalence of large-scale pretrained vision-language models (VLMs), such as CLIP, soft-prompt tuning has become a popular method for adapting these models to various downstream tasks. However, few works delve into the inherent properties of learnable soft-prompt vectors, specifically the im…

2024

NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video Reconstruction

NeurIPS 2024oral

Reconstruction of static visual stimuli from non-invasion brain activity fMRI achieves great success, owning to advanced deep learning models such as CLIP and Stable Diffusion. However, the research on fMRI-to-video reconstruction remains limited since decoding the spatiotemporal perception of conti…

2024

Online Efficient Safety-Critical Control for Mobile Robots in Unknown Dynamic Multi-Obstacle Environments

IROS 2024poster

This paper proposes a LiDAR-based goal-seeking and exploration framework, addressing the efficiency of online obstacle avoidance in unstructured environments populated with static and moving obstacles. This framework addresses two significant challenges associated with traditional dynamic control ba…

Cited by 5SourceScholar
2024

Parallelizing Linear Transformers with the Delta Rule over Sequence Length

NeurIPS 2024poster

Transformers with linear attention (i.e., linear transformers) and state-space models have recently been suggested as a viable linear-time alternative to transformers with softmax attention. However, these models still underperform transformers especially on tasks that require in-context retrieval.…

Cited by 51SourcePDFScholar
2024

Personalized Federated Learning for Cross-City Traffic Prediction

IJCAI 2024poster

Traffic prediction plays an important role in urban computing. However, many cities face data scarcity due to low levels of urban development. Although many approaches transfer knowledge from data-rich cities to data-scarce cities, the centralized training paradigm cannot uphold data privacy. For th…

2024

Planning First, Question Second: An LLM-Guided Method for Controllable Question Generation

ACL 2024findings

In the field of education, for better assessment of students’ abilities, generated questions often need to meet experts’ requirements, indicating the need for controllable question generation (CQG). However, current CQG methods mainly focus on difficulty control, neglecting the control of question c…

2024

Profiling Power Consumption in Low-Speed Autonomous Guided Vehicles

RA-L 2024

The increasing demand for automation has led to a rise in the use of low-speed Autonomous guided vehicles (AGVs). However, AGVs rely on batteries for their power source, which limits their operational time and affects their overall performance. To optimize their energy usage and enhance their batter

Cited by 7SourceScholar
2024

Question-guided Knowledge Graph Re-scoring and Injection for Knowledge Graph Question Answering

EMNLP 2024finding

Knowledge graph question answering (KGQA) involves answering natural language questions by leveraging structured information stored in a knowledge graph. Typically, KGQA initially retrieve a targeted subgraph from a large-scale knowledge graph, which serves as the basis for reasoning models to addre…

2024

RGBD-based Image Goal Navigation with Pose Drift: A Topo-metric Graph based Approach

ICRA 2024poster

Image-goal navigation in unknown environments with sensor error is of considerable difficulty for autonomous robots. In this paper, we propose a drift-resisting topo-metric graph to map the environment and localize the robot using only relative poses. The error-sharing mechanism under this represent…

Cited by 1SourceScholar
2024

Real-Time Adaptive Safety-Critical Control with Gaussian Processes in High-Order Uncertain Models

ICRA 2024poster

This paper presents an adaptive online learning framework for systems with uncertain parameters to ensure safety-critical control in non-stationary environments. Our approach consists of two phases. The initial phase is centered on a novel sparse Gaussian process (GP) framework. We first integrate a…

Cited by 4SourceScholar
2024

Rethinking Guidance Information to Utilize Unlabeled Samples: A Label Encoding Perspective

ICML 2024poster

Empirical Risk Minimization (ERM) is fragile in scenarios with insufficient labeled samples. A vanilla extension of ERM to unlabeled samples is Entropy Minimization (EntMin), which employs the soft-labels of unlabeled samples to guide their learning. However, EntMin emphasizes prediction discriminab…

2024

Robust Singing Voice Transcription Serves Synthesis

ACL 2024long

Note-level Automatic Singing Voice Transcription (AST) converts singing recordings into note sequences, facilitating the automatic annotation of singing datasets for Singing Voice Synthesis (SVS) applications. Current AST methods, however, struggle with accuracy and robustness when used for practica…

2024

RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language Models

NeurIPS 2024poster

Recent works show that assembling multiple off-the-shelf large language models (LLMs) can harness their complementary abilities. To achieve this, routing is a promising method, which learns a router to select the most suitable LLM for each query. However, existing routing models are ineffective when…

2024

SDAC: A Multimodal Synthetic Dataset for Anomaly and Corner Case Detection in Autonomous Driving

AAAI 2024technical

Nowadays, closed-set perception methods for autonomous driving perform well on datasets containing normal scenes. However, they still struggle to handle anomalies in the real world, such as unknown objects that have never been seen while training. The lack of public datasets to evaluate the model pe…

Cited by 3SourcePDFScholar
2024

SecureSQL: Evaluating Data Leakage of Large Language Models as Natural Language Interfaces to Databases

EMNLP 2024finding

With the widespread application of Large Language Models (LLMs) in Natural Language Interfaces to Databases (NLIDBs), concerns about security issues in NLIDBs have been increasing gradually. However, research on sensitive data leakage in NLIDBs is relatively limited. Therefore, we propose a benchmar…

Cited by 4SourcePDFScholar
2024

Seed-Guided Fine-Grained Entity Typing in Science and Engineering Domains

AAAI 2024technical

Accurately typing entity mentions from text segments is a fundamental task for various natural language processing applications. Many previous approaches rely on massive human-annotated data to perform entity typing. Nevertheless, collecting such data in highly specialized science and engineering do…

2024

Selective Prompting Tuning for Personalized Conversations with LLMs

ACL 2024findings

In conversational AI, personalizing dialogues with persona profiles and contextual understanding is essential. Despite large language models’ (LLMs) improved response coherence, effective persona integration remains a challenge. In this work, we first study two common approaches for personalizing LL…

2024

Self-Sensing Feedback Control of an Electrohydraulic Robotic Shoulder

ICRA 2024poster

The human shoulder, with its glenohumeral joint, tendons, ligaments, and muscles, allows for the execution of complex tasks with precision and efficiency. However, current robotic shoulder designs lack the compliance and compactness inherent in their biological counterparts. A major limitation of th…

Cited by 1SourceScholar
2024

SplattingAvatar: Realistic Real-Time Human Avatars with Mesh-Embedded Gaussian Splatting

CVPR 2024poster

We present SplattingAvatar a hybrid 3D representation of photorealistic human avatars with Gaussian Splatting embedded on a triangle mesh which renders over 300 FPS on a modern GPU and 30 FPS on a mobile device. We disentangle the motion and appearance of a virtual human with explicit mesh geometry…

2024

StyleSinger: Style Transfer for Out-of-Domain Singing Voice Synthesis

AAAI 2024technical

Style transfer for out-of-domain (OOD) singing voice synthesis (SVS) focuses on generating high-quality singing voices with unseen styles (such as timbre, emotion, pronunciation, and articulation skills) derived from reference singing voice samples. However, the endeavor to model the intricate nuanc…

2024

TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control

EMNLP 2024main

Zero-shot singing voice synthesis (SVS) with style transfer and style control aims to generate high-quality singing voices with unseen timbres and styles (including singing method, emotion, rhythm, technique, and pronunciation) from audio and text prompts. However, the multifaceted nature of singing…

2024

Time-Varying LoRA: Towards Effective Cross-Domain Fine-Tuning of Diffusion Models

NeurIPS 2024poster

Large-scale diffusion models are adept at generating high-fidelity images and facilitating image editing and interpolation. However, they have limitations when tasked with generating images in dynamic, evolving domains. In this paper, we introduce Terra, a novel Time-varying low-rank adapter that of…

2024

USM-SCD: Multilingual Speaker Change Detection Based on Large Pretrained Foundation Models

ICASSP 2024accepted

We introduce a multilingual speaker change detection model (USM-SCD) that can simultaneously detect speaker turns and perform ASR for 96 languages. This model is adapted from a speech foundation model trained on a large quantity of supervised and unsupervised data, demonstrating the utility of fine-…

Cited by 0SourceScholar
2023

A Quantum Kernel Learning Approach to Acoustic Modeling for Spoken Command Recognition

ICASSP 2023accepted

We propose a quantum kernel learning (QKL) framework to address the inherent data sparsity issues often encountered in training large-scare acoustic models in low-resource scenarios. We project acoustic features based on classical-to-quantum feature encoding. Different from existing quantum convolut…

Cited by 11SourceScholar
2023

Accelerating RNN-T Training and Inference Using CTC Guidance

ICASSP 2023accepted

We propose a novel method to accelerate training and inference process of recurrent neural network transducer (RNN-T) based on the guidance from a co-trained connectionist temporal classification (CTC) model. We made a key assumption that if an encoder embedding frame is classified as a blank frame…

Cited by 0SourceScholar
2023

Bi-LRFusion: Bi-Directional LiDAR-Radar Fusion for 3D Dynamic Object Detection

CVPR 2023poster

LiDAR and Radar are two complementary sensing approaches in that LiDAR specializes in capturing an object's 3D shape while Radar provides longer detection ranges as well as velocity hints. Though seemingly natural, how to efficiently combine them for improved feature representation is still unclear.…

2023

CAENet: Using Collaborative Attention Transformer and Add-Boost Strategy for Single Image Deraining

ICASSP 2023accepted

In recent years, the Convolutional Neural Network (CNN) based deraining methods have achieved remarkable results. However, these methods rarely used the long-range context information, and thus could not effectively restore the regions damaged by dense rain streaks. Moreover, the rain streaks in an…

Cited by 0SourceScholar
2023

Chain-of-Skills: A Configurable Model for Open-Domain Question Answering

ACL 2023long

The retrieval model is an indispensable component for real-world knowledge-intensive tasks, e.g., open-domain question answering (ODQA). As separate retrieval skills are annotated for different datasets, recent work focuses on customized methods, limiting the model transfer- ability and scalability.…

Cited by 27SourcePDFScholar
2023

CluB: Cluster Meets BEV for LiDAR-Based 3D Object Detection

NeurIPS 2023poster

Currently, LiDAR-based 3D detectors are broadly categorized into two groups, namely, BEV-based detectors and cluster-based detectors. BEV-based detectors capture the contextual information from the Bird's Eye View (BEV) and fill their center voxels via feature diffusion with a stack of convolution l…

Cited by 6SourcePDFScholar
2023

Comparison of Soft and Hard Target RNN-T Distillation for Large-Scale ASR

ICASSP 2023accepted

Knowledge distillation is an effective machine learning technique to transfer knowledge from a teacher model to a smaller student model, especially with unlabeled data. In this paper, we focus on knowledge distillation for the RNN-T model, which is widely used in state-of-the-art (SoTA) automatic sp…

Cited by 0SourceScholar
2023

Conditional Adapters: Parameter-efficient Transfer Learning with Fast Inference

NeurIPS 2023poster

We propose Conditional Adapter (CoDA), a parameter-efficient transfer learning method that also improves inference efficiency. CoDA generalizes beyond standard adapter approaches to enable a new way of balancing speed and accuracy using conditional computation. Starting with an existing dense pretra…

Cited by 63SourcePDFScholar
2023

Continuous-Time LiDAR-Inertial-Vehicle Odometry Method with Lateral Acceleration Constraint

ICRA 2023poster

In this paper, we propose a continuous-time-based LiDAR-inertial-vehicle odometry method, which can tightly fuse the data from Light Detection And Ranging (LiDAR), inertial measurement units (IMU), and vehicle measurements. The lateral acceleration constraint is further added to trajectory estimatio…

Cited by 3SourceScholar
2023

Denoising Pre-training for Machine Translation Quality Estimation with Curriculum Learning

AAAI 2023technical

Quality estimation (QE) aims to assess the quality of machine translations when reference translations are unavailable. QE plays a crucial role in many real-world applications of machine translation. Because labeled QE data are usually limited in scale, recent research, such as DirectQE, pre-trains…

2023

Edgeformers: Graph-Empowered Transformers for Representation Learning on Textual-Edge Networks

ICLR 2023poster

Edges in many real-world social/information networks are associated with rich text information (e.g., user-user communications or user-product reviews). However, mainstream network representation learning models focus on propagating and aggregating node attributes, lacking specific designs to utiliz…

2023

Effective Structured Prompting by Meta-Learning and Representative Verbalizer

ICML 2023poster

Prompt tuning for pre-trained masked language models (MLM) has shown promising performance in natural language processing tasks with few labeled examples. It tunes a prompt for the downstream task, and a verbalizer is used to bridge the predicted token and label prediction. Due to the limited traini…

2023

Efficient Domain Adaptation for Speech Foundation Models

ICASSP 2023accepted

Foundation models (FMs), that are trained on broad data at scale and are adaptable to a wide range of downstream tasks, have brought large interest in the research community. Benefiting from the diverse data sources such as different modalities, languages and application domains, foundation models h…

Cited by 0SourceScholar