← Search

Ming Li

205 accepted papers

2026

4DPC$^2$hat: Towards Dynamic Point Cloud Understanding with Failure-Aware Bootstrapping

ICML 2026poster

Point clouds provide a compact and expressive representation of 3D objects, and have recently been integrated into multimodal large language models (MLLMs). However, existing methods primarily focus on static objects, while understanding dynamic point cloud sequences remains largely unexplored. This…

Cited by 0SourceScholar
2026

A Unified Spectral-Spatial Framework for GNNs: Balancing Over-Smoothing and Over-Squashing

IJCAI 2026

Over-smoothing (OSM) and over-squashing (OSQ) are two fundamental phenomena that limit the performance of Graph Neural Networks (GNNs), yet a unified spectral-spatial understanding of these phenomena remains underexplored. In this paper, we adopt polynomial spectral filters as an analytical tool to

Cited by 0Scholar
2026

ADC-GNN: Adaptive Dual-level Collaborative Graph Neural Networks for Graph Classification

IJCAI 2026

Most existing Graph Neural Networks (GNNs) rely on the node-level message passing or attention mechanisms to propagate and extract useful information. Although recent advances attempt to move beyond purely the node-level propagation by constructing high-level representations, these approaches are of

Cited by 0Scholar
2026

AISHELL6-WHISPER: A CHINESE MANDARIN AUDIO-VISUAL WHISPER SPEECH DATASET WITH SPEECH RECOGNITION BASELINES

ICASSP 2026poster

Whisper speech recognition is crucial not only for ensuring privacy in sensitive communications but also for providing a critical communication bridge for patients under vocal restraint and enabling discrete interaction in noise-sensitive environments. The development of Chinese mandarin audio-visua…

Cited by 0SourcePDFScholar
2026

ARBench: Algorithmic Reasoner or API Alchemist? Evaluating LLMs Beyond API Calls

AAAI 2026technical

Large Language Models (LLMs) have demonstrated impressive capabilities in code generation. Like human programmers, LLMs tend to call high-level APIs and libraries to program efficiently. However, this shortcut may hinder LLMs from learning the essential algorithm reasoning, leading instead to rote m

Cited by 0SourcePDFScholar
2026

Accessible, Realistic, and Fair Evaluation of Positive-Unlabeled Learning Algorithms

ICLR 2026poster

Positive-unlabeled (PU) learning is a weakly supervised binary classification problem, in which the goal is to learn a binary classifier from only positive and unlabeled data, without access to negative data. In recent years, many PU learning algorithms have been developed to improve model performan…

Cited by 0SourceScholar
2026

Amplifying Discrepancies: Exploiting Macro and Micro Inconsistencies for Image Manipulation Localization

AAAI 2026technical

The rapid development of image manipulation technologies poses significant challenges to multimedia forensics, especially in accurate localization of manipulated regions. Existing methods often fail to fully explore the intrinsic discrepancies between manipulated and authentic regions, resulting in

Cited by 0SourcePDFScholar
2026

AnomalyPainter: Vision-Language-Diffusion Synergy for Realistic and Diverse Unseen Industrial Anomaly Synthesis

AAAI 2026technical

Visual anomaly detection is limited by the lack of sufficient anomaly data. While existing anomaly synthesis methods have made remarkable progress, achieving both realism and diversity in synthesis remains a major obstacle. To address this, we propose AnomalyPainter, a novel framework that breaks th

Cited by 0SourcePDFScholar
2026

CP-FREEZER: Latency Attacks Against Vehicular Cooperative Perception

AAAI 2026technical

Cooperative perception (CP) enhances situational awareness of connected and autonomous vehicles by exchanging and combining messages from multiple agents. While prior work has explored adversarial integrity attacks that degrade detection accuracy, little is known about CP

Cited by 0SourcePDFScholar
2026

Can We Build Scene Graphs, Not Classify Them? FlowSG: Progressive Image-Conditioned Scene Graph Generation with Flow Matching

CVPR 2026

Scene Graph Generation (SGG) unifies object localization and visual relationship reasoning by predicting boxes and subject-predicate-object triples. Yet most pipelines treat SGG as a one-shot, deterministic classification instead of a genuine progressive, generative task. We propose FlowSG, which re

Cited by 0SourceScholar
2026

CompSpoof: A Dataset and Joint Learning Framework for Component-Level Audio Anti-spoofing Countermeasures

ICASSP 2026poster

Component-level audio Spoofing (Comp-Spoof) targets a new form of audio manipulation where only specific components of a signal, such as speech or environmental sound, are forged or substituted while other components remain genuine. Existing anti-spoofing datasets and methods treat an utterance or a…

Cited by 0SourcePDFScholar
2026

Cross-Modal Attention Calibration for LVLM Hallucination Mitigation

CVPR 2026

Large vision-language models (LVLMs) have shown remarkable capabilities in visual-language understanding. Despite their success, LVLMs still suffer from generating hallucinations in complex generation tasks, leading to inconsistencies between visual inputs and generated content. To address this issu

Cited by 0SourceScholar
2026

CyC3D: Fine-grained Controllable 3D Generation via Cycle Consistency Regularization

AAAI 2026technical

Despite the remarkable progress of 3D generation, achieving controllability, i.e., ensuring consistency between generated 3D content and input conditions like edge and depth, remains a significant challenge. Existing methods often struggle to maintain accurate alignment, leading to noticeable discre

Cited by 0SourcePDFScholar
2026

D-GVIO: A Buffer-Driven and Efficient Decentralized GNSS-Visual-Inertial State Estimator for Multi-Agent Systems

ICRA 2026poster

Cooperative localization is essential for swarm applications like collaborative exploration and search-and-rescue missions. However, maintaining real-time capability, robustness, and computational efficiency on resource-constrained platforms presents significant challenges. To address these challeng…

2026

DataCube: A Video Retrieval Platform via Natural Language Semantic Profiling

IJCAI 2026

Large-scale video repositories are increasingly available for modern video understanding and generation tasks. However, transforming raw videos into high-quality, task-specific datasets remains costly and inefficient. We present DataCube, an intelligent platform for automatic video processing, multi

Cited by 0Scholar
2026

DepthMesh: A Dual-End Complementary Online Depth Estimation and Mesh Reconstruction

ICRA 2026poster

We present a novel dual-end complementary method for online depth estimation and mesh reconstruction, termed DepthMesh. Unlike most existing state-of-the-art methods that produce either only depth online or surface mesh offline, our method tightly couples online multiview depth estimation and Trunca…

Cited by 0Scholar
2026

Dynamic-Static Synergistic Selection Method for Candidate Code Solutions with Generated Test Cases

AAAI 2026technical

Large language models (LLMs) show significant improvement in code generation. A common practice is sampling multiple candidate codes to increase the likelihood of producing an accurate solution. However, effectively identifying the best candidate from the pool is a significant challenge. Although ex

Cited by 0SourcePDFScholar
2026

EventFlash: Towards Efficient MLLMs for Event-Based Vision

ICLR 2026poster

Event-based multimodal large language models (MLLMs) enable robust perception in high-speed and low-light scenarios, addressing key limitations of frame-based MLLMs. However, current event-based MLLMs often rely on dense image-like processing paradigms, overlooking the spatiotemporal sparsity of eve…

Cited by 0SourcecodeScholar
2026

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning

ICML 2026poster

Theory of Mind (ToM) is a must-acquire skill for modern foundation model systems to operate effectively and safely in the real world. Recent works have explored honing ToM via post-training; however, we show that such progress is confounded by a pervasive “shortcut” issue: tasks can reach up to 99% …

Cited by 0SourceScholar
2026

GI-GCN: Global Interacted Graph Convolutional Networks via Dominant Sets for Graph Classification

ICML 2026poster

Graph Convolutional Networks (GCNs) are defined based on aggregating the node information of adjacent nodes, that are usually treated as equally important as each other, limiting the representational power of existing GCNs for graph classification. To address this shortcoming, we propose a novel Glo…

Cited by 0SourceScholar
2026

H$^2$CL: Heterogeneity-Aware Hypergraph Contrastive Learning for Robust Representation

ICML 2026poster

In recent years, hypergraph contrastive learning methods have gained widespread attention due to their excellent performance in processing high-order structural data. However, traditional hypergraph learning method often assume that neighboring nodes are homogeneous, which can lead to the mixing of …

Cited by 0SourceScholar
2026

Harnessing Multiple Large Language Models: A Survey on LLM Ensemble

IJCAI 2026

LLM Ensemble---which involves the comprehensive use of multiple large language models (LLMs), each aimed at handling user queries during downstream inference, to benefit from their individual strengths---has gained substantial attention recently. The widespread availability of LLMs, coupled with the

Cited by 0Scholar
2026

Heterophily-aware Contrastive Learning for Heterophilic Hypergraphs

AAAI 2026technical

Hypergraph neural networks (HNNs) have emerged as powerful tools for modeling high-order relationships in complex systems. However, most existing HNNs are designed under the assumption of homophily, which does not hold in many real-world scenarios where connected nodes often exhibit diverse semantic

Cited by 0SourcePDFScholar
2026

High-Pass Matters: Theoretical Insights and Sheaflet-Based Design for Hypergraph Neural Networks

AAAI 2026technical

Hypergraph neural networks (HGNNs) have shown great potential in modeling higher-order relationships among multiple entities. However, most existing HGNNs primarily emphasize low-pass filtering while neglecting the role of high-frequency information. In this work, we present a theoretical investigat

Cited by 0SourcePDFScholar
2026

HyperAim: Hypergraph Contrastive Learning with Adaptive Multi-frequency Filters

AAAI 2026technical

Unsupervised hypergraph representation learning has recently gained traction for its ability to model complex high-order interactions without requiring labeled data. However, existing contrastive learning methods typically overlook the frequency diversity inherent in hypergraph signals. To address t

Cited by 0SourcePDFScholar
2026

HyperGOOD: Towards Out-of-Distribution Detection in Hypergraphs

AAAI 2026technical

Out-of-distribution (OOD) detection plays a critical role in ensuring the robustness of machine learning models in open-world settings. While extensive efforts have been made in vision, language, and graph domains, the challenge of OOD detection in hypergraph-structured data remains unexplored. In t

Cited by 0SourcePDFScholar
2026

HyperGait: Unleashing the Power of Parsing for Gait Recognition in the Wild via Hypergraph

CVPR 2026

In recent years, the gait parsing sequence has become increasingly popular due to its higher information entropy than the binary silhouette and the keypoint-based skeleton. However, existing parsing-based gait recognition methods have not fully explored the complex, non-linear relationships between

Cited by 0SourceScholar
2026

HyperNoRA: Hyperedge Prediction via Node-Level Relation-Aware Self-Supervised Hypergraph Learning

AAAI 2026technical

Hyperedge prediction plays a critical role in high-order relational modeling with hypergraphs, yet most existing methods primarily focus on sampling strategies or local aggregation within candidate hyperedges. These approaches often overlook global structural dependencies that are essential for lear

Cited by 0SourcePDFScholar
2026

Learning Brain Representation with Hierarchical Visual Embeddings

ICLR 2026poster

Decoding visual representations from brain signals has attracted significant attention in both neuroscience and artificial intelligence. However, the degree to which brain signals truly encode visual information remains unclear. Current visual decoding approaches explore various brain–image alignmen…

Cited by 0SourceScholar
2026

Learning from Noisy Preferences: A Semi-Supervised Learning Approach to Direct Preference Optimization

ICLR 2026poster

Human visual preferences are inherently multi-dimensional, encompassing aspects of aesthetics, detail fidelity, and semantic alignment. However, existing open-source preference datasets provide only single, holistic annotations, resulting in severe label noise—images that excel in some dimensions (e…

Cited by 0SourceScholar
2026

Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language Models

ICML 2026poster

Current research on video hallucination mitigation primarily focuses on isolated error types, leaving *compositional* hallucinations—arising from incorrect reasoning over multiple interacting spatial and temporal factors largely underexplored. We introduce **OmniVCHall**, a benchmark designed to sys…

Cited by 0SourceScholar
2026

MOES-Pred: Molecular Structural Representation Learning by Adaptive Energy-Sentinel Vibration for Generalized Property Prediction

ICML 2026poster

Molecular property prediction from 3D structures is fundamentally constrained by the scarcity of labeled data. To address this challenge, researchers have adapted various self-supervised pre-training methods from computer vision and natural language processing; however, these approaches often neglec…

Cited by 0SourceScholar
2026

Mind Your Margin and Boundary: Are Your Distilled Datasets Truly Robust?

ICML 2026oral

Dataset distillation (DD) compresses a large training set into a small synthetic set for efficient training, but most DD methods optimize only clean accuracy and leave robustness uncontrolled. Recent robust DD methods improve robustness, yet they often suffer from a poor accuracy–robustness trade-of…

Cited by 0SourceScholar
2026

Modeling Item-Level Dynamic Variability with Residual Diffusion for Bundle Recommendation

AAAI 2026technical

Existing solutions for bundle recommendation (BR) have achieved remarkable effectiveness for predicting the user’s preference for prebuilt bundles. However, bundle-item (B-I) affiliation will vary dynamically in real scenarios. For ex ample, a bundle themed as ‘casual outfit’ may add ‘hat’ or re

Cited by 0SourcePDFScholar
2026

Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following

CVPR 2026

Large multimodal models (LMMs) are increasingly adopted as judges in multimodal evaluation systems due to their strong instruction following and consistency with human preferences. However, their ability to follow diverse, fine-grained evaluation criteria remains underexplored. We develop Multi-Crit

Cited by 0SourcecodeScholar
2026

Multi-Granular Graph Learning with Fine-Grained Behavioral Pattern Awareness for Session-Based Recommendation

AAAI 2026technical

Session-based recommendation aims to predict users’ next actions by modeling their ongoing interaction sequences, particularly in scenarios where long-term user profiles are unavailable. While existing methods have achieved promising results by leveraging sequential and graph-based structures, they

Cited by 0SourcePDFScholar
2026

NI-Tex: Non-isometric Image-based Garment Texture Generation

CVPR 2026

Existing industrial 3D garment meshes already cover most real-world clothing geometries, yet their texture diversity remains limited. To acquire more realistic textures, generative methods are often used to extract Physically-based Rendering (PBR) textures and materials from large collections of wil

Cited by 1SourcecodeScholar
2026

OddGridBench: Exposing the Lack of Fine-Grained Visual Discrepancy Sensitivity in Multimodal Large Language Models

CVPR 2026

Multimodal large language models (MLLMs) have achieved remarkable performance across a wide range of vision-language tasks. However, their ability in low-level visual perception, particularly in detecting fine-grained visual discrepancies, remains underexplored and lacks systematic analysis.In this

Cited by 0SourcecodeScholar
2026

On the Predictive Power of Representation Dispersion in Language Models

ICLR 2026poster

We show that a language model’s ability to predict text is tightly linked to the breadth of its embedding space: models that spread their contextual representations more widely tend to achieve lower perplexity. Concretely, we find that representation dispersion—the average pairwise cosine distance a…

Cited by 0SourcecodeScholar
2026

Optimization and Robustness-Informed Membership Inference Attacks for LLMs

AAAI 2026technical

The proliferation of Large Language Models (LLMs) has raised concerns over training data privacy. Membership Inference Attacks (MIA), aiming to identify whether specific data was used for training, pose significant privacy risks. However, existing MIA methods struggle to address the scale and comple

Cited by 0SourcePDFScholar
2026

Permutation Equivariant Framelet-based Hypergraph Neural Networks

AAAI 2026technical

Hypergraphs provide a natural and expressive framework for modeling high-order relationships, enabling the representation of group-wise interactions beyond pairwise connections. While hypergraph neural networks (HNNs) have shown promise for learning on such structures, existing models often rely on

Cited by 0SourcePDFScholar
2026

PyVision-RL: Forging Open Agentic Vision Models via RL

ICML 2026poster

Reinforcement learning for agentic multimodal models often suffers from interaction collapse, where models learn to reduce tool usage and multi-turn reasoning, limiting the benefits of agentic behavior. We introduce PyVision-RL, a reinforcement learning framework for open-weight multimodal models th…

Cited by 0SourceScholar
2026

Random Selection Reveals Implicit Knowledge Consensus in Code Generation

ICML 2026poster

Training large language models for code generation requires selecting high-quality data from solution pools where each problem admits multiple correct implementations. Conventional studies on data selection hold that sophisticated strategies that employ various optimization objectives, such as diver…

Cited by 0SourceScholar
2026

ReWeaver: Towards Simulation-Ready and Topology-Accurate Garment Reconstruction

CVPR 2026

High-quality 3D garment reconstruction plays a crucial role in mitigating the sim-to-real gap in applications such as digital avatars, virtual try-on and robotic manipulation. However, existing garment reconstruction methods typically rely on unstructured representations, such as 3D Gaussian Splats,

Cited by 0SourcecodeScholar
2026

Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry

ICLR 2026poster

Large language models (LLMs) are widely used as reference-free evaluators via prompting, but this “LLM-as-a-Judge” paradigm is costly, opaque, and sensitive to prompt design. In this work, we investigate whether smaller models can serve as efficient evaluators by leveraging internal representations…

Cited by 0SourcecodeScholar
2026

Risk Awareness Injection: Calibrating Vision-Language Models for Safety without Compromising Utility

ICML 2026poster

Vision language models (VLMs) extend the reasoning capabilities of large language models (LLMs) to cross-modal settings, yet remain highly vulnerable to multimodal jailbreak attacks. Existing defenses predominantly rely on safety fine-tuning or \textit{aggressive} token manipulations, incurring subs…

Cited by 0SourceScholar
2026

SFCLTA: Spectral Fusion Contrastive Learning with Topology-Adaptive Graph Augmentation

ICML 2026poster

Graph Neural Networks (GNNs) have achieved remarkable successes in graph analysis due to the Message-Passing (MP) mechanism, yet they struggle with heterophilic graphs where connected nodes often have distinct labels or dissimilar attributes. Graph Contrastive Learning (GCL) serves as a promising ap…

Cited by 0SourceScholar
2026

SSHPool: The Separated Subgraph-based Hierarchical Pooling

AAAI 2026technical

In this paper, we develop a novel local graph pooling method, namely the Separated Subgraph-based Hierarchical Pooling (SSHPool), for graph classification. We commence by assigning the nodes of a sample graph into different clusters, resulting in a family of separated subgraphs. We individually empl

Cited by 0SourcePDFScholar
2026

SciEducator: Scientific Video Understanding and Educating via Deming-Cycle Multi-Agent System

CVPR 2026

Recent advancements in multimodal large language models (MLLMs) and video agent systems have significantly improved general video understanding. However, when applied to scientific video understanding and educating--a domain that demands external professional knowledge integration and rigorous step-

Cited by 0SourceScholar
2026

Self-Forcing++: Towards Minute-Scale High-Quality Video Generation

ICLR 2026poster

Diffusion models have revolutionized image and video generation, achieving unprecedented visual quality. However, their reliance on transformer architectures incurs prohibitively high computational costs, particularly when extending generation to long videos. Recent work has explored autoregressive…

Cited by 0SourcecodeScholar
2026

Self-Supervised Hypergraph Learning with Substructure Awareness for Hyperedge Prediction

AAAI 2026technical

Hyperedge prediction plays a central role in hypergraph learning, enabling the inference of high-order relations among multiple entities. However, existing methods often rely on a simplistic flat set assumption, treating candidate hyperedges as unstructured collections of nodes and neglecting their

Cited by 0SourcePDFScholar
2026

StyleTailor: Towards Personalized Fashion Styling via Hierarchical Negative Feedback

AAAI 2026technical

The advancement of intelligent agents has revolutionized problem-solving across diverse domains, yet solutions for personalized fashion styling remain underexplored, which holds immense promise for promoting shopping experiences. In this work, we present StyleTailor, the first collaborative agent fr

Cited by 0SourcePDFScholar
2026

Swift-SVD: Theoretical Optimality Meets Practical Efficiency in Low-Rank LLM Compression

ICML 2026poster

The deployment of Large Language Models is constrained by the memory and bandwidth demands of static weights and dynamic Key-Value cache. SVD-based compression provides a hardware-friendly solution to reduce these costs. However, existing methods suffer from two key limitations: some are suboptimal …

Cited by 0SourceScholar
2026

Towards Hierarchy–Uniformity Equilibrium: Recovering Semantic Depth in Hypergraph Contrastive Learning

ICML 2026oral

Hypergraph contrastive learning is an effective paradigm for representation learning on higher-order relational data, yet existing methods largely ignore that hyperedges link nodes with multi-level semantics. Standard contrastive objectives emphasize instance discrimination via hyperspherical unifor…

Cited by 0SourceScholar
2026

UniF$^2$ace: A $\underline{Uni}$fied $\underline{F}$ine-grained $\underline{Face}$ Understanding and Generation Model

ICLR 2026poster

Unified multimodal models (UMMs) have emerged as a powerful paradigm in fundamental cross-modality research, demonstrating significant potential in both image understanding and generation. However, existing research in the face domain primarily faces two challenges: **(1) fragmentation development**…

Cited by 0SourcecodeScholar
2026

Unified Latent Space for Understanding and Generation via Semantic Auto-encoder

CVPR 2026

Latent generative modeling has emerged as the dominant paradigm for Diffusion Transformers (DiT), where a pretrained autoencoder compresses image pixels into a latent space to facilitate the diffusion process. Recently, the use of semantic encoders within autoencoders (AEs) has gained attention, yet

Cited by 0SourceScholar
2026

XSpecMesh: Quality-Preserving Auto-Regressive Mesh Generation Acceleration via Multi-Head Speculative Decoding

ICML 2026poster

Current auto-regressive models can generate high-quality, topologically precise meshes; however, they necessitate thousands—or even tens of thousands—of next-token predictions during inference, resulting in substantial latency. We introduce XSpecMesh, a quality-preserving acceleration method for aut…

Cited by 0SourceScholar
2026

Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion

CVPR 2026

Realistic 3D city generation is fundamental to a wide range of applications, including virtual reality and digital twins. However, most existing methods rely on training a single diffusion model, which limits their ability to generate personalized and boundless city-scale scenes. In this paper, we p

Cited by 0SourceScholar
2025

ADFormer: Aggregation Differential Transformer for Passenger Demand Forecasting

IJCAI 2025

Passenger demand forecasting helps optimize vehicle scheduling, thereby improving urban efficiency. Recently, attention-based methods have been used to adequately capture the dynamic nature of spatio-temporal data. However, existing methods that rely on heuristic masking strategies cannot fully adap

2025

AKBR: Learning Adaptive Kernel-based Representations for Graph Classification

IJCAI 2025

In this paper, we propose a new model to learn Adaptive Kernel-based Representations (AKBR) for graph classification. Unlike state-of-the-art R-convolution graph kernels that are defined by merely counting any pair of isomorphic substructures between graphs and cannot provide an end-to-end learning

2025

ATLAS: Agent Tuning via Learning Critical Steps

ACL 2025finding

Large Language Model (LLM) agents have demonstrated remarkable generalization capabilities across multi-domain tasks. Existing agent tuning approaches typically employ supervised finetuning on entire expert trajectories. However, behavior-cloning of full trajectories can introduce expert bias and we…

Cited by 0SourcePDFScholar
2025

All Roads Lead to Rome: Exploring Edge Distribution Shifts for Heterophilic Graph Learning

IJCAI 2025

Heterophilic graph neural networks (GNNs) have gained prominence for their ability to learn effective representations in graphs with diverse, attribute-aware relationships. While existing methods leverage attribute inference during message passing to improve performance, they often struggle with cha

Cited by 0SourcePDFScholar
2025

An End-to-End Simple Clustering Hierarchical Pooling Operation for Graph Learning Based on Top-K Node Selection

IJCAI 2025

Graph Neural Networks (GNNs) are powerful tools for graph learning, but one of the important challenges is how to effectively extract representations for graph-level tasks. In this paper, we propose an end-to-end Simple Clustering Hierarchical Pooling (SCHPool) operation, which is based on Top-K nod

2025

CPO: Condition Preference Optimization for Controllable Image Generation

NeurIPS 2025poster

To enhance controllability in text-to-image generation, ControlNet introduces image-based control signals, while ControlNet++ improves pixel-level cycle consistency between generated images and the input control signal. To avoid the prohibitive cost of back-propagating through the sampling process,…

Cited by 0SourcecodeScholar
2025

Codec-ASV: Exploring Neural Audio Codec For Speaker Representation Learning

ICASSP 2025accepted

Discrete speech representations have gained significant success in a variety of speech-related tasks. Among these, Neural Audio Codec (NAC), which serves as a compressed form of audio signals, have proven effective in speech AIGC applications. Moreover, we believe that the speaker information can be…

Cited by 0SourceScholar
2025

ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness

NeurIPS 2025poster

Color plays an important role in human perception and usually provides critical clues in visual reasoning. However, it is unclear whether and how vision-language models (VLMs) can perceive, understand, and leverage color as humans. This paper introduces ColorBench, an innovative benchmark meticulous…

Cited by 0SourcecodeScholar
2025

CrossAD: Time Series Anomaly Detection with Cross-scale Associations and Cross-window Modeling

NeurIPS 2025poster

Time series anomaly detection plays a crucial role in a wide range of real-world applications. Given that time series data can exhibit different patterns at different sampling granularities, multi-scale modeling has proven beneficial for uncovering latent anomaly patterns that may not be apparent at…

Cited by 0SourceScholar
2025

DEP-SLAM: A Dynamic Environment Perception SLAM System with Large Language Models

ICASSP 2025accepted

Inderscience is a global company, a dynamic leading independent journal publisher disseminates the latest research across the broad fields of science, engineering and technology; management, public and business administration; environment, ecological economics and sustainable development; computing,…

Cited by 0SourceScholar
2025

DHAKR: Learning Deep Hierarchical Attention-Based Kernelized Representations for Graph Classification

AAAI 2025technical

Graph-based representations are powerful tools for analyzing structured data. In this paper, we propose a novel model to learn Deep Hierarchical Attention-based Kernelized Representations (DHAKR) for graph classification. To this end, we commence by learning an assignment matrix to hierarchically ma…

Cited by 0SourcePDFScholar
2025

DHTAGK: Deep Hierarchical Transitive-Aligned Graph Kernels for Graph Classification

IJCAI 2025

In this paper, we propose a family of novel Deep Hierarchical Transitive-Aligned Graph Kernels (DHTAGK) for graph classification. To this end, we commence by developing a new Hierarchical Aligned Graph Auto-Encoder (HA-GAE) to construct transitive-aligned embedding graphs that encapsulate the struct

2025

DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data

EMNLP 2025

Large Language Models (LLMs) are increasingly aligned with human preferences through Reinforcement Learning from Human Feedback (RLHF). Among RLHF methods, Group Relative Policy Optimization (GRPO) has gained attention for its simplicity and strong performance, notably eliminating the need for a lea

Cited by 0SourcePDFScholar
2025

Deep Hypergraph Neural Networks with Tight Framelets

AAAI 2025technical

Hypergraphs provide a flexible framework for modeling high-order (complex) interactions among multiple entities, extending beyond traditional pairwise correlations in graph structures. However, deep hypergraph neural networks (HGNNs) often face the challenge of oversmoothing with increasing depth, s…

Cited by 1SourcePDFScholar
2025

EEE-Bench: A Comprehensive Multimodal Electrical And Electronics Engineering Benchmark

CVPR 2025poster

Recent studies on large language models (LLMs) and large multimodal models (LMMs) have demonstrated promising skills in various domains including science and mathematics. However, their capability in more challenging and real-world related scenarios like engineering has not been systematically studi…

Cited by 2SourcePDFScholar
2025

EFTViT: Efficient Federated Training of Vision Transformers with Masked Images on Resource-Constrained Clients

ICCV 2025poster

Federated learning research has recently shifted from Convolutional Neural Networks (CNNs) to Vision Transformers (ViTs) due to their superior capacity. ViTs training demands higher computational resources due to the lack of 2D inductive biases inherent in CNNs. However, efficient federated training…

Cited by 0SourcePDFScholar
2025

ENAHPool: The Edge-Node Attention-based Hierarchical Pooling for Graph Neural Networks

ICML 2025poster

Graph Neural Networks (GNNs) have emerged as powerful tools for graph learning, and one key challenge arising in GNNs is the development of effective pooling operations for learning meaningful graph representations. In this paper, we propose a novel Edge-Node Attention-based Hierarchical Pooling (EN…

Cited by 0SourcePDFScholar
2025

EduLLM: Leveraging Large Language Models and Framelet-Based Signed Hypergraph Neural Networks for Student Performance Prediction

ICML 2025poster

The growing demand for personalized learning underscores the importance of accurately predicting students' future performance to support tailored education and optimize instructional strategies. Traditional approaches predominantly focus on temporal modeling using historical response records and lea…

Cited by 0SourcePDFScholar
2025

EventGPT: Event Stream Understanding with Multimodal Large Language Models

CVPR 2025poster

Event cameras capture visual information as asynchronous pixel change streams, excelling in challenging lighting and high-dynamic scenarios. Existing multimodal large language models (MLLMs) concentrate on natural RGB images, failing in scenarios where event data fits better. In this paper, we intro…

Cited by 3SourcePDFScholar
2025

Explicit Spatial Hint and Implicit Logits Relation: Distilling Heterogeneous Knowledge From Vision Transformer to CNN

ICASSP 2025accepted

A lightweight Convolutional Neural Network (CNN) typically requires knowledge transfer from a large powerful network before it is employed in resource-limited edge devices. Vision Transformer (ViT) possesses an unparalleled capability for global modeling but remains largely unexplored in Knowledge D…

Cited by 0SourceScholar
2025

Exploring the Over-smoothing Problem of Graph Neural Networks for Graph Classification: An Entropy-based Viewpoint

IJCAI 2025

The over-smoothing has emerged as a major challenge in the development of Graph Neural Networks (GNNs). While existing state-of-the-art methods effectively mitigate the diminishing distance between nodes and improve the performance of node classification, they tend to be elusive for graph-level task

2025

FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic Manipulation

ICCV 2025poster

Vision-Language-Action (VLA) models have significantly advanced robotic manipulation by enabling robots to interpret language instructions for task execution. However, training these models often relies on large-scale user-specific data, raising concerns about privacy and security, which in turn lim…

Cited by 0SourcePDFScholar
2025

HA-SCN: Learning Hierarchical Aligned Subtree Convolutional Networks for Graph Classification

IJCAI 2025

In this paper, we propose a Hierarchical Aligned Subtree Convolutional Network (HA-SCN) for graph classification. Our idea is to transform graphs of arbitrary sizes into fixed-sized aligned graphs and construct a normalized K-layer m-ary subtree for each node in the aligned graphs. By sliding convol

2025

HiFi-Portrait: Zero-shot Identity-preserved Portrait Generation with High-fidelity Multi-face Fusion

CVPR 2025poster

Recent advancements in diffusion-based technologies have made significant strides, particularly in identity-preserved portrait generation (IPG). However, when using multiple reference images from the same ID, existing methods typically produce lower-fidelity portraits and struggle to customize face…

Cited by 0SourcePDFScholar
2025

HyperMixup: Hypergraph-Augmented with Higher-order Information Mixup

NeurIPS 2025poster

Hypergraphs offer a natural paradigm for modeling complex systems with multi-way interactions. Hypergraph neural networks (HGNNs) have demonstrated remarkable success in learning from such higher-order relational data. While such higher-order modeling enhances relational reasoning, the effectiveness…

Cited by 0SourceScholar
2025

HyperNear: Unnoticeable Node Injection Attacks on Hypergraph Neural Networks

ICML 2025poster

With the growing adoption of Hypergraph Neural Networks (HNNs) to model higher-order relationships in complex data, concerns about their security and robustness have become increasingly important. However, current security research often overlooks the unique structural characteristics of hypergraph…

2025

Inter3D: A Benchmark and Strong Baseline for Human-Interactive 3D Object Reconstruction

IJCAI 2025

Recent advancements in implicit 3D reconstruction methods, e.g., neural rendering fields and Gaussian splatting, have primarily focused on novel view synthesis of static or dynamic objects with continuous motion states. However, these approaches struggle to efficiently model a human-interactive obje

2025

L3A: Label-Augmented Analytic Adaptation for Multi-Label Class Incremental Learning

ICML 2025poster

Class-incremental learning (CIL) enables models to learn new classes continually without forgetting previously acquired knowledge. Multi-label CIL (MLCIL) extends CIL to a real-world scenario where each sample may belong to multiple classes, introducing several challenges: label absence, which leads…

2025

MATCH: Modality-Calibrated Hypergraph Fusion Network for Conversational Emotion Recognition

IJCAI 2025

Multimodal emotion recognition aims to identify emotions by integrating multimodal features derived from spoken utterances. However, existing work often neglects the calibration of conversational entities, focusing mainly on extracting potential intra- or cross-modal information. This leads to the u

Cited by 0SourcePDFScholar
2025

MDP3: A Training-free Approach for List-wise Frame Selection in Video-LLMs

ICCV 2025poster

Video large language models (Video-LLMs) have made significant progress in understanding videos. However, processing multiple frames leads to lengthy visual token sequences, presenting challenges such as the limited context length cannot accommodate the entire video, and the inclusion of irrelevant…

2025

ML-GOOD: Towards Multi-Label Graph Out-Of-Distribution Detection

AAAI 2025technical

The out-of-distribution (OOD) detection on graph-structured data is crucial for deploying graph neural networks securely in open-world scenarios. However, existing methods have overlooked the prevalent scenario of multi-label classification in real-world applications. In this work, we investigate th…

2025

Mosaic-IT: Cost-Free Compositional Data Synthesis for Instruction Tuning

ACL 2025finding

Finetuning large language models with a variety of instruction-response pairs has enhanced their capability to understand and follow instructions. Current instruction tuning primarily relies on teacher models or human intervention to generate and refine the instructions and responses for training, w…

2025

Multi-Reward as Condition for Instruction-based Image Editing

ICLR 2025poster

High-quality training triplets (instruction, original image, edited image) are essential for instruction-based image editing. Predominant training datasets (e.g., InsPix2Pix) are created using text-to-image generative models (e.g., Stable Diffusion, DALL-E) which are not trained for image editing. A…

2025

MultiNet: Adaptive Multi-Viewed Subgraph Convolutional Networks for Graph Classification

NeurIPS 2025poster

The problem of over-smoothing has emerged as a fundamental issue for Graph Convolutional Networks (GCNs). While existing efforts primarily focus on enhancing the discriminability of node representations for node classification, they tend to overlook the over-smoothing at the graph level, significant…

Cited by 0SourceScholar
2025

PVChat: Personalized Video Chat with One-Shot Learning

ICCV 2025poster

Video large language models (ViLLMs) excel in general video understanding, e.g., recognizing activities like talking and eating, but struggle with identity-aware comprehension, such as "Wilson is receiving chemotherapy" or "Tom is discussing with Sarah", limiting their applicability in smart healthc…

Cited by 0SourcePDFScholar
2025

ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges

ICCV 2025poster

As multi-modal large language models (MLLMs) frequently exhibit errors when solving scientific problems, evaluating the validity of their reasoning processes is critical for ensuring reliability and uncovering fine-grained model weaknesses. Since human evaluation is laborious and costly, prompting M…

2025

ResoFilter: Fine-grained Synthetic Data Filtering for Large Language Models through Data-Parameter Resonance Analysis

NAACL 2025findings

Large language models (LLMs) have shown remarkable effectiveness across various domains, with data augmentation methods utilizing GPT for synthetic data generation becoming prevalent. However, the quality and utility of augmented data remain questionable, and current methods lack clear metrics for e…

2025

Revisiting Chain-of-Thought in Code Generation: Do Language Models Need to Learn Reasoning before Coding?

ICML 2025poster

Large Language Models (LLMs) have demonstrated exceptional performance in code generation, becoming increasingly vital for software engineering and development. Recently, Chain-of-Thought (CoT) has proven effective for complex tasks by prompting LLMs to reason step-by-step and provide a final answer…

Cited by 0SourcePDFScholar
2025

RuleR: Improving LLM Controllability by Rule-based Data Recycling

NAACL 2025short

Large language models (LLMs) still lack delicate controllability over their responses, which is critical to enhancing their performance and the user experience. However, curating supervised fine-tuning (SFT) datasets to improve LLM controllability usually relies on human experts or proprietary LLMs,…

2025

SPADE: Spatial-Aware Denoising Network for Open-vocabulary Panoptic Scene Graph Generation with Long- and Local-range Context Reasoning

ICCV 2025poster

Panoptic Scene Graph Generation (PSG) integrates instance segmentation with relation understanding to capture pixel-level structural relationships in complex scenes. Although recent approaches leveraging pre-trained vision-language models (VLMs) have significantly improved performance in the open-vo…

Cited by 0SourcePDFScholar
2025

Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking

NeurIPS 2025poster

The explosive growth of generative video models has amplified the demand for reliable copyright preservation of AI-generated content. Despite its popularity in image synthesis, invisible generative watermarking remains largely underexplored in video generation. To address this gap, we propose Safe-S…

Cited by 0SourceScholar
2025

Sekai: A Video Dataset towards World Exploration

NeurIPS 2025poster

Video generation techniques have made remarkable progress, promising to be the foundation of interactive world exploration. However, existing video generation datasets are not well-suited for world exploration training as they suffer from some limitations: limited locations, short duration, static s…

Cited by 0SourceScholar
2025

Self-Enhanced Reasoning Training: Activating Latent Reasoning in Small Models for Enhanced Reasoning Distillation

ICASSP 2025accepted

The rapid advancement of large language models (LLMs) has significantly enhanced their reasoning abilities, enabling increasingly complex tasks. However, these capabilities often diminish in smaller, more computationally efficient models like GPT-2. Recent research shows that reasoning distillation…

Cited by 0SourceScholar
2025

Semantic Shift Estimation via Dual-Projection and Classifier Reconstruction for Exemplar-Free Class-Incremental Learning

ICML 2025poster

Exemplar-Free Class-Incremental Learning (EFCIL) aims to sequentially learn from distinct categories without retaining exemplars but easily suffers from catastrophic forgetting of learned knowledge. While existing EFCIL methods leverage knowledge distillation to alleviate forgetting, they still face…

2025

SiMHand: Mining Similar Hands for Large-Scale 3D Hand Pose Pre-training

ICLR 2025poster

We present a framework for pre-training of 3D hand pose estimation from in-the-wild hand images sharing with similar hand characteristics, dubbed SiMHand. Pre-training with large-scale images achieves promising results in various tasks, but prior methods for 3D hand pose pre-training have not fully…

2025

SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image Editing

ICCV 2025poster

Due to the challenges of manually collecting accurate editing data, existing datasets are typically constructed using various automated methods, leading to noisy supervision signals caused by the mismatch between editing instructions and original-edited image pairs. Recent efforts attempt to improve…

2025

Synthetic-to-Real Self-supervised Robust Depth Estimation via Learning with Motion and Structure Priors

CVPR 2025poster

Self-supervised depth estimation from monocular cameras in diverse outdoor conditions, such as daytime, rain, and nighttime, is challenging due to the difficulty of learning universal representations and the severe lack of labeled real-world adverse data.Previous methods either rely on synthetic inp…

2025

Test-Time Graph Neural Dataset Search With Generative Projection

ICML 2025poster

In this work, we address the test-time adaptation challenge in graph neural networks (GNNs), focusing on overcoming the limitations in flexibility and generalization inherent in existing data-centric approaches. To this end, we propose a novel research problem, test-time graph neural dataset search,…

Cited by 0SourcePDFScholar
2025

Text2World: Benchmarking Large Language Models for Symbolic World Model Generation

ACL 2025finding

Recently, there has been growing interest in leveraging large language models (LLMs) to generate symbolic world models from textual descriptions. Although LLMs have been extensively explored in the context of world modeling, prior studies encountered several challenges, including evaluation randomne…

2025

To Think or Not To Think: A Study of Thinking in Rule-Based Visual Reinforcement Fine-Tuning

NeurIPS 2025spotlight

This paper investigates the role of explicit thinking process in rule-based reinforcement fine-tuning (RFT) for multi-modal large language models (MLLMs). We first extend \textit{Thinking-RFT} to image classification task, using verifiable rewards for fine-tuning~(FT). Experiments show {Thinking-RFT…

Cited by 0SourceScholar
2025

Understanding the Thinking Process of Reasoning Models: A Perspective from Schoenfeld’s Episode Theory

EMNLP 2025

While Large Reasoning Models (LRMs) generate extensive chain-of-thought reasoning, we lack a principled framework for understanding how these thoughts are structured. In this paper, we introduce a novel approach by applying Schoenfeld’s Episode Theory, a classic cognitive framework for human mathema

2025

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

NeurIPS 2025poster

Reinforcement learning (RL) has shown great effectiveness for fine-tuning large language models (LLMs) using tasks that are challenging yet easily verifiable, such as math reasoning or code generation. However, extending this success to visual perception in vision–language models (VLMs) has been imp…

Cited by 0SourcecodeScholar
2025

WMarkGPT: Watermarked Image Understanding via Multimodal Large Language Models

ICML 2025poster

Invisible watermarking is widely used to protect digital images from unauthorized use. Accurate assessment of watermarking efficacy is crucial for advancing algorithmic development. However, existing statistical metrics, such as PSNR, rely on access to original images, which are often unavailable in…

2025

What Happened in LLMs Layers when Trained for Fast vs. Slow Thinking: A Gradient Perspective

ACL 2025long

What makes a difference in the post-training of LLMs? We investigate the training patterns of different layers in large language models (LLMs) through the lens of the gradient. We are specifically interested in how fast vs. slow thinking affects the layer-wise gradients, given the recent popularity…

2025

When Hypergraph Meets Heterophily: New Benchmark Datasets and Baseline

AAAI 2025technical

Hypergraph neural networks (HNNs) have shown promise in handling tasks characterized by high-order correlations, achieving notable success across various applications. However, there has been limited focus on heterophilic hypergraph learning (HHL), in contrast to the increasing attention given to gr…

Cited by 1SourcePDFScholar
2025

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models

AAAI 2025technical

The target of video moment retrieval (VMR) is predicting temporal spans within a video that semantically match a given linguistic query. Existing VMR methods based on multimodal large language models (MLLMs) overly rely on expensive high-quality datasets and time-consuming fine-tuning. Although some…

Cited by 2SourcePDFScholar
2024

A Dual-Path Framework with Frequency-and-Time Excited Network for Anomalous Sound Detection

ICASSP 2024accepted

In contrast to human speech, machine-generated sounds of the same type often exhibit consistent frequency characteristics and discernible temporal periodicity. However, leveraging these dual attributes in anomaly detection remains relatively under-explored. In this paper, we propose an automated dua…

Cited by 0SourceScholar
2024

A Multi-Scale Convolutional Hybrid Attention Residual Network for Enhancing Underwater Image and Identifying Underwater Multi-Scene Sea Cucumber

RA-L 2024

At present, the use of underwater robots to replace underwater manual work is a future development direction. The complex and changeable underwater environment brings great difficulties to the operation of robots. In order to improve the problem of color distortion and degradation of sea cucumber im

Cited by 2SourceScholar
2024

Adaptive Chroma Block Vector Derivation from Luma for Screen Content Coding

ICASSP 2024accepted

Intra Block Copy (IBC) and Intra Template Matching Prediction (IntraTMP) are two efficient algorithms to sufficiently exploit the correlation in the same picture. Block Vector (BV) is used to represent the displacement between the current block and its reference within the same picture. The BV infor…

Cited by 0SourceScholar
2024

Can LLMs Speak For Diverse People? Tuning LLMs via Debate to Generate Controllable Controversial Statements

ACL 2024findings

Making LLMs speak for different, especially minority groups of people, and generate statements supporting their diverse or even controversial perspectives is critical to creating an inclusive environment. However, existing LLMs lack sufficient controllability to the stance of their generated content…

2024

Efficient Personal Voice Activity Detection with Wake Word Reference Speech

ICASSP 2024accepted

Personal voice activity detection (PVAD) is gradually used in speech assistants. Traditional PVAD schemes extract the target speaker’s embedding from existing query reference speech through a pre-trained speaker verification model. Consequently, the performance of the PVAD model may suffer if the qu…

Cited by 0SourceScholar
2024

Efficient and Stable Offline-to-online Reinforcement Learning via Continual Policy Revitalization

IJCAI 2024poster

In offline Reinforcement Learning (RL), the pre-trained policies are utilized for initialization and subsequent online fine-tuning. However, existing methods suffer from instability and low sample efficiency compared to pure online learning. This paper identifies these limitations stemming from dire…

2024

From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

NAACL 2024long

In the realm of Large Language Models (LLMs), the balance between instruction data quality and quantity is a focal point. Recognizing this, we introduce a self-guided methodology for LLMs to autonomously discern and select cherry samples from open-source datasets, effectively minimizing manual curat…

2024

HC-GAE: The Hierarchical Cluster-based Graph Auto-Encoder for Graph Representation Learning

NeurIPS 2024poster

Graph Auto-Encoders (GAEs) are powerful tools for graph representation learning. In this paper, we develop a novel Hierarchical Cluster-based GAE (HC-GAE), that can learn effective structural characteristics for graph data analysis. To this end, during the encoding process, we commence by utilizing…

Cited by 1SourcePDFScholar
2024

How Universal Polynomial Bases Enhance Spectral Graph Neural Networks: Heterophily, Over-smoothing, and Over-squashing

ICML 2024poster

Spectral Graph Neural Networks (GNNs), alternatively known as *graph filters*, have gained increasing prevalence for heterophily graphs. Optimal graph filters rely on Laplacian eigendecomposition for Fourier transform. In an attempt to avert prohibitive computations, numerous polynomial filters have…

2024

Joint Inference of Speaker Diarization and ASR with Multi-Stage Information Sharing

ICASSP 2024accepted

In this paper, we introduce a novel approach that unifies Automatic Speech Recognition (ASR) and speaker diarization in a cohesive framework. Utilizing the synergies between the two tasks, our method effectively extracts speaker-specific information from the lower layers of a pretrained Conformer-ba…

Cited by 2SourceScholar
2024

Leveraging Biases in Large Language Models: "bias-kNN" for Effective Few-Shot Learning

ICASSP 2024accepted

Large Language Models (LLMs) have shown significant promise in various applications, including zero-shot and few-shot learning. However, their performance can be hampered by inherent biases. Instead of traditionally sought methods that aim to minimize or correct these biases, this study introduces a…

Cited by 0SourceScholar
2024

Multi-Objective Progressive Clustering for Semi-Supervised Domain Adaptation in Speaker Verification

ICASSP 2024accepted

Utilizing the pseudo-labeling algorithm with large-scale unlabeled data becomes crucial for semi-supervised domain adaptation in speaker verification tasks. In this paper, we propose a novel pseudo-labeling method named Multi-objective Progressive Clustering (MoPC), specifically designed for semi-su…

Cited by 0SourceScholar
2024

QBMK: Quantum-based Matching Kernels for Un-attributed Graphs

ICML 2024spotlight

In this work, we develop a new Quantum-based Matching Kernel (QBMK) for un-attributed graphs, by computing the kernel-based similarity between the quantum Shannon entropies of aligned vertices through the Continuous-time Quantum Walk (CTQW). The theoretical analysis reveals that the proposed QBMK ke…

Cited by 0SourcePDFScholar
2024

Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual Conformer

ICASSP 2024accepted

In recent years, neural network-based Wake Word Spotting achieves good performance on clean audio samples but struggles in noisy environments. Audio-Visual Wake Word Spotting (AVWWS) receives lots of attention because visual lip movement information is not affected by complex acoustic scenes. Previo…

Cited by 0SourceScholar
2024

Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning

ACL 2024findings

Instruction tuning is critical to large language models (LLMs) for achieving better instruction following and task adaptation capabilities but its success heavily relies on the training data quality. Many recent methods focus on improving the data quality but often overlook the compatibility of the…

2024

SlideSpeech: A Large Scale Slide-Enriched Audio-Visual Corpus

ICASSP 2024accepted

Multi-Modal automatic speech recognition (ASR) techniques aim to leverage additional modalities to improve the performance of speech recognition systems. While existing approaches primarily focus on video or contextual information, the utilization of extra supplementary textual information has been…

Cited by 0SourceScholar
2024

Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning

ACL 2024long

Instruction tuning is critical to improve LLMs but usually suffers from low-quality and redundant data. Data filtering for instruction tuning has proved important in improving both the efficiency and performance of the tuning process. But it also leads to extra cost and computation due to the involv…

2024

Vision-Language Model Fine-Tuning via Simple Parameter-Efficient Modification

EMNLP 2024main

Recent advances in fine-tuning Vision-Language Models (VLMs) have witnessed the success of prompt tuning and adapter tuning, while the classic model fine-tuning on inherent parameters seems to be overlooked. It is believed that fine-tuning the parameters of VLMs with few-shot samples corrupts the pr…

2024

Visual Pivoting Unsupervised Multimodal Machine Translation in Low-Resource Distant Language Pairs

EMNLP 2024finding

Unsupervised multimodal machine translation (UMMT) aims to leverage vision information as a pivot between two languages to achieve better performance on low-resource language pairs. However, there is presently a challenge: how to handle alignment between distant language pairs (DLPs) in UMMT. To thi…

2024

Voxblink: A Large Scale Speaker Verification Dataset on Camera

ICASSP 2024accepted

In this paper, we introduce a large-scale and high-quality audiovisual speaker verification dataset, named VoxBlink. We propose an innovative and robust automatic audio-visual data mining pipeline to curate this dataset, which contains 1.45M utterances from 38K speakers. Due to the inherent nature o…

Cited by 0SourceScholar
2023

AlignDet: Aligning Pre-training and Fine-tuning in Object Detection

ICCV 2023poster

The paradigm of large-scale pre-training followed by downstream fine-tuning has been widely employed in various object detection algorithms. In this paper, we reveal discrepancies in data, model, and task between the pre-training and fine-tuning procedure in existing practices, which implicitly limi…

Cited by 22PDFcodeScholar
2023

Capturing the Long-Distance Dependency in the Control Flow Graph via Structural-Guided Attention for Bug Localization

IJCAI 2023poster

To alleviate the burden of software maintenance, bug localization, which aims to automatically locate the buggy source files based on the bug report, has drawn significant attention in the software mining community. Recent studies indicate that the program structure in source code carries more seman…

Cited by 6SourcePDFScholar
2023

Cooperative and Adversarial Learning: Co-enhancing Discriminability and Transferability in Domain Adaptation

AAAI 2023technical

Discriminability and transferability are two goals of feature learning for domain adaptation (DA), as we aim to find the transferable features from the source domain that are helpful for discriminating the class label in the target domain. Modern DA approaches optimize discriminability and transfera…

2023

Exploring Universal Singing Speech Language Identification Using Self-Supervised Learning Based Front-End Features

ICASSP 2023accepted

Despite the great performance of language identification (LID), there is a lack of large-scale singing LID databases to support the research of singing language identification (SLID). This paper proposed a over 3200 hours dataset used for singing language identification, called Slingua. As the basel…

Cited by 0SourceScholar
2023

FreeSeg: Unified, Universal and Open-Vocabulary Image Segmentation

CVPR 2023poster

Recently, open-vocabulary learning has emerged to accomplish segmentation for arbitrary categories of text-based descriptions, which popularizes the segmentation system to more general-purpose application scenarios. However, existing methods devote to designing specialized architectures or parameter…

2023

How Powerful are Shallow Neural Networks with Bandlimited Random Weights?

ICML 2023poster

We investigate the expressive power of depth-2 bandlimited random neural networks. A random net is a neural network where the hidden layer parameters are frozen with random assignment, and only the output layer parameters are trained by loss minimization. Using random weights for a hidden layer is a…

Cited by 10SourcePDFScholar
2023

Identifying Source Speakers for Voice Conversion Based Spoofing Attacks on Speaker Verification Systems

ICASSP 2023accepted

An automatic speaker verification system aims to verify the speaker identity of a speech signal. However, a voice conversion system could manipulate a person’s speech signal to make it sound like another speaker’s voice and deceive the speaker verification system. Most countermeasures for voice conv…

Cited by 0SourceScholar
2023

PRCA: Fitting Black-Box Large Language Models for Retrieval Question Answering via Pluggable Reward-Driven Contextual Adapter

EMNLP 2023long main

The Retrieval Question Answering (ReQA) task employs the retrieval-augmented framework, composed of a retriever and generator. The generators formulate the answer based on the documents retrieved by the retriever. Incorporating Large Language Models (LLMs) as generators is beneficial due to their ad…

Cited by 0SourceScholar
2023

STPrivacy: Spatio-Temporal Privacy-Preserving Action Recognition

ICCV 2023poster

Existing methods of privacy-preserving action recognition (PPAR) mainly focus on frame-level (spatial) privacy removal through 2D CNNs. Unfortunately, they have two major drawbacks. First, they may compromise temporal dynamics in input videos, which are critical for accurate action recognition. Seco…

Cited by 24PDFScholar
2023

Stability and Generalization of lp-Regularized Stochastic Learning for GCN

IJCAI 2023poster

Graph convolutional networks (GCN) are viewed as one of the most popular representations among the variants of graph neural networks over graph data and have shown powerful performance in empirical experiments. That l2-based graph smoothing enforces the global smoothness of GCN, while (soft) l1-base…

Cited by 1SourcePDFScholar
2023

Target-Speaker Voice Activity Detection Via Sequence-to-Sequence Prediction

ICASSP 2023accepted

Target-speaker voice activity detection is currently a promising approach for speaker diarization in complex acoustic environments. This paper presents a novel Sequence-to-Sequence Target-Speaker Voice Activity Detection (Seq2Seq-TSVAD) method that can efficiently address the joint modeling of large…

Cited by 0SourceScholar
2023

The DKU Post-Challenge Audio-Visual Wake Word Spotting System for the 2021 MISP Challenge: Deep Analysis

ICASSP 2023accepted

This paper further explores our previous wake word spotting system ranked 2-nd in Track 1 of the MISP Challenge 2021. First, we investigate a robust unimodal approach based on 3D and 2D convolution and adopt the simple attention module (SimAM) for our system to improve performance. Second, we explor…

Cited by 0SourceScholar
2023

The WHU-Alibaba Audio-Visual Speaker Diarization System for the MISP 2022 Challenge

ICASSP 2023accepted

This paper describes the system developed by the WHU-Alibaba team for the Multimodal Information Based Speech Processing (MISP) 2022 Challenge. We extend the Sequence-to-Sequence Target-Speaker Voice Activity Detection framework to simultaneously detect multiple speakers’ voice activities from audio…

Cited by 0SourceScholar
2022

Cross-Channel Attention-Based Target Speaker Voice Activity Detection: Experimental Results for the M2met Challenge

ICASSP 2022accepted

DukeECE. As the highly overlapped speech exists in the dataset, we employ an x-vector-based target-speaker voice activity detection (TS-VAD) to find the overlap between speakers. Firstly, we separately train a single-channel model for each of the 8 channels and fuse the results. In addition, we also…

Cited by 32SourceScholar
2022

Discriminator-Guided Model-Based Offline Imitation Learning

CoRL 2022poster

Offline imitation learning (IL) is a powerful method to solve decision-making problems from expert demonstrations without reward labels. Existing offline IL methods suffer from severe performance degeneration under limited expert data. Including a learned dynamics model can potentially improve the s…

Cited by 22SourceScholar
2022

Few-Shot Non-Parametric Learning with Deep Latent Variable Model

NeurIPS 2022accept

Most real-world problems that machine learning algorithms are expected to solve face the situation with (1) unknown data distribution; (2) little domain-specific knowledge; and (3) datasets with limited annotation. We propose Non-Parametric learning by Compression with Latent Variables (NPC-LV), a l…

Cited by 14SourcePDFScholar
2022

MODE: Multi-View Omnidirectional Depth Estimation with 360° Cameras

ECCV 2022poster

"In this paper, we propose a two-stage omnidirectional depth estimation framework with multi-view 360-degree cameras. The framework first estimates the depth maps from different camera pairs via omnidirectional stereo matching and then fuses the depth maps to achieve robustness against mud spots, wa…

2022

Multi-Granularity Distillation Scheme towards Lightweight Semi-Supervised Semantic Segmentation

ECCV 2022poster

"Albeit with varying degrees of progress in the field of Semi-Supervised Semantic Segmentation, most of its recent successes are involved in unwieldy models and the lightweight solution is still not yet explored. We find that existing knowledge distillation techniques pay more attention to pixel-lev…

2022

SIG-VC: A Speaker Information Guided Zero-Shot Voice Conversion System for Both Human Beings and Machines

ICASSP 2022accepted

Nowadays, as more and more systems achieve good performance in traditional voice conversion (VC) tasks, people’s attention gradually turns to VC tasks under extreme conditions. In this paper, we propose a novel method for zero-shot voice conversion. We aim to obtain intermediate representations for…

Cited by 0SourceScholar
2022

Simple Attention Module Based Speaker Verification with Iterative Noisy Label Detection

ICASSP 2022accepted

Recently, the attention mechanism such as squeeze-and-excitation module (SE) and convolutional block attention module (CBAM) has achieved great success in deep learning-based speaker verification system. This paper introduces an alternative effective yet simple one, i.e., simple attention module (Si…

Cited by 0SourceScholar
2022

The DKU Audio-Visual Wake Word Spotting System for the 2021 MISP Challenge

ICASSP 2022accepted

This paper describes the system developed by the DKU team for the MISP Challenge 2021. We present a two-stage approach consisting of end-to-end neural networks for the audio-visual wake word spotting task. We first process audio and video data to give them a similar structure and then train two unim…

Cited by 0SourceScholar
2022

Towards Lightweight Applications: Asymmetric Enroll-Verify Structure for Speaker Verification

ICASSP 2022accepted

With the development of deep learning, automatic speaker verification has made considerable progress over the past few years. However, to design a lightweight and robust system with limited computational resources is still a challenging problem. Traditionally, a speaker verification system is symmet…

Cited by 0SourceScholar
2022

Unified Matrix Coding for NN Originated MIP in H.266/VVC

ICASSP 2022accepted

Matrix-based Intra Prediction (MIP) is an effective coding algorithm in H.266/Versatile Video Coding (VVC) which is originated by Neural Networks (NN). With the requirement of low complexity, MIP is conducted by a matrix-vector multiplication. To handle with the diversity of video content, 30 matric…

Cited by 0SourceScholar
2022

When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement Learning

NeurIPS 2022accept

Learning effective reinforcement learning (RL) policies to solve real-world complex tasks can be quite challenging without a high-fidelity simulation environment. In most cases, we are only given imperfect simulators with simplified dynamics, which inevitably lead to severe sim-to-real gaps in RL po…

2021

How Framelets Enhance Graph Neural Networks

ICML 2021spotlight

This paper presents a new approach for assembling graph neural networks based on framelet transforms. The latter provides a multi-scale representation for graph-structured data. We decompose an input graph into low-pass and high-pass frequencies coefficients for network training, which then defines…

2021

Multi-Task Dense Retrieval via Model Uncertainty Fusion for Open-Domain Question Answering

EMNLP 2021finding

Multi-task dense retrieval models can be used to retrieve documents from a common corpus (e.g., Wikipedia) for different open-domain question-answering (QA) tasks. However, Karpukhin et al. (2020) shows that jointly learning different QA tasks with one dense model is not always beneficial due to cor…

2021

Segatron: Segment-Aware Transformer for Language Modeling and Understanding

AAAI 2021technical

Transformers are powerful for sequence modeling. Nearly all state-of-the-art language models and pre-trained language models are based on the Transformer architecture. However, it distinguishes sequential tokens only with the token position index. We hypothesize that better contextual representation…

2021

Self-Supervised Geometric Features Discovery via Interpretable Attention for Vehicle Re-Identification and Beyond

ICCV 2021poster

To learn distinguishable patterns, most of recent works in vehicle re-identification (ReID) struggled to redevelop official benchmarks to provide various supervisions, which requires prohibitive human labors. In this paper, we seek to achieve the similar goal but do not involve more human efforts. T…

Cited by 60PDFcodeScholar
2021

Unsupervised Chunking as Syntactic Structure Induction with a Knowledge-Transfer Approach

EMNLP 2021finding

In this paper, we address unsupervised chunking as a new task of syntactic structure induction, which is helpful for understanding the linguistic structures of human languages as well as processing low-resource languages. We propose a knowledge-transfer approach that heuristically induces chunk labe…

2020

An Obstacle-crossing Strategy Based on the Fast Self-reconfiguration for Modular Sphere Robots

IROS 2020poster

This paper introduces an obstacle-crossing strategy, and the self-reconfiguration algorithm for a new class of modular robots called the rolling sphere, which can fit obstacles represented by cubes of different sizes due to the chain connection of multiple spheres. For the self-reconfiguration of th…

Cited by 19SourceScholar
2020

FreeBOT: A Freeform Modular Self-reconfigurable Robot with Arbitrary Connection Point - Design and Implementation

IROS 2020poster

This paper proposes a novel modular selfreconfigurable robot (MSRR) "FreeBOT", which can be connected freely at any point on other robots. FreeBOT is mainly composed of two parts: a spherical ferromagnetic shell and an internal magnet. The connection between the modules is genderless and instant, si…

Cited by 99SourceScholar
2020

On-chip integration of ultra-thin glass cantilever for physical property measurement activated by femtosecond laser impulse

IROS 2020poster

Under the excitation of acoustic radiation, the amount of energy absorbed and rebounded by cells have the relationship with mechanical properties, e.g. stiffness, shape, weight and so on. In this paper, a femtosecond laser-activated micro-detector is designed to convert this relationship into an ele…

Cited by 3SourceScholar
2020

Path Integral Based Convolution and Pooling for Graph Neural Networks

NeurIPS 2020poster

Graph neural networks (GNNs) extends the functionality of traditional neural networks to graph-structured data. Similar to CNNs, an optimized design of graph convolution and pooling is key to success. Borrowing ideas from physics, we propose a path integral based graph neural networks (PAN) for clas…

2020

Robot-to-Robot Relative Pose Estimation based on Semidefinite Relaxation Optimization

IROS 2020poster

In this paper, the 2D robot-to-robot relative pose (position and orientation) estimation problem based on ego-motion and noisy distance measurements is considered. We address this problem using an optimization-based method, which does not require complicated numerical analysis while yields no inferi…

Cited by 21SourceScholar
2020

Within-Sample Variability-Invariant Loss for Robust Speaker Recognition Under Noisy Environments

ICASSP 2020accepted

Despite the significant improvements in speaker recognition enabled by deep neural networks, unsatisfactory performance persists under noisy environments. In this paper, we train the speaker embedding network to learn the "clean" embedding of the noisy utterance. Specifically, the network is trained…

Cited by 0SourceScholar