← Search

Jing Liu

184 accepted papers

2026

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding

CVPR 2026

Processing long-form videos with Video Large Language Models (Video-LLMs) is computationally prohibitive. Current efficiency methods often compromise fine-grained perception through irreversible information disposal or inhibit long-range temporal modeling via rigid, predefined sparse patterns. This

Cited by 0SourceScholar
2026

BEE-RAG: Balanced Entropy Engineering for Retrieval-Augmented Generation

AAAI 2026technical

With the rapid advancement of large language models (LLMs), retrieval-augmented generation (RAG) has emerged as a critical approach to supplement the inherent knowledge limitations of LLMs. However, due to the typically large volume of retrieved information, RAG tends to operate with long context le

Cited by 0SourcePDFScholar
2026

Certain Head, Uncertain Tail: Expert-Sample for Test-Time Scaling in Fine-Grained MoE

ICML 2026poster

Test-time scaling improves LLM performance by generating multiple candidate solutions, yet token-level sampling requires temperature tuning that trades off diversity against stability. Fine-grained MoE, featuring hundreds of well-trained experts per layer and multi-expert activation per token, offer…

Cited by 0SourceScholar
2026

CompetitorFormer: Mitigating Query Conflicts for 3D Instance Segmentation via Competitive Strategy

CVPR 2026

Transformer-based approaches have recently become the dominant paradigm for 3D instance segmentation. These methods typically employ a multi-layer decoder that iteratively refines a set of learnable queries into instance mask predictions. However, we observe that multiple queries often target the sa

Cited by 0SourcecodeScholar
2026

DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning

ICML 2026poster

Recent advances in multimodal language models (MLLMs) have made thinking with images a dominant paradigm for multimodal reasoning. However, existing methods still fail to ensure evidence–answer consistency, where correct answers must be supported by correct visual evidence. To address this issue, we…

Cited by 0SourceScholar
2026

Divid: Disentangled Spatial-Temporal Modeling within LLMs for Temporally Grounded Video Understanding

ICLR 2026poster

Recent advances in Video LLMs have improved video understanding performance, but temporally grounded understanding in long-form videos remains challenging. Most models encode video frames into a flat sequence of visual tokens, which are then processed together with textual input by the LLM. While ef…

Cited by 0SourceScholar
2026

DynaSchedBench: Calibrated Dynamic Scheduling Benchmarks and Observability Paradox in LLM-based Scheduling Agents

ICML 2026poster

Progress in neural combinatorial optimization for Dynamic Flexible Job Shop Scheduling Problem (DFJSP) is currently hindered by a methodological tension: static benchmarks encourage benchmark overfitting, while uncalibrated generators obscure algorithmic difficulty with stochastic noise. To resolve …

Cited by 0SourceScholar
2026

FedGLoRA: Grassmann-Manifold Federated Learning via Dual LoRA for Large EEG Models

IJCAI 2026

Large EEG Models (LEMs) are drawing increasing attention in EEG, as large-scale pretraining yields transferable representations that improve generalization. As EEG research moves to real-world deployment, objectives and paradigms diversify, yielding increasingly heterogeneous and unevenly scaled dat

Cited by 0Scholar
2026

FossilWriter: Learning Hypergraph World Models with Latent Narratives for Creative Story Generation

IJCAI 2026

Creative story generation has achieved notable progress with large language models. Current methods construct narratives through hierarchical planning or incremental expansion. These approaches produce structurally complete stories but offer limited support for organic narrative development. Many fi

Cited by 0Scholar
2026

GeoX-Bench: Benchmarking Cross-View Geo-Localization and Pose Estimation Capabilities of Large Multimodal Models

AAAI 2026technical

Large multimodal models (LMMs) have demonstrated remarkable capabilities across a wide range of tasks, however their knowledge and abilities in the cross-view geo-localization and pose estimation domains remain unexplored, despite potential benefits for navigation, autonomous driving, outdoor roboti

Cited by 0SourcePDFScholar
2026

Harmonizing Real-Time Constraints and Long-Horizon Reasoning: An Asynchronous Agentic Framework for Dynamic Scheduling

IJCAI 2026

The Dynamic Flexible Job Shop Scheduling Problem (DFJSP) necessitates a trade-off between instant reaction to stochastic disturbances and global optimization of production goals. Conventional priority rules are insufficiently flexible to handle complex disruptions, whereas learning-based approaches

Cited by 0Scholar
2026

INT vs. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats

ICML 2026poster

Modern AI hardware, such as Nvidia's Blackwell architecture, is increasingly embracing low-precision floating-point (FP) formats to handle the pervasive activation outliers in Large Language Models (LLMs). Despite this industry trend, a unified comparison of FP and integer (INT) quantization across …

Cited by 0SourceScholar
2026

LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws

ICML 2026poster

Existing scaling laws for Large Language Models (LLMs), predominantly monotonic power laws, have successfully guided model development but fail to explain emerging non-monotonic phenomena such as catastrophic overtraining and quantization-induced degradation, where performance deteriorates despite i…

Cited by 0SourceScholar
2026

LatentLLM: Activation-Aware Transform to Multi-Head Latent Attention

AAAI 2026technical

Modern foundation models such as large language models (LLMs) require a massive amount of computational and memory resources. We propose a new framework to convert such LLMs into a reduced-dimension latent structure. Our method extends a local activation-aware tensor decomposition to a global attent

Cited by 0SourcePDFScholar
2026

MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering

ICML 2026spotlight

The evolution of Large Language Model (LLM) agents for software engineering (SWE) is constrained by the scarcity of verifiable datasets, a bottleneck stemming from the complexity of constructing executable environments across diverse languages. To address this, we introduce **MEnvAgent**, a **M**ult…

Cited by 0SourceScholar
2026

ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?

ICLR 2026poster

Omnidirectional images (ODIs) provide full 360$^{\circ} \times$ 180$^{\circ}$ view which are widely adopted in VR, AR and embodied intelligence applications. While multi-modal large language models (MLLMs) have demonstrated remarkable performance on conventional 2D image and video understanding benc…

Cited by 0SourceScholar
2026

OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs

AAAI 2026technical

Existing sparse attention methods primarily target inference-time acceleration by selecting critical tokens under predefined sparsity patterns. However, they often fail to bridge the training–inference gap and lack the capacity for fine-grained token selection across multiple dimensions—such as quer

Cited by 0SourcePDFScholar
2026

OpenT2M: No-frill Motion Generation with Open-source, Large-scale, High-quality Data

CVPR 2026

Text-to-motion (T2M) generation aims to create realistic human movements from text descriptions, with promising applications in animation and robotics. Despite recent progress, current T2M models perform poorly on unseen text descriptions due to the small scale and limited diversity of existing moti

Cited by 0SourceScholar
2026

Order Matters: Unveiling the Hidden Impact of Macro Placement Sequences via Proxy-Guided LLM Evolution

ICML 2026poster

Macro placement is a fundamental step in modern VLSI physical design, determining the solution quality of high-dimensional combinatorial optimization problems. Despite recent advancements in machine learning for spatial coordinate determination, the temporal dimension of placement sequencing remains…

Cited by 0SourceScholar
2026

QVGen: Pushing the Limit of Quantized Video Generative Models

ICLR 2026poster

Video diffusion models (DMs) have enabled high-quality video synthesis. Yet, their substantial computational and memory demands pose serious challenges to real-world deployment, even on high-end GPUs. As a commonly adopted solution, quantization has proven notable success in reducing cost for image…

Cited by 0SourcecodeScholar
2026

ROBUST MULTIMODAL REPRESENTATION LEARNING IN HEALTHCARE

ICASSP 2026poster

Medical multimodal representation learning aims to integrate heterogeneous data into unified patient representations to support clinical outcome prediction. However, real-world medical datasets commonly contain systematic biases from multiple sources, which poses significant challenges for medical m…

Cited by 0SourcePDFScholar
2026

ROBUST MULTIMODAL REPRESENTATION LEARNING IN HEALTHCARE

ICASSP 2026poster

Medical multimodal representation learning aims to integrate heterogeneous data into unified patient representations to support clinical outcome prediction. However, real-world medical datasets commonly contain systematic biases from multiple sources, which poses significant challenges for medical m…

Cited by 0SourcePDFScholar
2026

SFGA: Similarity-Constrained Fusion Learning for Unsupervised Anomaly Detection in Multiplex Graphs

AAAI 2026technical

Multiplex graphs are widely used to model multi-relational complex systems and play an important role in various real-world scenarios, such as financial systems and social networks. Hence, detecting anomalous samples in multiplex graph becomes crucial to ensure cybersecurity and stability. Although

Cited by 0SourcePDFScholar
2026

SimpleDiffusion: A Lightweight and Efficient Conditional Diffusion Model for Multi-Modal Salient Object Detection

AAAI 2026technical

Multi-modal salient object detection (MSOD), which integrates complementary modalities such as depth or thermal data, primarily faces two challenges: accurately preserving salient object details and effectively aligning cross-modal features. Recent advances in using Stable Diffusion to generate imag

Cited by 0SourcePDFScholar
2026

Sparsity Forcing: Reinforcing Token Sparsity of MLLMs

ICLR 2026poster

Sparse attention mechanisms aim to reduce computational overhead with minimal accuracy loss by selectively processing salient tokens. Despite their effectiveness, most methods merely exploit a model’s inherent sparsity and thus plateau at moderate budgets (about 50\% token reduction), with little he…

Cited by 0SourceScholar
2026

SurveilNav: Collaborative Object Goal Navigation with Robot and Surveillance System

ICRA 2026poster

With the growing deployment of surveillance systems in factories, offices, and homes, integrating them with robots offers a promising direction for collaborative and efficient task execution. However, existing approaches largely focus on single-robot scenarios and struggle with multi-view collaborat…

2026

Textual Self-Attention Network: Test-Time Preference Optimization Through Textual Gradient-Based Attention

AAAI 2026technical

Large Language Models (LLMs) have demonstrated remarkable generalization capabilities, but aligning their outputs with human preferences typically requires expensive supervised fine-tuning. Recent test-time methods leverage textual feedback to overcome this, but they often critique and revise a sing

Cited by 0SourcePDFScholar
2026

UrbanNav: Learning Language-Guided Embodied Urban Navigation from Web-Scale Human Trajectories

AAAI 2026technical

Navigating complex urban environments using natural language instructions poses significant challenges for embodied agents, including noisy language instructions, ambiguous spatial references, diverse landmarks, and dynamic street scenes. Current visual navigation methods are typically limited to si

Cited by 0SourcePDFScholar
2026

VisualPrompter: Semantic-Aware Prompt Optimization with Visual Feedback for Text-to-Image Synthesis

ICLR 2026poster

The notable gap between user-provided and model-preferred prompts poses a significant challenge for generating high-quality images with text-to-image models, compelling the need for prompt engineering. Current studies on prompt engineering can effectively enhance the style and aesthetics of generate…

Cited by 0SourcecodeScholar
2026

W-EDIT: A Wavelet-Based Frequency-Aware Framework for Text-Driven Image Editing

ICLR 2026poster

While recent advances in Diffusion Transformers (DiTs) have significantly advanced text-to-image generation, text-driven image editing remains challenging. Existing approaches either struggle to balance structural preservation with flexible modifications or require costly fine-tuning of large models…

Cited by 0SourceScholar
2025

A Robust Lifelong Multi-Agent Path Finding With Active Conflict Resolution and Decentralized Execution

RA-L 2025

Multi-Agent Path Finding (MAPF) focuses on navigating agents along cost-efficient and conflict-free paths. This letter investigates a challenging and practical MAPF variant, namely Robust Lifelong MAPF (RLMAPF), where agents sequentially receive tasks and effectively deal with uncertainties. In this

Cited by 5SourceScholar
2025

AR-Diffusion: Asynchronous Video Generation with Auto-Regressive Diffusion

CVPR 2025poster

The task of video generation requires synthesizing visually realistic and temporally coherent video frames. Existing methods primarily use asynchronous auto-regressive models or synchronous diffusion models to address this challenge. However, asynchronous auto-regressive models often suffer from inc…

2025

Ada-K Routing: Boosting the Efficiency of MoE-based LLMs

ICLR 2025poster

In the era of Large Language Models (LLMs), Mixture-of-Experts (MoE) architectures offer a promising approach to managing computational costs while scaling up model parameters. Conventional MoE-based LLMs typically employ static Top-K routing, which activates a fixed and equal number of experts for…

Cited by 1SourcePDFScholar
2025

An Easy Method for Extrinsic Calibration of Camera and Time-of-Flight Sensor

IROS 2025

A multi-zone (typically 8×8) time-of-flight (ToF) sensor offers a low-cost, low-power, and compact solution for range measurement, making it ideal for specialized robotic applications. However, its low resolution limits its usability. Pairing a ToF sensor with a camera enhances depth perception and

Cited by 0SourcecodeScholar
2025

An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

ACL 2025finding

Recently, there has been a growing trend of utilizing Large Language Model (LLM) to evaluate the quality of other LLMs. Many studies have fine-tuned judge models based on open-source LLMs for evaluation. While the fine-tuned judge models are claimed to achieve comparable evaluation capability with G…

2025

AutoSGNN: Automatic Propagation Mechanism Discovery for Spectral Graph Neural Networks

AAAI 2025technical

In real-world applications, spectral Graph Neural Networks (GNNs) are powerful tools for processing diverse types of graphs. However, a single GNN often struggles to handle different graph types—such as homogeneous and heterogeneous graphs—simultaneously. This challenge has led to the manual design…

2025

Breaking the Encoder Barrier for Seamless Video-Language Understanding

ICCV 2025poster

Most Video-Large Language Models (Video-LLMs) adopt an encoder-decoder framework, where a vision encoder extracts frame-wise features for processing by a language model. However, this approach incurs high computational costs, introduces resolution biases, and struggles to capture fine-grained multim…

Cited by 0SourcePDFScholar
2025

C-NAV: Towards Self-Evolving Continual Object Navigation in Open World

NeurIPS 2025poster

Embodied agents are expected to perform object navigation in dynamic, open-world environments. However, existing approaches typically rely on static trajectories and a fixed set of object categories during training, overlooking the real-world requirement for continual adaptation to evolving scenario…

Cited by 0SourcecodeScholar
2025

COAP: Memory-Efficient Training with Correlation-Aware Gradient Projection

CVPR 2025poster

Training large-scale neural networks in vision, and multimodal domains demands substantial memory resources, primarily due to the storage of optimizer states. While LoRA, a popular parameter-efficient method, reduces memory usage, it often suffers from suboptimal performance due to the constraints o…

Cited by 3SourcePDFScholar
2025

COSMO: Combination of Selective Memorization for Low-cost Vision-and-Language Navigation

ICCV 2025poster

Vision-and-Language Navigation (VLN) tasks have gained prominence within artificial intelligence research due to their potential application in fields like home assistants. Many contemporary VLN approaches, while based on transformer architectures, have increasingly incorporated additional component…

2025

CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments

ICASSP 2025accepted

Creating Automatic Speech Recognition (ASR) systems that are robust and resilient to classroom conditions is paramount to the development of AI tools to aid teachers and students. In this work, we study the efficacy of continued pretraining (CPT) in adapting Wav2vec2.0 to the classroom domain. We sh…

Cited by 0SourceScholar
2025

Channel Merging: Preserving Specialization for Merged Experts

AAAI 2025technical

Lately, the practice of utilizing task-specific fine-tuning has been implemented to improve the performance of large language models (LLM) in subsequent tasks. Through the integration of diverse LLMs, the overall competency of LLMs is significantly boosted. Nevertheless, traditional ensemble methods…

2025

Context-aware Dynamic Pruning for Speech Foundation Models

ICLR 2025poster

Foundation models, such as large language models, have achieved remarkable success in natural language processing and are evolving into models capable of handling multiple modalities. Listening ability, in particular, is crucial for many applications, leading to research on building speech foundatio…

Cited by 0SourcePDFScholar
2025

DGO-VINS: A Visual-Inertial SLAM for Dynamic Environments With Geometric Constraint and Adaptive State Optimization

RA-L 2025

Traditional SLAM performs well in static environments, but experiences degeneration of localization accuracy and stability in dynamic settings. To enhance performance in dynamic environments, this letter presents DGO-VINS, a real-time dynamic visual-inertial SLAM system based on geometric constraint

Cited by 3SourceScholar
2025

DiGradPatch: Black-Box Patch Attacks via Diffusion-Based Double Gradient and Sensitive Distribution Guidance

ICASSP 2025accepted

Deep neural networks have demonstrated vulnerabilities to black-box adversarial patch attacks in image analysis tasks, raising concerns about their robustness in safety-critical applications. Current methods typically rely on randomized search strategies to determine patch locations and apply unrest…

Cited by 0SourceScholar
2025

DiMSOD: A Diffusion-Based Framework for Multi-Modal Salient Object Detection

AAAI 2025technical

Multi-modal salient object detection (SOD) through the integration of additional data such as depth or thermal information has become a significant task in computer vision during recent years. Traditionally, the challenges of identifying salient objects in RGB, RGB-D (Depth), and RGB-T (Thermal) ima…

Cited by 0SourcePDFScholar
2025

Diffusion Feedback Helps CLIP See Better

ICLR 2025poster

Contrastive Language-Image Pre-training (CLIP), which excels at abstracting open-world representations across domains and modalities, has become a foundation for a variety of vision and multimodal tasks. However, recent studies reveal that CLIP has severe visual shortcomings, such as which can hardl…

2025

ECC: Synergizing Emotion, Cause and Commonsense for Empathetic Dialogue Generation

COLING 2025main

Empathy improves human-machine dialogue systems by enhancing the user’s experience. While traditional models have aimed to detect and express users’ emotions from dialogue history, they neglect the crucial and complex interactions among emotion, emotion causes, and commonsense. To address this, we i…

2025

Efficient Motion-Aware Video MLLM

CVPR 2025highlight

Most current video MLLMs rely on uniform frame sampling and image-level encoders, resulting in inefficient data processing and limited motion awareness. To address these challenges, we introduce EMA, an Efficient Motion-Aware video MLLM that utilizes compressed video structures as inputs. We propose…

Cited by 0SourcePDFScholar
2025

Exploring the Frontiers of Animation Video Generation in the Sora Era: Method, Dataset and Benchmark

IJCAI 2025

Animation has gained significant interest in the recent film and TV industry. Despite the success of advanced video generation models like Sora, Kling, and CogVideoX in generating natural videos, they lack the same effectiveness in handling animation videos. Evaluating animation video generation is

2025

FedCross: Intertemporal Federated Learning Under Evolutionary Games

AAAI 2025technical

Federated Learning (FL) mitigates privacy leakage in decentralized machine learning by allowing multiple clients to train collaboratively locally. However, dynamic mobile networks with high mobility, intermittent connectivity, and bandwidth limitation severely hinder model updates to the cloud serv…

Cited by 0SourcePDFScholar
2025

Few-Shot Learner Generalizes Across AI-Generated Image Detection

ICML 2025poster

Current fake image detectors trained on large synthetic image datasets perform satisfactorily on limited studied generative models. However, these detectors suffer a notable performance decline over unseen models. Besides, collecting adequate training data from online generative models is often expe…

2025

Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy Leakage

AAAI 2025technical

Fine-tuning large language models on private data for downstream applications poses significant privacy risks in potentially exposing sensitive information. Several popular community platforms now offer convenient distribution of a large variety of pre-trained models, allowing anyone to publish with…

Cited by 4SourcePDFScholar
2025

Gap Preserving Distillation by Building Bidirectional Mappings with A Dynamic Teacher

ICLR 2025poster

Knowledge distillation aims to transfer knowledge from a large teacher model to a compact student counterpart, often coming with a significant performance gap between them. Interestingly, we find that a too-large performance gap can hamper the training process. To alleviate this, we propose a **Gap…

Cited by 0SourcePDFScholar
2025

Graph Contrastive Learning with Joint Spectral Augmentation of Attribute and Topology

AAAI 2025technical

As an essential technique for Graph Contrastive Learning (GCL), Graph Augmentation (GA) improves the generalization capability of the GCLs by introducing different forms of the same graph. To ensure information integrity, existing GA strategies have been designed to simultaneously process the two ty…

Cited by 0SourcePDFScholar
2025

HarmoniCa: Harmonizing Training and Inference for Better Feature Caching in Diffusion Transformer Acceleration

ICML 2025poster

Diffusion Transformers (DiTs) excel in generative tasks but face practical deployment challenges due to high inference costs. Feature caching, which stores and retrieves redundant computations, offers the potential for acceleration. Existing learning-based caching, though adaptive, overlooks the imp…

2025

ID-Patch: Robust ID Association for Group Photo Personalization

CVPR 2025poster

The ability to synthesize personalized group photos and specify the positions of each identity offers immense creative potential. While such imagery can be visually appealing, it presents significant challenges for existing technologies. A persistent issue is identity (ID) leakage, where injected fa…

2025

Investigating the Factual Knowledge Boundary of Large Language Models with Retrieval Augmentation

COLING 2025main

Large language models (LLMs) have shown impressive prowess in solving a wide range of tasks with world knowledge. However, it remains unclear how well LLMs are able to perceive their factual knowledge boundaries, particularly under retrieval augmentation settings. In this study, we present the first…

2025

Learning Beyond Still Frames: Scaling Vision-Language Models with Video

ICCV 2025poster

High-quality image-text data is critical in enhancing Vision-Language Models (VLMs), but traditional image-based pretraining approaches face limitations. These methods are resource-intensive, relying on curated, high-quality interleaved data that is costly and challenging to collect at scale. Additi…

2025

M2OST: Many-to-one Regression for Predicting Spatial Transcriptomics from Digital Pathology Images

AAAI 2025technical

The advancement of Spatial Transcriptomics (ST) has facilitated the spatially-aware profiling of gene expressions based on histopathology images. Although ST data offers valuable insights into the micro-environment of tumors, its acquisition cost remains expensive. Therefore, directly predicting the…

2025

MiniVLN: Efficient Vision-and-Language Navigation by Progressive Knowledge Distillation

ICRA 2025

In recent years, Embodied Artificial Intelligence (Embodied AI) has advanced rapidly, yet the increasing size of models conflicts with the limited computational capabilities of Embodied AI platforms. To address this challenge, we aim to achieve both high model performance and practical deployability

Cited by 5SourceScholar
2025

Model Merging in Pre-training of Large Language Models

NeurIPS 2025poster

Model merging has emerged as a promising technique for enhancing large language models, though its application in large-scale pre-training remains relatively unexplored. In this paper, we present a comprehensive investigation of model merging techniques during the pre-training process. Through exten…

Cited by 0SourceScholar
2025

MotionCtrl: A Real-time Controllable Vision-Language-Motion Model

ICCV 2025poster

Human motion generation involves synthesizing coherent human motion sequences conditioned on diverse multimodal inputs and holds significant potential for real-world applications. Despite recent advancements, existing vision-language-motion models (VLMMs) remain limited in achieving this goal. In th…

2025

Multi-modal Salient Object Detection via a Unified Diffusion Model

ICASSP 2025accepted

Salient Object Detection (SOD) aims to identify and segment the most striking elements within an image. Salient object detection methods can be differentiated into several types according to the input data, such as RGB-D (Depth) and RGB-T (Thermal). Previous research primarily focused on saliency de…

Cited by 0SourceScholar
2025

Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

ICLR 2025poster

Video understanding is a crucial next step for multimodal large language models (MLLMs). Various benchmarks are introduced for better evaluating the MLLMs. Nevertheless, current video benchmarks are still inefficient for evaluating video models during iterative development due to the high cost of co…

2025

Numerical Pruning for Efficient Autoregressive Models

AAAI 2025technical

Transformers have emerged as the leading architecture in deep learning, proving to be versatile and highly effective across diverse domains beyond language and image processing. However, their impressive performance often incurs high computational costs due to their substantial model size. This pape…

Cited by 10SourcePDFScholar
2025

OmniFC: Rethinking Federated Clustering via Lossless and Secure Distance Reconstruction

NeurIPS 2025poster

Federated clustering (FC) aims to discover global cluster structures across decentralized clients without sharing raw data, making privacy preservation a fundamental requirement. There are two critical challenges: (1) privacy leakage during collaboration, and (2) robustness degradation due to aggreg…

Cited by 0SourceScholar
2025

QuartDepth: Post-Training Quantization for Real-Time Depth Estimation on the Edge

CVPR 2025poster

Monocular Depth Estimation (MDE) has emerged as a pivotal task in computer vision, supporting numerous real-world applications. However, deploying accurate depth estimation models on resource-limited edge devices, especially Application-Specific Integrated Circuits (ASICs), is challenging due to the…

2025

Rapid Autonomous Exploration of Large-Scale Environments for Ground Robots Based on Region Partitioning

ICRA 2025

Autonomous exploration in large environments often leads to inefficient long backtracking, as distant targets are prioritized over closer ones. In this work, a hierarchical planning method is proposed, which employs region partitioning to systematically address the aforementioned issue. The space is

Cited by 1SourceScholar
2025

RoleBreak: Character Hallucination as a Jailbreak Attack in Role-Playing Systems

COLING 2025main

Role-playing systems powered by large language models (LLMs) have become increasingly influential in emotional communication applications. However, these systems are susceptible to character hallucinations, where the model deviates from predefined character roles and generates responses that are inc…

2025

Scaling Omni-modal Pretraining with Multimodal Context: Advancing Universal Representation Learning Across Modalities

ICCV 2025poster

This work introduces Multimodal Context (MiCo), a scalable pretraining framework designed to advance omni-modal intelligence--an AI system capable of understanding and learning from multiple modalities to achieve universal representation learning. MiCo allows for efficient scaling of both the number…

2025

Seg-diffusion: Text-to-Image Diffusion Model for Open-Vocabulary Semantic Segmentation

ICASSP 2025accepted

Open-vocabulary semantic segmentation (OVSS) is a challenging computer vision task that labels each pixel within an image based on text descriptions. Recent advancements in OVSS are largely attributed to the increased model capacity. However, these models often struggle with unfamiliar images or uns…

Cited by 0SourceScholar
2025

SharpZO: Hybrid Sharpness-Aware Vision Language Model Prompt Tuning via Forward-Only Passes

NeurIPS 2025poster

Fine-tuning vision language models (VLMs) has achieved remarkable performance across various downstream tasks; yet, it requires access to model gradients through backpropagation (BP), making them unsuitable for memory-constrained, inference-only edge devices. To address this limitation, previous wo…

Cited by 0SourcecodeScholar
2025

TRAIL: Trust-Aware Client Scheduling for Semi-Decentralized Federated Learning

AAAI 2025technical

Due to the sensitivity of data, Federated Learning (FL) is employed to enable distributed machine learning while safeguarding data privacy and accommodating the requirements of various devices. However, in the context of semidecentralized FL, clients’ communication and training states are dynamic. T…

Cited by 0SourcePDFScholar
2025

TaskSimLF: Efficient Leader-Follower Multi-Agent Path Finding With Clustered Pickup and Delivery

RA-L 2025

Multi-Agent Path Finding (MAPF) aims at finding a set of conflict-free and cost-optimal paths for agents from pickup to delivery locations. Most existing MAPF research focus on exhaustively search for path set for the agents with conflict-free paths, which often results in high computational costs.

Cited by 1SourceScholar
2025

VRoPE: Rotary Position Embedding for Video Large Language Models

EMNLP 2025

Rotary Position Embedding (RoPE) has shown strong performance in text-based Large Language Models (LLMs), but extending it to video remains a challenge due to the intricate spatiotemporal structure of video frames. Existing adaptations, such as RoPE-3D, attempt to encode spatial and temporal dimensi

2025

ViPE: Visual Perception in Parameter Space for Efficient Video-Language Understanding

EMNLP 2025

Existing video-language models (Video-LLMs) typically rely on concatenating visual tokens with textual inputs for joint modeling. However, this token-level alignment leads to significant inefficiency, especially when scaling to long videos with dense visual inputs. In this work, we propose a video-t

Cited by 0SourcePDFScholar
2025

ZipVL: Accelerating Vision-Language Models through Dynamic Token Sparsity

ICCV 2025poster

The efficiency of large vision-language models (LVLMs) is constrained by the computational bottleneck of the attention mechanism during the prefill phase and the memory bottleneck of fetching the key-value (KV) cache in the decoding phase, particularly in scenarios involving high-resolution images o…

Cited by 0SourcePDFScholar
2024

Automated Loss function Search for Class-imbalanced Node Classification

ICML 2024poster

Class-imbalanced node classification tasks are prevalent in real-world scenarios. Due to the uneven distribution of nodes across different classes, learning high-quality node representations remains a challenging endeavor. The engineering of loss functions has shown promising potential in addressing…

Cited by 1SourcePDFScholar
2024

BASES: Large-scale Web Search User Simulation with Large Language Model based Agents

EMNLP 2024finding

Due to the excellent capacities of large language models (LLMs), it becomes feasible to develop LLM-based agents for reliable user simulation. Considering the scarcity and limit (e.g., privacy issues) of real user data, in this paper, we conduct large-scale user simulations for the web search scenar…

Cited by 15SourcePDFScholar
2024

Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with Human Intentions

ACL 2024findings

Visual grounding (VG) aims at locating the foreground entities that match the given natural language expression. Previous datasets and methods for classic VG task mainly rely on the prior assumption that the given expression must literally refer to the target object, which greatly impedes the practi…

2024

CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing

NeurIPS 2024poster

We introduce Condition-Aware Self-Supervised Learning Representation (CA-SSLR), a generalist conditioning model broadly applicable to various speech-processing tasks. Compared to standard fine-tuning methods that optimize for downstream models, CA-SSLR integrates language and speaker embeddings from…

Cited by 0SourcePDFScholar
2024

COSA: Concatenated Sample Pretrained Vision-Language Foundation Model

ICLR 2024poster

Due to the limited scale and quality of video-text training corpus, most vision-language foundation models employ image-text datasets for pretraining and primarily focus on modeling visually semantic representations while disregarding temporal semantic representations and correlations. To addres…

2024

ConEC: Earnings Call Dataset with Real-world Contexts for Benchmarking Contextual Speech Recognition

COLING 2024main

Knowing the particular context associated with a conversation can help improving the performance of an automatic speech recognition (ASR) system. For example, if we are provided with a list of in-context words or phrases — such as the speaker’s contacts or recent song playlists — during inference, w…

2024

DyHGDAT: Dynamic Hypergraph Dual Attention Network for multi-agent trajectory prediction

ICRA 2024poster

Modeling the interactions among agents based on their historical trajectories is key to precise multi-agent trajectory prediction. Hypergraph Convolutional Networks (HGCN) have become a proper choice for capturing high-order interactions among agents in this field. However, most existing works only…

Cited by 0SourceScholar
2024

EPL-VINS: Efficient Point-Line Fusion Visual-Inertial SLAM With LK-RG Line Tracking Method and 2-DoF Line Optimization

RA-L 2024

The performance of a visual SLAM system based on point features significantly diminishes in low-textured environments due to the challenges in extracting sufficient and reliable points. The fusion of line and point features improves SLAM system performance by providing additional visual constraints.

Cited by 16SourceScholar
2024

EfficientDM: Efficient Quantization-Aware Fine-Tuning of Low-Bit Diffusion Models

ICLR 2024spotlight

Diffusion models have demonstrated remarkable capabilities in image synthesis and related generative tasks. Nevertheless, their practicality for low-latency real-world applications is constrained by substantial computational costs and latency issues. Quantization is a dominant way to compress and ac…

2024

Feature-Constrained and Attention-Conditioned Distillation Learning for Visual Anomaly Detection

ICASSP 2024accepted

Visual anomaly detection in computer vision is an essential one-class classification and segmentation problem. The student-teacher (S-T) approach has proven effective in addressing this challenge. However, previous studies based on S-T underutilize the feature representations learned by the teacher…

Cited by 0SourceScholar
2024

Graph Disentangled Contrastive Learning with Personalized Transfer for Cross-Domain Recommendation

AAAI 2024technical

Cross-Domain Recommendation (CDR) has been proven to effectively alleviate the data sparsity problem in Recommender System (RS). Recent CDR methods often disentangle user features into domain-invariant and domain-specific features for efficient cross-domain knowledge transfer. Despite showcasing rob…

Cited by 18SourcePDFScholar
2024

Hot-Fixing Wake Word Recognition for End-to-End ASR Via Neural Model Reprogramming

ICASSP 2024accepted

This paper proposes two novel variants of neural reprogramming to enhance wake word recognition in streaming end-to-end ASR models without updating model weights. The first, "trigger-frame reprogramming", prepends the input speech feature sequence with the learned trigger-frames of the target wake w…

Cited by 4SourceScholar
2024

LLM as Copilot for Coarse-grained Vision-and-Language Navigation

ECCV 2024poster

"Vision-and-Language Navigation (VLN) involves guiding an agent through indoor environments using human-provided textual instructions. Coarse-grained VLN, with short and high-level instructions, has gained popularity as it closely mirrors real-world scenarios. However, a significant challenge is the…

Cited by 9SourcePDFScholar
2024

MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

NeurIPS 2024poster

A critical approach for efficiently deploying computationally demanding large language models (LLMs) is Key-Value (KV) caching. The KV cache stores key-value states of previously generated tokens, significantly reducing the need for repetitive computations and thereby lowering latency in autoregress…

Cited by 44SourcePDFScholar
2024

Pretrained Optimization Model for Zero-Shot Black Box Optimization

NeurIPS 2024poster

Zero-shot optimization involves optimizing a target task that was not seen during training, aiming to provide the optimal solution without or with minimal adjustments to the optimizer. It is crucial to ensure reliable and robust performance in various applications. Current optimizers often struggle…

2024

QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language Models

ICLR 2024poster

Large Language Models (LLMs) have demonstrated unparalleled efficacy in natural language processing. However, their high computational demands and memory overheads hinder their broad deployment. To address this, two quantization strategies emerge, including Quantization-Aware Training (QAT) and Post…

2024

REAR: A Relevance-Aware Retrieval-Augmented Framework for Open-Domain Question Answering

EMNLP 2024main

Considering the limited internal parametric knowledge, retrieval-augmented generation (RAG) has been widely used to extend the knowledge scope of large language models (LLMs). Despite the extensive efforts on RAG research, in existing methods, LLMs cannot precisely assess the relevance of retrieved…

2024

SC-Tune: Unleashing Self-Consistent Referential Comprehension in Large Vision Language Models

CVPR 2024poster

Recent trends in Large Vision Language Models (LVLMs) research have been increasingly focusing on advancing beyond general image understanding towards more nuanced object-level referential comprehension. In this paper we present and delve into the self-consistency capability of LVLMs a crucial aspec…

2024

Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question Answering

EMNLP 2024main

While large pre-trained visual-language models have shown promising results on traditional visual question answering benchmarks, it is still challenging for them to answer complex VQA problems which requires diverse world knowledge. Motivated by the research of retrieval-augmented generation in the…

2024

Self-Evaluation of Large Language Model based on Glass-box Features

EMNLP 2024finding

The proliferation of open-source Large Language Models (LLMs) underscores the pressing need for evaluation methods. Existing works primarily rely on external evaluators, focusing on training and prompting strategies. However, a crucial aspect – model-aware glass-box features – is overlooked. In this…

2024

Signed Graph Neural Ordinary Differential Equation for Modeling Continuous-Time Dynamics

AAAI 2024technical

Modeling continuous-time dynamics constitutes a foundational challenge, and uncovering inter-component correlations within complex systems holds promise for enhancing the efficacy of dynamic modeling. The prevailing approach of integrating graph neural networks with ordinary differential equations h…

2024

Soft Knowledge Prompt: Help External Knowledge Become a Better Teacher to Instruct LLM in Knowledge-based VQA

ACL 2024long

LLM has achieved impressive performance on multi-modal tasks, which have received ever-increasing research attention. Recent research focuses on improving prediction performance and reliability (e.g., addressing the hallucination problem). They often prepend relevant external knowledge to the input…

2024

TFMQ-DM: Temporal Feature Maintenance Quantization for Diffusion Models

CVPR 2024highlight

The Diffusion model a prevalent framework for image generation encounters significant challenges in terms of broad applicability due to its extended inference times and substantial memory requirements. Efficient Post-training Quantization (PTQ) is pivotal for addressing these issues in traditional m…

2024

Temporal Adaptive RGBT Tracking with Modality Prompt

AAAI 2024technical

RGBT tracking has been widely used in various fields such as robotics, surveillance processing, and autonomous driving. Existing RGBT trackers fully explore the spatial information between the template and the search region and locate the target based on the appearance matching results. However, the…

Cited by 32SourcePDFScholar
2024

Text Prompt with Normality Guidance for Weakly Supervised Video Anomaly Detection

CVPR 2024poster

Weakly supervised video anomaly detection (WSVAD) is a challenging task. Generating fine-grained pseudo-labels based on weak-label and then self-training a classifier is currently a promising solution. However since the existing methods use only RGB visual modality and the utilization of category te…

Cited by 35SourcePDFScholar
2024

The Promises and Pitfalls of Using Language Models to Measure Instruction Quality in Education

NAACL 2024long

Assessing instruction quality is a fundamental component of any improvement efforts in the education system. However, traditional manual assessments are expensive, subjective, and heavily dependent on observers’ expertise and idiosyncratic factors, preventing teachers from getting timely and frequen…

Cited by 6SourcePDFScholar
2024

Unveiling Parts Beyond Objects: Towards Finer-Granularity Referring Expression Segmentation

CVPR 2024poster

Referring expression segmentation (RES) aims at segmenting the foreground masks of the entities that match the descriptive natural language expression. Previous datasets and methods for classic RES task heavily rely on the prior assumption that one expression must refer to object-level targets. In t…

2024

ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification

NeurIPS 2024poster

KV cache stores key and value states from previous tokens to avoid re-computation, yet it demands substantial storage space, especially for long sequences. Adaptive KV cache compression seeks to discern the saliency of tokens, preserving vital information while aggressively compressing those of l…

2023

A Novel Efficient Multi-View Traffic-Related Object Detection Framework

ICASSP 2023accepted

With the rapid development of intelligent transportation system applications, a tremendous amount of multi-view video data has emerged to enhance vehicle perception. However, performing video analytics efficiently by exploiting the spatial-temporal redundancy from video data remains challenging. Acc…

Cited by 0SourceScholar
2023

A Survey on Efficient Training of Transformers

IJCAI 2023poster

Recent advances in Transformers have come with a huge requirement on computing resources, highlighting the importance of developing efficient training techniques to make Transformer training faster, at lower cost, and to higher accuracy by the efficient use of computation and memory resources. This…

2023

A Thorough Examination on Zero-shot Dense Retrieval

EMNLP 2023long findings

Recent years have witnessed the significant advance in dense retrieval (DR) based on powerful pre-trained language models (PLM). DR models have achieved excellent performance in several benchmark datasets, while they are shown to be not as competitive as traditional sparse retrieval models (e.g., BM…

Cited by 0SourceScholar
2023

AIDE: A Vision-Driven Multi-View, Multi-Modal, Multi-Tasking Dataset for Assistive Driving Perception

ICCV 2023poster

Driver distraction has become a significant cause of severe traffic accidents over the past decade. Despite the growing development of vision-driven driver monitoring systems, the lack of comprehensive perception datasets restricts road safety and traffic security. In this paper, we present an AssIs…

Cited by 54PDFcodeScholar
2023

Anomalous Signal Detection for Cyber-Physical Systems Using Interpretable Causal Neural Network

ICASSP 2023accepted

Anomalous signal detection aims to detect unknown abnormal signals of machines from normal signals. However, building effective and interpretable anomaly detection models for safety-critical cyber-physical systems (CPS) is rather difficult due to the unidentified system noise and extremely intricate…

Cited by 0SourceScholar
2023

BiViT: Extremely Compressed Binary Vision Transformers

ICCV 2023poster

Model binarization can significantly compress model size, reduce energy consumption, and accelerate inference through efficient bit-wise operations. Although binarizing convolutional neural networks have been extensively studied, there is little work on exploring binarization of vision Transformers…

Cited by 44PDFScholar
2023

Boosting Verified Training for Robust Image Classifications via Abstraction

CVPR 2023poster

This paper proposes a novel, abstraction-based, certified training method for robust image classifiers. Via abstraction, all perturbed images are mapped into intervals before feeding into neural networks for training. By training on intervals, all the perturbed images that are mapped to the same int…

2023

Dual-Attention Neural Transducers for Efficient Wake Word Spotting in Speech Recognition

ICASSP 2023accepted

We present dual-attention neural biasing, an architecture designed to boost Wake Words (WW) recognition and improve inference time latency on speech recognition tasks. This architecture enables a dynamic switch for its runtime compute paths by exploiting WW spotting to select which branch of its att…

Cited by 6SourceScholar
2023

Dynamic Focus-Aware Positional Queries for Semantic Segmentation

CVPR 2023poster

The DETR-like segmentors have underpinned the most recent breakthroughs in semantic segmentation, which end-to-end train a set of queries representing the class prototypes or target segments. Recently, masked attention is proposed to restrict each query to only attend to the foreground regions predi…

2023

GAN-Based Robust Motion Planning for Mobile Robots Against Localization Attacks

RA-L 2023

Motion planning (MP) is essential but challenging for mobile robots. Most of the existing MP methods, at each instant, compute an action based on the states of the robot and the surrounding obstacles, assuming that the robot's localization module is attack-free. Unfortunately, the localization modul

Cited by 9SourceScholar
2023

GLOBER: Coherent Non-autoregressive Video Generation via GLOBal Guided Video DecodER

NeurIPS 2023poster

Video generation necessitates both global coherence and local realism. This work presents a novel non-autoregressive method GLOBER, which first generates global features to obtain comprehensive global guidance and then synthesizes video frames based on the global features to generate coherent videos…

2023

How2comm: Communication-Efficient and Collaboration-Pragmatic Multi-Agent Perception

NeurIPS 2023poster

Multi-agent collaborative perception has recently received widespread attention as an emerging application in driving scenarios. Despite the advancements in previous efforts, challenges remain due to various noises in the perception procedure, including communication redundancy, transmission delay,…

2023

Less Learn Shortcut: Analyzing and Mitigating Learning of Spurious Feature-Label Correlation

IJCAI 2023poster

Recent research has revealed that deep neural networks often take dataset biases as a shortcut to make decisions rather than understand tasks, leading to failures in real-world applications. In this study, we focus on the spurious correlation between word features and labels that models learn from t…

2023

LoTE-Animal: A Long Time-span Dataset for Endangered Animal Behavior Understanding

ICCV 2023poster

Understanding and analyzing animal behavior is increasingly essential to protect endangered animal species. However, the application of advanced computer vision techniques in this regard is minimal, which boils down to lacking large and diverse datasets for training deep models. To break the deadloc…

Cited by 18PDFcodeScholar
2023

MOSO: Decomposing MOtion, Scene and Object for Video Prediction

CVPR 2023poster

Motion, scene and object are three primary visual components of a video. In particular, objects represent the foreground, scenes represent the background, and motion traces their dynamics. Based on this insight, we propose a two-stage MOtion, Scene and Object decomposition framework (MOSO) for video…

2023

MSN-net: Multi-Scale Normality Network for Video Anomaly Detection

ICASSP 2023accepted

Existing unsupervised video anomaly detection methods often suffer from performance degradation due to the overgeneralization of deep models. In this paper, we propose a simple yet effective Multi-Scale Normality network (MSN-net) that uses hierarchical memories to learn multi-level prototypical spa…

Cited by 0SourceScholar
2023

March in Chat: Interactive Prompting for Remote Embodied Referring Expression

ICCV 2023poster

Many Vision-and-Language Navigation (VLN) tasks have been proposed in recent years, from room-based to object-based and indoor to outdoor. The REVERIE (Remote Embodied Referring Expression) is interesting since it only provides high-level instructions to the agent, which are closer to human commands…

Cited by 39PDFcodeScholar
2023

OmniAvatar: Geometry-Guided Controllable 3D Head Synthesis

CVPR 2023poster

We present OmniAvatar, a novel geometry-guided 3D head synthesis model trained from in-the-wild unstructured images that is capable of synthesizing diverse identity-preserved 3D heads with compelling dynamic details under full disentangled control over camera poses, facial expressions, head shapes,…

Cited by 29SourcePDFScholar
2023

PTQD: Accurate Post-Training Quantization for Diffusion Models

NeurIPS 2023poster

Diffusion models have recently dominated image synthesis and other related generative tasks. However, the iterative denoising process is expensive in computations at inference time, making diffusion models less practical for low-latency and scalable real-world applications. Post-training quantizati…

2023

Procter: Pronunciation-Aware Contextual Adapter For Personalized Speech Recognition In Neural Transducers

ICASSP 2023accepted

End-to-End (E2E) automatic speech recognition (ASR) systems used in voice assistants often have difficulties recognizing infrequent words personalized to the user, such as names and places. Rare words often have non-trivial pronunciations, and in such cases, human knowledge in the form of a pronunci…

Cited by 16SourceScholar
2023

Robust Acoustic And Semantic Contextual Biasing In Neural Transducers For Speech Recognition

ICASSP 2023accepted

Attention-based contextual biasing approaches have shown significant improvements in the recognition of generic and/or personal rare-words in End-to-End Automatic Speech Recognition (E2E ASR) systems like neural transducers. These approaches employ crossattention to bias the model towards specific c…

Cited by 24SourceScholar
2023

Spatio-Temporal Domain Awareness for Multi-Agent Collaborative Perception

ICCV 2023poster

Multi-agent collaborative perception as a potential application for vehicle-to-everything communication could significantly improve the perception performance of autonomous vehicles over single-agent perception. However, several challenges remain in achieving pragmatic information sharing in this em…

Cited by 68PDFcodeScholar
2023

SwiftAvatar: Efficient Auto-Creation of Parameterized Stylized Character on Arbitrary Avatar Engines

AAAI 2023technical

The creation of a parameterized stylized character involves careful selection of numerous parameters, also known as the "avatar vectors" that can be interpreted by the avatar engine. Existing unsupervised avatar vector estimation methods that auto-create avatars for users, however, often fail to wor…

2023

TOME: A Two-stage Approach for Model-based Retrieval

ACL 2023long

Recently, model-based retrieval has emerged as a new paradigm in text retrieval that discards the index in the traditional retrieval model and instead memorizes the candidate corpora using model parameters. This design employs a sequence-to-sequence paradigm to generate document identifiers, which e…

2023

VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset

NeurIPS 2023poster

Vision and text have been fully explored in contemporary video-text foundational models, while other modalities such as audio and subtitles in videos have not received sufficient attention. In this paper, we resort to establish connections between multi-modality video tracks, including Vision, Audi…

2023

Video Event Restoration Based on Keyframes for Video Anomaly Detection

CVPR 2023poster

Video anomaly detection (VAD) is a significant computer vision problem. Existing deep neural network (DNN) based VAD methods mostly follow the route of frame reconstruction or frame prediction. However, the lack of mining and learning of higher-level visual features and temporal context relationship…

Cited by 112SourcePDFScholar
2023

WL-MSR: Watch and Listen for Multimodal Subtitle Recognition

ICASSP 2023accepted

Video subtitles could be defined as the combination of visualized subtitles in frames and textual content recognized from speech, which play a significant role in video understanding for both humans and machines. In this paper, we propose a novel Watch and Listen for Multimodal Subtitle Recognition…

Cited by 0SourceScholar
2022

CoPur: Certifiably Robust Collaborative Inference via Feature Purification

NeurIPS 2022accept

Collaborative inference leverages diverse features provided by different agents (e.g., sensors) for more accurate inference. A common setup is where each agent sends its embedded features instead of the raw data to the Fusion Center (FC) for joint prediction. In this setting, we consider the inferen…

Cited by 11SourcePDFScholar
2022

Contextual Adapters for Personalized Speech Recognition in Neural Transducers

ICASSP 2022accepted

Personal rare word recognition in end-to-end Automatic Speech Recognition (E2E ASR) models is a challenge due to the lack of training data. A standard way to address this issue is with shallow fusion methods at inference time. However, due to their dependence on external language models and the dete…

Cited by 0SourceScholar
2022

DuQM: A Chinese Dataset of Linguistically Perturbed Natural Questions for Evaluating the Robustness of Question Matching Models

EMNLP 2022main

In this paper, we focus on the robustness evaluation of Chinese Question Matching (QM) models. Most of the previous work on analyzing robustness issues focus on just one or a few types of artificial adversarial examples. Instead, we argue that a comprehensive evaluation should be conducted on natura…

2022

DuReader-Retrieval: A Large-scale Chinese Benchmark for Passage Retrieval from Web Search Engine

EMNLP 2022main

In this paper, we present DuReader-retrieval, a large-scale Chinese dataset for passage retrieval. DuReader-retrieval contains more than 90K queries and over 8M unique passages from a commercial search engine. To alleviate the shortcomings of other datasets and ensure the quality of our benchmark, w…

2022

DuReadervis: A Chinese Dataset for Open-domain Document Visual Question Answering

ACL 2022findings

Open-domain question answering has been used in a wide range of applications, such as web search and enterprise search, which usually takes clean texts extracted from various formats of documents (e.g., web pages, PDFs, or Word documents) as the information source. However, designing different text…

2022

Dynamic Local Aggregation Network with Adaptive Clusterer for Anomaly Detection

ECCV 2022poster

"Existing methods for anomaly detection based on memory-augmented autoencoder (AE) have the following drawbacks: (1) Establishing a memory bank requires additional memory space. (2) The fixed number of prototypes from subjective assumptions ignores the data feature differences and diversity. To over…

2022

EcoFormer: Energy-Saving Attention with Linear Complexity

NeurIPS 2022accept

Transformer is a transformative framework for deep learning which models sequential data and has achieved remarkable performance on a wide range of tasks, but with high computational and energy cost. To improve its efficiency, a popular choice is to compress the models via binarization which constra…

2022

Learning Task-Specific Representation for Video Anomaly Detection with Spatial-Temporal Attention

ICASSP 2022accepted

The automatic detection of abnormal events in surveillance videos with weak supervision has been formulated as a multiple instance learning task, which aims to localize the clips containing abnormal events temporally with the video-level labels. However, most existing methods rely on the features ex…

Cited by 0SourceScholar
2022

Less Is More: Pay Less Attention in Vision Transformers

AAAI 2022technical

Transformers have become one of the dominant architectures in deep learning, particularly as a powerful alternative to convolutional neural networks (CNNs) in computer vision. However, Transformer training and inference in previous works can be prohibitively expensive due to the quadratic complexity…

2022

Look, Listen and Pay More Attention: Fusing Multi-Modal Information for Video Violence Detection

ICASSP 2022accepted

Violence detection is an essential and challenging problem in the computer vision community. Most existing works focus on single modal data analysis, which is not effective when multi-modality is available. Therefore, we propose a two-stage multi-modal information fusion method for violence detectio…

Cited by 0SourceScholar
2022

Multi-Task RNN-T with Semantic Decoder for Streamable Spoken Language Understanding

ICASSP 2022accepted

End-to-end Spoken Language Understanding (E2E SLU) has attracted increasing interest due to its advantages of joint optimization and low latency when compared to traditionally cascaded pipelines. Existing E2E SLU models usually follow a two-stage configuration where an Automatic Speech Recognition (…

Cited by 0SourceScholar
2021

A Novel Method to Solve Neural Knapsack Problems

ICML 2021spotlight

0-1 knapsack is of fundamental importance across many fields. In this paper, we present a game-theoretic method to solve 0-1 knapsack problems (KPs) where the number of items (products) is large and the values of items are not predetermined but decided by an external value assignment function (e.g.,…

Cited by 10SourcePDFScholar
2021

Consistent-Separable Feature Representation for Semantic Segmentation

AAAI 2021technical

Cross-entropy loss combined with softmax is one of the most commonly used supervision components in most existing segmentation methods. The softmax loss is typically good at optimizing the inter-class difference, but not good at reducing the intra-class variation, which can be suboptimal for semanti…

Cited by 3SourcePDFScholar
2021

DuReader_robust: A Chinese Dataset Towards Evaluating Robustness and Generalization of Machine Reading Comprehension in Real-World Applications

ACL 2021short

Machine reading comprehension (MRC) is a crucial task in natural language processing and has achieved remarkable advancements. However, most of the neural MRC models are still far from robust and fail to generalize well in real-world applications. In order to comprehensively verify the robustness an…

2021

HAIR: Hierarchical Visual-Semantic Relational Reasoning for Video Question Answering

ICCV 2021poster

Relational reasoning is at the heart of video question answering. However, existing approaches suffer from several common limitations: (1) they only focus on either object-level or frame-level relational reasoning, and fail to integrate the both; and (2) they neglect to leverage semantic knowledge f…

Cited by 62PDFcodeScholar
2021

Measuring Conversational Uptake: A Case Study on Student-Teacher Interactions

ACL 2021long

In conversation, uptake happens when a speaker builds on the contribution of their interlocutor by, for example, acknowledging, repeating or reformulating what they have said. In education, teachers’ uptake of student contributions has been linked to higher student achievement. Yet measuring and imp…

2021

RocketQA: An Optimized Training Approach to Dense Passage Retrieval for Open-Domain Question Answering

NAACL 2021long

In open-domain question answering, dense passage retrieval has become a new paradigm to retrieve relevant passages for finding answers. Typically, the dual-encoder architecture is adopted to learn dense representations of questions and passages for semantic matching. However, it is difficult to effe…

2021

RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-ranking

EMNLP 2021main

In various natural language processing tasks, passage retrieval and passage re-ranking are two key procedures in finding and ranking relevant information. Since both the two procedures contribute to the final performance, it is important to jointly optimize them in order to achieve mutual improvemen…

2021

Scalable Vision Transformers With Hierarchical Pooling

ICCV 2021poster

The recently proposed Visual image Transformers (ViT) with pure attention have achieved promising performance on image recognition tasks, such as image classification. However, the routine of the current ViT model is to maintain a full-length patch sequence during inference, which is redundant and l…

Cited by 186PDFcodeScholar
2020

Generative Low-bitwidth Data Free Quantization

ECCV 2020poster

Neural network quantization is an effective way to compress deep models and improve their execution latency and energy efficiency, so that they can be deployed on mobile or embedded devices. Existingquantization methods require original data for calibration or fine-tuning to get better performance.…

2020

Latent Regularized Generative Dual Adversarial Network For Abnormal Detection

IJCAI 2020poster

With the development of adversarial attack in deep learning, it is critical for abnormal detector to not only discover the out-of-distribution samples but also provide defence against the adversarial attacker. Since few previous universal detector is known to work well on both tasks, we consider aga…

Cited by 0SourcePDFScholar
2020

Learning Progressive Joint Propagation for Human Motion Prediction

ECCV 2020poster

Despite the great progress in human motion prediction, it remains a challenging task due to the complicated structural dynamics of human behaviors. In this paper, we address this problem in three aspects. First, to capture the long-range spatial correlations and temporal dependencies, we apply a tra…

Cited by 197SourcePDFScholar
2020

Non-Autoregressive Image Captioning with Counterfactuals-Critical Multi-Agent Learning

IJCAI 2020poster

Most image captioning models are autoregressive, i.e. they generate each word by conditioning on previously generated words, which leads to heavy latency during inference. Recently, non-autoregressive decoding has been proposed in machine translation to speed up the inference time by generating all…

Cited by 0SourcePDFScholar
2020

Normalized and Geometry-Aware Self-Attention Network for Image Captioning

CVPR 2020poster

Self-attention (SA) network has shown profound value in image captioning. In this paper, we improve SA from two aspects to promote the performance of image captioning. First, we propose Normalized Self-Attention (NSA), a reparameterization of SA that brings the benefits of normalization inside SA. W…

Cited by 284PDFScholar
2020

Not only Look, but also Listen: Learning Multimodal Violence Detection under Weak Supervision

ECCV 2020poster

but also Listen: Learning Multimodal Violence Detection under Weak Supervision","Violence detection has been studied in computer vision for years. However, previous work are either superficial, e.g., classification of short-clips, and the single scenario, or undersupplied, e.g., the single modality,…

2020

Residual Attention Network for Wavelet Domain Super-Resolution

ICASSP 2020accepted

Single-image super-resolution plays an important role in computer vision area. However, previous works using convolutional neural networks perform badly when reconstructing high frequency details, result in over-smooth and lacking of textural information in the output. At the same time, super-resolu…

Cited by 0SourceScholar
2018

Deep Uniqueness-Aware Hashing for Fine-Grained Multi-Label Image Retrieval

ICASSP 2018accepted

Deep supervised hashing methods for multi-label image retrieval have achieved great success nowadays. However, these methods only take the similarity between the database images and the query images into account, but they ignore the uniqueness of the database images when deciding on their rankings.…

Cited by 0SourceScholar
2018

Discrimination-aware Channel Pruning for Deep Neural Networks

NeurIPS 2018poster

Channel pruning is one of the predominant approaches for deep model compression. Existing pruning methods either train from scratch with sparsity constraints on channels, or minimize the reconstruction error between the pre-trained feature maps and the compressed ones. Both strategies suffer from s…

2017

A-Lamp: Adaptive Layout-Aware Multi-Patch Deep Convolutional Neural Network for Photo Aesthetic Assessment

CVPR 2017poster

Deep convolutional neural networks (CNN) have recently been shown to generate promising results for aesthetics assessment. However, the performance of these deep CNN methods is often compromised by the constraint that the neural network only takes the fixed-size input. To accommodate this requiremen…

Cited by 260PDFScholar
2016

Principal components analysis-based visual saliency detection

ICASSP 2016accepted

In this paper, a novel patch-wise saliency detection algorithm is proposed based on Principal Component Analysis (PCA). As a powerful statistical procedure in data analysis, PCA are fully exploited to convert color space and produce compact patch representation. Specifically, images are first conver…

Cited by 0SourceScholar