← Search

Jun Liu

229 accepted papers

2026

AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection

AAAI 2026technical

Universal visual anomaly detection aims to identify anomalies from novel or unseen vision domains without additional fine-tuning, which is critical in open scenarios. Recent studies have demonstrated that pre-trained vision-language models like CLIP exhibit strong generalization with just zero or a

Cited by 0SourcePDFScholar
2026

Beyond Layer-Wise Merging: Chain-of-Merging for Vision-Language Models

CVPR 2026

While model merging has demonstrated remarkable success for large language models (LLMs), its application to vision-language models (VLMs) remains largely underexplored. Recent methods attempt to enhance VLM reasoning capabilities by integrating specialized LLM parameters through layer-wise merging.

Cited by 0SourceScholar
2026

Beyond Sample-Level Forgetting: Improving Reliability in Multimodal Unlearning

ICML 2026poster

Multimodal unlearning aims to eliminate specific data from pretrained multimodal models, which offers significant advantages in data privacy and model efficiency. Current methods struggle to achieve the desired properties of effectiveness, reliability and locality, due to the complex interdependency…

Cited by 0SourceScholar
2026

Counterfactual Occlusion-Aware Learning via Visibility Intervention for LiDAR Anomaly Detection

ICML 2026poster

LiDAR point cloud anomaly detection is critical for autonomous system safety, yet most existing methods rely only on visible measurements, overlooking occlusion as a structured consequence of the LiDAR sensing process. We argue that anomalies are characterized not only by what is observed, but also …

Cited by 0SourceScholar
2026

DiffGraph: An Automated Agent-driven Model Merging Framework for In-the-Wild Text-to-Image Generation

CVPR 2026

The rapid growth of the text-to-image (T2I) community has fostered a thriving online ecosystem of expert models, which are variants of pretrained diffusion models specialized for diverse generative capabilities. Yet, existing model merging methods remain limited in fully leveraging abundant online e

Cited by 0SourceScholar
2026

Experience Transfer for Multimodal LLM Agents in Minecraft Game

CVPR 2026

Multimodal LLM agents operating in complex game environments must continually reuse past experience to solve new tasks efficiently. In this work, we propose Echo, a transfer-oriented memory framework that enables agents to derive actionable knowledge from prior interactions rather than treating memo

Cited by 0SourceScholar
2026

Exploring Category-level Articulated Object Pose Tracking on SE(3) Manifolds

AAAI 2026technical

Articulated objects are prevalent in daily life and robotic manipulation tasks. However, compared to rigid objects, pose tracking for articulated objects remains an underexplored problem due to their inherent kinematic constraints. To address these challenges, this work proposes a novel point-pair-b

Cited by 0SourcePDFScholar
2026

From Detection to Diagnosis: Advancing Hallucination Analysis with Automated Data Synthesis

AAAI 2026technical

Hallucinations in Large Language Models (LLMs), defined as the generation of content inconsistent with facts or context, represent a core obstacle to their reliable deployment in critical domains. Current research primarily focuses on binary "detection" approaches that, while capable of identifying

Cited by 0SourcePDFScholar
2026

GuidedBridge: Training-freely Improving Bridge Models with Prior Guidance

ICML 2026poster

Guidance methods, e.g., classifier-free guidance (CFG) and auto-guidance (AG), have distinctively improved noise-to-data diffusion generation results. Recently, bridge models have been proposed, which present a data-to-data sampling process to exploit instructive information from clean prior represe…

Cited by 0SourceScholar
2026

MAPS: Multi-Agent Personality Shaping for Collaborative Reasoning

AAAI 2026technical

Collaborative reasoning with multiple agents offers the potential for more robust and diverse problem-solving. However, existing approaches often suffer from homogeneous agent behaviors and lack of reflective and rethinking capabilities. We propose Multi-Agent Personality Shaping ((MAPS), a novel fr

Cited by 0SourcePDFScholar
2026

MARS: Multi-Agent Adaptive Reasoning with Socratic Guidance for Automated Prompt Optimization

AAAI 2026technical

Large language models (LLMs) typically operate in a question-answering paradigm, where the quality of the input prompt critically affects the response. Automated Prompt Optimization (APO) aims to overcome the cognitive biases of manually crafted prompts and explore a broader prompt design space. How

Cited by 0SourcePDFScholar
2026

Near-Field Driven Origami-Based Bio-Inspired Jellyfish Robot

ICRA 2026poster

The development of bio-inspired jellyfish robots holds significant benefits for autonomous aquatic systems due to jellyfish’s efficient water jet propulsion. However, the current design of jellyfish robots still faces challenges in balancing high biological fidelity with the demands of lightweight, …

Cited by 0Scholar
2026

Neurodynamics-Driven Coupled Neural P Systems for Multi-Focus Image Fusion

CVPR 2026

Multi-focus image fusion (MFIF) is a crucial technique in image processing, with a key challenge being the generation of decision maps with precise boundaries. However, traditional methods based on heuristic rules and deep learning methods with black-box networks are difficult to generate high-quali

Cited by 0SourcecodeScholar
2026

PerfGuard: A Performance-Aware Agent for Visual Content Generation

ICLR 2026poster

The advancement of Large Language Model (LLM)-powered agents has enabled automated task processing through reasoning and tool invocation capabilities. However, existing frameworks often operate under the idealized assumption that tool executions are invariably successful, relying solely on textual d…

Cited by 0SourcecodeScholar
2026

Roots Beneath the Cut: Uncovering the Risk of Concept Revival in Pruning-Based Unlearning for Diffusion Models

CVPR 2026

Pruning-based unlearning has recently emerged as a fast, training-free, and data-independent approach to remove undesired concepts from diffusion models. It promises high efficiency and robustness, offering an attractive alternative to traditional fine-tuning or editing-based unlearning. However, in

Cited by 0SourcecodeScholar
2026

SAGE: A Dataflow-Native Framework for Modular, Controllable, and Transparent LLM-Augmented Reasoning

ICML 2026poster

LLM applications increasingly execute as end-to-end inference pipelines that couple generation with retrieval, stateful memory, context refinement, and tool use under strict tail-latency and SLO constraints. Today, these stages are often stitched together as RPC-connected services, obscuring cross-s…

Cited by 0SourceScholar
2026

SkelHCC: A Hyperbolic CLIP-Driven Cache Adaptation Framework for Skeleton-based One-Shot Action Recognition

ICML 2026poster

Skeleton-based action recognition aims to understand human behaviors from body joint sequences and is especially challenging in the one-shot setting, where only a single labeled exemplar is available for each novel action. A key challenge is learning representations that capture the hierarchical and…

Cited by 0SourceScholar
2026

SketchVL: Policy Optimization via Fine-Grained Credit Assignment for Chart Understanding and More

CVPR 2026

Charts are high-density visual carriers of complex data and medium for information extraction and analysis. Due to the need for precise and complex visual reasoning, automated chart understanding poses a significant challenge to existing Multimodal Large Language Models (MLLMs). Many MLLMs trained w

Cited by 0SourceScholar
2026

Translating Signals to Languages for sEMG-Based Activity Recognition

CVPR 2026

Surface electromyography (sEMG) signal-based activity recognition has attracted increasing research attention in recent years. To develop accurate sEMG signal-based activity recognizers, numerous approaches have been proposed. Some studies focus on designing larger and more expressive model architec

Cited by 0SourceScholar
2026

UniF$^2$ace: A $\underline{Uni}$fied $\underline{F}$ine-grained $\underline{Face}$ Understanding and Generation Model

ICLR 2026poster

Unified multimodal models (UMMs) have emerged as a powerful paradigm in fundamental cross-modality research, demonstrating significant potential in both image understanding and generation. However, existing research in the face domain primarily faces two challenges: **(1) fragmentation development**…

Cited by 0SourcecodeScholar
2026

WPT: World-to-Policy Transfer via Online World Model Distillation

CVPR 2026

Recent years have witnessed remarkable progress in world models, which primarily aim to capture the spatiotemporal correlations between an agent's actions and the evolving environment. However, existing approaches often suffer from tight runtime coupling or depend on offline reward signals, resultin

Cited by 0SourceScholar
2026

YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection

CVPR 2026

Existing Real-Time Object Detection (RTOD) methods commonly adopt YOLO-like architectures for their favorable trade-off between accuracy and speed. However, these models rely on static dense computation that applies uniform processing to all inputs, misallocating representational capacity and comput

Cited by 0SourcecodeScholar
2025

An Image-like Diffusion Method for Human-Object Interaction Detection

CVPR 2025poster

Human-object interaction (HOI) detection often faces high levels of ambiguity and indeterminacy, as the same interaction can appear vastly different across different human-object pairs. Additionally, the indeterminacy can be further exacerbated by issues such as occlusions and cluttered backgrounds.…

Cited by 0SourcePDFScholar
2025

Blind Noisy Image Deblurring Using Residual Guidance Strategy

ICCV 2025poster

Blind deblurring is an ill-posed inverse problem that involves recovering both the clear image and the blur kernel from a single blurry image. In real photography, longer exposure time results in lots of noise in the blurry image. Although existing blind deblurring methods produce satisfactory resul…

Cited by 0SourcePDFScholar
2025

Boosting Skeleton-based Zero-Shot Action Recognition with Training-Free Test-Time Adaptation

NeurIPS 2025poster

We introduce Skeleton-Cache, the first training-free test-time adaptation framework for skeleton-based zero-shot action recognition (SZAR), aimed at improving model generalization to unseen actions during inference. Skeleton-Cache reformulates inference as a lightweight retrieval process over a non-…

Cited by 0SourcecodeScholar
2025

Boundary Probing for Input Privacy Protection When Using LMM Services

ICCV 2025poster

Alongside the rapid development of Large Multimodal Models (LMMs) like GPT-4V, privacy concerns also rise. As LMMs are commonly deployed as cloud services, users are typically required to upload their personal images and videos to the cloud to access these services, raising great concerns about visu…

Cited by 0SourcePDFScholar
2025

CMMLoc: Advancing Text-to-PointCloud Localization with Cauchy-Mixture-Model Based Framework

CVPR 2025poster

The goal of point cloud localization based on linguistic description is to identify a 3D position using textual description in large urban environments, which has potential applications in various fields, such as determining the location for vehicle pickup or goods delivery. Ideally, for a textual d…

2025

CPCF: A Cross-Prompt Contrastive Framework for Referring Multimodal Large Language Models

ICML 2025poster

Referring MLLMs extend conventional multimodal large language models by allowing them to receive referring visual prompts and generate responses tailored to the indicated regions. However, these models often suffer from suboptimal performance due to incorrect responses tailored to misleading areas a…

Cited by 0SourcePDFScholar
2025

Causal-R: A Causal-Reasoning Geometry Problem Solver for Optimized Solution Exploration

NeurIPS 2025poster

The task of geometry problem solving has been a long-standing focus in the automated mathematics community and draws growing attention due to its complexity for both symbolic and neural models. Although prior studies have explored various effective approaches for enhancing problem solving performanc…

Cited by 0SourceScholar
2025

ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding

NeurIPS 2025poster

Charts are high-density visualization carriers for complex data, serving as a crucial medium for information extraction and analysis. Automated chart understanding poses significant challenges to existing multimodal large language models (MLLMs) due to the need for precise and complex visual reasoni…

Cited by 0SourcecodeScholar
2025

CoFFT: Chain of Foresight-Focus Thought for Visual Language Models

NeurIPS 2025poster

Despite significant advances in Vision Language Models (VLMs), they remain constrained by the complexity and redundancy of visual input. When images contain large amounts of irrelevant information, VLMs are susceptible to interference, thus generating excessive task-irrelevant reasoning processes or…

Cited by 0SourceScholar
2025

Convex Combination Star Shape Prior for Data-driven Image Semantic Segmentation

CVPR 2025poster

Multi-center star shape is a prevalent object shape feature, which has proven effective in model-based image segmentation methods. However, the shape field function induced by the multi-center star shape is non-smooth, and directly applying it to the data-driven image segmentation network architectu…

Cited by 0SourcePDFScholar
2025

Cross-modal Collaborative Representation Learning for Text-to-Image Person Retrieval

IJCAI 2025

Text-to-image person retrieval (TIPR) aims to find images of the same identity that match a given text description. Current TIPR methods mainly focus on mining the association between images and texts, ignoring their potential complementarity. Besides, existing matching losses treat all positive pai

Cited by 0SourcePDFScholar
2025

Debate on Graph: A Flexible and Reliable Reasoning Framework for Large Language Models

AAAI 2025technical

Large Language Models (LLMs) may suffer from hallucinations in real-world applications due to the lack of relevant knowledge. In contrast, knowledge graphs encompass extensive, multi-relational structures that store a vast array of symbolic facts. Consequently, integrating LLMs with knowledge graphs…

2025

Deconfound Semantic Shift and Incompleteness in Incremental Few-shot Semantic Segmentation

AAAI 2025technical

Incremental few-shot semantic segmentation (IFSS) expands segmentation capacity of the trained model to segment new-class images with few samples. However, semantic meanings may shift from background to object class or vice versa during incremental learning. Moreover, new-class samples often lack re…

Cited by 0SourcePDFScholar
2025

Deliberation on Priors: Trustworthy Reasoning of Large Language Models on Knowledge Graphs

NeurIPS 2025poster

Knowledge graph-based retrieval-augmented generation seeks to mitigate hallucinations in Large Language Models (LLMs) caused by insufficient or outdated knowledge. However, existing methods often fail to fully exploit the prior knowledge embedded in knowledge graphs (KGs), particularly their structu…

Cited by 0SourcecodeScholar
2025

DiffIP: Representation Fingerprints for Robust IP Protection of Diffusion Models

ICCV 2025poster

Intellectual property (IP) protection for diffusion models is a critical concern, given the significant resources and time required for their development. To effectively safeguard the IP of diffusion models, a key step is enabling the comparison of unique identifiers (fingerprints) between suspect a…

Cited by 0SourcePDFScholar
2025

Diffuse&Refine: Intrinsic Knowledge Generation and Aggregation for Incremental Object Detection

IJCAI 2025

Incremental Object Detection(IOD) targets at progressively extending capability of object detectors to recognize new classes. However, representation confusion between old and new classes leads to catastrophic forgetting. To alleviate this problem, we propose DiffKA, with intrinsic knowledge generat

Cited by 0SourcePDFScholar
2025

Diffusion Models are Good Unsupervised Class-agnostic Shape Part Segmentators

ICASSP 2025accepted

Shape part segmentation is a critical task in computer graphics and robotics. However, traditional supervised methods rely heavily on large amounts of labeled data, which poses significant challenges in many real-world scenarios where such data is often scarce or difficult to obtain. To address this…

Cited by 0SourceScholar
2025

Dual-Interrelated Diffusion Model for Few-Shot Anomaly Image Generation

CVPR 2025poster

The performance of anomaly inspection in industrial manufacturing is constrained by the scarcity of anomaly data. To overcome this challenge, researchers have started employing anomaly generation approaches to augment the anomaly dataset. However, existing anomaly generation methods suffer from limi…

2025

Enhanced 3D LiDAR Features TLG: Multi-Feature Fusion and LiDAR Inertial Odometry Applications

RA-L 2025

In order to address the shortcomings of limited point cloud feature description capability, insufficient real-time performance, and severe feature homogenization in the field of 3D LiDAR SLAM. In this letter, a novel feature extraction and matching method TLG is proposed. First, the method makes ful

Cited by 4SourceScholar
2025

EvoChart: A Benchmark and a Self-Training Approach Towards Real-World Chart Understanding

AAAI 2025technical

Chart understanding enables automated data analysis for humans, which requires models to achieve highly accurate visual comprehension. While existing Visual Language Models (VLMs) have shown progress in chart understanding, the lack of high-quality training data and comprehensive evaluation benchmar…

2025

Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models

CVPR 2025poster

Despite inheriting security measures from underlying language models, Vision-Language Models (VLMs) may still be vulnerable to safety alignment issues. Through empirical analysis, we uncover two critical findings: scenario-matched images can significantly amplify harmful outputs, and contrary to com…

Cited by 1SourcePDFScholar
2025

FairSMOE: Mitigating Multi-Attribute Fairness Problem with Sparse Mixture-of-Experts

IJCAI 2025

Real‐world datasets usually contain multiple attributes, making it essential to ensure fairness across all of them simultaneously. However, different attributes may vary in difficulty, and no existing approaches have effectively addressed this issue. Consequently, an attribute‐adaptive strategy is n

Cited by 0SourcePDFScholar
2025

GUICourse: From General Vision Language Model to Versatile GUI Agent

ACL 2025long

Utilizing Graphic User Interfaces (GUIs) for human-computer interaction is essential for accessing various digital tools. Recent advancements in Vision Language Models (VLMs) reveal significant potential for developing versatile agents that assist humans in navigating GUIs. However, current VLMs fac…

2025

GaussianBlock: Building Part-Aware Compositional and Editable 3D Scene by Primitives and Gaussians

ICLR 2025poster

Recently, with the development of Neural Radiance Fields and Gaussian Splatting, 3D reconstruction techniques have achieved remarkably high fidelity. However, the latent representations learnt by these methods are highly entangled and lack interpretability. In this paper, we propose a novel part-awa…

Cited by 0SourcePDFScholar
2025

Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced Reasoning

ACL 2025long

Advancing LLM reasoning skills has captivated wide interest. However, current post-training techniques rely heavily on supervisory signals, such as outcome supervision or auxiliary reward models, which face the problem of scalability and high annotation costs. This motivates us to enhance LLM reason…

2025

Global-Semantic Alignment Distillation for Partial Multi-view Classification

AAAI 2025technical

Partial multi-view classification (PMvC) poses a significant challenge due to the incomplete nature of multi-view data, which complicates effective information fusion and accurate classification. Existing PMvC methods typically rely on heuristic evaluations of view informativeness to achieve global…

Cited by 0SourcePDFScholar
2025

Harmony in Divergence: Towards Fast, Accurate, and Memory-efficient Zeroth-order LLM Fine-tuning

NeurIPS 2025poster

Large language models (LLMs) excel across various tasks, but standard first-order (FO) fine-tuning demands considerable memory, significantly limiting real-world deployment. Recently, zeroth-order (ZO) optimization stood out as a promising memory-efficient training paradigm, avoiding backward passes…

Cited by 0SourcecodeScholar
2025

Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression Comprehension

AAAI 2025technical

In this work, we address the challenging task of Generalized Referring Expression Comprehension (GREC). Compared to the classic Referring Expression Comprehension (REC) that focuses on single-target expressions, GREC extends the scope to a more practical setting by further encompassing no-target and…

Cited by 1SourcePDFScholar
2025

Hierarchical Divide-and-Conquer Grouping for Classification Adaptation of Pre-Trained Models

ICCV 2025poster

Existing adaptation methods of pre-trained vision-language models like CLIP often rely on base-class samples during fine-tuning, introducing systematic biases that distort decision boundaries and degrade performance on novel classes. In this work, we break new ground by proposing a hierarchical divi…

Cited by 0SourcePDFScholar
2025

Hierarchical Optimization via LLM-Guided Objective Evolution for Mobility-on-Demand Systems

NeurIPS 2025poster

Online ride-hailing platforms aim to deliver efficient mobility-on-demand services, often facing challenges in balancing dynamic and spatially heterogeneous supply and demand. Existing methods typically fall into two categories: reinforcement learning (RL) approaches, which suffer from data ineffici…

Cited by 0SourceScholar
2025

Interactive Evolution: A Neural-Symbolic Self-Training Framework For Large Language Models

ACL 2025long

One of the primary driving forces contributing to the superior performance of Large Language Models (LLMs) is the extensive availability of human-annotated natural language data, which is used for alignment fine-tuning. This inspired researchers to investigate self-training methods to mitigate the e…

2025

Making Every Step Effective: Jailbreaking Large Vision-Language Models Through Hierarchical KV Equalization

EMNLP 2025

In the realm of large vision-language models (LVLMs), adversarial jailbreak attacks serve as a red-teaming approach to identify safety vulnerabilities of these models and their associated defense mechanisms. However, we identify a critical limitation: not every adversarial optimization step leads to

Cited by 0SourcePDFScholar
2025

Metapath and Hypergraph Structure-based Multi-Channel Graph Contrastive Learning for Student Performance Prediction

IJCAI 2025

Considerable attention has been paid to predicting student performance on exercises. The performance of prior studies is determined by the quality of the trait features of students and exercises. Nevertheless, most of the prior study primarily examines simple pairwise interactions in learning trait

2025

Mutual Effort for Efficiency: A Similarity-based Token Pruning for Vision Transformers in Self-Supervised Learning

ICLR 2025poster

Self-supervised learning (SSL) offers a compelling solution to the challenge of extensive labeled data requirements in traditional supervised learning. With the proven success of Vision Transformers (ViTs) in supervised tasks, there is increasing interest in adapting them for SSL frameworks. However…

Cited by 0SourcePDFScholar
2025

Optimized Design and Calibration of a Human-Eye-Sized Active Binocular Vision System Based on Spherical Parallel Mechanism

RA-L 2025

The Active Binocular Vision System (ABVS), resembling the human eye, demonstrates potential for improving visual perception in robotic systems, especially in dynamic and complex environments. In this letter, we present an optimized design of a three degree-of-freedom (DoF) Active Monocular Vision Sy

Cited by 2SourceScholar
2025

POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation

CVPR 2025poster

Existing LVLM-based reasoning segmentation methods often suffer from imprecise segmentation results and hallucinations in their text responses. This paper introduces POPEN, a novel framework designed to address these issues and achieve improved results. POPEN includes a preference-based optimization…

Cited by 2SourcePDFScholar
2025

Performing Defocus Deblurring by Modeling its Formation Process

ICCV 2025poster

Single image defocus deblurring (SIDD) is a challenging task that aims to recover an all-in-focus image from a defocused one. In this paper, we make the observation that a defocused image can be viewed as a blend of illuminated blobs based on fundamental imaging principles, and the defocus blur in t…

Cited by 0SourcePDFScholar
2025

PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning

ACL 2025long

Large language models demonstrate remarkable capabilities across various domains, especially mathematics and logic reasoning. However, current evaluations overlook physics-based reasoning - a complex task requiring physics theorems and constraints. We present PhysReason, a 1,200-problem benchmark co…

Cited by 0SourcePDFScholar
2025

RMath: A Logic Reasoning-Focused Datasets Toward Mathematical Multistep Reasoning Tasks

AAAI 2025technical

Mathematical reasoning ability objectively reflects a language model's understanding of implicit knowledge in contexts, with logic being a prerequisite for exploring, articulating and establishing effective reasoning. Large language models (LLMs) have shown great potential in complex reasoning tasks…

2025

Recognizing Actions from Robotic View for Natural Human-Robot Interaction

ICCV 2025poster

Natural Human-Robot Interaction (N-HRI) requires robots to recognize human actions at varying distances and states, regardless of whether the robot itself is in motion or stationary. This setup is more flexible and practical than conventional human action recognition tasks. However, existing benchma…

2025

ResoFilter: Fine-grained Synthetic Data Filtering for Large Language Models through Data-Parameter Resonance Analysis

NAACL 2025findings

Large language models (LLMs) have shown remarkable effectiveness across various domains, with data augmentation methods utilizing GPT for synthetic data generation becoming prevalent. However, the quality and utility of augmented data remain questionable, and current methods lack clear metrics for e…

2025

RoRA: Efficient Fine-Tuning of LLM with Reliability Optimization for Rank Adaptation

ICASSP 2025accepted

Fine-tuning helps large language models (LLM) recover degraded information and enhance task performance. Although Low-Rank Adaptation (LoRA) is widely used and effective for fine-tuning, we have observed that its scaling factor can limit or even reduce performance as the rank size increases. To addr…

Cited by 0SourceScholar
2025

Semantic-oriented Visual Prompt Learning for Class Incremental Learning

ICASSP 2025accepted

Class-incremental learning (CIL) enables models to continuously learn new classes while addressing catastrophic forgetting. With the introduction of pre-trained models, new tuning paradigms have emerged for CIL. This paper revisits parameter-efficient fine-tuning (PEFT) methods in the context of inc…

Cited by 0SourceScholar
2025

Stray Intrusive Outliers-Based Feature Selection on Intra-Class Asymmetric Instance Distribution or Multiple High-Density Clusters

ICML 2025poster

For data with intra-class Asymmetric instance Distribution or Multiple High-density Clusters (ADMHC), outliers are real and have specific patterns for data classification, where the class body is necessary and difficult to identify. Previous Feature Selection (FS) methods score features based on all…

2025

Throwing Planning Diffusion: A Solution to Learning and Planning of Robotic Throwing

IROS 2025

Dynamic manipulation enables efficient interaction tasks, such as throwing, which rely on finding one or more high-quality trajectories from the initial state to the goal state. While model-free learning methods have been used to acquire efficient robot manipulation configurations, traditional plann

Cited by 0SourceScholar
2025

Toward Adaptive Large Language Models Structured Pruning via Hybrid-grained Weight Importance Assessment

AAAI 2025technical

Structured pruning for large language models (LLMs) has garnered significant academic interest due to its ability to efficiently compress and accelerate LLMs by eliminating redundant weight groups at a coarse-grained granularity. Current structured pruning methods for LLMs typically depend on a sing…

2025

Towards Explicit Geometry-Reflectance Collaboration for Generalized LiDAR Segmentation in Adverse Weather

CVPR 2025poster

Existing LiDAR semantic segmentation models often suffer from decreased accuracy when exposed to adverse weather conditions. Recent methods addressing this issue focus on enhancing training data through weather simulation or universal augmentation techniques. However, few works have studied the nega…

Cited by 0SourcePDFScholar
2025

Ultra Lightweight Singing Melody Extraction via Combination of Convolution and MLP

ICASSP 2025accepted

Singing melody extraction serves as an important foundation in the realm of music information retrieval (MIR). Although fully convolutional neural networks (CNNs) are commonly employed for singing melody extraction, they are constrained by inductive biases and face challenges in establishing long ra…

Cited by 0SourceScholar
2025

VA-AR: Learning Velocity-Aware Action Representations with Mixture of Window Attention

AAAI 2025technical

Action recognition is a crucial task in artificial intelligence, with significant implications across various domains. We initially perform a comprehensive analysis of seven prominent action recognition methods across five widely-used datasets. This analysis reveals a critical, yet previously overlo…

2025

VProChart: Answering Chart Question Through Visual Perception Alignment Agent and Programmatic Solution Reasoning

AAAI 2025technical

Charts are widely used for data visualization across various fields, including education, research, and business. Chart Question Answering (CQA) is an emerging task focused on the automatic interpretation and reasoning of data presented in charts. However, chart images are inherently difficult to in…

2025

fairGNN-WOD: Fair Graph Learning Without Complete Demographics

IJCAI 2025

Graph Neural Networks (GNNs) have excelled in diverse applications due to their outstanding predictive performance, yet they often overlook fairness considerations, prompting numerous recent efforts to address this societal concern. However, most fair GNNs assume complete demographics by design, whi

Cited by 0SourcePDFScholar
2025

𝜙-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and Exploitation

ACL 2025long

Inference-time optimization scales computation to derive deliberate reasoning steps for effective performance. While previous search-based strategies address the short-sightedness of auto-regressive generation, the vast search space leads to excessive exploration and insufficient exploitation. To st…

2024

A Magnetic Catheter With Force Sensing Capability Toward Interventional Surgery

RA-L 2024

Magnetically actuated medical instruments could greatly facilitate minimally invasive surgery (MIS). For example, onboard magnetic materials help catheters travel across tortuous lumens and reach difficult-to-access sites inside human bodies, guided by a controlled magnetic field (MF). However, perm

Cited by 4SourceScholar
2024

A Semantic Mention Graph Augmented Model for Document-Level Event Argument Extraction

COLING 2024main

Document-level Event Argument Extraction (DEAE) aims to identify arguments and their specific roles from an unstructured document. The advanced approaches on DEAE utilize prompt-based methods to guide pre-trained language models (PLMs) in extracting arguments from input documents. They mainly concen…

2024

Addressing Background Context Bias in Few-Shot Segmentation through Iterative Modulation

CVPR 2024poster

Existing few-shot segmentation methods usually extract foreground prototypes from support images to guide query image segmentation. However different background contexts of support and query images can cause their foreground features to be misaligned. This phenomenon known as background context bias…

Cited by 17SourcePDFScholar
2024

Automated Non-invasive Analysis of Motile Sperms Using Cross-scale Guidance Network

ICRA 2024poster

Unbiased measurement of sperm morphometric and motility parameters is essential for assessing fertility potential and guiding visual feedback for microrobotic manipulation. Automated analysis of multiple sperms and selection of an optimal sperm is crucial for in vitro fertilisation treatment such as…

Cited by 0SourceScholar
2024

COSMIC: Compress Satellite Image Efficiently via Diffusion Compensation

NeurIPS 2024poster

With the rapidly increasing number of satellites in space and their enhanced capabilities, the amount of earth observation images collected by satellites is exceeding the transmission limits of satellite-to-ground links. Although existing learned image compression solutions achieve remarkable perfor…

Cited by 1SourcePDFScholar
2024

CoG-DQA: Chain-of-Guiding Learning with Large Language Models for Diagram Question Answering

CVPR 2024poster

Diagram Question Answering (DQA) is a challenging task requiring models to answer natural language questions based on visual diagram contexts. It serves as a crucial basis for academic tutoring technical support and more practical applications. DQA poses significant challenges such as the demand for…

Cited by 6SourcePDFScholar
2024

Correlation Matching Transformation Transformers for UHD Image Restoration

AAAI 2024technical

This paper proposes UHDformer, a general Transformer for Ultra-High-Definition (UHD) image restoration. UHDformer contains two learning spaces: (a) learning in high-resolution space and (b) learning in low-resolution space. The former learns multi-level high-resolution features and fuses low-high fe…

2024

DifAttack: Query-Efficient Black-Box Adversarial Attack via Disentangled Feature Space

AAAI 2024technical

This work investigates efficient score-based black-box adversarial attacks with high Attack Success Rate (ASR) and good generalizability. We design a novel attack method based on a Disentangled Feature space, called DifAttack, which differs significantly from the existing ones operating over the ent…

2024

E-GPS: Explainable Geometry Problem Solving via Top-Down Solver and Bottom-Up Generator

CVPR 2024poster

Geometry Problem Solving has drawn growing attention recently due to its application prospects in intelligent education field. However existing methods are still inadequate to meet the needs of practical application suffering from the following limitations: 1) explainability is not ensured which is…

Cited by 5SourcePDFScholar
2024

Echoes of the Past: Boosting Long-tail Recognition via Reflective Learning

ECCV 2024oral

"In real-world scenarios, where knowledge distributions exhibit long-tail. Humans manage to master knowledge uniformly across imbalanced distributions, a feat attributed to their diligent practices of reviewing, summarizing, and correcting errors. Motivated by this learning process, we propose a nov…

2024

Few-Shot Anomaly-Driven Generation for Anomaly Classification and Segmentation

ECCV 2024poster

"Anomaly detection is a practical and challenging task due to the scarcity of anomaly samples in industrial inspection. Some existing anomaly detection methods address this issue by synthesizing anomalies with noise or external data. However, there is always a large semantic gap between synthetic an…

2024

Generated and Pseudo Content guided Prototype Refinement for Few-shot Point Cloud Segmentation

NeurIPS 2024spotlight

Few-shot 3D point cloud semantic segmentation aims to segment query point clouds with only a few annotated support point clouds. Existing prototype-based methods learn prototypes from the 3D support set to guide the segmentation of query point clouds. However, they encounter the challenge of low pro…

Cited by 1SourcePDFScholar
2024

InstructGIE: Towards Generalizable Image Editing

ECCV 2024poster

"Recent advances in image editing have been driven by the development of denoising diffusion models, marking a significant leap forward in this field. Despite these advances, the generalization capabilities of recent image editing approaches remain constrained. In response to this challenge, our stu…

2024

LA-LIO: Robust Localizability-Aware LiDAR-Inertial Odometry for Challenging Scenes

IROS 2024poster

Modern robotic systems are increasingly deployed in complex and diverse environments, and reliable localization under challenging conditions becomes crucial for the safe and efficient operation of these systems. The odometry based on LiDAR is prone to system collapse caused by computational divergen…

Cited by 1SourceScholar
2024

LAFA: Multimodal Knowledge Graph Completion with Link Aware Fusion and Aggregation

AAAI 2024technical

Recently, an enormous amount of research has emerged on multimodal knowledge graph completion (MKGC), which seeks to extract knowledge from multimodal data and predict the most plausible missing facts to complete a given multimodal knowledge graph (MKG). However, existing MKGC approaches largely ign…

Cited by 14SourcePDFScholar
2024

LEAD: Exploring Logit Space Evolution for Model Selection

CVPR 2024poster

The remarkable success of "pretrain-then-finetune" paradigm has led to a proliferation of available pre-trained models for vision tasks. This surge presents a significant challenge in efficiently choosing the most suitable pre-trained models for downstream tasks. The critical aspect of this challeng…

Cited by 0SourcePDFScholar
2024

LLaFS: When Large Language Models Meet Few-Shot Segmentation

CVPR 2024poster

This paper proposes LLaFS the first attempt to leverage large language models (LLMs) in few-shot segmentation. In contrast to the conventional few-shot segmentation methods that only rely on the limited and biased information from the annotated support images LLaFS leverages the vast prior knowledge…

Cited by 44SourcePDFScholar
2024

LTGC: Long-tail Recognition via Leveraging LLMs-driven Generated Content

CVPR 2024poster

Long-tail recognition is challenging because it requires the model to learn good representations from tail categories and address imbalances across all categories. In this paper we propose a novel generative and fine-tuning framework LTGC to handle long-tail recognition via leveraging generated cont…

Cited by 16SourcePDFScholar
2024

Learning Task-Aware Language-Image Representation for Class-Incremental Object Detection

AAAI 2024technical

Class-incremental object detection (CIOD) is a real-world desired capability, requiring an object detector to continuously adapt to new tasks without forgetting learned ones, with the main challenge being catastrophic forgetting. Many methods based on distillation and replay have been proposed to al…

Cited by 5SourcePDFScholar
2024

Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering

NeurIPS 2024poster

Audio-Visual Question Answering (AVQA) is a complex multi-modal reasoning task, demanding intelligent systems to accurately respond to natural language queries based on audio-video input pairs. Nevertheless, prevalent AVQA approaches are prone to overlearning dataset biases, resulting in poor robust…

2024

MWSIS: Multimodal Weakly Supervised Instance Segmentation with 2D Box Annotations for Autonomous Driving

AAAI 2024technical

Instance segmentation is a fundamental research in computer vision, especially in autonomous driving. However, manual mask annotation for instance segmentation is quite time-consuming and costly. To address this problem, some prior works attempt to apply weakly supervised manner by exploring 2D or 3…

2024

MatchDet: A Collaborative Framework for Image Matching and Object Detection

AAAI 2024technical

Image matching and object detection are two fundamental and challenging tasks, while many related applications consider them two individual tasks (i.e. task-individual). In this paper, a collaborative framework called MatchDet (i.e. task-collaborative) is proposed for image matching and object detec…

Cited by 0SourcePDFScholar
2024

Mixed Geometry Message and Trainable Convolutional Attention Network for Knowledge Graph Completion

AAAI 2024technical

Knowledge graph completion (KGC) aims to study the embedding representation to solve the incompleteness of knowledge graphs (KGs). Recently, graph convolutional networks (GCNs) and graph attention networks (GATs) have been widely used in KGC tasks by capturing neighbor information of entities. Howev…

Cited by 10SourcePDFScholar
2024

Mutualreg: Mutual Learning for Unsupervised Medical Image Registration

ICASSP 2024accepted

Recently, self-training strategies have shown outstanding performance in the unsupervised medical image registration field. These strategies use their own network to generate pseudo-displacement fields (PFs) to supervise network training. However, limited diversity and accuracy of these PFs hinder t…

Cited by 0SourceScholar
2024

PathReasoner: Modeling Reasoning Path with Equivalent Extension for Logical Question Answering

ACL 2024long

Logical reasoning task has attracted great interest since it was proposed. Faced with such a task, current competitive models, even large language models (e.g., ChatGPT and PaLM 2), still perform badly. Previous promising LMs struggle in logical consistency modeling and logical structure perception.…

Cited by 2SourcePDFScholar
2024

Physics-Informed Neural Network Policy Iteration: Algorithms, Convergence, and Verification

ICML 2024poster

Solving nonlinear optimal control problems is a challenging task, particularly for high-dimensional problems. We propose algorithms for model-based policy iterations to solve nonlinear optimal control problems with convergence guarantees. The main component of our approach is an iterative procedure…

Cited by 14SourcePDFScholar
2024

QGEval: Benchmarking Multi-dimensional Evaluation for Question Generation

EMNLP 2024main

Automatically generated questions often suffer from problems such as unclear expression or factual inaccuracies, requiring a reliable and comprehensive evaluation of their quality. Human evaluation is widely used in the field of question generation (QG) and serves as the gold standard for automatic…

2024

Symbol-LLM: Towards Foundational Symbol-centric Interface For Large Language Models

ACL 2024long

Although Large Language Models (LLMs) demonstrate remarkable ability in processing and generating human-like text, they do have limitations when it comes to comprehending and expressing world knowledge that extends beyond the boundaries of natural language(e.g., chemical molecular formula). Injectin…

2024

Towards Physical World Backdoor Attacks against Skeleton Action Recognition

ECCV 2024poster

"Skeleton Action Recognition (SAR) has attracted significant interest for its efficient representation of the human skeletal structure. Despite its advancements, recent studies have raised security concerns in SAR models, particularly their vulnerability to adversarial attacks. However, such strateg…

Cited by 3SourcePDFScholar
2024

UPAM: Unified Prompt Attack in Text-to-Image Generation Models Against Both Textual Filters and Visual Checkers

ICML 2024poster

Text-to-Image (T2I) models have raised security concerns due to their potential to generate inappropriate or harmful images. In this paper, we propose UPAM, a novel framework that investigates the robustness of T2I models from the attack perspective. Unlike most existing attack methods that focus on…

Cited by 4SourcePDFScholar
2024

When Phrases Meet Probabilities: Enabling Open Relation Extraction with Cooperating Large Language Models

ACL 2024long

Current clustering-based open relation extraction (OpenRE) methods usually apply clustering algorithms on top of pre-trained language models. However, this practice has three drawbacks. First, embeddings from language models are high-dimensional and anisotropic, so using simple metrics to calculate…

2023

A Characteristic Function-Based Method for Bottom-Up Human Pose Estimation

CVPR 2023poster

Most recent methods formulate the task of human pose estimation as a heatmap estimation problem, and use the overall L2 loss computed from the entire heatmap to optimize the heatmap prediction. In this paper, we show that in bottom-up human pose estimation where each heatmap often contains multiple…

Cited by 9SourcePDFScholar
2023

Bi-Directional Feature Fusion Generative Adversarial Network for Ultra-High Resolution Pathological Image Virtual Re-Staining

CVPR 2023poster

The cost of pathological examination makes virtual re-staining of pathological images meaningful. However, due to the ultra-high resolution of pathological images, traditional virtual re-staining methods have to divide a WSI image into patches for model training and inference. Such a limitation lead…

Cited by 11SourcePDFScholar
2023

Chaotic World: A Large and Challenging Benchmark for Human Behavior Understanding in Chaotic Events

ICCV 2023poster

Understanding and analyzing human behaviors (actions and interactions of people), voices, and sounds in chaotic events is crucial in many applications, e.g., crowd management, emergency response services. Different from human behaviors in daily life, human behaviors in chaotic events are generally d…

Cited by 4PDFcodeScholar
2023

Clustered-patch Element Connection for Few-shot Learning

IJCAI 2023poster

Weak feature representation problem has influenced the performance of few-shot classification task for a long time. To alleviate this problem, recent researchers build connections between support and query instances through embedding patch features to generate discriminative representations. However…

2023

Continual Semantic Segmentation With Automatic Memory Sample Selection

CVPR 2023poster

Continual Semantic Segmentation (CSS) extends static semantic segmentation by incrementally introducing new classes for training. To alleviate the catastrophic forgetting issue in CSS, a memory buffer that stores a small number of samples from the previous classes is constructed for replay. However,…

Cited by 57SourcePDFScholar
2023

Diagram Visual Grounding: Learning to See with Gestalt-Perceptual Attention

IJCAI 2023poster

Diagram visual grounding aims to capture the correlation between language expression and local objects in the diagram, and plays an important role in the applications like textbook question answering and cross-modal retrieval. Most diagrams consist of several colors and simple geometries. This resul…

2023

DiffPose: Toward More Reliable 3D Pose Estimation

CVPR 2023poster

Monocular 3D human pose estimation is quite challenging due to the inherent ambiguity and occlusion, which often lead to high uncertainty and indeterminacy. On the other hand, diffusion models have recently emerged as an effective tool for generating high-quality images from noise. Inspired by their…

2023

Diffusion-based Image Translation with Label Guidance for Domain Adaptive Semantic Segmentation

ICCV 2023poster

Translating images from a source domain to a target domain for learning target models is one of the most common strategies in domain adaptive semantic segmentation (DASS). However, existing methods still struggle to preserve semantically-consistent local details between the original and translated i…

Cited by 32PDFScholar
2023

GPTR: Gestalt-Perception Transformer for Diagram Object Detection

AAAI 2023technical

Diagram object detection is the key basis of practical applications such as textbook question answering. Because the diagram mainly consists of simple lines and color blocks, its visual features are sparser than those of natural images. In addition, diagrams usually express diverse knowledge, in whi…

Cited by 6SourcePDFScholar
2023

Heterogeneous Diversity Driven Active Learning for Multi-Object Tracking

ICCV 2023poster

The existing one-stage multi-object tracking (MOT) algorithms have achieved satisfactory performance benefiting from a large amount of labeled data. However, acquiring plenty of laborious annotated frames is not practical in real applications. To reduce the cost of human annotations, we propose Hete…

Cited by 6PDFScholar
2023

Instance and Category Supervision are Alternate Learners for Continual Learning

ICCV 2023poster

Continual Learning (CL) is the constant development of complex behaviors by building upon previously acquired skills. Yet, current CL algorithms tend to incur class-level forgetting as the label information is often quickly overwritten by new knowledge. This motivates attempts to mine instance-level…

Cited by 2PDFScholar
2023

Joint Attribute and Model Generalization Learning for Privacy-Preserving Action Recognition

NeurIPS 2023poster

Privacy-Preserving Action Recognition (PPAR) aims to transform raw videos into anonymous ones to prevent privacy leakage while maintaining action clues, which is an increasingly important problem in intelligent vision applications. Despite recent efforts in this task, it is still challenging to deal…

Cited by 4SourcePDFScholar
2023

LMC: Large Model Collaboration with Cross-assessment for Training-Free Open-Set Object Recognition

NeurIPS 2023poster

Open-set object recognition aims to identify if an object is from a class that has been encountered during training or not. To perform open-set object recognition accurately, a key challenge is how to reduce the reliance on spurious-discriminative features. In this paper, motivated by that different…

2023

MDCS: More Diverse Experts with Consistency Self-distillation for Long-tailed Recognition

ICCV 2023poster

Recently, multi-expert methods have led to significant improvements in long-tail recognition (LTR). We summarize two aspects that need further enhancement to contribute to LTR boosting: (1) More diverse experts; (2) Lower model variance. However, the previous methods didn't handle them well. To this…

Cited by 16PDFcodeScholar
2023

Meta Compositional Referring Expression Segmentation

CVPR 2023poster

Referring expression segmentation aims to segment an object described by a language expression from an image. Despite the recent progress on this task, existing models tackling this task may not be able to fully capture semantics and visual representations of individual concepts, which limits their…

Cited by 34SourcePDFScholar
2023

MixPro: Data Augmentation with MaskMix and Progressive Attention Labeling for Vision Transformer

ICLR 2023poster

The recently proposed data augmentation TransMix employs attention labels to help visual transformers (ViT) achieve better robustness and performance. However, TransMix is deficient in two aspects: 1) The image cropping method of TransMix may not be suitable for vision transformer. 2) At the early s…

2023

Optimal Mixed-ADC Arrangement for DOA Estimation Via CRB Using ULA

ICASSP 2023accepted

We consider a mixed analog-to-digital converter (ADC) based architecture for direction of arrival (DOA) estimation using a uniform linear array (ULA). We derive the Cramér-Rao bound (CRB) of the DOA under the optimal time-varying threshold, and find that the asymptotic CRB is related to the arrangem…

Cited by 0SourceScholar
2023

Rethinking Gradient Projection Continual Learning: Stability / Plasticity Feature Space Decoupling

CVPR 2023poster

Continual learning aims to incrementally learn novel classes over time, while not forgetting the learned knowledge. Recent studies have found that learning would not forget if the updated gradient is orthogonal to the feature space. However, previous approaches require the gradient to be fully ortho…

Cited by 29SourcePDFScholar
2023

STPrivacy: Spatio-Temporal Privacy-Preserving Action Recognition

ICCV 2023poster

Existing methods of privacy-preserving action recognition (PPAR) mainly focus on frame-level (spatial) privacy removal through 2D CNNs. Unfortunately, they have two major drawbacks. First, they may compromise temporal dynamics in input videos, which are critical for accurate action recognition. Seco…

Cited by 24PDFScholar
2023

SpatialFormer: Semantic and Target Aware Attentions for Few-Shot Learning

AAAI 2023technical

Recent Few-Shot Learning (FSL) methods put emphasis on generating a discriminative embedding features to precisely measure the similarity between support and query sets. Current CNN-based cross-attention approaches generate discriminative representations via enhancing the mutually semantic similar r…

2023

Synthesize, Prompt and Transfer: Zero-shot Conversational Question Generation with Pre-trained Language Model

ACL 2023long

Conversational question generation aims to generate questions that depend on both context and conversation history. Conventional works utilizing deep learning have shown promising results, but heavily rely on the availability of large-scale annotated conversations. In this paper, we introduce a more…

Cited by 10SourcePDFScholar
2023

System-Status-Aware Adaptive Network for Online Streaming Video Understanding

CVPR 2023poster

Recent years have witnessed great progress in deep neural networks for real-time applications. However, most existing works do not explicitly consider the general case where the device's state and the available resources fluctuate over time, and none of them investigate or address the impact of vary…

2023

TECHS: Temporal Logical Graph Networks for Explainable Extrapolation Reasoning

ACL 2023long

Extrapolation reasoning on temporal knowledge graphs (TKGs) aims to forecast future facts based on past counterparts. There are two main challenges: (1) incorporating the complex information, including structural dependencies, temporal dynamics, and hidden logical rules; (2) implementing differentia…

Cited by 49SourcePDFScholar
2023

Token Boosting for Robust Self-Supervised Visual Transformer Pre-Training

CVPR 2023poster

Learning with large-scale unlabeled data has become a powerful tool for pre-training Visual Transformers (VTs). However, prior works tend to overlook that, in real-world scenarios, the input data may be corrupted and unreliable. Pre-training VTs on such corrupted data can be challenging, especially…

Cited by 6SourcePDFScholar
2023

Towards Optimal Design of Dielectric Elastomer Actuators Using a Graph Neural Network Encoder

RA-L 2023

Dielectric elastomer actuators (DEAs), a type of “artificial muscles”, can generate significant deformations and offer speedy responses when exposed to voltage. Owing to their high electromechanical conversion efficiency and great flexibility, they have been extensively used in soft robot applicatio

Cited by 6SourceScholar
2023

Uncertainty-Aware Unsupervised Image Deblurring With Deep Residual Prior

CVPR 2023poster

Non-blind deblurring methods achieve decent performance under the accurate blur kernel assumption. Since the kernel uncertainty (i.e. kernel error) is inevitable in practice, semi-blind deblurring is suggested to handle it by introducing the prior of the kernel (or induced) error. However, how to de…

Cited by 18SourcePDFScholar
2022

Animal Kingdom: A Large and Diverse Dataset for Animal Behavior Understanding

CVPR 2022oral

Understanding animals' behaviors is significant for a wide range of applications. However, existing animal behavior datasets have limitations in multiple aspects, including limited numbers of animal classes, data samples and provided tasks, and also limited variations in environmental conditions and…

Cited by 101PDFcodeScholar
2022

Decoupling Classifier for Boosting Few-shot Object Detection and Instance Segmentation

NeurIPS 2022accept

This paper focus on few-shot object detection~(FSOD) and instance segmentation~(FSIS), which requires a model to quickly adapt to novel classes with a few labeled instances. The existing methods severely suffer from bias classification because of the missing label issue which naturally exists in an…

2022

Dynamic Spatio-Temporal Specialization Learning for Fine-Grained Action Recognition

ECCV 2022poster

"The goal of fine-grained action recognition is to successfully discriminate between action categories with subtle differences. To tackle this, we derive inspiration from the human visual system which contains specialized regions in the brain that are dedicated towards handling specific tasks. We de…

Cited by 30SourcePDFScholar
2022

ERA: Expert Retrieval and Assembly for Early Action Prediction

ECCV 2022poster

"Early action prediction aims to successfully predict the class label of an action before it is completely performed. This is a challenging task because the beginning stages of different actions can be very similar, with only minor subtle differences for discrimination. In this paper, we propose a n…

Cited by 29SourcePDFScholar
2022

En-Compactness: Self-Distillation Embedding & Contrastive Generation for Generalized Zero-Shot Learning

CVPR 2022poster

Generalized zero-shot learning (GZSL) requires a classifier trained on seen classes that can recognize objects from both seen and unseen classes. Due to the absence of unseen training samples, the classifier tends to bias towards seen classes. To mitigate this problem, feature generation based model…

Cited by 89PDFScholar
2022

GradAuto: Energy-Oriented Attack on Dynamic Neural Networks

ECCV 2022poster

"Dynamic neural networks could adapt their structures or parameters based on different inputs. By reducing the computation redundancy for certain samples, it can greatly improve the computational efficiency without compromising the accuracy. In this paper, we investigate the robustness of dynamic ne…

2022

IGFormer: Interaction Graph Transformer for Skeleton-Based Human Interaction Recognition

ECCV 2022poster

"Human interaction recognition is very important in many applications. One crucial cue in recognizing an interaction is the interactive body parts. In this work, we propose a novel Interaction Graph Transformer (IGFormer) network for skeleton-based interaction recognition via modeling the interactiv…

Cited by 48SourcePDFScholar
2022

Improving the Reliability for Confidence Estimation

ECCV 2022poster

"Confidence estimation, a task that aims to evaluate the trustworthiness of the model’s prediction output during deployment, has received lots of research attention recently, due to its importance for the safe deployment of deep models. Previous works have outlined two important qualities that a rel…

Cited by 13SourcePDFScholar
2022

Inductive Relation Prediction with Logical Reasoning Using Contrastive Representations

EMNLP 2022main

Relation prediction in knowledge graphs (KGs) aims at predicting missing relations in incomplete triples, whereas the dominant embedding paradigm has a restriction on handling unseen entities during testing. In the real-world scenario, the inductive setting is more common because entities in the tra…

Cited by 21SourcePDFScholar
2022

MatchPrompt: Prompt-based Open Relation Extraction with Semantic Consistency Guided Clustering

EMNLP 2022main

Relation clustering is a general approach for open relation extraction (OpenRE). Current methods have two major problems. One is that their good performance relies on large amounts of labeled and pre-defined relational instances for pre-training, which are costly to acquire in reality. The other is…

2022

Meta Spatio-Temporal Debiasing for Video Scene Graph Generation

ECCV 2022poster

"Video scene graph generation (VidSGG) aims to parse the video content into scene graphs, which involves modeling the spatio-temporal contextual information in the video. However, due to the long-tailed training data in datasets, the generalization performance of existing VidSGG models can be affect…

Cited by 32SourcePDFScholar
2022

Neural Lyapunov Control of Unknown Nonlinear Systems with Stability Guarantees

NeurIPS 2022accept

Learning for control of dynamical systems with formal guarantees remains a challenging task. This paper proposes a learning framework to simultaneously stabilize an unknown nonlinear system with a neural controller and learn a neural Lyapunov function to certify a region of attraction (ROA) for the…

2022

Qrelation: an Agent Relation-Based Approach for Multi-Agent Reinforcement Learning Value Function Factorization

ICASSP 2022accepted

The Centralized Training with Decentralized Execution paradigm (CTDE), which trains policies centrally with additional information, is important for Multi-Agent Reinforcement Learning (MARL). For CTDE, value function factorization methods make use of state during training and factorize the value fun…

Cited by 0SourceScholar
2022

REMOTE: Reinforced Motion Transformation Network for Semi-supervised 2D Pose Estimation in Videos

AAAI 2022technical

Existing approaches for 2D pose estimation in videos often require a large number of dense annotations, which are costly and labor intensive to acquire. In this paper, we propose a semi-supervised REinforced MOtion Transformation nEtwork (REMOTE) to leverage a few labeled frames and temporal pose va…

Cited by 13SourcePDFScholar
2022

ResQ: A Residual Q Function-based Approach for Multi-Agent Reinforcement Learning Value Factorization

NeurIPS 2022accept

The factorization of state-action value functions for Multi-Agent Reinforcement Learning (MARL) is important. Existing studies are limited by their representation capability, sample efficiency, and approximation error. To address these challenges, we propose, ResQ, a MARL value function factorizatio…

Cited by 23SourcePDFScholar
2022

Robust Image Forgery Detection Over Online Social Network Shared Images

CVPR 2022oral

The increasing abuse of image editing softwares, such as Photoshop and Meitu, causes the authenticity of digital images questionable. Meanwhile, the widespread availability of online social networks (OSNs) makes them the dominant channels for transmitting forged images to report fake news, propagate…

Cited by 85PDFcodeScholar
2022

Simulation Data Driven Design Optimization for Reconfigurable Soft Gripper System

RA-L 2022

In the soft gripper design work, most of the designs such as gripping width and the design of finger actuator are purely based on experience, and repeated trial-and-error. In most scenarios, the designed actuators cannot achieve the best/optimized grasping performance with a specific design type. Th

Cited by 15SourceScholar
2022

Uncertainty Modeling for Out-of-Distribution Generalization

ICLR 2022poster

Though remarkable progress has been achieved in various vision tasks, deep neural networks still suffer obvious performance degradation when tested in out-of-distribution scenarios. We argue that the feature statistics (mean and standard deviation), which carry the domain characteristics of the trai…

2022

Versatile Motion Generation of Magnetic Origami Spring Robots in the Uniform Magnetic Field

RA-L 2022

Magnetic soft robots have attracted widespread attention for their untethered, remotely operated, and compliant deformation characteristics. Earlier work has demonstrated magnetic origami robots' diverse locomotion capabilities. This letter will focus on the motion generation and open-loop control o

Cited by 22SourceScholar
2022

tSF: Transformer-Based Semantic Filter for Few-Shot Learning

ECCV 2022poster

"Few-Shot Learning (FSL) alleviates the data shortage challenge via embedding discriminative target-aware features among plenty seen (base) and few unseen (novel) labeled samples. Most feature embedding modules in recent FSL methods are specially designed for corresponding learning tasks (e.g., clas…

2021

A Unified 3D Human Motion Synthesis Model via Conditional Variational Auto-Encoder

ICCV 2021poster

We present a unified and flexible framework to address the generalized problem of 3D motion synthesis that covers the tasks of motion prediction, completion, interpolation, and spatial-temporal recovery. Since these tasks have different input constraints and various fidelity and diversity requiremen…

Cited by 80PDFScholar
2021

Distributed Resilient Submodular Action Selection in Adversarial Environments

RA-L 2021

In this letter, we consider a distributed submodular maximization problem for multi-robot systems when attacked by adversaries. One of the major challenges for multi-robot systems is to increase resilience against failures or attacks. This is particularly important for distributed systems under atta

Cited by 29SourceScholar
2021

Dynamic tracking for microrobot with active magnetic sensor array

ICRA 2021poster

Accurate position feedback in a wide range is critical for medical microrobotics and robot-assisted examinations, such as colonoscopy, bronchoscopy and capsule endoscopy examination. Among the many modalities of positioning feedback, magnetic tracking is a preferable method due to the unique advanta…

Cited by 11SourceScholar
2021

Else-Net: Elastic Semantic Network for Continual Action Recognition From Skeleton Data

ICCV 2021poster

We address continual action recognition from skeleton sequence, which aims to learn a recognition model over time from a continuous stream of skeleton data. This task is very important in changing environment. Due to catastrophic forgetting problems of deep neural networks and large discrepancies be…

Cited by 56PDFScholar
2021

Generalizable Person Re-Identification With Relevance-Aware Mixture of Experts

CVPR 2021poster

Domain generalizable (DG) person re-identification (ReID) is a challenging problem because we cannot access any unseen target domain data during training. Almost all the existing DG ReID methods follow the same pipeline where they use a hybrid dataset from multiple source domains for training, and t…

Cited by 160PDFScholar
2021

IDM: An Intermediate Domain Module for Domain Adaptive Person Re-ID

ICCV 2021poster

Unsupervised domain adaptive person re-identification (UDA re-ID) aims at transferring the labeled source domain's knowledge to improve the model's discriminability on the unlabeled target domain. From a novel perspective, we argue that the bridging between the source and target domains can be utili…

Cited by 164PDFcodeScholar
2021

Interaction via Bi-Directional Graph of Semantic Region Affinity for Scene Parsing

ICCV 2021poster

In this work, we devote to address the challenging problem of scene parsing. Previous methods, though capture context to exploit global clues, handle scene parsing as a pixel-independent task. However, it is well known that pixels in an image are highly correlated with each other, especially those f…

Cited by 19PDFScholar
2021

Interventional Video Grounding With Dual Contrastive Learning

CVPR 2021poster

Video grounding aims to localize a moment from an untrimmed video for a given textual query. Existing approaches focus more on the alignment of visual and language stimuli with various likelihood-based matching or regression strategies, i.e., P(Y|X). Consequently, these models may suffer from spurio…

Cited by 174PDFcodeScholar
2021

Neighborhood Spatial Aggregation based Efficient Uncertainty Estimation for Point Cloud Semantic Segmentation

ICRA 2021poster

Uncertainty estimation for point cloud semantic segmentation is to quantify the confidence degree for the predicted label of points, which is essential for decision-making tasks. This paper proposes a neighborhood spatial aggregation based method, NSA-MC dropout, to achieve efficient uncertainty est…

Cited by 4SourcecodeScholar
2021

Person30K: A Dual-Meta Generalization Network for Person Re-Identification

CVPR 2021poster

Recently, person re-identification (ReID) has vastly benefited from the surging waves of data-driven methods. However, these methods are still not reliable enough for real-world deployments, due to the insufficient generalization capability of the models learned on existing benchmarks that have limi…

Cited by 78PDFScholar
2021

SUTD-TrafficQA: A Question Answering Benchmark and an Efficient Network for Video Reasoning Over Traffic Events

CVPR 2021poster

Traffic event cognition and reasoning in videos is an important task that has a wide range of applications in intelligent transportation, assisted driving, and autonomous vehicles. In this paper, we create a novel dataset, SUTD-TrafficQA (Traffic Question Answering), which takes the form of video QA…

Cited by 106PDFcodeScholar
2021

Skeleton Cloud Colorization for Unsupervised 3D Action Representation Learning

ICCV 2021poster

Skeleton-based human action recognition has attracted increasing attention in recent years. However, most of the existing works focus on supervised learning which requiring a large number of annotated action sequences that are often expensive to collect. We investigate unsupervised representation le…

Cited by 123PDFScholar
2021

UAV-Human: A Large Benchmark for Human Behavior Understanding With Unmanned Aerial Vehicles

CVPR 2021poster

Human behavior understanding with unmanned aerial vehicles (UAVs) is of great significance for a wide range of applications, which simultaneously brings an urgent demand of large, challenging, and comprehensive benchmarks for the development and evaluation of UAV-based models. However, existing benc…

Cited by 268PDFcodeScholar
2021

WSUIE: Weakly Supervised Underwater Image Enhancement for Improved Visual Perception

RA-L 2021

Underwater images inevitably suffer from degradation and blur due to the scattering and absorption of light as it propagates through the water, which hinders the development of underwater visual perception. Existing deep underwater image enhancement methods mainly rely on the strong supervision of a

Cited by 31SourceScholar
2020

Anomaly Detection with Training Data in Hyperspectral Imagery

ICASSP 2020accepted

In this paper, we investigate the anomaly detection problem for multi-pixel targets in hyperspectral imagery when training data are available. We derive the generalized likelihood ratio test and obtain its analytical expressions of the probability of false alarm and probability of detection. The per…

Cited by 0SourceScholar
2020

Collaborative Learning of Gesture Recognition and 3D Hand Pose Estimation with Multi-Order Feature Analysis

ECCV 2020poster

Gesture recognition and 3D hand pose estimation are two highly correlated tasks, yet they are often handled separately. In this paper, we present a novel collaborative learning network for joint gesture recognition and 3D hand pose estimation. The proposed network exploits joint-aware features that…

Cited by 57SourcePDFScholar
2020

Cross-Stained Segmentation from Renal Biopsy Images Using Multi-Level Adversarial Learning

ICASSP 2020accepted

Segmentation from renal pathological images is a key step in automatic analyzing the renal histological characteristics. However, the performance of models varies significantly in different types of stained datasets due to the appearance variations. In this paper, we design a robust and flexible mod…

Cited by 0SourceScholar