← Search

Fei Wang

159 accepted papers

2026

A Novel Upper Limb Rehabilitation Framework Based on Dual-Arm Robotics for Therapist-Like Traction Training

RA-L 2026

In this letter, we propose a novel upper limb rehabilitation framework based on dual-arm robotics for therapist- like traction training. Prioritizing patient safety, an 8-DOF kinematic model of the upper limb is derived to evaluate the reachable workspace of the palm center and proximal forearm duri

Cited by 1SourceScholar
2026

A Novel Upper Limb Rehabilitation Framework Based on Dual-Arm Robotics for Therapist-Like Traction Training

ICRA 2026poster

In this letter, we propose a novel upper limb rehabilitation framework based on dual-arm robotics for therapist-like traction training. Prioritizing patient safety, an 8-DOF kinematic model of the upper limb complex is derived to evaluate the reachable workspace of the end-of-arm and forearm during …

Cited by 0SourceScholar
2026

APT: Affine Prototype-Timestamp for Time Series Forecasting Under Distribution Shift

AAAI 2026technical

Time series forecasting under distribution shift remains challenging, as existing deep learning models often rely on local statistical normalization (e.g., mean and variance) that fails to capture global distribution shift. Methods like RevIN and its variants attempt to decouple distribution and pat

Cited by 2SourcePDFScholar
2026

DropoutTS: Sample-Adaptive Dropout for Robust Time Series Forecasting

ICML 2026poster

Deep time series models are vulnerable to noisy data ubiquitous in real-world applications. Existing robustness strategies either prune data or rely on costly prior quantification, failing to balance effectiveness and efficiency. In this paper, we introduce DropoutTS, a model-agnostic plugin that sh…

Cited by 0SourceScholar
2026

EdgeGrasp: Enhancing Edge Perception for 7-DoF Grasping Pose Estimation in Cluttered Scenes

ICRA 2026poster

Estimating 7-DoF grasping poses (6-DoF with gripper width) in cluttered scenes is a critical challenge for robotic manipulation. In such environments, object edges often contain many promising grasp candidates, but relying solely on incomplete single-view point cloud to infer them is difficult. Whil…

Cited by 0Scholar
2026

GRAPHPL: LEVERAGING GNN FOR EFFICIENT AND ROBUST MODALITIES IMPUTATION IN PATCHWORK LEARNING

ICASSP 2026poster

Current research on distributed multi-modal learning typically assumes that clients can access complete information across all modalities, which may not hold in practice. In this paper, we explore patchwork learning, in which the modalities available to different clients vary, and the objective is t…

Cited by 0SourcePDFScholar
2026

Localizing, Structuring, and Rendering: Bridging 3D and 2D Vision-Language-Action Models for Robotic Manipulation

CVPR 2026

Robotic manipulation in complex 3D environments requires unifying spatial reasoning with intuitive visual perception, which is a capability that current Vision-Language-Action paradigms address separately. While 3D VLAs excel in geometric and physical reasoning, they lack intuitive, image-level unde

Cited by 0SourcecodeScholar
2026

MSP: Probabilistically Consistent Multi-Scale Action Generation

ICML 2026spotlight

In robotic imitation learning, accurately modeling the multimodality and temporal correlations of long-horizon action sequences remains challenging. Long-horizon tasks require preserving global task intent while executing precise low-level control; otherwise, local errors can accumulate and lead to …

Cited by 0SourceScholar
2026

MimicTalker: A Multimodal Interactive and Memory-Enhanced Framework for Real-Time Dyadic 3D Head Generation

CVPR 2026

Dyadic interactive head generation aims to synthesize realistic head motions that respond both verbally and non-verbally to an interlocutor in real-time conversation. The existing works often focus on offline scenarios, and struggle with a shallow understanding of the multimodal conversational conte

Cited by 0SourceScholar
2026

Multimodal Protein Language Models for Enzyme Kinetic Parameters: From Substrate Recognition to Conformational Adaptation

CVPR 2026

Predicting enzyme kinetic parameters quantifies how efficiently an enzyme catalyzes a specific substrate under defined biochemical conditions. Canonical parameters such as the turnover number (k_\text cat ), Michaelis constant (K_\text m ), and inhibition constant (K_\text i ) depend jointly on the

Cited by 0SourceScholar
2026

PULSE: Generative Phase Evolution for Non-Stationary Time Series Forecasting

ICML 2026poster

Time series forecasting under non-stationarity faces a fundamental tension between capturing stable representations and adapting to distribution shifts. Existing methods implicitly rely on static historical assumptions, leading to a critical failure mode we term Phase Amnesia, where models become bl…

Cited by 0SourceScholar
2026

Parameterized Prompt for Incremental Object Detection

CVPR 2026

Recent studies have demonstrated that incorporating trainable prompts into pretrained models enables effective incremental learning. However, the application of prompts in incremental object detection (IOD) remains underexplored. Our study reveals that existing prompts-pool-based approaches assume d

Cited by 0SourcecodeScholar
2026

RayI2P: Learning Rays for Image-to-Point Cloud Registration

ICLR 2026poster

Image-to-point cloud registration aims to estimate the 6-DoF camera pose of a query image relative to a 3D point cloud map. Existing methods fall into two categories: matching-free methods regress pose directly using geometric priors, but lack fine-grained supervision and struggle with precise align…

Cited by 0SourceScholar
2026

Robust Admittance Control of an Electric Underwater Manipulator for Precise Motion and Safe Contact Inspection of Hydraulic Structures

ICRA 2026poster

Inspection of hydraulic structures is crucial for ensuring the reliability and safety of infrastructures. Although underwater manipulators are essential tools, existing systems often lack sufficient compliance and safe interaction capabilities. This study develops a novel underwater manipulator syst…

Cited by 0SourceScholar
2026

SAVE: A Generalizable Framework for Multi-Condition Single-Cell Generation with Gene Block Attention

ICLR 2026poster

Modeling single-cell gene expression across diverse biological and technical conditions is essential for understanding cellular states and simulating unobserved scenarios. We present SAVE, a unified generative framework for multi-condition single-cell modeling. SAVE combines a variational autoencode…

Cited by 0SourcecodeScholar
2026

Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation

ICML 2026poster

Vision-language-action (VLA) models typically rely on large-scale real-world videos, whereas simulated data, despite being inexpensive and highly parallelizable to collect, often suffers from a substantial visual domain gap and limited environmental diversity, resulting in weak real-world generaliza…

Cited by 0SourceScholar
2026

Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation

ICLR 2026poster

Camera-centric understanding and generation are two cornerstones of spatial intelligence, yet they are typically studied in isolation. We present Puffin, a unified camera-centric multimodal model that extends spatial awareness along the camera dimension. Puffin integrates language regression and dif…

Cited by 0SourcecodeScholar
2026

Zeus: Towards Tuning-Free Foundation Model for Time Series Analysis

ICML 2026poster

We present Zeus, a unified tuning-free Time Series Foundation Model (TSFM) that delivers superior performance across diverse analysis tasks without any task-specific fine-tuning. Unlike prior studies that primarily focus on zero-shot forecasting but require task-specific tuning for other tasks, Zeus…

Cited by 0SourceScholar
2025

Aligning to Constraints for Data-Efficient Language Model Customization

NAACL 2025findings

General-purpose language models (LMs) are aligned to diverse user intents, but fall short when it comes to specific applications. While finetuning is the default method for customized alignment, human annotations are often unavailable in various customization scenarios. Based on the observation that…

Cited by 0SourcePDFScholar
2025

Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language Models

ACL 2025long

Retrieval augmented generation (RAG), while effectively integrating external knowledge to address the inherent limitations of large language models (LLMs), can be hindered by imperfect retrieval that contain irrelevant, misleading, or even malicious information. Previous studies have rarely connecte…

Cited by 0SourcePDFScholar
2025

Benchmarking Vision Language Model Unlearning via Fictitious Facial Identity Dataset

ICLR 2025poster

Machine unlearning has emerged as an effective strategy for forgetting specific information in the training data. However, with the increasing integration of visual data, privacy concerns in Vision Language Models (VLMs) remain underexplored. To address this, we introduce Facial Identity Unlearning…

2025

CodexGraph: Bridging Large Language Models and Code Repositories via Code Graph Databases

NAACL 2025long

Large Language Models (LLMs) excel in stand-alone code tasks like HumanEval and MBPP, but struggle with handling entire code repositories. This challenge has prompted research on enhancing LLM-codebase interaction at a repository scale. Current solutions rely on similarity-based retrieval or manual…

2025

Conditional Causal Representation Learning for Heterogeneous Single-cell RNA Data Integration and Prediction

IJCAI 2025

Single-cell sequencing technology provides deep insights into gene activity at the individual cell level, facilitating the study of gene regulatory mechanisms. However, observed gene expression are often influenced by confounding factors such as batch effects, perturbations, and spatial position, wh

2025

DARN: An Attention-Based Neural Network Using Residual Blocks for Sleep Micro-Events Detection

ICASSP 2025accepted

Sleep micro-events are crucial indicators of sleep quality and neurological health. However, traditional sleep micro-events detectors often suffer from low precision, which leads to a high rate of false positives and misclassification. This affects the reliability of sleep research and clinical diag…

Cited by 0SourceScholar
2025

Democratizing Clinical Risk Prediction with Cross-Cohort Cross-Modal Knowledge Transfer

NeurIPS 2025poster

Clinical risk prediction plays a crucial role in early disease detection and personalized intervention. While recent models increasingly incorporate multimodal data, their development typically assumes access to large-scale, multimodal datasets and substantial computational resources. In practice, h…

Cited by 0SourceScholar
2025

Diffusion-based Realistic Listening Head Generation via Hybrid Motion Modeling

CVPR 2025highlight

Listening head generation aims to synthesize non-verbal responsive listening head videos that naturally react to a certain speaker, for which, both realistic head movements, expressive facial expressions, and high visual qualities are expected. Previous approaches typically follow a two-stage pipeli…

Cited by 0SourcePDFScholar
2025

EfficientSleepNet: A Novel Lightweight End-to-End Model for Automated Sleep Staging on Single-Channel EEG

ICASSP 2025accepted

Sleep staging is critical for evaluating sleep quality and regulating sleep patterns. While deep learning has shown potential for automatically scoring sleep stages from raw signals, many existing models are overly complex, computationally intensive, and rely on future information, limiting their us…

Cited by 0SourceScholar
2025

Fine-Grained 3D Gaussian Head Avatars Modeling from Static Captures via Joint Reconstruction and Registration

ICCV 2025poster

Recently, 3D head avatar modeling based on 3D Gaussians has demonstrated significant advantages in rendering quality and efficiency, given sufficient data. Some efforts have begun to train prior models on large datasets to develop generalizable 3D Gaussian head avatar modeling methods. Unfortunately…

Cited by 0SourcePDFScholar
2025

From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning

NAACL 2025long

Motivated by in-context learning (ICL) capabilities of Large Language Models (LLMs), multimodal LLMs with additional visual modality are also exhibited with similar ICL abilities when multiple image-text pairs are provided as demonstrations. However, relatively less work has been done to investigate…

2025

Geometric Feature Embedding for Effective 3D Few-Shot Class Incremental Learning

ICML 2025poster

3D few-shot class incremental learning (FSCIL) aims to learn new point cloud categories from limited samples while preventing the forgetting of previously learned categories. This research area significantly enhances the capabilities of self-driving vehicles and computer vision systems. Existing 3D…

2025

HiLoTs: High-Low Temporal Sensitive Representation Learning for Semi-Supervised LiDAR Segmentation in Autonomous Driving

CVPR 2025poster

LiDAR point cloud semantic segmentation plays a crucial role in autonomous driving. In recent years, semi-supervised methods have gained popularity due to their significant reduction in annotation labor and time costs. Current semi-supervised methods typically focus on point cloud spatial distributi…

2025

Human-guided robotic-assistance handheld continuum medical robot system

IROS 2025

Nowadays, laparoscopic surgery procedures face a trade-off between expensive, complex robotic systems and manual instruments with limited functionality. Fully robotic solutions offer precision but lack portability and intuitive control, while manual tools rely solely on the surgeon’s dexterity, limi

Cited by 0SourceScholar
2025

IRT-Router: Effective and Interpretable Multi-LLM Routing via Item Response Theory

ACL 2025long

Large language models (LLMs) have demonstrated exceptional performance across a wide range of natural language tasks. However, selecting the optimal LLM to respond to a user query often necessitates a delicate balance between performance and cost. While powerful models deliver better results, they c…

2025

Layer as Puzzle Pieces: Compressing Large Language Models through Layer Concatenation

NeurIPS 2025poster

Large Language Models (LLMs) excel at natural language processing tasks, but their massive size leads to high computational and storage demands. Recent works have sought to reduce their model size through layer-wise structured pruning. However, they tend to ignore retaining the capabilities in the p…

Cited by 0SourceScholar
2025

LayerIF: Estimating Layer Quality for Large Language Models using Influence Functions

NeurIPS 2025poster

Pretrained Large Language Models (LLMs) achieve strong performance across a wide range of tasks, yet exhibit substantial variability in the various layers' training quality with respect to specific downstream applications, limiting their downstream performance. It is therefore critical to estimate l…

Cited by 0SourceScholar
2025

Leveraging Global Stereo Consistency for Category-Level Shape and 6D Pose Estimation from Stereo Images

CVPR 2025poster

Stereo-based category-level shape and 6D pose estimation methods have the potential to generalize to a wider range of materials than RGBD methods, which often suffer from depth measurement errors. However, without explicit depth from two views, parameters to be estimated can become inherently entan…

Cited by 0SourcePDFScholar
2025

Local Causal Discovery for Structural Evidence of Direct Discrimination

AAAI 2025technical

Identifying the causal pathways of unfairness is a critical objective for improving policy design and algorithmic decision-making. Prior work in causal fairness analysis often requires knowledge of the causal graph, hindering practical applications in complex or low-knowledge domains. Moreover, glob…

2025

MMAD: Multi-label Micro-Action Detection in Videos

ICCV 2025poster

Human body actions are an important form of non-verbal communication in social interactions. This paper specifically focuses on a subset of body actions known as micro-actions, which are subtle, low-intensity body movements with promising applications in human emotion analysis. In real-world scenari…

2025

MPQ-DM: Mixed Precision Quantization for Extremely Low Bit Diffusion Models

AAAI 2025technical

Diffusion models have received wide attention in generation tasks. However, the expensive computation cost prevents the application of diffusion models in resource-constrained scenarios. Quantization emerges as a practical solution that significantly saves storage and computation by reducing the bit…

2025

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

ICLR 2025poster

We introduce MuirBench, a comprehensive benchmark that focuses on robust multi-image understanding capabilities of multimodal LLMs. MuirBench consists of 12 diverse multi-image tasks (e.g., scene understanding, ordering) that involve 10 categories of multi-image relations (e.g., multiview, temporal…

2025

On the Integration of Spatial-Temporal Knowledge: A Lightweight Approach to Atmospheric Time Series Forecasting

NeurIPS 2025poster

Transformers have gained attention in atmospheric time series forecasting (ATSF) for their ability to capture global spatial-temporal correlations. However, their complex architectures lead to excessive parameter counts and extended training times, limiting their scalability to large-scale forecasti…

Cited by 0SourceScholar
2025

PatchScaler: An Efficient Patch-Independent Diffusion Model for Image Super-Resolution

ICCV 2025poster

While diffusion models significantly improve the perceptual quality of super-resolved images, they usually require a large number of sampling steps, resulting in high computational costs and long inference times. Recent efforts have explored reasonable acceleration schemes by reducing the number of…

2025

SMARTraj$^2$: A Stable Multi-City Adaptive Method for Multi-View Spatio-Temporal Trajectory Representation Learning

NeurIPS 2025poster

Spatio-temporal trajectory representation learning plays a crucial role in various urban applications such as transportation systems, urban planning, and environmental monitoring. Existing methods can be divided into single-view and multi-view approaches, with the latter offering richer representati…

Cited by 0SourcecodeScholar
2025

Selective Learning for Deep Time Series Forecasting

NeurIPS 2025poster

Benefiting from high capacity for capturing complex temporal patterns, deep learning (DL) has significantly advanced time series forecasting (TSF). However, deep models tend to suffer from severe overfitting due to the inherent vulnerability of time series to noise and anomalies. The prevailing DL p…

Cited by 0SourceScholar
2025

SiQA: A Large Multi-Modal Question Answering Model for Structured Images Based on RAG

ICASSP 2025accepted

Existing Large Multimodal Models (LMMs) demonstrate excellent performance in handling visual tasks in everyday scenarios. However, they still face challenges in understanding structured images, such as flowcharts and organizational charts, which are characterized by text-rich and complex hierarchica…

Cited by 0SourceScholar
2025

SudoLM: Learning Access Control of Parametric Knowledge with Authorization Alignment

ACL 2025long

Existing preference alignment is a one-size-fits-all alignment mechanism, where the part of the large language model (LLM) parametric knowledge with non-preferred features is uniformly blocked to all the users. However, this part of knowledge can be useful to advanced users whose expertise qualifies…

Cited by 0SourcePDFScholar
2025

Temporal-Frequency State Space Duality: An Efficient Paradigm for Speech Emotion Recognition

ICASSP 2025accepted

Speech Emotion Recognition (SER) plays a critical role in enhancing user experience within human-computer interaction. However, existing methods are overwhelmed by temporal domain analysis, overlooking the valuable envelope structures of the frequency domain that are equally important for robust emo…

Cited by 0SourceScholar
2024

A Density-driven Iterative Prototype Optimization for Transductive Few-shot Learning

IJCAI 2024poster

Few-shot learning (FSL) poses a considerable challenge since it aims to improve the model generalization ability with limited labeled data. Previous works usually attempt to construct class-specific prototypes and then predict novel classes using these prototypes. However, the feature distribution r…

2024

CLAP: Collaborative Adaptation for Patchwork Learning

ICLR 2024spotlight

In this paper, we investigate a new practical learning scenario, where the data distributed in different sources/clients are typically generated with various modalities. Existing research on learning from multi-source data mostly assume that each client owns the data of all modalities, which may lar…

Cited by 1SourcePDFScholar
2024

CM-AVAE: Cross-Modal Adversarial Variational Autoencoder for Visual-to-Tactile Data Generation

RA-L 2024

Vibration acceleration signals allow humans to perceive the surface characteristics of textures during tool-surface interactions. However, acquiring acceleration signals requires a specialized system, which is relatively expensive. Conversely, visual images are more accessible than acceleration sign

Cited by 7SourceScholar
2024

Cognitive Overload: Jailbreaking Large Language Models with Overloaded Logical Thinking

NAACL 2024findings

While large language models (LLMs) have demonstrated increasing power, they have also called upon studies on their vulnerabilities. As representatives, jailbreak attacks can provoke harmful or unethical responses from LLMs, even after safety alignment. In this paper, we investigate a novel category…

2024

Contrastive Instruction Tuning

ACL 2024findings

Instruction tuning has been used as a promising approach to improve the performance of large language models (LLMs) on unseen tasks. However, current LLMs exhibit limited robustness to unseen instructions, generating inconsistent outputs when the same instruction is phrased with slightly varied form…

2024

Conversational Question Answering with Language Models Generated Reformulations over Knowledge Graph

ACL 2024findings

Conversational question answering (ConvQA) over knowledge graphs (KGs) involves answering multi-turn natural language questions about information contained in a KG. State-of-the-art methods of ConvQA often struggle with inexplicit question-answer pairs. These inputs are easy for human beings to unde…

Cited by 0SourcePDFScholar
2024

Cross-Modal Registration Using Adaptive Modeling in Infrastructure-based Vehicle Localization*

ICRA 2024poster

Infrastructure-based vehicle localization, in comparison to single-agent approaches, offers several advantages including reduced system cost, extended perception range, enhanced data fusion capabilities, and energy savings. Many conventional approaches impose limitations on the types of objects due…

Cited by 1SourceScholar
2024

Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models

EMNLP 2024main

Data are crucial element in large language model (LLM) alignment. Recent studies have explored using LLMs for efficient data collection. However, LLM-generated data often suffers from quality issues, with underrepresented or absent aspects and low-quality datapoints. To address these problems, we pr…

2024

Deceptive Semantic Shortcuts on Reasoning Chains: How Far Can Models Go without Hallucination?

NAACL 2024long

Despite the high performances of large language models (LLMs) across numerous benchmarks, recent research has unveiled their suffering from hallucinations and unfaithful reasoning. This work studies a type of hallucination induced by semantic associations. We investigate to what extent LLMs take sho…

2024

Dynamic Frequency Domain Graph Convolutional Network for Traffic Forecasting

ICASSP 2024accepted

Complex spatial dependencies in transportation networks make traffic prediction extremely challenging. Much existing work is devoted to learning dynamic graph structures among sensors, and the strategy of mining spatial dependencies from traffic data, known as data-driven, tends to be an intuitive a…

Cited by 0SourceScholar
2024

EulerMormer: Robust Eulerian Motion Magnification via Dynamic Filtering within Transformer

AAAI 2024technical

Video Motion Magnification (VMM) aims to break the resolution limit of human visual perception capability and reveal the imperceptible minor motion that contains valuable information in the macroscopic domain. However, challenges arise in this task due to photon noise inevitably introduced by photog…

2024

Frequency Decoupling for Motion Magnification via Multi-Level Isomorphic Architecture

CVPR 2024poster

Video Motion Magnification (VMM) aims to reveal subtle and imperceptible motion information of objects in the macroscopic world. Prior methods directly model the motion field from the Eulerian perspective by Representation Learning that separates shape and texture or Multi-domain Learning from phase…

2024

From Shortcuts to Triggers: Backdoor Defense with Denoised PoE

NAACL 2024long

Language models are often at risk of diverse backdoor attacks, especially data poisoning. Thus, it is important to investigate defense solutions for addressing them. Existing backdoor defense methods mainly focus on backdoor attacks with explicit triggers, leaving a universal defense against various…

2024

Instructional Fingerprinting of Large Language Models

NAACL 2024long

The exorbitant cost of training Large language models (LLMs) from scratch makes it essential to fingerprint the models to protect intellectual property via ownership authentication and to ensure downstream users and developers comply with their license terms (eg restricting commercial use). In this…

2024

Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models

NAACL 2024long

We investigate security concerns of the emergent instruction tuning paradigm, that models are trained on crowdsourced datasets with task instructions to achieve superior performance. Our studies demonstrate that an attacker can inject backdoors by issuing very few malicious instructions (~1000 token…

2024

LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing

EMNLP 2024main

Claim: This work is not advocating the use of LLMs for paper (meta-)reviewing. Instead, wepresent a comparative analysis to identify and distinguish LLM activities from human activities. Two research goals: i) Enable better recognition of instances when someone implicitly uses LLMs for reviewing act…

2024

Local Discovery by Partitioning: Polynomial-Time Causal Discovery Around Exposure-Outcome Pairs

UAI 2024poster

Causal discovery is crucial for causal inference in observational studies, as it can enable the identification of *valid adjustment sets* (VAS) for unbiased effect estimation. However, global causal discovery is notoriously hard in the nonparametric setting, with exponential time and sample complex…

2024

MassSpecGym: A benchmark for the discovery and identification of molecules

NeurIPS 2024spotlight

The discovery and identification of molecules in biological and environmental samples is crucial for advancing biomedical and chemical sciences. Tandem mass spectrometry (MS/MS) is the leading technique for high-throughput elucidation of molecular structures. However, decoding a molecular structure…

2024

Monotonic Paraphrasing Improves Generalization of Language Model Prompting

EMNLP 2024finding

Performance of large language models (LLMs) may vary with different prompts or instructions of even the same task. One commonly recognized factor for this phenomenon is the model’s familiarity with the given prompt or instruction, which is typically estimated by its perplexity. However, finding the…

2024

Multi-View Subspace Clustering With Consensus Graph Contrastive Learning

ICASSP 2024accepted

A significant challenge in multi-view clustering lies in the comprehensive extraction of consistency and complementary information from heterogeneous multi-view data. Numerous methods employ contrastive learning techniques to explore the information between views. However, the basic contrastive lear…

Cited by 0SourceScholar
2024

Person-in-WiFi 3D: End-to-End Multi-Person 3D Pose Estimation with Wi-Fi

CVPR 2024poster

Wi-Fi signals in contrast to cameras offer privacy protection and occlusion resilience for some practical scenarios such as smart homes elderly care and virtual reality. Recent years have seen remarkable progress in the estimation of single-person 2D pose single-person 3D pose and multi-person 2D po…

Cited by 15SourcePDFScholar
2024

Self-Improvement Programming for Temporal Knowledge Graph Question Answering

COLING 2024main

Temporal Knowledge Graph Question Answering (TKGQA) aims to answer questions with temporal intent over Temporal Knowledge Graphs (TKGs). The core challenge of this task lies in understanding the complex semantic information regarding multiple types of time constraints (e.g., before, first) in questi…

Cited by 9SourcePDFScholar
2024

Towards Disease-Aware Self-Supervised Dynamic Brain Network Learning For Mental Diagnosis

ICASSP 2024accepted

The dynamic brain network learning methods ignored the separation of redundant disease-irrelevant information, resulting in the model only achieving suboptimal diagnosis results. Meanwhile, the supervised learning scheme inevitably suffers from poor generalization due to the limited data. To address…

Cited by 0SourceScholar
2024

Unified Insights: Harnessing Multi-modal Data for Phenotype Imputation via View Decoupling

NeurIPS 2024poster

Phenotype imputation plays a crucial role in improving comprehensive and accurate medical evaluation, which in turn can optimize patient treatment and bolster the reliability of clinical research. Despite the adoption of various techniques, multi-modal biological data, which can provide crucial insi…

Cited by 0SourcePDFScholar
2024

mDPO: Conditional Preference Optimization for Multimodal Large Language Models

EMNLP 2024main

Direct preference optimization (DPO) has shown to be an effective method for large language model (LLM) alignment. Recent works have attempted to apply DPO to multimodal scenarios but have found it challenging to achieve consistent improvement. Through a comparative experiment, we identify the uncon…

2023

A Causal View of Entity Bias in (Large) Language Models

EMNLP 2023long findings

Entity bias widely affects pretrained (large) language models, causing them to rely on (biased) parametric knowledge to make unfaithful predictions. Although causality-inspired methods have shown great potential to mitigate entity bias, it is hard to precisely estimate the parameters of underlying c…

Cited by 0SourcecodeScholar
2023

Boosting Novel Category Discovery Over Domains with Soft Contrastive Learning and All in One Classifier

ICCV 2023oral

Unsupervised domain adaptation (UDA) has proven to be highly effective in transferring knowledge from a label-rich source domain to a label-scarce target domain. However, the presence of additional novel categories in the target domain has led to the development of open-set domain adaptation (ODA) a…

Cited by 19PDFcodeScholar
2023

DamoFD: Digging into Backbone Design on Face Detection

ICLR 2023poster

Face detection (FD) has achieved remarkable success over the past few years, yet, these leaps often arrive when consuming enormous computation costs. Moreover, when considering a realistic situation, i.e., building a lightweight face detector under a computation-scarce scenario, such heavy computati…

2023

Dense Retrieval as Indirect Supervision for Large-space Decision Making

EMNLP 2023long findings

Many discriminative natural language understanding (NLU) tasks have large label spaces. Learning such a process of large-space decision making is particularly challenging due to the lack of training instances per label and the difficulty of selection among many fine-grained labels. Inspired by dense…

Cited by 0SourcecodeScholar
2023

Design and Control of a Snake Robot With a Gripper for Inspection and Maintenance in Narrow Spaces

RA-L 2023

This letter presents a snake robot with a gripper for inspection and maintenance in narrow spaces. The proposed robot has a gripper equipped with a camera and a laser distance sensor that can inspect the surroundings and grasp objects. The control methods of the proposed robot consist of three parts

Cited by 7SourceScholar
2023

Exploring Language-Agnostic Speech Representations Using Domain Knowledge for Detecting Alzheimer's Dementia

ICASSP 2023accepted

We explore ways to use speech data to screen for indications of Alzheimer’s dementia (AD). In particular, we describe our approach to the ICASSP 2023 Signal Processing Grand Challenge, which involves extrapolating from models learned from English speech samples, to Greek speech samples, to determine…

Cited by 0SourceScholar
2023

FairLISA: Fair User Modeling with Limited Sensitive Attributes Information

NeurIPS 2023poster

User modeling techniques profile users' latent characteristics (e.g., preference) from their observed behaviors, and play a crucial role in decision-making. Unfortunately, traditional user models may unconsciously capture biases related to sensitive attributes (e.g., gender) from behavior data, even…

2023

Improving Factuality of Abstractive Summarization without Sacrificing Summary Quality

ACL 2023short

Improving factual consistency of abstractive summarization has been a widely studied topic. However, most of the prior works on training factuality-aware models have ignored the negative effect it has on summary quality. We propose {pasted macro ‘MODEL’}name (i.e. Effective Factual Summarization), a…

2023

InfoDiffusion: Representation Learning Using Information Maximizing Diffusion Models

ICML 2023poster

While diffusion models excel at generating high-quality samples, their latent variables typically lack semantic meaning and are not suitable for representation learning. Here, we propose InfoDiffusion, an algorithm that augments diffusion models with low-dimensional latent variables that capture hig…

Cited by 41SourcePDFScholar
2023

Knowledge Diffusion for Distillation

NeurIPS 2023poster

The representation gap between teacher and student is an emerging topic in knowledge distillation (KD). To reduce the gap and improve the performance, current methods often resort to complicated training schemes, loss functions, and feature alignments, which are task-specific and feature-specific. I…

2023

Knowledge-Graph Augmented Music Representation for Genre Classification

ICASSP 2023accepted

In this paper, we propose KGenre, a knowledge-embedded music representation learning framework for improved genre classification. We construct the knowledge graph from the metadata in the open-source FMA-medium and OpenMIC-2018 datasets, with no extra information/effort required. KGenre then mines t…

Cited by 0SourceScholar
2023

Learning Attention from Attention: Efficient Self-Refinement Transformer for Face Super-Resolution

IJCAI 2023poster

Recently, Transformer-based architecture has been introduced into face super-resolution task due to its advantage in capturing long-range dependencies. However, these approaches tend to integrate global information in a large searching region, which neglect to focus on the most relevant information…

2023

Local-to-Global Registration for Bundle-Adjusting Neural Radiance Fields

CVPR 2023poster

Neural Radiance Fields (NeRF) have achieved photorealistic novel views synthesis; however, the requirement of accurate camera poses limits its application. Despite analysis-by-synthesis extensions for jointly learning neural 3D representations and registering camera frames exist, they are susceptibl…

Cited by 77SourcePDFScholar
2023

Masked Distillation with Receptive Tokens

ICLR 2023poster

Distilling from the feature maps can be fairly effective for dense prediction tasks since both the feature discriminability and localization information can be well transferred. However, not every pixel contributes equally to the performance, and a good student should learn from what really matters…

2023

Robust Natural Language Understanding with Residual Attention Debiasing

ACL 2023findings

Natural language understanding (NLU) models often suffer from unintended dataset biases. Among bias mitigation methods, ensemble-based debiasing methods, especially product-of-experts (PoE), have stood out for their impressive empirical success. However, previous ensemble-based debiasing methods typ…

2023

SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation

CVPR 2023poster

Generating talking head videos through a face image and a piece of speech audio still contains many challenges. i.e., unnatural head movement, distorted expression, and identity modification. We argue that these issues are mainly caused by learning from the coupled 2D motion fields. On the other han…

2023

SimMatchV2: Semi-Supervised Learning with Graph Consistency

ICCV 2023poster

Semi-Supervised image classification is one of the most fundamental problem in computer vision, which significantly reduces the need for human labor. In this paper, we introduce a new semi-supervised learning algorithm - SimMatchV2, which formulates various consistency regularizations between labele…

Cited by 13PDFcodeScholar
2023

UV Volumes for Real-Time Rendering of Editable Free-View Human Performance

CVPR 2023poster

Neural volume rendering enables photo-realistic renderings of a human performer in free-view, a critical task in immersive VR/AR applications. But the practice is severely limited by high computational costs in the rendering process. To solve this problem, we propose the UV Volumes, a new approach t…

2022

A Keypoint-Based Global Association Network for Lane Detection

CVPR 2022poster

Lane detection is a challenging task that requires predicting complex topology shapes of lane lines and distinguishing different types of lanes simultaneously. Earlier works follow a top-down roadmap to regress predefined anchors into various shapes of lane lines, which lacks enough flexibility to f…

Cited by 154PDFcodeScholar
2022

BiSyn-GAT+: Bi-Syntax Aware Graph Attention Network for Aspect-based Sentiment Analysis

ACL 2022findings

Aspect-based sentiment analysis (ABSA) is a fine-grained sentiment analysis task that aims to align aspects and corresponding sentiments for aspect-specific sentiment polarity inference. It is challenging because a sentence may contain multiple aspects or complicated (e.g., conditional, coordinating…

2022

Contrastive Graph Structure Learning via Information Bottleneck for Recommendation

NeurIPS 2022accept

Graph convolution networks (GCNs) for recommendations have emerged as an important research topic due to their ability to exploit higher-order neighbors. Despite their success, most of them suffer from the popularity bias brought by a small number of active users and popular items. Also, a real-worl…

Cited by 72SourcePDFScholar
2022

Data Agnostic Filter Gating For Efficient Deep Networks

ICASSP 2022accepted

Filter pruning is essential for deploying a well-trained CNN model on edge computation devices with a target computation budget (e.g., FLOPs). Current filter pruning methods mainly focus on leveraging feature maps to analyze the importance of filters, and prune those with less impact on the value of…

Cited by 0SourceScholar
2022

Deep Recurrent Neural Network with Multi-Scale Bi-directional Propagation for Video Deblurring

AAAI 2022technical

The success of the state-of-the-art video deblurring methods stems mainly from implicit or explicit estimation of alignment among the adjacent frames for latent video restoration. However, due to the influence of the blur effect, estimating the alignment information from the blurry adjacent frames i…

2022

Does Your Model Classify Entities Reasonably? Diagnosing and Mitigating Spurious Correlations in Entity Typing

EMNLP 2022main

Entity typing aims at predicting one or more words that describe the type(s) of a specific mention in a sentence. Due to shortcuts from surface patterns to annotated entity labels and biased training, existing entity typing models are subject to the problem of spurious correlations. To comprehensive…

2022

DyRep: Bootstrapping Training With Dynamic Re-Parameterization

CVPR 2022poster

Structural re-parameterization (Rep) methods achieve noticeable improvements on simple VGG-style networks. Despite the prevalence, current Rep methods simply re-parameterize all operations into an augmented network, including those that rarely contribute to the model's performance. As such, the pric…

Cited by 42PDFcodeScholar
2022

Face2Exp: Combating Data Biases for Facial Expression Recognition

CVPR 2022poster

Facial expression recognition (FER) is challenging due to the class imbalance caused by data collection. Existing studies tackle the data bias problem using only labeled facial expression dataset. Orthogonal to existing FER methods, we propose to utilize large unlabeled face recognition (FR) dataset…

Cited by 129PDFcodeScholar
2022

GreedyNASv2: Greedier Search With a Greedy Path Filter

CVPR 2022poster

Training a good supernet in one-shot NAS methods is difficult since the search space is usually considerably huge (e.g., 13^ 21 ). In order to enhance the supernet's evaluation ability, one greedy strategy is to sample good paths, and let the supernet lean towards the good ones and ease its evaluati…

Cited by 22PDFScholar
2022

Green Hierarchical Vision Transformer for Masked Image Modeling

NeurIPS 2022accept

We present an efficient approach for Masked Image Modeling (MIM) with hierarchical Vision Transformers (ViTs), allowing the hierarchical ViTs to discard masked patches and operate only on the visible ones. Our approach consists of three key designs. First, for window attention, we propose a Group Wi…

2022

HEAD: HEtero-Assists Distillation for Heterogeneous Object Detectors

ECCV 2022poster

"Conventional knowledge distillation (KD) methods for object detection mainly concentrate on homogeneous teacher-student detectors. However, the design of a lightweight detector for deployment is often significantly different from a high-capacity detector. Thus, we investigate KD among heterogeneous…

2022

IDPT: Interconnected Dual Pyramid Transformer for Face Super-Resolution

IJCAI 2022poster

Face Super-resolution (FSR) task works for generating high-resolution (HR) face images from the corresponding low-resolution (LR) inputs, which has received a lot of attentions because of the wide application prospects. However, due to the diversity of facial texture and the difficulty of reconstruc…

Cited by 19SourcePDFScholar
2022

Learning Where To Learn in Cross-View Self-Supervised Learning

CVPR 2022poster

Self-supervised learning (SSL) has made enormous progress and largely narrowed the gap with the supervised ones, where the representation learning is mainly guided by a projection into an embedding space. During the projection, current methods simply adopt uniform aggregation of pixels for embedding…

Cited by 47PDFcodeScholar
2022

MogFace: Towards a Deeper Appreciation on Face Detection

CVPR 2022poster

Benefiting from the pioneering design of generic object detectors, significant achievements have been made in the field of face detection. Typically, the architectures of the backbone, feature pyramid layer, and detection head module within the face detector all assimilate the excellent experience f…

Cited by 36PDFcodeScholar
2022

Robust (Controlled) Table-to-Text Generation with Structure-Aware Equivariance Learning

NAACL 2022long

Controlled table-to-text generation seeks to generate natural language descriptions for highlighted subparts of a table. Previous SOTA systems still employ a sequence-to-sequence generation method, which merely captures the table as a linear structure and is brittle when table layouts change. We see…

2022

Salience Allocation as Guidance for Abstractive Summarization

EMNLP 2022main

Abstractive summarization models typically learn to capture the salient information from scratch implicitly.Recent literature adds extractive summaries as guidance for abstractive summarization models to provide hints of salient content and achieves better performance.However, extractive summaries a…

2022

SimMatch: Semi-Supervised Learning With Similarity Matching

CVPR 2022poster

Learning with few labeled data has been a longstanding problem in the computer vision and machine learning research community. In this paper, we introduced a new semi-supervised learning framework, SimMatch, which simultaneously considers semantic similarity and instance similarity. In SimMatch, the…

Cited by 274PDFcodeScholar
2022

Uncertainty-Aware Learning against Label Noise on Imbalanced Datasets

AAAI 2022technical

Learning against label noise is a vital topic to guarantee a reliable performance for deep neural networks.Recent research usually refers to dynamic noise modeling with model output probabilities and loss values, and then separates clean and noisy samples.These methods have gained notable success. H…

Cited by 50SourcePDFScholar
2022

ViTAS: Vision Transformer Architecture Search

ECCV 2022poster

"Vision transformers (ViTs) inherited the success of NLP but their structures have not been sufficiently investigated and optimized for visual tasks. One of the simplest solutions is to directly search the optimal one via the widely used neural architecture search (NAS) in CNNs. However, we empirica…

2021

Addressing Algorithmic Disparity and Performance Inconsistency in Federated Learning

NeurIPS 2021poster

Federated learning (FL) has gain growing interests for its capability of learning from distributed data sources collectively without the need of accessing the raw data samples across different sources. So far FL research has mostly focused on improving the performance, how the algorithmic disparity…

2021

Adversarial Example Detection Using Latent Neighborhood Graph

ICCV 2021poster

Detection of adversarial examples with high accuracy is critical for the security of deployed deep neural network-based models. We present the first graph-based adversarial detection method that constructs a Latent Neighborhood Graph (LNG) around an input example to determine if the input example is…

Cited by 74PDFScholar
2021

BCNet: Searching for Network Width With Bilaterally Coupled Network

CVPR 2021poster

Searching for a more compact network width recently serves as an effective way of channel pruning for the deployment of convolutional neural networks (CNNs) under hardware constraints. To fulfill the searching, a one-shot supernet is usually leveraged to efficiently evaluate the performance \wrt dif…

Cited by 42PDFScholar
2021

Collaborative Spatial-Temporal Modeling for Language-Queried Video Actor Segmentation

CVPR 2021poster

Language-queried video actor segmentation aims to predict the pixel-level mask of the actor which performs the actions described by a natural language query in the target frames. Existing methods adopt 3D CNNs over the video clip as a general encoder to extract a mixed spatio-temporal feature for th…

Cited by 58PDFScholar
2021

K-shot NAS: Learnable Weight-Sharing for NAS with K-shot Supernets

ICML 2021spotlight

In one-shot weight sharing for NAS, the weights of each operation (at each layer) are supposed to be identical for all architectures (paths) in the supernet. However, this rules out the possibility of adjusting operation weights to cater for different paths, which limits the reliability of the evalu…

Cited by 48SourcePDFScholar
2021

Learning To Restore Hazy Video: A New Real-World Dataset and a New Method

CVPR 2021poster

Most of the existing deep learning-based dehazing methods are trained and evaluated on the image dehazing datasets, where the dehazed images are generated by only exploiting the information from the corresponding hazy ones. On the other hand, the video dehazing algorithms, which can acquire more sat…

Cited by 103PDFScholar
2021

Locally Free Weight Sharing for Network Width Search

ICLR 2021spotlight

Searching for network width is an effective way to slim deep neural networks with hardware budgets. With this aim, a one-shot supernet is usually leveraged as a performance evaluator to rank the performance \wrt~different width. Nevertheless, current methods mainly follow a manually fixed weight sha…

Cited by 45SourcePDFScholar
2021

MFPN-6D : Real-time One-stage Pose Estimation of Objects on RGB Images

ICRA 2021poster

6D pose estimation of objects is an important part of robot grasping. The latest research trend on 6D pose estimation is to train a deep neural network to directly predict the 2D projection position of the 3D key points from the image, establish the corresponding relationship, and finally use Pespec…

Cited by 14SourceScholar
2021

Prioritized Architecture Sampling With Monto-Carlo Tree Search

CVPR 2021poster

One-shot neural architecture search (NAS) methods significantly reduce the search cost by considering the whole search space as one network, which only needs to be trained once. However, current methods select each operation independently without considering previous layers. Besides, the historical…

Cited by 66PDFcodeScholar
2021

ReSSL: Relational Self-Supervised Learning with Weak Augmentation

NeurIPS 2021poster

Self-supervised Learning (SSL) including the mainstream contrastive learning has achieved great success in learning visual representations without data annotations. However, most of methods mainly focus on the instance level information (\ie, the different augmented images of the same instance shoul…

2021

Reformulating HOI Detection As Adaptive Set Prediction

CVPR 2021poster

Determining which image regions to concentrate is critical for Human-Object Interaction (HOI) detection. Conventional HOI detectors focus on either detected human and object pairs or pre-defined interaction locations, which limits learning of the effective features. In this paper, we reformulate HOI…

Cited by 182PDFcodeScholar
2021

Table-based Fact Verification With Salience-aware Learning

EMNLP 2021finding

Tables provide valuable knowledge that can be used to verify textual statements. While a number of works have considered table-based fact verification, direct alignments of tabular data with tokens in textual statements are rarely available. Moreover, training a generalized fact verification model r…

2021

Towards Improving the Consistency, Efficiency, and Flexibility of Differentiable Neural Architecture Search

CVPR 2021poster

Most differentiable neural architecture search methods construct a super-net for search and derive a target-net as its sub-graph for evaluation. There exists a significant gap between the architectures in search and evaluation. As a result, current methods suffer from an inconsistent, inefficient, a…

Cited by 57PDFScholar
2021

Weakly Supervised Contrastive Learning

ICCV 2021poster

Unsupervised visual representation learning has gained much attention from the computer vision community because of the recent achievement of contrastive learning. Most of the existing contrastive learning frameworks adopt the instance discrimination as the pretext task, which treating every single…

Cited by 152PDFcodeScholar
2020

A Real-Time Cross-Modality Correlation Filtering Method for Referring Expression Comprehension

CVPR 2020poster

Referring expression comprehension aims to localize the object instance described by a natural language expression. Current referring expression methods have achieved good performance. However, none of them is able to achieve real-time inference without accuracy drop. The reason for the relatively s…

Cited by 240PDFScholar
2020

Agree to Disagree: Adaptive Ensemble Knowledge Distillation in Gradient Space

NeurIPS 2020poster

Distilling knowledge from an ensemble of teacher models is expected to have a more promising performance than that from a single one. Current methods mainly adopt a vanilla average rule, i.e., to simply take the average of all teacher losses for training the student network. However, this approach t…

2020

CentripetalNet: Pursuing High-Quality Keypoint Pairs for Object Detection

CVPR 2020poster

Keypoint-based detectors have achieved pretty-well performance. However, incorrect keypoint matching is still widespread and greatly affects the performance of the detector. In this paper, we propose CentripetalNet which uses centripetal shift to pair corner keypoints from the same instance. Centrip…

Cited by 220PDFcodeScholar
2020

GreedyNAS: Towards Fast One-Shot NAS With Greedy Supernet

CVPR 2020poster

Training a supernet matters for one-shot neural architecture search (NAS) methods since it serves as a basic performance estimator for different architectures (paths). Current methods mainly hold the assumption that a supernet should give a reasonable ranking over all paths. They thus treat all path…

Cited by 188PDFScholar
2020

ISTA-NAS: Efficient and Consistent Neural Architecture Search by Sparse Coding

NeurIPS 2020poster

Neural architecture search (NAS) aims to produce the optimal sparse solution from a high-dimensional space spanned by all candidate connections. Current gradient-based NAS methods commonly ignore the constraint of sparsity in the search phase, but project the optimized solution onto a sparse one by…

2020

Local Correlation Consistency for Knowledge Distillation

ECCV 2020poster

Sufficient knowledge extraction from the teacher network plays a critical role in the knowledge distillation task to improve the performance of the student network. Existing methods mainly focus on the consistency of instance-level features and their relationships, but neglect the local features and…

Cited by 62SourcePDFScholar
2020

Multi-Scale Boosted Dehazing Network With Dense Feature Fusion

CVPR 2020poster

In this paper, we propose a Multi-Scale Boosted Dehazing Network with Dense Feature Fusion based on the U-Net architecture. The proposed method is designed based on two principles, boosting and error feedback, and we show that they are suitable for the dehazing problem. By incorporating the Strength…

Cited by 1034PDFcodeScholar
2020

PPDM: Parallel Point Detection and Matching for Real-Time Human-Object Interaction Detection

CVPR 2020poster

We propose a single-stage Human-Object Interaction (HOI) detection method that has outperformed all existing methods on HICO-DET dataset at 37 fps on a single Titan XP GPU. It is the first real-time HOI detection method. Conventional HOI detection methods are composed of two stages, i.e., human-obje…

Cited by 341PDFcodeScholar
2019

Deep Comprehensive Correlation Mining for Image Clustering

ICCV 2019poster

Recent developed deep unsupervised methods allow us to jointly learn representation and cluster unlabelled data. These deep clustering methods %like DAC start with mainly focus on the correlation among samples, e.g., selecting high precision pairs to gradually tune the feature representation, which…

Cited by 242PDFcodeScholar
2019

Discriminative Feature Learning With Consistent Attention Regularization for Person Re-Identification

ICCV 2019poster

Person re-identification (Re-ID) has undergone a rapid development with the blooming of deep neural network. Most methods are very easily affected by target misalignment and background clutter in the training process. In this paper, we propose a simple yet effective feedforward attention network to…

Cited by 133PDFScholar
2019

Distributed Power Allocation for Spectral Coexisting Multistatic Radar and Communication Systems Based on Stackelberg Game

ICASSP 2019accepted

This paper studies the problem of Stackelberg game based distributed power allocation for spectral coexisting multistatic radar and communication systems. The strategy aims to minimize the radiated power of each radar by optimizing transmit power allocation for a desired signal-to-interference-plus-…

Cited by 0SourceScholar
2019

Glyce: Glyph-vectors for Chinese Character Representations

NeurIPS 2019poster

It is intuitive that NLP tasks for logographic languages like Chinese should benefit from the use of the glyph information in those languages. However, due to the lack of rich pictographic evidence in glyphs and the weak generalization ability of standard computer vision models on character data, a…

2019

Person-in-WiFi: Fine-Grained Person Perception Using WiFi

ICCV 2019poster

Fine-grained person perception such as body segmentation and pose estimation has been achieved with many 2D and 3D sensors such as RGB/depth cameras, radars (e.g. RF-Pose), and LiDARs. These solutions require 2D images, depth maps or 3D point clouds of person bodies as input. In this paper, we take…

Cited by 215PDFcodeScholar
2018

Backpropagation with Callbacks: Foundations for Efficient and Expressive Differentiable Programming

NeurIPS 2018poster

Training of deep learning models depends on gradient descent and end-to-end differentiation. Under the slogan of differentiable programming, there is an increasing demand for efficient automatic gradient computation for emerging network architectures that incorporate dynamic control flow, especially…

2018

Eigendecomposition-free Training of Deep Networks with Zero Eigenvalue-based Losses

ECCV 2018poster

Many classical Computer Vision problems, such as essential matrix computation and pose estimation from 3D to 2D correspondences, can be solved by finding the eigenvector corresponding to the smallest, or zero, eigenvalue of a matrix representing a linear system. Incorporating this in deep learning f…

Cited by 54SourcePDFScholar
2018

The Devil of Face Recognition is in the Noise

ECCV 2018poster

The growing scale of face recognition datasets empowers us to train strong convolutional networks for face recognition. While a variety of architectures and loss functions have been devised, we still have a limited understanding of the source and consequence of label noise inherent in existing datas…

2017

Residual Attention Network for Image Classification

CVPR 2017spotlight

In this work, we propose "Residual Attention Network", a convolutional neural network using attention mechanism which can incorporate with state-of-art feed forward network architecture in an end-to-end training fashion. Our Residual Attention Network is built by stacking Attention Modules which gen…

Cited by 3644PDFScholar
2015

Security information factor based low probability of identification in distributed multiple-radar system

ICASSP 2015accepted

In this study, the problem of low probability of identification (LPID) performance improvement for distributed multiple-radar system (DMRS) is addressed. Firstly, we propose security information factor originating from secrecy capacity to evaluate the LPID performance for DMRS, and derive an explici…

Cited by 0SourceScholar