← Search

Gang Li

77 accepted papers

2026

ACAVCAPS: ENABLING LARGE-SCALE TRAINING FOR FINE-GRAINED AND DIVERSE AUDIO UNDERSTANDING

ICASSP 2026poster

General audio understanding is a fundamental goal for large audio-language models, with audio captioning serving as a cornerstone task for their development. However, progress in this domain is hindered by existing datasets, which lack the scale and descriptive granularity required to train truly ve…

Cited by 0SourcePDFScholar
2026

Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly Detection

CVPR 2026

Zero-shot (ZS) 3D anomaly detection is crucial for reliable industrial inspection, as it enables detecting and localizing defects without requiring any target-category training data. Existing approaches render 3D point clouds into 2D images and leverage pre-trained Vision-Language Models (VLMs) for

Cited by 0SourceScholar
2026

CUARewardBench: Benchmark for Evaluating Reward Models on Computer-using Agent Trajectories

ICML 2026poster

Computer-using agents (CUAs) enable task completion through natural interaction with operating systems and software interfaces. While script-based verifiers are widely adopted for evaluation, they suffer from limited scalability and inability to provide step-wise assessment. Reward models offer prom…

Cited by 0SourceScholar
2026

Certain Head, Uncertain Tail: Expert-Sample for Test-Time Scaling in Fine-Grained MoE

ICML 2026poster

Test-time scaling improves LLM performance by generating multiple candidate solutions, yet token-level sampling requires temperature tuning that trades off diversity against stability. Fine-grained MoE, featuring hundreds of well-trained experts per layer and multi-expert activation per token, offer…

Cited by 0SourceScholar
2026

FedCure: Mitigating Participation Bias in Semi-Asynchronous Federated Learning with Non-IID Data

AAAI 2026technical

While semi-asynchronous federated learning (SAFL) combines the efficiency of synchronous training with the flexibility of asynchronous updates, it inherently suffers from participation bias, which is further exacerbated by non-IID data distributions. More importantly, hierarchical architecture shift

Cited by 0SourcePDFScholar
2026

FedScar: Correcting Geometric Bias for Flatness-Consistent Federated Learning

ICML 2026poster

Federated Learning (FL) often suffers from degraded generalization under statistical heterogeneity, where client updates systematically deviate from the global objective. While recent Sharpness-Aware Minimization (SAM) methods promote locally flat solutions, they implicitly assume that local flatnes…

Cited by 0SourceScholar
2026

FedVeer: Self-Adaptive Skew Estimation for Robust Federated Learning

ICML 2026poster

Federated Learning (FL) enables collaborative model training across decentralized clients, but its performance often degrades under non-IID data distributions, particularly in the presence of data skew. Existing approaches mitigate this issue by estimating client skew via kernel density estimation o…

Cited by 0SourceScholar
2026

Fusing Pixels and Genes: Spatially-Aware Learning in Computational Pathology

ICLR 2026poster

Recent years have witnessed remarkable progress in multimodal learning within computational pathology. Existing models primarily rely on vision and language modalities; however, language alone lacks molecular specificity and offers limited pathological supervision, leading to representational bottle…

Cited by 0SourcecodeScholar
2026

Hilbert Curve-Based Attention Enabling Topology-Preserving Image Tensor Representation for Semantic Segmentation Network

CVPR 2026

Drone-based building defect segmentation remains challenging due to complex surface textures and illumination variations. We propose TPSegformer, a topology-preserving segmentation framework that mitigates mis-segmentation in such scenarios. Its decoder incorporates a Hilbert curve-based topology-pr

Cited by 0SourcecodeScholar
2026

Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning

ICLR 2026poster

Reinforcement learning (RL) is the dominant paradigm for sharpening strategic tool use capabilities of LLMs on long-horizon, sparsely-rewarded agent tasks, yet it faces a fundamental challenge of exploration-exploitation trade-off. Existing studies stimulate exploration through the lens of policy en…

Cited by 0SourcecodeScholar
2026

Learning from Disagreement: A Group Decision Simulation Framework for Robust Medical Image Segmentation

ICASSP 2026poster

Medical image segmentation annotation suffers from inter-rater variability (IRV) due to differences in annotators' expertise and the inherent blurriness of medical images. Standard approaches that simply average expert labels are flawed, as they discard the valuable clinical uncertainty revealed in…

Cited by 0SourcePDFScholar
2026

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks

ICML 2026poster

While large audio-language models have advanced open-ended audio understanding, they still fall short of nuanced human-level comprehension. This gap persists largely because current benchmarks, limited by data annotations and evaluation metrics, fail to reliably distinguish between generic and highl…

Cited by 0SourceScholar
2026

MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation

ICLR 2026poster

Precise recognition, editing, and generation of molecules are essential prerequisites for both chemists and AI systems tackling various chemical tasks. We present MolLangBench, a comprehensive benchmark designed to evaluate fundamental molecule-language interface tasks: language-prompted molecular s…

Cited by 0SourcecodeScholar
2026

OPTION: An Online Pricing Strategy for Asynchronous Federated Learning Against Free-Riding Attacks

AAAI 2026technical

Asynchronous Federated Learning (AFL) is acclaimed for accelerating collaborative training on heterogeneous systems by eliminating the wait for stragglers. While current solutions focus on improving convergence amidst update delays, they neglect how delayed aggregation fosters free-riding attacks, a

Cited by 0SourcePDFScholar
2026

One Skill, Many Websites: Learning Generalizable Skills Through Polymorphic Abstraction

ICLR 2026poster

Large language models (LLMs) are moving beyond static uses and are now powering agents that learn during their interaction with external environments. For example, agents can learn reusable skills while navigating web pages or toggling new tools. However, existing methods for skill learning often cr…

Cited by 0SourcecodeScholar
2026

OursFed: Provable Group Fairness-Aware Federated Learning Against Distrust and Fragility

AAAI 2026technical

With the increasing application of high-stakes decisionmaking application in Federated Learning (FL), ensuring fairness across different populations to prevent biases against certain groups has become crucial. However, achieving group fairness (GF) in FL presents a formidable challenge due to its de

Cited by 0SourcePDFScholar
2026

TOPOGRAPH: Topology-Preserving Graph Reduction with Adaptive Structure for Persistent Homology

AAAI 2026technical

Topological Data Analysis (TDA) provides artificial intelligence (AI) systems with mathematically rigorous geometric descriptors through Persistent Homology (PH), capturing essential shape characteristics in high-dimensional data. Yet, PH’s combinatorial complexity and sensitivity to outliers hinder

Cited by 0SourcePDFScholar
2026

WARC-Bench: Web Archive based Benchmark for GUI Subtask Executions

ICLR 2026poster

Training web agents to navigate complex, real-world websites requires them to master subtasks—short-horizon interactions on multiple UI components (e.g., choosing the correct date in a date picker, or scrolling in a container to extract information). We introduce WARC-Bench (Web Archive Benchmark),…

Cited by 0SourceScholar
2025

Align-A-Video: Deterministic Reward Tuning of Image Diffusion Models for Consistent Video Editing

CVPR 2025poster

Due to control limitations in the denoising process and the lack of training, zero-shot video editing methods often struggle to meet user instructions, resulting in generated videos that are visually unappealing and fail to fully satisfy expectations. To address this problem, we propose Align-A-Vide…

Cited by 0SourcePDFScholar
2025

Bridging Training and Execution via Dynamic Directed Graph-Based Communication in Cooperative Multi-Agent Systems

AAAI 2025technical

Multi-agent systems must learn to communicate and understand interactions between agents to achieve cooperative goals in partially observed tasks. However, existing approaches lack a dynamic directed communication mechanism and rely on global states, thus diminishing the role of communication in cen…

2025

Can Language Models Capture Human Writing Preferences for Domain-Specific Text Summarization?

ACL 2025finding

With the popularity of large language models and their high-quality text generation capabilities, researchers are using them as auxiliary tools for text summary writing. Although summaries generated by these large language models are smooth and capture key information sufficiently, the quality of th…

2025

DaringFed: A Dynamic Bayesian Persuasion Pricing for Online Federated Learning Under Two-sided Incomplete Information

IJCAI 2025

Online Federated Learning (OFL) is a real-time learning paradigm that sequentially executes parameter aggregation immediately for each random arriving client. To motivate clients to participate in OFL, it is crucial to offer appropriate incentives to offset the training resource consumption. However

Cited by 0SourcePDFScholar
2025

Denoising Trajectory Biases for Zero-Shot AI-Generated Image Detection

NeurIPS 2025poster

The rapid advancement of generative models has led to the widespread emergence of highly realistic synthetic images, making the detection of AI-generated content increasingly critical. In particular, diffusion models have recently achieved unprecedented levels of visual fidelity, further raising con…

Cited by 0SourceScholar
2025

DisCO: Reinforcing Large Reasoning Models with Discriminative Constrained Optimization

NeurIPS 2025poster

The recent success and openness of DeepSeek-R1 have brought widespread attention to Group Relative Policy Optimization (GRPO) as a reinforcement learning method for large reasoning models (LRMs). In this work, we analyze the GRPO objective under a binary reward setting and reveal an inherent limitat…

Cited by 0SourcecodeScholar
2025

Exploring and Detecting Self-disclosure in Multi-modal posts on Chinese Social Media

EMNLP 2025

Self-disclosure can provide psychological comfort and social support, but it also carries the risk of unintentionally revealing sensitive information, leading to serious privacy concerns. Research on self-disclosure in Chinese multimodal contexts remains limited, lacking high-quality corpora, analys

2025

Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation

CVPR 2025highlight

Text-to-image (T2I) generation has made significant advances in recent years, but challenges still remain in the generation of perceptual artifacts, misalignment with complex prompts, and safety. The prevailing approach to address these issues involves collecting human feedback on generated images,…

Cited by 2SourcePDFScholar
2025

Incentivizing Reasoning for Advanced Instruction-Following of Large Language Models

NeurIPS 2025poster

Existing large language models (LLMs) face challenges of following complex instructions, especially when multiple constraints are present and organized in paralleling, chaining, and branching structures. One intuitive solution, namely chain-of-thought (CoT), is expected to universally improve capabi…

Cited by 0SourcecodeScholar
2025

Incentivizing Truthful Language Models via Peer Elicitation Games

NeurIPS 2025poster

Large Language Models (LLMs) have demonstrated strong generative capabilities but remain prone to inconsistencies and hallucinations. We introduce Peer Elicitation Games (PEG), a training-free, game-theoretic framework for aligning LLMs through a peer elicitation mechanism involving a generator and…

Cited by 0SourcecodeScholar
2024

CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models

NeurIPS 2024poster

Artificial intelligence has significantly impacted medical applications, particularly with the advent of Medical Large Vision Language Models (Med-LVLMs), sparking optimism for the future of automated and personalized healthcare. However, the trustworthiness of Med-LVLMs remains unverified, posing s…

2024

CartoonDiff: Training-free Cartoon Image Generation with Diffusion Transformer Models

ICASSP 2024accepted

Image cartoonization has attracted significant interest in the field of image generation. However, most of the existing image cartoonization techniques require re-training models using images of cartoon style. In this paper, we present CartoonDiff, a novel training-free sampling approach which gener…

Cited by 0SourceScholar
2024

Crosstalk-Free Impedance-Separating Array Measurement for Iontronic Tactile Sensors*

ICRA 2024poster

Iontronic tactile sensors are promising to measure spatial-temporal contact information with high performance. However, no suitable measuring method has been presented, due to issues with crosstalk and non-negligible equivalent resistance. Hence, this study presents an impedance-separating method, w…

Cited by 0SourceScholar
2024

Detecting Change Intervalswith Isolation Distributional Kernel (Abstract Reprint)

IJCAI 2024poster

Detecting abrupt changes in data distribution is one of the most significant tasks in streaming data analysis. Although many unsupervised Change-Point Detection (CPD) methods have been proposed recently to identify those changes, they still suffer from missing subtle changes, poor scalability, or/an…

Cited by 0SourcePDFScholar
2024

FPGNet: Single Image Deraining with High-Frequency Channel and Frequency Domain Prior Guidance

ICASSP 2024accepted

In recent years, deep learning methods have shown promising results in Single Image Deraining (SID). However, these methods still suffer from unsatisfactory residual rain streaks, primarily due to the absence of image priors embedding and limitations in modeling capacity. In this paper, we propose a…

Cited by 0SourceScholar
2024

On Scaling Up a Multilingual Vision and Language Model

CVPR 2024poster

We explore the boundaries of scaling up a multilingual vision and language model both in terms of size of the components and the breadth of its training task mixture. Our model achieves new levels of performance on a wide-range of varied and complex tasks including multiple image-based captioning an…

Cited by 8SourcePDFScholar
2024

Parrot: Pareto-optimal Multi-Reward Reinforcement Learning Framework for Text-to-Image Generation

ECCV 2024oral

"Recent works have demonstrated that using reinforcement learning (RL) with multiple quality rewards can improve the quality of generated images in text-to-image (T2I) generation. However, manually adjusting reward weights poses challenges and may cause over-optimization in certain metrics. To solve…

Cited by 22SourcePDFScholar
2024

PediatricsGPT: Large Language Models as Chinese Medical Assistants for Pediatric Applications

NeurIPS 2024poster

Developing intelligent pediatric consultation systems offers promising prospects for improving diagnostic efficiency, especially in China, where healthcare resources are scarce. Despite recent advances in Large Language Models (LLMs) for Chinese medicine, their performance is sub-optimal in pediatri…

2024

RULE: Reliable Multimodal RAG for Factuality in Medical Vision Language Models

EMNLP 2024main

The recent emergence of Medical Large Vision Language Models (Med-LVLMs) has enhanced medical diagnosis. However, current Med-LVLMs frequently encounter factual issues, often generating responses that do not align with established medical facts. Retrieval-Augmented Generation (RAG), which utilizes e…

2024

Radardiff: Improving Sea Clutter Suppression Using Diffusion Models for Radar Images

ICASSP 2024accepted

Marine radar is employed across multiple fields, notably in navigation, meteorology, defense, and security. Marine radar images are highly sensitive to sea clutter, highlighting the crucial importance of sea clutter suppression in radar image processing. However, existing algorithms for sea clutter…

Cited by 0SourceScholar
2024

Rich Human Feedback for Text-to-Image Generation

CVPR 2024poster

Recent Text-to-Image (T2I) generation models such as Stable Diffusion and Imagen have made significant progress in generating high-resolution images based on text descriptions. However many generated images still suffer from issues such as artifacts/implausibility misalignment with text descriptions…

2024

SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery Detection

NeurIPS 2024poster

Detection of face forgery videos remains a formidable challenge in the field of digital forensics, especially the generalization to unseen datasets and common perturbations. In this paper, we tackle this issue by leveraging the synergy between audio and visual speech elements, embarking on a novel a…

Cited by 1SourcePDFScholar
2024

StraightPCF: Straight Point Cloud Filtering

CVPR 2024poster

Point cloud filtering is a fundamental 3D vision task which aims to remove noise while recovering the underlying clean surfaces. State-of-the-art methods remove noise by moving noisy points along stochastic trajectories to the clean surfaces. These methods often require regularization within the tra…

2024

SumSurvey: An Abstractive Dataset of Scientific Survey Papers for Long Document Summarization

ACL 2024findings

With the popularity of large language models (LLMs) and their ability to handle longer input documents, there is a growing need for high-quality long document summarization datasets. Although many models already support 16k input, current lengths of summarization datasets are inadequate, and salient…

2024

UniAR: A Unified model for predicting human Attention and Responses on visual content

NeurIPS 2024poster

Progress in human behavior modeling involves understanding both implicit, early-stage perceptual behavior, such as human attention, and explicit, later-stage behavior, such as subjective preferences or likes. Yet most prior research has focused on modeling implicit and explicit human behavior in iso…

Cited by 2SourcePDFScholar
2023

$\rm A^2Q$: Aggregation-Aware Quantization for Graph Neural Networks

ICLR 2023poster

As graph data size increases, the vast latency and memory consumption during inference pose a significant challenge to the real-world deployment of Graph Neural Networks (GNNs). While quantization is a powerful approach to reducing GNNs complexity, most previous works on GNNs quantization fail to ex…

2023

A Zero-Shot Language Agent for Computer Control with Structured Reflection

EMNLP 2023long findings

Large language models (LLMs) have shown increasing capacity at planning and executing a high-level goal in a live computer environment (e.g. MiniWoB++). To perform a task, recent works often require a model to learn from trace examples of the task via either supervised learning or few/many-shot prom…

Cited by 0SourceScholar
2023

Do We Need an Encoder-Decoder to Model Dynamical Systems on Networks?

IJCAI 2023poster

As deep learning gains popularity in modelling dynamical systems, we expose an underappreciated misunderstanding relevant to modelling dynamics on networks. Strongly influenced by graph neural networks, latent vertex embeddings are naturally adopted in many neural dynamical network models. However,…

2023

HFMRE: Constructing Huffman Tree in Bags to Find Excellent Instances for Distantly Supervised Relation Extraction

EMNLP 2023long findings

Since the introduction of distantly supervised relation extraction methods, numerous approaches have been developed, the most representative of which is multi-instance learning (MIL). To find reliable features that are most representative of multi-instance bags, aggregation strategies such as AVG (a…

Cited by 0SourceScholar
2023

IterativePFN: True Iterative Point Cloud Filtering

CVPR 2023poster

The quality of point clouds is often limited by noise introduced during their capture process. Consequently, a fundamental 3D vision task is the removal of noise, known as point cloud filtering or denoising. State-of-the-art learning based methods focus on training neural networks to infer filtered…

2023

Maximization of Average Precision for Deep Learning with Adversarial Ranking Robustness

NeurIPS 2023spotlight

This paper seeks to address a gap in optimizing Average Precision (AP) while ensuring adversarial robustness, an area that has not been extensively explored to the best of our knowledge. AP maximization for deep learning has widespread applications, particularly when there is a significant imbalance…

2023

PLay: Parametrically Conditioned Layout Generation using Latent Diffusion

ICML 2023poster

Layout design is an important task in various design fields, including user interfaces, document, and graphic design. As this task requires tedious manual effort by designers, prior works have attempted to automate this process using generative models, but commonly fell short of providing intuitive…

Cited by 31SourcePDFScholar
2023

TEFISTA-NET: GTD Parameter Estimation of Low-Frequency Ultra- Wideband Radar via Model-Based Deep Learning

ICASSP 2023accepted

The geometrical theory of diffraction (GTD) has been widely investigated to describe the target scattering behaviors with the low-frequency ultra-wideband (LFW) radar. In this paper, we propose a new model-based deep learning method for GTD parameter estimation. The proposed method is designed by un…

Cited by 0SourceScholar
2022

DTG-SSOD: Dense Teacher Guidance for Semi-Supervised Object Detection

NeurIPS 2022accept

The Mean-Teacher (MT) scheme is widely adopted in semi-supervised object detection (SSOD). In MT, sparse pseudo labels, offered by the final predictions of the teacher (e.g., after Non Maximum Suppression (NMS) post-processing), are adopted for the dense supervision for the student via hand-crafted…

Cited by 28SourcePDFScholar
2022

Knowledge Distillation for Object Detection via Rank Mimicking and Prediction-Guided Feature Imitation

AAAI 2022technical

Knowledge Distillation (KD) is a widely-used technology to inherit information from cumbersome teacher models to compact student models, consequently realizing model compression and acceleration. Compared with image classification, object detection is a more complex task, and designing specific KD m…

Cited by 101SourcePDFScholar
2022

Multi-block-Single-probe Variance Reduced Estimator for Coupled Compositional Optimization

NeurIPS 2022accept

Variance reduction techniques such as SPIDER/SARAH/STORM have been extensively studied to improve the convergence rates of stochastic non-convex optimization, which usually maintain and update a sequence of estimators for a single function across iterations. What if we need to track multiple functi…

Cited by 22SourcePDFScholar
2022

PalQuant: Accelerating High-Precision Networks on Low-Precision Accelerators

ECCV 2022poster

"Recently low-precision deep learning accelerators (DLAs) have become popular due to their advantages in chip area and energy consumption, yet the low-precision quantized models on these DLAs bring in severe accuracy degradation. One way to achieve both high accuracy and efficient inference is to de…

2022

PseCo: Pseudo Labeling and Consistency Training for Semi-Supervised Object Detection

ECCV 2022poster

"In this paper, we delve into two key techniques in Semi-Supervised Object Detection (SSOD), namely pseudo labeling and consistency training. We observe that these two techniques currently neglect some important properties of object detection, hindering efficient learning on unlabeled data. Specific…

2022

SemMAE: Semantic-Guided Masking for Learning Masked Autoencoders

NeurIPS 2022accept

Recently, significant progress has been made in masked image modeling to catch up to masked language modeling. However, unlike words in NLP, the lack of semantic decomposition of images still makes masked autoencoding (MAE) different between vision and language. In this paper, we explore a potential…

2022

When AUC meets DRO: Optimizing Partial AUC for Deep Learning with Non-Convex Convergence Guarantee

ICML 2022spotlight

In this paper, we propose systematic and efficient gradient-based methods for both one-way and two-way partial AUC (pAUC) maximization that are applicable to deep learning. We propose new formulations of pAUC surrogate objectives by using the distributionally robust optimization (DRO) to define the…

Cited by 37SourcePDFScholar
2021

An Integrated High-dexterity Cooperative Robotic Assistant for Intraocular Micromanipulation*

ICRA 2021

Retinal surgeons are required to manipulate multiple surgical instruments in a confined intraocular space, while the instruments are constrained at the small incisions made on the sclera. Furthermore, physiological hand tremor can affect the precision of the instrument motion. The Steady-Hand Eye Ro

Cited by 3SourceScholar
2021

Learnable Fourier Features for Multi-dimensional Spatial Positional Encoding

NeurIPS 2021poster

Attentional mechanisms are order-invariant. Positional encoding is a crucial component to allow attention-based deep model architectures such as Transformer to address sequences or images where the position of information matters. In this paper, we propose a novel positional encoding method based on…

Cited by 112SourcePDFScholar
2021

Towards Safe In Situ Needle Manipulation for Robot Assisted Lumbar Injection in Interventional MRI

IROS 2021poster

Lumbar injection is an image-guided procedure performed manually for diagnosis and treatment of lower back pain and leg pain. Previously, we have developed and verified an MR-Conditional robotic solution to assisting the needle insertion process. Drawing on our clinical experiences, a virtual remote…

Cited by 6SourceScholar
2020

A Fully Actuated Body-Mounted Robotic Assistant for MRI-Guided Low Back Pain Injection

ICRA 2020poster

This paper reports the development of a fully actuated body-mounted robotic assistant for MRI-guided low back pain injection. The robot is designed with a 4-DOF needle alignment module and a 2-DOF remotely actuated needle driver module. The 6-DOF fully actuated robot can operate inside the scanner b…

Cited by 23SourceScholar
2020

An Optimized Tilt Mechanism for a New Steady-Hand Eye Robot

IROS 2020poster

Robot-assisted vitreoretinal surgery can filter surgeons' hand tremors and provide safe, accurate tool manipulation. In this paper, we report the design, optimization, and evaluation of a novel tilt mechanism for a new Steady-Hand Eye Robot (SHER). The new tilt mechanism features a four-bar linkage…

Cited by 12SourceScholar
2020

Distributed Detection of Sparse Signals with 1-Bit Data in Two-Level Two-Degree Tree-Structured Sensor Networks

ICASSP 2020accepted

In this paper, we present a new detector for the detection of sparse stochastic signals using 1-bit data in two-level two- degree tree-structured sensor networks (2L-2D TSNs). Related prior work mostly concentrates on parallel sensor networks (PSNs). However, PSNs may sometime become impractical in…

Cited by 0SourceScholar
2020

Fully Actuated Body-Mounted Robotic System for MRI-Guided Lower Back Pain Injections: Initial Phantom and Cadaver Studies

RA-L 2020

This paper reports the improved design, system integration, and initial experimental evaluation of a fully actuated body-mounted robotic system for real-time MRI-guided lower back pain injections. The 6-DOF robot is composed of a 4-DOF needle alignment module and a 2-DOF remotely actuated needle dri

Cited by 19SourceScholar
2018

Training Binary Weight Networks via Semi-Binary Decomposition

ECCV 2018poster

Recently binary weight networks have attracted lots of attentions due to their high computational efficiency and small parameter size. Yet they still suffer from large accuracy drops because of their limited representation capacity. In this paper, we propose a novel semi-binary decomposition method…

Cited by 23SourcePDFScholar