← Search

Fei Zhang

16 accepted papers

2026

From Gradient Volume to Shapley Fairness: Towards Fair Multi-Task Learning

ICLR 2026poster

Multi-task learning often suffers from gradient conflicts, leading to unfair optimization and degraded overall performance. To address this, we present SVFair, a Shapley value-based framework for fair gradient aggregation. We propose two scalable geometric conflict metrics: VolDet, a gram determinan…

Cited by 0SourceScholar
2025

Comprehensive Performance Optimization of a Novel Anthropomorphic Dexterous Hand

RA-L 2025

Grasp and in-hand manipulation of dexterous hands are essential for daily operational tasks. In this study, an optimized design method for a novel anthropomorphic dexterous hand is proposed, which improves the comprehensive performance of grasp and manipulability of the dexterous hand. The constrain

Cited by 1SourceScholar
2025

ConText: Driving In-context Learning for Text Removal and Segmentation

ICML 2025poster

This paper presents the first study on adapting the visual in-context learning (V-ICL) paradigm to optical character recognition tasks, specifically focusing on text removal and segmentation. Most existing V-ICL generalists employ a reasoning-as-reconstruction approach: they turn to using a straight…

2025

Sim-to-Real Transfer of Automatic Extinguishing Strategy for Firefighting Robots

RA-L 2025

The automatic extinguishing strategy (AES) is the core of the decision-making system for intelligent firefighting robots. Inspired by the fire extinguishing action of firefighters, designing a vision-based end-to-end AES aligns with human intuition. However, the cost of training agents to learn AES

Cited by 3SourceScholar
2024

Adaptive Video Watermarking with Perceptual Guarantee and Efficiency Optimization

ICASSP 2024accepted

Existing video watermarking embeds robust watermarks in each frame of the video for copyright protection and tracking. However, just as any content written on a blank paper is easily perceived, embedding watermarks in the texture-poor frames impairs imperceptibility. Common geometric attacks such as…

Cited by 0SourceScholar
2024

Audio-Visual Segmentation via Unlabeled Frame Exploitation

CVPR 2024poster

Audio-visual segmentation (AVS) aims to segment the sounding objects in video frames. Although great progress has been witnessed we experimentally reveal that current methods reach marginal performance gain within the use of the unlabeled frames leading to the underutilization issue. To fully explor…

Cited by 10SourcePDFScholar
2024

Exploring Consistent Spatio-Temporal Distortion and Stable 3-D DCT Coefficients for Robust Blind Video Watermarking

ICASSP 2024accepted

With the rapid development of mobile Internet and video applications, robust video watermarking technology has become a focal area of research for protecting and tracking intellectual property rights in digital media. An important characteristic of video is that it has both spatial and temporal prop…

Cited by 0SourceScholar
2024

Probabilistic Conformal Distillation for Enhancing Missing Modality Robustness

NeurIPS 2024poster

Multimodal models trained on modality-complete data are plagued with severe performance degradation when encountering modality-missing data. Prevalent cross-modal knowledge distillation-based methods precisely align the representation of modality-missing data and that of its modality-complete counte…

2023

AttrSeg: Open-Vocabulary Semantic Segmentation via Attribute Decomposition-Aggregation

NeurIPS 2023poster

Open-vocabulary semantic segmentation is a challenging task that requires segmenting novel object categories at inference time. Recent works explore vision-language pre-training to handle this task, but suffer from unrealistic assumptions in practical scenarios, i.e., low-quality textual category n…

2023

Dimensional Optimization and Anti-Disturbance Analysis of an Upgraded Feed Mechanism in FAST

ICRA 2023poster

Five-hundred-meter aperture spherical radio telescope (FAST) is a very famous large-scale scientific facility with excellent performance for astronomical observation in the world, but it currently fails to observe the center of the Milky Way Galaxy due to the limited observation angle that is affect…

Cited by 2SourceScholar
2023

Monte Carlo Linear Clustering with Single-Point Supervision is Enough for Infrared Small Target Detection

ICCV 2023poster

Single-frame infrared small target (SIRST) detection aims at separating small targets from clutter backgrounds on infrared images. Recently, deep learning based methods have achieved promising performance on SIRST detection, but at the cost of a large amount of training data with expensive pixel-lev…

Cited by 56PDFcodeScholar
2023

Uncovering Prototypical Knowledge for Weakly Open-Vocabulary Semantic Segmentation

NeurIPS 2023poster

This paper studies the problem of weakly open-vocabulary semantic segmentation (WOVSS), which learns to segment objects of arbitrary classes using mere image-text pairs. Existing works turn to enhance the vanilla vision transformer by introducing explicit grouping recognition, i.e., employing severa…

Cited by 29SourcePDFScholar
2022

Design and Stiffness Analysis of a Novel 7-DOF Cable-Driven Manipulator

RA-L 2022

The lightweight design of robots is an important factor in the field of human-robot interaction. Thus, a lightweight 7-DOF cable-driven manipulator is proposed and manufactured, including a shoulder joint with 3-DOF, an elbow joint with 1-DOF, and a wrist joint with 3-DOF. The offset design of the s

Cited by 26SourceScholar
2022

Exploiting Class Activation Value for Partial-Label Learning

ICLR 2022poster

Partial-label learning (PLL) solves the multi-class classification problem, where each training instance is assigned a set of candidate labels that include the true label. Recent advances showed that PLL can be compatible with deep neural networks, which achieved state-of-the-art performance. Howeve…

Cited by 59SourcePDFScholar
2021

Complementary Patch for Weakly Supervised Semantic Segmentation

ICCV 2021poster

Weakly Supervised Semantic Segmentation (WSSS) based on image-level labels has been greatly advanced by exploiting the outputs of Class Activation Map (CAM) to generate the pseudo labels for semantic segmentation. However, CAM merely discovers seeds from a small number of regions, which may be insuf…

Cited by 172PDFcodeScholar