← Search

Fei zhao

27 accepted papers

2026

CompBench: Benchmarking Complex Instruction-guided Image Editing

CVPR 2026

While real-world applications increasingly demand intricate scene manipulation, existing instruction-guided image editing benchmarks often oversimplify task complexity and lack comprehensive, fine-grained instructions. To bridge this gap, we introduce CompBench, a large-scale benchmark specifically

Cited by 0SourcecodeScholar
2026

Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training

ICML 2026poster

Determining an effective data mixture is a key factor in Large Language Model (LLM) pre-training, where models must balance general competence with proficiency on hard tasks such as math and code. However, identifying an optimal mixture remains an open challenge, as existing approaches either rely o…

Cited by 0SourceScholar
2026

GePBench: Evaluating Fundamental Geometric Perception for Multimodal Large Language Models

ICML 2026poster

Geometric shapes play important roles in both physical world and human cognition. While multimodal large language models (MLLMs) have made significant advancements in visual understanding, their abilities to recognize geometric shapes and their spatial relationships, which we term geometric percepti…

Cited by 0SourceScholar
2026

Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

ICLR 2026poster

DeepSeek-R1-Zero has successfully demonstrated the emergence of reasoning capabilities in LLMs purely through Reinforcement Learning (RL). Inspired by this breakthrough, we explore how RL can be utilized to enhance the reasoning capability of MLLMs. However, direct training with RL struggles to act…

Cited by 0SourcecodeScholar
2025

Decouple Distortion from Perception: Region Adaptive Diffusion for Extreme-low Bitrate Perception Image Compression

CVPR 2025poster

Leveraging the generative power of diffusion models, generative image compression has achieved impressive perceptual fidelity even at extremely low bitrates. However, current methods often neglect the non-uniform complexity of images, limiting their ability to balance global perceptual quality with…

Cited by 0SourcePDFScholar
2025

Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification

ICLR 2025poster

Multimodal Large Language Models (MLLMs) have achieved remarkable success in vision understanding, reasoning, and interaction. However, the inference computation and memory increase progressively with the generation of output tokens during decoding, directly affecting the efficacy of MLLMs. Existing…

2025

RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios

EMNLP 2025

While various multimodal multi-image evaluation datasets have been emerged, but these datasets are primarily based on English, and there has yet to be a Chinese multi-image dataset. To fill this gap, we introduce RealBench, the first Chinese multimodal multi-image dataset, which contains 9393 sample

2025

SNS-Bench: Defining, Building, and Assessing Capabilities of Large Language Models in Social Networking Services

ICML 2025poster

With the rapid advancement of Social Networking Services (SNS), the need for intelligent and efficient interaction within diverse platforms has become more crucial. Large Language Models (LLMs) play an important role in SNS as they possess the potential to revolutionize user experience, content gene…

Cited by 0SourcePDFScholar
2024

A Tri-Dynamic Preprocessing Framework for UGC Video Compression

ICASSP 2024accepted

In recent years, user generated content (UGC) has become the dominant force in internet traffic. However, UGC videos exhibit a higher degree of variability and diverse characteristics compared to traditional encoding test videos. This variance challenges the effectiveness of data-driven machine lear…

Cited by 0SourceScholar
2024

An Ergo-Interactive Framework for Human-Robot Collaboration Via Learning From Demonstration

RA-L 2024

This work presents an ergonomic and interactive human-robot collaboration (HRC) framework, through which new collaborative skills are extracted from a one-shot human demonstration and learned through Riemannian dynamic movement primitives (DMP). The proposed framework responds to human-robot interac

Cited by 10SourceScholar
2024

EFUF: Efficient Fine-Grained Unlearning Framework for Mitigating Hallucinations in Multimodal Large Language Models

EMNLP 2024main

Multimodal large language models (MLLMs) have attracted increasing attention in the past few years, but they may still generate descriptions that include objects not present in the corresponding images, a phenomenon known as object hallucination. To eliminate hallucinations, existing methods manuall…

2024

Multi-Contact Cartesian Null-Space Impedance Control for the Anthropomorphic Manipulator Without Knowledge of Force Locations

RA-L 2024

There is still a lack of null-space impedance control defined in Cartesian space that is suitable for multipoint contact and does not require knowledge of the force locations. To address this problem, this letter demonstrates a type of Cartesian null-space impedance control for the anthropomorphic m

Cited by 2SourceScholar
2024

MultiSQL: A Schema-Integrated Context-Dependent Text2SQL Dataset with Diverse SQL Operations

ACL 2024findings

Text2SQL is a task that translates natural language into SQL statements. Context-dependent Text2SQL offers a more natural database interaction by simulating dialogues between users and databases, with CoSQL and SparC as representative datasets. Yet, these datasets struggle to accurately replicate re…

2024

Simple-Rotation Angle/Axis Representations Based Second-Order Impedance Control

RA-L 2024

Since the difference in angular velocity is used as the derivative of the orientation error in the classical impedance control, there is no longer a form of the second-order differential equation (SODE), and there is non-linearity in the classical impedance control, which limits applications. To add

Cited by 0SourceScholar
2023

Collision-Free Motion Generation Based on Stochastic Optimization and Composite Signed Distance Field Networks of Articulated Robot

RA-L 2023

Safe robot motion generation is critical for practical applications from manufacturing to homes. In this work, we proposed a stochastic optimization-based motion generation method to generate collision-free and time-optimal motion for the articulated robot represented by composite signed distance fi

Cited by 30SourceScholar
2023

M2DF: Multi-grained Multi-curriculum Denoising Framework for Multimodal Aspect-based Sentiment Analysis

EMNLP 2023long main

Multimodal Aspect-based Sentiment Analysis (MABSA) is a fine-grained Sentiment Analysis task, which has attracted growing research interests recently. Existing work mainly utilizes image information to improve the performance of MABSA task. However, most of the studies overestimate the importance of…

Cited by 0SourcecodeScholar
2023

Measuring Your ASTE Models in The Wild: A Diversified Multi-domain Dataset For Aspect Sentiment Triplet Extraction

ACL 2023findings

Aspect Sentiment Triplet Extraction (ASTE) is widely used in various applications. However, existing ASTE datasets are limited in their ability to represent real-world scenarios, hindering the advancement of research in this area. In this paper, we introduce a new dataset, named DMASTE, which is man…

2022

Label-Driven Denoising Framework for Multi-Label Few-Shot Aspect Category Detection

EMNLP 2022finding

Multi-Label Few-Shot Aspect Category Detection (FS-ACD) is a new sub-task of aspect-based sentiment analysis, which aims to detect aspect categories accurately with limited training instances. Recently, dominant works use the prototypical network to accomplish this task, and employ the attention mec…

2022

Learning from Adjective-Noun Pairs: A Knowledge-enhanced Framework for Target-Oriented Multimodal Sentiment Classification

COLING 2022main

Target-oriented multimodal sentiment classification (TMSC) is a new subtask of aspect-based sentiment analysis, which aims to determine the sentiment polarity of the opinion target mentioned in a (sentence, image) pair. Recently, dominant works employ the attention mechanism to capture the correspon…

2021

A Framework for Autonomous Impedance Regulation of Robots Based on Imitation Learning and Optimal Control

RA-L 2021

In this work, we propose a framework to address the autonomous impedance regulation problem of robots in a class of constrained manipulation tasks. In this framework, a human arm endpoint stiffness model is used to extract the task stiffness geometry along the constrained trajectory, which is then e

Cited by 39SourceScholar
2021

Unified Approach for Hybrid Motion Control of MOCA Based on Weighted Whole-Body Cartesian Impedance Formulation

RA-L 2021

This work presents a unified approach for hybrid motion control of the Mobile Collaborative Robotic Assistant (MOCA). The objective is to develop a loco-manipulation controller, enabling various couplings of the arm and the mobile base movements, and particularly their purely decoupled motions. The

Cited by 26SourceScholar
2021

Unit Selection Synthesis Based Data Augmentation for Fixed Phrase Speaker Verification

ICASSP 2021accepted

Data augmentation is commonly used to help build a robust speaker verification system, especially in limited-resource case. However, conventional data augmentation methods usually focus on the diversity of acoustic environment, leaving the lexicon variation neglected. For text dependent speaker veri…

Cited by 0SourceScholar
2019

A Robust Text-independent Speaker Verification Method Based on Speech Separation and Deep Speaker

ICASSP 2019accepted

Recently, deep neural networks (DNNs) have achieved incredible performance in speaker verification. However, most of which remains sensitive to environment noise. In this paper, we propose an end-to-end speaker verification framework to enhance the robustness against background noise. The proposed f…

Cited by 0SourceScholar
2019

A Teleoperation Interface for Loco-Manipulation Control of Mobile Collaborative Robotic Assistant

RA-L 2019

This letter presents a novel teleoperation interface that enables remote loco-manipulation control of a MObile Collaborative robotic Assistant (MOCA). MOCA is a new research platform developed at the Istituto Italiano di Tecnologia (IIT), which is composed of a lightweight manipulator arm, a Pisa/II

Cited by 91SourceScholar