← Search

Zhaofeng He

27 accepted papers

2026

BEYOND FACE SWAPPING: A DIFFUSION-BASED DIGITAL HUMAN BENCHMARK FOR MULTIMODAL DEEPFAKE DETECTION

ICASSP 2026poster

In recent years, the explosive advancement of deepfake technology has posed a critical and escalating threat to public security: diffusion-based digital human generation. Unlike traditional face manipulation methods, such models can generate highly realistic videos with consistency via multimodal co…

Cited by 0SourcePDFScholar
2026

Detoxifying Large Language Models via Localized Feature Editing with Sparse Autoencoders

IJCAI 2026

Large Language Models (LLMs) powerful generative capabilities also pose significant risks, underscoring the need for effective detoxification methods to ensure safer deployment. Due to the polysemantic nature of LLM neurons, recent neuron intervention methods inevitably entangle unrelated concepts,

Cited by 0Scholar
2026

From Chaos to Cure: A Prefix Heuristics Guided Model-Agnostic Adaptive Detoxification Framework

AAAI 2026technical

The impressive performance of large language models (LLMs) also brings inherent toxicity risks, prompting the need for effective detoxification to support responsible deployment. Prevailing methods generally follow an inflexible model-specific fashion, addressing only individual models or model fami

Cited by 0SourcePDFScholar
2026

Harnessing Chain-of-Thought Reasoning in Multimodal Large Language Models for Face Anti-Spoofing

CVPR 2026

Face Anti-Spoofing (FAS) typically depends on a single visual modality when defending against presentation attacks such as print attacks, screen replays, and 3D masks, resulting in limited generalization across devices, environments, and attack types. Meanwhile, Multimodal Large Language Models (MLL

Cited by 0SourcecodeScholar
2026

UniDoorManip: Learning Universal Door Manipulation Policy Over Large-Scale and Diverse Door Manipulation Environments

ICRA 2026poster

Learning a universal manipulation policy encompassing doors with diverse categories, geometries and mechanisms, is crucial for future embodied agents to effectively work in complex and broad real-world scenarios. Due to the limited datasets and unrealistic simulation environments, previous studies f…

2025

A*-Thought: Efficient Reasoning via Bidirectional Compression for Low-Resource Settings

NeurIPS 2025poster

Large Reasoning Models (LRMs) achieve superior performance by extending the thought length. However, a lengthy thinking trajectory leads to reduced efficiency. Most of the existing methods are stuck in the assumption of overthinking and attempt to reason efficiently by compressing the Chain-of-Thoug…

Cited by 0SourcecodeScholar
2025

AdaManip: Adaptive Articulated Object Manipulation Environments and Policy Learning

ICLR 2025poster

Articulated object manipulation is a critical capability for robots to perform various tasks in real-world scenarios. Composed of multiple parts connected by joints, articulated objects are endowed with diverse functional mechanisms through complex relative motions. For example, a safe consists of a…

Cited by 4SourcePDFScholar
2025

Adaptive Articulated Object Manipulation On The Fly with Foundation Model Reasoning and Part Grounding

ICCV 2025poster

Articulated objects pose diverse manipulation challenges for robots. Since their internal structures are not directly observable, robots must adaptively explore and refine actions to generate successful manipulation trajectories. While existing works have attempted cross-category generalization in a…

Cited by 0SourcePDFScholar
2025

Eye Movements as Images: A Multimodal Framework for Eye Movements Representation

ICASSP 2025accepted

Eye movements are increasingly popular for enhancing natural language processing and modeling individual states. Although specialized methods have been developed to represent eye movements for various tasks, effectively modeling the complex dynamics of eye movements and the heterogeneity with stimul…

Cited by 0SourceScholar
2025

GaussianEnhancer: A General Rendering Enhancer for Gaussian Splatting

ICASSP 2025accepted

Gaussian Splatting (GS) methods, including 3DGS and 2DGS, have demonstrated exceptional performance in real-time novel view synthesis (NVS), emerging as a transformative technology in the fields of explicit rendering and computer graphics. However, GS-based methods still face challenges in rendering…

Cited by 0SourceScholar
2025

Improving Food Recognition with Retrieval-Augmented and Domain-Adaptive LVLMs

ICASSP 2025accepted

Food recognition is pivotal in enhancing intelligent food recommendation systems and nutritional management, contributing to balanced diets and overall health. Although Large Vision-Language Models (LVLMs) have demonstrated impressive performances across various domains, their performance on the foo…

Cited by 0SourceScholar
2025

MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs

ICLR 2025poster

Current multimodal misinformation detection (MMD) methods often assume a single source and type of forgery for each sample, which is insufficient for real-world scenarios where multiple forgery sources coexist. The lack of a benchmark for mixed-source misinformation has hindered progress in this fie…

Cited by 0SourcePDFScholar
2025

SecDecoding: Steerable Decoding for Safer LLM Generation

EMNLP 2025

Large language models (LLMs) have achieved remarkable performance across diverse tasks, yet ensuring output safety remains a fundamental challenge. Existing defense methods often suffer from limited generalization, high computational overhead, or significant utility degradation. In this work, we pre

2025

Select-Then-Decompose: From Empirical Analysis to Adaptive Selection Strategy for Task Decomposition in Large Language Models

EMNLP 2025

Large language models (LLMs) have demonstrated remarkable reasoning and planning capabilities, driving extensive research into task decomposition. Existing task decomposition methods focus primarily on memory, tool usage, and feedback mechanisms, achieving notable success in specific domains, but th

2024

Enhancing Short-and Long-Term Sea Surface Temperature Forecasting with a Static and Dynamic Learnable Personalized Graph Convolution Network

ICASSP 2024accepted

Sea surface temperature (SST) plays an important role in our Earth’s atmosphere, wielding significant influence over both local and global climates and profoundly impacting ecosystems. However, this task presents unique challenges due to the inherent complexity and uncertainty within ocean systems.…

Cited by 0SourceScholar
2024

HyperMoE: Towards Better Mixture of Experts via Transferring Among Experts

ACL 2024long

The Mixture of Experts (MoE) for language models has been proven effective in augmenting the capacity of models by dynamically routing each input token to a specific subset of experts for processing. Despite the success, most existing methods face a challenge for balance between sparsity and the ava…

2024

InstaStyle: Inversion Noise of a Stylized Image is Secretly a Style Adviser

ECCV 2024poster

"Stylized text-to-image generation focuses on creating images from textual descriptions while adhering to a style specified by reference images. However, subtle style variations within different reference images can hinder the model from accurately learning the target style. In this paper, we propos…

2024

LLMArena: Assessing Capabilities of Large Language Models in Dynamic Multi-Agent Environments

ACL 2024long

Recent advancements in large language models (LLMs) have revealed their potential for achieving autonomous agents possessing human-level intelligence. However, existing benchmarks for evaluating LLM Agents either use static datasets, potentially leading to data leakage or focus only on single-agent…

2024

Localize, Understand, Collaborate: Semantic-Aware Dragging via Intention Reasoner

NeurIPS 2024poster

Flexible and accurate drag-based editing is a challenging task that has recently garnered significant attention. Current methods typically model this problem as automatically learning "how to drag" through point dragging and often produce one deterministic estimation, which presents two key limitati…

2024

MORE-3S:Multimodal-based Offline Reinforcement Learning with Shared Semantic Spaces

COLING 2024main

Drawing upon the intuition that aligning different modalities to the same semantic embedding space would allow models to understand states and actions more easily, we propose a new perspective to the offline reinforcement learning (RL) challenge. More concretely, we transform it into a supervised le…

2024

MQE: Unleashing the Power of Interaction with Multi-agent Quadruped Environment

IROS 2024poster

The advent of deep reinforcement learning (DRL) has significantly advanced the field of robotics, particularly in the control and coordination of quadruped robots. However, the complexity of real-world tasks often necessitates the deployment of multi-robot systems capable of sophisticated interactio…

Cited by 4SourcecodeScholar
2023

An Adaptive Prompt Generation Framework for Task-oriented Dialogue System

EMNLP 2023long findings

The de facto way of utilizing black-box large language models (LLMs) to perform various downstream tasks is prompting. However, obtaining suitable prompts for specific tasks is still a challenging problem. While existing LLM-based methods demonstrate promising performance in task-oriented dialogue (…

Cited by 0SourceScholar
2023

ChatEdit: Towards Multi-turn Interactive Facial Image Editing via Dialogue

EMNLP 2023long main

This paper explores interactive facial image editing through dialogue and presents the ChatEdit benchmark dataset for evaluating image editing and conversation abilities in this context. ChatEdit is constructed from the CelebA-HQ dataset, incorporating annotated multi-turn dialogues corresponding to…

Cited by 0SourceScholar
2023

Learning-to-Rank Meets Language: Boosting Language-Driven Ordering Alignment for Ordinal Classification

NeurIPS 2023poster

We present a novel language-driven ordering alignment method for ordinal classification. The labels in ordinal classification contain additional ordering relations, making them prone to overfitting when relying solely on training data. Recent developments in pre-trained vision-language models inspir…