← Search

Yibo Hu

17 accepted papers

2025

DSFormer: Deformable Pointformer for 3D Salient Object Detection

ICASSP 2025accepted

Due to the irregularity of 3D point clouds, it is extremely challenging to detect the most salient objects from them and segment contours. DSFormer is proposed for 3D salient object detection, addressing challenges like small objects, multiple objects and complex backgrounds. It employs an encoder-d…

Cited by 0SourceScholar
2025

Ensuring Force Safety in Vision-Guided Robotic Manipulation via Implicit Tactile Calibration

CoRL 2025poster

In unstructured environments, robotic manipulation tasks involving objects with constrained motion trajectories—such as door opening—often experience discrepancies between the robot's vision-guided end-effector trajectory and the object's constrained motion path. Such discrepancies generate uninten…

Cited by 0SourceScholar
2025

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

ICML 2025poster

Existing efforts to align multimodal large language models (MLLMs) with human preferences have only achieved progress in narrow areas, such as hallucination reduction, but remain limited in practical applicability and generalizability. To this end, we introduce **MM-RLHF**, a dataset containing **12…

Cited by 13SourcePDFScholar
2025

ReCoT: Reflective Self-Correction Training for Mitigating Confirmation Bias in Large Vision-Language Models

ICCV 2025poster

Recent advancements in Large Vision-Language Models (LVLMs) have greatly improved their ability to understand both visual and text information. However, a common problem in LVLMs is confirmation bias, where models tend to repeat previous assumptions and follow earlier viewpoints instead of reflectin…

Cited by 0SourcePDFScholar
2024

Leveraging Codebook Knowledge with NLI and ChatGPT for Zero-Shot Political Relation Classification

ACL 2024long

Is it possible accurately classify political relations within evolving event ontologies without extensive annotations? This study investigates zero-shot learning methods that use expert knowledge from existing annotation codebook, and evaluates the performance of advanced ChatGPT (GPT-3.5/4) and a n…

2023

ChatEdit: Towards Multi-turn Interactive Facial Image Editing via Dialogue

EMNLP 2023long main

This paper explores interactive facial image editing through dialogue and presents the ChatEdit benchmark dataset for evaluating image editing and conversation abilities in this context. ChatEdit is constructed from the CelebA-HQ dataset, incorporating annotated multi-turn dialogues corresponding to…

Cited by 0SourceScholar
2022

ConfliBERT: A Pre-trained Language Model for Political Conflict and Violence

NAACL 2022long

Analyzing conflicts and political violence around the world is a persistent challenge in the political science and policy communities due in large part to the vast volumes of specialized text needed to monitor conflict and violence on a global scale. To help advance research in political science, we…

2022

Controllable Fake Document Infilling for Cyber Deception

EMNLP 2022finding

Recent works in cyber deception study how to deter malicious intrusion by generating multiple fake versions of a critical document to impose costs on adversaries who need to identify the correct information. However, existing approaches are context-agnostic, resulting in sub-optimal and unvaried out…

2021

CM-NAS: Cross-Modality Neural Architecture Search for Visible-Infrared Person Re-Identification

ICCV 2021poster

Visible-Infrared person re-identification (VI-ReID) aims to match cross-modality pedestrian images, breaking through the limitation of single-modality person ReID in dark environment. In order to mitigate the impact of large modality discrepancy, existing works manually design various two-stream arc…

Cited by 159PDFcodeScholar
2021

Dive Into Ambiguity: Latent Distribution Mining and Pairwise Uncertainty Estimation for Facial Expression Recognition

CVPR 2021poster

Due to the subjective annotation and the inherent inter-class similarity of facial expressions, one of key challenges in Facial Expression Recognition (FER) is the annotation ambiguity. In this paper, we proposes a solution, named DMUE, to address the problem of annotation ambiguity from two perspec…

Cited by 291PDFcodeScholar
2021

Multidimensional Uncertainty-Aware Evidential Neural Networks

AAAI 2021technical

Traditional deep neural networks (NNs) have significantly contributed to the state-of-the-art performance in the task of classification under various application domains. However, NNs have not considered inherent uncertainty in data associated with the class probabilities where misclassification un…

2020

Hierarchical Face Aging through Disentangled Latent Characteristics

ECCV 2020poster

Current age datasets lie in a long-tailed distribution, which brings difficulties to describe the aging mechanism for the imbalance ages. To alleviate it, we design a novel facial age prior to guide the aging mechanism modeling. To explore the age effects on facial images, we propose a Disentangled…

Cited by 25SourcePDFScholar
2020

TF-NAS: Rethinking Three Search Freedoms of Latency-Constrained Differentiable Neural Architecture Search

ECCV 2020poster

With the flourish of differentiable neural architecture search (NAS), automatically searching latency-constrained architectures gives a new perspective to reduce human labor and expertise. However, the searched architectures are usually suboptimal in accuracy and may have large jitters around the ta…

2019

Dual Variational Generation for Low Shot Heterogeneous Face Recognition

NeurIPS 2019spotlight

Heterogeneous Face Recognition (HFR) is a challenging issue because of the large domain discrepancy and a lack of heterogeneous data. This paper considers HFR as a dual generation problem, and proposes a novel Dual Variational Generation (DVG) framework. It generates large-scale new paired heterogen…

2019

M2FPA: A Multi-Yaw Multi-Pitch High-Quality Dataset and Benchmark for Facial Pose Analysis

ICCV 2019poster

Facial images in surveillance or mobile scenarios often have large view-point variations in terms of pitch and yaw angles. These jointly occurred angle variations make face recognition challenging. Current public face databases mainly consider the case of yaw variations. In this paper, a new large-s…

Cited by 45PDFcodeScholar
2018

Learning a High Fidelity Pose Invariant Model for High-resolution Face Frontalization

NeurIPS 2018poster

Face frontalization refers to the process of synthesizing the frontal view of a face from a given profile. Due to self-occlusion and appearance distortion in the wild, it is extremely challenging to recover faithful results and preserve texture details in a high-resolution. This paper proposes a Hi…

Cited by 113SourcePDFScholar