← Search

Yuhang Zhang

23 accepted papers

2026

AirSim360: A Panoramic Simulation Platform within Drone View

CVPR 2026

The field of 360-degree omnidirectional understanding has been receiving increasing attention for advancing spatial intelligence. However, the lack of large-scale and diverse data remains a major limitation. In this work, we propose AirSim360, a simulation platform for omnidirectional data from aeri

Cited by 0SourcecodeScholar
2026

FedBRICK: Structural Bias Aware Heterogeneous Foundation Model Federated Tuning

AAAI 2026technical

Model-heterogeneous federated tuning (MHFT) enables the privacy-preserving fine-tuning of foundation models in heterogeneous systems by allowing clients and the server to adopt different model architectures. Depth partial training—where each client updates only a subset of the model

Cited by 0SourcePDFScholar
2026

Learning Underwater Image Enhancement Iteratively Without Reference Images

AAAI 2026technical

Since high-fidelity reference images are difficult to obtain in real underwater scenes, most deep models trained by synthetic paired data cannot match real-world data exactly. In this paper, we propose an unsupervised training framework for underwater image enhancement (UIE) by leveraging an iterati

Cited by 0SourcePDFScholar
2026

Weaving Graph over Tokens: Contextualizing Structured Sequences for LLMs

ICML 2026poster

Generative Graph Language Models (GLMs) must reconcile topology with causal language modeling. Linearization obscures multi-hop connectivity, while encoder-based methods bottleneck token-level reasoning during generation. Viewing context modeling as a form of message passing, we introduce **Weaver**…

Cited by 0SourceScholar
2025

Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants

NeurIPS 2025poster

Faces and humans are crucial elements in social interaction and are widely included in everyday photos and videos. Therefore, a deep understanding of faces and humans will enable multi-modal assistants to achieve improved response quality and broadened application scope. Currently, the multi-modal a…

Cited by 0SourcecodeScholar
2025

Lumina-T2X: Scalable Flow-based Large Diffusion Transformer for Flexible Resolution Generation

ICLR 2025spotlight

Sora unveils the potential of scaling Diffusion Transformer (DiT) for generating photorealistic images and videos at arbitrary resolutions, aspect ratios, and durations, yet it still lacks sufficient implementation details. In this paper, we introduce the Lumina-T2X family -- a series of Flow-based…

2025

RJE: A Retrieval-Judgment-Exploration Framework for Efficient Knowledge Graph Question Answering with LLMs

EMNLP 2025

Knowledge graph question answering (KGQA) aims to answer natural language questions using knowledge graphs.Recent research leverages large language models (LLMs) to enhance KGQA reasoning, but faces limitations: retrieval-based methods are constrained by the quality of retrieved information, while a

Cited by 0SourcePDFScholar
2025

SpaceDet: A Large-scale Space-based Image Dataset and RSO Detection for Space Situational Awareness

IJCAI 2025

Space situational awareness (SSA) plays an imperative role in maintaining safe space operations, especially given the increasingly congested space traffic around the Earth. Space-based SSA offers a flexible and lightweight solution compared to traditional ground-based SSA. With advanced machine lear

2024

Beyond Traditional Threats: A Persistent Backdoor Attack on Federated Learning

AAAI 2024technical

Backdoors on federated learning will be diluted by subsequent benign updates. This is reflected in the significant reduction of attack success rate as iterations increase, ultimately failing. We use a new metric to quantify the degree of this weakened backdoor effect, called attack persistence. Give…

2024

Bio-Inspired Pupal-Mode Actuator with Ultra-Crossing Capability for Soft Robots

ICRA 2024poster

Robot-assisted Natural Orifice Translu-minal Endoscopic Surgery (NOTES) represents a paradigm shift in surgical practice, significantly mini-mizing patient morbidity. However, the variability of inner diameter and the inter-luminal crossing within the luminal tracts lead to challenge for effective r…

Cited by 0SourceScholar
2024

FT-AED: Benchmark Dataset for Early Freeway Traffic Anomalous Event Detection

NeurIPS 2024poster

Early and accurate detection of anomalous events on the freeway, such as accidents, can improve emergency response and clearance. However, existing delays and mistakes from manual crash reporting records make it a difficult problem to solve. Current large-scale freeway traffic datasets are not desig…

2024

Faceptor: A Generalist Model for Face Perception

ECCV 2024oral

"With the comprehensive research conducted on various face analysis tasks, there is a growing interest among researchers to develop a unified approach to face perception. Existing methods mainly discuss unified representation and training, which lack task extensibility and application efficiency. To…

2024

Generalizable Facial Expression Recognition

ECCV 2024poster

"SOTA facial expression recognition (FER) methods fail on test sets that have domain gaps with the train set. Recent domain adaptation FER methods need to acquire labeled or unlabeled samples of target domains to fine-tune the FER model, which might be infeasible in real-world deployment. In this pa…

2024

Open-Set Facial Expression Recognition

AAAI 2024technical

Facial expression recognition (FER) models are typically trained on datasets with a fixed number of seven basic classes. However, recent research works (Cowen et al. 2021; Bryant et al. 2022; Kollias 2023) point out that there are far more expressions than the basic ones. Thus, when these models are…

Cited by 4SourcePDFScholar
2024

UMG-CLIP: A Unified Multi-Granularity Vision Generalist for Open-World Understanding

ECCV 2024poster

"Vision-language foundation models, represented by Contras-tive Language-Image Pre-training (CLIP), have gained increasing attention for jointly understanding both vision and textual tasks. However, existing approaches primarily focus on training models to match global image representations with tex…

2023

Enhancing Generalization of Universal Adversarial Perturbation through Gradient Aggregation

ICCV 2023poster

Deep neural networks are vulnerable to universal adversarial perturbation (UAP), an instance-agnostic perturbation capable of fooling the target model for most samples. Compared to instance-specific adversarial examples, UAP is more challenging as it needs to generalize across various samples and mo…

Cited by 30PDFcodeScholar
2023

Leave No Stone Unturned: Mine Extra Knowledge for Imbalanced Facial Expression Recognition

NeurIPS 2023poster

Facial expression data is characterized by a significant imbalance, with most collected data showing happy or neutral expressions and fewer instances of fear or disgust. This imbalance poses challenges to facial expression recognition (FER) models, hindering their ability to fully understand various…

Cited by 24SourcePDFScholar
2022

Learn from All: Erasing Attention Consistency for Noisy Label Facial Expression Recognition

ECCV 2022poster

"Noisy label Facial Expression Recognition (FER) is more challenging than traditional noisy label classification tasks due to the inter-class similarity and the annotation ambiguity. Recent works mainly tackle this problem by filtering out large-loss samples. In this paper, we explore dealing with n…

2022

One-Bit Active Query With Contrastive Pairs

CVPR 2022poster

How to achieve better results with fewer labeling costs remains a challenging task. In this paper, we present a new active learning framework, which for the first time incorporates contrastive learning into recently proposed one-bit supervision. Here one-bit supervision denotes a simple Yes or No qu…

Cited by 9PDFcodeScholar
2021

Relative Uncertainty Learning for Facial Expression Recognition

NeurIPS 2021poster

In facial expression recognition (FER), the uncertainties introduced by inherent noises like ambiguous facial expressions and inconsistent labels raise concerns about the credibility of recognition results. To quantify these uncertainties and achieve good performance under noisy data, we regard unce…

2017

Rotation invariance through structured sparsity for robust hyperspectral image classification

ICASSP 2017accepted

Sparse representation based classification has gained popularity with geospatial image analysis in general and hyperspectral image analysis in particular. A central idea with such classification approaches is that a test pixel (spectral reflectance vector) can be sparsely represented in a training d…

Cited by 0SourceScholar