← Search

Huadong Ma

34 accepted papers

2026

HDRMovieformer: A Transformer Framework and Benchmark for Cinematic SDR-to-HDR Conversion

AAAI 2026technical

With the growing prevalence of HDR-capable cinema venues such as Cinity LED theaters, there is an increasing demand to convert existing Standard Dynamic Range (SDR) films into High Dynamic Range (HDR) formats for theatrical presentation. However, existing SDR-to-HDR conversion methods are primarily

Cited by 0SourcePDFScholar
2026

Improving Batch Normalization with Test-Time Adaptation for Robust Object Detection in Self-Driving

AAAI 2026technical

In open real-world autonomous driving scenarios, challenges such as sensor failure and extreme weather hinder the generalization of current autonomous driving perception models to these unseen domain, due to the domain shifts between the test and training data. As the parameter scale of autonomous d

Cited by 0SourcePDFScholar
2026

Regression over Classification: Assessing Image Aesthetics via Multimodal Large Language Models

AAAI 2026technical

Image Aesthetics Assessment (IAA) evaluates visual quality through user-centered perceptual analysis and can guide various applications. Recent advances in Multimodal Large Language Models (MLLMs) have sparked interest in adapting them for IAA. However, two critical limitations persist in applying M

Cited by 0SourcePDFScholar
2026

Selective Diffusion Distillation for Real-World High-Scale Image Super-Resolution

AAAI 2026technical

High-scale image super-resolution (SR) has become increasingly important with the rapid growth of mobile devices and high-resolution displays. However, current SR methods primarily focus on lower scales and generalize poorly to high-scale scenarios due to severe information loss and complex real-wor

Cited by 0SourcePDFScholar
2025

From Abyssal Darkness to Blinding Glare: A Benchmark on Extreme Exposure Correction in Real World

ICCV 2025poster

Exposure correction aims to restore over/under-exposed images to well-exposed ones using a single network. However, existing methods mainly handle non-extreme exposure conditions and struggle with the severe luminance and texture loss caused by extreme exposure. Through a thorough investigation, we…

2025

Rethinking Personalized Aesthetics Assessment: Employing Physique Aesthetics Assessment as An Exemplification

CVPR 2025highlight

The Personalized Aesthetics Assessment (PAA) aims to accurately predict an individual's unique perception of aesthetics. With the surging demand for customization, PAA enables applications to generate personalized outcomes by aligning with individual aesthetic preferences. The prevailing PAA paradig…

2025

Synergistic Tensor and Pipeline Parallelism

NeurIPS 2025poster

In the machine learning system, the hybrid model parallelism combining tensor parallelism (TP) and pipeline parallelism (PP) has become the dominant solution for distributed training of Large Language Models~(LLMs) and Multimodal LLMs (MLLMs). However, TP introduces significant collective communicat…

Cited by 0SourcecodeScholar
2025

T2SG: Traffic Topology Scene Graph for Topology Reasoning in Autonomous Driving

CVPR 2025poster

Understanding the traffic scenes and then generating high-definition (HD) maps present significant challenges in autonomous driving. In this paper, we defined a novel \underline T raffic \underline T opology \underline S cene \underline G raph (\text T ^2\text SG ), a unified scene graph explicitl…

2025

Towards Efficient Object Re-Identification with a Novel Cloud-Edge Collaborative Framework

AAAI 2025technical

Object re-identification (ReID) is committed to searching for objects of the same identity across cameras, and its real-world deployment is gradually increasing. Current ReID methods assume that the deployed system follows the centralized processing paradigm, i.e., all computations are conducted in…

Cited by 0SourcePDFScholar
2025

VIoTGPT: Learning to Schedule Vision Tools Towards Intelligent Video Internet of Things

AAAI 2025technical

Video Internet of Things (VIoT) has shown full potential in collecting an unprecedented volume of video data. How to schedule the domain-specific perceiving models and analyze the collected videos uniformly, efficiently, and especially intelligently to accomplish complicated tasks is challenging. To…

2025

mmHIU: a human-to-human interaction understanding system based on mmWave sensing

ICASSP 2025accepted

Human-to-human interaction understanding(HIU) plays a significant role in both physical and mental health of individuals in their daily lives. Nowadays, many of HIU tasks are based on visual information, which can compromise individuals’ privacy in daily life. In this paper, we propose mmHIU, a priv…

Cited by 0SourceScholar
2024

Continuous Optical Zooming: A Benchmark for Arbitrary-Scale Image Super-Resolution in Real World

CVPR 2024poster

Most current arbitrary-scale image super-resolution (SR) methods has commonly relied on simulated data generated by simple synthetic degradation models (e.g. bicubic downsampling) at continuous various scales thereby falling short in capturing the complex degradation of real-world images. This limit…

2024

Decomposed Vector-Quantized Variational Autoencoder for Human Grasp Generation

ECCV 2024poster

"Generating realistic human grasps is a crucial yet challenging task for applications involving object manipulation in computer graphics and robotics. Existing methods often struggle with generating fine-grained realistic human grasps that ensure all fingers effectively interact with objects, as the…

2024

ELTA: An Enhancer against Long-Tail for Aesthetics-oriented Models

ICML 2024poster

Real-world datasets often exhibit long-tailed distributions, compromising the generalization and fairness of learning-based models. This issue is particularly pronounced in Image Aesthetics Assessment (IAA) tasks, where such imbalance is difficult to mitigate due to a severe distribution mismatch be…

Cited by 4SourcePDFScholar
2024

Region-Aware Exposure Consistency Network for Mixed Exposure Correction

AAAI 2024technical

Exposure correction aims to enhance images suffering from improper exposure to achieve satisfactory visual effects. Despite recent progress, existing methods generally mitigate either overexposure or underexposure in input images, and they still struggle to handle images with mixed exposure, i.e., o…

2024

Rethinking No-reference Image Exposure Assessment from Holism to Pixel: Models, Datasets and Benchmarks

NeurIPS 2024poster

The past decade has witnessed an increasing demand for enhancing image quality through exposure, and as a crucial prerequisite in this endeavor, Image Exposure Assessment (IEA) is now being accorded serious attention. However, IEA encounters two persistent challenges that remain unresolved over the…

2024

SGFormer: Semantic Graph Transformer for Point Cloud-Based 3D Scene Graph Generation

AAAI 2024technical

In this paper, we propose a novel model called SGFormer, Semantic Graph TransFormer for point cloud-based 3D scene graph generation. The task aims to parse a point cloud-based scene into a semantic structural graph, with the core challenge of modeling the complex global structure. Existing methods b…

2024

Weakly-Supervised Temporal Action Localization by Inferring Salient Snippet-Feature

AAAI 2024technical

Weakly-supervised temporal action localization aims to locate action regions and identify action categories in untrimmed videos simultaneously by taking only video-level labels as the supervision. Pseudo label generation is a promising strategy to solve the challenging problem, but the current metho…

2023

Dancing in the Dark: A Benchmark towards General Low-light Video Enhancement

ICCV 2023poster

Low-light video enhancement is a challenging task with broad applications. However, current research in this area is limited by the lack of high-quality benchmark datasets. To address this issue, we design a camera system and collect a high-quality low-light video dataset with multiple exposures and…

Cited by 21PDFcodeScholar
2023

Disentangled Counterfactual Learning for Physical Audiovisual Commonsense Reasoning

NeurIPS 2023poster

In this paper, we propose a Disentangled Counterfactual Learning (DCL) approach for physical audiovisual commonsense reasoning. The task aims to infer objects’ physics commonsense based on both video and audio input, with the main challenge is how to imitate the reasoning ability of humans. Most of…

2023

Thinking Image Color Aesthetics Assessment: Models, Datasets and Benchmarks

ICCV 2023poster

We present a comprehensive study on a new task named image color aesthetics assessment (ICAA), which aims to assess color aesthetics based on human perception. ICAA is important for various applications such as imaging measurement and image analysis. However, due to the highly diverse aesthetic pref…

Cited by 22PDFcodeScholar
2023

Unsupervised Self-Driving Attention Prediction via Uncertainty Mining and Knowledge Embedding

ICCV 2023poster

Predicting attention regions of interest is an important yet challenging task for self-driving systems. Existing methodologies rely on large-scale labeled traffic datasets that are labor-intensive to obtain. Besides, the huge domain gap between natural scenes and traffic scenes in current datasets a…

Cited by 6PDFcodeScholar
2023

You Do Not Need Additional Priors or Regularizers in Retinex-Based Low-Light Image Enhancement

CVPR 2023poster

Images captured in low-light conditions often suffer from significant quality degradation. Recent works have built a large variety of deep Retinex-based networks to enhance low-light images. The Retinex-based methods require decomposing the image into reflectance and illumination components, which i…

Cited by 63SourcePDFScholar
2022

Defending Against Universal Attack Via Curvature-Aware Category Adversarial Training

ICASSP 2022accepted

Adversarial training can defend against universal adversarial perturbation (UAP) by injecting corresponding adversarial samples during training. However, adversarial samples used by existing methods, such as UAP, inevitably include excessive perturbations related to other categories due to its inher…

Cited by 0SourceScholar
2022

Global-Local Feature Enhancement Network for Robust Object Detection using mmWave Radar and Camera

ICASSP 2022accepted

Object detection with camera has achieved promising results using deep learning methods, but it suffers degraded performance under adverse conditions (e.g., foggy weather, poor illumination). To remedy this, some recent studies resort to leveraging the complementary mmWave radar, which is less affec…

Cited by 0SourceScholar
2021

m-Activity: Accurate and Real-Time Human Activity Recognition Via Millimeter Wave Radar

ICASSP 2021accepted

Natural human activity recognition (HAR) via millimeter wave (mmWave) sensing is a key to the human-computer interaction (HCI), e.g., activity assistance and living state monitoring. Prior work has shown the feasibility of HAR by utilizing mmWave radar, but it falls short of two real-world issues: p…

Cited by 0SourceScholar
2018

Joint License Plate Super-Resolution and Recognition in One Multi-Task Gan Framework

ICASSP 2018accepted

License plate recognition (LPR) plays an important role in intelligent transport systems. The existed LPR systems are mostly based on hand-crafted methods for detection, segmentation, and recognition, which cannot accurately recognize the license plate in unconstrained surveillance environments. In…

Cited by 0SourceScholar