← Search

Mingliang Zhou

16 accepted papers

2026

DPGF-Net: Dual-Prior Guided Fusion Network for Joint Assessment of Perceptual Quality and Semantic Consistency in AI-Generated Images

CVPR 2026

The development of AI-generated technology requires effective image quality assessment (AGIQA) methods to jointly evaluate visual quality and text-content alignment, ensuring that the generated content is both visually appealing and faithful to the user's instructions. Nevertheless, visual degradati

Cited by 0SourcecodeScholar
2026

HAIC: Humanoid Agile Object Interaction Control via Dynamics-Aware World Model

RSS 2026poster

Humanoid robots exhibit significant potential for executing complex whole-body interaction tasks in unstructured environments. While recent advancements in Human-Object Interaction (HOI) have been substantial, prevailing methodologies predominantly address the manipulation of fully actuated objects,…

Cited by 0SourceScholar
2026

Language-Guided One-Step Diffusion Model for Nighttime Flare Removal

CVPR 2026

Nighttime photography is susceptible to flare caused by strong light sources, which degrades visual quality and disrupts structural information required by downstream vision tasks. Existing nighttime flare removal methods generally lack semantic priors for flare-occluded regions and thus tend to int

Cited by 0SourceScholar
2025

A Dynamic Learning Strategy for Dempster-Shafer Theory with Applications in Classification and Enhancement

NeurIPS 2025poster

Effective modelling of uncertain information is crucial for quantifying uncertainty. Dempster–Shafer evidence (DSE) theory is a widely recognized approach for handling uncertain information. However, current methods often neglect the inherent a priori information within data during modelling, and im…

Cited by 0SourceScholar
2025

Image Quality Assessment: Investigating Causal Perceptual Effects with Abductive Counterfactual Inference

CVPR 2025poster

Existing full-reference image quality assessment (FR-IQA) methods often fail to capture the complex causal mechanisms that underlie human perceptual responses to image distortions, limiting their ability to generalize across diverse scenarios. In this paper, we propose an FR-IQA method based on abdu…

Cited by 0SourcePDFScholar
2024

Optimization Based Dynamic Skateboarding of Quadrupedal Robot

ICRA 2024poster

Robot skateboarding is a novel and challenging task for legged robots. Accurately modeling the dynamics of dual floating bases and developing effective planning and control methods present significant complexities in accomplishing skateboarding behavior. This paper focuses on enabling the quadrupeda…

Cited by 1SourceScholar
2024

State Estimation Transformers for Agile Legged Locomotion

IROS 2024poster

We propose a state estimation method that can accurately predict the robot’s privileged states to push the limits of quadruped robots in executing advanced skills such as jumping in the wild. In particular, we present the State Estimation Transformers (SET), an architecture that casts the state esti…

Cited by 1SourceScholar
2023

Joint Robust Representation And Generalization Enhancement For Cross-Modality Person Re-Identification

ICASSP 2023accepted

Cross-modality person re-identification (cm-ReID) aims to match pedestrian images from visible and infrared cameras. Most existing methods ignore data bias due to different cameras and views and overlook the strong dependence between feature maps that hinders modal alignment. In this paper, we propo…

Cited by 0SourceScholar
2023

Run and Catch: Dynamic Object-Catching of Quadrupedal Robots

IROS 2023poster

Quadrupedal robots are performing increasingly more real-world capabilities, but are primarily limited to locomotion tasks. To expand their task-level abilities of object acquisition, i.e., run-to-catch as frisbee catching for dogs, this paper developed a control pipeline using stereo vision for leg…

Cited by 2SourceScholar
2021

Deep Adversarial Quantization Network for Cross-Modal Retrieval

ICASSP 2021accepted

In this paper, we propose a seamless multimodal binary learning method for cross-modal retrieval. First, we utilize adversarial learning to learn modality-independent representations of different modalities. Second, we formulate loss function through the Bayesian approach, which aims to jointly maxi…

Cited by 0SourceScholar
2021

Distribution-Aware Hierarchical Weighting Method for Deep Metric Learning

ICASSP 2021accepted

In this paper, we propose distribution-aware hierarchical weighting (DHW) method for deep metric learning. First, we formulate the distributions of different classes according to the form of gaussian curves, and update distributions as the training process. Second, depending on the learnable distrib…

Cited by 0SourceScholar
2021

Joint Reinforcement Learning and Game Theory Bitrate Control Method for 360-Degree Dynamic Adaptive Streaming

ICASSP 2021accepted

A joint reinforcement learning (RL) and game theory method is presented for segment-level continuous bitrate selection and tile-level bitrate allocation in tile-based 360-degree streaming to increase users’ quality of experience (QoE). First, a viewpoint prediction method based on single-user (SU) v…

Cited by 0SourceScholar
2020

Hybrid Active Contour Driven by Double-Weighted Signed Pressure Force for Image Segmentation

ICASSP 2020accepted

In this paper, we proposed a novel hybrid active contour driven by double-weighted signed pressure force method for image segmentation. First, the Legendre polynomials and global information are integrated into the signed pressure force (SPF) function and a coefficient is applied to weight the effec…

Cited by 0SourceScholar
2020

Intra Frame Rate Control for Versatile Video Coding with Quadratic Rate-Distortion Modelling

ICASSP 2020accepted

With numerous coding tools adopted in the forthcoming Versatile Video Coding (VVC) standard, much less work has been dedicated to study the corresponding Rate-Distortion (R-D) characteristics. This paper proposes a new quadratic R-D model for Versatile Video Coding. In particular, based on the propo…

Cited by 0SourceScholar
2020

Regularized Attentive Capsule Network for Overlapped Relation Extraction

COLING 2020main

Distantly supervised relation extraction has been widely applied in knowledge base construction due to its less requirement of human efforts. However, the automatically established training datasets in distant supervision contain low-quality instances with noisy words and overlapped relations, intro…

Cited by 9SourcePDFScholar