← Search

Qingmin Liao

46 accepted papers

2026

AgentExpt: Automating AI Experiment Design with LLM-based Resource Retrieval Agent

ICML 2026poster

In modern AI research, baseline and dataset selection is a high-stakes decision in experimental design. It operationalizes a research idea into a concrete evaluation protocol and largely determines the validity and comparability of empirical conclusions. However, making appropriate choices is increa…

Cited by 0SourceScholar
2026

AgentSwift: Efficient LLM Agent Design via Value-Guided Hierarchical Search

AAAI 2026technical

Large language model (LLM) agents have demonstrated strong capabilities across diverse domains, yet automated agent design remains a significant challenge. Current automated agent design approaches are often constrained by limited search spaces that primarily optimize workflows but fail to integrate

Cited by 0SourcePDFScholar
2026

Beyond Skeletons: Learning Animation Directly from Driving Videos with Same2X Training Strategy

ICLR 2026poster

Human image animation aims to generate a video from a static reference image, guided by pose information extracted from a driving video. Existing approaches often rely on pose estimators to extract intermediate representations, but such signals are prone to errors under occlusion or complex poses. B…

Cited by 0SourcecodeScholar
2026

EgoTactile: Learning Grasp Pressure for Everyday Objects from Egocentric Video

ICML 2026spotlight

Estimating full-hand grasp pressure from egocentric video is critical for immersive VR and robotic manipulation, yet dense tactile sensing often relies on intrusive hardware. Existing vision-based methods predominantly rely on planar surfaces or fingertip contacts, failing to generalize to complex 3…

Cited by 0SourcecodeScholar
2026

Generative Adaptation of Dynamics to Environmental Shifts via Weight-space Diffusion

ICML 2026poster

Data-driven dynamics prediction often fails under environmental shifts, while traditional fine-tuning remains computationally prohibitive for hardware-constrained or data-scarce applications. We propose DynaDiff, a generative meta-learning framework that transitions the paradigm from gradient-based …

Cited by 0SourceScholar
2026

MARSS: Radar Semantic Segmentation via Modular Attention and State Space Models

CVPR 2026

Radar semantic segmentation (RSS) is critical for robust perception in adverse conditions, but poses unique challenges: radar frequency maps are highly anisotropic, multi-scale, sparse and noisy. Conventional CNN or Transformer architectures, designed for camera images, fail to account for these cha

Cited by 0SourceScholar
2026

RAR: Reversing Visual Attention Re-Sinking for Unlocking Potential in Multimodal Large Language Models

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have achieved remarkable success in vision-language tasks, yet they frequently exhibit suboptimal output layers, where intermediate decoder layers outperform the final ones, signaling underutilized model capacity. In this work, we delve into the root causes a…

Cited by 0SourceScholar
2026

UVU: Improving Multimodal Understanding via Vision-Language Unified Autoregressive Paradigm

CVPR 2026

Despite remarkable advancements in multimodal large language models (MLLMs), their fine-grained visual understanding is constrained by a primary reliance on sparse textual supervision. Existing efforts to introduce visual supervision typically do so during post-training, when visual representations

Cited by 0SourceScholar
2026

UrbanMLLM: Joint Learning of Cross-view Imagery for Urban Understanding

ICML 2026poster

Comprehensive urban understanding requires integrating macroscopic spatial structure with fine-grained street-level semantics. However, existing urban Multimodal Large Language Models (MLLMs) primarily rely on satellite imagery, limiting their ability to capture detailed urban appearance and cross-v…

Cited by 0SourceScholar
2026

WeightFlow: Learning Stochastic Dynamics via Evolving Weight of Neural Network

AAAI 2026technical

Modeling stochastic dynamics from discrete observations is a key interdisciplinary challenge. Existing methods often fail to estimate the continuous evolution of probability densities from trajectories or face the curse of dimensionality. To address these limitations, we presents a novel paradigm:

Cited by 0SourcePDFScholar
2025

DM-Adapter: Domain-Aware Mixture-of-Adapters for Text-Based Person Retrieval

AAAI 2025technical

Text-based person retrieval (TPR) has gained significant attention as a fine-grained and challenging task that closely aligns with practical applications. Tailoring CLIP to person domain is now a emerging research topic due to the abundant knowledge of vision-language pretraining, but challenges sti…

2025

LLM-Explorer: A Plug-in Reinforcement Learning Policy Exploration Enhancement Driven by Large Language Models

NeurIPS 2025spotlight

Policy exploration is critical in reinforcement learning (RL), where existing approaches include $\epsilon$-greedy, Gaussian process, etc. However, these approaches utilize preset stochastic processes and are indiscriminately applied in all kinds of RL tasks without considering task-specific feature…

Cited by 0SourcecodeScholar
2025

Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network

ICML 2025poster

Reinforcement learning (RL) for continuous control often requires large amounts of online interaction data. Value-based RL methods can mitigate this burden by offering relatively high sample efficiency. Some studies further enhance sample efficiency by incorporating offline demonstration data to “…

Cited by 0SourcePDFScholar
2025

Pose Magic: Efficient and Temporally Consistent Human Pose Estimation with a Hybrid Mamba-GCN Network

AAAI 2025technical

Current state-of-the-art (SOTA) methods in 3D Human Pose Estimation (HPE) are primarily based on Transformers. However, existing Transformer-based 3D HPE backbones often encounter a trade-off between accuracy and computational efficiency. To resolve the above dilemma, in this work, we leverage recen…

Cited by 3SourcePDFScholar
2025

Predicting the Energy Landscape of Stochastic Dynamical System via Physics-informed Self-supervised Learning

ICLR 2025poster

Energy landscapes play a crucial role in shaping dynamics of many real-world complex systems. System evolution is often modeled as particles moving on a landscape under the combined effect of energy-driven drift and noise-induced diffusion, where the energy governs the long-term motion of the partic…

2025

SAP-SLAM: Semantic-Assisted Perception SLAM with 3D Gaussian Splatting

ICRA 2025

The integration of 3D Gaussians has introduced a novel scene representation in Simultaneous Localization and Mapping (SLAM), characterized by explicit representation and differentiable rendering capabilities that enhance scene reconstruction and understanding. However, most current SLAM systems only

Cited by 0SourceScholar
2025

What Can RL Bring to VLA Generalization? An Empirical Study

NeurIPS 2025poster

Large Vision-Language Action (VLA) models have shown significant potential for embodied AI. However, their predominant training via supervised fine-tuning (SFT) limits generalization due to susceptibility to compounding errors under distribution shifts. Reinforcement learning (RL) offers a path to…

Cited by 0SourcecodeScholar
2024

AdaPKC: PeakConv with Adaptive Peak Receptive Field for Radar Semantic Segmentation

NeurIPS 2024poster

Deep learning-based radar detection technology is receiving increasing attention in areas such as autonomous driving, UAV surveillance, and marine monitoring. Among recent efforts, PeakConv (PKC) provides a solution that can retain the peak response characteristics of radar signals and play the char…

2024

Clip-Based Synergistic Knowledge Transfer for text-based Person Retrieval

ICASSP 2024accepted

Text-based Person Retrieval (TPR) aims to retrieve the target person images given a textual query. The primary challenge lies in bridging the substantial gap between vision and language modalities, especially when dealing with limited large-scale datasets. In this paper, we introduce a CLIP-based Sy…

Cited by 0SourceScholar
2024

EconAgent: Large Language Model-Empowered Agents for Simulating Macroeconomic Activities

ACL 2024long

The advent of artificial intelligence has led to a growing emphasis on data-driven modeling in macroeconomics, with agent-based modeling (ABM) emerging as a prominent bottom-up simulation paradigm. In ABM, agents (*e.g.*, households, firms) interact within a macroeconomic environment, collectively g…

2024

IRGen: Generative Modeling for Image Retrieval

ECCV 2024poster

"While generative modeling has become prevalent across numerous research fields, its integration into the realm of image retrieval remains largely unexplored and underjustified. In this paper, we present a novel methodology, reframing image retrieval as a variant of generative modeling and employing…

2024

Long-term Detection and Monitory of Chinese Urban Village Using Satellite Imagery

IJCAI 2024poster

Urban villages are areas filled with rural-like improvised structures in Chinese cities, usually housing the most vulnerable groups. Under the guidance of the Sustainable Development Goals (SDGs), the Chinese government initiated renewal and redevelopment projects, underscoring the meticulous mapp…

2024

Reschedule Diffusion-based Bokeh Rendering

IJCAI 2024poster

Bokeh rendering for images shot with small apertures has drawn much attention in practice. Very recently people start to explore diffusion models for bokeh rendering, aiming to leverage the models' surging power of image generation. However, we can clearly observe two big issues with the images rend…

2024

Taming Diffusion Prior for Image Super-Resolution with Domain Shift SDEs

NeurIPS 2024poster

Diffusion-based image super-resolution (SR) models have attracted substantial interest due to their powerful image restoration capabilities. However, prevailing diffusion models often struggle to strike an optimal balance between efficiency and performance. Typically, they either neglect to exploit…

2024

UV-SAM: Adapting Segment Anything Model for Urban Village Identification

AAAI 2024technical

Urban villages, defined as informal residential areas in or around urban centers, are characterized by inadequate infrastructures and poor living conditions, closely related to the Sustainable Development Goals (SDGs) on poverty, adequate housing, and sustainable cities. Traditionally, governments h…

2023

Dynamic Ensemble of Low-Fidelity Experts: Mitigating NAS “Cold-Start”

AAAI 2023technical

Predictor-based Neural Architecture Search (NAS) employs an architecture performance predictor to improve the sample efficiency. However, predictor-based NAS suffers from the severe ``cold-start'' problem, since a large amount of architecture-performance data is required to get a working predictor.…

2023

Flowpose: Conditional Normalizing Flows for 3D Human Pose and Shape Estimation from Monocular Videos

ICASSP 2023accepted

Human motion modeling is essential for video-based 3D human pose and shape estimation. Most existing methods model human motion by learning a deterministic mapping from the input videos to the human body parameters, while the uncertainties such as occlusions and depth ambiguities are ignored. To add…

Cited by 0SourceScholar
2023

Robust Content-Variant Reference Image Quality Assessment Via Similar Patch Matching

ICASSP 2023accepted

Although image quality assessment (IQA) methods have achieved remarkable success in the past decades, full-reference IQA is limited to reference images, while no-reference IQA has relatively poor performance. To boost the performance of IQA models in the no-reference scenario, a new class of IQA met…

Cited by 0SourceScholar
2022

Coarse-to-Fine Embedded PatchMatch and Multi-Scale Dynamic Aggregation for Reference-Based Super-resolution

AAAI 2022technical

Reference-based super-resolution (RefSR) has made significant progress in producing realistic textures using an external reference (Ref) image. However, existing RefSR methods obtain high-quality correspondence matchings consuming quadratic computation resources with respect to the input size, limit…

2022

Efficient Non-local Contrastive Attention for Image Super-resolution

AAAI 2022technical

Non-Local Attention (NLA) brings significant improvement for Single Image Super-Resolution (SISR) by leveraging intrinsic feature correlation in natural images. However, NLA gives noisy information large weights and consumes quadratic computation resources with respect to the input size, limiting it…

2022

Pose-Invariant Face Recognition via Adaptive Angular Distillation

AAAI 2022technical

Pose-invariant face recognition is a practically useful but challenging task. This paper introduces a novel method to learn pose-invariant feature representation without normalizing profile faces to frontal ones or learning disentangled features. We first design a novel strategy to learn pose-invari…

Cited by 3SourcePDFScholar
2022

SCS-Co: Self-Consistent Style Contrastive Learning for Image Harmonization

CVPR 2022poster

Image harmonization aims to achieve visual consistency in composite images by adapting a foreground to make it compatible with a background. However, existing methods always only use the real image as the positive sample to guide the training, and at most introduce the corresponding composite image…

Cited by 53PDFcodeScholar
2021

Group Fisher Pruning for Practical Network Compression

ICML 2021spotlight

Network compression has been widely studied since it is able to reduce the memory and computation cost during inference. However, previous methods seldom deal with complicated structures like residual connections, group/depth-wise convolution and feature pyramid network, where channels of multiple l…

2021

Towards Impartial Multi-task Learning

ICLR 2021poster

Multi-task learning (MTL) has been widely used in representation learning. However, naively training all tasks simultaneously may lead to the partial training issue, where specific tasks are trained more adequately than others. In this paper, we propose to learn multiple tasks impartially. Specifica…

Cited by 198SourcePDFScholar
2018

Full-Reference Quality Assessment of Contrast Changed Images Based on Local Linear Model

ICASSP 2018accepted

This paper presents a new full-reference method to assess the quality of contrast changed images. In this method, we employ a linear model to describe the relationship between local patches of reference images and contrast changed images. With parameters of this model, three quality measures conside…

Cited by 0SourceScholar
2017

Illumination-robust face recognition with Block-based Local Contrast Patterns

ICASSP 2017accepted

This paper proposes a novel facial image representation Block-based Local Contrast Patterns (BLCP) for illumination-robust face recognition. This method is based on an effective texture descriptor local contrast patterns (LCP). We use the directed and undirected difference masks to calculate three t…

Cited by 0SourceScholar
2017

Locality Sensitive Hashing based deepmatching for optical flow estimation

ICASSP 2017accepted

DeepMatching (DM) is one of the state-of-art matching algorithms to compute quasi-dense correspondences between images. Recent optical flow methods use DeepMatching to find initial image correspondences and achieves outstanding performance. However, the key building block of DeepMatching, the correl…

Cited by 0SourceScholar
2017

Wavelet-based single image super-resolution with an overall enhancement procedure

ICASSP 2017accepted

In this paper, we address the problem of generating a super-resolution image based on a dictionary of low- and high-resolution exemplars from a single input image in wavelet domain with a overall enhancement procedure. Most methods extract different kinds of features in low-resolution image and high…

Cited by 0SourceScholar
2016

An efficient anomaly detection approach in surveillance video based on oriented GMM

ICASSP 2016accepted

The detection and localization of abnormal activities are considered in this work. An efficient approach called oriented G-MM(OGMM) is proposed. The approach uses optical flow as low-level feature and quantizes the orientation of optical flow into 8 sections. In training stage, the approach will lea…

Cited by 0SourceScholar
2016

Face recognition with local contourlet combined patterns

ICASSP 2016accepted

This paper proposes a novel face image descriptor called local contourlet combined patterns (LCCP), based on the Non-Subsampled Contourlet Transform (NSCT), for face recognition. NSCT is a multiresolution analysis tool and can capture image information at multiple scales, orientations, and frequency…

Cited by 0SourceScholar
2016

Saliency detection based on integration of central bias, reweighting and multi-scale for superpixels

ICASSP 2016accepted

Saliency detection has been a significant problem in computer vision and helpful to object detection. In this paper, we propose a new computational saliency detection model under the Bayesian framework. First, central bias and the reweighting of the salient regions in the convex hull are applied to…

Cited by 0SourceScholar