← Search

Weisi Lin

53 accepted papers

2026

DipGuava: Disentangling Personalized Gaussian Features for 3D Head Avatars from Monocular Video

AAAI 2026technical

While recent 3D head avatar creation methods attempt to animate facial dynamics, they often fail to capture personalized details, limiting realism and expressiveness. To fill this gap, we present DipGuava (Disentangled and Personalized Gaussian UV Avatar), a novel 3D Gaussian head avatar creation me

Cited by 0SourcePDFScholar
2026

Image Quality Assessment for Embodied AI

ICLR 2026poster

Embodied AI has developed rapidly in recent years, but it is still mainly deployed in laboratories, with various distortions in the Real-world limiting its application. Traditionally, Image Quality Assessment (IQA) methods are applied to predict human preferences for distorted images; however, there…

Cited by 0SourcecodeScholar
2026

MachaGrasp: Morphology-Aware Cross-Embodiment Dexterous Hand Articulation Generation for Grasping

ICRA 2026poster

Dexterous grasping with multi-fingered hands remains challenging due to high-dimensional articulations and the cost of optimization-based pipelines. Existing end-to-end methods require training on large-scale datasets for specific hands, limiting their ability to generalize across different embodime…

2026

Nighttime Flare Removal via Wavelet-Guided and Gated-Enhanced Spatial-Frequency Fusion Network

AAAI 2026technical

Nighttime flares, caused by complex scattering and reflections from artificial light sources, significantly degrade image quality and hinder downstream visual tasks. Existing deflare networks usually struggle to jointly capture and fuse latent spatial and frequency features. In this paper, we propos

Cited by 0SourcePDFScholar
2026

QD-PCQA: Quality-Aware Domain Adaptation for Point Cloud Quality Assessment

CVPR 2026

No-Reference Point Cloud Quality Assessment (NR-PCQA) still struggles with generalization, primarily due to the scarcity of annotated point cloud datasets. Since the Human Visual System (HVS) drives perceptual quality assessment independently of media types, prior knowledge on quality learned from i

Cited by 0SourcecodeScholar
2026

R4-CGQA: Retrieval-based Vision Language Models for Computer Graphics Image Quality Assessment

CVPR 2026

Immersive Computer Graphics (CGs) rendering has become ubiquitous in modern daily life. However, comprehensively evaluating CG quality remains challenging for two reasons: (1) existing CG datasets lack systematic descriptions of rendering quality; and (2) existing CG quality assessment methods canno

Cited by 0SourcecodeScholar
2026

SCALING AUDIO-VISUAL QUALITY ASSESSMENT DATASET VIA CROWDSOURCING

ICASSP 2026oral

Audio-visual quality assessment (AVQA) research has been stalled by limitations of existing datasets: they are typically small in scale, with insufficient diversity in content and quality, and annotated only with overall scores. These shortcomings provide limited support for model development and mu…

Cited by 0SourcePDFScholar
2026

Semantic Contact Fields for Category-Level Generalizable Tool Manipulation

RSS 2026poster

Generalizing tool manipulation requires both semantic planning and precise physical control. Modern generalist robot policies, such as Vision-Language-Action (VLA) models, often lack the high-fidelity physical grounding required for contact-rich tool manipulation. Conversely, existing contact-aware …

Cited by 0SourceScholar
2026

VisualScore: Learning Holistic Visual Quality Scores via Multi-Task Reasoning

ICML 2026poster

Image quality assessment (IQA) is inherently multi-mage quality assessment (IQA) is inherently multi-dimensional, yet existing reward models are typically limited to a single task and become unstable when extended to multi-task settings. In particular, heterogeneous reward scales and variances acros…

Cited by 0SourceScholar
2025

A-Bench: Are LMMs Masters at Evaluating AI-generated Images?

ICLR 2025poster

How to accurately and efficiently assess AI-generated images (AIGIs) remains a critical challenge for generative models. Given the high costs and extensive time commitments required for user studies, many researchers have turned towards employing large multi-modal models (LMMs) as AIGI evaluators, t…

2025

Deep Learning Based Topography Aware Gas Source Localization with Mobile Robot

ICRA 2025

Gas source localization in complex environments is critical for applications such as environmental monitoring, industrial safety, and disaster response. Traditional methods often struggle with the challenges posed by a lack of environmental topography integration, especially when interactions betwee

Cited by 0SourceScholar
2025

Explore the Hallucination on Low-level Perception for MLLMs

ICASSP 2025accepted

The rapid development of Multi-modality Large Language Models (MLLMs) has significantly influenced various aspects of industry and daily life, showcasing impressive capabilities in visual perception and understanding. However, these models also exhibit hallucinations, which limit their reliability a…

Cited by 0SourceScholar
2025

Feature Coding in the Era of Large Models: Dataset, Test Conditions, and Benchmark

ICCV 2025poster

Large models have achieved remarkable performance across various tasks, yet they incur significant computational costs and privacy concerns during both training and inference. Distributed deployment has emerged as a potential solution, but it necessitates the exchange of intermediate information bet…

2025

Image Quality Assessment: From Human to Machine Preference

CVPR 2025highlight

Image Quality Assessment (IQA) based on human subjective preferences has undergone extensive research in the past decades. However, with the development of communication protocols, the visual data consumption volume of machines has gradually surpassed that of humans. For machines, the preference dep…

2025

Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression

NeurIPS 2025poster

Large Language Models (LLMs) have demonstrated remarkable capabilities but typically require extensive computational resources and memory for inference. Post-training quantization (PTQ) can effectively reduce these demands by storing weights in lower bit-width formats. However, standard uniform quan…

Cited by 0SourcecodeScholar
2025

Q-Bench-Video: Benchmark the Video Quality Understanding of LMMs

CVPR 2025poster

With the rising interest in research on Large Multi-modal Models (LMMs) for video understanding, many studies have emphasized general video comprehension capabilities, neglecting the systematic exploration into video quality understanding. To address this oversight, we introduce Q-Bench-Video in thi…

2025

Robust-PIFu: Robust Pixel-aligned Implicit Function for 3D Human Digitalization from a Single Image

ICLR 2025poster

Existing methods for 3D clothed human digitalization perform well when the input image is captured in ideal conditions that assume the lack of any occlusion. However, in reality, images may often have occlusion problems such as incomplete observation of the human subject's full body, self-occlusion…

Cited by 0SourcePDFScholar
2025

Unveiling the Invisible: Reasoning Complex Occlusions Amodally with AURA

ICCV 2025poster

Amodal segmentation aims to infer the complete shape of occluded objects, even when the occluded region's appearance is unavailable. However, current amodal segmentation methods lack the capability to interact with users through text input and struggle to understand or reason about implicit and comp…

2024

3DFG-PIFu: 3D Feature Grids for Human Digitization from Sparse Views

ECCV 2024poster

"Pixel-aligned implicit models, such as Multi-view PIFu, DeepMultiCap, DoubleField, and SeSDF, are well-established methods for reconstructing a clothed human from sparse views. However, given V images, these models would only combine features from these images in a point-wise and localized manner.…

2024

Adaptive Image Quality Assessment via Teaching Large Multimodal Model to Compare

NeurIPS 2024spotlight

While recent advancements in large multimodal models (LMMs) have significantly improved their abilities in image quality assessment (IQA) relying on absolute quality rating, how to transfer reliable relative quality comparison outputs to continuous perceptual quality scores remains largely unexplore…

2024

Boosting Image Quality Assessment through Efficient Transformer Adaptation with Local Feature Enhancement

CVPR 2024poster

Image Quality Assessment (IQA) constitutes a fundamental task within the field of computer vision yet it remains an unresolved challenge owing to the intricate distortion conditions diverse image contents and limited availability of data. Recently the community has witnessed the emergence of numerou…

2024

Fine Structure-Aware Sampling: A New Sampling Training Scheme for Pixel-Aligned Implicit Models in Single-View Human Reconstruction

AAAI 2024technical

Pixel-aligned implicit models, such as PIFu, PIFuHD, and ICON, are used for single-view clothed human reconstruction. These models need to be trained using a sampling training scheme. Existing sampling training schemes either fail to capture thin surfaces (e.g. ears, fingers) or cause noisy artefact…

2024

Iterative Token Evaluation and Refinement for Real-World Super-resolution

AAAI 2024technical

Real-world image super-resolution (RWSR) is a long-standing problem as low-quality (LQ) images often have complex and unidentified degradations. Existing methods such as Generative Adversarial Networks (GANs) or continuous diffusion models present their own issues including GANs being difficult to t…

2024

PeVL: Pose-Enhanced Vision-Language Model for Fine-Grained Human Action Recognition

CVPR 2024poster

Recent progress in Vision-Language (VL) foundation models has revealed the great advantages of cross-modality learning. However due to a large gap between vision and text they might not be able to sufficiently utilize the benefits of cross-modality information. In the field of human action recogniti…

Cited by 3SourcePDFScholar
2024

Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

ICML 2024poster

The explosion of visual content available online underscores the requirement for an accurate machine assessor to robustly evaluate scores across diverse types of visual contents. While recent studies have demonstrated the exceptional potentials of large multi-modality models (LMMs) on a wide range o…

2024

Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

ICLR 2024spotlight

The rapid evolution of Multi-modality Large Language Models (MLLMs) has catalyzed a shift in computer vision from specialized models to general-purpose foundation models. Nevertheless, there is still an inadequacy in assessing the abilities of MLLMs on **low-level visual perception and understanding…

2024

Q-Instruct: Improving Low-level Visual Abilities for Multi-modality Foundation Models

CVPR 2024poster

Multi-modality large language models (MLLMs) as represented by GPT-4V have introduced a paradigm shift for visual perception and understanding tasks that a variety of abilities can be achieved within one foundation model. While current MLLMs demonstrate primary low-level visual abilities from the id…

2024

R-Cyclic Diffuser: Reductive and Cyclic Latent Diffusion for 3D Clothed Human Digitalization

CVPR 2024poster

Recently the authors of Zero-1-to-3 demonstrated that a latent diffusion model pretrained with Internet-scale data can not only address the single-view 3D object reconstruction task but can even attain SOTA results in it. However when applied to the task of single-view 3D clothed human reconstructio…

2023

Efficient Joint Optimization of Layer-Adaptive Weight Pruning in Deep Neural Networks

ICCV 2023poster

In this paper, we propose a novel layer-adaptive weight-pruning approach for Deep Neural Networks (DNNs) that addresses the challenge of optimizing the output distortion minimization while adhering to a target pruning ratio constraint. Our approach takes into account the collective influence of all…

Cited by 29PDFcodeScholar
2023

Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives

ICCV 2023poster

The rapid increase in user-generated-content (UGC) videos calls for the development of effective video quality assessment (VQA) algorithms. However, the objective of the UGC-VQA problem is still ambiguous and can be viewed from two perspectives: the technical perspective, measuring the perception of…

Cited by 162PDFcodeScholar
2023

GCFAgg: Global and Cross-View Feature Aggregation for Multi-View Clustering

CVPR 2023poster

Multi-view clustering can partition data samples into their categories by learning a consensus representation in unsupervised way and has received more and more attention in recent years. However, most existing deep clustering methods learn consensus representation or view-specific representations f…

2022

CMUA-Watermark: A Cross-Model Universal Adversarial Watermark for Combating Deepfakes

AAAI 2022technical

Malicious applications of deepfakes (i.e., technologies generating target facial attributes or entire faces from facial images) have posed a huge threat to individuals' reputation and security. To mitigate these threats, recent studies have proposed adversarial watermarks to combat deepfake models,…

2022

FAST-VQA: Efficient End-to-End Video Quality Assessment with Fragment Sampling

ECCV 2022poster

"Current deep video quality assessment (VQA) methods are usually with high computational costs when evaluating high-resolution videos. This cost hinders them from learning better video-quality-related representations via end-to-end training. Existing approaches typically consider naive sampling to r…

2022

IntegratedPIFu: Integrated Pixel Aligned Implicit Function for Single-View Human Reconstruction

ECCV 2022poster

"We propose IntegratedPIFu, a new pixel-aligned implicit model that builds on the foundation set by PIFuHD. IntegratedPIFu shows how depth and human parsing information can be predicted and capitalized upon in a pixel-aligned implicit model. In addition, IntegratedPIFu introduces depth-oriented samp…

2022

S-PIFu: Integrating Parametric Human Models with PIFu for Single-view Clothed Human Reconstruction

NeurIPS 2022accept

We present three novel strategies to incorporate a parametric body model into a pixel-aligned implicit model for single-view clothed human reconstruction. Firstly, we introduce ray-based sampling, a novel technique that transforms a parametric model into a set of highly informative, pixel-aligned 2D…

2021

Low Resolution Information Also Matters: Learning Multi-Resolution Representations for Person Re-Identification

IJCAI 2021poster

As a prevailing task in video surveillance and forensics field, person re-identification (re-ID) aims to match person images captured from non-overlapped cameras. In unconstrained scenarios, person images often suffer from the resolution mismatch problem, i.e., Cross-Resolution Person Re-ID. To over…

Cited by 30SourcePDFScholar
2019

Cascaded Parallel Filtering for Memory-Efficient Image-Based Localization

ICCV 2019poster

Image-based localization (IBL) aims to estimate the 6DOF camera pose for a given query image. The camera pose can be computed from 2D-3D matches between a query image and Structure-from-Motion (SfM) models. Despite recent advances in IBL, it remains difficult to simultaneously resolve the memory con…

Cited by 38PDFcodeScholar
2019

Towards Robust Curve Text Detection With Conditional Spatial Expansion

CVPR 2019poster

It is challenging to detect curve texts due to their irregular shapes and varying sizes. In this paper, we first investigate the deficiency of the existing curve detection methods and then propose a novel Conditional Spatial Expansion (CSE) mechanism to improve the performance of curve detection. In…

Cited by 105PDFScholar
2018

Image Quality Assessment Based Label Smoothing in Deep Neural Network Learning

ICASSP 2018accepted

For many computer vision problems, deep neural networks are trained and validated based on the assumption that the input images are pristine (i.e., artifact-free). However, digital images are subject to a wide range of distortions in real application scenarios, while the practical issues regarding i…

Cited by 0SourceScholar
2018

Learning Markov Clustering Networks for Scene Text Detection

CVPR 2018poster

A novel framework named Markov Clustering Network (MCN) is proposed for fast and robust scene text detection. MCN predicts instance-level bounding boxes by firstly converting an image into a Stochastic Flow Graph (SFG) and then performing Markov Clustering on this graph. Our method can detect text o…

Cited by 139SourcePDFScholar
2016

Aspect Ratio Similarity (ARS) for image retargeting quality assessment

ICASSP 2016accepted

During the past few years, there have been various kinds of content-aware image retargeting methods proposed for image resizing. However, the lack of effective objective retargeting quality metric limits the further development of image retargeting. Different from the traditional image quality asses…

Cited by 0SourceScholar
2016

Enhanced just noticeable difference model with visual regularity consideration

ICASSP 2016accepted

Just noticeable difference (JND) reveals the visibility of our human visual system (HVS), below which changes cannot be perceived by the human. Though dozens of JND estimation models have been introduced during the past decade, how to accurately estimate the JND thresholds for different content regi…

Cited by 0SourceScholar
2015

B-SHOT: A binary feature descriptor for fast and efficient keypoint matching on 3D point clouds

IROS 2015poster

In this paper, we introduce the very first ‘binary’ 3D feature descriptor, B-SHOT, for fast and efficient keypoint matching on 3D point clouds. We propose a binary quantization method that converts a real valued vector to a binary vector. We apply this method on a state-of-the-art 3D feature descrip…

Cited by 89SourceScholar
2015

Multi-task rank learning for image quality assessment

ICASSP 2015accepted

In practice, multiple types of distortions are associated with an image quality degradation process. The existing machine learning (ML) based image quality assessment (IQA) approaches generally established a unified model for all distortion types, or each model is trained independently for each dist…

Cited by 0SourceScholar