← Search

Guangrun Wang

32 accepted papers

2026

Human-Centric Open-Future Task Discovery: Formulation, Benchmark, and Scalable Tree-Based Search

AAAI 2026technical

Recent progress in robotics and embodied AI is largely driven by Large Multimodal Models (LMMs). However, a key challenge remains underexplored: how can we advance LMMs to discover tasks that assist humans in open-future scenarios, where human intentions are highly concurrent and dynamic. In this wo

Cited by 0SourcePDFScholar
2026

VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modeling

CVPR 2026

Vision-language-action (VLA) models achieve strong in-distribution performance but degrade sharply under novel camera viewpoints and visual perturbations. We show that this brittleness primarily arises from misalignment in Spatial Modeling, rather than Physical Modeling. To address this, we propose

Cited by 0SourceScholar
2025

Anima2: Cross-Species Animal Animation through Image-to-Video Synthesis with Subject Alignment

ICASSP 2025accepted

Recent video editing advancements rely on accurate pose sequences to animate human actors. However, these efforts are not suitable for cross-species animation due to pose misalignment between species (for example, the poses of a cat differ greatly from that of a pig due to their distinct body struct…

Cited by 0SourceScholar
2025

VTON 360: High-Fidelity Virtual Try-On from Any Viewing Direction

CVPR 2025poster

Virtual Try-On (VTON) is a transformative technology in e-commerce and fashion design, enabling realistic digital visualization of clothing on individuals. In this work, we propose VTON 360, a novel 3D VTON method that addresses the open challenge of achieving high-fidelity VTON that supports any-vi…

Cited by 1SourcePDFScholar
2024

AlignMiF: Geometry-Aligned Multimodal Implicit Field for LiDAR-Camera Joint Synthesis

CVPR 2024highlight

Neural implicit fields have been a de facto standard in novel view synthesis. Recently there exist some methods exploring fusing multiple modalities within a single field aiming to share implicit features from different modalities to enhance reconstruction performance. However these modalities often…

2024

GaussCtrl: Multi-View Consistent Text-Driven 3D Gaussian Splatting Editing

ECCV 2024poster

"We propose , a text-driven method to edit a 3D scene reconstructed by the 3D Gaussian Splatting (3DGS). Our method first renders a collection of images by using the 3DGS and edits them by using a pre-trained 2D diffusion model (ControlNet) based on the input prompt, which is then used to optimise t…

2024

Language-Free Compositional Action Generation via Decoupling Refinement

ICASSP 2024accepted

Composing simple actions into complex actions is crucial yet challenging. Existing methods largely rely on language annotations to discern composable latent semantics, which is costly and labor-intensive. In this study, we introduce a novel framework to generate compositional actions without languag…

Cited by 0SourceScholar
2024

Making Large Language Models Better Planners with Reasoning-Decision Alignment

ECCV 2024oral

"Data-driven approaches for autonomous driving (AD) have been widely adopted in the past decade but are confronted with dataset bias and uninterpretability. Inspired by the knowledge-driven nature of human driving, recent approaches explore the potential of large language models (LLMs) to improve un…

Cited by 12SourcePDFScholar
2024

NeRF-VPT: Learning Novel View Representations with Neural Radiance Fields via View Prompt Tuning

AAAI 2024technical

Neural Radiance Fields (NeRF) have garnered remarkable success in novel view synthesis. Nonetheless, the task of generating high-quality images for novel views persists as a critical challenge. While the existing efforts have exhibited commendable progress, capturing intricate details, enhancing tex…

2024

WildVidFit: Video Virtual Try-On in the Wild via Image-Based Controlled Diffusion Models

ECCV 2024poster

"Video virtual try-on aims to generate realistic sequences that maintain garment identity and adapt to a person’s pose and body shape in source videos. Traditional image-based methods, relying on warping and blending, struggle with complex human movements and occlusions, limiting their effectiveness…

Cited by 11SourcePDFScholar
2023

LAW-Diffusion: Complex Scene Generation by Diffusion with Layouts

ICCV 2023poster

Thanks to the rapid development of diffusion models, unprecedented progress has been witnessed in image synthesis. Prior works mostly rely on pre-trained linguistic models, but a text is often too abstract to properly specify all the spatial properties of an image, e.g., the layout configuration of…

Cited by 14PDFScholar
2023

MixReorg: Cross-Modal Mixed Patch Reorganization is a Good Mask Learner for Open-World Semantic Segmentation

ICCV 2023poster

Recently, semantic segmentation models trained with image-level text supervision have shown promising results in challenging open-world scenarios. However, these models still face difficulties in learning fine-grained semantic alignment at the pixel level and predicting accurate object masks. To add…

Cited by 19PDFScholar
2023

Tem-Adapter: Adapting Image-Text Pretraining for Video Question Answer

ICCV 2023poster

Video-language pre-trained models have shown remarkable success in guiding video question-answering (VideoQA) tasks. However, due to the length of video sequences, training large-scale video-based models incurs considerably higher costs than training image-based ones. This motivates us to leverage t…

Cited by 17PDFcodeScholar
2023

ViewCo: Discovering Text-Supervised Segmentation Masks via Multi-View Semantic Consistency

ICLR 2023poster

Recently, great success has been made in learning visual representations from text supervision, facilitating the emergence of text-supervised semantic segmentation. However, existing works focus on pixel grouping and cross-modal semantic alignment, while ignoring the correspondence among multiple au…

2022

Automated Progressive Learning for Efficient Training of Vision Transformers

CVPR 2022poster

Recent advances in vision Transformers (ViTs) have come with a voracious appetite for computing power, high-lighting the urgent need to develop efficient training methods for ViTs. Progressive learning, a training scheme where the model capacity grows progressively during training, has started showi…

Cited by 49PDFcodeScholar
2022

Beyond Fixation: Dynamic Window Visual Transformer

CVPR 2022poster

Recently, a surge of interest in visual transformers is to reduce the computational cost by limiting the calculation of self-attention to a local window. Most current work uses a fixed single-scale window for modeling by default, ignoring the impact of window size on model performance. However, this…

Cited by 41PDFcodeScholar
2022

Semantic-Aware Auto-Encoders for Self-Supervised Representation Learning

CVPR 2022poster

The resurgence of unsupervised learning can be attributed to the remarkable progress of self-supervised learning, which includes generative (G) and discriminative (D) models. In computer vision, the mainstream self-supervised learning algorithms are D models. However, designing a D model could be ov…

Cited by 11PDFcodeScholar
2022

Structure-Preserving 3D Garment Modeling with Neural Sewing Machines

NeurIPS 2022accept

3D Garment modeling is a critical and challenging topic in the area of computer vision and graphics, with increasing attention focused on garment representation learning, garment reconstruction, and controllable garment manipulation, whereas existing methods were constrained to model garments under…

Cited by 16SourcePDFScholar
2021

BossNAS: Exploring Hybrid CNN-Transformers With Block-Wisely Self-Supervised Neural Architecture Search

ICCV 2021poster

A myriad of recent breakthroughs in hand-crafted neural architectures for visual recognition have highlighted the urgent need to explore hybrid architectures consisting of diversified building blocks. Meanwhile, neural architecture search methods are surging with an expectation to reduce human effor…

Cited by 142PDFcodeScholar
2021

EfficientBERT: Progressively Searching Multilayer Perceptron via Warm-up Knowledge Distillation

EMNLP 2021finding

Pre-trained language models have shown remarkable results on various NLP tasks. Nevertheless, due to their bulky size and slow inference speed, it is hard to deploy them on edge devices. In this paper, we have a critical insight that improving the feed-forward network (FFN) in BERT has a higher gain…

2021

Pi-NAS: Improving Neural Architecture Search by Reducing Supernet Training Consistency Shift

ICCV 2021poster

Recently proposed neural architecture search (NAS) methods co-train billions of architectures in a supernet and estimate their potential accuracy using the network weights detached from the supernet. However, the ranking correlation between the architectures' predicted accuracy and their actual capa…

Cited by 22PDFcodeScholar
2021

Solving Inefficiency of Self-Supervised Representation Learning

ICCV 2021poster

Self-supervised learning (especially contrastive learning) has attracted great interest due to its huge potential in learning discriminative representations in an unsupervised manner. Despite the acknowledged successes, existing contrastive learning methods suffer from very low learning efficiency,…

Cited by 66PDFcodeScholar
2020

Block-Wisely Supervised Neural Architecture Search With Knowledge Distillation

CVPR 2020poster

Neural Architecture Search (NAS), aiming at automatically designing network architectures by machines, is expected to bring about a new revolution in machine learning. Despite these high expectation, the effectiveness and efficiency of existing NAS solutions are unclear, with some recent works going…

Cited by 244PDFcodeScholar
2020

EagleEye: Fast Sub-net Evaluation for Efficient Neural Network Pruning

ECCV 2020poster

Finding out the computational redundant part of a trained Deep Neural Network (DNN) is the key question that pruning algorithms target on. Many algorithms try to predict model performance of the pruned sub-nets by introducing various evaluation methods. But they are either inaccurate or very complic…

2020

Smoothing Adversarial Domain Attack and P-Memory Reconsolidation for Cross-Domain Person Re-Identification

CVPR 2020poster

Most of the existing person re-identification (re-ID) methods achieve promising accuracy in a supervised manner, but they assume the identity labels of the target domain is available. This greatly limits the scalability of person re-ID in real-world scenarios. Therefore, the current person re-ID com…

Cited by 85PDFScholar
2020

Transferable, Controllable, and Inconspicuous Adversarial Attacks on Person Re-identification With Deep Mis-Ranking

CVPR 2020oral

The success of DNNs has driven the extensive applications of person re-identification (ReID) into a new era. However, whether ReID inherits the vulnerability of DNNs remains unexplored. To examine the robustness of ReID systems is rather important because the insecurity of ReID systems may cause sev…

Cited by 106PDFcodeScholar
2018

Kalman Normalization: Normalizing Internal Representations Across Network Layers

NeurIPS 2018poster

As an indispensable component, Batch Normalization (BN) has successfully improved the training of deep neural networks (DNNs) with mini-batches, by normalizing the distribution of the internal representation for each hidden layer. However, the effectiveness of BN would diminish with the scenario of…

Cited by 31SourcePDFScholar
2017

Learning Object Interactions and Descriptions for Semantic Image Segmentation

CVPR 2017poster

Recent advanced deep convolutional networks (CNNs) achieved great successes in many computer vision tasks, because of their compelling learning complexity and the presences of large-scale labeled data. However, as obtaining per-pixel annotations is expensive, performances of CNNs in semantic image s…

Cited by 58PDFScholar
2016

Deep Structured Scene Parsing by Learning With Image Descriptions

CVPR 2016oral

This paper addresses the problem of structured scene parsing, i.e., parsing the input scene into a configuration including hierarchical semantic objects with their interaction relations. We propose a deep architecture consisting of two networks: i) a convolutional neural network (CNN) extracting the…

Cited by 40PDFScholar