← Search

QIANQIAN WANG

76 accepted papers

2026

A Hybrid Magnetic Actuation System for Hybrid Microrobotic Targeted Delivery

ICRA 2026poster

Magnetic microrobots hold great promise for biomedical applications. However, achieving flexible magnetic field adjustment with a magnetic actuation system (MAS) to actuate diverse microrobots remains a significant challenge. In this work, we propose an Electromagnetic-Permanent Magnet Actuation (EP…

Cited by 0Scholar
2026

Debiased and Denoised Projection Learning for Incomplete Multi-view Clustering

ICLR 2026poster

Multi-view clustering achieves outstanding performance but relies on the assumption of complete multi-view samples. However, certain views may be partially unavailable due to failures during acquisition or storage, resulting in distribution shifts across views. Although some incomplete multi-view cl…

Cited by 0SourceScholar
2026

Discriminative Graph Embedding Framework via Label-Free Marginal Fisher Analysis

AAAI 2026technical

Marginal Fisher Analysis (MFA) is a classical dimensionality reduction (DR) method that leverages dual graphs to capture intra-class compactness and inter-class separability. However, MFA’s reliance on high-quality labels limits its practical application. For another, existing unsupervised DR method

Cited by 0SourcePDFScholar
2026

E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training

CVPR 2026

Self-supervised pre-training has driven rapid progress in foundation models for language, 2D images, and video, yet remains largely unexplored for learning 3D-aware representations from multi-view images. In this paper, we present E-RayZer, a self-supervised 3D vision model that learns geometrically

Cited by 0SourcecodeScholar
2026

Emergent Outlier View Rejection in Visual Geometry Grounded Transformers

CVPR 2026

Reliable 3D reconstruction from in-the-wild image collections is often hindered by noisy images--irrelevant inputs with little or no view overlap with others. While traditional Structure-from-Motion pipelines handle such cases through geometric verification and outlier rejection, feed-forward 3D rec

Cited by 0SourcecodeScholar
2026

Federated Incomplete Multi-View Clustering with Tensorized Low-Rank Constraint

AAAI 2026technical

Federated Multi-View Clustering has gained increasing attention for its ability to discover complementary clustering structures of distributed multi-view data while preserving data privacy. However, real-world clients often only have access to partial views, and the view incompleteness poses great c

Cited by 0SourcePDFScholar
2026

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly

CVPR 2026

The emergence of Large Vision-Language Models (LVLMs) has significantly advanced video understanding capabilities. However, existing benchmarks focus predominantly on coarse-grained tasks such as action segmentation, classification, captioning, and retrieval. Furthermore, these benchmarks often rely

Cited by 0SourcecodeScholar
2026

Gaze-Based Teleoperation with Intent Inference Model for Robotic Manipulators

ICRA 2026poster

Eye gaze-based control interfaces provide a non-invasive means of enhancing human-robot collaboration for activities of daily living and can reduce the cognitive burden on operators performing complex tasks. Eye gaze has traditionally been used for "gaze triggering," where fixating on an object acti…

Cited by 0Scholar
2026

Learning to Control the Whole-Body Shape of a Soft Robotic Arm in Unknown Situations

ICRA 2026poster

Control of soft robots is considered one of the key elements in achieving their intelligence. However, it faces challenging problems such as nonlinear dynamics, highly deformable structures, and operation in unpredictable situations. Numerous methods have been proposed to overcome these challenges, …

Cited by 0Scholar
2026

MFmamba: A Multi-function Network for Panchromatic Image Resolution Restoration Based on State-Space Model

AAAI 2026technical

Remote sensing images are becoming increasingly widespread in military, earth resource exploration. Because of the limitation of a single sensor, we can obtain high spatial resolution grayscale panchromatic (PAN) images and low spatial resolution color multispectral (MS) images. Therefore, an import

Cited by 0SourcePDFScholar
2026

Rethinking Cross-Modal Anchor Alignment for Mitigating Error Accumulation

CVPR 2026

Mitigating noisy correspondence in cross-modal matching poses a serious challenge due to the problem of error accumulation. Existing methods primarily attribute this accumulation to errors caused by noisy sample pairs. However, a novel source of error from clean sample pairs (also termed anchor pair

Cited by 0SourceScholar
2025

A Simple yet Effective Hypergraph Clustering Network

IJCAI 2025

Hypergraph Clustering has gained significant attention due to its capability of capturing high order structural information. Among different approaches, contrastive learning-based methods leverage self-supervised learning and data augmentation, exhibiting impressive performance. However, most of the

Cited by 0SourcePDFScholar
2025

Bridging Viewpoint Gaps: Geometric Reasoning Boosts Semantic Correspondence

CVPR 2025poster

Finding semantic correspondences between images is a challenging problem in computer vision, particularly under significant viewpoint changes. Previous methods rely on semantic features from pre-trained 2D models like Stable Diffusion and DINOv2, which often struggle to extract viewpoint-invariant f…

Cited by 0SourcePDFScholar
2025

Continuous 3D Perception Model with Persistent State

CVPR 2025poster

We present a unified framework capable of solving a broad range of 3D tasks. Our approach features a stateful recurrent model that continuously updates its state representation with each new observation. Given a stream of images, this evolving state can be used to generate metric-scale pointmaps (pe…

2025

Contrastive Multi-view Subspace Clustering via Tensor Transformers Autoencoder

AAAI 2025technical

Multi-view clustering aims to identify consistent and complementary information across multiple views to partition data into clusters, emerging as a popular unsupervised method for multi-view data analysis. However, existing methods often design view-specific encoders to extract distinct features fr…

Cited by 0SourcePDFScholar
2025

Deep Multi-modal Graph Clustering via Graph Transformer Network

AAAI 2025technical

Current deep multi-modal graph clustering methods primarily rely on Graph Neural Network (GNN) to fully exploit attribute features and graph structures, including message propagation and low-dimensional feature embedding. However, these methods lack further exploration of graph structural informatio…

Cited by 0SourcePDFScholar
2025

DreamLight: Towards Harmonious and Consistent Image Relighting

NeurIPS 2025poster

We introduce a model named DreamLight for universal image relighting in this work, which can seamlessly composite subjects into a new background while maintaining aesthetic uniformity in terms of lighting and color tone. The background can be specified by natural images (image-based relighting) or g…

Cited by 0SourceScholar
2025

Efficient Multi-view Clustering via Reinforcement Contrastive Learning

IJCAI 2025

Contrastive multi-view clustering has demonstrated remarkable potential in complex data analysis, yet existing approaches face two critical challenges: difficulty in constructing high-quality positive and negative pairs and high computational overhead due to static optimization strategies. To addres

Cited by 0SourcePDFScholar
2025

Enhanced Unsupervised Discriminant Dimensionality Reduction for Nonlinear Data

IJCAI 2025

Linear Discriminant Analysis (LDA) is a classical supervised dimensionality reduction algorithm. However, LDA focuses more on global structure and overly depends on reliable data labels. For data with outliers and nonlinear structures, LDA cannot effectively capture the true structure of the data. M

Cited by 0SourcePDFScholar
2025

Fair Incomplete Multi-View Clustering via Distribution Alignment

IJCAI 2025

Incomplete multi-view clustering (IMVC) extracts consistent and complementary information from multi-source/modality data with missing views, aiming to partition the data into different clusters. It can effectively address the problem of unsupervised multi-source data analysis in complex environment

Cited by 0SourcePDFScholar
2025

Haptic Feedback Control Strategy for Microswarm Navigation in Flowing Environments

IROS 2025

Swarming microrobots offer great promise for targeted delivery in biofluidic environments. However, current approaches insufficiently utilize the operator’s perceptual awareness and interactive decision-making capabilities. This work proposes a real-time navigation and control strategy with haptic f

Cited by 0SourceScholar
2025

Hypergraph Clustering Network with Partial Attribute Imputation

ICCV 2025poster

Existing hypergraph clustering methods typically assume that node attributes are fully available. However, in real-world scenarios, missing node attributes are common for the sake of privacy or due to data noise. While some approaches attempt to handle missing attributes in traditional graphs, they…

Cited by 0SourcePDFScholar
2025

Long-Distance Delivery of Collective Cell Microrobots Driven by Mobile Magnetic Actuation System

IROS 2025

Collective microrobots enable controlled batch delivery, showing promising application in the biomedical field. However, significant challenges remain in achieving long-distance delivery of collective microrobots in dynamic environments. This study proposes a magnetic actuation strategy for deliveri

Cited by 0SourceScholar
2025

MegaSaM: Accurate, Fast and Robust Structure and Motion from Casual Dynamic Videos

CVPR 2025award

We present a system that allows for accurate, fast, and robust estimation of camera parameters and depth maps from casual monocular videos of dynamic scenes. Most conventional structure from motion and monocular SLAM techniques assume input videos that feature predominantly static scenes with large…

Cited by 18SourcePDFScholar
2025

Reinforcement Learning-Based Microrobotic Swarm Navigation and Obstacle Avoidance in Partially Observable Environments

IROS 2025

Microrobotic swarms have shown promising features due to their collective and flexible behaviours, while achieving precise swarm control and autonomous navigation in complex environments remains a challenge. Here, we propose a Transformer-based reinforcement learning strategy that integrates Proxima

Cited by 0SourceScholar
2025

Scalable Federated One-Step Multi-View Clustering with Tensorized Regularization

AAAI 2025technical

Multi-view clustering (MVC) methods have garnered considerable attention within centralized data frameworks. However, real-world multi-view data are often collected and stored by different organizations, complicating the practical deployment of MVC and motivating the emergence of federated multi-vie…

Cited by 0SourcePDFScholar
2025

Segment Any Motion in Videos

CVPR 2025poster

Moving object segmentation is a crucial task for achieving a high-level understanding of visual scenes and has numerous downstream applications. Humans can effortlessly segment moving objects in videos. Previous work has largely relied on optical flow to provide motion cues; however, this approach o…

2025

Selective Motion Control of Cell Microrobots in Three-Dimensional Space

IROS 2025

Magnetic microrobots are showing great potential in micromanipulation due to the capability of motion control under external fields. However, achieving selective control of magnetic microrobots in three-dimensional (3D) space using global magnetic fields still presents a challenge. In this work, we

Cited by 0SourceScholar
2025

Shape of Motion: 4D Reconstruction from a Single Video

ICCV 2025poster

Monocular dynamic reconstruction is a challenging and long-standing vision problem due to the highly ill-posed nature of the task. Existing approaches depend on templates, are effective only in quasi-static scenes, or fail to model 3D motion explicitly. We introduce a method for reconstructing gener…

Cited by 0SourcePDFScholar
2025

St4RTrack: Simultaneous 4D Reconstruction and Tracking in the World

ICCV 2025poster

Dynamic 3D reconstruction and point tracking in videos are typically treated as separate tasks, despite their deep connection. We propose St4RTrack, a feed-forward frame- work that simultaneously reconstructs and tracks dynamic video content in a world coordinate frame from RGB in- puts. This is ach…

Cited by 0SourcePDFScholar
2025

Unified K-Means Clustering with Label-Guided Manifold Learning

ICML 2025poster

K-Means clustering is a classical and effective unsupervised learning method attributed to its simplicity and efficiency. However, it faces notable challenges, including sensitivity to random initial centroid selection, a limited ability to discover the intrinsic manifold structures within nonlinear…

Cited by 0SourcePDFScholar
2024

"ByteEdit: Boost, Comply and Accelerate Generative Image Editing"

ECCV 2024poster

"Recent advancements in diffusion-based generative image editing have sparked a profound revolution, reshaping the landscape of image outpainting and inpainting tasks. Despite these strides, the field grapples with inherent challenges, including: i) inferior quality; ii) poor consistency; iii) insuf…

Cited by 6SourcePDFScholar
2024

Efficient Federated Multi-View Clustering with Integrated Matrix Factorization and K-Means

IJCAI 2024poster

Multi-view clustering is a popular unsupervised multi-view learning method. Real-world multi-view data are often distributed across multiple entities, presenting a challenge for performing multi-view clustering. Federated learning provides a solution by enabling multiple entities to collaboratively…

Cited by 1SourcePDFScholar
2024

Embedded Feature Selection on Graph-Based Multi-View Clustering

AAAI 2024technical

Recently, anchor graph-based multi-view clustering has been proven to be highly efficient for large-scale data processing. However, most existing anchor graph-based clustering methods necessitate post-processing to obtain clustering labels and are unable to effectively utilize the information within…

Cited by 5SourcePDFScholar
2024

Federated Multi-View Clustering via Tensor Factorization

IJCAI 2024poster

Multi-view clustering is an effective method to process massive unlabeled multi-view data. Since data of different views may be collected and held by different parties, it becomes impractical to train a multi-view clustering model in a centralized way, for the sake of privacy. However, federated mul…

Cited by 1SourcePDFScholar
2024

Model-Free Control of Magnetic Microrobotic Swarm for On-Demand Pattern Spreading

RA-L 2024

Introducing collective control strategies to microrobots provides a promising tool for enhancing the capability of individual microrobots and promoting the manipulation performance. However, controlling numerous micro-agents brings challenges to the precise control and deployment. Herein, an on-dema

Cited by 8SourceScholar
2024

Partial Multi-View Clustering via Self-Supervised Network

AAAI 2024technical

Partial multi-view clustering is a challenging and practical research problem for data analysis in real-world applications, due to the potential data missing issue in different views. However, most existing methods have not fully explored the correlation information among various incomplete views. I…

Cited by 7SourcePDFScholar
2024

Reconstruction Weighting Principal Component Analysis with Fusion Contrastive Learning

IJCAI 2024poster

Principal component analysis (PCA) is a popular unsupervised dimensionality reduction method to extract the principal components of data. However, there are two problems with the existing PCA: (1) Traditional PCA methods treat each sample equally and ignore sample differences. (2) They fail to extra…

2024

Robot See Robot Do: Imitating Articulated Object Manipulation with Monocular 4D Reconstruction

CoRL 2024poster

Humans can learn to manipulate new objects by simply watching others; providing robots with the ability to learn from such demonstrations would enable a natural interface specifying new behaviors. This work develops Robot See Robot Do (RSRD), a method for imitating articulated object manipulation fr…

Cited by 15SourcecodeScholar
2024

SpatialTracker: Tracking Any 2D Pixels in 3D Space

CVPR 2024highlight

Recovering dense and long-range pixel motion in videos is a challenging problem. Part of the difficulty arises from the 3D-to-2D projection process leading to occlusions and discontinuities in the 2D motion domain. While 2D motion can be intricate we posit that the underlying 3D motion can often be…

2023

Centerless Multi-View K-means Based on the Adjacency Matrix

AAAI 2023technical

Although K-Means clustering has been widely studied due to its simplicity, these methods still have the following fatal drawbacks. Firstly, they need to initialize the cluster centers, which causes unstable clustering performance. Secondly, they have poor performance on non-Gaussian datasets. Inspir…

2023

Doppelgangers: Learning to Disambiguate Images of Similar Structures

ICCV 2023oral

We consider the visual disambiguation task of determining whether a pair of visually similar images depict the same or distinct 3D surfaces (e.g., the same or opposite sides of a symmetric building). Illusory image matches, where two images observe distinct but visually similar 3D surfaces, can be c…

Cited by 37PDFcodeScholar
2023

Emergent Correspondence from Image Diffusion

NeurIPS 2023poster

Finding correspondences between images is a fundamental problem in computer vision. In this paper, we show that correspondence emerges in image diffusion models without any explicit supervision. We propose a simple strategy to extract this implicit knowledge out of diffusion networks as image featur…

2023

Neural Scene Chronology

CVPR 2023poster

In this work, we aim to reconstruct a time-varying 3D model, capable of rendering photo-realistic renderings with independent control of viewpoint, illumination, and time, from Internet photos of large-scale landmarks. The core challenges are twofold. First, different types of temporal changes, such…

2023

Orthogonal Non-negative Tensor Factorization based Multi-view Clustering

NeurIPS 2023poster

Multi-view clustering (MVC) based on non-negative matrix factorization (NMF) and its variants have attracted much attention due to their advantages in clustering interpretability. However, existing NMF-based multi-view clustering methods perform NMF on each view respectively and ignore the impact of…

Cited by 32SourcePDFScholar
2023

Tracking Everything Everywhere All at Once

ICCV 2023oral

We present a new test-time optimization method for estimating dense and long-range motion from a video sequence. Prior optical flow or particle video tracking algorithms typically operate within limited temporal windows, struggling to track through occlusions and maintain global consistency of estim…

Cited by 170PDFcodeScholar
2022

3D Moments From Near-Duplicate Photos

CVPR 2022poster

We introduce 3D Moments, a new computational photography effect. As input we take a pair of near-duplicate photos, i.e., photos of moving subjects from similar viewpoints, common in people's photo collections. As output, we produce a video that smoothly interpolates the scene motion from the first p…

Cited by 18PDFcodeScholar
2022

InfiniteNature-Zero: Learning Perpetual View Generation of Natural Scenes from Single Images

ECCV 2022poster

"We present a method for learning to generate unbounded flythrough videos of natural scenes starting from a single view. This capability is learned from a collection of single photographs, without requiring camera poses or even multiple views of each scene. To achieve this, we propose a novel self-s…

2022

Neural 3D Scene Reconstruction With the Manhattan-World Assumption

CVPR 2022oral

This paper addresses the challenge of reconstructing 3D indoor scenes from multi-view images. Many previous works have shown impressive reconstruction results on textured objects, but they still have difficulty in handling low-textured planar regions, which are common in indoor scenes. An approach t…

Cited by 188PDFcodeScholar
2022

Neural Rays for Occlusion-Aware Image-Based Rendering

CVPR 2022poster

We present a new neural representation, called Neural Ray (NeuRay), for the novel view synthesis task. Recent works construct radiance fields from image features of input views to render novel view images, which enables the generalization to new scenes. However, due to occlusions, a 3D point may be…

Cited by 234PDFcodeScholar
2022

Real-Time Navigation of an Untethered Miniature Robot Using Mobile Ultrasound Imaging and Magnetic Actuation Systems

RA-L 2022

Remotely actuated small-scale robots have shown promising capabilities in targeted delivery, micromanipulation, and biosensing. However, challenges remain in imaging and control of a robot in a large workspace, especially when conducting targeted navigation in hard-to-reach and tortuous regions of a

Cited by 21SourceScholar
2021

Animatable Neural Radiance Fields for Modeling Dynamic Human Bodies

ICCV 2021poster

This paper addresses the challenge of reconstructing an animatable human model from a multi-view video. Some recent works have proposed to decompose a non-rigidly deforming scene into a canonical neural radiance field and a set of deformation fields that map observation-space points to the canonical…

Cited by 514PDFcodeScholar
2021

IBRNet: Learning Multi-View Image-Based Rendering

CVPR 2021poster

We present a method that synthesizes novel views of complex scenes by interpolating a sparse set of nearby views. The core of our method is a network architecture that includes a multilayer perceptron and a ray transformer that estimates radiance and volume density at continuous 5D locations (3D spa…

Cited by 956PDFScholar
2021

Magnetic Control of a Steerable Guidewire Under Ultrasound Guidance Using Mobile Electromagnets

RA-L 2021

Endovascular surgery has become a popular minimally invasive approach to diagnose and treat various vascular diseases. However, manipulating conventional passive guidewires and catheters still has technical challenges, such as long duration and undesired trauma. In addition, radiation exposure induc

Cited by 61SourceScholar
2021

Neural Body: Implicit Neural Representations With Structured Latent Codes for Novel View Synthesis of Dynamic Humans

CVPR 2021poster

This paper addresses the challenge of novel view synthesis for a human performer from a very sparse set of camera views. Some recent works have shown that learning implicit neural representations of 3D scenes achieves remarkable view synthesis quality given dense input views. However, the representa…

Cited by 862PDFcodeScholar
2021

Parallel Actuation of Nanorod Swarm and Nanoparticle Swarm to Different Targets

ICRA 2021poster

After years of development, various swarms of robots have been proposed for many complicated tasks, such as forming patterns, cooperative locomotion, and adapting to different environments. However, controlling microrobotic swarms is still a challenging task owing to the lacking of integrated device…

Cited by 1SourceScholar
2021

PhySG: Inverse Rendering With Spherical Gaussians for Physics-Based Material Editing and Relighting

CVPR 2021poster

We present an end-to-end inverse rendering pipeline that includes a fully differentiable renderer, and can reconstruct geometry, materials, and illumination from scratch from a set of images. Our rendering framework represents specular BRDFs and environmental illumination using mixtures of spherical…

Cited by 365PDFScholar
2021

Ultrasound Doppler Imaging and Navigation of Collective Magnetic Cell Microrobots in Blood

ICRA 2021poster

We propose ultrasound Doppler imaging and magnetic navigation of collective cell microrobots in whole blood. Cell microrobots are cultured using stem cells and iron microparticles, they have spheroidal structures and can be actuated under external magnetic fields. A collective of cell microrobots ca…

Cited by 3SourceScholar
2020

Hidden Footprints: Learning Contextual Walkability from 3D Human Trails

ECCV 2020poster

Predicting where people can walk in a scene is important for many tasks, including autonomous driving systems and human behavior analysis. Yet learning a computational model for this purpose is challenging due to semantic ambiguity and a lack of labeled data: current datasets only have labels on whe…

2020

Learning Feature Descriptors using Camera Pose Supervision

ECCV 2020poster

Recent research on learned visual descriptors has shown promising improvements in correspondence estimation, a key component of many 3D vision tasks. However, existing descriptor learning frameworks typically require ground-truth correspondences between feature points for training, which are challen…

Cited by 201SourcePDFScholar
2020

Multi-View Attribute Graph Convolution Networks for Clustering

IJCAI 2020poster

Graph neural networks (GNNs) have made considerable achievements in processing graph-structured data. However, existing methods can not allocate learnable weights to different nodes in the neighborhood and lack of robustness on account of neglecting both node attributes and graph reconstruction. Mor…

Cited by 0SourcePDFScholar
2020

Reconfigurable Magnetic Microswarm for Thrombolysis under Ultrasound Imaging

ICRA 2020poster

We propose thrombolysis using a magnetic nanoparticle microswarm with tissue plasminogen activator (tPA) under ultrasound imaging. The microswarm is generated in blood using an oscillating magnetic field and can be navigated with locomotion along both the long and short axis. By modulating the input…

Cited by 31SourceScholar
2019

Characterizing Nanoparticle Swarms With Tuneable Concentrations for Enhanced Imaging Contrast

RA-L 2019

Microrobots capable of performing targeted delivery are promising for biomedical applications. Due to the restriction of their small size and volume, however, the real-time <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">in vivo</i> imaging strategy

Cited by 28SourceScholar
2019

Magnetic-Needle-Assisted Micromanipulation of Dynamically Self-Assembled Magnetic Droplets for Cargo Transportation

IROS 2019poster

Dynamic self-assembly is treated as a promising approach for generating a robotic swarm to perform coordinated tasks, and the assembled pattern can be tuned by regulating the energy input. However, location of a dynamically assembled pattern is hard to be determined, especially under global fields,…

Cited by 3SourceScholar
2018

Magnetic Navigation of a Rotating Colloidal Swarm Using Ultrasound Images

IROS 2018poster

Microrobots are considered as promising tools for biomedical applications. However, the imaging of them becomes challenges in order to be further applied on in vivo environments. Here we report the magnetic navigation of a paramagnetic nanoparticle-based swarm using ultrasound images. The swarm can…

Cited by 32SourceScholar
2017

A Miniature Flexible-Link Magnetic Swimming Robot With Two Vibration Modes: Design, Modeling and Characterization

RA-L 2017

In this letter, we report a miniature swimming robot that consists of a nonmagnetic V-shaped head, an I-shaped tail attached with a magnet and a flexible link for interconnection of the rigid head and tail. The use of a flexible link enables the robot to exhibit two different vibration modes by tuni

Cited by 28SourceScholar