← Search

Tao Guan

16 accepted papers

2026

MIDAS: Multi-Image Dispersion and Semantic Reconstruction for Jailbreaking MLLMs

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have achieved remarkable performance but remain vulnerable to jailbreak attacks that can induce harmful content and undermine their secure deployment. Previous studies have shown that introducing additional inference steps, which disrupt security attention, c…

Cited by 0SourcecodeScholar
2026

PATexGS: Perceptual-Adaptive Texture Scheduling for Visual Coherence in Textured Gaussian Splatting

AAAI 2026technical

3D Gaussian Splatting (3DGS) has emerged as a mainstream solution for real-time rendering and high-fidelity novel view synthesis. Building on this foundation, methods based on Textured Gaussians further improve the expression ability by incorporating explicit texture mapping into Gaussians. However,

Cited by 0SourcePDFScholar
2025

Frequency-Aware Density Control via Reparameterization for High-Quality Rendering of 3D Gaussian Splatting

AAAI 2025technical

By adaptively controlling the density and generating more Gaussians in regions with high-frequency information, 3D Gaussian Splatting (3DGS) can better represent scene details. From the signal processing perspective, representing details usually needs more Gaussians with relatively smaller scales. H…

2025

IndoorGS: Geometric Cues Guided Gaussian Splatting for Indoor Scene Reconstruction

CVPR 2025poster

3D Gaussian Splatting (3DGS) has shown impressive performance in scene reconstruction, offering high rendering quality and rapid rendering speed with short training time. However, it often yields unsatisfactory results when applied to indoor scenes due to its poor ability to learn geometries without…

Cited by 0SourcePDFScholar
2025

Instant GaussianImage: A Generalizable and Self-Adaptive Image Representation via 2D Gaussian Splatting

ICCV 2025poster

Implicit Neural Representation (INR) has demonstrated remarkable advances in the field of image representation but demands substantial GPU resources. GaussianImage recently pioneered the use of Gaussian Splatting to mitigate this cost, however, the slow training process limits its practicality, and…

2024

ReinforceNS: Reinforcement Learning-based Multi-start Neighborhood Search for Solving the Traveling Thief Problem

IJCAI 2024poster

The Traveling Thief Problem (TTP) is a challenging combinatorial optimization problem with broad practical applications. TTP combines two NP-hard problems: the Traveling Salesman Problem (TSP) and Knapsack Problem (KP). While a number of machine learning and deep learning based algorithms have been…

Cited by 0SourcePDFScholar
2023

Adaptive Patch Deformation for Textureless-Resilient Multi-View Stereo

CVPR 2023poster

In recent years, deep learning-based approaches have shown great strength in multi-view stereo because of their outstanding ability to extract robust visual features. However, most learning-based methods need to build the cost volume and increase the receptive field enormously to get a satisfactory…

2023

C2F2NeUS: Cascade Cost Frustum Fusion for High Fidelity and Generalizable Neural Surface Reconstruction

ICCV 2023poster

There is an emerging effort to combine the two popular 3D frameworks using Multi-View Stereo (MVS) and Neural Implicit Surfaces (NIS) with a specific focus on the few-shot / sparse view setting. In this paper, we introduce a novel integration scheme that combines the multi-view stereo with neural si…

Cited by 10PDFScholar
2020

Adversarial Style Mining for One-Shot Unsupervised Domain Adaptation

NeurIPS 2020poster

We aim at the problem named One-Shot Unsupervised Domain Adaptation. Unlike traditional Unsupervised Domain Adaptation, it assumes that only one unlabeled target sample can be available when learning to adapt. This setting is realistic but more challenging, in which conventional adaptation approache…

2020

Mesh-Guided Multi-View Stereo With Pyramid Architecture

CVPR 2020poster

Multi-view stereo (MVS) aims to reconstruct 3D geometry of the target scene by using only information from 2D images. Although much progress has been made, it still suffers from textureless regions. To overcome this difficulty, we propose a mesh-guided MVS method with pyramid architecture, which mak…

Cited by 38PDFcodeScholar
2019

P-MVSNet: Learning Patch-Wise Matching Confidence Aggregation for Multi-View Stereo

ICCV 2019poster

Learning-based methods are demonstrating their strong competitiveness in estimating depth for multi-view stereo reconstruction in recent years. Among them the approaches that generate cost volumes based on the plane-sweeping algorithm and then use them for feature matching have shown to be very prom…

Cited by 259PDFScholar
2019

Significance-Aware Information Bottleneck for Domain Adaptive Semantic Segmentation

ICCV 2019poster

For unsupervised domain adaptation problems, the strategy of aligning the two domains in latent feature space through adversarial learning has achieved much progress in image classification, but usually fails in semantic segmentation tasks in which the latent representations are overcomplex. In this…

Cited by 262PDFScholar
2019

Taking a Closer Look at Domain Shift: Category-Level Adversaries for Semantics Consistent Domain Adaptation

CVPR 2019oral

We consider the problem of unsupervised domain adaptation in semantic segmentation. The key in this campaign consists in reducing the domain shift, i.e., enforcing the data distributions of the two domains to be similar. A popular strategy is to align the marginal distribution in the feature space t…

Cited by 933PDFcodeScholar
2018

Macro-Micro Adversarial Network for Human Parsing

ECCV 2018poster

In human parsing, the pixel-wise classification loss has drawbacks in its low-level local inconsistency and high-level semantic inconsistency. The introduction of the adversarial network tackles the two problems using a single discriminator. However, the two types of parsing inconsistency are genera…

2015

Interactive on-device Mobile Landmark Recognition with compact binary codes

ICASSP 2015accepted

Interactive mobile vision applications, such as Mobile Landmark Recognition (MLR), have recently attracted ever increasing research attention due to the exponential growth of mobile devices. However, the recognition accuracy retains as a bottleneck hesitating the proliferation of such applications.…

Cited by 0SourceScholar