← Search

Heng Guo

25 accepted papers

2026

MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement Learning

ICLR 2026poster

Medical Vision-Language Models (VLMs) hold immense promise for complex clinical tasks, but their reasoning capabilities are often constrained by text-only paradigms that fail to ground inferences in visual evidence. This limitation not only curtails performance on tasks requiring fine-grained visual…

Cited by 0SourcecodeScholar
2026

Photon: Speedup Volume Understanding with Efficient Multimodal Large Language Models

ICLR 2026poster

Multimodal large language models are promising for clinical visual question answering tasks, but scaling to 3D imaging is hindered by high computational costs. Prior methods often rely on 2D slices or fixed-length token compression, disrupting volumetric continuity and obscuring subtle findings. We…

Cited by 0SourcecodeScholar
2026

Seeing Through the Rain: Resolving High-Frequency Conflicts in Deraining and Super-Resolution via Diffusion Guidance

AAAI 2026technical

Clean images are crucial for visual tasks such as small object detection, especially at high resolutions. However, real-world images are often degraded by adverse weather, and weather restoration methods may sacrifice high-frequency details critical for analyzing small objects. A natural solution is

Cited by 0SourcePDFScholar
2026

TOWARDS PRIVACY-PRESERVING FINE-GRAINED VISUAL CLASSIFICATION VIA HIERARCHICAL LEARNING FROM LABEL PROPORTIONS

ICASSP 2026poster

In recent years, Fine-Grained Visual Classification (FGVC) has achieved impressive recognition accuracy, despite minimal inter-class variations. However, existing methods heavily rely on instance-level labels, making them impractical in privacy-sensitive scenarios such as medical image analysis. Thi…

Cited by 0SourcePDFScholar
2025

MFogHub: Bridging Multi-Regional and Multi-Satellite Data for Global Marine Fog Detection and Forecasting

CVPR 2025poster

Deep learning approaches for marine fog detection and forecasting have outperformed traditional methods, demonstrating significant scientific and practical importance. However, the limited availability of open-source datasets remains a major challenge. Existing datasets, often focused on a single re…

2025

PIDSR: Complementary Polarized Image Demosaicing and Super-Resolution

CVPR 2025poster

Polarization cameras can capture multiple polarized images with different polarizer angles in a single shot, bringing convenience to polarization-based downstream tasks. However, their direct outputs are color-polarization filter array (CPFA) raw images, requiring demosaicing to reconstruct full-res…

2025

PMNI: Pose-free Multi-view Normal Integration for Reflective and Textureless Surface Reconstruction

CVPR 2025poster

Reflective and textureless surfaces remain a challenge in multi-view 3D reconstruction. Both camera pose calibration and shape reconstruction often fail due to insufficient or unreliable cross-view visual features. To address these issues, we present PMNI (Pose-free Multi-view Normal Integration), a…

2025

PolGS: Polarimetric Gaussian Splatting for Fast Reflective Surface Reconstruction

ICCV 2025poster

Efficient shape reconstruction for surfaces with complex reflectance properties is crucial for real-time virtual reality. While 3D Gaussian Splatting (3DGS)-based methods offer fast novel view rendering by leveraging their explicit surface representation, their reconstruction quality lags behind tha…

Cited by 0SourcePDFScholar
2025

PolarAnything: Diffusion-based Polarimetric Image Synthesis

ICCV 2025poster

Polarization images facilitate image enhancement and 3D reconstruction tasks, but the limited accessibility of polarization cameras hinders their broader application. This gap drives the need for synthesizing photorealistic polarization images. The existing polarization simulator Mitsuba relies on a…

Cited by 0SourcePDFScholar
2025

Polarimetric Neural Field via Unified Complex-Valued Wave Representation

ICCV 2025poster

Polarization has found applications in various computer vision tasks by providing additional physical cues. However, due to the limitations of current imaging systems, polarimetric parameters are typically stored in discrete form, which is non-differentiable and limits their applicability in polariz…

Cited by 0SourcePDFScholar
2025

Self-Supervised Selective-Guided Diffusion Model for Old-Photo Face Restoration

NeurIPS 2025poster

Old-photo face restoration poses significant challenges due to compounded degradations such as breakage, fading, and severe blur. Existing pre-trained diffusion-guided methods either rely on explicit degradation priors or global statistical guidance, which struggle with localized artifacts or face c…

Cited by 0SourcecodeScholar
2025

Towards a Comprehensive, Efficient and Promptable Anatomic Structure Segmentation Model Using 3D Whole-Body CT Scans

AAAI 2025technical

Segment anything model (SAM) demonstrates strong generalization ability on natural image segmentation. However, its direct adaptation in medical image segmentation tasks shows significant performance drops. It also requires an excessive number of prompt points to obtain a reasonable accuracy. Althou…

2024

CycleINR: Cycle Implicit Neural Representation for Arbitrary-Scale Volumetric Super-Resolution of Medical Data

CVPR 2024poster

In the realm of medical 3D data such as CT and MRI images prevalent anisotropic resolution is characterized by high intra-slice but diminished inter-slice resolution. The lowered resolution between adjacent slices poses challenges hindering optimal viewing experiences and impeding the development of…

Cited by 2SourcePDFScholar
2024

DiLiGenRT: A Photometric Stereo Dataset with Quantified Roughness and Translucency

CVPR 2024poster

Photometric stereo faces challenges from non-Lambertian reflectance in real-world scenarios. Systematically measuring the reliability of photometric stereo methods in handling such complex reflectance necessitates a real-world dataset with quantitatively controlled reflectances. This paper introduce…

2024

NeRSP: Neural 3D Reconstruction for Reflective Objects with Sparse Polarized Images

CVPR 2024poster

We present NeRSP a Neural 3D reconstruction technique for Reflective surfaces with Sparse Polarized images. Reflective surface reconstruction is extremely challenging as specular reflections are view-dependent and thus violate the multiview consistency for multiview stereo. On the other hand sparse…

Cited by 7SourcePDFScholar
2024

Representing Topological Self-Similarity Using Fractal Feature Maps for Accurate Segmentation of Tubular Structures

ECCV 2024poster

"Accurate segmentation of long and thin tubular structures is required in a wide variety of areas such as biology, medicine, and remote sensing. The complex topology and geometry of such structures often pose significant technical challenges. A fundamental property of such structures is their topolo…

2024

SfPUEL: Shape from Polarization under Unknown Environment Light

NeurIPS 2024poster

Shape from polarization (SfP) benefits from advancements like polarization cameras for single-shot normal estimation, but its performance heavily relies on light conditions. This paper proposes SfPUEL, an end-to-end SfP method to jointly estimate surface normal and material under unknown environment…

2023

Anatomical Invariance Modeling and Semantic Alignment for Self-supervised Learning in 3D Medical Image Analysis

ICCV 2023oral

Self-supervised learning (SSL) has recently achieved promising performance for 3D medical image analysis tasks. Most current methods follow existing SSL paradigm originally designed for photographic or natural images, which cannot explicitly and thoroughly exploit the intrinsic similar anatomical st…

Cited by 29PDFcodeScholar
2023

DiLiGenT-Pi: Photometric Stereo for Planar Surfaces with Rich Details - Benchmark Dataset and Beyond

ICCV 2023poster

Photometric stereo aims to recover detailed surface shapes from images captured under varying illuminations. However, existing real-world datasets primarily focus on evaluating photometric stereo for general non-Lambertian reflectances and feature bulgy shapes that have a certain height. As shape de…

Cited by 12PDFcodeScholar
2023

Non-Lambertian Multispectral Photometric Stereo via Spectral Reflectance Decomposition

IJCAI 2023poster

Multispectral photometric stereo (MPS) aims at recovering the surface normal of a scene from a single-shot multispectral image captured under multispectral illuminations. Existing MPS methods adopt the Lambertian reflectance model to make the problem tractable, but it greatly limits their applicatio…

Cited by 9SourcePDFScholar
2023

ReLeaPS : Reinforcement Learning-based Illumination Planning for Generalized Photometric Stereo

ICCV 2023poster

Illumination planning in photometric stereo aims to find a balance between tween surface normal estimation accuracy and image capturing efficiency by selecting optimal light configurations. It depends on factors such as the unknown shape and general reflectance of the target object, global illuminat…

Cited by 2PDFScholar
2022

Improving Certified Robustness via Statistical Learning with Logical Reasoning

NeurIPS 2022accept

Intensive algorithmic efforts have been made to enable the rapid improvements of certificated robustness for complex ML models recently. However, current robustness certification methods are only able to certify under a limited perturbation radius. Given that existing pure data-driven statistical ap…

2021

Multispectral Photometric Stereo for Spatially-Varying Spectral Reflectances: A Well Posed Problem?

CVPR 2021poster

Multispectral photometric stereo (MPS) aims at recovering the surface normal of a scene from a single-shot multispectral image, which is known as an ill-posed problem. To make the problem well-posed, existing MPS methods rely on restrictive assumptions, such as shape prior, surfaces having a monochr…

Cited by 14PDFcodeScholar
2019

X2CT-GAN: Reconstructing CT From Biplanar X-Rays With Generative Adversarial Networks

CVPR 2019poster

Computed tomography (CT) can provide a 3D view of the patient's internal organs, facilitating disease diagnosis, but it incurs more radiation dose to a patient and a CT scanner is much more cost prohibitive than an X-ray machine too. Traditional CT reconstruction methods require hundreds of X-ray pr…

Cited by 297PDFcodeScholar