← Search

Long Sun

10 accepted papers

2026

DSGCR: Decomposed Spectral Geometry-Aware Cross-Modal Semantic Representation for 3D Visual Grounding

ICML 2026poster

3D visual grounding encompassing 3D referring expression comprehension (3DREC) and segmentation (3DRES) requires robust cross-modal representation to achieve fine-grained semantic alignment and precise geometric reasoning. However, most methods employ unimodal pre-trained encoders that transfer visu…

Cited by 0SourceScholar
2026

PortraitSR: Artist-Inspired Prior Learning for Progressive Face Super-Resolution

AAAI 2026technical

Face super-resolution (FSR) aims to reconstruct high-resolution (HR) face images from low-resolution (LR) inputs. While recent methods have advanced this task through architectural innovations and generative modeling, but they often leads to semantically inconsistent structures and unrealistic textu

Cited by 0SourcePDFScholar
2026

RECS4R: Bridging Semantics and Geometry for Referring Remote Sensing Interpretation

CVPR 2026

Referring expression comprehension and segmentation (RECS) task plays a vital role in remote sensing due to its high efficiency in multi-tasking. However, RECS has reached a performance bottleneck rooted in representational insufficiency, primarily due to cross-task representational fragmentation in

Cited by 0SourcecodeScholar
2026

SOAR: Semi-Supervised Open-Vocabulary Aerial Object Detection via Dual-Aware Enhanced Prior Denoising

AAAI 2026technical

Open-Vocabulary Object Detection (OVOD) shows promise in remote sensing (RS), but due to its unique value, there are challenges such as the predominance of background regions, sparse labels, limited semantic information, and difficulties in semi-supervised training. To tackle these challenges, we pr

Cited by 0SourcePDFScholar
2026

STCDiT: Spatio-Temporally Consistent Diffusion Transformer for High-Quality Video Super-Resolution

CVPR 2026

We present STCDiT, a video super-resolution framework built upon a pre-trained video diffusion model, aiming to restore structurally faithful and temporally stable videos from degraded inputs, even under complex camera motions. The main challenges lie in maintaining temporal stability during reconst

Cited by 0SourceScholar
2025

Efficient Video Super-Resolution for Real-time Rendering with Decoupled G-buffer Guidance

CVPR 2025poster

Latency is a key driver for real-time rendering applications, making super-resolution techniques increasingly popular to accelerate rendering processes. In contrast to existing methods that directly concatenate low-resolution frames and G-buffers as input without discrimination, we develop an asymme…

2025

TCTformer: Long-term forecasting with dual attention transformers

ICASSP 2025accepted

In the field of multi-variable long-term time series (MLTS) prediction, many deep learning models have been developed, and Transformer-based models have received widespread attention for their ability to capture the complex interactions between sequences. However, as the sequence length increases an…

Cited by 0SourceScholar
2024

SMFANet: A Lightweight Self-Modulation Feature Aggregation Network for Efficient Image Super-Resolution

ECCV 2024poster

"Transformer-based restoration methods achieve significant performance as the self-attention (SA) of the Transformer can explore non-local information for better high-resolution image reconstruction. However, the key dot-product SA requires substantial computational resources, which limits its appli…

2023

Spatially-Adaptive Feature Modulation for Efficient Image Super-Resolution

ICCV 2023poster

Although deep learning-based solutions have achieved impressive reconstruction performance in image super-resolution (SR), these models are generally large, with complex architectures, making them incompatible with low-power devices with many computational and memory constraints. To overcome these c…

Cited by 142PDFcodeScholar