← Search

Song Tang

20 accepted papers

2026

Consistent Text-to-Image Generation via Scene De-Contextualization

ICLR 2026poster

Consistent text-to-image (T2I) generation seeks to produce identity-preserving images of the same subject across diverse scenes, yet it often fails due to a phenomenon called identity (ID) shift. Previous methods have tackled this issue, but typically rely on the unrealistic assumption of knowing al…

Cited by 0SourcecodeScholar
2026

Graph Domain Adaptation via Homophily-Agnostic Reconstructing Structure

AAAI 2026technical

Graph Domain Adaptation (GDA) transfers knowledge from labeled source graphs to unlabeled target graphs, addressing the challenge of label scarcity. However, existing GDA methods typically assume that both source and target graphs exhibit homophily, leading existing methods to perform poorly when he

Cited by 0SourcePDFScholar
2026

PSDesigner: Automated Graphic Design with a Human-Like Creative Workflow

CVPR 2026

Graphic design is a creative and innovative process that plays a crucial role in applications such as e-commerce and advertising. However, developing an automated design system that can faithfully translate user intentions into editable design files remains an open challenge. Although recent studies

Cited by 0SourcecodeScholar
2026

SAGA: Structural Aggregation Guided Alignment with Dynamic View and Neighborhood Order Selection for Multiview Graph Domain Adaptation

ICLR 2026poster

Graph domain adaptation (GDA) transfers knowledge from a labeled source graph to an unlabeled target graph to alleviate label scarcity. In multi-view graphs, the challenge of mitigating domain shift is constrained by structural information across various views. Moreover, within each view, structures…

Cited by 0SourcecodeScholar
2026

VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing

ICML 2026poster

Despite the rapid advancements in Vision-Language Models (VLMs), a critical gap remains in their ability to handle structured, controllable diagrammatic tasks essential for professional workflows, as existing methods predominantly rely on pixel-based synthesis which operates in probabilistic pixel s…

Cited by 0SourceScholar
2025

Domain Adaptive Diabetic Retinopathy Grading with Model Absence and Flowing Data

CVPR 2025poster

Domain shift (the difference between source and target domains) poses a significant challenge in clinical applications, e.g., Diabetic Retinopathy (DR) grading. Despite considering certain clinical requirements, like source data privacy, conventional transfer methods are predominantly model-centered…

2025

Multimodal Causal Reasoning for UAV Object Detection

NeurIPS 2025poster

Unmanned Aerial Vehicle (UAV) object detection faces significant challenges due to complex environmental conditions and different imaging conditions. These factors introduce significant changes in scale and appearance, particularly for small objects that occupy limited pixels and exhibit limited inf…

Cited by 0SourceScholar
2025

Proxy Denoising for Source-Free Domain Adaptation

ICLR 2025oral

Source-Free Domain Adaptation (SFDA) aims to adapt a pre-trained source model to an unlabeled target domain with no access to the source data. Inspired by the success of large Vision-Language (ViL) models in many applications, the latest research has validated ViL's benefit for SFDA by using their p…

2025

Pseudo Visible Feature Fine-Grained Fusion for Thermal Object Detection

CVPR 2025poster

Thermal object detection is a critical task in various fields, such as surveillance and autonomous driving. Current state-of-the-art (SOTA) models always leverage a prior Thermal-To-Visible (T2V) translation model to obtain visible spectrum information, followed by a cross-modality aggregation modul…

2025

SAFormer: Spatially Adaptive Transformer for Efficient and Multi-Resolution Occupancy Prediction

IROS 2025

Accurate and efficient 3D scene understanding from multi-view images remains a fundamental challenge in autonomous driving. Existing methods often struggle with high-dimensional features, leading to excessive computational costs and memory usage. In this paper, we present SAFormer, a novel transform

Cited by 0SourceScholar
2025

Self-Prompting Analogical Reasoning for UAV Object Detection

AAAI 2025technical

Unmanned Aerial Vehicle Object Detection (UAVOD) presents unique challenges due to varying altitudes, dynamic backgrounds, and the small size of objects. Traditional detection methods often struggle with these challenges, as they typically rely on visual feature only and fail to extract the semantic…

Cited by 0SourcePDFScholar
2025

UnrealLLM: Towards Highly Controllable and Interactable 3D Scene Generation by LLM-powered Procedural Content Generation

ACL 2025finding

The creation of high-quality 3D scenes is essential for applications like video games and simulations, yet automating this process while retaining the benefits of Procedural Content Generation (PCG) remains challenging. In this paper, we introduce UnrealLLM, a novel multi-agent framework that connec…

Cited by 0SourcePDFScholar
2024

Cloud Object Detector Adaptation by Integrating Different Source Knowledge

NeurIPS 2024poster

We propose to explore an interesting and promising problem, Cloud Object Detector Adaptation (CODA), where the target domain leverages detections provided by a large cloud model to build a target detector. Despite with powerful generalization capability, the cloud model still cannot achieve error-fr…

Cited by 3SourcePDFScholar
2024

Source-Free Domain Adaptation with Frozen Multimodal Foundation Model

CVPR 2024poster

Source-Free Domain Adaptation (SFDA) aims to adapt a source model for a target domain with only access to unlabeled target training data and the source model pretrained on a supervised source domain. Relying on pseudo labeling and/or auxiliary supervision conventional methods are inevitably error-pr…

2023

SwinLSTM: Improving Spatiotemporal Prediction Accuracy using Swin Transformer and LSTM

ICCV 2023poster

Integrating CNNs and RNNs to capture spatiotemporal dependencies is a prevalent strategy for spatiotemporal prediction tasks. However, the property of CNNs to learn local spatial information decreases their efficiency in capturing spatiotemporal dependencies, thereby limiting their prediction accura…

Cited by 75PDFcodeScholar
2023

Weakly Supervised Referring Expression Grounding via Target-Guided Knowledge Distillation

ICRA 2023poster

Weakly supervised referring expression grounding aims to train a model without the manual labels between image regions and referring expressions during the training phase. Current predominant models often adopt deep structures to reconstruct the region-expression correspondence. A crucial deficiency…

Cited by 4SourcecodeScholar
2021

Model Adaptation through Hypothesis Transfer with Gradual Knowledge Distillation

IROS 2021poster

The ability to adapt their perception to changing environments is a core characterization of intelligent robots. At present, Unsupervised Domain Adaptation (UDA) methods are used to address this problem where the adaptation task is formulated as a transfer problem from a well-described scenario (sou…

Cited by 21SourceScholar
2019

PointNetGPD: Detecting Grasp Configurations from Point Sets

ICRA 2019poster

In this paper, we propose an end-to-end grasp evaluation model to address the challenging problem of localizing robot grasp configurations directly from the point cloud. Compared to recent grasp evaluation metrics that are based on handcrafted depth features and a convolutional neural network (CNN),…

Cited by 444SourcecodeScholar
2019

Visual Domain Adaptation Exploiting Confidence-Samples

IROS 2019poster

Domain adaptation methods are used to address a problem, in which train scenario (source domain) and test scenario (target domain) are different. The existing methods mainly perform adaptation via reducing domain discrepancy from the view of a probability distribution. However, the idea of probabili…

Cited by 6SourceScholar