← Search

Gang Xiong

19 accepted papers

2026

From Scale to Speed: Adaptive Test-Time Scaling for Image Editing

CVPR 2026

Image Chain-of-Thought (Image-CoT) is a test-time scaling paradigm that improves image generation by extending inference time. Most Image-CoT methods focus on text-to-image (T2I) generation. Unlike T2I generation, image editing is goal-directed: the solution space is constrained by the source image

Cited by 0SourceScholar
2026

Manipulation Intention Understanding for Zero-Shot Composed Image Retrieval

AAAI 2026technical

Zero-shot Composed Image Retrieval (ZS-CIR) involves diverse tasks with varied visual manipulation intents across domains, scenes, objects, and attributes. A key challenge is that existing datasets contain limited intent-relevant annotations, making it hard for models to infer human intent from text

Cited by 0SourcePDFScholar
2026

SegQuant: A Semantics-Aware and Generalizable Quantization Framework for Diffusion Models

CVPR 2026

Diffusion models have demonstrated exceptional generative capabilities but are computationally intensive, posing significant challenges for deployment in resource-constrained or latency-sensitive environments.Quantization offers an effective means to reduce model size and computational cost, with po

Cited by 0SourcecodeScholar
2026

Threat2Traffic: Multi-Agent Environment Synthesis for Malware Traffic Generation from Threat Intelligence

ICML 2026poster

Data-driven cybersecurity research is fundamentally constrained by the scarcity of labeled datasets, yet acquiring authentic, large-scale malware traffic remains bottlenecked by obsolescent public datasets, unscalable manual construction, and inflexible sandboxes that fail to satisfy the sample-spec…

Cited by 0SourcecodeScholar
2025

ANASETC: Automatic Neural Architecture Search for Encrypted Traffic Classification

ICASSP 2025accepted

The widespread adoption of encrypted network protocols has made traffic encryption ubiquitous, creating substantial challenges for network management and security. This paper introduces a novel encrypted traffic classification system, ANASETC, which combines traffic burst features with Neural Archit…

Cited by 0SourceScholar
2025

Efficient Non-Sequential Relational Modeling for Temporal Knowledge Graph Link Predictions

ICASSP 2025accepted

Temporal Knowledge Graphs (TKGs) are being widely explored to predict the future for they record multi-relational knowledge and the happening time of real-life facts. Existing works learn sequential patterns to infer the future from past facts in TKGs for predictions. Although achieving promising re…

Cited by 0SourceScholar
2025

Improving Embeddings by Refining Meanings for Temporal Knowledge Graph Link Predictions

ICASSP 2025accepted

Temporal Knowledge Graphs (TKGs) represent real-life facts using entities, relational types, and timestamps where relational types state the semantic scenario of facts. Current methods learn embeddings by merging facts of multiple types (e.g. sport and family) for predictions. Such embeddings associ…

Cited by 0SourceScholar
2025

Missing Target-Relevant Information Prediction with World Model for Accurate Zero-Shot Composed Image Retrieval

CVPR 2025poster

Zero-Shot Composed Image Retrieval (ZS-CIR) involves diverse tasks with a broad range of visual content manipulation intent across domain, scene, object, and attribute. The key challenge for ZS-CIR tasks is to modify a reference image according to manipulation text to accurately retrieve a target im…

2025

ProAPO: Progressively Automatic Prompt Optimization for Visual Classification

CVPR 2025poster

Vision-language models (VLMs) have made significant progress in image classification by training with large-scale paired image-text data. Their performances largely depend on the prompt quality. While recent methods show that visual descriptions generated by large language models (LLMs) enhance the…

2025

Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval

CVPR 2025highlight

Composed Image Retrieval (CIR) aims to retrieve target images that closely resemble a reference image while integrating user-specified textual modifications, thereby capturing user intent more accurately. Existing training-free zero-shot CIR (ZS-CIR) methods often employ a two-stage process: they fi…

2025

Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining

ICLR 2025poster

A significant aspiration of offline reinforcement learning (RL) is to develop a generalist agent with high capabilities from large and heterogeneous datasets. However, prior approaches that scale offline RL either rely heavily on expert trajectories or struggle to generalize to diverse unseen tasks.…

2025

Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning

NeurIPS 2025poster

Process reward model (PRM) has been proven effective in test-time scaling of LLM on challenging reasoning tasks. However, the reward hacking induced by PRM hinders its successful applications in reinforcement fine-tuning. We find the primary cause of reward hacking induced by PRM is that: the canoni…

Cited by 0SourcecodeScholar
2024

Context-I2W: Mapping Images to Context-Dependent Words for Accurate Zero-Shot Composed Image Retrieval

AAAI 2024technical

Different from the Composed Image Retrieval task that requires expensive labels for training task-specific models, Zero-Shot Composed Image Retrieval (ZS-CIR) involves diverse tasks with a broad range of visual content manipulation intent that could be related to domain, scene, object, and attribute…

2024

D-PBS: Dueling Priority-Based Search for Multiple Nonholonomic Robots Motion Planning in Congested Environments

RA-L 2024

This letter focuses on the multiple nonholonomic robots motion planning (MRMP) problem in congested and complex environments, where the complexity escalates dramatically with the increase in the number of robots, frequently leading to deadlocks. We present the <italic xmlns:mml="http://www.w3.org/19

Cited by 11SourceScholar
2024

Evaluate Geometry of Radiance Fields with Low-Frequency Color Prior

AAAI 2024technical

A radiance field is an effective representation of 3D scenes, which has been widely adopted in novel-view synthesis and 3D reconstruction. It is still an open and challenging problem to evaluate the geometry, i.e., the density field, as the ground-truth is almost impossible to obtain. One alternativ…

2024

RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences

ICML 2024spotlight

Preference-based Reinforcement Learning (PbRL) circumvents the need for reward engineering by harnessing human preferences as the reward signal. However, current PbRL methods excessively depend on high-quality feedback from domain experts, which results in a lack of robustness. In this paper, we pre…

2024

SC-Tune: Unleashing Self-Consistent Referential Comprehension in Large Vision Language Models

CVPR 2024poster

Recent trends in Large Vision Language Models (LVLMs) research have been increasingly focusing on advancing beyond general image understanding towards more nuanced object-level referential comprehension. In this paper we present and delve into the self-consistency capability of LVLMs a crucial aspec…

2022

GraphFit: Learning Multi-Scale Graph-Convolutional Representation for Point Cloud Normal Estimation

ECCV 2022poster

"We propose a precise and efficient normal estimation method that can deal with noise and nonuniform density for unstructured 3D point clouds. Unlike existing approaches that directly take patches and ignore the local neighborhood relationships, which make them susceptible to challenging regions suc…

2019

A GPU Based Parallel Genetic Algorithm for the Orientation Optimization Problem in 3D Printing

ICRA 2019poster

The choice of model orientation is a very important issue in Additive Manufacturing (AM). In this paper, the model orientation problem is formulated as a multi-objective optimization problem, aiming at minimizing the building time, the surface quality, and the supporting area. Then we convert the pr…

Cited by 16SourceScholar