← Search

Peng Jiang

39 accepted papers

2026

Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step Generation

CVPR 2026

Video generation has recently emerged as a central task in the field of generative AI. However, the substantial computational cost inherent in video synthesis makes model distillation a critical technique for efficient deployment. Despite its significance, there is a scarcity of methods specifically

Cited by 0SourceScholar
2026

Align³GR: Unified Multi-Level Alignment for LLM-based Generative Recommendation

AAAI 2026technical

Large Language Models (LLMs) demonstrate significant advantages in leveraging structured world knowledge and multi-step reasoning capabilities. However, fundamental challenges arise when transforming LLMs into real-world recommendation systems due to semantic and behavioral misalignment. To bridge

Cited by 7SourcePDFScholar
2026

AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generation

CVPR 2026

Short-form videos have become a primary medium for digital advertising, requiring scalable and efficient content creation. However, current workflows and AI tools remain disjoint and modality-specific, leading to high production costs and low overall efficiency. To address this issue, we propose Aut

Cited by 0SourcecodeScholar
2026

Fairness-Aware Design for Contextual Experiments: Guaranteeing Reliability and Equity in Heterogeneous Subgroups

AAAI 2026technical

Experimental design is critical for evidence-based decision-making in healthcare, marketing, and public policy. However, designing efficient experiments across heterogeneous subgroups presents significant challenges. Existing methods often optimize for statistical power or overall sample efficiency,

Cited by 0SourcePDFScholar
2026

Generic Adversarial Attack Framework Against Graph-based Vertical Federated Learning

AAAI 2026technical

Graph-based vertical federated learning (GVFL) enables multiple parties to collaboratively train and infer over aligned nodes, where each party contributes its own local embedding derived from different attributes and adjacency relations. Adversarial inputs injected by an attacker can skew the joint

Cited by 0SourcePDFScholar
2026

Learning to Rank by Directly Optimizing Full-Order Probabilities

ICML 2026poster

Learning to rank can be cast as a probabilistic modeling problem over permutations, where the goal is to estimate the likelihood of an observed total ordering of items. This formulation naturally involves full-order probabilities of the form $\mathbb{P}(\mathrm{z}_1 < \cdots < \mathrm{z}_n)$, whose …

Cited by 0SourceScholar
2026

Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning

CVPR 2026

We present Narrative Weaver, a novel framework that addresses a fundamental challenge in generative AI: achieving controllable, long-range, and consistent visual content generation. While existing models excel at generating high-fidelity short-form visual content, they struggle to maintain narrative

Cited by 0SourcecodeScholar
2026

Phase-Aware Mixture of Experts for Agentic Reinforcement Learning

ICML 2026poster

Reinforcement learning (RL) has equipped LLM agents with a strong ability to solve complex tasks. However, existing RL methods normally use a single policy network, causing simplicity bias where simple tasks occupy most parameters and dominate gradient updates, leaving insufficient capacity for comp…

Cited by 0SourceScholar
2026

Stylos: Multi-View 3D Stylization with Single-Forward Gaussian Splatting

ICLR 2026poster

We present Stylos, a single-forward 3D Gaussian framework for 3D style transfer that operates on unposed content, from a single image to a multi-view collection, conditioned on a separate reference style image. Stylos synthesizes a stylized 3D Gaussian scene without per-scene optimization or precomp…

Cited by 0SourcecodeScholar
2026

Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM Encoders

ICML 2026poster

Recent progress in text-to-image (T2I) diffusion models (DMs) has enabled high-quality visual synthesis from diverse textual prompts. Yet, most existing T2I DMs, even those equipped with large language model (LLM)-based text encoders, remain text-pixel mappers -- they employ LLMs merely as text enco…

Cited by 0SourceScholar
2025

D&M: Enriching E-commerce Videos with Sound Effects by Key Moment Detection and SFX Matching

AAAI 2025technical

Videos showcasing specific products are increasingly important for E-commerce. Key moments naturally exist as the first appearance of a specific product, presentation of its distinctive features, the presence of a buying link, etc. Adding proper sound effects (SFX) to such moments, or video decorati…

Cited by 0SourcePDFScholar
2025

Improving Preference Alignment of LLM with Inference-Free Self-Refinement

EMNLP 2025

Large language models (LLMs) develop the in-context learning capability through pretraining and instruction tuning, enabling task adaptation without parameter updates. Self-refinement is a manifestation of this capability, which allows LLMs to iteratively refine the output using self-generated feedb

2025

LEARN: Knowledge Adaptation from Large Language Model to Recommendation for Practical Industrial Application

AAAI 2025technical

Contemporary recommendation systems predominantly rely on ID embedding to capture latent associations among users and items. However, this approach overlooks the wealth of semantic information embedded within textual descriptions of items, leading to suboptimal performance and poor generalizations.…

2025

LLM-Powered User Simulator for Recommender System

AAAI 2025technical

User simulators can rapidly generate a large volume of timely user behavior data, providing a testing platform for reinforcement learning-based recommender systems, thus accelerating their iteration and optimization. However, prevalent user simulators generally suffer from significant limitations, i…

2025

Learning Monotonic Probabilities with a Generative Cost Model

ICML 2025poster

In many machine learning tasks, it is often necessary for the relationship between input and output variables to be monotonic, including both strictly monotonic and implicitly monotonic relationships. Traditional methods for maintaining monotonicity mainly rely on construction or regularization tech…

2025

Learning Multiple User Distributions for Recommendation via Guided Conditional Diffusion

AAAI 2025technical

Recommender systems are increasingly prevalent to provide personalized suggestions and enhance user satisfaction. Typical recommendation models encode users and items as embeddings, and generate recommendations by assessing the similarity between these embeddings. Despite their effectiveness, these…

2025

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

ICML 2025poster

We introduce Orthus, a unified multimodal model that excels in generating interleaved images and text from mixed-modality inputs by simultaneously handling discrete text tokens and continuous image features under the \textbf{AR} modeling principle. The continuous treatment of visual signals minimize…

Cited by 8SourcePDFScholar
2025

RevPv8: Reverse Information for Road Space and Lane Line Segmentation in Highway Surveillance Scene

ICASSP 2025accepted

Road space and lane line segmentation are vital tasks in intelligent traffic surveillance, yet they receive less attention compared to similar tasks in autonomous driving. In highway surveillance, segmenting occluded road areas and lane lines is crucial for comprehensive perception, though it adds c…

Cited by 0SourceScholar
2025

Structured Spectral Reasoning for Frequency-Adaptive Multimodal Recommendation

NeurIPS 2025poster

Multimodal recommendation aims to integrate collaborative signals with heterogeneous content such as visual and textual information, but remains challenged by modality-specific noise, semantic inconsistency, and unstable propagation over user–item graphs. These issues are often exacerbated by naive…

Cited by 0SourceScholar
2025

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization

ICCV 2025poster

This paper presents the Semantic-aWarE spatial-tEmporal Tokenizer (SweetTok), a novel video tokenizer to overcome the limitations in current video tokenization methods for compacted yet effective discretization. Unlike previous approaches that process flattened local visual patches via direct discre…

Cited by 0SourcePDFScholar
2024

Coupling Self-Supervised and Supervised Contrastive Learning for Multiple Classification of Cervical Cytological Whole Slide Images

ICASSP 2024accepted

Cervical cytologic whole slide image (WSI) multiple classificaton (grading) is a challenging task. Current studies typically ignore the unbalanced data distribution and require multi-class annotations to learn cell features for WSI grading, which largely suffers from label noise. In this paper, we d…

Cited by 0SourceScholar
2023

Centimeter-Scale Underwater Robot With High-Speed Inspired by Jellyfish

RA-L 2023

Centimeter-scale underwater robots have important applications in underwater resource exploration, environmental monitoring, equipment fault diagnosis, and military applications. However, developing small underwater robots remains a challenge because of the limitation of miniaturized structures. In

Cited by 13SourceScholar
2023

Classifying Pathological Images Based on Multi-Instance Learning and End-to-End Attention Pooling

ICASSP 2023accepted

In order to address the issue that previous deep learning methods for classifying pathological images cannot adaptively learn features, we propose an end-to-end attention pooling method based on a multi-instance learning patch scoring model. Our method integrates feature extraction and classificatio…

Cited by 0SourceScholar
2023

DDN: Dynamic Aggregation Enhanced Dual-Stream Network for Medical Image Classification

ICASSP 2023accepted

Convolutional Neural Networks (CNNs) have become the de facto approach for medical image classification in recent years. However, the deficiency of convolutional operations in extracting global features has limited the further improvement of this task. Vision Transformers (ViTs) can model long-range…

Cited by 0SourceScholar
2023

LGVIT: Local-Global Vision Transformer for Breast Cancer Histopathological Image Classification

ICASSP 2023accepted

Breast cancer histopathological image classification has made great progress with the use of Convolutional Neural Networks (CNNs). However, due to the limited receptive field, CNNs have difficulty in learning the global information of breast cancer histopathological images, hindering the further imp…

Cited by 0SourceScholar
2023

ResAct: Reinforcing Long-term Engagement in Sequential Recommendation with Residual Actor

ICLR 2023poster

Long-term engagement is preferred over immediate engagement in sequential recommendation as it directly affects product operational metrics such as daily active users (DAUs) and dwell time. Meanwhile, reinforcement learning (RL) is widely regarded as a promising framework for optimizing long-term en…

Cited by 30SourcePDFScholar
2023

State Regularized Policy Optimization on Data with Dynamics Shift

NeurIPS 2023poster

In many real-world scenarios, Reinforcement Learning (RL) algorithms are trained on data with dynamics shift, i.e., with different underlying environment dynamics. A majority of current methods address such issue by training context encoders to identify environment parameters. Data with dynamics shi…

Cited by 17SourcePDFScholar
2022

Exposing and Exploiting Fine-Grained Block Structures for Fast and Accurate Sparse Training

NeurIPS 2022accept

Sparse training is a popular technique to reduce the overhead of training large models. Although previous work has shown promising results for nonstructured sparse models, it is still unclear whether a sparse model with structural constraints can be trained from scratch to high accuracy. In this wor…

Cited by 20SourcePDFScholar
2022

GALAXY: A Generative Pre-trained Model for Task-Oriented Dialog with Semi-supervised Learning and Explicit Policy Injection

AAAI 2022technical

Pre-trained models have proved to be powerful in enhancing task-oriented dialog systems. However, current pre-training methods mainly focus on enhancing dialog understanding and generation tasks while neglecting the exploitation of dialog policy. In this paper, we propose GALAXY, a novel pre-trained…

2021

Scribble-Supervised Semantic Segmentation by Uncertainty Reduction on Neural Representation and Self-Supervision on Neural Eigenspace

ICCV 2021poster

Scribble-supervised semantic segmentation has gained much attention recently for its promising performance without high-quality annotations. Due to the lack of supervision, confident and consistent predictions are usually hard to obtain. Typically, people handle these problems by either adopting an…

Cited by 48PDFcodeScholar
2018

A Linear Speedup Analysis of Distributed Deep Learning with Sparse and Quantized Communication

NeurIPS 2018poster

The large communication overhead has imposed a bottleneck on the performance of distributed Stochastic Gradient Descent (SGD) for training deep neural networks. Previous works have demonstrated the potential of using gradient sparsification and quantization to reduce the communication cost. Howeve…

Cited by 251SourcePDFScholar
2018

DifNet: Semantic Segmentation by Diffusion Networks

NeurIPS 2018poster

Deep Neural Networks (DNNs) have recently shown state of the art performance on semantic segmentation tasks, however, they still suffer from problems of poor boundary localization and spatial fragmented predictions. The difficulties lie in the requirement of making dense predictions from a long path…

Cited by 36SourcePDFScholar