← Search

Nong Sang

53 accepted papers

2026

DenseGRPO: From Sparse to Dense Reward for Flow Matching Model Alignment

ICLR 2026poster

Recent GRPO-based approaches built on flow matching models have shown remarkable improvements in human preference alignment for text-to-image generation. Nevertheless, they still suffer from the sparse reward problem: the terminal reward of the entire denoising trajectory is applied to all intermedi…

Cited by 0SourceScholar
2026

HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation

CVPR 2026

Recent unified models integrate understanding experts (e.g., LLMs) with generative experts (e.g., diffusion models), achieving strong multimodal performance. However, recent advanced methods such as BAGEL and LMFusion follow the Mixture-of-Transformers (MoT) paradigm, adopting a symmetric design tha

Cited by 9SourceScholar
2026

Learning to Tell Apart: Weakly Supervised Video Anomaly Detection via Disentangled Semantic Alignment

AAAI 2026technical

Recent advancements in weakly-supervised video anomaly detection have achieved remarkable performance by applying the multiple instance learning paradigm based on multimodal foundation models such as CLIP to highlight anomalous instances and classify categories. However, their objectives may tend to

Cited by 0SourcePDFScholar
2026

SceneDirector: Bridging Explicit Geometry and Generative Priors for Unified Driving Scene Editing

ICML 2026poster

Validating autonomous driving systems requires diverse scenarios, yet real-world data collection is biased and costly. Editing existing driving logs offers a scalable solution, but simultaneously editing objects and ego-trajectory—termed unified editing—remains challenging. Current methods face an i…

Cited by 0SourceScholar
2026

Selecting Samples on Graphs: A Unified Dataset Pruning Framework for Lossless Training Acceleration

ICML 2026poster

The rapid growth of modern training datasets has significantly increased computational cost, motivating dataset pruning(DP) methods which retain only a subset of informative samples to reduce training cost. Existing pruning criteria typically rely on either intrinsic signals that assess samples inde…

Cited by 0SourceScholar
2026

Semantic-Guided Global-Local Collaborative Prompt Learning for Few-Shot Class Incremental Learning

CVPR 2026

Few-Shot Class-Incremental Learning (FSCIL) poses a critical challenge in machine learning, requiring models to continuously integrate novel classes with limited samples while preserving knowledge of previously seen classes. While existing FSCIL approaches have demonstrated promising results, they s

Cited by 0SourceScholar
2025

Adaptive Prototype Replay for Class Incremental Semantic Segmentation

AAAI 2025technical

Class incremental semantic segmentation (CISS) aims to segment new classes during continual steps while preventing the forgetting of old knowledge. Existing methods alleviate catastrophic forgetting by replaying distributions of previously learned classes using stored prototypes or features. However…

2025

Continual Gaussian Mixture Distribution Modeling for Class Incremental Semantic Segmentation

NeurIPS 2025poster

Class incremental semantic segmentation (CISS) enables a model to continually segment new classes from non-stationary data while preserving previously learned knowledge. Recent top-performing approaches are prototype-based methods that assign a prototype to each learned class to reproduce previous k…

Cited by 0SourceScholar
2025

Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity

CVPR 2025highlight

How can we enable models to comprehend video anomalies occurring over varying temporal scales and contexts?Traditional Video Anomaly Understanding (VAU) methods focus on frame-level anomaly prediction, often missing the interpretability of complex and diverse real-world anomalies. Recent multimodal…

2025

L-Man: A Large Multi-modal Model Unifying Human-centric Tasks

AAAI 2025technical

Large language models (LLMs) have recently shown notable progress in unifying various visual tasks with an open-ended form. However, when transferred to human-centric tasks, despite their remarkable multi-modal understanding ability in general domains, they lack further human-related domain knowledg…

Cited by 0SourcePDFScholar
2025

MP-Mat: A 3D-and-Instance-Aware Human Matting and Editing Framework with Multiplane Representation

ICLR 2025poster

Human instance matting aims to estimate an alpha matte for each human instance in an image, which is challenging as it easily fails in complex cases requiring disentangling mingled pixels belonging to multiple instances along hairy and thin boundary structures. In this work, we address this by intro…

2025

Partial Forward Blocking: A Novel Data Pruning Paradigm for Lossless Training Acceleration

ICCV 2025poster

The ever-growing size of training datasets enhances the generalization capability of machine learning models but also incurs exorbitant computational costs. Existing data pruning approaches aim to accelerate training by removing those less important samples. However, they often rely on gradients or…

Cited by 0SourcePDFScholar
2025

ReID5o: Achieving Omni Multi-modal Person Re-identification in a Single Model

NeurIPS 2025poster

In real-word scenarios, person re-identification (ReID) expects to identify a person-of-interest via the descriptive query, regardless of whether the query is a single modality or a combination of multiple modalities. However, existing methods and datasets remain constrained to limited modalities, f…

Cited by 0SourcecodeScholar
2025

Structural Pruning via Spatial-aware Information Redundancy for Semantic Segmentation

AAAI 2025technical

In recent years, semantic segmentation has flourished in various applications. However, the high computational cost remains a significant challenge that hinders its further adoption. The filter pruning method for structured network slimming offers a direct and effective solution for the reduction o…

2025

StyleSRN: Scene Text Image Super-Resolution with Text Style Embedding

ICCV 2025poster

Scene text image super-resolution (STISR) focuses on enhancing the clarity and readability of low-resolution text images. Existing methods often rely on text probability distribution priors derived from text recognizers to guide the super-resolution process. While effective in capturing general stru…

2025

Towards Reliable and Holistic Visual In-Context Learning Prompt Selection

NeurIPS 2025poster

Visual In-Context Learning (VICL) has emerged as a prominent approach for adapting visual foundation models to novel tasks, by effectively exploiting contextual information embedded in in-context examples, which can be formulated as a global ranking problem of potential candidates. Current VICL meth…

Cited by 0SourceScholar
2025

VideoLucy: Deep Memory Backtracking for Long Video Understanding

NeurIPS 2025poster

Recent studies have shown that agent-based systems leveraging large language models (LLMs) for key information retrieval and integration have emerged as a promising approach for long video understanding. However, these systems face two major challenges. First, they typically perform modeling and rea…

Cited by 0SourceScholar
2024

A Recipe for Scaling up Text-to-Video Generation with Text-free Videos

CVPR 2024poster

Diffusion-based text-to-video generation has witnessed impressive progress in the past year yet still falls behind text-to-image generation. One of the key reasons is the limited scale of publicly available data (e.g. 10M video-text pairs in WebVid10M vs. 5B image-text pairs in LAION) considering th…

Cited by 37SourcePDFScholar
2024

Cross-video Identity Correlating for Person Re-identification Pre-training

NeurIPS 2024poster

Recent researches have proven that pre-training on large-scale person images extracted from internet videos is an effective way in learning better representations for person re-identification. However, these researches are mostly confined to pre-training at the instance-level or single-video trackle…

2024

HR-Pro: Point-Supervised Temporal Action Localization via Hierarchical Reliability Propagation

AAAI 2024technical

Point-supervised Temporal Action Localization (PSTAL) is an emerging research direction for label-efficient learning. However, current methods mainly focus on optimizing the network either at the snippet-level or the instance-level, neglecting the inherent reliability of point annotations at both le…

2024

Hierarchical Spatio-temporal Decoupling for Text-to-Video Generation

CVPR 2024poster

Despite diffusion models having shown powerful abilities to generate photorealistic images generating videos that are realistic and diverse still remains in its infancy. One of the key reasons is that current methods intertwine spatial content and temporal dynamics together leading to a notably incr…

2024

Open-Vocabulary Semantic Segmentation with Image Embedding Balancing

CVPR 2024poster

Open-vocabulary semantic segmentation is a challenging task which requires the model to output semantic masks of an image beyond a close-set vocabulary. Although many efforts have been made to utilize powerful CLIP models to accomplish this task they are still easily overfitting to training classes…

2024

PLIP: Language-Image Pre-training for Person Representation Learning

NeurIPS 2024poster

Language-image pre-training is an effective technique for learning powerful representations in general domains. However, when directly turning to person representation learning, these general pre-training methods suffer from unsatisfactory performance. The reason is that they neglect critical person…

2024

Real-Time Exposure Correction via Collaborative Transformations and Adaptive Sampling

CVPR 2024poster

Most of the previous exposure correction methods learn dense pixel-wise transformations to achieve promising results but consume huge computational resources. Recently Learnable 3D lookup tables (3D LUTs) have demonstrated impressive performance and efficiency for image enhancement. However these me…

2024

SCTNet: Single-Branch CNN with Transformer Semantic Information for Real-Time Segmentation

AAAI 2024technical

Recent real-time semantic segmentation methods usually adopt an additional semantic branch to pursue rich long-range context. However, the additional branch incurs undesirable computational overhead and slows inference speed. To eliminate this dilemma, we propose SCTNet, a single branch CNN with tra…

2024

UFineBench: Towards Text-based Person Retrieval with Ultra-fine Granularity

CVPR 2024poster

Existing text-based person retrieval datasets often have relatively coarse-grained text annotations. This hinders the model to comprehend the fine-grained semantics of query texts in real scenarios. To address this problem we contribute a new benchmark named UFineBench for text-based person retrieva…

2023

Disentangling Spatial and Temporal Learning for Efficient Image-to-Video Transfer Learning

ICCV 2023poster

Recently, large-scale pre-trained language-image models like CLIP have shown extraordinary capabilities for understanding spatial contents, but naively transferring such models to video recognition still suffers from unsatisfactory temporal modelling capabilities. Existing methods insert tunable str…

Cited by 29PDFcodeScholar
2023

Lookup Table meets Local Laplacian Filter: Pyramid Reconstruction Network for Tone Mapping

NeurIPS 2023poster

Tone mapping aims to convert high dynamic range (HDR) images to low dynamic range (LDR) representations, a critical task in the camera imaging pipeline. In recent years, 3-Dimensional LookUp Table (3D LUT) based methods have gained attention due to their ability to strike a favorable balance between…

2023

MoLo: Motion-Augmented Long-Short Contrastive Learning for Few-Shot Action Recognition

CVPR 2023poster

Current state-of-the-art approaches for few-shot action recognition achieve promising performance by conducting frame-level matching on learned visual features. However, they generally suffer from two limitations: i) the matching procedure between local frames tends to be inaccurate due to the lack…

2023

Towards General Low-Light Raw Noise Synthesis and Modeling

ICCV 2023poster

Modeling and synthesizing low-light raw noise is a fundamental problem for computational photography and image processing applications. Although most recent works have adopted physics-based models to synthesize noise, the signal-independent noise in low-light conditions is far more complicated and v…

Cited by 16PDFcodeScholar
2022

Applying Deep Learning to Known-Plaintext Attack on Chaotic Image Encryption Schemes

ICASSP 2022accepted

In this paper, we demonstrate that traditional chaotic encryption schemes are vulnerable to the known-plaintext attack (KPA) with deep learning. Considering the decryption process as image restoration based on deep learning, we apply Convolutional Neural Network to perform known-plaintext attack on…

Cited by 0SourceScholar
2022

Hierarchical Feature Embedding for Visual Tracking

ECCV 2022poster

"Features extracted by existing tracking methods may contain instance- and category-level information. However, it usually occurs that either instance- or category-level information uncontrollably dominates the feature embeddings depending on the training data distribution, since the two types of in…

2022

Hybrid Relation Guided Set Matching for Few-Shot Action Recognition

CVPR 2022poster

Current few-shot action recognition methods reach impressive performance by learning discriminative features for each video via episodic training and designing various temporal alignment strategies. Nevertheless, they are limited in that (a) learning individual features without considering the entir…

Cited by 121PDFcodeScholar
2022

Learning From Untrimmed Videos: Self-Supervised Video Representation Learning With Hierarchical Consistency

CVPR 2022poster

Natural videos provide rich visual contents for self-supervised learning. Yet most existing approaches for learning spatio-temporal representations rely on manually trimmed videos, leading to limited diversity in visual patterns and limited performance gain. In this work, we aim to learn representat…

Cited by 20PDFScholar
2022

Learning a Condensed Frame for Memory-Efficient Video Class-Incremental Learning

NeurIPS 2022accept

Recent incremental learning for action recognition usually stores representative videos to mitigate catastrophic forgetting. However, only a few bulky videos can be stored due to the limited memory. To address this problem, we propose FrameMaker, a memory-efficient video class-incremental learning…

Cited by 20SourcePDFScholar
2022

Multi-Centroid Representation Network for Domain Adaptive Person Re-ID

AAAI 2022technical

Recently, many approaches tackle the Unsupervised Domain Adaptive person re-identification (UDA re-ID) problem through pseudo-label-based contrastive learning. During training, a uni-centroid representation is obtained by simply averaging all the instance features from a cluster with the same pseudo…

Cited by 72SourcePDFScholar
2021

Decoupled and Memory-Reinforced Networks: Towards Effective Feature Learning for One-Step Person Search

AAAI 2021technical

The goal of person search is to localize and match query persons from scene images. For high efficiency, one-step methods have been developed to jointly handle the pedestrian detection and identification sub-tasks using a single network. There are two major challenges in the current one-step approac…

2021

Lite-HRNet: A Lightweight High-Resolution Network

CVPR 2021poster

We present an efficient high-resolution network, Lite-HRNet, for human pose estimation. We start by simply applying the efficient shuffle block in ShuffleNet to HRNet (high-resolution network), yielding stronger performance over popular lightweight networks, such as MobileNet, ShuffleNet, and Small…

Cited by 503PDFcodeScholar
2021

OadTR: Online Action Detection With Transformers

ICCV 2021poster

Most recent approaches for online action detection tend to apply Recurrent Neural Network (RNN) to capture long-range temporal structure. However, RNN suffers from non-parallelism and gradient vanishing, hence it is hard to be optimized. In this paper, we propose a new encoder-decoder framework base…

Cited by 154PDFcodeScholar
2021

Self-Supervised Learning for Semi-Supervised Temporal Action Proposal

CVPR 2021poster

Self-supervised learning presents a remarkable performance to utilize unlabeled data for various video tasks. In this paper, we focus on applying the power of self-supervised methods to improve semi-supervised action proposal generation. Particularly, we design a Self-supervised Semi-supervised Temp…

Cited by 82PDFcodeScholar
2021

Temporal Context Aggregation Network for Temporal Action Proposal Refinement

CVPR 2021poster

Temporal action proposal generation aims to estimate temporal intervals of actions in untrimmed videos, which is a challenging yet important task in the video understanding field. The proposals generated by current methods still suffer from inaccurate temporal boundaries and inferior confidence used…

Cited by 166PDFScholar
2021

Weakly Supervised Person Search With Region Siamese Networks

ICCV 2021poster

Supervised learning is dominant in person search, but it requires elaborate labeling of bounding boxes and identities. Large-scale labeled training data is often difficult to collect, especially for person identities. A natural question is whether a good person search model can be trained without th…

Cited by 30PDFScholar
2021

Weakly Supervised Text-Based Person Re-Identification

ICCV 2021poster

The conventional text-based person re-identification methods heavily rely on identity annotations. However, this labeling process is costly and time-consuming. In this paper, we consider a more practical setting called weakly supervised text-based person re-identification, where only the text-image…

Cited by 40PDFcodeScholar
2020

Adversarial Semantic Data Augmentation for Human Pose Estimation

ECCV 2020poster

Human pose estimation is the task of localizing body keypoints from still images. The state-of-the-art methods suffer from insufficient examples of challenging cases such as symmetric appearance, heavy occlusion and nearby person. To enlarge the amounts of challenging cases, previous methods augment…

2020

Do Not Disturb Me: Person Re-identification Under the Interference of Other Pedestrians

ECCV 2020poster

In the conventional person Re-ID setting, it is assumed that cropped images are the person images within the bounding box for each individual. However, in a crowded scene, off-shelf-detectors may generate bounding boxes involving multiple people, where the large proportion of background pedestrians…

2019

Re-ID Driven Localization Refinement for Person Search

ICCV 2019poster

Person search aims at localizing and identifying a query person from a gallery of uncropped scene images. Different from person re-identification (re-ID), its performance also depends on the localization accuracy of a pedestrian detector. The state-of-the-art methods train the detector individually,…

Cited by 162PDFcodeScholar
2018

BiSeNet: Bilateral Segmentation Network for Real-time Semantic Segmentation

ECCV 2018poster

Semantic segmentation requires both rich spatial information and sizeable receptive field. However, modern approaches usually compromise spatial resolution to achieve real-time inference speed, which leads to poor performance. In this paper, we address this dilemma with a novel Bilateral Segmentatio…

Cited by 2734SourcePDFScholar
2018

Learning a Discriminative Feature Network for Semantic Segmentation

CVPR 2018poster

Most existing methods of semantic segmentation still suffer from two aspects of challenges: intra-class inconsistency and inter-class indistinction. To tackle these two problems, we propose a Discriminative Feature Network (DFN), which contains two sub-networks: Smooth Network and Border Network. Sp…

Cited by 987SourcePDFScholar
2018

Learning a Discriminative Prior for Blind Image Deblurring

CVPR 2018poster

We present an effective blind image deblurring method based on a data-driven discriminative prior. Our work is motivated by the fact that a good image prior should favor clear images over blurred images. To obtain such an image prior for deblurring, we formulate the image prior as a binary classifie…

Cited by 195SourcePDFScholar