← Search

Sam Kwong

27 accepted papers

2026

Optimizing Agentic Reasoning with Retrieval via Synthetic Semantic Information Gain Reward

ICML 2026poster

Agentic reasoning enables large reasoning models (LRMs) to dynamically acquire external knowledge, but yet optimizing the retrieval process remains challenging due to the lack of dense, principled reward signals. In this paper, we introduce *InfoReasoner*, a unified framework that incentivizes effec…

Cited by 0SourceScholar
2026

WaterFlow: Explicit Physics-Prior Rectified Flow for Underwater Saliency Mask Generation

ICASSP 2026poster

Underwater Salient Object Detection (USOD) faces significant challenges, including underwater image quality degradation and domain gaps. Existing methods tend to ignore the physical principles of underwater imaging or simply treat degradation phenomena in underwater images as interference factors th…

Cited by 0SourcePDFScholar
2026

When Privacy Meets Recovery: The Overlooked Half of Surrogate-Driven Privacy Preservation for MLLM Editing

AAAI 2026technical

Privacy leakage in Multimodal Large Language Models (MLLMs) has long been an intractable problem. Existing studies, though effectively obscure private information in MLLMs, often overlook the evaluation of authenticity and recovery quality of user privacy. To this end, this work uniquely focuses on

Cited by 0SourcePDFScholar
2025

AI-generated Image Quality Assessment in Visual Communication

AAAI 2025technical

Assessing the quality of artificial intelligence-generated images (AIGIs) plays a crucial role in their application in real-world scenarios. However, traditional image quality assessment (IQA) algorithms primarily focus on low-level visual perception, while existing IQA works on AIGIs overemphasize…

2025

An Efficient Sample Utilization Method for Deep Learning Based on Class Uncertainty

ICASSP 2025accepted

Deep learning has achieved success across many domains when sufficient training samples are available. However, the commonly used mini-batch stochastic gradient descent (SGD) training paradigm treats each sample equally, resulting in massive computational waste on samples that are easily identifiabl…

Cited by 0SourceScholar
2025

CP-Guard: Malicious Agent Detection and Defense in Collaborative Bird’s Eye View Perception

AAAI 2025technical

Collaborative Perception (CP) has shown a promising technique for autonomous driving, where multiple connected and autonomous vehicles (CAVs) share their perception information to enhance the overall perception performance and expand the perception range. However, in CP, ego CAV needs to receive mes…

Cited by 3SourcePDFScholar
2025

Decoupled Motion Expression Video Segmentation

CVPR 2025poster

Motion expression video segmentation aims to segment objects based on input motion descriptions. Compared with traditional referring video object segmentation, it focuses on motion and multi-object expressions and is more challenging. Previous works achieved it by simply injecting text information i…

Cited by 0SourcePDFScholar
2025

Distribution-Aligned Decoding for Efficient LLM Task Adaptation

NeurIPS 2025poster

Adapting billion-parameter language models to a downstream task is still costly, even with parameter-efficient fine-tuning (PEFT). We re-cast task adaptation as output-distribution alignment: the objective is to steer the output distribution toward the task distribution directly during decoding rath…

Cited by 0SourceScholar
2025

Joint Semantic and Rendering Enhancements in 3D Gaussian Modeling with Anisotropic Local Encoding

ICCV 2025poster

Recent works propose extending 3DGS with semantic feature vectors for simultaneous semantic segmentation and image rendering. However, these methods often treat the semantic and rendering branches separately, relying solely on 2D supervision while ignoring the 3D Gaussian geometry. Moreover, current…

2025

MR-FIQA: Face Image Quality Assessment with Multi-Reference Representations from Synthetic Data Generation

ICCV 2025poster

Recent advancements in Face Image Quality Assessment (FIQA) models trained on real large-scale face datasets are pivotal in guaranteeing precise face recognition in unrestricted scenarios. Regrettably, privacy concerns lead to the discontinuation of real datasets, underscoring the pressing need for…

2025

SEFE: Superficial and Essential Forgetting Eliminator for Multimodal Continual Instruction Tuning

ICML 2025poster

Multimodal Continual Instruction Tuning (MCIT) aims to enable Multimodal Large Language Models (MLLMs) to incrementally learn new tasks without catastrophic forgetting, thus adapting to evolving requirements. In this paper, we explore the forgetting caused by such incremental training, categorizing…

2024

CLIB-FIQA: Face Image Quality Assessment with Confidence Calibration

CVPR 2024poster

Face Image Quality Assessment (FIQA) is pivotal for guaranteeing the accuracy of face recognition in unconstrained environments. Recent progress in deep quality-fitting-based methods that train models to align with quality anchors has shown promise in FIQA. However these methods heavily depend on a…

2024

ColNeRF: Collaboration for Generalizable Sparse Input Neural Radiance Field

AAAI 2024technical

Neural Radiance Fields (NeRF) have demonstrated impressive potential in synthesizing novel views from dense input, however, their effectiveness is challenged when dealing with sparse input. Existing approaches that incorporate additional depth or semantic supervision can alleviate this issue to an e…

2024

Diving into Underwater: Segment Anything Model Guided Underwater Salient Instance Segmentation and A Large-scale Dataset

ICML 2024poster

With the breakthrough of large models, Segment Anything Model (SAM) and its extensions have been attempted to apply in diverse tasks of computer vision. Underwater salient instance segmentation is a foundational and vital step for various underwater vision tasks, which often suffer from low segmenta…

2023

Saving 100x Storage: Prototype Replay for Reconstructing Training Sample Distribution in Class-Incremental Semantic Segmentation

NeurIPS 2023poster

Existing class-incremental semantic segmentation (CISS) methods mainly tackle catastrophic forgetting and background shift, but often overlook another crucial issue. In CISS, each step focuses on different foreground classes, and the training set for a single step only includes images containing pix…

2023

Troubleshooting Ethnic Quality Bias with Curriculum Domain Adaptation for Face Image Quality Assessment

ICCV 2023poster

Face Image Quality Assessment (FIQA) lays the foundation for ensuring the stability and accuracy of face recognition systems. However, existing FIQA methods mainly formulate quality relationships within the training set to yield quality scores, ignoring the generalization problem caused by ethnic qu…

Cited by 9PDFcodeScholar
2023

WaterMask: Instance Segmentation for Underwater Imagery

ICCV 2023poster

Underwater image instance segmentation is a fundamental and critical step in underwater image analysis and understanding. However, the paucity of general multiclass instance segmentation datasets has impeded the development of instance segmentation studies for underwater images. In this paper, we pr…

Cited by 35PDFcodeScholar
2021

Joint Reinforcement Learning and Game Theory Bitrate Control Method for 360-Degree Dynamic Adaptive Streaming

ICASSP 2021accepted

A joint reinforcement learning (RL) and game theory method is presented for segment-level continuous bitrate selection and tile-level bitrate allocation in tile-based 360-degree streaming to increase users’ quality of experience (QoE). First, a viewpoint prediction method based on single-user (SU) v…

Cited by 0SourceScholar
2020

Intra Frame Rate Control for Versatile Video Coding with Quadratic Rate-Distortion Modelling

ICASSP 2020accepted

With numerous coding tools adopted in the forthcoming Versatile Video Coding (VVC) standard, much less work has been dedicated to study the corresponding Rate-Distortion (R-D) characteristics. This paper proposes a new quadratic R-D model for Versatile Video Coding. In particular, based on the propo…

Cited by 0SourceScholar
2020

Just Noticeable Distortion Based Perceptually Lossless Intra Coding

ICASSP 2020accepted

Perceptual video coding plays a very important role in video codec optimization aiming at removing the perceptual redundancies in video content. In this paper, a just noticeable distortion (JND) guided perceptually lossless coding framework is proposed for Versatile Video Coding (VVC) intra coding.…

Cited by 0SourceScholar
2020

Light Field Spatial Super-Resolution via Deep Combinatorial Geometry Embedding and Structural Consistency Regularization

CVPR 2020poster

Light field (LF) images acquired by hand-held devices usually suffer from low spatial resolution as the limited sampling resources have to be shared with the angular dimension. LF spatial super-resolution (SR) thus becomes an indispensable part of the LF camera processing pipeline. The high-dimensio…

Cited by 188PDFcodeScholar
2020

PUGeo-Net: A Geometry-centric Network for 3D Point Cloud Upsampling

ECCV 2020poster

In this paper, we propose a novel deep neural network based method, called PUGeo-Net, for upsampling 3D point clouds. PUGeo-Net incorporates discrete differential geometry into deep learning elegantly by learning the first and second fundamental forms that are able to fully represent the local geome…

2020

Zero-Reference Deep Curve Estimation for Low-Light Image Enhancement

CVPR 2020poster

The paper presents a novel method, Zero-Reference Deep Curve Estimation (Zero-DCE), which formulates light enhancement as a task of image-specific curve estimation with a deep network. Our method trains a lightweight deep network, DCE-Net, to estimate pixel-wise and high-order curves for dynamic ran…

Cited by 2093PDFcodeScholar
2019

Fast Coding Unit Decision for Intra Screen Content Coding Based on Ensemble Learning

ICASSP 2019accepted

The Screen Content Coding (SCC) is an extension of High Efficiency Video Coding (HEVC), and it achieves significant improvement on compression ratio. However, the obtained coding efficiency is at the cost of high computational complexity. In this paper, to reduce the computation complexity, we propo…

Cited by 0SourceScholar
2019

Learning to Explore Intrinsic Saliency for Stereoscopic Video

CVPR 2019poster

The human visual system excels at biasing the stereoscopic visual signals by the attention mechanisms. Traditional methods relying on the low-level features and depth relevant information for stereoscopic video saliency prediction have fundamental limitations. For example, it is cumbersome to model…

Cited by 6PDFScholar