← Search

Richang Hong

47 accepted papers

2026

Benchmarking Trustworthiness in Multimodal LLMs for Video Understanding

AAAI 2026technical

Recent advancements in multimodal large language models for video understanding (videoLLMs) have enhanced their capacity to process complex spatiotemporal data. However, challenges such as factual inaccuracies, harmful content, biases, hallucinations, and privacy risks compromise their reliability.

Cited by 0SourcePDFScholar
2026

Debate over Mixed-knowledge: A Robust Multi-Agent Reasoning Framework for Incomplete Knowledge Graph Question Answering

AAAI 2026technical

Knowledge Graph Question Answering (KGQA) aims to improve factual accuracy by leveraging structured knowledge. However, real-world Knowledge Graphs (KGs) are often incomplete, leading to the problem of Incomplete KGQA (IKGQA). A common solution is to incorporate external data to fill knowledge gaps,

Cited by 0SourcePDFScholar
2026

DragNeXt: Rethinking Drag-Based Image Editing

AAAI 2026technical

Drag-Based Image Editing (DBIE), which allows users to manipulate images by directly dragging objects within them, has recently attracted much attention from the community. However, it faces two key challenges: (i) point-based drag is often highly ambiguous and difficult to align with user intention

Cited by 0SourcePDFScholar
2026

ProSafePrune: Projected Safety Pruning for Mitigating Over-Refusal in LLMs

ICLR 2026poster

Large Language Models (LLMs) excel in various domains, but their safe deployment faces the challenge of balancing safety and utility. Existing alignment strategies often strengthen refusal mechanisms to reduce harmful outputs, but harmless instructions with superficial risky words are mistakenly rej…

Cited by 0SourcecodeScholar
2026

RecCocktail: A Generalizable and Efficient Framework for LLM-Based Recommendation

AAAI 2026technical

Large Language Models (LLMs) have achieved remarkable success in recent years, owing to their impressive generalization capabilities and rich world knowledge. To capitalize on the potential of using LLMs as recommender systems, mainstream approaches typically focus on two paradigms. The first paradi

Cited by 0SourcePDFScholar
2026

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control

ICML 2026poster

Safety alignment remains brittle under domain shift and noisy preference supervision. Existing robust alignment methods predominantly focus on data uncertainty in alignment data, while being less effective at addressing failures caused by optimization-induced fragility. In this work, we revisit robu…

Cited by 0SourceScholar
2026

Sparse-Scale Transformer with Bidirectional Awareness for Time Series Forecasting

AAAI 2026technical

Time series forecasting (TSF) plays a crucial role in many real-world applications, such as weather prediction and economic planning. While Transformer-based models have shown strong capabilities in modeling long-range dependencies, effectively capturing the multi-scale temporal dynamics inherent in

Cited by 0SourcePDFScholar
2026

Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!

ICLR 2026poster

Achieving streaming, fine-grained control over the outputs of autoregressive video diffusion models remains challenging, making it difficult to ensure that they consistently align with user expectations. To bridge this gap, we propose \textbf{stReaming drag-oriEnted interactiVe vidEo manipuLation (R…

Cited by 0SourceScholar
2026

Your Classifier Can Do More: Towards Balancing the Gaps in Classification, Robustness, and Generation

CVPR 2026

Joint Energy-based Models (JEMs) are well known for their ability to unify classification and generation within a single framework. Despite their promising generative and discriminative performance, their robustness remains far inferior to adversarial training (AT), which, conversely, achieves stron

Cited by 0SourcecodeScholar
2025

Analyzing and Reducing Catastrophic Forgetting in Parameter Efficient Tuning

ICASSP 2025accepted

Existing continual learning works explored strategies like memory replay, regularization, and parameter isolation, but little analysis was conducted on the optimization behavior of LLMs’ continual fine-tuning. In this work, we investigate the geometric connections of different minima along the conti…

Cited by 0SourceScholar
2025

CARE Transformer: Mobile-Friendly Linear Visual Transformer via Decoupled Dual Interaction

CVPR 2025highlight

Recently, large efforts have been made to design efficient linear-complexity visual Transformers. However, current linear attention models are generally unsuitable to be deployed in resource-constrained mobile devices, due to suffering from either few efficiency gains or significant accuracy drops.…

Cited by 0SourcePDFScholar
2025

Discrete to Continuous: Generating Smooth Transition Poses from Sign Language Observations

CVPR 2025poster

Generating continuous sign language videos from discrete segments is challenging due to the need for smooth transitions that preserve natural flow and meaning. Traditional approaches that simply concatenate isolated signs often result in abrupt transitions, disrupting video coherence. To address thi…

Cited by 0SourcePDFScholar
2025

EgoBlind: Towards Egocentric Visual Assistance for the Blind

NeurIPS 2025poster

We present EgoBlind, the first egocentric VideoQA dataset collected from blind individuals to evaluate the assistive capabilities of contemporary multimodal large language models (MLLMs). EgoBlind comprises 1,392 first-person videos from the daily lives of blind and visually impaired individuals. It…

Cited by 0SourcecodeScholar
2025

GT-Mean Loss: A Simple Yet Effective Solution for Brightness Mismatch in Low-Light Image Enhancement

ICCV 2025poster

Low-light image enhancement (LLIE) aims to improve the visual quality of images captured under poor lighting conditions. In supervised LLIE tasks, there exists a significant yet often overlooked inconsistency between the overall brightness of an enhanced image and its ground truth counterpart, refer…

Cited by 0SourcePDFScholar
2025

High-Fidelity Stereoscopic Image Rain Removal with Texture Integrity and Disparity Consistency

ICASSP 2025accepted

This paper tackles the challenge of stereoscopic image rain removal by focusing on enhancing texture integrity and disparity consistency. Existing stereoscopic rain removal techniques often fall short due to 1) disruptions in texture coherence caused by complex rain streaks, and 2) inaccuracies in d…

Cited by 0SourceScholar
2025

Learnable Feature Patches and Vectors for Boosting Low-light Image Enhancement without External Knowledge

ICCV 2025poster

A major challenge in Low-Light Image Enhancement (LLIE) is its ill-posed nature: low-light images often lack sufficient information to align with normal-light ones (e.g., not all training data can be fully fitted to the ground truth). Numerous studies have attempted to bridge the gap between low- an…

Cited by 0SourcePDFScholar
2025

Linguistics-Vision Monotonic Consistent Network for Sign Language Production

ICASSP 2025accepted

Sign Language Production (SLP) aims to generate sign videos corresponding to spoken language sentences, where the conversion of sign Glosses to Poses (G2P) is the key step. Due to the cross-modal semantic gap and the lack of word-action correspondence labels for strong supervision alignment, the SLP…

Cited by 0SourceScholar
2025

Moderating the Generalization of Score-based Generative Model

ICCV 2025poster

Score-based Generative Models (SGMs) have demonstrated remarkable generalization capabilities, e.g. generating unseen, but natural data. However, the greater the generalization power, the more likely the unintended generalization, and the more dangerous the abuse. Despite these concerns, research on…

2025

RAGG: Retrieval-Augmented Grasp Generation Model

AAAI 2025technical

Intent-based grasp generation inherently involves challenges such as manipulation ambiguity and modality gaps. To address these, we propose a novel Retrieval-Augmented Grasp Generation model (RAGG). Our key insight is that when humans manipulate new objects, they initially mimic the interaction patt…

Cited by 0SourcePDFScholar
2025

Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval

NeurIPS 2025poster

Recent progress in text–video retrieval has been largely driven by contrastive learning. However, existing methods often overlook the effect of the modality gap, which causes anchor representations to undergo in-place optimization (i.e., optimization tension) that limits their alignment capacity. Mo…

Cited by 0SourceScholar
2025

Revisiting Audio-Visual Segmentation with Vision-Centric Transformer

CVPR 2025poster

Audio-Visual Segmentation (AVS) aims to segment sound-producing objects in video frames based on the associated audio signal. Prevailing AVS methods typically adopt an audio-centric Transformer architecture, where object queries are derived from audio features. However, audio-centric Transformers su…

Cited by 0SourcePDFScholar
2025

SURE: Safety Understanding and Reasoning Enhancement for Multimodal Large Language Models

EMNLP 2025

Multimodal large language models (MLLMs) demonstrate impressive capabilities by integrating visual and textual information. However, the incorporation of visual modalities also introduces new and complex safety risks, rendering even the most advanced models vulnerable to sophisticated jailbreak atta

2025

Sign-IDD: Iconicity Disentangled Diffusion for Sign Language Production

AAAI 2025technical

Sign Language Production (SLP) aims to generate semantically consistent sign videos from textual statements, where the conversion from textual glosses to sign poses (G2P) is a crucial step. Existing G2P methods typically treat sign poses as discrete three-dimensional coordinates and directly fit the…

2025

Text Proxy: Decomposing Retrieval from a 1-to-N Relationship into N 1-to-1 Relationships for Text-Video Retrieval

AAAI 2025technical

Text-video retrieval (TVR) has seen substantial advancements in recent years, fueled by the utilization of pre-trained models and large language models (LLMs). Despite these advancements, achieving accurate matching in TVR remains challenging due to inherent disparities between video and textual mod…

2025

Unveiling Uncertainty: A Deep Dive into Calibration and Performance of Multimodal Large Language Models

COLING 2025main

Multimodal large language models (MLLMs) combine visual and textual data for tasks like image captioning and visual question answering. Proper uncertainty calibration is crucial but challenging for reliable use in areas like healthcare and autonomous driving. This paper investigates several MLLMs, f…

2024

Doubly Abductive Counterfactual Inference for Text-based Image Editing

CVPR 2024poster

We study text-based image editing (TBIE) of a single image by counterfactual inference because it is an elegant formulation to precisely address the requirement: the edited image should retain the fidelity of the original one. Through the lens of the formulation we find that the crux of TBIE is that…

2024

Few-shot Learner Parameterization by Diffusion Time-steps

CVPR 2024poster

Even when using large multi-modal foundation models few-shot learning is still challenging -- if there is no proper inductive bias it is nearly impossible to keep the nuanced class attributes while removing the visually prominent attributes that spuriously correlate with class labels. To this end we…

2024

Gradient-Aware Logit Adjustment Loss for Long-Tailed Classifier

ICASSP 2024accepted

In the real-world setting, data often follows a long-tailed distribution, where head classes contain significantly more training samples than tail classes. Consequently, models trained on such data tend to be biased toward head classes. The medium of this bias is imbalanced gradients, which include…

Cited by 0SourceScholar
2024

Towards Unified Multimodal Editing with Enhanced Knowledge Collaboration

NeurIPS 2024spotlight

The swift advancement in Multimodal LLMs (MLLMs) also presents significant challenges for effective knowledge editing. Current methods, including intrinsic knowledge editing and external knowledge resorting, each possess strengths and weaknesses, struggling to balance the desired properties of relia…

2023

Disentangling Cognitive Diagnosis with Limited Exercise Labels

NeurIPS 2023poster

Cognitive diagnosis is an important task in intelligence education, which aims at measuring students’ proficiency in specific knowledge concepts. Given a fully labeled exercise-concept matrix, most existing models focused on mining students' response records for cognitive diagnosis. Despite their su…

2023

Fair Representation Learning for Recommendation: A Mutual Information Perspective

AAAI 2023technical

Recommender systems have been widely used in recent years. By exploiting historical user-item interactions, recommender systems can model personalized potential interests of users and have been widely applied to a wide range of scenarios. Despite their impressive performance, most of them may be sub…

Cited by 20SourcePDFScholar
2022

Deep Color Consistent Network for Low-Light Image Enhancement

CVPR 2022poster

Low-light image enhancement focus on refining the illumination and keep naturalness to obtain the normal-light image. Current low-light image enhancement methods can well improve the illumination. However, there is still color difference between the enhanced image and the ground-truth image. To alle…

Cited by 165PDFcodeScholar
2022

Multi-scale Spatial Representation Learning via Recursive Hermite Polynomial Networks

IJCAI 2022poster

Multi-scale representation learning aims to leverage diverse features from different layers of Convolutional Neural Networks (CNNs) for boosting the feature robustness to scale variance. For dense prediction tasks, two key properties should be satisfied: the high spatial variance across convolutiona…

2022

Vibration-Based Uncertainty Estimation for Learning from Limited Supervision

ECCV 2022poster

"We investigate the problem of estimating uncertainty for training data, so that deep neural networks can make use of the results for learning from limited supervision. However, both prediction probability and entropy estimate uncertainty from the instantaneous information. In this paper, we present…

Cited by 4SourcePDFScholar
2020

Creating Something From Nothing: Unsupervised Knowledge Distillation for Cross-Modal Hashing

CVPR 2020poster

In recent years, cross-modal hashing (CMH) has attracted increasing attentions, mainly because its potential ability of mapping contents from different modalities, especially in vision and language, into the same space, so that it becomes efficient in cross-modal data retrieval. There are two main f…

Cited by 151PDFScholar
2020

Real-World Person Re-Identification via Degradation Invariance Learning

CVPR 2020poster

Person re-identification (Re-ID) in real-world scenarios usually suffers from various degradation factors, e.g., low-resolution, weak illumination, blurring and adverse weather. On the one hand, these degradations lead to severe discriminative information loss, which significantly obstructs identity…

Cited by 89PDFScholar
2019

Adaptive Transfer Network for Cross-Domain Person Re-Identification

CVPR 2019poster

Recent deep learning based person re-identification approaches have steadily improved the performance for benchmarks, however they often fail to generalize well from one domain to another. In this work, we propose a novel adaptive transfer network (ATNet) for effective cross-domain person re-identif…

Cited by 350PDFScholar
2018

Interleaved Structured Sparse Convolutional Neural Networks

CVPR 2018poster

In this paper, we study the problem of designing efficient convolutional neural network architectures with the interest in eliminating the redundancy in convolution kernels. In addition to structured sparse kernels, low-rank kernels and the product of low-rank kernels,the product of structured spars…

Cited by 160SourcePDFScholar
2018

Multi-Cue Correlation Filters for Robust Visual Tracking

CVPR 2018poster

In recent years, many tracking algorithms achieve impressive performance via fusing multiple types of features, however, most of them fail to fully explore the context among the adopted multiple features and the strength of them. In this paper, we propose an efficient multi-cue analysis framework fo…

2016

Cascaded Interactional Targeting Network for Egocentric Video Analysis

CVPR 2016poster

Knowing how hands move and what object is being manipulated are two key sub-tasks for analyzing first-person (egocentric) action. However, lack of fully annotated hand data as well as imprecise foreground segmentation make either sub-task challenging. This work aims to explicitly address these two i…

Cited by 68PDFScholar
2015

Interaction Part Mining: A Mid-Level Approach for Fine-Grained Action Recognition

CVPR 2015poster

Modeling human-object interactions and manipulating motions lies in the heart of fine-grained action recognition. Previous methods heavily rely on explicit detection of the object being interacted, which requires intensive human labour on object annotation. To bypass this constraint and achieve bett…

Cited by 103SourcePDFScholar