← Search

BO HU

39 accepted papers

2026

Enhancing Persona Following at Decoding Time via Dynamic Importance Estimation for Role-Playing Agents

ICLR 2026poster

The utility of Role-Playing Language Agents in sociological research is growing alongside the adoption of Large Language Models. For realism in social simulation, these agents must adhere to their personas defined by character profiles, yet existing strategies—static prompt engineering or costly fin…

Cited by 0SourceScholar
2026

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding

CVPR 2026

Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation of the output (e.g. the answer to a question), without evaluation of inter- mediate reasoning steps, and most provide answers only in the text doma

Cited by 0SourcecodeScholar
2026

Self-guided Semantic Inspection for Zero-Shot Composed Image Retrieval

CVPR 2026

Zero-Shot Composed Image Retrieval (ZS-CIR) aims to retrieve target images using a composed query of a reference image and a textual modification, without relying on triplet-based supervision. As the two inputs describe related but semantically unaligned information, the key challenge lies in interp

Cited by 0SourcecodeScholar
2026

Size Transferability of Graph Convolutional Networks across Sparsity: A Generalized Graphon Perspective

ICML 2026poster

Size transfer scales Graph Convolutional Networks (GCNs) by applying models trained on sampled subgraphs to larger target graphs. However, existing theoretical guarantees are typically confined to dense graphs or restricted sparsity regimes, failing to cover the arbitrary sparsity of real-world netw…

Cited by 0SourceScholar
2025

A Multi-annotated and Multi-modal Dataset for Wide-angle Video Quality Assessment

ICASSP 2025accepted

Wide-angle video is favored for its wide viewing angle and ability to capture a large area of scenery, making it an ideal choice for sports and adventure recording. However, wide-angle video is prone to deformation, exposure and other distortions, resulting in poor video quality and affecting the pe…

Cited by 0SourceScholar
2025

A Two-Stage AIGC Image Quality Assessment with T2I Correspondence and Visual Perception

ICASSP 2025accepted

Image quality assessment (IQA) of artificial intelligence-generated content (AIGC) has recently attracted significant research attention. Unlike general-purpose IQA, which primarily focuses on evaluating image content, AIGCIQA often requires addressing both the Text-to-Image (T2I) correspondence and…

Cited by 0SourceScholar
2025

AGIAA-2K: A Fine-grained Dataset for Aesthetic and Alignment Evaluation of AI-Generated Images

ICASSP 2025accepted

With the advancement of AI-generated content technologies, AI-generated images (AGIs) have become increasingly influential in artistic creation and visual communication. However, the aesthetic quality of AGIs varies significantly due to technical limitations and the influence of user input, undersco…

Cited by 0SourceScholar
2025

Adaptive Multimodal Fusion: Dynamic Attention Allocation for Intent Recognition

AAAI 2025technical

In recent years, deep multimodal learning has seen significant advancements. However, there remains a lack of multimodal fusion methods capable of dynamically adjusting the weighting of information both within and across modalities based on input samples. In the domain of multimodal intent recogniti…

2025

Bidirectional Reference Image Quality Assessment via Content-Quality Correlation Modeling

ICASSP 2025accepted

The emphasis on no-reference image quality assessment has often overshadowed the significance of Full-Reference Image Quality Assessment (FR-IQA), which generally better reflects human contrastive perception mechanism. However, FRIQA presents challenges in obtaining content-aligned reference images.…

Cited by 0SourceScholar
2025

DSRAG: A Double-Stream Retrieval-Augmented Generation Framework for Countless Intent Detection

NAACL 2025industry

Current intent detection work experiments with minor intent categories. However, in real-world scenarios of data analysis dialogue systems, intents are composed of combinations of numerous metrics and dimensions, resulting in countless intents and posing challenges for the language model. The retrie…

2025

EventMG: Efficient Multilevel Mamba-Graph Learning for Spatiotemporal Event Representation

NeurIPS 2025poster

Event cameras offer unique advantages in scenarios involving high speed, low light, and high dynamic range, yet their asynchronous and sparse nature poses significant challenges to efficient spatiotemporal representation learning. Specifically, despite notable progress in the field, effectively mode…

Cited by 0SourceScholar
2025

Graph Mixture of Experts and Memory-augmented Routers for Multivariate Time Series Anomaly Detection

AAAI 2025technical

Multivariate time series (MTS) anomaly detection is a critical task that involves identifying abnormal patterns or events in data that consist of multiple interrelated time series. In order to better model the complex interdependence between entities and the various inherent characteristics of each…

2025

Nonlinear Viscoelastic Model-based Deformation Optimization for Robotic Micropuncture in Retinal Vein Cannulation

IROS 2025

Micropuncture is a critical step in drug injection during retinal vein cannulation (RVC) surgery. Minimizing deformation during the micropuncture process is beneficial to reduce mechanical damage. However, this goal is challenging due to the viscoelastic characteristics of retinal tissue. In this pa

Cited by 0SourceScholar
2025

PathwiseRAG: Multi-Dimensional Exploration and Integration Framework

EMNLP 2025

Conventional retrieval-augmented generation(RAG) systems employ rigid retrieval strategies that create: (1) knowledge blind spots across domain boundaries, (2) reasoning fragmentation when processing interdependent concepts, and (3) contradictions from conflicting evidence sources. Motivated by thes

Cited by 0SourcePDFScholar
2025

Solid-SQL: Enhanced Schema-linking based In-context Learning for Robust Text-to-SQL

COLING 2025main

Recently, large language models (LLMs) have significantly improved the performance of text-to-SQL systems. Nevertheless, many state-of-the-art (SOTA) approaches have overlooked the critical aspect of system robustness. Our experiments reveal that while LLM-driven methods excel on standard datasets,…

2025

Toward Structured Knowledge Reasoning: Contrastive Retrieval-Augmented Generation on Experience

ACL 2025finding

Large language models (LLMs) achieve strong performance on plain text tasks but underperform on structured data like tables and databases. Potential challenges arise from their underexposure during pre-training and rigid text-to-structure transfer mechanisms. Unlike humans who seamlessly apply learn…

Cited by 0SourcePDFScholar
2025

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding

ICCV 2025accepted

Human intelligence requires both correctness and robustness, with the former being foundational for the latter. In video understanding, correctness ensures the accurate interpretation of visual content, and robustness maintains consistent performance in challenging conditions. Despite advances in vi…

Cited by 0SourcePDFScholar
2024

Decomposition for Enhancing Attention: Improving LLM-based Text-to-SQL through Workflow Paradigm

ACL 2024findings

In-context learning of large-language models (LLMs) has achieved remarkable success in the field of natural language processing, while extensive case studies reveal that the single-step chain-of-thought prompting approach faces challenges such as attention diffusion and inadequate performance in com…

2024

Design of Human-Machine Compatible Ankle Rehabilitation Robot Based on Equivalent Human Ankle Model

RA-L 2024

In this letter, a human–machine compatible ankle rehabilitation robot (HMCARR) is proposed to help stroke patients with motion dysfunction recover their motor function. The HMCARR can make the human ankle center-of-rotation (H-CoR) and the ankle rehabilitation robot center-of-rotation (R-CoR) coinci

Cited by 6SourceScholar
2024

EGSST: Event-based Graph Spatiotemporal Sensitive Transformer for Object Detection

NeurIPS 2024poster

Event cameras provide exceptionally high temporal resolution in dynamic vision systems due to their unique event-driven mechanism. However, the sparse and asynchronous nature of event data makes frame-based visual processing methods inappropriate. This study proposes a novel framework, Event-based G…

Cited by 0SourcePDFScholar
2024

Gradual Residuals Alignment: A Dual-Stream Framework for GAN Inversion and Image Attribute Editing

AAAI 2024technical

GAN-based image attribute editing firstly leverages GAN Inversion to project real images into the latent space of GAN and then manipulates corresponding latent codes. Recent inversion methods mainly utilize additional high-bit features to improve image details preservation, as low-bit codes cannot f…

Cited by 3SourcePDFScholar
2024

Learning Cortico-Muscular Dependence through Orthonormal Decomposition of Density Ratios

NeurIPS 2024poster

The cortico-spinal neural pathway is fundamental for motor control and movement execution, and in humans it is typically studied using concurrent electroencephalography (EEG) and electromyography (EMG) recordings. However, current approaches for capturing high-level and contextual connectivity betwe…

2024

RESEMO: A Benchmark Chinese Dataset for Studying Responsive Emotion from Social Media Content

ACL 2024findings

On social media platforms, users’ emotions are triggered when they encounter particular content from other users,where such emotions are different from those that spontaneously emerged, owing to the “responsive” nature. Analyzing the aforementioned responsive emotions from user interactions is a tas…

Cited by 0SourcePDFScholar
2023

A Soma Segmentation Benchmark in Full Adult Fly Brain

CVPR 2023poster

Neuron reconstruction in a full adult fly brain from high-resolution electron microscopy (EM) data is regarded as a cornerstone for neuroscientists to explore how neurons inspire intelligence. As the central part of neurons, somas in the full brain indicate the origin of neurogenesis and neural func…

2023

End-to-end Aspect-based Sentiment Analysis with Combinatory Categorial Grammar

ACL 2023findings

End-to-end Aspect-based Sentiment Analysis (EASA) is a natural language processing (NLP) task that involves extracting aspect terms and identifying the sentiments for them, which provides a fine-grained level of text analysis and thus requires a deep understanding of the running text. Many previous…

2022

Boundary-Aware Bias Loss for Transformer-Based Aerial Image Segmentation Model

ICASSP 2022accepted

Inspired by the tremendous success of the transformer-based model in natural language processing (NLP), many efforts introduce the transformer-based model into the image processing tasks. However, naive transformer models have to down-sample the image resolution to satisfy computational restrictions…

Cited by 0SourceScholar
2022

Tenrec: A Large-scale Multipurpose Benchmark Dataset for Recommender Systems

NeurIPS 2022accept

Existing benchmark datasets for recommender systems (RS) either are created at a small scale or involve very limited forms of user feedback. RS models evaluated on such datasets often lack practical values for large-scale real-world applications. In this paper, we describe Tenrec, a novel and publ…

2021

Domain Generalization under Conditional and Label Shifts via Variational Bayesian Inference

IJCAI 2021poster

In this work, we propose a domain generalization (DG) approach to learn on several labeled source domains and transfer knowledge to a target domain that is inaccessible in training. Considering the inherent conditional and label shifts, we would expect the alignment of p(x|y) and p(y). However, the…

Cited by 33SourcePDFScholar
2021

Regularized Recovery by Multi-Order Partial Hypergraph Total Variation

ICASSP 2021accepted

Capturing complex high-order interactions among data is an important task in many scenarios. A common way to model high-order interactions is to use hypergraphs whose topology can be mathematically represented by tensors. Existing methods use a fixed-order tensor to describe the topology of the whol…

Cited by 0SourceScholar
2021

Subtype-aware Unsupervised Domain Adaptation for Medical Diagnosis

AAAI 2021technical

Recent advances in unsupervised domain adaptation (UDA) show that transferable prototypical learning presents a powerful means for class conditional alignment, which encourages the closeness of cross-domain class centroids. However, the cross-domain inner-class compactness and the underlying fine-gr…

2020

Speeding up Very Fast Decision Tree with Low Computational Cost

IJCAI 2020poster

Very Fast Decision Tree (VFDT) is one of the most widely used online decision tree induction algorithms, and it provides high classification accuracy with theoretical guarantees. In VFDT, the split-attempt operation is essential for leaf-split. It is computation-intensive since it computes the heuri…