← Search

Chong Li

20 accepted papers

2026

DeepGB-TB: A Risk-Balanced Cross-Attention Gradient-Boosted Convolutional Network for Rapid, Interpretable Tuberculosis Screening

AAAI 2026technical

Large-scale tuberculosis (TB) screening is limited by the high cost and operational complexity of traditional diagnostics, creating a need for artificial-intelligence solutions. We propose DeepGB-TB, a non-invasive system that instantly assigns TB risk scores using only cough audio and basic demogra

Cited by 0SourcePDFScholar
2026

Subspace-Aware Graph Construction and Contrastive Alignment for Multimodal Recommendation with Large Language Models

AAAI 2026technical

Multimedia content offers additional context for recommender systems to better understand user interests. Existing studies on multimodal recommendation primarily focus on constructing item-item semantic graphs. However, most of these methods capture only shallow semantic structures based on feature

Cited by 0SourcePDFScholar
2025

Discovering Semantic Subdimensions through Disentangled Conceptual Representations

EMNLP 2025

Understanding the core dimensions of conceptual semantics is fundamental to uncovering how meaning is organized in language and the brain. Existing approaches often rely on predefined semantic dimensions that offer only broad representations, overlooking finer conceptual distinctions. This paper pro

Cited by 0SourcePDFScholar
2025

Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model

ACL 2025finding

The curse of multilinguality phenomenon is a fundamental problem of multilingual Large Language Models (LLMs), where the competition between massive languages results in inferior performance. It mainly comes from limited capacity and negative transfer between dissimilar languages. To address this is…

Cited by 0SourcePDFScholar
2025

Hybrid Data-Model-Driven External Force Estimation for Manipulators via Generalized Momentum-Based Third-Order Observer*

IROS 2025

Accurate dynamic modeling and external force estimation are crucial for high-precision robot control and applications. However, model incompleteness and external disturbances inevitably lead to a residual between the actual joint torque and the torque calculated by the identified dynamic model. To a

Cited by 0SourceScholar
2025

LLM×MapReduce: Simplified Long-Sequence Processing using Large Language Models

ACL 2025long

We propose a training-free framework that enables large language models (LLMs) to effectively process long texts, using a divide-and-conquer strategy for comprehensive document understanding.The proposed LLM×MapReduce framework splits the entire document into several chunks for LLMs to read and then…

Cited by 0SourcePDFScholar
2025

SEP: A General Lossless Compression Framework with Semantics Enhancement and Multi-Stream Pipelines

IJCAI 2025

Deep-learning-based lossless compression is of immense importance in real-world applications, such as cold data persistence, sensor data collection, and astronomical data transmission. However, existing compressors typically model data using single-byte symbols as tokens, which makes it hard to capt

2024

Improving In-context Learning of Multilingual Generative Language Models with Cross-lingual Alignment

NAACL 2024long

Multilingual generative models obtain remarkable cross-lingual in-context learning capabilities through pre-training on large-scale corpora. However, they still exhibit a performance bias toward high-resource languages and learn isolated distributions of multilingual sentence representations, which…

2024

NeuroPictor: Refining fMRI-to-Image Reconstruction via Multi-individual Pretraining and Multi-level Modulation

ECCV 2024poster

"Recent fMRI-to-image approaches mainly focused on associating fMRI signals with specific conditions of pre-trained diffusion models. These approaches, while producing high-quality images, capture only a limited aspect of the complex information in fMRI signals and offer little detailed control over…

2024

VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time

NeurIPS 2024oral

We introduce VASA, a framework for generating lifelike talking faces with appealing visual affective skills (VAS) given a single static image and a speech audio clip. Our premiere model, VASA-1, is capable of not only generating lip movements that are exquisitely synchronized with the audio, but als…

Cited by 92SourcePDFScholar
2024

X-Instruction: Aligning Language Model in Low-resource Languages with Self-curated Cross-lingual Instructions

ACL 2024findings

Large language models respond well in high-resource languages like English but struggle in low-resource languages. It may arise from the lack of high-quality instruction following data in these languages. Directly translating English samples into these languages can be a solution but unreliable, lea…

2023

FreeEnricher: Enriching Face Landmarks without Additional Cost

AAAI 2023technical

Recent years have witnessed significant growth of face alignment. Though dense facial landmark is highly demanded in various scenarios, e.g., cosmetic medicine and facial beautification, most works only consider sparse face alignment. To address this problem, we present a framework that can enrich l…

Cited by 3SourcePDFScholar
2022

Frame-Wise Action Representations for Long Videos via Sequence Contrastive Learning

CVPR 2022poster

Prior works on action representation learning mainly focus on designing various architectures to extract the global representations for short video clips. In contrast, many practical applications such as video alignment have strong demand for learning dense representations for long videos. In this p…

Cited by 53PDFcodeScholar
2021

ADNet: Leveraging Error-Bias Towards Normal Direction in Face Alignment

ICCV 2021poster

The recent progress of CNN has dramatically improved face alignment performance. However, few works have paid attention to the error-bias with respect to error distribution of facial landmarks. In this paper, we investigate the error-bias issue in face alignment, where the distributions of landmark…

Cited by 69PDFcodeScholar
2021

Exploration and Exploitation: Two Ways to Improve Chinese Spelling Correction Models

ACL 2021short

A sequence-to-sequence learning with neural networks has empirically proven to be an effective framework for Chinese Spelling Correction (CSC), which takes a sentence with some spelling errors as input and outputs the corrected one. However, CSC models may fail to correct spelling errors covered by…

2018

Constrained Optimization Based Low-Rank Approximation of Deep Neural Networks

ECCV 2018poster

We present COBLA---Constrained Optimization Based Low-rank Approximation---a systematic method of finding an optimal low-rank approximation of a trained convolutional neural network, subject to constraints in the number of multiply-accumulate (MAC) operations and the memory footprint. COBLA optimall…

2018

Improving Colorectal Polyp Classification Based on Physical Examination Data - An Ensemble Learning Approach

RA-L 2018

Colorectal cancer is a common type of cancer. Due to the alarming incidence and mortality rate, it has received increasing attention on early detection and treatment. Colorectal polyps form and grow at initial stages of most colorectal cancer cases. Due to rather stringent medical resource availabil

Cited by 6SourceScholar