← Search

Liqing Gao

8 accepted papers

2026

iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models

ICLR 2026poster

Recent methods have made notable progress in accelerating Large Vision-Language Models (LVLMs) by exploiting the inherent redundancy in visual inputs. Most existing approaches, however, focus narrowly on reducing image tokens before or within the Large Language Model (LLM) stage to lower computation…

Cited by 0SourcecodeScholar
2025

GReg: Geometry-Aware Region Refinement for Sign Language Video Generation

ICCV 2025poster

Sign Language Video Generation (SLVG) aims to transform sign language sequences into natural and fluent sign language videos. Existing SLVG methods lack geometric modeling of human anatomical structures, leading to anatomically implausible and temporally inconsistent generation. To address these cha…

Cited by 0SourcePDFScholar
2024

Dynamic Spatial-Temporal Aggregation for Skeleton-Aware Sign Language Recognition

COLING 2024main

Skeleton-aware sign language recognition (SLR) has gained popularity due to its ability to remain unaffected by background information and its lower computational requirements. Current methods utilize spatial graph modules and temporal modules to capture spatial and temporal features, respectively.…

2023

Continuous Sign Language Recognition With Correlation Network

CVPR 2023poster

Human body trajectories are a salient cue to identify actions in video. Such body trajectories are mainly conveyed by hands and face across consecutive frames in sign language. However, current methods in continuous sign language recognition(CSLR) usually process frames independently to capture fram…

2023

Self-Emphasizing Network for Continuous Sign Language Recognition

AAAI 2023technical

Hand and face play an important role in expressing sign language. Their features are usually especially leveraged to improve system performance. However, to effectively extract visual representations and capture trajectories for hands and face, previous methods always come at high computations with…

2022

Temporal Lift Pooling for Continuous Sign Language Recognition

ECCV 2022poster

"Pooling methods are necessities for modern neural networks for increasing receptive fields and lowering down computational costs. However, commonly used hand-crafted pooling approaches, e.g. max pooling and average pooling, may not well preserve discriminative features. While many researchers have…

2020

Key Action and Joint CTC-Attention based Sign Language Recognition

ICASSP 2020accepted

Sign Language Recognition (SLR) translates sign language video into natural language. In practice, sign language video, owning a large number of redundant frames, is necessary to be selected the essential. However, unlike common video that describes actions, sign language video is characterized as c…

Cited by 0SourceScholar