← Search

Qingyu Zhang

12 accepted papers

2026

AI-Salesman: Towards Reliable Large Language Model Driven Telemarketing

AAAI 2026technical

Goal-driven persuasive dialogue, exemplified by applications like telemarketing, requires sophisticated multi-turn planning and strict factual faithfulness, which remains a significant challenge for even state-of-the-art Large Language Models (LLMs). A lack of task-specific data often limits previou

Cited by 0SourcePDFScholar
2026

FisherPoser: Human Motion Estimation from Sparse Observations with Hierarchical Region-Wise Fisher-Matrix Uncertainty Modeling

CVPR 2026

Full-body motion estimation from sparse VR observations is an inherently under-constrained problem, with only three 6-DoF trackers (HMD and controllers) available to infer a full skeletal pose. To address this ambiguity, we introduce a probabilistic framework that models joint orientations as distri

Cited by 0SourceScholar
2026

VideoSeeker: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

ICML 2026poster

Existing multimodal large language models for long-video understanding predominantly rely on uniform sampling and single-turn inference, limiting their ability to identify sparse yet critical evidence amid extensive redundancy. We introduce VideoSeeker, a novel framework that supports iterative disc…

Cited by 13SourceScholar
2025

ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

ACL 2025finding

As Large Language Models (LLMs) continue to advance, their computational overhead has increased significantly. In this study, we identify notable redundancy across the layers of LLMs, where some layers contribute minimally to the overall network functionality. To quantify this, we introduce a metric…

Cited by 0SourcePDFScholar
2025

StreamForest: Efficient Online Video Understanding with Persistent Event Memory

NeurIPS 2025spotlight

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in video understanding. However, their effectiveness in real-time streaming scenarios remains limited due to storage constraints of historical visual features and insufficient real-time spatiotemporal reasoning. To a…

Cited by 0SourceScholar
2025

exUMI: Extensible Robot Teaching System with Action-aware Task-agnostic Tactile Representation

CoRL 2025poster

Tactile-aware robot learning faces critical challenges in data collection and representation due to data scarcity and sparsity, and the absence of force feedback in existing systems. To address these limitations, we introduce a tactile robot learning system with both hardware and algorithm innovatio…

Cited by 0SourceScholar
2024

Base of RoPE Bounds Context Length

NeurIPS 2024poster

Position embedding is a core component of current Large Language Models (LLMs). Rotary position embedding (RoPE), a technique that encodes the position information with a rotation matrix, has been the de facto choice for position embedding in many LLMs, such as the Llama series. RoPE has been furthe…

Cited by 10SourcePDFScholar
2024

MEVTR: A Multilingual Model Enhanced with Visual Text Representations

COLING 2024main

The goal of multilingual modelling is to generate multilingual text representations for various downstream tasks in different languages. However, some state-of-the-art pre-trained multilingual models perform poorly on many low-resource languages due to the lack of representation space and model capa…

2023

Magnet Array-Actuated Steerable Flexible Robot with Beacon- TfmUltrasonic Position Sensing for Robotic Neurosurgery

IROS 2023

A magnetically controlled steerable robot with the capability of flexible navigation and intraoperative ultrasonic imaging and position sensing in neurosurgery is introduced in this paper. The robot system uses a piezoelectric transducer as the ultrasonic imaging beacon and applies a permanent magne

Cited by 1SourceScholar
2023

Magnet Array-Actuated Steerable Flexible Robot with Beacon-TFM Ultrasonic Position Sensing for Robotic Neurosurgery

IROS 2023poster

A magnetically controlled steerable robot with the capability of flexible navigation and intraoperative ultrasonic imaging and position sensing in neurosurgery is introduced in this paper. The robot system uses a piezoelectric transducer as the ultrasonic imaging beacon and applies a permanent magne…

Cited by 0SourceScholar
2020

Few-shot Visual Learning with Contextual Memory and Fine-grained Calibration

IJCAI 2020poster

Few-shot learning aims to learn a model that can be readily adapted to new unseen classes (concepts) by accessing one or few examples. Despite the successful progress, most of the few-shot learning approaches, concentrating on either global or local characteristics of examples, still suffer from wea…

Cited by 0SourcePDFScholar