← Search

Shengjie Zhao

11 accepted papers

2026

DEPO: Dual-Efficiency Preference Optimization for LLM Agents

AAAI 2026technical

Recent advances in large language models (LLMs) have greatly improved their reasoning and decision-making abilities when deployed as agents. Richer reasoning, however, often comes at the cost of longer chain of thought (CoT), hampering interaction efficiency in real-world scenarios. Nevertheless, th

Cited by 0SourcePDFScholar
2026

GenCape: Structure-Inductive Generative Modeling for Category-Agnostic Pose Estimation

ICLR 2026poster

Category-agnostic pose estimation (CAPE) aims to localize keypoints on query images from arbitrary categories, using only a few annotated support examples for guidance. Recent approaches either treat keypoints as isolated entities or rely on manually defined skeleton priors, which are costly to anno…

Cited by 0SourceScholar
2026

SiMO: Single-Modality-Operable Multimodal Collaborative Perception

ICLR 2026poster

Collaborative perception integrates multi-agent perspectives to enhance the sensing range and overcome occlusion issues. While existing multimodal approaches leverage complementary sensors to improve performance, they are highly prone to failure—especially when a key sensor like LiDAR is unavailable…

Cited by 0SourcecodeScholar
2026

Uncovering Pretraining Code in LLMs: A Syntax-Aware Attribution Approach

AAAI 2026technical

As large language models (LLMs) become increasingly capable, concerns over the unauthorized use of copyrighted and licensed content in their training data have grown, especially in the context of code. Open-source code, often protected by open source licenses (e.g, GPL), poses legal and ethical chal

Cited by 0SourcePDFScholar
2025

From Imitation to Introspection: Probing Self-Consciousness in Language Models

ACL 2025finding

Self-consciousness, the introspection of one’s existence and thoughts, represents a high-level cognitive process. As language models advance at an unprecedented pace, a critical question arises: Are these models becoming self-conscious? Drawing upon insights from psychological and neural science, th…

2025

Point4Bit: Post Training 4-bit Quantization for Point Cloud 3D Detection

NeurIPS 2025poster

Voxel-based 3D object detectors have achieved remarkable performance in point cloud perception, yet their high computational and memory demands pose significant challenges for deployment on resource-constrained edge devices. Post-training quantization (PTQ) provides a practical means to compress mod…

Cited by 0SourceScholar
2025

ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models

NeurIPS 2025poster

Recent Vision-Language Models (VLMs) have shown strong performance in general-purpose visual understanding and reasoning, but their ability to comprehend the visual grammar of movie shots remains underexplored and insufficiently evaluated. To bridge this gap, we present \textbf{ShotBench}, a dedicat…

Cited by 0SourceScholar
2024

CLEAR: Can Language Models Really Understand Causal Graphs?

EMNLP 2024finding

Causal reasoning is a cornerstone of how humans interpret the world. To model and reason about causality, causal graphs offer a concise yet effective solution. Given the impressive advancements in language models, a crucial question arises: can they really understand causal graphs? To this end, we p…

2024

SEA-GNN: Sequence Extension Augmented Graph Neural Network for Sequential Recommendation

ICASSP 2024accepted

Sequential recommendation aims to anticipate the next preference of users by examining their recent interactions. Recently, graph neural networks (GNNs) have been widely utilized in sequential recommendation, but existing schemes focus on interactions within individual sequences and tend to connect…

Cited by 0SourceScholar