← Search

Sheng Guo

32 accepted papers

2026

IMPACT: Behavioral Intention-Aware Multimodal Trajectory Prediction With Adaptive Context Trimming

RA-L 2026

This paper presents a unified framework that jointly predicts behavioral intentions and vectorized occupancy, leveraging them as priors to dynamically prune context information during trajectory decoding, thereby enhancing prediction accuracy, interpretability, and efficiency. While most prior work

Cited by 4SourceScholar
2026

IMPACT: Behavioral Intention-Aware Multimodal Trajectory Prediction with Adaptive Context Trimming

ICRA 2026poster

This paper presents a unified framework that jointly predicts behavioral intentions and vectorized occupancy, leveraging them as priors to dynamically prune context information during trajectory decoding, thereby enhancing prediction accuracy, interpretability, and efficiency. While most prior work …

2026

SafeSearch: Automated Red-Teaming of LLM-Based Search Agents

ICML 2026poster

Search agents connect LLMs to the Internet, enabling them to access broader and more up-to-date information. However, this also introduces a new threat surface: unreliable search results can mislead agents into producing unsafe outputs. Real-world incidents and our two in-the-wild observations show …

Cited by 0SourceScholar
2026

Stop Unnecessary Reflection: Training LRMs for Efficient Reasoning with Adaptive Reflection and Length Coordinated Penalty

ICLR 2026poster

Large Reasoning Models (LRMs) have demonstrated remarkable performance on complex reasoning tasks by employing test-time scaling. However, they often generate over-long chains-of-thought that, driven by substantial reflections such as repetitive self-questioning and circular reasoning, lead to high…

Cited by 0SourcecodeScholar
2026

Trimming the Fat: Redundancy-Aware Acceleration Framework for DGNNs

AAAI 2026technical

Temporal graphs are essential for modeling complex real-world systems, such as social interactions, financial transactions, and recommendation systems, but the high computational cost and model complexity of dynamic graph neural networks (DGNNs) pose significant challenges for practical deployment.

Cited by 0SourcePDFScholar
2025

Enhancing Document Understanding with Group Position Embedding: A Novel Approach to Incorporate Layout Information

ICLR 2025poster

Recent advancements in document understanding have been dominated by leveraging large language models (LLMs) and multimodal large models. However, enabling LLMs to comprehend complex document layouts and structural information often necessitates intricate network modifications or costly pre-training…

2025

LeTS: Learning to Think-and-Search via Process-and-Outcome Reward Hybridization

EMNLP 2025

Large language models (LLMs) have demonstrated impressive capabilities in reasoning with the emergence of reasoning models like OpenAI-o1 and DeepSeek-R1. Recent research focuses on integrating reasoning capabilities into the realm of retrieval-augmented generation (RAG) via outcome-supervised reinf

2025

MobileViCLIP: An Efficient Video-Text Model for Mobile Devices

ICCV 2025poster

Efficient lightweight neural networks have received increasing attention due to their faster reasoning speed and easier deployment on mobile devices. However, existing video models still focus on the larger ViT architecture, and few works attempt to build efficient architecture. Since many efficient…

2025

OmniKV: Dynamic Context Selection for Efficient Long-Context LLMs

ICLR 2025poster

During the inference phase of Large Language Models (LLMs) with long context, a substantial portion of GPU memory is allocated to the KV cache, with memory usage increasing as the sequence length grows. To mitigate the GPU memory footprint associate with KV cache, some previous studies have discarde…

2024

3D Affordance Keypoint Detection for Robotic Manipulation

IROS 2024poster

This paper presents a novel approach for affordance-informed robotic manipulation by introducing 3D keypoints to enhance the understanding of object parts’ functionality. The proposed approach provides direct information about what the potential use of objects is, as well as guidance on where and ho…

Cited by 0SourceScholar
2024

Mind the Boundary: Coreset Selection via Reconstructing the Decision Boundary

ICML 2024poster

Existing paradigms of pushing the state of the art require exponentially more training data in many fields. Coreset selection seeks to mitigate this growing demand by identifying the most efficient subset of training data. In this paper, we delve into geometry-based coreset methods and preliminarily…

Cited by 11SourcePDFScholar
2024

Strain-based Modeling of Rod-driven Soft Continuum Robots with Co-located Embedded Sensors

IROS 2024poster

Rod-driven soft robots (RDSR) with a well-balanced performance in terms of perception, precision, and intelligence have a great potential for application. Mathematical description and predicted sensing of deformable soft bodies are crucial to achieve controllable and intelligent behaviors of these r…

Cited by 0SourceScholar
2023

CoMAE: Single Model Hybrid Pre-training on Small-Scale RGB-D Datasets

AAAI 2023technical

Current RGB-D scene recognition approaches often train two standalone backbones for RGB and depth modalities with the same Places or ImageNet pre-training. However, the pre-trained depth network is still biased by RGB-based models which may result in a suboptimal solution. In this paper, we present…

2023

Design, Simulation and Kinematic Verification of a Multi-Loop Ankle-Foot Prosthetic Mechanism

RA-L 2023

Inspired by the bionic characteristics of ankle and calf skeletal muscles, a novel ankle-foot prosthesis (AFP) with variable stiffness mechanisms (VSMs) is proposed to assist transtibial amputees to restore ankle plantarflexion-dorsiflexion. The prosthesis is designed in the form of a spring-loaded

Cited by 3SourceScholar
2023

Efficient Training of Large-Scale Industrial Fault Diagnostic Models through Federated Opportunistic Block Dropout

AAAI 2023technical

Artificial intelligence (AI)-empowered industrial fault diagnostics is important in ensuring the safe operation of industrial applications. Since complex industrial systems often involve multiple industrial plants (possibly belonging to different companies or subsidiaries) with sensitive data collec…

Cited by 7SourcePDFScholar
2023

Learning 3D-Aware Image Synthesis With Unknown Pose Distribution

CVPR 2023poster

Existing methods for 3D-aware image synthesis largely depend on the 3D pose distribution pre-estimated on the training set. An inaccurate estimation may mislead the model into learning faulty geometry. This work proposes PoF3D that frees generative radiance fields from the requirements of 3D pose pr…

2023

MHSCNET: A Multimodal Hierarchical Shot-Aware Convolutional Network for Video Summarization

ICASSP 2023accepted

Video summarization is an essential problem in signal processing, which intends to produce a concise summary of the original video. Existing video summarization approaches regard the task as a keyframe selection problem and generally construct the frame-wise representation by combining the long-rang…

Cited by 0SourceScholar
2023

PDPP:Projected Diffusion for Procedure Planning in Instructional Videos

CVPR 2023highlight

In this paper, we study the problem of procedure planning in instructional videos, which aims to make goal-directed plans given the current visual observations in unstructured real-life videos. Previous works cast this problem as a sequence planning problem and leverage either heavy intermediate vis…

2023

StageInteractor: Query-based Object Detector with Cross-stage Interaction

ICCV 2023poster

Previous object detectors make predictions based on dense grid points or numerous preset anchors. Most of these detectors are trained with one-to-many label assignment strategies. On the contrary, recent query-based object detectors are based a sparse set of learnable queries refined by a series of…

Cited by 12PDFcodeScholar
2022

Cross-Architecture Self-Supervised Video Representation Learning

CVPR 2022poster

In this paper, we present a new cross-architecture contrastive learning (CACL) framework for self-supervised video representation learning. CACL consists of a 3D CNN and a video transformer which are used in parallel to generate diverse positive pairs for contrastive learning. This allows the model…

Cited by 31PDFcodeScholar
2022

Design and Analysis of a Novel Variable Stiffness Continuum Robot With Built-in Winding-Styled Ropes

RA-L 2022

Continuum robots driven by rods have a wide range of applications, such as detection and maintenance tasks in unstructured environments. However, their inherent nature of flexibility also limits their function. Thus, variable stiffness mechanisms for continuum robots have consistently attracted the

Cited by 34SourceScholar
2022

Design and Experimental Characterization of a Push-Pull Flexible Rod-Driven Soft-Bodied Robot

RA-L 2022

Soft robots with a well-balanced performance in terms of dexterity, accuracy, and payload have a great potential for application. Balancing safe human-robot interaction with operation performance enables the use of soft robot in biomedical fields, among others, such as surgery, rehabilitation and el

Cited by 30SourceScholar
2022

InsCLR: Improving Instance Retrieval with Self-Supervision

AAAI 2022technical

This work aims at improving instance retrieval with self-supervision. We find that fine-tuning using the recently developed self-supervised learning (SSL) methods, such as SimCLR and MoCo, fails to improve the performance of instance retrieval. In this work, we identify that the learnt representatio…

2022

RGL: A Simple yet Effective Relation Graph Augmented Prompt-based Tuning Approach for Few-Shot Learning

NAACL 2022findings

Pre-trained language models (PLMs) can provide a good starting point for downstream applications. However, it is difficult to generalize PLMs to new tasks given a few labeled samples. In this work, we show that Relation Graph augmented Learning (RGL) can improve the performance of few-shot natural l…

2021

Unchain the Search Space with Hierarchical Differentiable Architecture Search

AAAI 2021technical

Differentiable architecture search (DAS) has made great progress in searching for high-performance architectures with reduced computational cost. However, DAS-based methods mainly focus on searching for a repeatable cell structure, which is then stacked sequentially in multiple stages to form the n…

2020

Representation Sharing for Fast Object Detector Search and Beyond

ECCV 2020poster

Region Proposal Network (RPN) provides strong support for handling the scale variation of objects in two-stage object detection. For one-stage detectors which do not have RPN, it is more demanding to have powerful sub-networks capable of directly capturing objects of unknown sizes. To enhance such c…

2020

V4D: 4D Convolutional Neural Networks for Video-level Representation Learning

ICLR 2020poster

Most existing 3D CNN structures for video representation learning are clip-based methods, and do not consider video-level temporal evolution of spatio-temporal features. In this paper, we propose Video-level 4D Convolutional Neural Networks, namely V4D, to model the evolution of long-range spatio-te…

Cited by 123SourceScholar
2019

Decoupling Category-wise Independence and Relevance with Self-attention for Multi-label Image Classification

ICASSP 2019accepted

Multi-label image classification has achieved remarkable progress thanks to deep convolutional neural networks (CNNs). In this paper, we propose a Decouple Network (DecoupleNet) which is an end-to-end CNN-based framework able to trade off class-level feature independence and relevance during trainin…

Cited by 0SourceScholar
2019

Label-PEnet: Sequential Label Propagation and Enhancement Networks for Weakly Supervised Instance Segmentation

ICCV 2019poster

Weakly-supervised instance segmentation aims to detect and segment object instances precisely, given image-level labels only. Unlike previous methods which are composed of multiple offline stages, we propose Sequential Label Propagation and Enhancement Networks (referred as Label-PEnet) that progres…

Cited by 67PDFScholar
2018

CurriculumNet: Weakly Supervised Learning from Large-Scale Web Images

ECCV 2018poster

We present a simple yet efficient approach capable of training deep neural networks on large-scale weakly-supervised web images, which are crawled rawly from the Internet by using text queries, without any human annotation. We develop a principled learning strategy by leveraging curriculum learning,…