← Search

Shu Wang

22 accepted papers

2026

ArchRAG: Attributed Community-based Hierarchical Retrieval-Augmented Generation

AAAI 2026technical

Retrieval-Augmented Generation (RAG) has proven effective in integrating external knowledge into large language models (LLMs) for solving question-answer (QA) tasks. The state-of-the-art RAG approaches often use the graph data as the external data since they capture the rich semantic information and

Cited by 0SourcePDFScholar
2026

EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn Editing

ICLR 2026poster

Instruction-based image editing has advanced rapidly, yet reliable and interpretable evaluation remains a bottleneck. Current protocols either (i) depend on paired reference images—resulting in limited coverage and inheriting biases from prior generative models—or (ii) rely *solely* on zero-shot vis…

Cited by 0SourcecodeScholar
2026

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models

CVPR 2026

Vision-Language-Action (VLA) models have recently enabled robotic manipulation by grounding visual and linguistic cues into actions. However, most VLAs assume the Markov property, relying only on the current observation and thus suffering from temporal myopia that degrades long-horizon coherence. In

Cited by 0SourcecodeScholar
2026

Integrated Exploration and Sequential Manipulation on Scene Graph with LLM-Based Situated Replanning

ICRA 2026poster

In partially known environments, robots must combine exploration to gather information with task planning for efficient execution. To address this challenge, we propose EPoG, an Exploration-based sequential manipulation Planning framework on Graph-based representations. EPoG integrates a graph-based…

2026

Read the Room: Video Social Reasoning with Mental-Physical Causal Chains

ICLR 2026poster

``Read the room,'' or the ability to infer others' mental states from subtle social cues, is a hallmark of human social intelligence but remains a major challenge for current AI systems. Existing social reasoning datasets are limited in complexity, scale, and coverage of mental states, falling short…

Cited by 0SourcecodeScholar
2026

TS-PEFT: Unveiling Token-Level Redundancy in Parameter-Efficient Fine-Tuning

IJCAI 2026

Current Parameter-Efficient Fine-Tuning (PEFT) methods typically operate under an implicit assumption: once a target module is selected, every token passing through it contributes equally to the downstream task and requires a parameter update. In this paper, we challenge this convention by revealing

Cited by 0Scholar
2025

Explore the Reasoning Capability of LLMs in the Chess Testbed

NAACL 2025short

Reasoning is a central capability of human intelligence. In recent years, with the advent of large-scale datasets, pretrained large language models have emerged with new capabilities, including reasoning. However, these models still struggle with long-term, complex reasoning tasks, such as playing c…

Cited by 1SourcePDFScholar
2025

Keep the Balance: A Parameter-Efficient Symmetrical Framework for RGB+X Semantic Segmentation

CVPR 2025poster

Multimodal semantic segmentation is a critical challenge in computer vision, with early methods suffering from high computational costs and limited transferability due to full fine-tuning of RGB-based pre-trained parameters. Recent studies, while leveraging additional modalities as supplementary pro…

Cited by 0SourcePDFScholar
2025

LVPTrack: High Performance Domain Adaptive UAV Tracking with Label Aligned Visual Prompt Tuning

AAAI 2025technical

Visual object tracking is essentially crucial for unmanned aerial vehicles (UAVs). Despite the substantial progress, most of the existing UAV trackers are designed for well-conditioned daytime data, while for the scenarios in challenging weather condition, e.g. foggy or nighttime environment, the tr…

Cited by 0SourcePDFScholar
2025

Latent Thought Models with Variational Bayes Inference-Time Computation

ICML 2025poster

We propose a novel class of language models, Latent Thought Models (LTMs), which incorporate explicit latent thought vectors that follow an explicit prior model in latent space. These latent thought vectors guide the autoregressive generation of ground tokens through a Transformer decoder. Training…

2025

Look Both Ways and No Sink: Converting LLMs into Text Encoders without Training

ACL 2025long

Recent advancements have demonstrated the advantage of converting pretrained large language models into powerful text encoders by enabling bidirectional attention in transformer layers. However, existing methods often require extensive training on large-scale datasets, posing challenges in low-resou…

2025

MetricGrids: Arbitrary Nonlinear Approximation with Elementary Metric Grids based Implicit Neural Representation

CVPR 2025highlight

This paper presents MetricGrids, a novel grid-based neural representation that combines elementary metric grids in various metric spaces to approximate complex nonlinear signals. While grid-based representations are widely adopted for their efficiency and scalability, the existing feature grids with…

2025

Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing

ICML 2025poster

*Warning: Contains harmful model outputs.* Despite significant advancements, the propensity of Large Language Models (LLMs) to generate harmful and unethical content poses critical challenges. Measuring value alignment of LLMs becomes crucial for their regulation and responsible deployment. Althoug…

Cited by 5SourcePDFScholar
2025

Semi-Supervised State-Space Model with Dynamic Stacking Filter for Real-World Video Deraining

CVPR 2025poster

Significant progress has been made in video restoration under rainy conditions over the past decade, largely propelled by advancements in deep learning. Nevertheless, existing methods that depend on paired data struggle to generalize effectively to real-world scenarios, primarily due to the disparit…

Cited by 0SourcePDFScholar
2025

ToolGen: Unified Tool Retrieval and Calling via Generation

ICLR 2025poster

As large language models (LLMs) advance, their inability to autonomously execute tasks by directly interacting with external tools remains a critical limitation. Traditional methods rely on inputting tool descriptions as context, which is constrained by context length and requires separate, often in…

2025

Ultra-High-Definition Dynamic Multi-Exposure Image Fusion via Infinite Pixel Learning

AAAI 2025technical

With the continuous improvement of device imaging resolution, the popularity of Ultra-High-Definition (UHD) images is increasing. Unfortunately, existing methods for fusing multi-exposure images in dynamic scenes are designed for low-resolution images, which makes them inefficient for generating hig…

Cited by 0SourcePDFScholar
2024

Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action Representations

IROS 2024poster

Autonomous robotic systems capable of learning novel manipulation tasks are poised to transform industries from manufacturing to service automation. However, current methods (e.g., VIP and R3M) still face significant hurdles, notably the domain gap among robotic embodiments and the sparsity of succe…

Cited by 15SourcecodeScholar
2024

Fall Prediction by a Spatio-Temporal Multi-Channel Causal Model from Wearable Sensors Data

ICASSP 2024accepted

Predicting human falls from wearable devices is a complex task due to the inherent diversity and causality of multivariate physical changes, where each instance exhibits a unique style of motion events and their spatio-temporal causal dependencies. Consequently, we propose a multichannel causal mode…

Cited by 0SourceScholar
2024

LLM3: Large Language Model-based Task and Motion Planning with Motion Failure Reasoning

IROS 2024poster

Conventional Task and Motion Planning (TAMP) approaches rely on manually designed interfaces connecting symbolic task planning with continuous motion generation. These domain-specific and labor-intensive modules are limited in addressing emerging tasks in real-world settings. Here, we present LLM3,…

Cited by 45SourcecodeScholar
2024

Predicting Fall Events by a Spatio-Temporal Topological Network with Multiple Wearable Sensors

ICASSP 2024accepted

A key challenge in sensor-based fall prediction is the fact that a fall event can often occur in various configurations of fall poses together with their own spatio-temporal dependencies. This leads us to define a spatio-temporal model to explicitly characterize these internal configurations of pose…

Cited by 0SourceScholar
2022

T-NGA: Temporal Network Grafting Algorithm for Learning to Process Spiking Audio Sensor Events

ICASSP 2022accepted

Spiking silicon cochlea sensors encode sound as an asynchronous stream of spikes from different frequency channels. The lack of labeled training datasets for spiking cochleas makes it difficult to train deep neural networks on the outputs of these sensors. This work proposes a self-supervised method…

Cited by 0SourceScholar