← Search

Nan Wang

32 accepted papers

2026

Active Perception Driven Tactile Modeling of Deformable Objects for Robots

RA-L 2026

Despite extensive research on tactile sensing, exploiting it to model and interpret the external world remains largely underexplored. We propose a unified framework that leverages vision-based tactile sensors to construct fine-grained object models, enhancing situational understanding. This frame wo

Cited by 0SourceScholar
2026

DGGT: Feedforward 4D Reconstruction of Dynamic Driving Scenes using Unposed Images

CVPR 2026

Autonomous driving needs fast, scalable 4D reconstruction and re-simulation for training and evaluation, yet most methods for dynamic driving scenes still rely on per-scene optimization, known camera calibration, or short frame windows, making them slow and impractical. We revisit this problem from

Cited by 0SourcecodeScholar
2026

Diffusion Knows Transparency: Repurposing Video Diffusion for Transparent Object Depth and Normal Estimation

ICRA 2026poster

Transparent objects remain notoriously hard for perception systems: refraction, reflection and transmission break the assumptions behind stereo, ToF and purely discriminative monocular depth, causing holes and temporally unstable estimates. Our key observation is that modern video diffusion models a…

2026

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding

CVPR 2026

While Multimodal Large Language Models (MLLMs) excel at single-image understanding, they exhibit significantly degraded performance in multi-image reasoning scenarios. Multi-image reasoning presents fundamental challenges including complex inter-relationships between images and scattered critical in

Cited by 0SourceScholar
2026

ORV: 4D Occupancy-centric Robot Video Generation

CVPR 2026

Recent embodied intelligence suffers from data scarcity, while conventional simulators lack visual realism. Controllable video generation is emerging as a promising data engine, yet current action-conditioned methods still fall short: generated videos are limited in fidelity and temporal consistency

Cited by 0SourcecodeScholar
2026

SparseCam4D: Spatio-Temporally Consistent 4D Reconstruction from Sparse Cameras

CVPR 2026

High-quality 4D reconstruction enables photorealistic and immersive rendering of the dynamic real world. However, unlike static scenes that can be fully captured with a single camera, high-quality dynamic scenes typically require dense arrays of tens or even hundreds of synchronized cameras. Depende

Cited by 0SourcecodeScholar
2025

A2ATS: Retrieval-Based KV Cache Reduction via Windowed Rotary Position Embedding and Query-Aware Vector Quantization

ACL 2025finding

Long context large language models (LLMs) pose significant challenges for efficient serving due to the large memory footprint and high access overhead of KV cache.Retrieval-based KV cache reduction methods can mitigate these challenges, typically by offloading the complete KV cache to CPU and retrie…

2025

AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark

ACL 2025long

Evaluation plays a crucial role in the advancement of information retrieval (IR) models. However, current benchmarks, which are based on predefined domains and human-labeled data, face limitations in addressing evaluation needs for emerging domains both cost-effectively and efficiently. To address t…

2025

Breaking the Discretization Barrier of Continuous Physics Simulation Learning

NeurIPS 2025poster

The modeling of complicated time-evolving physical dynamics from partial observations is a long-standing challenge. Particularly, observations can be sparsely distributed in a seemingly random or unstructured manner, making it difficult to capture highly nonlinear features in a variety of scientific…

Cited by 0SourcecodeScholar
2025

CT Image Prediction Of PD-1 Gastric Cancer Patients Based On The PLSG Framework

ICASSP 2025accepted

Survival prediction in PD-1 inhibitor patients has received extensive attention in recent years. Existing diffusion models generally focus blurring on key lesion regions, and the masks are weakly matched to CT images during the sampling process, resulting in a relatively high risk of misdiagnosis. F…

Cited by 0SourceScholar
2025

Exploring the Domain-Invariant Flow Representation in Vision-Based Tactile Sensors for Omni-Hardness Perception

ICRA 2025

Vision-based tactile sensors have recently gained prominence due to their superior resolution and ability to capture multi-dimensional contact information. However, even when sensors share the same sensing principle, variations in production factors can lead to differences in the color patterns of t

Cited by 0SourceScholar
2025

MedEureka: A Medical Domain Benchmark for Multi-Granularity and Multi-Data-Type Embedding-Based Retrieval

NAACL 2025findings

Embedding-based retrieval (EBR), the mainstream approach in information retrieval (IR), aims to help users obtain relevant information and plays a crucial role in retrieval-augmented generation (RAG) techniques of large language models (LLMs). Numerous methods have been proposed to significantly imp…

2025

MolParser: End-to-end Visual Recognition of Molecule Structures in the Wild

ICCV 2025poster

In recent decades, chemistry publications and patents have increased rapidly. A significant portion of key information is embedded in molecular structure figures, complicating large-scale literature searches and limiting the application of large language models in fields such as biology, chemistry,…

Cited by 0SourcePDFScholar
2025

One View, Many Worlds: Single-Image to 3D object Meets Generative Domain Randomization for One-Shot 6D Pose Estimation

CoRL 2025oral

Estimating the 6D pose of arbitrary objects from a single reference image is a critical yet challenging task in robotics, especially considering the long-tail distribution of real-world instances. While category-level and model-based approaches have achieved notable progress, they remain limited in…

Cited by 0SourceScholar
2025

PUGS: Zero-Shot Physical Understanding with Gaussian Splatting

ICRA 2025

Current robotic systems can understand the categories and poses of objects well. But understanding physical properties like mass, friction, and hardness, in the wild, remains challenging. We propose a new method that reconstructs 3D objects using the Gaussian splatting representation and predicts va

Cited by 11SourcecodeScholar
2025

Passive Adaptive Object Prehension, Retention, and Release With a Mechanically Intelligent Gripper

RA-L 2025

This letter presents a mechanically intelligent gripper that is capable of passive and adaptive object prehension, passive object retention, and passive object release. Passive adaptive prehension is achieved through a compliant linkage with two Fin Ray fingers that enclose an object when the grippe

Cited by 0SourceScholar
2025

RE0: Recognize Everything with 3D Zero-Shot Instance Segmentation

ICRA 2025

Recognizing objects in the 3D world is a significant challenge for robotics. Due to the lack of high-quality 3D data, directly training a general-purpose segmentation model in 3D is almost infeasible. Meanwhile, vision foundation models (VFM) have revolutionized the 2D computer vision field with out

Cited by 1SourcecodeScholar
2025

StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation

CVPR 2025poster

Recent advances in large reconstruction and generative models have significantly improved scene reconstruction and novel view generation. However, due to compute limitations, each inference with these large models is confined to a small area, making long-range consistent scene generation challenging…

Cited by 1SourcePDFScholar
2025

Unifying Appearance Codes and Bilateral Grids for Driving Scene Gaussian Splatting

NeurIPS 2025poster

Neural rendering techniques, including NeRF and Gaussian Splatting (GS), rely on photometric consistency to produce high-quality reconstructions. However, in real-world driving scenarios, it is challenging to guarantee perfect photometric consistency in acquired images. Appearance codes have been wi…

Cited by 0SourcecodeScholar
2024

Extracting Financial Events from Raw Texts via Matrix Chunking

COLING 2024main

Event Extraction (EE) is widely used in the Chinese financial field to provide valuable structured information. However, there are two key challenges for Chinese financial EE in application scenarios. First, events need to be extracted from raw texts, which sets it apart from previous works like the…

Cited by 1SourcePDFScholar
2024

Omnidirectional Dense SLAM for Back-to-back Fisheye Cameras

ICRA 2024poster

We propose a real-time visual-inertial dense SLAM system that utilizes the online data streams from back-to-back dual fisheye cameras setup, providing 360◦ coverage of the environment. Firstly, we employ a sliding-window-based front-end to estimate real-time poses from the binocular fisheye images a…

Cited by 2SourceScholar
2024

Prompt Space Optimizing Few-shot Reasoning Success with Large Language Models

NAACL 2024findings

Prompt engineering is an essential technique for enhancing the abilities of large language models (LLMs) by providing explicit and specific instructions. It enables LLMs to excel in various tasks, such as arithmetic reasoning, question answering, summarization, relation extraction, machine translati…

2024

Revisiting Graph-Based Fraud Detection in Sight of Heterophily and Spectrum

AAAI 2024technical

Graph-based fraud detection (GFD) can be regarded as a challenging semi-supervised node binary classification task. In recent years, Graph Neural Networks (GNN) have been widely applied to GFD, characterizing the anomalous possibility of a node by aggregating neighbor information. However, fraud gra…

2023

COFFEE: Counterfactual Fairness for Personalized Text Generation in Explainable Recommendation

EMNLP 2023long main

As language models become increasingly integrated into our digital lives, Personalized Text Generation (PTG) has emerged as a pivotal component with a wide range of applications. However, the bias inherent in user written text, often used for PTG model training, can inadvertently associate different…

Cited by 0SourceScholar
2023

Multi-Objective Intrinsic Reward Learning for Conversational Recommender Systems

NeurIPS 2023poster

Conversational Recommender Systems (CRS) actively elicit user preferences to generate adaptive recommendations. Mainstream reinforcement learning-based CRS solutions heavily rely on handcrafted reward functions, which may not be aligned with user intent in CRS tasks. Therefore, the design of task-sp…

Cited by 1SourcePDFScholar
2022

An MRC Framework for Semantic Role Labeling

COLING 2022main

Semantic Role Labeling (SRL) aims at recognizing the predicate-argument structure of a sentence and can be decomposed into two subtasks: predicate disambiguation and argument labeling. Prior work deals with these two tasks independently, which ignores the semantic connection between the two tasks. I…

2022

IMO^3: Interactive Multi-Objective Off-Policy Optimization

IJCAI 2022poster

Most real-world optimization problems have multiple objectives. A system designer needs to find a policy that trades off these objectives to reach a desired operating point. This problem has been studied extensively in the setting of known objective functions. However, we consider a more practical b…

Cited by 4SourcePDFScholar
2022

VIP-SLAM: An Efficient Tightly-Coupled RGB-D Visual Inertial Planar SLAM

ICRA 2022poster

In this paper, we propose a tightly-coupled SLAM system fused with RGB, Depth, IMU and structured plane information. Traditional sparse points based SLAM systems always maintain a mass of map points to model the environment. Huge number of map points bring us a high computational complexity, making…

Cited by 30SourceScholar
2020

Summarizing Medical Conversations via Identifying Important Utterances

COLING 2020main

Summarization is an important natural language processing (NLP) task in identifying key information from text. For conversations, the summarization systems need to extract salient contents from spontaneous utterances by multiple speakers. In a special task-oriented scenario, namely medical conversat…