← Search

Jingwei Xu

23 accepted papers

2026

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding

ICLR 2026poster

We introduce FLARE, a family of vision language models (VLMs) with a fully vision-language alignment and integration paradigm. Unlike existing approaches that rely on single MLP projectors for modality alignment and defer cross-modal interaction to LLM decoding, FLARE achieves deep, dynamic integrat…

Cited by 0SourcecodeScholar
2026

Long-Context Attention Benchmark: From Kernel Efficiency to Distributed Context Parallelism

ICLR 2026poster

Transformer-based large language models (LLMs) have achieved remarkable success, yet their standard attention mechanism incurs quadratic computation and memory costs with respect to sequence length, posing a major bottleneck for long-context training. Prior work tackles this challenge along two dire…

Cited by 0SourcecodeScholar
2026

Pointer-CAD: Unifying B-Rep and Command Sequences via Pointer-based Edges & Faces Selection

CVPR 2026

Constructing computer-aided design (CAD) models is labor-intensive but essential for engineering and manufacturing. Recent advances in Large Language Models (LLMs) have inspired the LLM-based CAD generation by representing CAD as command sequences. But these methods struggle in practical scenarios b

Cited by 0SourcecodeScholar
2025

3D StreetUnveiler with Semantic-aware 2DGS - a simple baseline

ICLR 2025poster

Unveiling an empty street from crowded observations captured by in-car cameras is crucial for autonomous driving. However, removing all temporarily static objects, such as stopped vehicles and standing pedestrians, presents a significant challenge. Unlike object-centric 3D inpainting, which relies o…

Cited by 0SourcePDFScholar
2025

Loquetier: A Virtualized Multi-LoRA Framework for Unified LLM Fine-tuning and Serving

NeurIPS 2025poster

Low-Rank Adaptation (LoRA) has become a widely adopted parameter-efficient fine-tuning (PEFT) technique for adapting large language models (LLMs) to downstream tasks. While prior work has explored strategies for integrating LLM training and serving, there still remains a gap in unifying fine-tuning…

Cited by 0SourcecodeScholar
2024

ADMap: Anti-disturbance Framework for Vectorized HD Map Construction

ECCV 2024poster

"In the field of autonomous driving, online High-definition (HD) map construction is crucial for planning tasks. Recent studies have developed several high-performance HD map construction models to meet the demand. However, the point sequences generated by recent HD map construction models are jitte…

2024

Mainlobe Deceptive Jammer Suppression Using FDA-MIMO Radar in the Presence of Multipath Propagation

ICASSP 2024accepted

This paper aims to suppress mainlobe deceptive jammers considering the multipath effect in a frequency diverse array-multiple-input multiple-output (FDA-MIMO) radar. At the problem formulation stage, the overall received signal including the true target, main-lobe deceptive jammers, and burst jammin…

Cited by 0SourceScholar
2023

Learning with Logical Constraints but without Shortcut Satisfaction

ICLR 2023top-25%

Recent studies have started to explore the integration of logical knowledge into deep learning via encoding logical constraints as an additional loss function. However, existing approaches tend to vacuously satisfy logical constraints through shortcuts, failing to fully exploit the knowledge. In thi…

2023

Neuro-symbolic Learning Yielding Logical Constraints

NeurIPS 2023poster

Neuro-symbolic systems combine the abilities of neural perception and logical reasoning. However, end-to-end learning of neuro-symbolic systems is still an unsolved challenge. This paper proposes a natural framework that fuses neural network training, symbol grounding, and logical constraint synthes…

2023

Softened Symbol Grounding for Neuro-symbolic Systems

ICLR 2023poster

Neuro-symbolic learning generally consists of two separated worlds, i.e., neural network training and symbolic constraint solving, whose success hinges on symbol grounding, a fundamental problem in AI. This paper presents a novel, softened symbol grounding process, bridging the gap between the two…

2022

A Deep Learning Dataloader with Shared Data Preparation

NeurIPS 2022accept

Executing a family of Deep Neural Networks (DNNs) training jobs on the same or similar datasets in parallel is typical in current deep learning scenarios. It is time-consuming and resource-intensive because each job repetitively prepares (i.e., loads and preprocesses) the data independently, causing…

Cited by 9SourcePDFScholar
2021

Bilevel Online Adaptation for Out-of-Domain Human Mesh Reconstruction

CVPR 2021poster

This paper considers a new problem of adapting a pre-trained model of human mesh reconstruction to out-of-domain streaming videos. However, most previous methods based on the parametric SMPL model underperform in new domains with unexpected, domain-specific attributes, such as camera parameters, len…

Cited by 63PDFcodeScholar
2021

PyTouch: A Machine Learning Library for Touch Processing

ICRA 2021poster

With the increased availability of rich tactile sensors, there is an an equally proportional need for open-source and integrated software capable of efficiently and effectively processing raw touch measurements into high-level signals that can be used for control and decision-making. In this paper,…

Cited by 28SourcecodeScholar
2021

Skeleton2Mesh: Kinematics Prior Injected Unsupervised Human Mesh Recovery

ICCV 2021poster

In this paper, we decouple unsupervised human mesh recovery into the well-studied problems of unsupervised 3D pose estimation, and human mesh recovery from estimated 3D skeletons, focusing on the latter task. The challenges of the latter task are two folds: (1) pose failure (i.e., pose mismatching -…

Cited by 29PDFcodeScholar
2021

Synthesizing Long-Term 3D Human Motion and Interaction in 3D Scenes

CVPR 2021poster

Synthesizing 3D human motion plays an important role in many graphics applications as well as understanding human activity. While many efforts have been made on generating realistic and natural human motion, most approaches neglect the importance of modeling human-scene interactions and affordances.…

Cited by 145PDFScholar
2021

Towards Alleviating the Modeling Ambiguity of Unsupervised Monocular 3D Human Pose Estimation

ICCV 2021poster

In this work, we study the ambiguity problem in the task of unsupervised 3D human pose estimation from 2D counterpart. On one hand, without explicit annotation, the scale of 3D pose is difficult to be accurately captured (scale ambiguity). On the other hand, one 2D pose might correspond to multiple…

Cited by 49PDFScholar
2020

Deep Kinematics Analysis for Monocular 3D Human Pose Estimation

CVPR 2020poster

For monocular 3D pose estimation conditioned on 2D detection, noisy/unreliable input is a key obstacle in this task. Simple structure constraints attempting to tackle this problem, e.g., symmetry loss and joint angle limit, could only provide marginal improvements and are commonly treated as auxilia…

Cited by 233PDFScholar
2020

Hierarchical Style-based Networks for Motion Synthesis

ECCV 2020poster

Generating diverse and natural behaviors is one of the long-standing goals for creating intelligent characters in the animated world. In this paper, we propose an unsupervised method for generating long-range, diverse and plausible behaviors to achieve a specific goal location. Our proposed method l…

Cited by 35SourcePDFScholar
2016

TC: Throughput centric successive cancellation decoder hardware implementation for polar codes

ICASSP 2016accepted

This paper presents a hardware architecture of fast simplified successive cancellation (fast-SSC) algorithm for polar codes, which significantly reduces the decoding latency and dramatically increases the throughput. Algorithmically, fast-SSC algorithm suffers from the fact that its decoder scheduli…

Cited by 0SourceScholar