← Search

Qi WANG

117 accepted papers

2026

$\phi$-Balancing for Mixture-of-Experts Training

ICML 2026poster

Mixture-of-Experts (MoE) models rely on balanced expert utilization to fully realize their scalability. However, existing load-balancing methods are largely heuristic and operate on mini-batch assignment statistics, introducing bias relative to population-level objectives. We propose $\phi$-balancin…

Cited by 0SourceScholar
2026

$\sigma$: Sigmoid Modulation for Ultra High Resolution Diffusion

ICML 2026poster

While Diffusion Transformers (DiTs) have revolutionized high-fidelity image synthesis, the prohibitive computational costs of training at ultra-high resolutions necessitate robust inference-time extrapolation. Existing extrapolation methods typically operate under a *scale-agnostic* assumption, trea…

Cited by 0SourceScholar
2026

AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMs

AAAI 2026technical

Visual abductive reasoning (VAR) is a challenging task that requires AI systems to infer the most likely explanation for incomplete visual observations. While recent MLLMs develop strong general-purpose multimodal reasoning capabilities, they remain fall short in abductive inference, as compared to

Cited by 2SourcePDFScholar
2026

Beyond Prompt Degradation: Prototype-guided Dual-pool Prompting for Incremental Object Detection

CVPR 2026

Incremental Object Detection (IOD) aims to continuously learn new object categories without forgetting previously learned ones. Recently, prompt-based methods have gained popularity for their replay-free design and parameter efficiency. However, due to prompt coupling and prompt drift, these methods

Cited by 0SourcecodeScholar
2026

Beyond Retraining: Training-Free Unknown Class Filtering for Source-Free Open Set Domain Adaptation of Vision–Language Models

AAAI 2026technical

Vision-language models (VLMs) have gained widespread attention for their strong zero-shot capabilities across numerous downstream tasks. However, these models assume that each test image’s class label is drawn from a predefined label set and lack a reliable mechanism to reject samples from emerging

Cited by 0SourcePDFScholar
2026

Bridging Dynamics and Data: A Unified Diffusion Framework for Mechanistically-Informed Epidemic Forecasting

ICML 2026poster

Reliable epidemic forecasting is critical for public health decision-making yet remains challenging due to data sparsity and the non-stationary nature of disease dynamics. While recent hybrid models attempt to integrate mechanistic principles with data-driven approaches, they often relegate mechanis…

Cited by 0SourceScholar
2026

CoGrad3D: Spatially-Coupled Timestep Optimization with Orthogonal Gradient Fusion for 3D Generation

AAAI 2026technical

Score Distillation Sampling has driven recent advances in text-to-3D generation. However, current approaches often fail to produce 3D assets that are both rich in detail and consistent across viewpoints. These limitations primarily arise from imbalanced guidance on fine-grained details and an overde

Cited by 0SourcePDFScholar
2026

Content-Aware Mamba for Learned Image Compression

ICLR 2026poster

Recent Learned image compression (LIC) leverages Mamba-style state-space models (SSMs) for global receptive fields with linear complexity. However, the standard Mamba adopts content-agnostic, predefined raster (or multi-directional) scans under strict causality. This rigidity hinders its ability to…

Cited by 0SourcecodeScholar
2026

DiL: Discrete-anchored Representation Alignment for Semi-Supervised Continual Learning

ICML 2026poster

Leveraging the unlabeled stream is crucial yet challenging in Semi-Supervised Continual Learning (SSCL) under continual class expansion. Existing SSCL methods typically enforce dense pseudo-label consistency and indiscriminate distillation on unlabeled data, which can reinforce errors and intensify …

Cited by 0SourceScholar
2026

EiGS: Event-Informed 3D Deblur Reconstruction With Gaussian Splatting

RA-L 2026

Neural Radiance Fields (NeRF) have significantly advanced photorealistic novel view synthesis. Recently, 3D Gaus sian Splatting has emerged as a promising technique with faster training and rendering speeds. However, both methods rely heavily on clear images and precise camera poses, limiting perfor

Cited by 0SourceScholar
2026

EiGS: Event-Informed 3D Deblur Reconstruction with Gaussian Splatting

ICRA 2026poster

Neural Radiance Fields (NeRF) have significantly advanced photorealistic novel view synthesis. Recently, 3D Gaussian Splatting has emerged as a promising technique with faster training and rendering speeds. However, both methods rely heavily on clear images and precise camera poses, limiting perform…

Cited by 0SourceScholar
2026

Exploring Generalizable Remote Sensing Change Detection via Low-Rank Exchange Adaptation of Vision Foundation Model

AAAI 2026technical

Remote sensing change detection (CD) has achieved remarkable progress in recent years. However, little attention has been paid to generalizable change detection (GCD) methods that can effectively generalize to unseen scenarios or domains beyond the training distribution. The major challenges in GCD

Cited by 0SourcePDFScholar
2026

Goal-Driven Reward by Video Diffusion Models for Reinforcement Learning

CVPR 2026

Reinforcement Learning (RL) has achieved remarkable success in various domains, yet it often relies on carefully designed programmatic reward functions to guide agent behavior. Designing such reward functions can be challenging and may not generalize well across different tasks. To address this limi

Cited by 0SourceScholar
2026

HISE-KT: Synergizing Heterogeneous Information Networks and LLMs for Explainable Knowledge Tracing with Meta-Path Optimization

AAAI 2026technical

Knowledge Tracing (KT) aims to mine students’ evolving knowledge states and predict their future question-answering performance. Existing methods based on heterogeneous information networks (HINs) are prone to introducing noises due to manual or random selection of meta-paths and lack necessary qua

Cited by 0SourcePDFScholar
2026

Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs

CVPR 2026

While continual visual instruction tuning (CVIT) has shown promise in adapting multimodal large language models (MLLMs), existing studies predominantly focus on models without safety alignment. This critical oversight ignores the fact that real-world MLLMs inherently require such mechanisms to mitig

Cited by 0SourcecodeScholar
2026

Inconsistency Biases in Dynamic Data Pruning

ICLR 2026poster

Dynamic data pruning accelerates training by focusing on informative samples. However, comparing importance scores across different model states introduces inconsistency (score context drift), and variable selection rates bias gradient dynamics over time (temporal gradient bias). We introduce RePB (…

Cited by 0SourcecodeScholar
2026

MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos

ICRA 2026poster

We aim to enable humanoid robots to efficiently solve new manipulation tasks from a few video examples. In-context learning (ICL) is a promising framework for achieving this goal due to its test-time data efficiency and rapid adaptability. However, current ICL methods rely on labor-intensive teleope…

2026

MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular Videos

CVPR 2026

Motion capture now underpins content creation far beyond digital humans, yet most pipelines remain species- or template-specific. We formalize this gap as Category-Agnostic Motion Capture (CAMoCap): given a monocular video and an arbitrary rigged 3D asset as a prompt, the goal is to reconstruct a ro

Cited by 0SourcecodeScholar
2026

OmniZip: Learning a Unified and Lightweight Lossless Compressor for Multi-Modal Data

CVPR 2026

Lossless compression is essential for efficient data storage and transmission. Although learning-based lossless compressors achieve strong results, most of them are designed for a single modality, leading to redundant compressor deployments in multi-modal settings. Designing a unified multi-modal co

Cited by 0SourcecodeScholar
2026

PIMRL: Physics-Informed Multi-Scale Recurrent Learning for Burst-Sampled Spatiotemporal Dynamics

AAAI 2026technical

Deep learning has shown strong potential in modeling complex spatiotemporal dynamics. However, most existing methods depend on densely and uniformly sampled data, which is often unavailable in practice due to sensor and cost limitations. In many real-world settings, such as mobile sensing and physic

Cited by 0SourcePDFScholar
2026

Reasoning via Implicit Self-supervised Emergence for Instruction Segmentation

AAAI 2026technical

We challenge the assumption that complex instruction-guided segmentation tasks necessitate equally complex and explicit supervision. This paper introduces RISE (Reasoning via Implicit Self-supervised Emergence), a framework that learns intricate compositional reasoning, spanning spatial relations to

Cited by 0SourcePDFScholar
2026

Slender3D: Curve-Guided Multi-View Reconstruction of Slender Structures

AAAI 2026technical

Although geometric reconstruction of general objects from images has made remarkable progress in recent years, slender structures remain largely underexplored, despite their critical importance in engineering, biomedical, and agricultural applications. To bridge this gap, we propose a dedicated 2DGS

Cited by 0SourcePDFScholar
2026

Symphony-MoE: Harmonizing Disparate Pre-trained Models into a Coherent Mixture-of-Experts

AAAI 2026technical

Mixture-of-Experts (MoE) models enable scalable performance by activating large parameter sets sparsely, minimizing computational overhead. To mitigate the prohibitive cost of training MoEs from scratch, recent work employs upcycling, reusing a single pre-trained dense model by replicating its feed-

Cited by 0SourcePDFScholar
2026

TAPO: Dynamic Teacher and Perturbed Answer Injection for Policy Optimization

AAAI 2026technical

Reinforcement learning (RL) has emerged as a powerful framework to improve the reasoning performance of large language models (LLMs), with approaches such as Group Relative Policy Optimization (GRPO) showing promising results. However, GRPO and its variants struggle with collapsed groups (i.e., all-

Cited by 0SourcePDFScholar
2025

CLIP-driven View-aware Prompt Learning for Unsupervised Vehicle Re-identification

AAAI 2025technical

With the emergence of vision-language pre-trained models, such as CLIP, some textual prompts have been gradually introduced recently into re-identification (Re-ID) tasks to obtain considerably robust multimodal information. However, most textual descriptions based on vehicle Re-ID tasks only contain…

Cited by 0SourcePDFScholar
2025

CarPlanner: Consistent Auto-regressive Trajectory Planning for Large-Scale Reinforcement Learning in Autonomous Driving

CVPR 2025poster

Trajectory planning is vital for autonomous driving, ensuring safe and efficient navigation in complex environments. While recent learning-based methods, particularly reinforcement learning (RL), have shown promise in specific scenarios, RL planners struggle with training inefficiencies and managing…

2025

Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement Learning

ICCV 2025poster

Training visual reinforcement learning (RL) in practical scenarios presents a significant challenge, i.e., RL agents suffer from low sample efficiency in environments with variations. While various approaches have attempted to alleviate this issue by disentangled representation learning, these metho…

Cited by 0SourcePDFScholar
2025

DreamGen: Unlocking Generalization in Robot Learning through Video World Models

CoRL 2025poster

In this work, we unlock new capabilities in robot learning from neural trajectories, synthetic robot data generated from video world models. Our proposed recipe is simple, but powerful: we take the most recent state-of-the-art video generative models (world models), adapt them to the target robot em…

Cited by 0SourcecodeScholar
2025

ExtPose: Robust and Coherent Pose Estimation by Extending ViTs

ICML 2025poster

Vision Transformers (ViT) are remarkable at 3D pose estimation, yet they still encounter certain challenges. One issue is that the popular ViT architecture for pose estimation is limited to images and lacks temporal information. Another challenge is that the prediction often fails to maintain pixel…

Cited by 0SourcePDFScholar
2025

FLARE: Robot Learning with Implicit World Modeling

CoRL 2025poster

We introduce **F**uture **LA**tent **R**presentation Alignm**E**nt (**FLARE**), a novel framework that integrates predictive world modeling into robot policy learning. By aligning features from a diffusion transformer with latent embeddings of future observations, **FLARE** enables a diffusion trans…

Cited by 0SourceScholar
2025

From One to More: Contextual Part Latents for 3D Generation

ICCV 2025poster

To generate 3D objects, early research focused on multi-view-driven approaches relying solely on 2D renderings. Recently, the 3D native latent diffusion paradigm has demonstrated superior performance in 3D generation, because it fully leverages the geometric information provided in ground truth 3D d…

2025

GTDE: Grouped Training with Decentralized Execution for Multi-agent Actor-Critic

AAAI 2025technical

The rapid advancement of multi-agent reinforcement learning (MARL) has given rise to diverse training paradigms to learn the policies of each agent in the multi-agent system. The paradigms of decentralized training and execution (DTDE) and centralized training with decentralized execution (CTDE) hav…

2025

H3D-DGS: Exploring Heterogeneous 3D Motion Representation for Deformable 3D Gaussian Splatting

NeurIPS 2025poster

Dynamic scene reconstruction poses a persistent challenge in 3D vision. Deformable 3D Gaussian Splatting has emerged as an effective method for this task, offering real-time rendering and high visual fidelity. This approach decomposes a dynamic scene into a static representation in a canonical space…

Cited by 0SourceScholar
2025

Inverse Rendering using Multi-Bounce Path Tracing and Reservoir Sampling

ICLR 2025poster

We introduce MIRReS, a novel two-stage inverse rendering framework that jointly reconstructs and optimizes explicit geometry, materials, and lighting from multi-view images. Unlike previous methods that rely on implicit irradiance fields or oversimplified ray tracing, our method begins with an initi…

Cited by 0SourcePDFScholar
2025

Leveraging Pretrained Diffusion Models for Zero-Shot Part Assembly

IJCAI 2025

3D part assembly aims to understand part relationships and predict their 6-DoF poses to construct realistic 3D shapes, addressing the growing demand for autonomous assembly, which is crucial for robots. Existing methods mainly estimate the transformation of each part by training neural networks unde

2025

MicroEdit: Neuron-level Knowledge Disentanglement and Localization in Lifelong Model Editing

EMNLP 2025

Large language models (LLMs) require continual knowledge updates to keep pace with the evolving world. While various model editing methods have been proposed, most face critical challenges in the context of lifelong learning due to two fundamental limitations: (1) Edit Overshooting - parameter updat

2025

MultiPDENet: PDE-embedded Learning with Multi-time-stepping for Accelerated Flow Simulation

ICML 2025poster

Solving partial differential equations (PDEs) by numerical methods meet computational cost challenge for getting the accurate solution since fine grids and small time steps are required. Machine learning can accelerate this process, but struggle with weak generalizability, interpretability, and data…

Cited by 0SourcePDFScholar
2025

Open-World Reinforcement Learning over Long Short-Term Imagination

ICLR 2025oral

Training visual reinforcement learning agents in a high-dimensional open world presents significant challenges. While various model-based methods have improved sample efficiency by learning interactive world models, these agents tend to be “short-sighted”, as they are typically trained on short snip…

2025

PeSANet: Physics-encoded Spectral Attention Network for Simulating PDE-Governed Complex Systems

IJCAI 2025

Accurately modeling and forecasting complex systems governed by partial differential equations (PDEs) is crucial in various scientific and engineering domains. However, traditional numerical methods struggle in real-world scenarios due to incomplete or unknown physical laws. Meanwhile, machine learn

2025

PhyMPGN: Physics-encoded Message Passing Graph Network for spatiotemporal PDE systems

ICLR 2025spotlight

Solving partial differential equations (PDEs) serves as a cornerstone for modeling complex dynamical systems. Recent progresses have demonstrated grand benefits of data-driven neural-based models for predicting spatiotemporal dynamics (e.g., tremendous speedup gain compared with classical numerical…

Cited by 4SourcePDFScholar
2025

RHYTHM: Reasoning with Hierarchical Temporal Tokenization for Human Mobility

NeurIPS 2025poster

Predicting human mobility is inherently challenging due to complex long-range dependencies and multi-scale periodic behaviors. To address this, we introduce RHYTHM (Reasoning with Hierarchical Temporal Tokenization for Human Mobility), a unified framework that leverages large language models (LLMs)…

Cited by 0SourcecodeScholar
2025

SMoLoRA: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction Tuning

ICCV 2025poster

Visual instruction tuning (VIT) enables multimodal large language models (MLLMs) to effectively handle a wide range of vision tasks by framing them as language-based instructions. Building on this, continual visual instruction tuning (CVIT) extends the capability of MLLMs to incrementally learn new…

2025

Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets

ICML 2025poster

Large language models (LLMs) have shown great potential as general-purpose AI assistants across various domains. To fully leverage this potential in specific applications, many companies provide fine-tuning API services, enabling users to upload their own data for LLM customization. However, fine-tu…

2025

TAD-E2E: A Large-scale End-to-end Autonomous Driving Dataset

ICCV 2025poster

End-to-end autonomous driving technology has recently become a focal point of research and application in autonomous driving. State-of-the-art (SOTA) methods are often trained and evaluated on the NuScenes dataset. However, the NuScenes dataset, introduced in 2019 for 3D perception tasks, faces seve…

Cited by 0SourcePDFScholar
2025

VehicleMAE: View-asymmetry Mutual Learning for Vehicle Re-identification Pre-training via Masked AutoEncoders

ICCV 2025poster

Large-scale pre-training technology has achieved remarkable performance in diversified object re-identification (Re-ID) downstream tasks. Nevertheless, to our best knowledge, the pre-training model specifically for vehicle Re-ID, which focuses on tackling the challenge of multi-view variations, has…

Cited by 0SourcePDFScholar
2025

VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning

NeurIPS 2025poster

Reinforcement fine-tuning (RFT) has shown great promise in achieving humanlevel reasoning capabilities of Large Language Models (LLMs), and has recently been extended to MLLMs. Nevertheless, reasoning about videos, which is a fundamental aspect of human intelligence, remains a persistent challenge d…

Cited by 0SourcecodeScholar
2024

A Bi-Pyramid Multimodal Fusion Method for the Diagnosis Of Bipolar Disorders

ICASSP 2024accepted

Previous research on the diagnosis of Bipolar disorder has mainly focused on resting-state functional magnetic resonance imaging. However, their accuracy can not meet the requirements of clinical diagnosis. Efficient multimodal fusion strategies have great potential for applications in multimodal da…

Cited by 0SourceScholar
2024

Cross-Sentence Gloss Consistency for Continuous Sign Language Recognition

AAAI 2024technical

Continuous sign language recognition (CSLR) aims to recognize gloss sequences from continuous sign videos. Recent works enhance the gloss representation consistency by mining correlations between visual and contextual modules within individual sentences. However, there still remain much richer corre…

Cited by 2SourcePDFScholar
2024

DailyDVS-200: A Comprehensive Benchmark Dataset for Event-Based Action Recognition

ECCV 2024poster

"Neuromorphic sensors, specifically event cameras, revolutionize visual data acquisition by capturing pixel intensity changes with exceptional dynamic range, minimal latency, and energy efficiency, setting them apart from conventional frame-based cameras. The distinctive capabilities of event camera…

2024

EMIE-MAP: Large-Scale Road Surface Reconstruction Based on Explicit Mesh and Implicit Encoding

ECCV 2024poster

"Road surface reconstruction plays a vital role in autonomous driving systems, enabling road lane perception and high-precision mapping. Recently, neural implicit encoding has achieved remarkable results in scene representation, particularly in the realistic rendering of scene textures. However, it…

Cited by 8SourcePDFScholar
2024

Error-aware Sampling in Adaptive Shells for Neural Surface Reconstruction

IJCAI 2024poster

Neural implicit surfaces with signed distance functions (SDFs) achieve superior quality in 3D geometry reconstruction. However, training SDFs is time-consuming because it requires a great number of samples to calculate accurate weight distributions and a considerable amount of samples sampled from t…

2024

FDC-NeRF: Learning Pose-Free Neural Radiance Fields with Flow-Depth Consistency

ICASSP 2024accepted

Learning neural radiance fields (NeRF) without camera poses has been widely studied. However, recent methods lack explicit and effective supervision for pose estimation, resulting in ambiguous optimization of camera pose and NeRF geometry during joint training, particularly in scenarios involving la…

Cited by 0SourceScholar
2024

FLDM-VTON: Faithful Latent Diffusion Model for Virtual Try-on

IJCAI 2024poster

Despite their impressive generative performance, latent diffusion model-based virtual try-on (VTON) methods lack faithfulness to crucial details of the clothes, such as style, pattern, and text. To alleviate these issues caused by the diffusion stochastic nature and latent supervision, we propose a…

Cited by 6SourcePDFScholar
2024

Generalizable Thermal-based Depth Estimation via Pre-trained Visual Foundation Model

ICRA 2024poster

Depth estimation is a crucial task in computer vision, applicable to various domains such as 3D reconstruction, robotics, and autonomous driving. In particular, thermal-based depth estimation has unique advantages, including night-time vision. However, the existing depth estimation method remains ch…

Cited by 0SourceScholar
2024

HDPNERF: Hybrid Depth Priors for Neural Radiance Fields from Sparse Input Views

ICASSP 2024accepted

Neural Radiance Field (NeRF) shows a high prospect in the task of novel view synthesis. However, performance degrades drastically under limited input views since NeRF heavily relies on a large number of images to fit the geometry in scenes. Recent efforts focus on introducing extra constraints to im…

Cited by 0SourceScholar
2024

Improving Learned Video Compression by Exploring Spatial Redundancy

ICASSP 2024accepted

Learned video compression has developed rapidly and shown promising rate-distortion performance recently. Existing works have made great progress on removing temporal redundancy between inter-frames, while neglecting spatial redundancy within a frame. In this paper, we propose to explore spatial red…

Cited by 0SourceScholar
2024

Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement Learning

NeurIPS 2024poster

Training offline RL models using visual inputs poses two significant challenges, *i.e.*, the overfitting problem in representation learning and the overestimation bias for expected future rewards. Recent work has attempted to alleviate the overestimation bias by encouraging conservative behaviors. T…

2024

MuChin: A Chinese Colloquial Description Benchmark for Evaluating Language Models in the Field of Music

IJCAI 2024poster

The rapidly evolving multimodal Large Language Models (LLMs) urgently require new benchmarks to uniformly evaluate their performance on understanding and textually describing music. However, due to semantic gaps between Music Information Retrieval (MIR) algorithms and human understanding, discrepanc…

2024

P$^2$C$^2$Net: PDE-Preserved Coarse Correction Network for efficient prediction of spatiotemporal dynamics

NeurIPS 2024poster

When solving partial differential equations (PDEs), classical numerical methods often require fine mesh grids and small time stepping to meet stability, consistency, and convergence conditions, leading to high computational cost. Recently, machine learning has been increasingly utilized to solve PDE…

Cited by 5SourcePDFScholar
2024

PEP: Policy-Embedded Trajectory Planning for Autonomous Driving

RA-L 2024

Autonomous driving demands proficient trajectory planning to ensure safety and comfort. This letter introduces Policy-Embedded Planner (PEP), a novel framework that enhances closed-loop performance of imitation learning (IL) based planners by embedding a neural policy for sequential ego pose generat

Cited by 8SourceScholar
2024

Resource-Aware Federated Self-Supervised Learning with Global Class Representations

NeurIPS 2024poster

Due to the heterogeneous architectures and class skew, the global representation models training in resource-adaptive federated self-supervised learning face with tricky challenges: $\textit{deviated representation abilities}$ and $\textit{inconsistent representation spaces}$. In this work, we are…

Cited by 0SourcePDFScholar
2024

ScreenAgent: A Vision Language Model-driven Computer Control Agent

IJCAI 2024poster

Large Language Models (LLM) can invoke a variety of tools and APIs to complete complex tasks. The computer, as the most powerful and universal tool, could potentially be controlled by a trained LLM agent. Powered by the computer, we can hopefully build a more generalized agent to assist humans in va…

2024

Similarity Knowledge Distillation with Calibrated Mask

ICASSP 2024accepted

In this paper, we propose a novel and efficient method for knowledge distillation, which is structurally simple and requires negligible computation overhead. Our method includes three modules. The first module is the calibrated mask, which avoids the teacher model’s incorrect representation to distu…

Cited by 0SourceScholar
2023

Bridge the Inference Gaps of Neural Processes via Expectation Maximization

ICLR 2023poster

The neural process (NP) is a family of computationally efficient models for learning distributions over functions. However, it suffers from under-fitting and shows suboptimal performance in practice. Researchers have primarily focused on incorporating diverse structural inductive biases, e.g. attent…

2023

Centimeter-Scale Underwater Robot With High-Speed Inspired by Jellyfish

RA-L 2023

Centimeter-scale underwater robots have important applications in underwater resource exploration, environmental monitoring, equipment fault diagnosis, and military applications. However, developing small underwater robots remains a challenge because of the limitation of miniaturized structures. In

Cited by 13SourceScholar
2023

Hierarchical Attention Network for Planning-Informed Multi-Agent Trajectory Prediction

IROS 2023poster

The accurate prediction of the neighboring vehicles' trajectories affects the security of autonomous driving vehicles. However, it is challenging for existing methods to anticipating the trajectories of vehicles in the vicinity due to the uncertainty of driving behaviors and the complex interaction…

Cited by 3SourceScholar
2023

RenderIH: A Large-Scale Synthetic Dataset for 3D Interacting Hand Pose Estimation

ICCV 2023poster

The current interacting hand (IH) datasets are relatively simplistic in terms of background and texture, with hand joints being annotated by a machine annotator, which may result in inaccuracies, and the diversity of pose distribution is limited. However, the variability of background, pose distribu…

Cited by 19PDFcodeScholar
2023

Towards Stable Human Pose Estimation via Cross-View Fusion and Foot Stabilization

CVPR 2023poster

Towards stable human pose estimation from monocular images, there remain two main dilemmas. On the one hand, the different perspectives, i.e., front view, side view, and top view, appear the inconsistent performances due to the depth ambiguity. On the other hand, foot posture plays a significant rol…

Cited by 5SourcePDFScholar
2023

Weakly-Supervised Scene-Specific Crowd Counting Using Real-Synthetic Hybrid Data

ICASSP 2023accepted

Due to the domain gap between the public large-scale datasets and actual scenes, the crowd counting models trained on the common datasets have a significant performance degradation when applying in practical applications. To address the above issue, one of the solution is to label additional data fr…

Cited by 0SourceScholar
2022

A Speech-driven Sign Language Avatar Animation System for Hearing Impaired Applications

IJCAI 2022poster

Sign language is the communication language used in hearing impaired community. Recently, the research of sign language production has made great progress but still need to cope with some critical challenges. In this paper, we propose a system-level scheme and push forward the implementation of sign…

Cited by 6SourcePDFScholar
2022

AttExplainer: Explain Transformer via Attention by Reinforcement Learning

IJCAI 2022poster

Transformer and its variants, built based on attention mechanisms, have recently achieved remarkable performance in many NLP tasks. Most existing works on Transformer explanation tend to reveal and utilize the attention matrix with human subjective intuitions in a qualitative manner. However, the hu…

2022

BiP-Net: Bidirectional Perspective Strategy Based Arbitrary-Shaped Text Detection Network

ICASSP 2022accepted

Detecting irregular-shaped text instances is the main challenge for text detection. Existing approaches can be roughly divided into top-down and bottom-up perspective methods. The former encodes text contours into unified units, which always fails to fit highly curved text contours. The latter repre…

Cited by 0SourceScholar
2022

DR.VIC: Decomposition and Reasoning for Video Individual Counting

CVPR 2022poster

Pedestrian counting is a fundamental tool for understanding pedestrian patterns and crowd flow analysis. Existing works (e.g., image-level pedestrian counting, crossline crowd counting et al.) either only focus on the image-level counting or are constrained to the manual annotation of lines. In this…

Cited by 28PDFcodeScholar
2022

Global Evolution Neural Network for Segmentation of Remote Sensing Images

ICASSP 2022accepted

The popular convolutional neural networks (CNNs) have been successfully used in very high-resolution remote sensing image semantic segmentation. However, these networks often suffer from performance limitations. First, although deeper networks usually provide better feature representation, they may…

Cited by 0SourceScholar
2022

Model-based Meta Reinforcement Learning using Graph Structured Surrogate Models and Amortized Policy Search

ICML 2022spotlight

Reinforcement learning is a promising paradigm for solving sequential decision-making problems, but low data efficiency and weak generalization across tasks are bottlenecks in real-world applications. Model-based meta reinforcement learning addresses these issues by learning dynamics and leveraging…

Cited by 29SourcePDFScholar
2021

Lightweight Non-Local Network for Image Super-Resolution

ICASSP 2021accepted

The popular deep convolutional networks used for image super-resolution (SR) reconstruction often increase the network depth and employ attention mechanism to improve image reconstruction effect. However, these networks suffer from two problems. The first is the deeper network easily causes higher c…

Cited by 0SourceScholar
2021

PI-Net: An End-to-End Deep Neural Network for Bidirectionally and Directly Fusing Point Clouds With Images

RA-L 2021

We present a novel network, PI-Net, for the fusion between point clouds and images in this letter. Most existing fusion methods project point clouds into pseudo images and then fuse the pseudo and RGB images with 2D CNNs. To get rid of structuring the pseudo images as the preprocessing, we propose a

Cited by 6SourceScholar
2020

Doubly Stochastic Variational Inference for Neural Processes with Hierarchical Latent Variables

ICML 2020poster

Neural processes (NPs) constitute a family of variational approximate models for stochastic processes with promising properties in computational efficiency and uncertainty quantification. These processes use neural networks with latent variable inputs to induce a predictive distribution. However, th…

Cited by 47SourcePDFScholar
2020

IQ-STAN: Image Quality Guided Spatio-Temporal Attention Network for License Plate Recognition

ICASSP 2020accepted

License plate recognition (LPR) is one of the essential components in intelligent transportation systems. Although the image processing algorithms for LPR have been extensively studied in the past several years, the recognition performance is still not satisfactory especially in unconstrained comple…

Cited by 0SourceScholar
2020

KALM: Key Area Localization Mechanism for Abnormality Detection in Musculoskeletal Radiographs

ICASSP 2020accepted

Recently abnormality detection in musculoskeletal radio-graphs has attracted many attentions. For abnormality detection, it is crucial to locate the most important area in the musculoskeletal radiographs. To achieve this goal, we propose a key area localization mechanism (KALM) for abnormality detec…

Cited by 0SourceScholar
2020

Robust Rank Constrained Sparse Learning: A Graph-Based Method for Clustering

ICASSP 2020accepted

Graph-based clustering is an advanced clustering techniuqe, which partitions the data according to an affinity graph. However, the graph quality affects the clustering results to a large extent, and it is difficult to construct a graph with high quality, especially for data with noises and outliers.…

Cited by 0SourceScholar
2020

Unsupervised Semantic Aggregation and Deformable Template Matching for Semi-Supervised Learning

NeurIPS 2020poster

Unlabeled data learning has attracted considerable attention recently. However, it is still elusive to extract the expected high-level semantic feature with mere unsupervised learning. In the meantime, semi-supervised learning (SSL) demonstrates a promising future in leveraging few samples. In this…

2020

Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences

CVPR 2020poster

In this paper, we consider a novel task, Spatio-Temporal Video Grounding for Multi-Form Sentences (STVG). Given an untrimmed video and a declarative/interrogative sentence depicting an object, STVG aims to localize the spatio-temporal tube of the queried object. STVG has two challenging settings: (1…

Cited by 134PDFcodeScholar
2017

Embedding structured contour and location prior in siamesed fully convolutional networks for road detection

ICRA 2017poster

Road detection from the perspective of moving vehicles is a challenging issue in autonomous driving. Recently, many deep learning methods spring up for this task because they can extract high-level local features to find road regions from raw RGB data, such as Convolutional Neural Networks (CNN) and…

Cited by 307SourceScholar