← Search

Qiang Wang

91 accepted papers

2026

CHESS: Chebyshev Spectral Synthesis for Trajectory Condensation

ICML 2026poster

Learning from continuous-time trajectories requires modeling multivariate sensor measurements generated by underlying physical or dynamical processes. Under extreme data compression and heterogeneous sampling, directly optimizing synthetic signals as discrete sample values becomes fundamentally misa…

Cited by 0SourceScholar
2026

CTBench: Cryptocurrency Time Series Generation Benchmark

ICLR 2026poster

Synthetic time series are vital for data augmentation, stress testing, and prototyping in quantitative finance. Yet in cryptocurrency markets, characterized by 24/7 trading, extreme volatility, and rapid regime shifts, existing Time Series Generation (TSG) methods and benchmarks often fall short, je…

Cited by 0SourcecodeScholar
2026

Class Incremental Medical Image Segmentation via Prototype-Guided Calibration and Dual-Aligned Distillation

AAAI 2026technical

Class incremental medical image segmentation (CIMIS) aims to preserve knowledge of previously learned classes while learning new ones without relying on old-class annotations. However, existing methods 1) either adopt one-size-fits-all strategies that treat all spatial regions and feature channels e

Cited by 0SourcePDFScholar
2026

CoPE: A Framework for Optimizing Coordination between Planning and Execution in LLM-based Agents

ICML 2026poster

Fine-tuning Large Language Models (LLMs) as autonomous agents on domain-specific data has emerged as a promising paradigm for tackling interactive, real-world tasks. However, existing studies have overlooked the critical coordination between long-term planning and multi-step execution in optimizing …

Cited by 0SourceScholar
2026

D-FCGS: Feedforward Compression of Dynamic Gaussian Splatting for Free-Viewpoint Videos

AAAI 2026technical

Free-Viewpoint Video (FVV) enables immersive 3D experiences, but efficient compression of dynamic 3D representation remains a major challenge. Existing dynamic 3D Gaussian Splatting methods couple reconstruction with optimization-dependent compression and customized motion formats, limiting generali

Cited by 0SourcePDFScholar
2026

EC-MVSNet: Enhanced Cascaded Multi-View Stereo with Cross-Scale Relevance Integration

AAAI 2026technical

Cascade-based multi-scale architectures are currently the mainstream in Multi-view Stereo (MVS), achieving a balance between computational efficiency and reconstruction accuracy. However, existing cascade MVS methods suffer from significant limitations in cross-scale information utilization, where d

Cited by 0SourcePDFScholar
2026

GMT: Effective Global Framework for Multi-Camera Multi-Target Tracking

CVPR 2026

Existing Multi-Camera Multi-Target (MCMT) tracking models typically adopt a two-stage framework, involving single-camera tracking followed by inter-camera tracking. However, in this paradigm, the use of multiple views is confined to recovering missed matches in the first stage, providing a limited c

Cited by 0SourcecodeScholar
2026

GOAL: Geometrically Optimal Alignment for Continual Generalized Category Discovery

AAAI 2026technical

Continual Generalized Category Discovery (C-GCD) requires identifying novel classes from unlabeled data while retaining knowledge of known classes over time. Existing methods typically update classifier weights dynamically, resulting in forgetting and inconsistent feature alignment. We propose GOAL,

Cited by 0SourcePDFScholar
2026

Is Parameter Isolation Better for Prompt-Based Continual Learning?

CVPR 2026

Prompt-based continual learning methods effectively mitigate catastrophic forgetting. However, most existing methods assign a fixed set of prompts to each task, completely isolating knowledge across tasks and resulting in suboptimal parameter utilization. To address this, we consider the practical n

Cited by 0SourceScholar
2026

Learning Like Humans: Analogical Concept Learning for Generalized Category Discovery

CVPR 2026

Generalized Category Discovery (GCD) seeks to uncover novel categories in unlabeled data while preserving recognition of known categories, yet prevailing visual-only pipelines and the loose coupling between supervised learning and discovery often yield brittle boundaries on fine-grained, look-alike

Cited by 0SourcecodeScholar
2026

PipeDiT: Accelerating Diffusion Transformers in Video Generation with Task Pipelining and Model Decoupling

AAAI 2026technical

Video generation has been advancing rapidly, and diffusion transformer (DiT) based models have demonstrated remarkable capabilities. However, their practical deployment is often hindered by slow inference speeds and high memory consumption. In this paper, we propose a novel pipelining framework name

Cited by 0SourcePDFScholar
2026

Predictive Regularization Against Visual Representation Degradation in Multimodal Large Language Models

CVPR 2026

While Multimodal Large Language Models (MLLMs) excel at vision-language tasks, the cost of their language-driven training on internal visual foundational competence remains unclear. In this paper, we conduct a detailed diagnostic analysis to unveil a pervasive issue: visual representation degradatio

Cited by 0SourceScholar
2026

SALR: Sparsity-Aware Low-Rank Representation for Efficient Fine-Tuning of Large Language Models

AAAI 2026technical

Adapting large pre-trained language models to downstream tasks often entails fine-tuning millions of parameters or deploying costly dense weight updates, which hinders their use in resource-constrained environments. Low-rank Adaptation (LoRA) reduces trainable parameters by factorizing weight update

Cited by 0SourcePDFScholar
2026

SPE-MVS: Spatial Position Encoding Enhanced Multi-View Stereo with Monocular Depth Priors

CVPR 2026

Learning-based Multi-View Stereo (MVS) methods have become the mainstream in the field, relying on the construction of cost Learning-based Multi-View Stereo (MVS) methods have become the mainstream in the field, relying on the construction of cost volumes through multi-view feature similarity comput

Cited by 0SourcecodeScholar
2026

STUR3D: Spatio-Temporal Unified Representation Learning for 3D Object Detection

CVPR 2026

Existing surrounding-view 3D object detectors initialize high-confidence queries using current 2D information, while leveraging historical 3D features as priors. However, such heavy reliance on 2D cues introduces spatio-temporal inconsistencies between 2D and 3D representations. Specifically, 2D cue

Cited by 0SourcecodeScholar
2026

Shared & Domain Self-Adaptive Experts with Frequency-Aware Discrimination for Continual Test-Time Adaptation

AAAI 2026technical

This paper focuses on the Continual Test-Time Adaptation (CTTA) task, aiming to enable an agent to continuously adapt to evolving target domains while retaining previously acquired domain knowledge for effective reuse when those domains reappear. Existing shared-parameter paradigms struggle to balan

Cited by 0SourcePDFScholar
2026

TIPS: Tiered Information-Rich Planning Strategy for Efficient UGV Autonomous Exploration

ICRA 2026poster

In this letter, we propose a tiered systematic framework to enhance the overall efficiency and environmental coverage of autonomous exploration for Autonomous Ground Vehicle (AGV) in complex environments with narrow regions. At the local level, we introduce a novel Multi-cause Triggering Sensor Mode…

Cited by 0Scholar
2026

VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing

ICML 2026poster

Despite the rapid advancements in Vision-Language Models (VLMs), a critical gap remains in their ability to handle structured, controllable diagrammatic tasks essential for professional workflows, as existing methods predominantly rely on pixel-based synthesis which operates in probabilistic pixel s…

Cited by 0SourceScholar
2026

WiTTA-Bench: Benchmarking Test-Time Adaptation for WiFi Sensing

CVPR 2026

WiFi sensing offers passive and privacy-preserving perception that complements vision-based sensing, but its performance degrades sharply under domain shifts caused by changes in environment, subjects, or hardware. This challenge is exacerbated in real-world deployments where source data are unavail

Cited by 0SourcecodeScholar
2025

All-Day Multi-Camera Multi-Target Tracking

CVPR 2025poster

The capability of tracking objects in low-light environments like nighttime is crucial for numerous real-world applications such as crowd behavior analysis and traffic scene understanding. However, previous Multi-Camera Multi-Target(MCMT) tracking methods are primarily focused on tracking during day…

2025

Applicability Analysis for Optical Cooperative Localization

IROS 2025

For optical cooperative localization, which employs optical beacons with prior features as cooperative targets, a fundamental prerequisite is to ensure that the beacons are always captured by the vision sensors during the entire localization process. In other words, there is an applicability issue o

Cited by 0SourceScholar
2025

CPRM: A LLM-based Continual Pre-training Framework for Relevance Modeling in Commercial Search

NAACL 2025industry

Relevance modeling between queries and items stands as a pivotal component in commercial search engines, directly affecting the user experience. Given the remarkable achievements of large language models (LLMs) in various natural language processing (NLP) tasks, LLM-based relevance modeling is gradu…

2025

Consistent Supervised-Unsupervised Alignment for Generalized Category Discovery

NeurIPS 2025poster

Generalized Category Discovery (GCD) focuses on classifying known categories while simultaneously discovering novel categories from unlabeled data. However, previous GCD methods face challenges due to inconsistent optimization objectives and category confusion. This leads to feature overlap and ulti…

Cited by 0SourceScholar
2025

DualCP: Rehearsal-Free Domain-Incremental Learning via Dual-Level Concept Prototype

AAAI 2025technical

Domain-Incremental Learning (DIL) enables vision models to adapt to changing conditions in real-world environments while maintaining the knowledge acquired from previous domains. Given privacy concerns and training time, Rehearsal-Free DIL (RFDIL) is more practical. Inspired by the incremental cogni…

Cited by 0SourcePDFScholar
2025

Dynamic Integration of Task-Specific Adapters for Class Incremental Learning

CVPR 2025poster

Non-exemplar Class Incremental Learning (NECIL) enables models to continuously acquire new classes without retraining from scratch and storing old task exemplars, addressing privacy and storage issues. However, the absence of data from earlier tasks exacerbates the challenge of catastrophic forgetti…

Cited by 2SourcePDFScholar
2025

Flexible Sharpness-Aware Personalized Federated Learning

AAAI 2025technical

Personalized federated learning (PFL) is a new paradigm to address the statistical heterogeneity problem in federated learning. Most existing PFL methods focus on leveraging global and local information such as model interpolation or parameter decoupling. However, these methods often overlook the ge…

2025

Fuse Before Transfer: Knowledge Fusion for Heterogeneous Distillation

ICCV 2025poster

Most knowledge distillation (KD) methods focus on teacher-student pairs with similar architectures, such as both being CNN models. The potential and flexibility of KD can be greatly improved by expanding it to Cross-Architecture KD (CAKD), where the knowledge of homogeneous and heterogeneous teacher…

2025

OAMaskFlow: Occlusion-Aware Motion Mask for Scene Flow

AAAI 2025technical

The scene flow estimation methods make significant progress by estimating pixel-wise 3D motion on implicitly learning a motion embedding using an end-to-end differentiable optimization framework. However, the motion embedding learned implicitly is insufficient for grouping pixels into rigid object i…

Cited by 0SourcePDFScholar
2025

ParZC: Parametric Zero-Cost Proxies for Efficient NAS

AAAI 2025technical

Recent advancements in Zero-shot Neural Architecture Search (NAS) highlight the ability of zero-cost proxies in identifying superior architecture. However, we identify a critical issue with current zero-cost proxies: they aggregate node-wise zero-cost statistics without considering that not all node…

Cited by 7SourcePDFScholar
2025

RA-NeRF: Robust Neural Radiance Field Reconstruction with Accurate Camera Pose Estimation under Complex Trajectories

IROS 2025

Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have emerged as powerful tools for 3D reconstruction and SLAM tasks. However, their performance depends heavily on accurate camera pose priors. Existing approaches attempt to address this issue by introducing external constraints but fal

Cited by 1SourceScholar
2025

SAFormer: Spatially Adaptive Transformer for Efficient and Multi-Resolution Occupancy Prediction

IROS 2025

Accurate and efficient 3D scene understanding from multi-view images remains a fundamental challenge in autonomous driving. Existing methods often struggle with high-dimensional features, leading to excessive computational costs and memory usage. In this paper, we present SAFormer, a novel transform

Cited by 0SourceScholar
2025

STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs

ICLR 2025poster

In this paper, we present the first structural binarization method for LLM compression to less than 1-bit precision. Although LLMs have achieved remarkable performance, their memory-bound nature during the inference stage hinders the adoption of resource-constrained devices. Reducing weights to 1-bi…

2025

TIPS: Tiered Information-Rich Planning Strategy for Efficient AGV Autonomous Exploration

RA-L 2025

In this letter, we propose a tiered systematic framework to enhance the overall efficiency and environmental coverage of autonomous exploration for Unmanned Ground Vehicle (UGV) in complex environments with narrow regions. At the local level, we introduce a novel Multi-cause Triggering Sensor Model

Cited by 0SourceScholar
2025

UnrealLLM: Towards Highly Controllable and Interactable 3D Scene Generation by LLM-powered Procedural Content Generation

ACL 2025finding

The creation of high-quality 3D scenes is essential for applications like video games and simulations, yet automating this process while retaining the benefits of Procedural Content Generation (PCG) remains challenging. In this paper, we introduce UnrealLLM, a novel multi-agent framework that connec…

Cited by 0SourcePDFScholar
2024

Beyond Prompt Learning: Continual Adapter for Efficient Rehearsal-Free Continual Learning

ECCV 2024poster

"The problem of Rehearsal-Free Continual Learning (RFCL) aims to continually learn new knowledge while preventing forgetting of the old knowledge, without storing any old samples and prototypes. The latest methods leverage large-scale pre-trained models as the backbone and use key-query matching to…

Cited by 12SourcePDFScholar
2024

CF-NeRF: Camera Parameter Free Neural Radiance Fields with Incremental Learning

AAAI 2024technical

Neural Radiance Fields have demonstrated impressive performance in novel view synthesis. However, NeRF and most of its variants still rely on traditional complex pipelines to provide extrinsic and intrinsic camera parameters, such as COLMAP. Recent works, like NeRFmm, BARF, and L2G-NeRF, directly tr…

Cited by 10SourcePDFScholar
2024

DOCTR: Disentangled Object-Centric Transformer for Point Scene Understanding

AAAI 2024technical

Point scene understanding is a challenging task to process real-world scene point cloud, which aims at segmenting each object, estimating its pose, and reconstructing its mesh simultaneously. Recent state-of-the-art method first segments each object and then processes them independently with multipl…

2024

DVI-SLAM: A Dual Visual Inertial SLAM Network

ICRA 2024poster

Recent deep learning based visual simultaneous localization and mapping (SLAM) methods have made significant progress. However, how to make full use of visual information as well as better integrate with inertial measurement unit (IMU) in visual SLAM has potential research value. This paper proposes…

Cited by 14SourceScholar
2024

Discovering Sparsity Allocation for Layer-wise Pruning of Large Language Models

NeurIPS 2024poster

In this paper, we present DSA, the first automated framework for discovering sparsity allocation schemes for layer-wise pruning in Large Language Models (LLMs). LLMs have become increasingly powerful, but their large parameter counts make them computationally expensive. Existing pruning methods fo…

Cited by 10SourcePDFScholar
2024

Discovering Syntactic Interaction Clues for Human-Object Interaction Detection

CVPR 2024poster

Recently Vision-Language Model (VLM) has greatly advanced the Human-Object Interaction (HOI) detection. The existing VLM-based HOI detectors typically adopt a hand-crafted template (e.g. a photo of a person [action] a/an [object]) to acquire text knowledge through the VLM text encoder. However such…

Cited by 5SourcePDFScholar
2024

Efficient Learning on Successive Test Time Augmentation

ICASSP 2024accepted

Test time augmentation (TTA) has been a promising tool for improving the robustness against out-of-distribution data at inference time. Recent TTA methods try to learn predictive transformations which are supposed to provide the best performance gain on each test sample. However, existing methods ar…

Cited by 0SourceScholar
2024

Identifying Expert Behavior in Offline Training Datasets Improves Behavioral Cloning of Robotic Manipulation Policies

RA-L 2024

This letter presents our solution for the Real Robot Challenge III <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> , aiming to address dexterous robotic manipulation tasks through learning from offline data. In this competition, participants wer

Cited by 13SourcecodeScholar
2024

Is Your HD Map Constructor Reliable under Sensor Corruptions?

NeurIPS 2024poster

Driving systems often rely on high-definition (HD) maps for precise environmental information, which is crucial for planning and navigation. While current HD map constructors perform well under ideal conditions, their resilience to real-world challenges, \eg, adverse weather and sensor failures, is…

Cited by 17SourcePDFScholar
2024

LPZero: Language Model Zero-cost Proxy Search from Zero

EMNLP 2024finding

Despite the outstanding performance, Neural Architecture Search (NAS) is criticized for massive computation. Recently, Zero-shot NAS has emerged as a promising approach by exploiting Zero-cost (ZC) proxies, which markedly reduce computational demands. Despite this, existing ZC proxies heavily rely o…

Cited by 2SourcePDFScholar
2024

Learning-Based Multimodal Control for a Supernumerary Robotic System in Human-Robot Collaborative Sorting

RA-L 2024

In this letter, a multi-modal learning and control framework is proposed for the control of a supernumerary robotic limb (SRL). The SRL is a wearable robotic arm designed to enhance the manipulation capabilities of its human user and extend the workspace by reaching greater heights. The multi-modal

Cited by 12SourceScholar
2024

Mask-Homo: Pseudo Plane Mask-Guided Unsupervised Multi-Homography Estimation

AAAI 2024technical

Homography estimation is a fundamental problem in computer vision. Previous works mainly focus on estimating either a single homography, or multiple homographies based on mesh grid division of the image. In practical scenarios, single homography is inadequate and often leads to a compromised result…

2024

MmAP: Multi-Modal Alignment Prompt for Cross-Domain Multi-Task Learning

AAAI 2024technical

Multi-Task Learning (MTL) is designed to train multiple correlated tasks simultaneously, thereby enhancing the performance of individual tasks. Typically, a multi-task network structure consists of a shared backbone and task-specific decoders. However, the complexity of the decoders increases with t…

Cited by 69SourcePDFScholar
2024

Pruner-Zero: Evolving Symbolic Pruning Metric From Scratch for Large Language Models

ICML 2024poster

Despite the remarkable capabilities, Large Language Models (LLMs) face deployment challenges due to their extensive size. Pruning methods drop a subset of weights to accelerate, but many of them require retraining, which is prohibitively expensive and computationally demanding. Recently, post-traini…

2024

Residual Denoising Diffusion Models

CVPR 2024poster

We propose residual denoising diffusion models (RDDM) a novel dual diffusion process that decouples the traditional single denoising diffusion process into residual diffusion and noise diffusion. This dual diffusion framework expands the denoising-based diffusion models initially uninterpretable for…

2024

Revisiting the Self-Consistency Challenges in Multi-Choice Question Formats for Large Language Model Evaluation

COLING 2024main

Multi-choice questions (MCQ) are a common method for assessing the world knowledge of large language models (LLMs), demonstrated by benchmarks such as MMLU and C-Eval. However, recent findings indicate that even top-tier LLMs, such as ChatGPT and GPT4, might display inconsistencies when faced with s…

Cited by 8SourcePDFScholar
2024

Robust Learning-Based Incipient Slip Detection Using the PapillArray Optical Tactile Sensor for Improved Robotic Gripping

RA-L 2024

The ability to detect slip, particularly incipient slip, enables robotic systems to take corrective measures to prevent a grasped object from being dropped. Therefore, slip detection can enhance the overall security of robotic gripping. However, accurately detecting incipient slip remains a signific

Cited by 9SourceScholar
2024

VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene Understanding

AAAI 2024technical

Large-scale pre-trained models have achieved remarkable success in various computer vision tasks. A standard approach to leverage these models is to fine-tune all model parameters for downstream tasks, which poses challenges in terms of computational and storage costs. Recently, inspired by Natural…

Cited by 63SourcePDFScholar
2023

BadTrack: A Poison-Only Backdoor Attack on Visual Object Tracking

NeurIPS 2023poster

Visual object tracking (VOT) is one of the most fundamental tasks in computer vision community. State-of-the-art VOT trackers extract positive and negative examples that are used to guide the tracker to distinguish the object from the background. In this paper, we show that this characteristic can b…

Cited by 6SourcePDFScholar
2023

Hybrid-Regressive Paradigm for Accurate and Speed-Robust Neural Machine Translation

ACL 2023findings

This work empirically confirms that non-autoregressive translation (NAT) is less robust in decoding batch size and hardware settings than autoregressive translation (AT). To address this issue, we demonstrate that prompting a small number of AT predictions can significantly reduce the performance ga…

2023

Improving Behavioural Cloning with Positive Unlabeled Learning

CoRL 2023poster

Learning control policies offline from pre-recorded datasets is a promising avenue for solving challenging real-world problems. However, available datasets are typically of mixed quality, with a limited number of the trajectories that we would consider as positive examples; i.e., high-quality demons…

Cited by 8SourceScholar
2023

Rethinking Disparity: A Depth Range Free Multi-View Stereo Based on Disparity

AAAI 2023technical

Existing learning-based multi-view stereo (MVS) methods rely on the depth range to build the 3D cost volume and may fail when the range is too large or unreliable. To address this problem, we propose a disparity-based MVS method based on the epipolar disparity flow (E-flow), called DispMVS, which in…

2023

SketchKnitter: Vectorized Sketch Generation with Diffusion Models

ICLR 2023top-25%

We show vectorized sketch generation can be identified as a reversal of the stroke deformation process. This relationship was established by means of a diffusion model that learns data distributions over the stroke-point locations and pen states of real human sketches. Given randomly scattered strok…

Cited by 31SourcePDFScholar
2023

Towards Reliable Neural Machine Translation with Consistency-Aware Meta-Learning

AAAI 2023technical

Neural machine translation (NMT) has achieved remarkable success in producing high-quality translations. However, current NMT systems suffer from a lack of reliability, as their outputs that are often affected by lexical or syntactic changes in inputs, resulting in large variations in quality. This…

2023

TrajectoryFormer: 3D Object Tracking Transformer with Predictive Trajectory Hypotheses

ICCV 2023poster

3D multi-object tracking (MOT) is vital for many applications including autonomous driving vehicles and service robots. With the commonly used tracking-by-detection paradigm, 3D MOT has made important progress in recent years. However, these methods only use the detection boxes of the current frame…

Cited by 15PDFcodeScholar
2023

Understanding and Improving the Robustness of Terminology Constraints in Neural Machine Translation

ACL 2023long

In this work, we study the robustness of two typical terminology translation methods: Placeholder (PH) and Code-Switch (CS), concerning (1) the number of constraints and (2) the target constraint length. We identify that existing terminology constraint test sets, such as IATE, Wiktionary, and TICO,…

2022

Adaptive Learning Attention Network for Underwater Image Enhancement

RA-L 2022

Underwater images suffer from color casts and low illumination due to the scattering and absorption of light as it propagates in water. These problems can interfere with underwater vision tasks, such as recognition and detection. We propose an adaptive learning attention network for underwater image

Cited by 112SourcecodeScholar
2022

Attention-guided RGB-D Fusion Network for Category-level 6D Object Pose Estimation

IROS 2022poster

This work focuses on estimating 6D poses and sizes of category-level objects from a single RGB-D image. How to exploit the complementary RGB and depth features plays an important role in this task yet remains an open question. Due to the large intra-category texture and shape variations, an object i…

Cited by 5SourceScholar
2022

DH-LC: Hierarchical Matching and Hybrid Bundle Adjustment Towards Accurate and Robust Loop Closure

IROS 2022poster

A loop closure module plays an important role in visual SLAM systems, which can reduce the accumulat-ed drift. This task faces the challenges of large viewpoint changes and expensive computational costs when optimizing the global map. This paper proposes DH-LC, a novel accurate and robust loop closu…

Cited by 0SourceScholar
2022

EASNet: Searching Elastic and Accurate Network Architecture for Stereo Matching

ECCV 2022poster

"Recent advanced studies have spent considerable human efforts on optimizing network architectures for stereo matching but hardly achieved both high accuracy and fast inference speed. To ease the workload in network design, neural architecture search (NAS) has been applied with great success to vari…

2022

FlowFormer: A Transformer Architecture for Optical Flow

ECCV 2022poster

"We introduce optical Flow transFormer, dubbed as FlowFormer, a transformer-based neural network architecture for learning optical flow. FlowFormer tokenizes the 4D cost volume built from an image pair, encodes the cost tokens into a cost memory with alternate-group transformer (AGT) layers in a nov…

2022

Learning Decoupled Retrieval Representation for Nearest Neighbour Neural Machine Translation

COLING 2022main

K-Nearest Neighbor Neural Machine Translation (kNNMT) successfully incorporates external corpus by retrieving word-level representations at test time. Generally, kNNMT borrows the off-the-shelf context representation in the translation task, e.g., the output of the last decoder layer, as the query v…

Cited by 5SourcePDFScholar
2022

Learning a Structured Latent Space for Unsupervised Point Cloud Completion

CVPR 2022oral

Unsupervised point cloud completion aims at estimating the corresponding complete point cloud of a partial point cloud in an unpaired manner. It is a crucial but challenging problem since there is no paired partial-complete supervision that can be exploited directly. In this work, we propose a novel…

Cited by 53PDFScholar
2022

Learning from Students: Online Contrastive Distillation Network for General Continual Learning

IJCAI 2022poster

The goal of General Continual Learning (GCL) is to preserve learned knowledge and learn new knowledge with constant memory from an infinite data stream where task boundaries are blurry. Distilling the model's response of reserved samples between the old and the new models is an effective way to achi…

2022

Runtime Safety Assurance for Learning-enabled Control of Autonomous Driving Vehicles

ICRA 2022poster

Providing safety guarantees for Autonomous Vehicle (AV) systems with machine-learning based controllers remains a challenging issue. In this work, we propose Simplex-Drive, a framework that can achieve runtime safety assurance for machine-learning enabled controllers of AVs. The proposed Simplex-Dri…

Cited by 25SourceScholar
2022

S2G2: Semi-Supervised Semantic Bird-Eye-View Grid-Map Generation Using a Monocular Camera for Autonomous Driving

RA-L 2022

Semantic bird-eye-view (BEV) grid map is a straightforward data representation for semantic environment perception. It can be conveniently integrated with downstream tasks, such as motion planning, trajectory prediction, etc. Most existing methods of semantic BEV grid-map generation adopt supervised

Cited by 17SourceScholar
2021

Accurate Visual-Inertial SLAM by Feature Re-identification

IROS 2021poster

Most of the state-of-the-art visual inertial SLAM methods pay less attention to 2D-2D and 3D-2D matching with more reliable features in a long time span, which easily results in continuous estimation drift. In this paper, we propose an efficient drift-free visual-inertial SLAM method by a pose guide…

Cited by 5SourceScholar
2021

Accurate Visual-Inertial SLAM by Manhattan Frame Re-identification

IROS 2021poster

Most of the state-of-the-art visual-inertial SLAM methods pay less attention to the scene structure of man-made environments. In this paper, based on the assumption of multiple local Manhattan worlds (MWs), we propose a Manhattan frame (MF) re-identification method to build relative rotation constra…

Cited by 8SourceScholar
2021

Dynamic Rebalancing Dockless Bike-Sharing System based on Station Community Discovery

IJCAI 2021poster

Influenced by the era of the sharing economy and mobile payment, Dockless Bike-Sharing System (Dockless BSS) is expanding in many major cities. The mobility of users constantly leads to supply and demand imbalance, which seriously affects the total profit and customer satisfaction. In this paper, we…

Cited by 6SourcePDFScholar
2021

EDNet: Efficient Disparity Estimation With Cost Volume Combination and Attention-Based Spatial Residual

CVPR 2021poster

Existing state-of-the-art disparity estimation works mostly leverage the 4D concatenation volume and construct a very deep 3D convolution neural network (CNN) for disparity regression, which is inefficient due to the high memory consumption and slow inference speed. In this paper, we propose a netwo…

Cited by 23PDFScholar
2021

Learning Generalized Intersection Over Union for Dense Pixelwise Prediction

ICML 2021spotlight

Intersection over union (IoU) score, also named Jaccard Index, is one of the most fundamental evaluation methods in machine learning. The original IoU computation cannot provide non-zero gradients and thus cannot be directly optimized by nowadays deep learning methods. Several recent works generaliz…

Cited by 33SourcePDFScholar
2021

UASNet: Uncertainty Adaptive Sampling Network for Deep Stereo Matching

ICCV 2021poster

Recent studies have shown that cascade cost volume can play a vital role in deep stereo matching to achieve high resolution depth map with efficient hardware usage. However, how to construct good cascade volume as well as effective sampling for them are still under in-depth study. Previous cascade-b…

Cited by 31PDFScholar
2020

FADNet: A Fast and Accurate Network for Disparity Estimation

ICRA 2020poster

Deep neural networks (DNNs) have achieved great success in the area of computer vision. The disparity estimation problem tends to be addressed by DNNs which achieve much better prediction accuracy in stereo matching than traditional hand-crafted feature based methods. On one hand, however, the desig…

Cited by 100SourcecodeScholar
2020

Force-Guided High-Precision Grasping Control of Fragile and Deformable Objects Using sEMG-Based Force Prediction

RA-L 2020

Regulating contact forces with high precision is crucial for grasping and manipulating fragile or deformable objects. We aim to utilize the dexterity of human hands to regulate the contact forces for robotic hands and exploit human sensory-motor synergies in a wearable and non-invasive way. We extra

Cited by 43SourceScholar
2020

Layer-Wise Multi-View Learning for Neural Machine Translation

COLING 2020main

Traditional neural machine translation is limited to the topmost encoder layer’s context representation and cannot directly perceive the lower encoder layers. Existing solutions usually rely on the adjustment of network architecture, making the calculation more complicated or introducing additional…

2020

Rethinking Performance Estimation in Neural Architecture Search

CVPR 2020poster

Neural architecture search (NAS) remains a challenging problem, which is attributed to the indispensable and time-consuming component of performance estimation (PE). In this paper, we provide a novel yet systematic rethinking of PE in a resource constrained regime, termed budgeted PE (BPE), which pr…

Cited by 35PDFcodeScholar
2019

Anchor Diffusion for Unsupervised Video Object Segmentation

ICCV 2019poster

Unsupervised video object segmentation has often been tackled by methods based on recurrent neural networks and optical flow. Despite their complexity, these kinds of approach tend to favour short-term temporal dependencies and are thus prone to accumulating inaccuracies, which cause drift over time…

Cited by 142PDFcodeScholar
2019

Fast Online Object Tracking and Segmentation: A Unifying Approach

CVPR 2019poster

In this paper we illustrate how to perform both visual object tracking and semi-supervised video object segmentation, in real-time, with a single simple approach. Our method, dubbed SiamMask, improves the offline training procedure of popular fully-convolutional Siamese approaches for object trackin…

Cited by 1768PDFScholar
2019

SiamRPN++: Evolution of Siamese Visual Tracking With Very Deep Networks

CVPR 2019oral

Siamese network based trackers formulate tracking as convolutional feature cross-correlation between target template and searching region. However, Siamese trackers still have accuracy gap compared with state-of-the-art algorithms and they cannot take advantage of feature from deep networks, such as…

Cited by 2748PDFScholar
2018

Distractor-aware Siamese Networks for Visual Object Tracking

ECCV 2018poster

Recently, Siamese networks have drawn great attention in visual tracking community because of their balanced accuracy and speed. However, features used in most Siamese tracking approaches can only discriminate foreground from the non-semantic backgrounds. The semantic backgrounds are always consider…

2018

Learning Attentions: Residual Attentional Siamese Network for High Performance Online Visual Tracking

CVPR 2018poster

Offline training for object tracking has recently shown great potentials in balancing tracking accuracy and speed. However, it is still difficult to adapt an offline trained model to a target tracked online. This work presents a Residual Attentional Siamese Network (RASNet) for high performance obje…

2018

Visual Tracking via Spatially Aligned Correlation Filters Network

ECCV 2018poster

Correlation filters based trackers rely on a periodic assumption of the search sample to efficiently distinguish the target from the background. This assumption however yields undesired boundary effects and restricts aspect ratios of search samples. To handle these issues, an end-to-end deep archite…

2017

Robust Object Tracking Based on Temporal and Spatial Deep Networks

ICCV 2017poster

Recently deep neural networks have been widely employed to deal with the visual tracking problem. In this work, we present a new deep architecture which incorporates the temporal and spatial information to boost the tracking performance. Our deep architecture contains three networks, a Feature Net,…

Cited by 60PDFScholar