← Search

zhe wang

110 accepted papers

2026

An Integrated Electrohydraulic Soft Robotic Fish With 3D Maneuverability and Autonomous Control

RA-L 2026

Soft robots enable compliant interaction with humans and the environment. However, their widespread deployment is constrained by significant challenges, including limited mobility and the inherent complexity of incorporating power and control systems into their bodies. In this paper, we present a fu

Cited by 0SourceScholar
2026

Decoding Multi-Finger Motions and Grasp Types with Grasp-Specific Models and Lightmyography Based Muscle-Machine Interfaces

ICRA 2026poster

Efficiently decoding human movement and/or intention is essential for controlling advanced prosthetic and robotic systems. Various muscle-machine interfaces have been researched for this purpose, including electromyography and lightmyography based interfaces. However, the decoding effectiveness of l…

Cited by 0Scholar
2026

Dens3R: A Foundation Model for 3D Geometry Prediction

ICLR 2026poster

Recent advances in dense 3D reconstruction have led to significant progress, yet achieving accurate unified geometric prediction remains a major challenge. Most existing methods are limited to predicting a single geometry quantity from input images. However, geometric quantities such as depth, surfa…

Cited by 0SourcecodeScholar
2026

Gradient as Conditions: Rethinking HOG for All-in-one Image Restoration

AAAI 2026technical

All-in-one image restoration (AIR) aims to address diverse degradations within a unified model by leveraging informative degradation conditions to guide the restoration process. However, existing methods often rely on implicitly learned priors, which may entangle feature representations and hinder p

Cited by 0SourcePDFScholar
2026

Keypoint-Based Dynamic Object 6-DoF Pose Tracking Via Event Camera

ICRA 2026poster

Accurate 6-DoF pose estimation of objects is critical for robots to perform precise manipulation tasks. However, for dynamic object pose estimation, conventional camera-based approaches face several major challenges, such as motion blur, sensor noise, and low-light limitation. To address these issue…

2026

LBA: Textual Hard-Label Adversarial Attack Under Low Query Budgets

IJCAI 2026

Generating high-quality adversarial texts with low query budgets remains a challenging problem in the hard-label scenario. Most existing approaches rely on greedy algorithms, where one position in the text is selected for substitution, followed by the substitutions of other positions. This local sea

Cited by 0Scholar
2026

Latent-Guided Reasoning: Empowering Small LLMs with Large-Model Thinking

ICLR 2026poster

Large Language Models (LLMs) have demonstrated remarkable capabilities in complex reasoning tasks, but their high computational costs limit their widespread practical application. We argue that this inefficiency arises from the tight coupling of high-level cognitive planning (devising the solution s…

Cited by 0SourceScholar
2026

Mixed Reality-Based, Immersive, Semi-Autonomous Robotic Telemanipulation for the Execution of Peg-In-Hole Tasks

ICRA 2026poster

Semi-autonomy in telemanipulation frameworks has the potential to reduce user cognitive load while preserving human perceptual oversight and decision-making capabilities. However, existing semi-autonomous telemanipulation systems are heavily dependent on calibration and hardware configurations, maki…

Cited by 0Scholar
2026

MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts

CVPR 2026

Recent advances in language and vision have demonstrated that scaling up model capacity consistently improves performance across diverse tasks.In 3D visual geometry reconstruction, large-scale training has likewise proven effective for learning versatile representations.However, further scaling of 3

Cited by 0SourcecodeScholar
2026

Rethinking Convergence in MoE Training: The Role of Routing Sparsity

ICML 2026poster

In Mixture-of-Experts (MoE) training, sparse routing, i.e., activating only the top-$K$ experts per token, is essential for balancing convergence speed and computational cost. However, existing works typically choose $K$ empirically, without theoretical guidance. To address this gap, we characterize…

Cited by 0SourceScholar
2026

The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable Reward

ICLR 2026poster

A central paradox in fine-tuning Large Language Models (LLMs) with Reinforcement Learning with Verifiable Reward (RLVR) is the frequent degradation of multi-attempt performance (Pass@k) despite improvements in single-attempt accuracy (Pass@1). This is often accompanied by catastrophic forgetting, wh…

Cited by 0SourceScholar
2026

TinyVPR: Distilling Correct and Confusing Knowledge for Lightweight Visual Place Recognition

ICRA 2026poster

Visual Place Recognition (VPR) is a key technology in autonomous driving, robotics, and augmented reality, requiring efficient and robust localization in large-scale environments. However, most existing methods rely on heavy deep models that are computationally expensive and difficult to deploy on e…

Cited by 0Scholar
2026

Variational Learning for Insertion-based Generation

ICML 2026spotlight

Non-monotonic sequence generation methods, such as masked diffusion models, provide a flexible alternative to left-to-right autoregressive modeling by allowing tokens to be generated in non-fixed and prescribed orders. Despite their practical advantages, most existing non-monotonic models are order-…

Cited by 0SourceScholar
2025

Confusion is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMs

EMNLP 2025

With the development of Large Language Models (LLMs), numerous efforts have revealed their vulnerabilities to jailbreak attacks. Although these studies have driven the progress in LLMs’ safety alignment, it remains unclear whether LLMs have internalized authentic knowledge to deal with real-world cr

2025

CoopDETR: A Unified Cooperative Perception Framework for 3D Detection via Object Query

ICRA 2025

Cooperative perception enhances the individual perception capabilities of autonomous vehicles (AVs) by providing a comprehensive view of the environment. However, balancing perception performance and transmission costs remains a significant challenge. Current approaches that transmit regionlevel fea

Cited by 9SourceScholar
2025

DeepMatch: Navigating the Complexities of Underwater Textures for Enhanced Keypoint Matching

ICASSP 2025accepted

Driven by demands for oceanic exploration, advancements in 3D visual tasks based on video frames are essential. Keypoint matching, essential for camera pose and motion, is hindered by the unique challenges of underwater imagery, such as sparse and repetitive textures. To tackle these issues, we intr…

Cited by 0SourceScholar
2025

Defining and Discovering Hyper-meta-paths for Heterogeneous Hypergraphs

NeurIPS 2025poster

Heterogeneous hypergraph is a kind of structural data that contains multiple types of nodes and multiple types of hyperedges. Each hyperedge type corresponds to a specific multi-ary relation (called hyper-relation) among subsets of nodes, which goes beyond traditional pair-wise relations in simple g…

Cited by 0SourcecodeScholar
2025

Enhancing NLU in Large Language Models Using Adversarial Noisy Instruction Tuning

AAAI 2025technical

Instruction tuning has emerged as an effective approach that notably improves large language models (LLMs) performance, showing particular promise in natural language generation tasks by producing more diverse, coherent, and task-relevant outputs. However, extending instruction tuning to natural lan…

Cited by 0SourcePDFScholar
2025

FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers

ICCV 2025poster

In light of recent breakthroughs in text-to-image (T2I) generation, particularly with diffusion transformers (DiT), subject-driven technologies are increasingly being employed for high-fidelity customized production that preserves subject identity from reference inputs, enabling thrilling design wor…

2025

IROAM: Improving Roadside Monocular 3D Object Detection Learning from Autonomous Vehicle Data Domain

ICRA 2025

In autonomous driving, The perception capabilities of the ego-vehicle can be improved with roadside sensors, which can provide a holistic view of the environment. However, existing monocular detection methods designed for vehicle cameras are not suitable for roadside cameras due to viewpoint domain

Cited by 0SourceScholar
2025

Implicit and Explicit Rule Injection for Complex Query Answering over Knowledge Graphs

ICASSP 2025accepted

Complex Query Answering over incomplete knowledge graphs is a fundamental yet challenging task. Existing methods based on a pretrained knowledge graph embedding model have achieved good performance. However, they ignore logical rules. Logical rules, as part of the conceptual layer in knowledge graph…

Cited by 0SourceScholar
2025

Learnable Retrieval Enhanced Visual-Text Alignment and Fusion for Radiology Report Generation

ICCV 2025poster

Automated radiology report generation is essential for improving diagnostic efficiency and reducing the workload of medical professionals. However, existing methods face significant challenges, such as disease class imbalance and insufficient cross-modal fusion. To address these issues, we propose t…

2025

Learning-Order Autoregressive Models with Application to Molecular Graph Generation

ICML 2025poster

Autoregressive models (ARMs) have become the workhorse for sequence generation tasks, since many problems can be modeled as next-token prediction. While there appears to be a natural ordering for text (i.e., left-to-right), for many data types, such as graphs, the canonical ordering is less obvious.…

Cited by 0SourcePDFScholar
2025

MamV2XCalib: V2X-based Target-less Infrastructure Camera Calibration with State Space Model

ICCV 2025poster

As cooperative systems that leverage roadside cameras to assist autonomous vehicle perception become increasingly widespread, large-scale precise calibration of infrastructure cameras has become a critical issue. Traditional manual calibration methods are often time-consuming, labor-intensive, and m…

2025

MultiAgentBench : Evaluating the Collaboration and Competition of LLM agents

ACL 2025long

Large Language Models (LLMs) have shown remarkable capabilities as autonomous agents; yet existing benchmarks either focus on single-agent tasks or are confined to narrow domains, failing to capture the dynamics of multi-agent coordination and competition. In this paper, we introduce MultiAgentBench…

2025

Online Iterative Self-Alignment for Radiology Report Generation

ACL 2025long

Radiology Report Generation (RRG) is an important research topic for relieving radiologists’ heavy workload. Existing RRG models mainly rely on supervised fine-tuning (SFT) based on different model architectures using data pairs of radiological images and corresponding radiologist-annotated reports.…

Cited by 0SourcePDFScholar
2025

PScalpel: A Machine Learning-based Guider for Protein Phase-Separating Behaviour Alteration

AAAI 2025technical

Missense mutations could affect the Liquid-Liquid Phase Separation (LLPS) propensity of proteins and lead to aberrant phase-separating behaviours, which are recently found to be associated with many diseases including Alzheimer's and cancer. However, the regulatory role of mutations in LLPS remains…

2025

PurpCode: Reasoning for Safer Code Generation

NeurIPS 2025poster

We introduce PurpCode, the first post-training recipe for training safe code reasoning models towards generating secure code and defending against malicious cyberactivities. PurpCode trains a reasoning model in two stages: (i) Rule Learning, which explicitly teaches the model to reference cybersafet…

Cited by 0SourceScholar
2025

Radiology Report Generation via Multi-objective Preference Optimization

AAAI 2025technical

Automatic Radiology Report Generation (RRG) is an important topic for alleviating the substantial workload of radiologists. Existing RRG approaches rely on supervised regression based on different architectures or additional knowledge injection, while the generated report may not align optimally wit…

Cited by 2SourcePDFScholar
2025

Renderworld: World Model with Self-Supervised 3D Label

ICRA 2025

End-to-end autonomous driving with vision-only is not only more cost-effective compared to LiDAR-vision fusion but also more reliable than traditional methods. To achieve a economical and robust purely visual autonomous driving system, we propose RenderWorld, a vision-only end-to-end autonomous driv

Cited by 47SourceScholar
2025

Rule-Guided Graph Neural Networks for Explainable Knowledge Graph Reasoning

AAAI 2025technical

The connections between symbolic rules and neural networks have been explored in various directions, including rule mining through neural networks and rule-based explanation for neural networks. These approaches allow symbolic rules to be extracted from neural network models, which offers explainabi…

2025

Semantic-guided Masked Mutual Learning for Multi-modal Brain Tumor Segmentation with Arbitrary Missing Modalities

AAAI 2025technical

Malignant brain tumors have become an aggressive and dangerous disease that leads to death worldwide. Multi-modal MRI data is crucial for accurate brain tumor segmentation, but missing modalities common in clinical practice can severely degrade the segmentation performance. While incomplete multi-mo…

Cited by 0SourcePDFScholar
2025

TacoDepth: Towards Efficient Radar-Camera Depth Estimation with One-stage Fusion

CVPR 2025award

Radar-Camera depth estimation aims to predict dense and accurate metric depth by fusing input images and Radar data. Model efficiency is crucial for this task in pursuit of real-time processing on autonomous vehicles and robotic platforms. However, due to the sparsity of Radar returns, the prevailin…

2025

TurboFuzzLLM: Turbocharging Mutation-based Fuzzing for Effectively Jailbreaking Large Language Models in Practice

NAACL 2025industry

Jailbreaking large-language models (LLMs) involves testing their robustness against adversarial prompts and evaluating their ability to withstand prompt attacks that could elicit unauthorized or malicious responses. In this paper, we present TurboFuzzLLM, a mutation-based fuzzing technique for effic…

2025

Wave-wise Discriminative Tracking by Phase-Amplitude Separation, Augmentation and Mixture

IJCAI 2025

Distinguishing key features in complex visual tasks is challenging. A novel approach treats image patches (tokens) as waves. By using both phase and amplitude, it captures richer semantics and specific invariances compared to pixel-based methods, and allows for feature fusion across regions for a ho

Cited by 0SourcePDFScholar
2024

AT4CTR: Auxiliary Match Tasks for Enhancing Click-Through Rate Prediction

AAAI 2024technical

Click-through rate (CTR) prediction is a vital task in industrial recommendation systems. Most existing methods focus on the network architecture design of the CTR model for better accuracy and suffer from the data sparsity problem. Especially in industrial recommendation systems, the widely applied…

Cited by 9SourcePDFScholar
2024

Advancing Medical Image Segmentation via Self-supervised Instance-adaptive Prototype Learning

IJCAI 2024poster

Medical Image Segmentation (MIS) plays a crucial role in medical therapy planning and robot navigation. Prototype learning methods in MIS focus on generating segmentation masks through pixel-to-prototype comparison. However, current approaches often overlook sample diversity by using a fixed prototy…

Cited by 0SourcePDFScholar
2024

Attention Calibration for Disentangled Text-to-Image Personalization

CVPR 2024poster

Recent thrilling progress in large-scale text-to-image (T2I) models has unlocked unprecedented synthesis quality of AI-generated content (AIGC) including image generation 3D and video composition. Further personalized techniques enable appealing customized production of a novel concept given only se…

2024

Bayesian Calibration of Win Rate Estimation with LLM Evaluators

EMNLP 2024main

Recent advances in large language models (LLMs) show the potential of using LLMs as evaluators for assessing the quality of text generations from LLMs. However, applying LLM evaluators naively to compare different systems can lead to unreliable results due to the inaccuracy and intrinsic bias of LLM…

2024

Cross-Modal Feature Distribution Calibration for Few-Shot Visual Question Answering

AAAI 2024technical

Few-shot Visual Question Answering (VQA) realizes few-shot cross-modal learning, which is an emerging and challenging task in computer vision. Currently, most of the few-shot VQA methods are confined to simply extending few-shot classification methods to cross-modal tasks while ignoring the spatial…

Cited by 3SourcePDFScholar
2024

EMIFF: Enhanced Multi-scale Image Feature Fusion for Vehicle-Infrastructure Cooperative 3D Object Detection

ICRA 2024poster

In autonomous driving, cooperative perception makes use of multi-view cameras from both vehicles and infrastructure, providing a global vantage point with rich semantic context of road conditions beyond a single vehicle viewpoint. Currently, two major challenges persist in vehicle-infrastructure coo…

Cited by 7SourcecodeScholar
2024

Enhancing Learning-Based Binary Code Similarity Detection Model through Adversarial Training with Multiple Function Variants

EMNLP 2024finding

Compared to identifying binary versions of the same function under different compilation options, existing Learning-Based Binary Code Similarity Detection (LB-BCSD) methods exhibit lower accuracy in recognizing functions with the same functionality but different implementations. To address this issu…

Cited by 0SourcePDFScholar
2024

Idempotence and Perceptual Image Compression

ICLR 2024spotlight

Idempotence is the stability of image codec to re-compression. At the first glance, it is unrelated to perceptual image compression. However, we find that theoretically: 1) Conditional generative model-based perceptual codec satisfies idempotence; 2) Unconditional generative model with idempotence c…

2024

M3: A Multi-Task Mixed-Objective Learning Framework for Open-Domain Multi-Hop Dense Sentence Retrieval

COLING 2024main

In recent research, contrastive learning has proven to be a highly effective method for representation learning and is widely used for dense retrieval. However, we identify that relying solely on contrastive learning can lead to suboptimal retrieval performance. On the other hand, despite many retri…

2024

MPOD123: One Image to 3D Content Generation Using Mask-enhanced Progressive Outline-to-Detail Optimization

CVPR 2024poster

Recent advancements in single image driven 3D content generation have been propelled by leveraging prior knowledge from pretrained 2D diffusion models. However the 3D content generated by existing methods often exhibits distorted outline shapes and inadequate details. To solve this problem we propos…

Cited by 1SourcePDFScholar
2024

Magicoder: Empowering Code Generation with OSS-Instruct

ICML 2024poster

We introduce Magicoder, a series of fully open-source (code, weights, and data) Large Language Models (LLMs) for code that significantly closes the gap with top code models while having no more than 7B parameters. Magicoder models are trained on 75K synthetic instruction data using **OSS-Instruct**,…

2024

Powerful Multidirectional Pneumatic Jumper With Lightweight Fabric Chambers and Buckling-Controllable Elastic Beams

RA-L 2024

Jumping is an advantageous locomotion method to extend the motion range, overcome obstacles, and adapt to unstructured environments; however, it is challenging for robotic jumpers to simultaneously maintain structural compliance and achieve powerful, multidirectional jumping. This letter introduces

Cited by 6SourceScholar
2024

RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding

CVPR 2024poster

We propose a lightweight and scalable Regional Point-Language Contrastive learning framework namely RegionPLC for open-world 3D scene understanding aiming to identify and recognize open-set objects and categories. Specifically based on our empirical studies we introduce a 3D-aware SFusion strategy t…

2024

S-DyRF: Reference-Based Stylized Radiance Fields for Dynamic Scenes

CVPR 2024poster

Current 3D stylization methods often assume static scenes which violates the dynamic nature of our real world. To address this limitation we present S-DyRF a reference-based spatio-temporal stylization method for dynamic neural radiance fields. However stylizing dynamic 3D scenes is inherently chall…

Cited by 4SourcePDFScholar
2024

Self-Supervised Class-Agnostic Motion Prediction with Spatial and Temporal Consistency Regularizations

CVPR 2024poster

The perception of motion behavior in a dynamic environment holds significant importance for autonomous driving systems wherein class-agnostic motion prediction methods directly predict the motion of the entire point cloud. While most existing methods rely on fully-supervised learning the manual labe…

2024

Semi-supervised Class-Agnostic Motion Prediction with Pseudo Label Regeneration and BEVMix

AAAI 2024technical

Class-agnostic motion prediction methods aim to comprehend motion within open-world scenarios, holding significance for autonomous driving systems. However, training a high-performance model in a fully-supervised manner always requires substantial amounts of manually annotated data, which can be bot…

2024

Simplified and Generalized Masked Diffusion for Discrete Data

NeurIPS 2024poster

Masked (or absorbing) diffusion is actively explored as an alternative to autoregressive models for generative modeling of discrete data. However, existing work in this area has been hindered by unnecessarily complex model formulations and unclear relationships between different perspectives, leadin…

2024

SparseLIF: High-Performance Sparse LiDAR-Camera Fusion for 3D Object Detection

ECCV 2024poster

"Sparse 3D detectors have received significant attention since the query-based paradigm embraces low latency without explicit dense BEV feature construction. However, these detectors achieve worse performance than their dense counterparts. In this paper, we find the key to bridging the performance g…

2024

nuCraft: Crafting High Resolution 3D Semantic Occupancy for Unified 3D Scene Understanding

ECCV 2024poster

"Existing benchmarks for 3D semantic occupancy prediction in autonomous driving are limited by low resolution (up to [512×512×40] with 0.2m voxel size) and inaccurate annotations, hindering the unification of 3D scene understanding through the occupancy representation. Moreover, previous methods can…

Cited by 3SourcePDFScholar
2023

Calibration-Free BEV Representation for Infrastructure Perception

IROS 2023poster

Effective BEV object detection on infrastructure can greatly improve traffic scene understanding and vehicle-to-infrastructure (V2I) cooperative perception. However, cameras installed on infrastructure have various postures, and previous BEV detection methods rely on accurate calibration, which is d…

Cited by 22SourceScholar
2023

ConQueR: Query Contrast Voxel-DETR for 3D Object Detection

CVPR 2023highlight

Although DETR-based 3D detectors simplify the detection pipeline and achieve direct sparse predictions, their performance still lags behind dense detectors with post-processing for 3D object detection from point clouds. DETRs usually adopt a larger number of queries than GTs (e.g., 300 queries v.s.…

2023

Efficient Joint Optimization of Layer-Adaptive Weight Pruning in Deep Neural Networks

ICCV 2023poster

In this paper, we propose a novel layer-adaptive weight-pruning approach for Deep Neural Networks (DNNs) that addresses the challenge of optimizing the output distortion minimization while adhering to a target pruning ratio constraint. Our approach takes into account the collective influence of all…

Cited by 29PDFcodeScholar
2023

Improving Interpretability via Explicit Word Interaction Graph Layer

AAAI 2023technical

Recent NLP literature has seen growing interest in improving model interpretability. Along this direction, we propose a trainable neural network layer that learns a global interaction graph between words and then selects more informative words using the learned word interactions. Our layer, we call…

2023

Summary on the Multimodal Information Based Speech Processing (MISP) 2022 Challenge

ICASSP 2023accepted

The Multimodal Information based Speech Processing (MISP) 2022 challenge aimed to enhance speech processing performance in harsh acoustic environments by leveraging additional modalities such as video or text. The challenge included two tracks: audio-visual speaker diarization (AVSD) and audio-visua…

Cited by 0SourceScholar
2023

The Multimodal Information Based Speech Processing (Misp) 2022 Challenge: Audio-Visual Diarization And Recognition

ICASSP 2023accepted

The Multi-modal Information based Speech Processing (MISP) challenge aims to extend the application of signal processing technology in specific scenarios by promoting the research into wake-up words, speaker diarization, speech recognition, and other technologies. The MISP2022 challenge has two trac…

Cited by 0SourceScholar
2023

Towards Trustworthy Multi-Label Sewer Defect Classification via Evidential Deep Learning

ICASSP 2023accepted

An automatic vision-based sewer inspection plays a key role of sewage system in a modern city. Recent advances focus on utilizing deep learning model to realize the sewer inspection system, benefiting from the capability of data-driven feature representation. However, the inherent uncertainty of sew…

Cited by 0SourceScholar
2023

Weakly Supervised Class-Agnostic Motion Prediction for Autonomous Driving

CVPR 2023poster

Understanding the motion behavior of dynamic environments is vital for autonomous driving, leading to increasing attention in class-agnostic motion prediction in LiDAR point clouds. Outdoor scenes can often be decomposed into mobile foregrounds and static backgrounds, which enables us to associate m…

Cited by 11SourcePDFScholar
2022

A Tensegrity-Based Inchworm-Like Robot for Crawling in Pipes With Varying Diameters

RA-L 2022

Most current in-pipe robots are usually designed for pipes of a specific size. In this letter, we propose a novel inchworm-like in-pipe robot based on the concept of tensegrity for moving in pipes with varying diameters. Firstly, a tensegrity-based robotic module capable of two kinds of shape change

Cited by 34SourceScholar
2022

Beyond Data Samples: Aligning Differential Networks Estimation with Scientific Knowledge

AISTATS 2022poster

Learning the differential statistical dependency network between two contexts is essential for many real-life applications, mostly in the high dimensional low sample regime. In this paper, we propose a novel differential network estimator that allows integrating various sources of knowledge beyond d…

2022

FreGAN: Exploiting Frequency Components for Training GANs under Limited Data

NeurIPS 2022accept

Training GANs under limited data often leads to discriminator overfitting and memorization issues, causing divergent training. Existing approaches mitigate the overfitting by employing data augmentations, model regularization, or attention mechanisms. However, they ignore the frequency bias of GANs…

2022

Learning Versatile Neural Architectures by Propagating Network Codes

ICLR 2022poster

This work explores how to design a single neural network capable of adapting to multiple heterogeneous vision tasks, such as image segmentation, 3D detection, and video recognition. This goal is challenging because both network architecture search (NAS) spaces and methods in different tasks are inco…

2022

PPT: Token-Pruned Pose Transformer for Monocular and Multi-View Human Pose Estimation

ECCV 2022poster

"Recently, the vision transformer and its variants have played an increasingly important role in both monocular and multi-view human pose estimation. Considering image patches as tokens, transformers can model the global dependencies within the entire image or across images from other views. However…

2022

RDO-Q: Extremely Fine-Grained Channel-Wise Quantization via Rate-Distortion Optimization

ECCV 2022poster

"Allocating different bit widths to different channels and quantizing them independently bring higher quantization precision and accuracy. Most of prior works use equal bit width to quantize all layers or channels, which is sub-optimal. On the other hand, it is very challenging to explore the hyperp…

Cited by 9SourcePDFScholar
2022

RigidFlow: Self-Supervised Scene Flow Learning on Point Clouds by Local Rigidity Prior

CVPR 2022poster

In this work, we focus on scene flow learning on point clouds in a self-supervised manner. A real-world scene can be well modeled as a collection of rigidly moving parts, therefore its scene flow can be represented as a combination of rigid motion of each part. Inspired by this observation, we propo…

Cited by 65PDFScholar
2022

ST-MAML : A stochastic-task based method for task-heterogeneous meta-learning

UAI 2022poster

Optimization-based meta-learning typically assumes tasks are sampled from a single distribution - an assumption that oversimplifies and limits the diversity of tasks that meta-learning can model. Handling tasks from multiple distributions is challenging for meta-learning because it adds ambiguity to…

Cited by 10SourcePDFScholar
2022

Self-supervised Heterogeneous Graph Pre-training Based on Structural Clustering

NeurIPS 2022accept

Recent self-supervised pre-training methods on Heterogeneous Information Networks (HINs) have shown promising competitiveness over traditional semi-supervised Heterogeneous Graph Neural Networks (HGNNs). Unfortunately, their performance heavily depends on careful customization of various strategies…

2022

Towards Efficient 3D Object Detection with Knowledge Distillation

NeurIPS 2022accept

Despite substantial progress in 3D object detection, advanced 3D detectors often suffer from heavy computation overheads. To this end, we explore the potential of knowledge distillation (KD) for developing efficient 3D object detectors, focusing on popular pillar- and voxel-based detectors. In the a…

2022

WaveGAN: Frequency-Aware GAN for High-Fidelity Few-Shot Image Generation

ECCV 2022poster

"Existing few-shot image generation approaches typically employ fusion-based strategies, either on the image or the feature level, to produce new images. However, previous approaches struggle to synthesize high-frequency signals with fine details, deteriorating the synthesis quality. To address this…

2021

ACMo: Angle-Calibrated Moment Methods for Stochastic Optimization

AAAI 2021technical

Stochastic gradient descent (SGD) is a widely used method for its outstanding generalization ability and simplicity. Adaptive gradient methods have been proposed to further accelerate the optimization process. In this paper, we revisit existing adaptive gradient optimization methods with a new inter…

2021

AdaStereo: A Simple and Efficient Approach for Adaptive Stereo Matching

CVPR 2021poster

Recently, records on stereo matching benchmarks are constantly broken by end-to-end disparity networks. However, the domain adaptation ability of these deep models is quite poor. Addressing such problem, we present a novel domain-adaptive pipeline called AdaStereo that aims to align multi-level repr…

Cited by 91PDFScholar
2021

Cross-Layer Distillation with Semantic Calibration

AAAI 2021technical

Recently proposed knowledge distillation approaches based on feature-map transfer validate that intermediate layers of a teacher model can serve as effective targets for training a student model to obtain better generalization ability. Existing studies mainly focus on particular representation forms…

2021

PC-HMR: Pose Calibration for 3D Human Mesh Recovery from 2D Images/Videos

AAAI 2021technical

The end-to-end Human Mesh Recovery (HMR) approach has been successfully used for 3D body reconstruction. However, most HMR-based frameworks reconstruct human body by directly learning mesh parameters from images or videos, while lacking explicit guidance of 3D human pose in visual data. As a result,…

Cited by 43SourcePDFScholar
2021

ST3D: Self-Training for Unsupervised Domain Adaptation on 3D Object Detection

CVPR 2021poster

We present a new domain adaptive self-training pipeline, named ST3D, for unsupervised domain adaptation on 3D object detection from point clouds. First, we pre-train the 3D detector on the source domain with our proposed random object scaling strategy for mitigating the negative effects of source do…

Cited by 249PDFcodeScholar
2020

A Generalized Training Approach for Multiagent Learning

ICLR 2020talk

This paper investigates a population-based training regime based on game-theoretic principles called Policy-Spaced Response Oracles (PSRO). PSRO is general in the sense that it (1) encompasses well-known algorithms such as fictitious play and double oracle as special cases, and (2) in principle appl…

Cited by 127SourcecodeScholar
2020

History-Gradient Aided Batch Size Adaptation for Variance Reduced Algorithms

ICML 2020poster

Variance-reduced algorithms, although achieve great theoretical performance, can run slowly in practice due to the periodic gradient estimation with a large batch of data. Batch-size adaptation thus arises as a promising approach to accelerate such algorithms. However, existing schemes either apply…

Cited by 20SourcePDFScholar
2020

Learning Depth-Guided Convolutions for Monocular 3D Object Detection

CVPR 2020poster

3D object detection from a single image without LiDAR is a challenging task due to the lack of accurate depth information. Conventional 2D convolutions are unsuitable for this task because they fail to capture local object and its scale information, which are vital for 3D object detection. To better…

Cited by 384PDFcodeScholar
2020

PV-RCNN: Point-Voxel Feature Set Abstraction for 3D Object Detection

CVPR 2020poster

We present a novel and high-performance 3D object detection framework, named PointVoxel-RCNN (PV-RCNN), for accurate 3D object detection from point clouds. Our proposed method deeply integrates both 3D voxel Convolutional Neural Network (CNN) and PointNet-based set abstraction to learn more discrimi…

Cited by 2428PDFcodeScholar
2020

Proximal Gradient Algorithm with Momentum and Flexible Parameter Restart for Nonconvex Optimization

IJCAI 2020poster

Various types of parameter restart schemes have been proposed for proximal gradient algorithm with momentum to facilitate their convergence in convex optimization. However, under parameter restart, the convergence of proximal gradient algorithm with momentum remains obscure in nonconvex optimization…

Cited by 0SourcePDFScholar
2020

Query Answering for Existential Rules via Efficient Datalog Rewriting

IJCAI 2020poster

Existential rules are an expressive ontology formalism for ontology-mediated query answering and thus query answering is of high complexity, while several tractable fragments have been identified. Existing systems based on first-order rewriting methods can lead to queries too large for DBMS to handl…

2020

SegVoxelNet: Exploring Semantic Context and Depth-aware Features for 3D Vehicle Detection from Point Cloud

ICRA 2020poster

3D vehicle detection based on point cloud is a challenging task in real-world applications such as autonomous driving. Despite significant progress has been made, we observe two aspects to be further improved. First, the semantic context information in LiDAR is seldom explored in previous works, whi…

Cited by 78SourceScholar
2020

ViTAA: Visual-Textual Attributes Alignment in Person Search by Natural Language

ECCV 2020poster

Person search by natural language aims at retrieving a specific person in a large-scale image pool that matches given textual descriptions. While most of the current methods treat the task as a holistic visual and textual feature matching one, we approach it from an attribute-aligning perspective th…

2019

Breast Cancer Detection Based on Merging Four Modes MRI Using Convolutional Neural Networks

ICASSP 2019accepted

The objective of the study is to develop a framework for automatic breast cancer detection with merging four imaging modes. Attempts were made for tumor classification and segmentation; using a multi-parametric Magnetic Resonance Imaging (MRI) method on breast tumors. MRI data of the breast were obt…

Cited by 0SourceScholar
2019

Improved Zeroth-Order Variance Reduced Algorithms and Analysis for Nonconvex Optimization

ICML 2019oral

Two types of zeroth-order stochastic algorithms have recently been designed for nonconvex optimization respectively based on the first-order techniques SVRG and SARAH/SPIDER. This paper addresses several important issues that are still open in these methods. First, all existing SVRG-type zeroth-orde…

2019

MaxpoolNMS: Getting Rid of NMS Bottlenecks in Two-Stage Object Detectors

CVPR 2019poster

Modern convolutional object detectors have improved the detection accuracy significantly, which in turn inspired the development of dedicated hardware accelerators to achieve real-time performance by exploiting inherent parallelism in the algorithm. Non-maximum suppression (NMS) is an indispensable…

Cited by 39PDFScholar
2019

Robust Multi-Modality Multi-Object Tracking

ICCV 2019poster

Multi-sensor perception is crucial to ensure the reliability and accuracy in autonomous driving system, while multi-object tracking (MOT) improves that by tracing sequential movement of dynamic objects. Most current approaches for multi-sensor multi-object tracking are either lack of reliability by…

Cited by 272PDFcodeScholar
2019

SpiderBoost and Momentum: Faster Variance Reduction Algorithms

NeurIPS 2019poster

SARAH and SPIDER are two recently developed stochastic variance-reduced algorithms, and SPIDER has been shown to achieve a near-optimal first-order oracle complexity in smooth nonconvex optimization. However, SPIDER uses an accuracy-dependent stepsize that slows down the convergence in practice, and…

Cited by 213SourcePDFScholar
2019

Stochastic Variance-Reduced Cubic Regularization for Nonconvex Optimization

AISTATS 2019poster

Cubic regularization (CR) is an optimization method with emerging popularity due to its capability to escape saddle points and converge to second-order stationary solutions for nonconvex optimization. However, CR encounters a high sample complexity issue for finite-sum problems with a large data siz…

Cited by 67SourcePDFScholar
2019

Training Multi-task Adversarial Network for Extracting Noise-robust Speaker Embedding

ICASSP 2019accepted

Under noisy environments, to achieve the robust performance of speaker recognition is still a challenging task. Motivated by the promising performance of multi-task training in a variety of image processing tasks, we explore the potential of multitask adversarial training for learning a noise-robust…

Cited by 0SourceScholar
2018

Convergence of Cubic Regularization for Nonconvex Optimization under KL Property

NeurIPS 2018spotlight

Cubic-regularized Newton's method (CR) is a popular algorithm that guarantees to produce a second-order stationary solution for solving nonconvex optimization problems. However, existing understandings of convergence rate of CR are conditioned on special types of geometrical properties of the object…

Cited by 28SourcePDFScholar
2016

Real-Time Action Recognition With Enhanced Motion Vector CNNs

CVPR 2016poster

The deep two-stream architecture exhibited excellent performance on video based action recognition. The most computationally expensive step in this approach comes from the calculation of optical flow which prevents it to be real-time. This paper accelerates this architecture by replacing optical flo…

Cited by 546PDFcodeScholar
2015

DeepID-Net: Deformable Deep Convolutional Neural Networks for Object Detection

CVPR 2015poster

In this paper, we propose deformable deep convolutional neural networks for generic object detection. This new deep learning object detection diagram has innovations in multiple aspects. In the proposed new deep architecture, a new deformation constrained pooling (def-pooling) layer models the defor…

Cited by 612SourcePDFScholar
2015

Linear prediction based comfort noise generation in the EVS codec

ICASSP 2015accepted

A Discontinuous transmission (DTX) system, which is widely adopted in speech codecs, is an important function for speech communication systems that can reduce the transmission bandwidth by at least a half. Within a DTX system, the comfort noise generation (CNG) plays a key role in the overall qualit…

Cited by 0SourceScholar
2015

Overview of the EVS codec architecture

ICASSP 2015accepted

The recently standardized 3GPP codec for Enhanced Voice Services (EVS) offers new features and improvements for low-delay real-time communication systems. Based on a novel, switched low-delay speech/audio codec, the EVS codec contains various tools for better compression efficiency and higher qualit…

Cited by 169SourceScholar