← Search

Hui Zhang

96 accepted papers

2026

Anomagic: Crossmodal Prompt-driven Zero-shot Anomaly Generation

AAAI 2026technical

We propose Anomagic, a zero-shot anomaly generation method that produces semantically coherent anomalies without requiring any exemplar anomalies. By unifying both visual and textual cues through a crossmodal prompt encoding scheme, Anomagic leverages rich contextual information to steer an inpaint

Cited by 0SourcePDFScholar
2026

ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation

ICLR 2026poster

Infographic charts are a powerful medium for communicating abstract data by combining visual elements (e.g., charts, images) with textual information. However, their visual and structural richness poses challenges for large vision-language models (LVLMs), which are typically trained on plain charts.…

Cited by 0SourcecodeScholar
2026

CoLC: Communication-Efficient Collaborative Perception with LiDAR Completion

CVPR 2026

Collaborative perception empowers autonomous agents to share complementary information and overcome perception limitations. While early fusion offers more perceptual complementarity and is inherently robust to model heterogeneity, its high communication cost has limited its practical deployment, pro

Cited by 0SourceScholar
2026

CreatiDesign: A Unified Multi-Conditional Diffusion Transformer for Creative Graphic Design

ICLR 2026poster

Graphic design plays a vital role in visual communication across advertising, marketing, and multimedia entertainment. Prior work has explored automated graphic design generation using diffusion models, aiming to streamline creative workflows and democratize design capabilities. However, complex gra…

Cited by 0SourcecodeScholar
2026

FedSDWC: Federated Synergistic Dual-Representation Weak Causal Learning for OOD

AAAI 2026technical

Amid growing demands for data privacy and advances in computational infrastructure, federated learning (FL) has emerged as a prominent distributed learning paradigm. Nevertheless, differences in data distribution (such as covariate and semantic shifts) severely affect its reliability in real-world d

Cited by 0SourcePDFScholar
2026

From Discriminative to Generative: A Diffusion-Based Paradigm for Multi-Agent Collaborative Perception

AAAI 2026technical

Collaborative perception leveraging intermediate feature fusion has emerged as a leading paradigm to significantly enhance the environmental perception capabilities of autonomous driving systems. However, existing methods typically rely on discriminative supervision guided by downstream tasks. This

Cited by 0SourcePDFScholar
2026

Hybrid Robust Collaborative Perception with LiDAR-4D Radar Fusion under Adverse Weather Conditions

CVPR 2026

Current collaborative perception systems have significantly improved 3D object detection performance. However, widely used LiDAR and camera systems often suffer performance degradation under adverse weather conditions. The weather-robust 4D radar provides a promising solution to address this challen

Cited by 0SourceScholar
2026

LLaVA-MS-PIT: Multi-Modal Schema-Guided Progressive Instruction Tuning for Multi-Modal Event Extraction

AAAI 2026technical

The proliferation of multi-modal data on the internet has intensified the need for structured event understanding across textual and visual modalities. However, existing multi-modal event extraction models suffer from three major limitations: the absence of explicit event schema guidance, coarse-gra

Cited by 0SourcePDFScholar
2026

Path Planning for Mobile Robots Based on Hybrid Sampling and Space-Optimized RRT

RA-L 2026

To enhance the efficiency and safety of mobile robot path planning, this letter proposes a hybrid sampling and space-optimized RRT (HB-RRT) method. First, a hybrid sampling strategy combining Gaussian sampling with parallel sampling is employed to replace conventional random sampling. During the sam

Cited by 0SourceScholar
2026

Primary Visual Cortex Inspired Point Cloud Analysis Framework

AAAI 2026technical

Despite significant advancements in point cloud analysis, reducing energy consumption and improving robustness remain understudied, largely due to the inherent limitations of Convolutional Neural Networks (CNNs). To address this, we take the cue from the primary visual cortex and propose a Dendritic

Cited by 0SourcePDFScholar
2026

Remember Me: Bridging the Long-Range Gap in LVLMs with Three-Step Inference-Only Decay Resilience Strategies

AAAI 2026technical

Large Vision-Language Models (LVLMs) have achieved impressive performance across a wide range of multimodal tasks. However, they still face critical challenges in modeling long-range dependencies under the usage of Rotary Positional Encoding (ROPE). Although it can facilitate precise modeling of tok

Cited by 0SourcePDFScholar
2026

SGDE: Self-supervised Geometry Degradation Estimation Framework for Coded Aperture Compressive Spectral Imaging

CVPR 2026

Coded Aperture Snapshot Spectral Imaging (CASSI) has emerged as a prominent technique for efficient hyperspectral imaging. However, the tight coupling between physical encoding and computational decoding makes CASSI highly sensitive to slight hardware misalignments, which can significantly degrade r

Cited by 0SourcecodeScholar
2026

StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answering

CVPR 2026

Knowledge-based Visual Question Answering (KVQA) requires models to ground entities in images and reason over factual knowledge. Recent work has introduced its implicit-knowledge variant, IK-KVQA, where a multimodal large language model (MLLM) is the sole knowledge source and answers are produced wi

Cited by 0SourceScholar
2026

Toward Robust Collaborative Perception under Adverse Weather Conditions Via Dual-Branch Network

ICRA 2026poster

Recent advances in collaborative perception systems have led to significant improvements in 3D object detection performance. While widely deployed LiDAR and camera systems often experience performance degradation under adverse weather conditions, weather-robust 4D radar offers a promising alternativ…

Cited by 0Scholar
2026

Towards High-Resolution 3D Anomaly Detection: A Scalable Dataset and Real-Time Framework for Subtle Industrial Defects

AAAI 2026technical

In industrial point cloud analysis, detecting subtle anomalies demands high-resolution spatial data, yet prevailing benchmarks emphasize low-resolution inputs. To address this disparity, we propose a scalable pipeline for generating realistic and subtle 3D anomalies. Employing this pipeline, we deve

Cited by 0SourcePDFScholar
2026

VGGS: VGGT-guided Gaussian Splatting for Efficient and Faithful Sparse-View Surface Reconstruction

AAAI 2026technical

Reconstructing a faithful geometric surface from sparse images remains a fundamental challenge in 3D computer vision. While recent methods have achieved remarkable progress, they still struggle to recover reliable geometry due to the lack of multi-view geometric cues, particularly in non-overlapping

Cited by 0SourcePDFScholar
2025

A Method for Constructing Building Structure Grid Map Based on a Climbing Algorithm

ICRA 2025

Aerial-terrestrial amphibious robots excel in search and rescue tasks in unstructured terrains but face challenges in autonomous navigation indoors. Traditional full-mapping methods can degrade global path planning performance, especially when semi-static obstacles shift, leading to suboptimal paths

Cited by 0SourceScholar
2025

A Privacy-Preserving Cross-Modal Retrieval Scheme Based on CLIP and Deep Hashing

ICASSP 2025accepted

With the massive growth of multimedia data, local devices gradually cannot meet the data processing needs, thus utilizing cloud server resources becomes better choice. To prevent privacy leakage, user data can only be stored in ciphertext. Existing cross-media retrieval schemes are only for plaintex…

Cited by 0SourceScholar
2025

AdaDiff: Adaptive Step Selection for Fast Diffusion Models

AAAI 2025technical

Diffusion models, as a type of generative model, have achieved impressive results in generating images and videos conditioned on textual conditions. However, the generation process of diffusion models involves denoising dozens of steps to produce photorealistic images/videos, which is computationall…

Cited by 0SourcePDFScholar
2025

Autonomous Subtask Generation for Indoor Search and Rescue Mission via Large-Language-Model and Behavior-Tree Integration

IROS 2025

The ability of autonomous subtask generation is important for robots to effectively cope with unforeseen situations during indoor search and rescue missions. While prior work mainly focused on improving individual low-level skills of the rescue robot, this paper proposes AutoExpand: a high-level fra

Cited by 0SourcecodeScholar
2025

BlockDance: Reuse Structurally Similar Spatio-Temporal Features to Accelerate Diffusion Transformers

CVPR 2025poster

Diffusion models have demonstrated impressive generation capabilities, particularly with recent advancements leveraging transformer architectures to improve both visual and artistic quality. However, Diffusion Transformers (DiTs) continue to encounter challenges related to low inference speed, prima…

Cited by 1SourcePDFScholar
2025

CoDTS: Enhancing Sparsely Supervised Collaborative Perception with a Dual Teacher-Student Framework

AAAI 2025technical

Current collaborative perception methods often rely on fully annotated datasets, which can be expensive to obtain in practical situations. To reduce annotation costs, some works adopt sparsely supervised learning techniques and generate pseudo labels for the missing instances. However, these methods…

Cited by 0SourcePDFScholar
2025

Coordinated Energy-Trajectory Economic Model Predictive Control for Autonomous Surface Vehicles under Disturbances

IROS 2025

The paper proposes a novel Economic Model Predictive Control (EMPC) scheme for Autonomous Surface Vehicles (ASVs) to simultaneously address path following accuracy and energy constraints under environmental disturbances. By formulating lateral deviations as energy-equivalent penalties in the cost fu

Cited by 0SourceScholar
2025

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation

ICCV 2025poster

Diffusion models have been recognized for their ability to generate images that are not only visually appealing but also of high artistic quality. As a result, Layout-to-Image (L2I) generation has been proposed to leverage region-specific positions and descriptions to enable more precise and control…

Cited by 0SourcePDFScholar
2025

Cross-Modal Interactive Perception Network with Mamba for Lung Tumor Segmentation in PET-CT Images

CVPR 2025poster

Lung cancer is a leading cause of cancer-related deaths globally. PET-CT is crucial for imaging lung tumors, providing essential metabolic and anatomical information, while it faces challenges such as poor image quality, motion artifacts, and complex tumor morphology. Deep learning-based models are…

2025

CrossAD: Time Series Anomaly Detection with Cross-scale Associations and Cross-window Modeling

NeurIPS 2025poster

Time series anomaly detection plays a crucial role in a wide range of real-world applications. Given that time series data can exhibit different patterns at different sampling granularities, multi-scale modeling has proven beneficial for uncovering latent anomaly patterns that may not be apparent at…

Cited by 0SourceScholar
2025

DATA: Domain-And-Time Alignment for High-Quality Feature Fusion in Collaborative Perception

ICCV 2025poster

Feature-level fusion shows promise in collaborative perception (CP) through balanced performance and communication bandwidth trade-off. However, its effectiveness critically relies on input feature quality. The acquisition of high-quality features faces domain gaps from hardware diversity and deploy…

Cited by 0SourcePDFScholar
2025

Enhancing Fairness in Gaussian Mixture Clustering through Impact Factor

ICASSP 2025accepted

Clustering is a common method used in machine learning to group sample points in a dataset. Gaussian Mixture Clustering (GMC) is a clustering method based on maximum likelihood estimation and expectation maximisation (EM) algorithms. Traditional GMC does not consider the fairness between different s…

Cited by 0SourceScholar
2025

Enhancing Multi-Channel Speech with Limited Microphones via Spherical Harmonic Transform

ICASSP 2025accepted

The performance of traditional beamforming algorithms is influenced by the number of microphones, with performance improving as the number increases. However, in practice, the number of microphones is often limited. In this paper, we propose a novel virtual microphone estimation method that combines…

Cited by 0SourceScholar
2025

Enpowering Your Pansharpening Models with Generalizability: Unified Distribution is All You Need

ICCV 2025poster

Existing deep learning-based models for remote sensing pansharpening exhibit exceptional performance on training datasets. However, due to sensor-specific characteristics and varying imaging conditions, these models suffer from substantial performance degradation when applied to unseen satellite dat…

2025

Face Relighting with Ratio Function for Explicit Geometric Representation

ICASSP 2025accepted

This paper addresses the problem of face relighting under varying illumination conditions. Lighting is a fundamental element in portrait photography that shapes the mood, geometry, and overall realism of the captured characters. Most previous studies have mainly treated relighting as a 2D generation…

Cited by 0SourceScholar
2025

Forget the Token and Pixel: Rethinking Gradient Ascent for Concept Unlearning in Multimodal Generative Models

ACL 2025finding

Gradient Ascent (GA) has emerged as a promising approach for concept unlearning in Multimodal Generative Models (MGMs), such as Multimodal Large Language Models (MLLMs) and Stable Diffusion Models (SDMs). Despite its effectiveness in removing undesired knowledge, GA leads to severe utility degradati…

Cited by 0SourcePDFScholar
2025

HGSFusion: Radar-Camera Fusion with Hybrid Generation and Synchronization for 3D Object Detection

AAAI 2025technical

Millimeter-wave radar plays a vital role in 3D object detection for autonomous driving due to its all-weather and all-lighting-condition capabilities for perception. However, radar point clouds suffer from pronounced sparsity and unavoidable angle estimation errors. To address these limitations, inc…

2025

Individual Fairness for Fuzzy C-Means Clustering

ICASSP 2025accepted

In the field of clustering algorithms, the Fuzzy CMeans algorithm stands out for its ability to deal with uncertainty by assigning membership degrees to data points. However, research on the fairness of Fuzzy C-Means algorithms has mainly focused on group fairness, with limited attention to individu…

Cited by 0SourceScholar
2025

Inverse-Free and Data-Driven Motion Tracking Control for Redundant Robot with Fuzzy Recurrent Neural Network

IROS 2025

Precise motion tracking control with unknown structural knowledge and noise disturbance for redundant robots remains a critical and unresolved challenge. This article proposes a novel data-driven fuzzy discrete recurrent neural network (D<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlin

Cited by 0SourcecodeScholar
2025

MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance

ICCV 2025poster

Recent advances in video generation have led to remarkable improvements in visual quality and temporal coherence. Upon this, trajectory-controllable video generation has emerged to enable precise object motion control through explicitly defined spatial paths.However, existing methods struggle with c…

Cited by 0SourcePDFScholar
2025

NCL-CIR: Noise-aware Contrastive Learning for Composed Image Retrieval

ICASSP 2025accepted

Composed Image Retrieval (CIR) seeks to find a target image using a multi-modal query, which combines an image with modification text to pinpoint the target. While recent CIR methods have shown promise, they mainly focus on exploring relationships between the query pairs (image and text) through dat…

Cited by 0SourceScholar
2025

Novel Data-Driven Repetitive Motion Control Scheme for Redundant Manipulators With Zeroing Neurodynamics

IROS 2025

Repetitive motion control of redundant manipulators typically requires precise kinematic models to construct Jacobian matrices. However, model-based approaches are inherently limited when manipulator parameters are unavailable or only partially known. This paper introduces a novel data-driven discre

Cited by 0SourceScholar
2025

Optimization Based Human-Guided Variable-Stiffness Visual Impedance Control for Contact-Rich Tasks

IROS 2025

In contact-rich tasks such as polishing and drilling, inevitable physical interactions often lead to task deviations due to interference, typically resulting in excessive contact forces and eventual task failure. To tackle these challenges, we propose an innovative human-guided visual-impedance cont

Cited by 0SourceScholar
2025

Parameterized Motion Planning for Aerial Manipulators in Contact with Unstructured Surfaces

IROS 2025

Motion planning for continuous contact-based aerial manipulators on complex unstructured surfaces remains a substantial challenge due to the sophisticated topology of unstructured surfaces. While direct planning in the high-dimensional configuration space manifolds faces efficiency limitations, simp

Cited by 0SourceScholar
2025

Safety-Aware Geometric Force-Impedance Control for Manipulators

IROS 2025

Since its inception, impedance control has emerged as a fundamental framework for robotic interaction control. Recent advancements in geometric impedance control have demonstrated certain advantages over traditional Cartesian impedance control. However, existing geometric impedance control approache

Cited by 0SourceScholar
2025

Seg2Any: Open-set Segmentation-Mask-to-Image Generation with Precise Shape and Semantic Control

NeurIPS 2025poster

Despite recent advances in diffusion models, top-tier text-to-image (T2I) models still struggle to achieve precise spatial layout control, *i.e.* accurately generating entities with specified attributes and locations. Segmentation-mask-to-image (S2I) generation has emerged as a promising solution by…

Cited by 0SourceScholar
2025

Unlocking Constraints: Source-Free Occlusion-Aware Seamless Segmentation

ICCV 2025poster

Panoramic image processing is essential for omni-context perception, yet faces constraints like distortions, perspective occlusions, and limited annotations. Previous unsupervised domain adaptation methods transfer knowledge from labeled pinhole data to unlabeled panoramic images, but they require a…

2025

VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

ICASSP 2025accepted

Recent TTS models with decoder-only Transformer architecture, such as SPEAR-TTS and VALL-E, achieve impressive naturalness and demonstrate the ability for zero-shot adaptation given a speech prompt. However, such decoder-only TTS models lack monotonic alignment constraints, sometimes leading to hall…

Cited by 0SourceScholar
2025

Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting

NeurIPS 2025poster

The increase in computing power and the necessity of AI-assisted decision-making boost the growing application of large language models (LLMs). Along with this, the potential retention of sensitive data of LLMs has spurred increasing research into machine unlearning. However, existing unlearning app…

Cited by 0SourceScholar
2024

Efficient Multi-Channel Speech Enhancement with Spherical Harmonics Injection for Directional Encoding

ICASSP 2024accepted

Multi-channel speech enhancement extracts speech using multiple microphones that capture spatial cues. Effectively utilizing directional information is therefore key. Deep learning shows great potential on multi-channel speech enhancement and often takes short-time Fourier Transform (STFT) as inputs…

Cited by 0SourceScholar
2024

HIMap: HybrId Representation Learning for End-to-end Vectorized HD Map Construction

CVPR 2024poster

Vectorized High-Definition (HD) map construction requires predictions of the category and point coordinates of map elements (e.g. road boundary lane divider pedestrian crossing etc.). State-of-the-art methods are mainly based on point-level representation learning for regressing accurate point coord…

Cited by 24SourcePDFScholar
2024

IBCA: An Intelligent Platform for Social Insurance Benefit Qualification Status Assessment

AAAI 2024technical

Social insurance benefits qualification assessment is an important task to ensure that retirees enjoy their benefits according to the regulations. It also plays a key role in curbing social security frauds. In this paper, we report the deployment of the Intelligent Benefit Certification and Analysis…

Cited by 0SourcePDFScholar
2024

Innovative Directional Encoding in Speech Processing: Leveraging Spherical Harmonics Injection for Multi-Channel Speech Enhancement

IJCAI 2024poster

Multi-channel speech enhancement leverages multiple microphones to extract target speech signals amid background noise. Effectively utilizing directional cues is key for robust enhancement. While deep learning shows promise for multi-channel speech processing, most methods operate on short-time Four…

2024

Is Your HD Map Constructor Reliable under Sensor Corruptions?

NeurIPS 2024poster

Driving systems often rely on high-definition (HD) maps for precise environmental information, which is crucial for planning and navigation. While current HD map constructors perform well under ideal conditions, their resilience to real-world challenges, \eg, adverse weather and sensor failures, is…

Cited by 17SourcePDFScholar
2024

MBFusion: A New Multi-modal BEV Feature Fusion Method for HD Map Construction

ICRA 2024poster

HD map construction is a fundamental and challenging task in autonomous driving to understand the surrounding environment. Recently, Camera-LiDAR BEV feature fusion methods have attracted increasing attention in HD map construction task, which can significantly boost the benchmark. However, existing…

Cited by 11SourceScholar
2024

MapDistill: Boosting Efficient Camera-based HD Map Construction via Camera-LiDAR Fusion Model Distillation

ECCV 2024poster

"Online high-definition (HD) map construction is an important and challenging task in autonomous driving. Recently, there has been a growing interest in cost-effective multi-view camera-based methods without relying on other sensors like LiDAR. However, these methods suffer from a lack of explicit d…

Cited by 14SourcePDFScholar
2024

SEG-Net: Deep Learning Grasping With a Soft Enveloping Gripper

RA-L 2024

The emergence of non-fingered soft bioinspired grippers poses a challenge for learning-based grasping control due to the lack of a model describing grasping robustness and a dataset for training. In this letter, we propose a comprehensive pipeline encompassing grasping evaluation, dataset generation

Cited by 0SourceScholar
2024

ToolEENet: Tool Affordance 6D Pose Estimation

IROS 2024poster

The exploration of robotic dexterous hands utilizing tools has recently attracted considerable attention. A significant challenge in this field is the precise awareness of a tool’s pose when grasped, as occlusion by the hand often degrades the quality of the estimation. Additionally, the tool’s over…

Cited by 2SourcecodeScholar
2024

UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding

AAAI 2024technical

The utilization of discrete speech tokens, divided into semantic tokens and acoustic tokens, has been proven superior to traditional acoustic feature mel-spectrograms in terms of naturalness and robustness for text-to-speech (TTS) synthesis. Recent popular models, such as VALL-E and SPEAR-TTS, allow…

2024

Validating Privacy-Preserving Face Recognition under a Minimum Assumption

CVPR 2024poster

The widespread use of cloud-based face recognition technology raises privacy concerns as unauthorized access to face images can expose personal information or be exploited for fraudulent purposes. In response privacy-preserving face recognition (PPFR) schemes have emerged to hide visual information…

2023

AlphaRoute: Large-Scale Coordinated Route Planning via Monte Carlo Tree Search

AAAI 2023technical

This paper proposes AlphaRoute, an AlphaGo inspired algorithm for coordinating large-scale routes, built upon graph attention reinforcement learning and Monte Carlo Tree Search (MCTS). We first partition the road network into regions and model large-scale coordinated route planning as a Markov game,…

Cited by 6SourcePDFScholar
2023

DEdgeNet: Extrinsic Calibration of Camera and LiDAR with Depth-discontinuous Edges

ICRA 2023poster

This paper addresses the problem of calibrating extrinsic parameter matrix between an RGB camera and a LiDAR. Multimodal sensing systems are essential for fully autonomous navigation platforms. A key pre-requisite for such a system is calibration between different sensors. As the two most widely equ…

Cited by 11SourceScholar
2023

Focused and Collaborative Feedback Integration for Interactive Image Segmentation

CVPR 2023poster

Interactive image segmentation aims at obtaining a segmentation mask for an image using simple user annotations. During each round of interaction, the segmentation result from the previous round serves as feedback to guide the user's annotation and provides dense prior information for the segmentati…

2023

Physically Realizable Natural-Looking Clothing Textures Evade Person Detectors via 3D Modeling

CVPR 2023poster

Recent works have proposed to craft adversarial clothes for evading person detectors, while they are either only effective at limited viewing angles or very conspicuous to humans. We aim to craft adversarial texture for clothes based on 3D modeling, an idea that has been used to craft rigid adversar…

2023

Prototypical Residual Networks for Anomaly Detection and Localization

CVPR 2023poster

Anomaly detection and localization are widely used in industrial manufacturing for its efficiency and effectiveness. Anomalies are rare and hard to collect and supervised models easily over-fit to these seen anomalies with a handful of abnormal samples, producing unsatisfactory performance. On the o…

Cited by 84SourcePDFScholar
2023

Retro-FPN: Retrospective Feature Pyramid Network for Point Cloud Semantic Segmentation

ICCV 2023poster

Learning per-point semantic features from the hierarchical feature pyramid is essential for point cloud semantic segmentation. However, most previous methods suffered from ambiguous region features or failed to refine per-point features effectively, which leads to information loss and ambiguous sema…

Cited by 15PDFcodeScholar
2022

Alignment-Learning Based Single-Step Decoding for Accurate and Fast Non-Autoregressive Speech Recognition

ICASSP 2022accepted

Non-autoregressive transformer (NAT) based speech recognition models have gained more and more attention since they perform faster inference speed compared with autoregressive counterparts, especially when the single-step decoding is applied. However, the single-step decoding process with length pre…

Cited by 0SourceScholar
2022

Alleviating the Loss-Metric Mismatch in Supervised Single-Channel Speech Enhancement

ICASSP 2022accepted

In this paper, we study the loss-metric mismatch problem of supervised single-channel speech enhancement system. Most of the existing speech enhancement systems achieve unsatisfying performance since their empirically selected loss functions have semantic gaps with the non-differentiable evaluation…

Cited by 0SourceScholar
2022

CSL: A Large-scale Chinese Scientific Literature Dataset

COLING 2022main

Scientific literature serves as a high-quality corpus, supporting a lot of Natural Language Processing (NLP) research. However, existing datasets are centered around the English language, which restricts the development of Chinese scientific NLP. In this work, we present CSL, a large-scale Chinese S…

2022

PaddleSpeech: An Easy-to-Use All-in-One Speech Toolkit

NAACL 2022system demonstrations

PaddleSpeech is an open-source all-in-one speech toolkit. It aims at facilitating the development and research of speech processing technologies by providing an easy-to-use command-line interface and a simple code structure. This paper describes the design philosophy and core architecture of PaddleS…

2022

Slot-VPS: Object-Centric Representation Learning for Video Panoptic Segmentation

CVPR 2022poster

Video Panoptic Segmentation (VPS) aims at assigning a class label to each pixel, uniquely segmenting and identifying all object instances consistently across all frames. Classic solutions usually decompose the VPS task into several sub-tasks and utilize multiple surrogates (e.g. boxes and masks, cen…

Cited by 30PDFcodeScholar
2021

A Variable Stiffness Actuator Based on Second-order Lever Mechanism and Its Manipulator Integration

ICRA 2021poster

This paper presents a new variable stiffness actuator based on a second-order lever mechanism which has wide stiffness regulation range. By employing a novel symmetric structure design and improving the load capacity of the stiffness regulation module, the proposed actuator also shows well performan…

Cited by 10SourceScholar
2021

Free-Form Description Guided 3D Visual Graph Network for Object Grounding in Point Cloud

ICCV 2021poster

3D object grounding aims to locate the most relevant target object in a raw point cloud scene based on a free-form language description. Understanding complex and diverse descriptions, and lifting them directly to a point cloud is a new and challenging topic due to the irregular and sparse nature of…

Cited by 100PDFcodeScholar
2021

Interaction via Bi-Directional Graph of Semantic Region Affinity for Scene Parsing

ICCV 2021poster

In this work, we devote to address the challenging problem of scene parsing. Previous methods, though capture context to exploit global clues, handle scene parsing as a pixel-independent task. However, it is well known that pixels in an image are highly correlated with each other, especially those f…

Cited by 19PDFScholar
2021

Joint Alignment Learning-Attention Based Model for Grapheme-to-Phoneme Conversion

ICASSP 2021accepted

Sequence-to-sequence attention-based models for grapheme-to-phoneme (G2P) conversion have gained significant interests. The attention-based encoder-decoder framework learns the mapping of input to output tokens by selectively focusing on relevant information, and has been shown well performance. How…

Cited by 0SourceScholar
2021

Learning Frequency-Aware Dynamic Network for Efficient Super-Resolution

ICCV 2021poster

Deep learning based methods, especially convolutional neural networks (CNNs) have been successfully applied in the field of single image super-resolution (SISR). To obtain better fidelity and visual quality, most of existing networks are of heavy design with massive computation. However, the computa…

Cited by 84PDFScholar
2021

Order Regularization on Ordinal Loss for Head Pose, Age and Gaze Estimation

AAAI 2021technical

Ordinal loss is widely used in solving regression problems with deep learning technologies. Its basic idea is to convert regression to classification while preserving the natural order. However, the order constraint is enforced only by ordinal label implicitly, leading to the real output values not…

Cited by 8SourcePDFScholar
2020

AutoSTR: Efficient Backbone Search for Scene Text Recognition

ECCV 2020poster

Scene text recognition (STR) is challenging due to the diversity of text instances and the complexity of scenes. However, no STR methods can adapt backbones to different diversities and complexities. In this work, inspired by the success of neural architecture search (NAS), we propose automated STR…

2018

Monocular Visual Odometry Scale Recovery Using Geometrical Constraint

ICRA 2018poster

Scale recovery is one of the essential problems for monocular visual odometry. The camera height is usually used as an absolute reference to recover the scale. In this case, the precision of scale recovery depends on the accuracy of the road region detection and road geometrical model calculation. I…

Cited by 31SourceScholar
2018

Training Supervised Speech Separation System to Improve STOI and PESQ Directly

ICASSP 2018accepted

Supervised speech separation methods train learning machine to cast the noisy speech to the target clean speech. Most of them use mean-square error (MSE) as loss function. However, MSE is not the perfect choice because it doesn't match the human auditory perception. Short-time objective intelligibil…

Cited by 0SourceScholar
2015

A pairwise algorithm for pitch estimation and speech separation using deep stacking network

ICASSP 2015accepted

Pitch information is an important cue for speech separation. However, pitch estimation in noisy condition is also a task as challenging as speech separation. In this paper, we propose a supervised learning architecture which combines these two problems concisely. The proposed algorithm is based on d…

Cited by 0SourceScholar
2015

The Common Self-Polar Triangle of Concentric Circles and Its Application to Camera Calibration

CVPR 2015poster

In projective geometry, the common self-polar triangle has often been used to discuss the position relationship of two planar conics. However, there are few researches on the properties of the common self-polar triangle, especially when the two planar conics are special conics. In this paper, we exp…

Cited by 63SourcePDFScholar