← Search

Xu Zhang

79 accepted papers

2026

Any2RSI: Controllable Remote Sensing Text-to-Image Generation via Any Control and Enriched Description

AAAI 2026technical

Recent advances in controllable text-to-image (T2I) generation have achieved impressive results in natural images, but remote sensing (RS) T2I remains challenging due to the unique nature of geospatial data. Existing methods struggle to integrate diverse spatial controls and model complex spatial re

Cited by 2SourcePDFScholar
2026

BEVFormer++: Temporal Amplified BEVformer with Explicit Parameter Prediction for Automatic Trajectory Prediction

IJCAI 2026

Vision-based trajectory prediction with BEV representations has achieved promising results, yet existing methods often suffer from limited temporal modeling and insufficient characterization of motion dynamics. To address these issues, we propose a temporally enhanced framework with explicit motion

Cited by 0Scholar
2026

ClearAIR: A Human-Visual-Perception-Inspired All-in-One Image Restoration

AAAI 2026technical

Recently, All-in-One image restoration (AiOIR) has advanced significantly, offering promising solutions for complex real-world degradations. However, most existing approaches heavily rely on degradation-specific representation learning, which can lead to oversmoothing and artifacts in the restored i

Cited by 0SourcePDFScholar
2026

DiGraphHal-Bench: Evaluating Multimodal Large Language Models on Complex Directed Graphs

CVPR 2026

While prior research on Multimodal Large Language Model (MLLM) hallucinations has primarily examined cross-modal inconsistencies in natural images, hallucination over complex graph structures remains underexplored.Concurrently, there is a lack of robust evaluation for fine-grained reasoning integrat

Cited by 0SourcecodeScholar
2026

GeoPurify: A Data-Efficient Geometric Distillation Framework for Open-Vocabulary 3D Segmentation

ICLR 2026poster

Recent attempts to transfer features from 2D Vision–Language Models (VLMs) to 3D semantic segmentation expose a persistent trade-off. Directly projecting 2D features into 3D yields noisy and fragmented predictions, whereas enforcing geometric coherence necessitates costly training pipelines and larg…

Cited by 0SourcecodeScholar
2026

Heuristic-inspired Reasoning Priors Facilitate Data-Efficient Referring Object Detection

CVPR 2026

Most referring object detection (ROD) models, especially the modern grounding detectors, are designed for data-rich conditions, yet many practical deployments, such as robotics, augmented reality, and other specialized domains, would face severe label scarcity. In such regimes, end-to-end grounding

Cited by 0SourcecodeScholar
2026

InterLight: Leveraging Intrinsic Illumination Priors for Low-Light Image Enhancement

IJCAI 2026

Low-Light Image Enhancement (LLIE) has long been a challenging problem in low-level vision, as insufficient illumination often leads to low contrast, detail loss, and noise. Recent studies show that deep learning-based Retinex theory can effectively decouple illumination and reflectance. However, ex

Cited by 0Scholar
2026

MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX

AAAI 2026technical

We introduce MAVERIX (Multimodal Audio-Visual Evaluation and Recognition IndeX), a unified benchmark to probe video understanding in multimodal LLMs, encompassing video, audio, and text inputs with human performance baselines. Although recent advancements in audiovisual models have shown substantial

Cited by 0SourcePDFScholar
2026

MM-ReCoder: Advancing Chart-to-Code Generation with Reinforcement Learning and Self-Correction

CVPR 2026

Multimodal Large Language Models (MLLMs) have recently demonstrated promising capabilities in multimodal coding tasks such as chart-to-code generation. However, existing methods primarily rely on supervised fine-tuning (SFT), which requires the model to learn code patterns through chart-code pairs b

Cited by 0SourceScholar
2026

MN-Diff: Diffusion Parameterized MoE-NCDE for Continuous Time Series Generation with Irregular Observations

ICML 2026poster

Time series generation (TSG) is widely used across domains, yet most existing methods assume regular sampling and fixed output resolutions. These assumptions are often violated in practice, where observations are irregular and sparse, while downstream applications require continuous and high-resolut…

Cited by 0SourceScholar
2026

Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image Reconstruction

ICLR 2026poster

Reconstructing visual stimuli from fMRI signals is a central challenge bridging machine learning and neuroscience. Recent diffusion-based methods typically map fMRI activity to a single neural embedding, using it as static guidance throughout the entire generation process. However, this fixed guidan…

Cited by 0SourcecodeScholar
2026

Personalized Federated Learning with Bidirectional Communication Compression via One-Bit Random Sketching

AAAI 2026technical

Federated Learning (FL) enables collaborative training across decentralized data, but faces key challenges of bidirectional communication overhead and client-side data heterogeneity. To address communication costs while embracing data heterogeneity, we propose pFed1BS, a novel personalized federate

Cited by 0SourcePDFScholar
2026

ShadeEdit: A Utility-Preserving and Defense-Evasive Knowledge Manipulation Attack in Federated LLMs

AAAI 2026technical

Recent studies reveal that adversaries can manipulate the internal knowledge of large language models (LLMs) on selected topics through model editing, causing attacker-specified harmful or biased outputs when queried about the edited content. Once such tampered LLMs are distributed, they can mislead

Cited by 0SourcePDFScholar
2026

Sortblock: Similarity-Aware Feature Reuse for Diffusion Model

AAAI 2026technical

Diffusion Transformers (DiTs) have demonstrated remarkable generative capabilities, particularly benefiting from Transformer architectures that enhance visual and artistic fidelity. However, their inherently sequential denoising process results in high inference latency, limiting their deployment in

Cited by 0SourcePDFScholar
2026

Taming Hierarchical Image Coding Optimization: A Spectral Regularization Perspective

ICLR 2026poster

Hierarchical coding offers distinct advantages for learned image compression by capturing multi-scale representations to support scale-wise modeling and enable flexible quality scalability, making it a promising alternative to single-scale models. However, its practical performance remains limited.…

Cited by 0SourceScholar
2026

UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction as Reasoning

ICLR 2026poster

GUI grounding, which maps natural-language instructions to actionable UI elements, is a core capability of GUI agents. Prior work largely treats instructions as a static proxy for user intent, overlooking the impact of instruction diversity on grounding performance. Through a careful investigation o…

Cited by 0SourcecodeScholar
2026

UniChange: Unifying Change Detection with Multimodal Large Language Model

CVPR 2026

Change detection (CD) is a fundamental task for monitoring and analysing land cover dynamics. While recent high performance models and high quality datasets have significantly advanced the field, a critical limitation persists. Current models typically acquire limited knowledge from single-type anno

Cited by 0SourcecodeScholar
2025

3D Shape Classification by Registration: Neural-Network-Free and Training-Free

ICASSP 2025accepted

Point cloud classification, crucial for discriminative 3D shape analysis, has witnessed significant progress through the application of deep learning. A significant research focus has been on aggregating local point cloud features. A key limitation of previous methods lies in their inherent opacity,…

Cited by 0SourceScholar
2025

A General Knowledge Injection Framework for ICD Coding

ACL 2025finding

ICD Coding aims to assign a wide range of medical codes to a medical text document, which is a popular and challenging task in the healthcare domain. To alleviate the problems of long-tail distribution and the lack of annotations of code-specific evidence, many previous works have proposed incorpora…

2025

A Lightweight Sparse Interaction Network for Time Series Forecasting

AAAI 2025technical

Recent work shows that linear models can outperform several transformer models in long-term time-series forecasting (TSF). However, instead of explicitly performing temporal interaction through self-attention, linear models implicitly perform it based on stacked MLP structures, which may be insuffic…

Cited by 0SourcePDFScholar
2025

AA-CLIP: Enhancing Zero-Shot Anomaly Detection via Anomaly-Aware CLIP

CVPR 2025poster

Anomaly detection (AD) identifies outliers for applications like defect and lesion detection. While CLIP shows promise for zero-shot AD tasks due to its strong generalization capabilities, its inherent Anomaly-Unawareness leads to limited discrimination between normal and abnormal features. To addre…

2025

Achieving More with Less: Additive Prompt Tuning for Rehearsal-Free Class-Incremental Learning

ICCV 2025poster

Class-incremental learning (CIL) enables models to learn new classes progressively while preserving knowledge of previously learned ones. Recent advances in this field have shifted towards parameter-efficient fine-tuning techniques, with many approaches building upon the framework that maintains a p…

Cited by 0SourcePDFScholar
2025

Adversarial Data Augmentation for Single Domain Generalization via Lyapunov Exponent-Guided Optimization

ICCV 2025poster

Single Domain Generalization (SDG) aims to develop models capable of generalizing to unseen target domains using only one source domain, a task complicated by substantial domain shifts and limited data diversity. Existing SDG approaches primarily rely on data augmentation techniques, which struggle…

2025

BoRe-Depth: Self-Supervised Monocular Depth Estimation with Boundary Refinement for Embedded Systems

IROS 2025

Depth estimation is one of the key technologies for realizing 3D perception in unmanned systems. Monocular depth estimation has been widely researched because of its low-cost advantage, but the existing methods face the challenges of poor depth estimation performance and blurred object boundaries on

Cited by 1SourcecodeScholar
2025

DAMON: A Dialogue-Aware MCTS Framework for Jailbreaking Large Language Models

EMNLP 2025

While large language models (LLMs) demonstrate remarkable capabilities across a wide range of tasks, they remain vulnerable to generating outputs that are potentially harmful. Red teaming, which involves crafting adversarial inputs to expose vulnerabilities, is a widely adopted approach for evaluati

2025

Enhancing Continuum Robot Mobility: Design and Control with Integrated Dual Rotational DOFs

IROS 2025

Continuum robots, known for their compliance in unstructured environments, face limitations due to the lack of rotational degrees of freedom (DOFs) about the backbone. This prevents them from compensating undesired torsional deformation and performing 6-DOF control of the end-effector, thereby restr

Cited by 0SourceScholar
2025

Generalization Performance of Ensemble Clustering: From Theory to Algorithm

ICML 2025poster

Ensemble clustering has demonstrated great success in practice; however, its theoretical foundations remain underexplored. This paper examines the generalization performance of ensemble clustering, focusing on generalization error, excess risk and consistency. We derive a convergence rate of general…

2025

Incorporating Legal Logic into Deep Learning: An Intelligent Approach to Probation Prediction

IJCAI 2025

Probation is a crucial institution in modern criminal law, embodying the principles of fairness and justice while contributing to the harmonious development of society. Despite its importance, the current Intelligent Judicial Assistant System (IJAS) lacks dedicated methods for probation prediction,

Cited by 0SourcePDFScholar
2025

MC-MKE: A Fine-Grained Multimodal Knowledge Editing Benchmark Emphasizing Modality Consistency

ACL 2025finding

Multimodal large language models (MLLMs) are prone to non-factual or outdated knowledge issues, highlighting the importance of knowledge editing. Many benchmark has been proposed for researching multimodal knowledge editing. However, previous benchmarks focus on limited scenarios due to the lack of…

Cited by 0SourcePDFScholar
2025

MMAG: Multimodal Learning for Mucus Anomaly Grading in Nasal Endoscopy via Semantic Attribute Prompting

EMNLP 2025

Accurate grading of rhinitis severity in nasal endoscopy relies heavily on the characterization of key secretion types, notably clear nasal discharge (CND) and purulent nasal secretion (PUS). However, both exhibit ambiguous appearance and high structural variability, posing challenges to automated g

Cited by 0SourcePDFScholar
2025

Optimizing Robustness and Accuracy in Mixture of Experts: A Dual-Model Approach

ICML 2025poster

Mixture of Experts (MoE) have shown remarkable success in leveraging specialized expert networks for complex machine learning tasks. However, their susceptibility to adversarial attacks presents a critical challenge for deployment in robust applications. This paper addresses the critical question of…

Cited by 0SourcePDFScholar
2025

Role-Specific Reward Design with Large Language Model for StarCraft II

ICASSP 2025accepted

Reward acts as a signal to guide the agent’s learning process in Reinforcement Learning (RL), evaluating and assigning rewards to the agent’s actions based on theiralignment with goals. Designing reward is challenging in multiagent environment such as StarCraft II benchmark since agents face credit…

Cited by 0SourceScholar
2025

SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioning

AAAI 2025technical

Controllable image semantic understanding tasks, such as captioning or segmentation, necessitate users to input a prompt (e.g., text or bounding boxes) to predict a unique outcome, presenting challenges such as high-cost prompt input or limited information output. This paper introduces a new task ``…

Cited by 0SourcePDFScholar
2025

SparseRecon: Neural Implicit Surface Reconstruction from Sparse Views with Feature and Depth Consistencies

ICCV 2025poster

Surface reconstruction from sparse views aims to reconstruct a 3D shape or scene from few RGB images. The latest methods are either generalization-based or overfitting-based. However, the generalization-based methods do not generalize well on views that were unseen during training, while the reconst…

2025

UFO: A UI-Focused Agent for Windows OS Interaction

NAACL 2025long

We introduce UFO, a UI-Fcused agent designed to fulfill user requests tailored to Windows OS applications by observing and analyzing the GUI and control information of these applications. UFO utilizes a hierarchical dual-agent framework that decomposes user requests using a divide-and-conquer approa…

2025

Ultra Lowrate Image Compression with Semantic Residual Coding and Compression-aware Diffusion

ICML 2025poster

Existing multimodal large model-based image compression frameworks often rely on a fragmented integration of semantic retrieval, latent compression, and generative models, resulting in suboptimal performance in both reconstruction fidelity and coding efficiency. To address these challenges, we propo…

Cited by 0SourcePDFScholar
2024

A Lightweight Hybrid Multi-Channel Speech Extraction System with Directional Voice Activity Detection

ICASSP 2024accepted

Although deep learning (DL) based end-to-end models have shown outstanding performance in multi-channel speech extraction, their practical applications on edge devices are restricted due to their high computational complexity. In this paper, we propose a hybrid system that can more effectively integ…

Cited by 0SourceScholar
2024

All-in-One Image Coding for Joint Human-Machine Vision with Multi-Path Aggregation

NeurIPS 2024poster

Image coding for multi-task applications, catering to both human perception and machine vision, has been extensively investigated. Existing methods often rely on multiple task-specific encoder-decoder pairs, leading to high overhead of parameter and bitrate usage, or face challenges in multi-objecti…

2024

Benchmarking Knowledge Boundary for Large Language Models: A Different Perspective on Model Evaluation

ACL 2024long

In recent years, substantial advancements have been made in the development of large language models, achieving remarkable performance across diverse tasks.To evaluate the knowledge ability of language models, previous studies have proposed lots of benchmarks based on question-answering pairs.We arg…

2024

DMPlug: A Plug-in Method for Solving Inverse Problems with Diffusion Models

NeurIPS 2024poster

Pretrained diffusion models (DMs) have recently been popularly used in solving inverse problems (IPs). The existing methods mostly interleave iterative steps in the reverse diffusion process and iterative steps to bring the iterates closer to satisfying the measurement constraint. However, such inte…

2024

DPA-Net: Structured 3D Abstraction from Sparse Views via Differentiable Primitive Assembly

ECCV 2024poster

"We present a differentiable rendering framework to learn structured 3D abstractions in the form of primitive assemblies from sparse RGB images capturing a 3D object. By leveraging differentiable volume rendering, our method does not require 3D supervision. Architecturally, our network follows the g…

Cited by 4SourcePDFScholar
2024

Design-Modeling and Control of a Novel Wearable Exoskeleton for Lower-Limb Enhancement

RA-L 2024

In this paper, a novel powered lower limb exoskeleton prototype called PTEXO for reducing user burden and enhancing following comfort is presented. The PTEXO is designed with a new control strategy, Enhanced Sensitivity Amplification Control (ESAC), and improves comfort of lower-limb locomotion thro

Cited by 5SourceScholar
2024

GTCRN: A Speech Enhancement Model Requiring Ultralow Computational Resources

ICASSP 2024accepted

While modern deep learning-based models have significantly outperformed traditional methods in the area of speech enhancement, they often necessitate a lot of parameters and extensive computational power, making them impractical to be deployed on edge devices in real-world applications. In this pape…

Cited by 0SourceScholar
2024

HVCLIP: High-dimensional Vector in CLIP for Unsupervised Domain Adaptation

ECCV 2024poster

"Recent advancement in the large-scale image-text pre-training model (such as CLIP) has significantly improved unsupervised domain adaptation (UDA) by leveraging the pre-trained knowledge to bridge the source and target domain gap. However, Catastrophic forgetting still remains to be the main challe…

Cited by 2SourcePDFScholar
2024

Image-Based Distributed Predictive Visual Servo Control for Cooperative Tracking of Multiple Fixed-Wing UAVs

RA-L 2024

This letter proposes a novel approach combining the distributed model predictive control (DMPC) and the image-based visual servoing (IBVS) for cooperative tracking problem of multiple fixed-wing Unmanned Aerial Vehicles (UAVs) equipped with pan-tilt cameras. In particular, the target is unknown and

Cited by 8SourceScholar
2024

Independent-Set Design of Experiments for Estimating Treatment and Spillover Effects under Network Interference

ICLR 2024poster

Interference is ubiquitous when conducting causal experiments over networks. Except for certain network structures, causal inference on the network in the presence of interference is difficult due to the entanglement between the treatment assignments and the interference levels. In this article, we…

Cited by 6SourcePDFScholar
2024

Long-Term Social Interaction Context: The Key to Egocentric Addressee Detection

ICASSP 2024accepted

As embodied agents learn to interact, it is crucial for them to understand when, what, and to whom they should respond. While advances in natrual language processing and speech technologies have enabled conversational agents to focus on what to respond, they still struggle to determine when and to w…

Cited by 0SourceScholar
2024

Low-Light Raw Image Enhancement on a Dataset Suffering Light Effects

ICASSP 2024accepted

Deep learning-based methods have achieved remarkable success in low-light image enhancement (LLIE). But most existing works are based on sRGB data and do not focus on the light effects in bright regions when enhancing low-light regions. This inevitably leads to excessive enhancement and saturation o…

Cited by 0SourceScholar
2024

Plug-In Diffusion Model for Sequential Recommendation

AAAI 2024technical

Pioneering efforts have verified the effectiveness of the diffusion models in exploring the informative uncertainty for recommendation. Considering the difference between recommendation and image synthesis tasks, existing methods have undertaken tailored refinements to the diffusion and reverse proc…

2024

Reduce Redundancy Then Rerank: Enhancing Code Summarization with a Novel Pipeline Framework

COLING 2024main

Code summarization is the task of automatically generating natural language descriptions from source code. Recently, pre-trained language models have gained significant popularity in code summarization due to their capacity to capture richer semantic representations of both code and natural language…

2024

Self-Adaptive Scale Handling for Forecasting Time Series with Scale Heterogeneity

ICASSP 2024accepted

Time series forecasting (TSF) is crucial in various fields and has gained extensive research. However, most studies are conducted based on TS data with scale homogeneity. This paper proposes a self-Adaptive Scale-handling (AS) module to improve the performance of forecasting TS with scale heterogene…

Cited by 0SourceScholar
2023

An Empirical Study of Instruction-tuning Large Language Models in Chinese

EMNLP 2023long findings

The success of ChatGPT validates the potential of large language models (LLMs) in artificial general intelligence (AGI). Subsequently, the release of LLMs has sparked the open-source community's interest in instruction-tuning, which is deemed to accelerate ChatGPT's replication process. However, r…

Cited by 0SourcecodeScholar
2023

Avoiding spurious correlations via logit correction

ICLR 2023poster

Empirical studies suggest that machine learning models trained with empirical risk minimization (ERM) often rely on attributes that may be spuriously correlated with the class labels. Such models typically lead to poor performance during inference for data lacking such correlations. In this work, we…

2023

FashionNTM: Multi-turn Fashion Image Retrieval via Cascaded Memory

ICCV 2023poster

Multi-turn textual feedback-based fashion image retrieval focuses on a real-world setting, where users can iteratively provide information to refine retrieval results until they find an item that fits all their requirements. In this work, we present a novel memory-based method, called FashionNTM, fo…

Cited by 9PDFcodeScholar
2023

G3R: A Graph-Guided Generate-and-Rerank Framework for Complex and Cross-domain Text-to-SQL Generation

ACL 2023findings

We present a framework called G3R for complex and cross-domain Text-to-SQL generation. G3R aims to address two limitations of current approaches: (1) The structure of the abstract syntax tree (AST) is not fully explored during the decoding process which is crucial for complex SQL generation; (2) Dom…

2023

MIL-Decoding: Detoxifying Language Models at Token-Level via Multiple Instance Learning

ACL 2023long

Despite advances in large pre-trained neural language models, they are prone to generating toxic language, which brings security risks to their applications. We introduce MIL-Decoding, which detoxifies language models at token-level by interpolating it with a trained multiple instance learning (MIL)…

2023

MST-Q: Micro Suction Tape Quadruped Robot With High Payload Capacity

RA-L 2023

Payload capacity is a crucial factor for climbing robots, as it directly affects their ability to carry and transport heavy loads during various climbing tasks. However, many dry adhesion-based legged robots prioritize foot design from a bionic perspective to accomplish various climbing tasks while

Cited by 14SourceScholar
2023

Towards Lightweight, Model-Agnostic and Diversity-Aware Active Anomaly Detection

ICLR 2023poster

Active Anomaly Discovery (AAD) is flourishing in the anomaly detection research area, which aims to incorporate analysts’ feedback into unsupervised anomaly detectors. However, existing AAD approaches usually prioritize the samples with the highest anomaly scores for user labeling, which hinders the…

Cited by 1SourcePDFScholar
2023

User-Controllable Arbitrary Style Transfer via Entropy Regularization

AAAI 2023technical

Ensuring the overall end-user experience is a challenging task in arbitrary style transfer (AST) due to the subjective nature of style transfer quality. A good practice is to provide users many instead of one AST result. However, existing approaches require to run multiple AST models or inference a…

2022

A Two-Step Backward Compatible Fullband Speech Enhancement System

ICASSP 2022accepted

Speech enhancement methods based on deep learning have surpassed traditional methods. While many of these new approaches are operating on the wideband (16kHz) sample rate, a new fullband (48kHz) speech enhancement system is proposed in this paper. Compared to the existing full-band systems that util…

Cited by 0SourceScholar
2022

Byzantine-tolerant federated Gaussian process regression for streaming data

NeurIPS 2022accept

In this paper, we consider Byzantine-tolerant federated learning for streaming data using Gaussian process regression (GPR). In particular, a cloud and a group of agents aim to collaboratively learn a latent function where some agents are subject to Byzantine attacks. We develop a Byzantine-tolerant…

Cited by 5SourcePDFScholar
2022

Code Generation From Flowcharts with Texts: A Benchmark Dataset and An Approach

EMNLP 2022finding

Currently, researchers focus on generating codes from the requirement documents. However, current approaches still perform poorly on some requirements needing complex problem-solving skills. In reality, to tackle such complex requirements, instead of directly translating requirement documents into c…

2022

Complicate Then Simplify: A Novel Way to Explore Pre-trained Models for Text Classification

COLING 2022main

With the development of pre-trained models (PTMs), the performance of text classification has been continuously improved by directly employing the features generated by PTMs. However such way might not fully explore the knowledge in PTMs as it is constrained by the difficulty of the task. Compared t…

2022

Large-Scale Video Panoptic Segmentation in the Wild: A Benchmark

CVPR 2022poster

In this paper, we present a new large-scale dataset for the video panoptic segmentation task, which aims to assign semantic classes and track identities to all pixels in a video. As the ground truth for this task is difficult to annotate, previous datasets for video panoptic segmentation are limited…

Cited by 100PDFcodeScholar
2022

Multi-Stage and Multi-Loss Training for Fullband Non-Personalized and Personalized Speech Enhancement

ICASSP 2022accepted

Deep learning-based wideband (16kHz) speech enhancement approaches have surpassed traditional methods. This work further extends the existing wideband systems to enable full-band (48kHz) speech enhancement while simultaneously ensuring automatic speech recognition compatibility and optionally, perso…

Cited by 0SourceScholar
2022

Personalized Federated Learning via Variational Bayesian Inference

ICML 2022spotlight

Federated learning faces huge challenges from model overfitting due to the lack of data and statistical diversity among clients. To address these challenges, this paper proposes a novel personalized federated learning method via Bayesian variational inference named pFedBayes. To alleviate the overfi…

Cited by 122SourcePDFScholar
2021

Generalized Relation Learning with Semantic Correlation Awareness for Link Prediction

AAAI 2021technical

Developing link prediction models to automatically complete knowledge graphs has recently been the focus of significant research interest. The current methods for the link prediction task have two natural problems: 1) the relation distributions in KGs are usually unbalanced, and 2) there are many un…

Cited by 18SourcePDFScholar
2020

An Optimal Symmetric Threshold Strategy for Remote Estimation Over The Collision Channel

ICASSP 2020accepted

A wireless sensing system with n sensors, observing independent and identically distributed continuous random variables with a symmetric probability density function, and one non-collocated estimator acting as a fusion center is considered. The sensors transmit information to the fusion center via a…

Cited by 0SourceScholar
2020

Intra-Correlation Encoding for Chinese Sentence Intention Matching

COLING 2020main

Sentence intention matching is vital for natural language understanding. Especially for Chinese sentence intention matching task, due to the ambiguity of Chinese words, semantic missing or semantic confusion are more likely to occur in the encoding process. Although the existing methods have enriche…

2019

Unsupervised Embedding Learning via Invariant and Spreading Instance Feature

CVPR 2019poster

This paper studies the unsupervised embedding learning problem, which requires an effective similarity measurement between samples in low-dimensional embedding space. Motivated by the positive concentrated and negative separated properties observed from category-wise supervised learning, we propose…

Cited by 760PDFcodeScholar
2017

Learning Discriminative and Transformation Covariant Local Feature Detectors

CVPR 2017poster

Robust covariant local feature detectors are important for detecting local features that are (1) discriminative of the image content and (2) can be repeatably detected at consistent locations when the image undergoes diverse transformations. Such detectors are critical for applications such as image…

Cited by 157PDFcodeScholar
2015

Fast Orthogonal Projection Based on Kronecker Product

ICCV 2015poster

We propose a family of structured matrices to speed up orthogonal projections for high-dimensional data commonly seen in computer vision applications. In this, a structured matrix is formed by the Kronecker product of a series of smaller orthogonal matrices. This achieves O(dlogd) computational comp…

Cited by 55PDFScholar