← Search

Qian Zhang

101 accepted papers

2026

Alignment between Brains and AI: Evidence for Convergent Evolution across Modalities, Scales and Training Trajectories

ICML 2026poster

Artificial and biological systems may evolve similar computational solutions despite fundamental differences in architecture and learning mechanisms—a form of convergent evolution. We provide large-scale evidence for this phenomenon through comprehensive analysis of alignment between human brain act…

Cited by 0SourceScholar
2026

Boosting the Robustness-Accuracy Trade-off of SNNs by Robust Temporal Self-Ensemble

AAAI 2026technical

Spiking Neural Networks (SNNs) offer a promising direction for energy-efficient and brain-inspired computing, yet their vulnerability to adversarial perturbations remains poorly understood. In this work, we revisit the adversarial robustness of SNNs through the lens of temporal ensembling, treating

Cited by 0SourcePDFScholar
2026

ClearAIR: A Human-Visual-Perception-Inspired All-in-One Image Restoration

AAAI 2026technical

Recently, All-in-One image restoration (AiOIR) has advanced significantly, offering promising solutions for complex real-world degradations. However, most existing approaches heavily rely on degradation-specific representation learning, which can lead to oversmoothing and artifacts in the restored i

Cited by 0SourcePDFScholar
2026

DevEvol: Benchmarking LLM Agents on Continuous Software Evolution

ICML 2026poster

Large Language Model (LLM) agents have demonstrated remarkable proficiency in solving isolated software engineering tasks. However, existing benchmarks predominantly evaluate static, independent issues, failing to reflect the continuous and sequentially dependent nature of real-world software evolut…

Cited by 0SourceScholar
2026

Di-BiLPS: Denoising induced Bidirectional Latent-PDE-Solver under Sparse Observations

ICML 2026poster

Partial differential equations (PDEs) are fundamental for modeling complex natural and physical phenomena. In many real-world applications, however, observational data are \textbf{extremely sparse}, which severely limits the applicability of both classical numerical solvers and existing neural appro…

Cited by 0SourceScholar
2026

EgoRoC: Towards Egocentric Robotic Control via Task-Agnostic Visual Alignment

CVPR 2026

Recent Vision-Language-Action (VLA) models map visual-textual inputs to robotic actions via end-to-end architectures, yet this approach entangles visual understanding with task-specific actions. This leads to an exhaustive collection of full operational sequences and parameter redundancy across task

Cited by 0SourceScholar
2026

Key Decision-Makers in Multi-Agent Debates: Who Holds the Power?

AAAI 2026technical

Recent studies on LLM agent scaling have highlighted the potential of Multi-Agent Debate (MAD) to enhance reasoning abilities. However, the critical aspect of role allocation strategies remains underexplored. In this study, we demonstrate that allocating roles with differing viewpoints to specific p

Cited by 0SourcePDFScholar
2026

LongStream: Long-Sequence Streaming Autoregressive Visual Geometry

CVPR 2026

Long-sequence streaming 3D reconstruction remains a significant open challenge. Existing autoregressive models often fail when processing long sequences because they anchor poses to the first frame, leading to attention decay, scale drift, and extrapolation errors. We introduce LongStream, a novel g

Cited by 0SourcecodeScholar
2026

MASPOB: Bandit-Based Prompt Optimization for Multi-Agent Systems with Graph Neural Networks

ICML 2026spotlight

Large Language Models (LLMs) have achieved significant success across a wide range of tasks, serving as the cognitive backbone for Multi-Agent Systems (MAS) designed to orchestrate complex practical workflows. Given that MAS performance is highly sensitive to input prompts and many deployment scenar…

Cited by 0SourceScholar
2026

OccTENS: 3D Occupancy World Model Via Temporal Next-Scale Prediction

ICRA 2026poster

In this paper, we propose OccTENS, a generative occupancy world model that enables controllable, high-fidelity long-term occupancy generation while maintaining computational efficiency. Different from visual generation, the occupancy world model must capture the fine-grained 3D geometry and dynamic …

2026

OccTENS: 3D Occupancy World Model via Temporal Next-Scale Prediction

RA-L 2026

In this paper, we propose OccTENS, a generative occupancy world model that enables controllable, high-fidelity long-term occupancy generation while maintaining computational efficiency. Different from visual generation, the occupancy world model must capture the fine-grained 3D geometry and dynamic

Cited by 7SourceScholar
2026

RS2-SAM2: Customized SAM2 for Referring Remote Sensing Image Segmentation

AAAI 2026technical

Referring Remote Sensing Image Segmentation (RRSIS) aims to segment target objects in remote sensing (RS) images based on textual descriptions. Although Segment Anything Model 2 (SAM2) has shown remarkable performance in various segmentation tasks, its application to RRSIS presents several challenge

Cited by 0SourcePDFScholar
2026

ResAD: Normalized Residual Trajectory Modeling for End-to-End Autonomous Driving

CVPR 2026

End-to-end autonomous driving (E2EAD) systems, which learn to predict future trajectories directly from sensor data, are fundamentally challenged by the inherent spatio-temporal imbalance of trajectory data. This imbalance creates a significant optimization burden, causing models to learn spurious c

Cited by 0SourcecodeScholar
2026

Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction

CVPR 2026

This paper addresses the task of large-scale 3D scene reconstruction from long video sequences. Recent feed-forward reconstruction models have shown promising results by directly regressing 3D geometry from RGB images without explicit 3D priors or geometric constraints. However, these methods often

Cited by 0SourcecodeScholar
2026

Towards Reliable Evaluation of Adversarial Robustness for Spiking Neural Networks

CVPR 2026

Spiking Neural Networks (SNNs) utilize spike-based activations to mimic the brain's energy-efficient information processing. However, the binary and discontinuous nature of spike activations causes vanishing gradients, making adversarial robustness evaluation via gradient descent unreliable. While i

Cited by 0SourcecodeScholar
2026

UV-RGS: Relightable 3D Gaussian Splatting from Unposed Views Under Varied Illuminations

AAAI 2026technical

The latest advancements in scene relighting have been predominantly driven by inverse rendering with 3D Gaussian Splatting (3DGS). However, existing methods remain overly reliant on precise camera parameters under static illumination conditions, which is prohibitively expensive and even impractical

Cited by 0SourcePDFScholar
2026

VADv2: End-to-End Autonomous Driving via Probabilistic Planning

ICLR 2026poster

Learning a human-like driving policy from large-scale driving demonstrations is promising, but the uncertainty and non-deterministic nature of planning make it challenging. Existing learning-based planning methods follow a deterministic paradigm to directly regress the action, failing to cope with t…

Cited by 0SourcecodeScholar
2026

VIL2C: Value-of-Information Aware Low-Latency Communication for Multi-Agent Reinforcement Learning

AAAI 2026technical

Inter-agent communication serves as an effective mechanism for enhancing performance in collaborative multi-agent reinforcement learning (MARL) systems. However, the inherent communication latency in practical systems induces both action decision delays and outdated information sharing, impeding MAR

Cited by 0SourcePDFScholar
2026

Write Where It Matters: Policy-Guided Watermarks for 3D Gaussian Splatting

CVPR 2026

Recent advances in 3D Gaussian Splatting (3DGS) enable photorealistic real-time rendering but also increase the risks of unauthorized copying and redistribution. Existing 3DGS watermarking methods typically rely on handcrafted thresholds or globally fixed hyperparameters to balance invisibility and

Cited by 0SourceScholar
2025

Boost 3D Reconstruction using Diffusion-based Monocular Camera Calibration

ICCV 2025poster

In this paper, we present DM-Calib, a diffusion-based approach for estimating pinhole camera intrinsic parameters from a single input image. Monocular camera calibration is essential for many 3D vision tasks. However, most existing methods depend on handcrafted assumptions or are constrained by limi…

2025

CityNavAgent: Aerial Vision-and-Language Navigation with Hierarchical Semantic Planning and Global Memory

ACL 2025long

Aerial vision-and-language navigation (VLN) — requiring drones to interpret natural language instructions and navigate complex urban environments — emerges as a critical embodied AI challenge that bridges human-robot interaction, 3D spatial reasoning, and real-world deployment. Although existing gro…

2025

ComDrive: Comfort-Oriented End-to-End Autonomous Driving

IROS 2025

We propose ComDrive: the first comfort-oriented end-to-end autonomous driving system to generate temporally consistent and comfortable trajectories. Recent studies have demonstrated that imitation learning-based planners and learning-based trajectory scorers can effectively generate and select safet

Cited by 14SourcecodeScholar
2025

DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving

CVPR 2025highlight

Recently, the diffusion model has emerged as a powerful generative technique for robotic policy learning, capable of modeling multi-mode action distributions. Leveraging its capability for end-to-end autonomous driving is a promising direction. However, the numerous denoising steps in the robotic di…

2025

Epona: Autoregressive Diffusion World Model for Autonomous Driving

ICCV 2025poster

Diffusion models have demonstrated exceptional visual quality in video generation, making them promising for autonomous driving world modeling. However, existing video diffusion-based world models struggle with flexible-length, long-horizon predictions and integrating trajectory planning. This is be…

2025

GoalFlow: Goal-Driven Flow Matching for Multimodal Trajectories Generation in End-to-End Autonomous Driving

CVPR 2025poster

We propose GoalFlow, an end-to-end autonomous driving method for generating high-quality multimodal trajectories. In autonomous driving scenarios, there is rarely a single suitable trajectory. Recent methods have increasingly focused on modeling multimodal trajectory distributions. However, they suf…

2025

Graph Neural Networks Meet Probabilistic Graphical Models: A Survey

ICASSP 2025accepted

Graphs are a powerful data structure for representing relational data, and Graph Neural Networks (GNNs) have emerged as effective tools for inference and learning on graph-structured data. Probabilistic Graphical Models (PGMs), which provide compact graphical representations of variable distribution…

Cited by 0SourceScholar
2025

HierPrompt: Zero-Shot Hierarchical Text Classification with LLM-Enhanced Prototypes

EMNLP 2025

Hierarchical Text Classification is a challenging task which classifies texts into categories arranged in a hierarchy. Zero‐Shot Hierarchical Text Classification (ZS-HTC) further assumes only the availability of hierarchical taxonomy, without any training data. Existing works of ZS-HTC are typically

Cited by 0SourcePDFScholar
2025

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation

ICCV 2025poster

Referring video object segmentation (RVOS) aims to segment objects in a video according to textual descriptions, which requires the integration of multimodal information and temporal dynamics perception. The Segment Anything Model 2 (SAM 2) has shown great effectiveness across various video segmenta…

2025

OccRWKV: Rethinking Efficient 3D Semantic Occupancy Prediction with Linear Complexity

ICRA 2025

3D semantic occupancy prediction networks have demonstrated remarkable capabilities in reconstructing the geometric and semantic structure of 3D scenes, providing crucial information for robot navigation and autonomous driving systems. However, due to their large overhead from dense network structur

Cited by 10SourcecodeScholar
2025

RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning

NeurIPS 2025poster

Existing end-to-end autonomous driving (AD) algorithms typically follow the Imitation Learning (IL) paradigm, which faces challenges such as causal confusion and an open-loop gap. In this work, we propose RAD, a 3DGS-based closed-loop Reinforcement Learning (RL) framework for end-to-end Autonomous D…

Cited by 0SourcecodeScholar
2025

SU-RGS: Relightable 3D Gaussian Splatting from Sparse Views under Unconstrained Illuminations

ICCV 2025poster

The latest advancements in scene relighting have been predominantly driven by inverse rendering with 3D Gaussian Splatting (3DGS). However, existing methods remain overly reliant on densely sampled images under static illumination conditions, which is prohibitively expensive and even impractical in…

Cited by 0SourcePDFScholar
2025

ScEdit: Script-based Assessment of Knowledge Editing

ACL 2025finding

Knowledge Editing (KE) has gained increasing attention, yet current KE tasks remain relatively simple. Under current evaluation frameworks, many editing methods achieve exceptionally high scores, sometimes nearing perfection. However, few studies integrate KE into real-world application scenarios (e…

2025

SynthDrive: Scalable Real2Sim2Real Sensor Simulation Pipeline for High-Fidelity Asset Generation and Driving Data Synthesis

IROS 2025

In the field of autonomous driving, sensor simulation is essential for generating rare and diverse scenarios that are difficult to capture in real-world environments. Current solutions fall into two categories: 1) CG-based methods, such as CARLA, which lack diversity and struggle to scale to the vas

Cited by 1SourceScholar
2025

ViG: Linear-complexity Visual Sequence Learning with Gated Linear Attention

AAAI 2025technical

Recently, linear complexity sequence modeling networks have achieved modeling capabilities similar to Vision Transformers on a variety of computer vision tasks, while using fewer FLOPs and less memory. However, their advantage in terms of actual runtime speed is not significant. To address this issu…

2025

Weak-to-Strong Generalization in Speech Recognition

ICASSP 2025accepted

To surpass human-level accuracy, speech recognition models must go beyond relying solely on human labels. To this end, we must build stronger models from weaker supervisors and this is the main goal in weak-to-strong generalization (WSG). WSG methods normally incorporate additional information into…

Cited by 0SourceScholar
2024

A Vision-Centric Approach for Static Map Element Annotation

ICRA 2024poster

The recent development of online static map element (a.k.a. HD Map) construction algorithms has raised a vast demand for data with ground truth annotations. However, available public datasets currently cannot provide high-quality training data regarding consistency and accuracy. To this end, we pres…

Cited by 3SourcecodeScholar
2024

An Efficient Algorithm for Multiuser Sum-Rate Maximization of Large-Scale Active RIS-Aided MIMO System

ICASSP 2024accepted

Active reconfigurable intelligent surface (RIS) is a new RIS architecture that can reflect and amplify communication signals. It can provide enhanced performance gain compared to the conventional passive RIS systems that can only reflect the signals. On the other hand, the design problem of active R…

Cited by 0SourceScholar
2024

Document Hashing with Multi-Grained Prototype-Induced Hierarchical Generative Model

EMNLP 2024finding

Document hashing plays a crucial role in large-scale information retrieval. However, existing unsupervised document hashing methods merely consider flat semantics of documents, resulting in the inability of preserving hierarchical semantics in hash codes. In this paper, we propose a hierarchical gen…

Cited by 1SourcePDFScholar
2024

Enhancing RAW-to-sRGB with Decoupled Style Structure in Fourier Domain

AAAI 2024technical

RAW to sRGB mapping, which aims to convert RAW images from smartphones into RGB form equivalent to that of Digital Single-Lens Reflex (DSLR) cameras, has become an important area of research. However, current methods often ignore the difference between cell phone RAW images and DSLR camera RGB image…

2024

Lane Graph as Path: Continuity-preserving Path-wise Modeling for Online Lane Graph Construction

ECCV 2024poster

"Online lane graph construction is a promising but challenging task in autonomous driving. Previous methods usually model the lane graph at the pixel or piece level, and recover the lane graph by pixel-wise or piece-wise connection, which breaks down the continuity of the lane and results in subopti…

2024

MARRGM: Learning Framework for Multi-Agent Reinforcement Learning via Reinforcement Recommendation and Group Modification

RA-L 2024

Sample usage efficiency is an important factor affecting the convergence speed of multi-agent deep reinforcement learning (MADRL) algorithms. Most existing experience replay (ER) methods manually select experience samples to update the agent's policy. It is difficult to give suitable and efficient e

Cited by 6SourceScholar
2024

Monte Carlo Self-Training for Speech Recognition

ICASSP 2024accepted

Self-training in the teacher-student framework generally suffers from the confirmation bias problem, where errors from the teacher are propagated to the student and hence get amplified with multiple iterations. In this paper, we present Monte Carlo Self-training where pseudo labels are generated by…

Cited by 0SourceScholar
2024

Neuro-Vision to Language: Enhancing Brain Recording-based Visual Reconstruction and Language Interaction

NeurIPS 2024poster

Decoding non-invasive brain recordings is pivotal for advancing our understanding of human cognition but faces challenges due to individual differences and complex neural signal representations. Traditional methods often require customized models and extensive trials, lacking interpretability in vis…

Cited by 3SourcePDFScholar
2024

Scale Optimization Using Evolutionary Reinforcement Learning for Object Detection on Drone Imagery

AAAI 2024technical

Object detection in aerial imagery presents a significant challenge due to large scale variations among objects. This paper proposes an evolutionary reinforcement learning agent, integrated within a coarse-to-fine object detection framework, to optimize the scale for more effective detection of obje…

2024

Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model

ICML 2024poster

Recently the state space models (SSMs) with efficient hardware-aware designs, i.e., the Mamba deep learning model, have shown great potential for long sequence modeling. Meanwhile building efficient and generic vision backbones purely upon SSMs is an appealing direction. However, representing visual…

2023

BAEFormer: Bi-Directional and Early Interaction Transformers for Bird's Eye View Semantic Segmentation

CVPR 2023poster

Bird's Eye View (BEV) semantic segmentation is a critical task in autonomous driving. However, existing Transformer-based methods confront difficulties in transforming Perspective View (PV) to BEV due to their unidirectional and posterior interaction mechanisms. To address this issue, we propose a n…

Cited by 25SourcePDFScholar
2023

BoxTeacher: Exploring High-Quality Pseudo Labels for Weakly Supervised Instance Segmentation

CVPR 2023poster

Labeling objects with pixel-wise segmentation requires a huge amount of human labor compared to bounding boxes. Most existing methods for weakly supervised instance segmentation focus on designing heuristic losses with priors from bounding boxes. While, we find that box-supervised methods can produc…

2023

Cross-Training: A Semi-Supervised Training Scheme for Speech Recognition

ICASSP 2023accepted

Semi-supervised training can be performed by jointly optimizing supervised and unsupervised losses. In many settings, supervised and unsupervised losses are inconsistent, and this inconsistency creates instability in training. As a solution, we propose cross-training: instead of training one network…

Cited by 3SourceScholar
2023

Interpretable Motion Planner for Urban Driving via Hierarchical Imitation Learning

IROS 2023poster

Learning-based approaches have achieved remarkable performance in the domain of autonomous driving. Leveraging the impressive ability of neural networks and large amounts of human driving data, complex patterns and rules of driving behavior can be encoded as a model to benefit the autonomous driving…

Cited by 5SourceScholar
2023

MapTR: Structured Modeling and Learning for Online Vectorized HD Map Construction

ICLR 2023top-25%

High-definition (HD) map provides abundant and precise environmental information of the driving scene, serving as a fundamental and indispensable component for planning in autonomous driving system. We present MapTR, a structured end-to-end Transformer for efficient online vectorized HD map construc…

2023

Multi-Granularity Archaeological Dating of Chinese Bronze Dings Based on a Knowledge-Guided Relation Graph

CVPR 2023poster

The archaeological dating of bronze dings has played a critical role in the study of ancient Chinese history. Current archaeology depends on trained experts to carry out bronze dating, which is time-consuming and labor-intensive. For such dating, in this study, we propose a learning-based approach t…

2023

Non-reversible Parallel Tempering for Deep Posterior Approximation

AAAI 2023technical

Parallel tempering (PT), also known as replica exchange, is the go-to workhorse for simulations of multi-modal distributions. The key to the success of PT is to adopt efficient swap schemes. The popular deterministic even-odd (DEO) scheme exploits the non-reversibility property and has successfully…

Cited by 6SourcePDFScholar
2023

VAD: Vectorized Scene Representation for Efficient Autonomous Driving

ICCV 2023poster

Autonomous driving requires a comprehensive understanding of the surrounding environment for reliable trajectory planning. Previous works rely on dense rasterized scene representation (e.g., agent occupancy and semantic map) to perform planning, which is computationally intensive and misses the inst…

Cited by 233PDFcodeScholar
2022

AdaptivePose: Human Parts as Adaptive Points

AAAI 2022technical

Multi-person pose estimation methods generally follow top-down and bottom-up paradigms, both of which can be considered as two-stage approaches thus leading to the high computation cost and low efficiency. Towards a compact and efficient pipeline for multi-person pose estimation task, in this paper,…

Cited by 26SourcePDFScholar
2022

AziNorm: Exploiting the Radial Symmetry of Point Cloud for Azimuth-Normalized 3D Perception

CVPR 2022poster

Studying the inherent symmetry of data is of great importance in machine learning. Point cloud, the most important data format for 3D environmental perception, is naturally endowed with strong radial symmetry. In this work, we exploit this radial symmetry via a divide-and-conquer strategy to boost 3…

Cited by 7PDFcodeScholar
2022

Contrastive Siamese Network for Semi-Supervised Speech Recognition

ICASSP 2022accepted

This paper introduces contrastive siamese (c-siam) network, an architecture for leveraging unlabeled acoustic data in speech recognition. c-siam is the first network that extracts high-level linguistic information from speech by matching outputs of two identical transformer encoders. It contains aug…

Cited by 17SourceScholar
2022

Cross-Domain Correlation Distillation for Unsupervised Domain Adaptation in Nighttime Semantic Segmentation

CVPR 2022poster

The performance of nighttime semantic segmentation is restricted by the poor illumination and a lack of pixel-wise annotation, which severely limit its application in autonomous driving. Existing works, e.g., using the twilight as the intermediate target domain to perform the adaptation from daytime…

Cited by 91PDFcodeScholar
2022

Cross-Image Relational Knowledge Distillation for Semantic Segmentation

CVPR 2022poster

Current Knowledge Distillation (KD) methods for semantic segmentation often guide the student to mimic the teacher's structured information generated from individual data samples. However, they ignore the global semantic relations among pixels across various images that are valuable for KD. This pap…

Cited by 247PDFcodeScholar
2022

Learning Quality-Aware Representation for Multi-Person Pose Regression

AAAI 2022technical

Off-the-shelf single-stage multi-person pose regression methods generally leverage the instance score (i.e., confidence of the instance localization) to indicate the pose quality for selecting the pose candidates. We consider that there are two gaps involved in existing paradigm: 1) The instance sco…

Cited by 17SourcePDFScholar
2022

Learning from the Target: Dual Prototype Network for Few Shot Semantic Segmentation

AAAI 2022technical

Due to the scarcity of annotated samples, the diversity between support set and query set becomes the main obstacle for few shot semantic segmentation. Most existing prototype-based approaches only exploit the prototype from the support feature and ignore the information from the query sample, faili…

Cited by 22SourcePDFScholar
2022

MixSKD: Self-Knowledge Distillation from Mixup for Image Recognition

ECCV 2022poster

"Unlike the conventional Knowledge Distillation (KD), Self-KD allows a network to learn knowledge from itself without any guidance from extra networks. This paper proposes to perform Self-KD from image Mixture (MixSKD), which integrates these two techniques into a unified framework. MixSKD mutually…

2022

OCTOANTS: A Heterogeneous Lightweight Intelligent Multi-Robot Collaboration System with Resource-constrained IoT Devices

IROS 2022poster

As the focus on highly intelligent robots continues, a problem that cannot be ignored has emerged: resource con-straints. Considering the game problem of resource limitation and the level of intelligence, we focus on lightweight intelligence. This work is a further refinement of our previous work, a…

Cited by 2SourceScholar
2022

Sparse Instance Activation for Real-Time Instance Segmentation

CVPR 2022poster

In this paper, we propose a conceptually novel, efficient, and fully convolutional framework for real-time instance segmentation. Previously, most instance segmentation methods heavily rely on object detection and perform mask prediction based on bounding boxes or dense centers. In contrast, we prop…

Cited by 182PDFcodeScholar
2022

Spatial-Context-Aware Deep Neural Network for Multi-Class Image Classification

ICASSP 2022accepted

Multi-label image classification is a fundamental but challenging task in computer vision. Over the past few decades, solutions exploring relationships between semantic labels have made great progress. However, the underlying spatial-contextual information of labels is under-exploited. To tackle thi…

Cited by 0SourceScholar
2022

Vision-based Uneven BEV Representation Learning with Polar Rasterization and Surface Estimation

CoRL 2022poster

In this work, we propose PolarBEV for vision-based uneven BEV representation learning. To adapt to the foreshortening effect of camera imaging, we rasterize the BEV space both angularly and radially, and introduce polar embedding decomposition to model the associations among polar grids. Polar gri…

Cited by 26SourcecodeScholar
2021

Addressing Domain Gap via Content Invariant Representation for Semantic Segmentation

AAAI 2021technical

The problem of unsupervised domain adaptation in semantic segmentation is a major challenge for numerous computer vision tasks because acquiring pixel-level labels is time-consuming with expensive human labor. A large gap exists among data distributions in different domains, which will cause severe…

Cited by 18SourcePDFScholar
2021

Conquering Textureless with RF-referenced Monocular Vision for MAV State Estimation

ICRA 2021poster

The versatile nature of agile micro aerial vehicles (MAVs) poses fundamental challenges to the design of robust state estimation in various complex environments. Achieving high-quality performance in textureless scenes is one of the missing pieces in the puzzle. Previously proposed solutions either…

Cited by 5SourcecodeScholar
2021

Deep Online Correction for Monocular Visual Odometry

ICRA 2021poster

In this work, we propose a novel deep online correction (DOC) framework for monocular visual odometry. The whole pipeline has two stages: First, depth maps and initial poses are obtained from convolutional neural networks (CNNs) trained in self-supervised manners. Second, the poses predicted by CNNs…

Cited by 27SourceScholar
2021

Hierarchical Aggregation for 3D Instance Segmentation

ICCV 2021poster

Instance segmentation on point clouds is a fundamental task in 3D scene perception. In this work, we propose a concise clustering-based framework named HAIS, which makes full use of spatial relation of points and point sets. Considering clustering-based methods may result in over-segmentation or und…

Cited by 193PDFcodeScholar
2021

Meta Learning for Support Recovery in High-dimensional Precision Matrix Estimation

ICML 2021spotlight

In this paper, we study meta learning for support (i.e., the set of non-zero entries) recovery in high-dimensional precision matrix estimation where we reduce the sufficient sample complexity in a novel task with the information learned from other auxiliary tasks. In our setup, each task has a diffe…

Cited by 5SourcePDFScholar
2021

Rethinking Soft Labels for Knowledge Distillation: A Bias–Variance Tradeoff Perspective

ICLR 2021poster

Knowledge distillation is an effective approach to leverage a well-trained network or an ensemble of them, named as the teacher, to guide the training of a student network. The outputs from the teacher network are used as soft labels for supervising the training of a new network. Recent studies (M…

2021

Robust Knowledge Transfer via Hybrid Forward on the Teacher-Student Model

AAAI 2021technical

When adopting deep neural networks for a new vision task, a common practice is to start with fine-tuning some off-the-shelf well-trained network models from the community. Since a new task may require training a different network architecture with new domain data, taking advantage of off-the-shelf m…

Cited by 13SourcePDFScholar
2021

Stacked Homography Transformations for Multi-View Pedestrian Detection

ICCV 2021poster

Multi-view pedestrian detection aims to predict a bird's eye view (BEV) occupancy map from multiple camera views. This task is confronted with two challenges: how to establish the 3D correspondences from views to the BEV map and how to assemble occupancy information across views. In this paper, we p…

Cited by 56PDFScholar
2021

THOR, Trace-based Hardware-driven Layer-Oriented Natural Gradient Descent Computation

AAAI 2021technical

It is well-known that second-order optimizer can accelerate the training of deep neural networks, however, the huge computation cost of second-order optimization makes it impractical to apply in real practice. In order to reduce the cost, many methods have been proposed to approximate a second-order…

Cited by 9SourcePDFScholar
2020

AugFPN: Improving Multi-Scale Feature Learning for Object Detection

CVPR 2020poster

Current state-of-the-art detectors typically exploit feature pyramid to detect objects at different scales. Among them, FPN is one of the representative works that build a feature pyramid by multi-scale features summation. However, the design defects behind prevent the multi-scale features from bein…

Cited by 589PDFcodeScholar
2020

Densely Connected Search Space for More Flexible Neural Architecture Search

CVPR 2020poster

Neural architecture search (NAS) has dramatically advanced the development of neural network design. We revisit the search space design in most previous NAS methods and find the number and widths of blocks are set manually. However, block counts and block widths determine the network scale (depth an…

Cited by 165PDFcodeScholar
2020

Fast Neural Network Adaptation via Parameter Remapping and Architecture Search

ICLR 2020poster

Deep neural networks achieve remarkable performance in many computer vision tasks. Most state-of-the-art~(SOTA) semantic segmentation and object detection approaches reuse neural network architectures designed for image classification as the backbone, commonly pre-trained on ImageNet. However, perfo…

Cited by 44SourcecodeScholar
2020

FasterSeg: Searching for Faster Real-time Semantic Segmentation

ICLR 2020poster

We present FasterSeg, an automatically designed semantic segmentation network with not only state-of-the-art performance but also faster speed than current methods. Utilizing neural architecture search (NAS), FasterSeg is discovered from a novel and broader search space integrating multi-resolution…

Cited by 255SourcecodeScholar
2020

Learning Where to Focus for Efficient Video Object Detection

ECCV 2020poster

Transferring existing image-based detectors to the video is non-trivial since the quality of frames is always deteriorated by part occlusion, rare pose, and motion blur. Previous approaches exploit to propagate and aggregate features across video frames by using optical flow-warping. However, direct…

2020

Temporal-Context Enhanced Detection of Heavily Occluded Pedestrians

CVPR 2020poster

State-of-the-art pedestrian detectors have performed promisingly on non-occluded pedestrians, yet they are still confronted by heavy occlusions. Although many previous works have attempted to alleviate the pedestrian occlusion issue, most of them rest on still images. In this paper, we exploit the l…

Cited by 69PDFScholar
2020

Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss

ICASSP 2020accepted

In this paper we present an end-to-end speech recognition model with Transformer encoders that can be used in a streaming speech recognition system. Transformer computation blocks based on self-attention are used to encode both audio and label sequences independently. The activations from both audio…

Cited by 0SourceScholar
2019

Progressive Sparse Local Attention for Video Object Detection

ICCV 2019poster

Transferring image-based object detectors to the domain of videos remains a challenging problem. Previous efforts mostly exploit optical flow to propagate features across frames, aiming to achieve a good trade-off between accuracy and efficiency. However, introducing an extra model to estimate optic…

Cited by 113PDFScholar
2019

RENAS: Reinforced Evolutionary Neural Architecture Search

CVPR 2019poster

Neural Architecture Search (NAS) is an important yet challenging task in network design due to its high computational consumption. To address this issue, we propose the Reinforced Evolutionary Neural Architecture Search (RENAS), which is an evolutionary method with reinforced mutation for NAS. Our m…

Cited by 154PDFScholar
2019

Semi-supervised Learning with Generative Adversarial Networks for Arabic Dialect Identification

ICASSP 2019accepted

Dialect Identification (DID) refers to the process of identifying different dialects within the same language class. Compared with more general language identification (LID), DID is a more challenging task because of the substantial similarity between dialects. For an i-vector based LID/DID, prior s…

Cited by 10SourceScholar
2019

View-Consistent 4D Light Field Superpixel Segmentation

ICCV 2019oral

Many 4D light field processing applications rely on superpixel segmentations, for which occlusion-aware view consistency is important. Yet, existing methods often enforce consistency by propagating clusters from a central view only, which can lead to inconsistent superpixels for non-central views. O…

Cited by 30PDFcodeScholar
2018

Mancs: A Multi-task Attentional Network with Curriculum Sampling for Person Re-identification

ECCV 2018poster

We propose a novel deep network called Mancs that solves the person re-identification problem from the following aspects: fully utilizing the attention mechanism for the person misalignment problem and properly sampling for the ranking loss to obtain more stable person representation. Technically, w…

Cited by 501SourcePDFScholar
2016

Gauss-Seidel based non-negative matrix factorization for gene expression clustering

ICASSP 2016accepted

Genome-wide expression data consists of millions of measurements towards large number of genes, and thus it is challenging for human beings to directly analyze such large-scale data. Clustering provides a more convenient way to analyze gene expression data because it can subdivide raw data into comp…

Cited by 0SourceScholar
2016

Joint information from nonlinear and linear features for spoofing detection: An i-vector/DNN based approach

ICASSP 2016accepted

Sustaining automatic speaker verification(ASV) systems from spoofing attacks remains an essential challenge, even if significant progress in ASV has been achieved in recent years. In this study, an automatic spoofing detection approach using an i-vector framework is proposed. Two approaches are used…

Cited by 20SourceScholar
2016

UTD-CRSS system for the NIST 2015 language recognition i-vector machine learning challenge

ICASSP 2016accepted

In this paper, we present the system developed by the Center for Robust Speech Systems (CRSS), University of Texas at Dallas, for the NIST 2015 language recognition i-vector machine learning challenge. Our system includes several subsystems, based on Linear Discriminant Analysis - Support Vector Mac…

Cited by 0SourceScholar
2015

Fine-Grained Change Detection of Misaligned Scenes With Varied Illuminations

ICCV 2015poster

Detecting fine-grained subtle changes among a scene is critically important in practice. Previous change detection methods, focusing on detecting large-scale significant changes, cannot do this well. This paper proposes a feasible end-to-end approach to this challenging problem. We start from active…

Cited by 43PDFScholar