← Search

Yan Li

90 accepted papers

2026

Active3D: Active High-Fidelity 3D Reconstruction via Multi-Level Uncertainty Quantification

AAAI 2026technical

In this paper, we present an active exploration framework for high-fidelity 3D reconstruction that incrementally builds a multi-level uncertainty space and selects next-best-views through an uncertainty-driven motion planner. We introduce a hybrid implicit–explicit representation that fuses neural

Cited by 0SourcePDFScholar
2026

AdapTok: Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Space

CVPR 2026

We propose AdapTok, an adaptive temporal causal video tokenizer that can flexibly allocate tokens for different frames based on video content. AdapTok is equipped with a block-wise masking strategy that randomly drops tail tokens of each block during training, and a block causal scorer to predict th

Cited by 0SourcecodeScholar
2026

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning

ICML 2026poster

Reinforcement Learning (RL) has become a cornerstone for improving the performance of Large Language Models (LLMs). However, its rollout phase constitutes a significant efficiency bottleneck, mainly arising from the long-tail bubbles across data parallel ranks, particularly in long-context scenarios…

Cited by 0SourceScholar
2026

CogniEdit: Dense Gradient Flow Optimization for Fine-Grained Image Editing

CVPR 2026

Instruction-based image editing with diffusion models has achieved impressive results, yet existing methods struggle with fine-grained instructions specifying precise attributes such as colors, positions, and quantities. While recent approaches employ Group Relative Policy Optimization (GRPO) for al

Cited by 0SourcecodeScholar
2026

DSAP: Enhancing Generalization in Goal-Conditioned Reinforcement Learning

AAAI 2026technical

Goal-conditioned Reinforcement Learning (RL) is a promising direction for training agents capable of tackling a variety of tasks. However, generalizing to new goals in different environments remains a central challenge for goal-conditioned RL agents. Existing methods often rely on state abstraction,

Cited by 0SourcePDFScholar
2026

DuetMerging: Synergizing Dynamic and Static Strategies for Mitigating Task Interference in Model Merging

CVPR 2026

Model merging offers a promising paradigm for consolidating multiple expert models into a single multitask architecture. However, its effectiveness is often hindered by task interference, where conflicting parameter updates from different tasks degrade performance. While dynamic, Mixture-of-Experts

Cited by 0SourceScholar
2026

EvoGM: Learning to Merge LLMs via Evolutionary Generative Optimization

ICML 2026poster

Evolutionary model merging provides a powerful framework for the automated, training-free composition of LLMs through parameter-space search. However, existing methods predominantly rely on stochastic, hand-crafted operators that overlook the underlying performance landscape of the coefficient space…

Cited by 0SourceScholar
2026

Expected Returns and Policy Inconsistency-Aware Offline Federated Deep Reinforcement Learning

ICML 2026poster

Offline Federated Deep Reinforcement Learning (FDRL) methods aggregate multiple client-side offline Deep Reinforcement Learning (DRL) models, each trained locally, to facilitate knowledge sharing while preserving privacy. Existing offline FDRL methods assign client weights during global aggregation …

Cited by 0SourceScholar
2026

Explore to Learn: Latent Exploration Through Disentangled Synergy Patterns for Reinforcement Learning in Overactuated Control

AAAI 2026technical

Control in high-dimensional action spaces remains a fundamental challenge in reinforcement learning (RL), primarily due to inefficient exploration of the action space. While recent methods attempt to guide exploration, they often fall short of achieving the agility and coordination exhibited in biol

Cited by 0SourcePDFScholar
2026

Few-Shot Hybrid Incremental Learning:Continually Learning under Data Scarcity and Task Uncertainty

CVPR 2026

The increasing complexity of real-world deployment requires intelligent agents to effectively adapt to non-stationary data streams with stochastic increments under data scarcity. We formally define this challenge as the Few-Shot Hybrid Incremental Learning (FSHIL) paradigm, which reveals a critical

Cited by 0SourceScholar
2026

Half-order Fine-Tuning for Diffusion Model: A Recursive Likelihood Ratio Optimizer

ICLR 2026oral

The probabilistic diffusion model (DM), generating content by inferencing through a recursive chain structure, has emerged as a powerful framework for visual generation. After pre-training on enormous data, the model needs to be properly aligned to meet requirements for downstream applications. How…

Cited by 0SourcecodeScholar
2026

JoPPO: Hierarchical Photography Assessment via Contrastive Joint Conditional Probabilistic Reinforcement Learning

CVPR 2026

With the advancement of Vision-Language Models (VLMs), employing VLM-as-a-Judge for visual evaluation has become a widely adopted metric in vision research. However, existing VLM-as-a-Judge approaches suffer from biased scoring outcomes with low discrimination and lack the capacity for unified multi

Cited by 0SourcecodeScholar
2026

Latent State-Predictive Exploration for Deep Reinforcement Learning

AAAI 2026technical

Reinforcement learning (RL) has achieved promising results in continuous control tasks, where efficient exploration of the state space is crucial for success. However, many recent RL approaches still struggle with sample inefficiency and insufficient exploration for long-horizon tasks, particularly

Cited by 0SourcePDFScholar
2026

Manipulating the Mind’s Eye: A-SAGE, the Attention-Based Attack on ViT Explainability

AAAI 2026technical

The rise of Vision Transformers (ViTs) as cornerstone models in safety-critical applications like autonomous driving and medical diagnosis has shifted the focus from pure accuracy to verifiable trustworthiness. However, the very mechanisms used to explain these models, their internal attention maps,

Cited by 0SourcePDFScholar
2026

Pixel-Perfect Puppetry: Precision-Guided Enhancement for Face Image and Video Editing

ICLR 2026poster

Preserving identity while precisely manipulating attributes is a central challenge in face editing for both images and videos. Existing methods often introduce visual artifacts or fail to maintain temporal consistency. We present **FlowGuide**, a unified framework that achieves fine-grained control…

Cited by 0SourceScholar
2026

ProCURE: Addressing the Programming Concept Understanding Gap for Code Generation in LLMs via Concept-Aware Consistency Learning

IJCAI 2026

Although Large Language Models (LLMs) excel at code generation, recent research reveals that they exhibit an insufficient grasp of core programming concepts, such as data flow and control flow. This limitation undermines their robustness when encountering variations in these concepts in practice; ho

Cited by 0Scholar
2026

Rethink Representation Learning for Questionnaire Data

AAAI 2026technical

Questionnaire data serve as a valuable resource across numerous scientific domains, offering insights into human behavior, health, and social trends. Traditional downsampling-based representation learning methods—such as standardization and one-hot encoding—reformat these data into tabular structure

Cited by 0SourcePDFScholar
2026

RiemanLine: Riemannian Manifold Representation of 3D Lines for Factor Graph Optimization

AAAI 2026technical

Minimal parametrization of 3D lines plays a critical role in camera localization and structural mapping. Existing representations in robotics and computer vision predominantly handle independent lines, overlooking structural regularities such as sets of parallel lines that are pervasive in man-made

Cited by 0SourcePDFScholar
2026

Scaling-Aware Adapter for Structure-Grounded LLM Reasoning

ICML 2026poster

Large language models (LLMs) enable reasoning over biomolecular structures, yet existing methods remain modality-specific and typically compress structural inputs via sequence-based tokenization or fixed-length query connectors. Such architectures either omit geometric grounding required to mitigate…

Cited by 0SourceScholar
2026

Selection, Reflection and Self-Refinement: Revisit Reasoning Tasks via a Causal Lens

ICLR 2026poster

Due to their inherent complexity, reasoning tasks have long been regarded as rigorous benchmarks for assessing the capabilities of machine learning models, especially large language models (LLMs). Although humans can solve these tasks with ease, existing models, even after extensive pre-training and…

Cited by 0SourcecodeScholar
2026

Sem-MoE: Semantic-aware Model-Data Collaborative Scheduling for Efficient MoE Inference

ICLR 2026poster

Prevailing LLM (Large Language Model) serving engines employ expert parallelism (EP) to implement multi-device inference of massive Mixture-of-Experts (MoE) models. However, the efficiency of expert parallel inference is largely bounded by inter-device communication, as EP embraces expensive all-to-…

Cited by 0SourceScholar
2026

SlerpFlow: Spherical Trajectory Correction for Rectified Flow Inversion

ICML 2026poster

Rectified-flow-based diffusion transformers, particularly FLUX, have demonstrated outstanding performance in high-quality image generation. However, achieving fast and accurate inversion—transforming images back to latent noise for faithful reconstruction and editing—remains a challenging bottleneck…

Cited by 0SourceScholar
2026

UniAPO: Unified Multimodal Automated Prompt Optimization

AAAI 2026technical

Prompting is fundamental to unlocking the full potential of large language models. To automate and enhance this process, automatic prompt optimization (APO) has been developed, demonstrating effectiveness primarily in text-only input scenarios. However, extending existing APO methods to multimodal t

Cited by 0SourcePDFScholar
2026

VirtueBench: Evaluating Trustworthiness under Uncertainty in Long Video Understanding

CVPR 2026

Recent Vision-Language Models (VLMs) have made remarkable progress in multimodal understanding tasks, yet their evaluation on long video understanding remains unreliable. Due to limited frame inputs, key frames necessary for answering the question may be missing from the model's input. However, mode

Cited by 0SourceScholar
2025

A General Representation-Based Approach to Multi-Source Domain Adaptation

ICML 2025poster

A central problem in unsupervised domain adaptation is determining what to transfer from labeled source domains to an unlabeled target domain. To handle high-dimensional observations (e.g., images), a line of approaches use deep learning to learn latent representations of the observations, which fac…

Cited by 0SourcePDFScholar
2025

A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation

EMNLP 2025

Transformer-based Large Language Models (LLMs) struggle with inputs exceeding their training context window due to positional out-of-distribution (O.O.D.) issues that disrupt attention. Existing solutions, including fine-tuning and training-free methods, face challenges like inefficiency, redundant

2025

AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment

CVPR 2025poster

Cross-modal alignment is crucial for multimodal representation fusion due to the inherent heterogeneity between modalities. While Transformer-based methods have shown promising results in modeling inter-modal relationships, their quadratic computational complexity limits their applicability to long-…

Cited by 1SourcePDFScholar
2025

Appearance- and Orientation-aware Fine-grained Rotated Ship Detection in High-Resolution Satellite Imagery

ICASSP 2025accepted

Ship detection using remote sensing imagery is a crucial research area with both military and civilian applications. However, it remains challenging due to limitations in current ship datasets, such as insufficient volume, incomplete annotations, and inaccuracies. Additionally, ships often exhibit a…

Cited by 0SourceScholar
2025

BILE: An Effective Behavior-based Latent Exploration Scheme for Deep Reinforcement Learning

IJCAI 2025

Efficient exploration of state spaces is critical for the success of deep reinforcement learning (RL). While many methods leverage exploration bonuses to encourage exploration instead of relying solely on extrinsic rewards, these bonus-based approaches often face challenges with learning efficiency

Cited by 0SourcePDFScholar
2025

CausalVerse: Benchmarking Causal Representation Learning with Configurable High-Fidelity Simulations

NeurIPS 2025spotlight

Causal Representation Learning (CRL) aims to uncover the data-generating process and identify the underlying causal variables and relations, whose evaluation remains inherently challenging due to the requirement of known ground-truth causal variables and causal structure. Existing evaluations often…

Cited by 0SourcecodeScholar
2025

ChuLo: Chunk-Level Key Information Representation for Long Document Understanding

ACL 2025finding

Transformer-based models have achieved remarkable success in various Natural Language Processing (NLP) tasks, yet their ability to handle long documents is constrained by computational limitations. Traditional approaches, such as truncating inputs, sparse self-attention, and chunking, attempt to mit…

2025

Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization

NeurIPS 2025poster

Preference optimization for diffusion models aims to align them with human preferences for images. Previous methods typically use Vision-Language Models (VLMs) as pixel-level reward models to approximate human preferences. However, when used for step-level preference optimization, these models face…

Cited by 0SourcecodeScholar
2025

EPI-Mamba: State Space Model for Semantic Segmentation from Light Fields

ICASSP 2025accepted

Global contextual dependency is of significance for semantic segmentation from light fields. However, previous works mostly exploit attention mechanisms to model spatial context dependency and angular context dependency separately, since a light field capture is very data-intensive. Considering that…

Cited by 0SourceScholar
2025

Efficient Diversity-based Experience Replay for Deep Reinforcement Learning

IJCAI 2025

Experience replay is widely used to improve learning efficiency in reinforcement learning by leveraging past experiences. However, existing experience replay methods, whether based on uniform or prioritized sampling, often suffer from low efficiency, particularly in real-world scenarios with high-di

Cited by 0SourcePDFScholar
2025

Heteroscedastic Bayesian Optimization-Based Dynamic PID Tuning for Accurate and Robust UAV Trajectory Tracking

IROS 2025

Unmanned Aerial Vehicles (UAVs) play an important role in various applications, where precise trajectory tracking is crucial. However, conventional control algorithms for trajectory tracking often exhibit limited performance due to the underactuated, nonlinear, and highly coupled dynamics of quadrot

Cited by 1SourceScholar
2025

LEARN: Knowledge Adaptation from Large Language Model to Recommendation for Practical Industrial Application

AAAI 2025technical

Contemporary recommendation systems predominantly rely on ID embedding to capture latent associations among users and items. However, this approach overlooks the wealth of semantic information embedded within textual descriptions of items, leading to suboptimal performance and poor generalizations.…

2025

Learning vector fields of differential equations on manifolds with geometrically constrained operator-valued kernels

ICLR 2025spotlight

We address the problem of learning ordinary differential equations (ODEs) on manifolds. Existing machine learning methods, particularly those using neural networks, often struggle with high computational demands. To overcome this issue, we introduce a geometrically constrained operator-valued kernel…

Cited by 0SourcePDFScholar
2025

M2EIT: Multi-Domain Mixture of Experts for Robust Neural Inertial Tracking

ICCV 2025poster

Inertial tracking (IT), independent of the environment and external infrastructure, has long been the ideal solution for providing location services to humans. Despite significant strides in inertial tracking empowered by deep learning, prevailing neural inertial tracking predominantly utilizes conv…

Cited by 0SourcePDFScholar
2025

MIDAS: Multi-level Intent, Domain, And Slot Knowledge Distillation for Multi-turn NLU

NAACL 2025findings

Although Large Language Models (LLMs) can generate coherent text, they often struggle to recognise user intent behind queries. In contrast, Natural Language Understanding (NLU) models interpret the purpose and key information of user input for responsive interactions. Existing NLU models typically m…

2025

MKD-YOLO: Multi-Scale and Knowledge-Distilling YOLO for Efficient PPE Compliance Detection

ICASSP 2025accepted

YOLO-based models are widely used for personal protective equipment (PPE) compliance detection due to their excellent detection performance and efficiency. However, most YOLO models are not competent for detection tasks in complex industrial scenarios such as remote surveillance and extremely small…

Cited by 0SourceScholar
2025

MUSE: Multi-Subject Unified Synthesis via Explicit Layout Semantic Expansion

ICCV 2025poster

Existing text-to-image diffusion models have demonstrated remarkable capabilities in generating high-quality images guided by textual prompts. However, achieving multi-subject compositional synthesis with precise spatial control remains a significant challenge. In this work, we address the task of l…

2025

Pipeline-Centered Neighboring Network for Deep Unfolding Pansharpening

ICASSP 2025accepted

Pansharpening technique is dedicated to enriching the spatial details of low-resolution multispectral images (LRMS) under the guidance of a panchromatic (PAN) image. With the guarantee of promising results, Transformer-based methods have enjoyed a high reputation in this field. However, to reduce co…

Cited by 0SourceScholar
2025

SCOUT: Teaching Pre-trained Language Models to Enhance Reasoning via Flow Chain-of-Thought

NeurIPS 2025poster

Chain-of-Thought (CoT) prompting improves the reasoning performance of large language models (LLMs) by encouraging step-by-step thinking. However, CoT-based methods depend on intermediate reasoning steps, which limits scalability and generalization. Recent work explores recursive reasoning, where L…

Cited by 0SourceScholar
2025

Towards Self-Refinement of Vision-Language Models with Triangular Consistency

NeurIPS 2025poster

Vision-Language Models (VLMs) integrate visual knowledge with the analytical capabilities of Large Language Models (LLMs) through supervised visual instruction tuning, using image-question-answer triplets. However, the potential of VLMs trained without supervised instruction remains largely unexplor…

Cited by 0SourcecodeScholar
2025

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding

IJCAI 2025

Visually Rich Document Understanding (VRDU) has emerged as a critical field in document intelligence, enabling automated extraction of key information from complex documents across domains such as medical, financial, and educational applications. However, form-like documents pose unique challenges d

Cited by 0SourcePDFScholar
2025

When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding

NeurIPS 2025poster

Large Multimodal Models (LMMs) have achieved impressive progress in visual perception and reasoning. However, when confronted with visually ambiguous or non-semantic scene text, they often struggle to accurately spot and understand the content, frequently generating semantically plausible yet visual…

Cited by 0SourceScholar
2024

Drag Anything: Motion Control for Anything using Entity Representation

ECCV 2024poster

"We introduce , which utilizes a entity representation to achieve motion control for any object in controllable video generation. Comparison to existing motion control methods, offers several advantages. Firstly, trajectory-based is more user-friendly for interaction, when acquiring other guidance s…

2024

Frequency Aware and Graph Fusion Network for Polyp Segmentation

ICASSP 2024accepted

Polyp segmentation plays a crucial role in the prevention of colon cancer. However, the diverse shapes of polyps and their similarity to normal areas in terms of color and texture make polyp segmentation a challenging task. Currently, most polyp segmentation methods solely focus on spatial domain fe…

Cited by 0SourceScholar
2024

Gradient and Brightness Guided Low-Light Enhancement with Attention-Based Self-Paced Learning

ICASSP 2024accepted

Low-light image enhancement aims to reconstruct images with insufficient illumination into visually appealing representations with natural brightness. While most existing methods tend to focus on enhancing illumination, they often overlook the restoration of finer details in the enhanced image. More…

Cited by 0SourceScholar
2024

Learning Multi-Dimensional Human Preference for Text-to-Image Generation

CVPR 2024poster

Current metrics for text-to-image models typically rely on statistical metrics which inadequately represent the real preference of humans. Although recent work attempts to learn these preferences via human annotated images they reduce the rich tapestry of human preference to a single overall score.…

Cited by 24SourcePDFScholar
2024

Toward Open Vocabulary Aerial Object Detection with CLIP-Activated Student-Teacher Learning

ECCV 2024poster

"An increasingly massive number of remote-sensing images spurs the development of extensible object detectors that can detect objects beyond training categories without costly collecting new labeled data. In this paper, we aim to develop open-vocabulary object detection (OVD) technique in aerial ima…

2023

Cross-Domain Product Representation Learning for Rich-Content E-Commerce

ICCV 2023poster

The proliferation of short video and live-streaming platforms has revolutionized how consumers engage in online shopping. Instead of browsing product pages, consumers are now turning to rich-content e-commerce, where they can purchase products through dynamic and interactive media like short videos…

Cited by 3PDFcodeScholar
2023

Cross-view Semantic Alignment for Livestreaming Product Recognition

ICCV 2023poster

Live commerce is the act of selling products online through livestreaming. The customer's diverse demands for online products introduces more challenges to Livestreaming Product Recognition. Previous works are either focus on fashion clothing data or subject to single-modal input, thus inconsistent…

Cited by 4PDFcodeScholar
2023

Robust Multi-Agent Reinforcement Learning via Adversarial Regularization: Theoretical Foundation and Stable Algorithms

NeurIPS 2023poster

Multi-Agent Reinforcement Learning (MARL) has shown promising results across several domains. Despite this promise, MARL policies often lack robustness and are therefore sensitive to small changes in their environment. This presents a serious concern for the real world deployment of MARL algorithms,…

2022

A Hybrid Causal Structure Learning Algorithm for Mixed-Type Data

AAAI 2022technical

Inferring the causal structure of a set of random variables is a crucial problem in many disciplines of science. Over the past two decades, various approaches have been pro- posed for causal discovery from observational data. How- ever, most of the existing methods are designed for either purely dis…

2022

Frequency-aware SGD for Efficient Embedding Learning with Provable Benefits

ICLR 2022poster

Embedding learning has found widespread applications in recommendation systems and natural language modeling, among other domains. To learn quality embeddings efficiently, adaptive learning rate algorithms have demonstrated superior empirical performance over SGD, largely accredited to their token-d…

Cited by 5SourcePDFScholar
2022

Generative Time Series Forecasting with Diffusion, Denoise, and Disentanglement

NeurIPS 2022accept

Time series forecasting has been a widely explored task of great importance in many applications. However, it is common that real-world time series data are recorded in a short time period, which results in a big gap between the deep model and the limited and noisy time series. In this work, we prop…

2022

Noise Regularizes Over-parameterized Rank One Matrix Recovery, Provably

AISTATS 2022poster

We investigate the role of noise in optimization algorithms for learning over-parameterized models. Specifically, we consider the recovery of a rank one matrix $Y^*\in R^{d\times d}$ from a noisy observation $Y$ using an over-parameterization model. Specifically, we parameterize the rank one matrix…

Cited by 0SourcePDFScholar
2022

Robust Single Image Dehazing Based on Consistent and Contrast-Assisted Reconstruction

IJCAI 2022poster

Single image dehazing as a fundamental low-level vision task, is essential for the development of robust intelligent surveillance system. In this paper, we make an early effort to consider dehazing robustness under variational haze density, which is a realistic while under-studied problem in the res…

Cited by 7SourcePDFScholar
2021

Noisy Gradient Descent Converges to Flat Minima for Nonconvex Matrix Factorization

AISTATS 2021poster

Numerous empirical evidences have corroborated the importance of noise in nonconvex optimization problems. The theory behind such empirical observations, however, is still largely unknown. This paper studies this fundamental problem through investigating the nonconvex rectangular matrix factorizatio…

Cited by 14SourcePDFScholar
2021

Pessimism Meets Invariance: Provably Efficient Offline Mean-Field Multi-Agent RL

NeurIPS 2021poster

Mean-Field Multi-Agent Reinforcement Learning (MF-MARL) is attractive in the applications involving a large population of homogeneous agents, as it exploits the permutation invariance of agents and avoids the curse of many agents. Most existing results only focus on online settings, in which agents…

2021

Query-Memory Re-Aggregation for Weakly-supervised Video Object Segmentation

AAAI 2021technical

Weakly-supervised video object segmentation (WVOS) is an emerging video task that can track and segment the target given a simple bounding box label. However, existing WVOS methods are still unsatisfied in either speed or accuracy, since they only use the exemplar frame to guide the prediction while…

Cited by 26SourcePDFScholar
2020

Bilinear Graph Neural Network with Neighbor Interactions

IJCAI 2020poster

Graph Neural Network (GNN) is a powerful model to learn representations and make predictions on graph data. Existing efforts on GNN have largely defined the graph convolution as a weighted sum of the features of the connected nodes to form the representation of the target node. Nevertheless, the ope…

2020

Deep Reinforcement Learning with Robust and Smooth Policy

ICML 2020poster

Deep reinforcement learning (RL) has achieved great empirical successes in various domains. However, the large search space of neural networks requires a large amount of data, which makes the current RL algorithms not sample efficient. Motivated by the fact that many environments with continuous sta…

Cited by 102SourcePDFScholar
2020

Manet: Multi-Scale Aggregated Network For Light Field Depth Estimation

ICASSP 2020accepted

We present a novel end-to-end network, MANet, for light field depth estimation. MANet is a parameter-effective and effi-cient multi-scale aggregated network, which is about 3 times smaller and 3 times faster than the current top-performing method Epinet. The MANet architecture is performed for estim…

Cited by 0SourceScholar
2020

Multi-Modality Cross Attention Network for Image and Sentence Matching

CVPR 2020poster

The key of image and sentence matching is to accurately measure the visual-semantic similarity between an image and a sentence. However, most existing methods make use of only the intra-modality relationship within each modality or the inter-modality relationship between image regions and sentence w…

Cited by 470PDFScholar
2020

Ordinal Learning for Emotion Recognition in Customer Service Calls

ICASSP 2020accepted

Approaches toward ordinal speech emotion recognition (SER) tasks are commonly based on the categorical classification algorithms, where the rank-order emotions are arbitrarily treated as independent categories. To employ the ordinal information between emotional ranks, we propose to model the ordina…

Cited by 0SourceScholar
2020

TEA: Temporal Excitation and Aggregation for Action Recognition

CVPR 2020poster

Temporal modeling is key for action recognition in videos. It normally considers both short-range motions and long-range aggregations. In this paper, we propose a Temporal Excitation and Aggregation (TEA) block, including a motion excitation (ME) module and a multiple temporal aggregation (MTA) modu…

Cited by 638PDFScholar
2019

Automatic Singing Evaluation without Reference Melody Using Bi-dense Neural Network

ICASSP 2019accepted

Automatic singing evaluation without reference melody has long been a difficult problem. This paper aims to pilot a novel data driven approach to tackle this artistic problem. We constructed a large scale dataset and designed an innovative Bi-Dense neural network which can address this task efficien…

Cited by 0SourceScholar
2019

The Speechtransformer for Large-scale Mandarin Chinese Speech Recognition

ICASSP 2019accepted

Attention-based sequence-to-sequence architectures have made great progress in the speech recognition task. The SpeechTransformer, a no-recurrence encoder-decoder architecture, has shown promising results on small-scale speech recognition data sets in previous works. In this paper, we focus on a lar…

Cited by 0SourceScholar
2019

Toward Understanding the Importance of Noise in Training Neural Networks

ICML 2019oral

Numerous empirical evidence has corroborated that the noise plays a crucial rule in effective and efficient training of deep neural networks. The theory behind, however, is still largely unknown. This paper studies this fundamental problem through training a simple two-layer convolutional neural net…

Cited by 106SourcePDFScholar
2019

Transductive Zero-Shot Learning with Visual Structure Constraint

NeurIPS 2019poster

To recognize objects of the unseen classes, most existing Zero-Shot Learning (ZSL) methods first learn a compatible projection function between the common semantic space and the visual space based on the data of source seen classes, then directly apply it to the target unseen classes. However, in re…

2018

Discriminative Learning of Latent Features for Zero-Shot Recognition

CVPR 2018poster

Zero-shot learning (ZSL) aims to recognize unseen image categories by learning an embedding space between image and semantic representations. For years, among existing works, it has been the center task to learn the proper mapping matrices aligning the visual and semantic space, whilst the importanc…

Cited by 196SourcePDFScholar
2015

Face Video Retrieval With Image Query via Hashing Across Euclidean Space and Riemannian Manifold

CVPR 2015poster

Retrieving videos of a specific person given his/her face image as query becomes more and more appealing for applications like smart movie fast-forwards and suspect searching. It also forms an interesting but challenging computer vision task, as the visual data to match, i.e., still image and video…

Cited by 82SourcePDFScholar
2015

Maintaining constant towing tension between cable ship and burying system under sea waves by hybrid FUZZY P + ID controller

IROS 2015poster

In this paper, we propose a hybrid FUZZY P + ID controller to stabilize the towing cable tension between a cable ship and a burying system. First, we develop the model of a winch system driven by valve-controlled hydraulic motors and evaluate the step responses yielded by the conventional PID and th…

Cited by 5SourceScholar
2015

Two Birds, One Stone: Jointly Learning Binary Code for Large-Scale Face Image Retrieval and Attributes Prediction

ICCV 2015poster

We address the challenging large-scale content-based face image retrieval problem, intended as searching images based on the presence of specific subject, given one face image of him/her. To this end, one natural demand is a supervised binary code learning method. While the learned codes might be di…

Cited by 65PDFScholar