← Search

Yiming Wang

102 accepted papers

2026

Asymmetric Multi-View Clustering with Hyperbolic Uncertainty Modeling

ICML 2026spotlight

Deep Multi-View Clustering (MVC) aims to extract a unified semantic consensus from diverse data sources without supervision. However, current approaches relying on flat Euclidean embeddings often fail to model data uncertainty, resulting in rigid alignment where high-quality views are forced to drif…

Cited by 0SourceScholar
2026

BulletTime: Decoupled Control of Time and Camera Pose for Video Generation

CVPR 2026

Emerging video diffusion models achieve high visual fidelity but fundamentally couple scene dynamics with camera motion, limiting their ability to provide precise spatial and temporal control. We introduce a 4D-controllable video diffusion framework that explicitly decouples scene dynamics from came

Cited by 0SourceScholar
2026

Covariance Volume Maximization for Embodied Latent Exploration in Deep Reinforcement Learning

ICML 2026poster

Efficient exploration remains a key challenge in deep reinforcement learning, especially for embodied agents operating in realistic environments with high-dimensional observations and complex dynamics. Recent latent exploration methods define bonuses in a learned latent space, but often struggle in …

Cited by 0SourceScholar
2026

DSAP: Enhancing Generalization in Goal-Conditioned Reinforcement Learning

AAAI 2026technical

Goal-conditioned Reinforcement Learning (RL) is a promising direction for training agents capable of tackling a variety of tasks. However, generalizing to new goals in different environments remains a central challenge for goal-conditioned RL agents. Existing methods often rely on state abstraction,

Cited by 0SourcePDFScholar
2026

Direct Simultaneous Translation Activation for Large Audio-Language Models

ICASSP 2026poster

Simultaneous speech-to-text translation (Simul-S2TT) aims to translate speech into target text in real time, outputting translations while receiving source speech input, rather than waiting for the entire utterance to be spoken. Simul-S2TT research often modifies model architectures to implement rea…

Cited by 0SourcePDFScholar
2026

E^2DT: Efficient and Effective Decision Transformer with Experience-Aware Sampling for Robotic Manipulation

ICRA 2026poster

In reinforcement learning (RL) for robotic manipulation, the Decision Transformer (DT) has emerged as an effective framework for addressing long-horizon tasks. However, DT’s performance depends heavily on the coverage of collected experiences. Without an active exploration mechanism, standard DT rel…

Cited by 0Scholar
2026

Efficient Encoder-Free Fourier-based 3D Large Multimodal Model

CVPR 2026

Large Multimodal Models (LMMs) that process 3D data typically rely on heavy, pretrained visual encoders to extract geometric features. While recent 2D LMMs have begun to eliminate such encoders for efficiency and scalability, extending this paradigm to 3D remains challenging due to the unordered and

Cited by 0SourceScholar
2026

Explore to Learn: Latent Exploration Through Disentangled Synergy Patterns for Reinforcement Learning in Overactuated Control

AAAI 2026technical

Control in high-dimensional action spaces remains a fundamental challenge in reinforcement learning (RL), primarily due to inefficient exploration of the action space. While recent methods attempt to guide exploration, they often fall short of achieving the agility and coordination exhibited in biol

Cited by 0SourcePDFScholar
2026

Latent State-Predictive Exploration for Deep Reinforcement Learning

AAAI 2026technical

Reinforcement learning (RL) has achieved promising results in continuous control tasks, where efficient exploration of the state space is crucial for success. However, many recent RL approaches still struggle with sample inefficiency and insufficient exploration for long-horizon tasks, particularly

Cited by 0SourcePDFScholar
2026

Obstruction Reasoning for Robotic Grasping

CVPR 2026

Successful robotic grasping in cluttered environments not only requires a model to visually ground a target object but also to reason about obstructions that must be cleared beforehand. While current vision-language embodied reasoning models show emergent spatial understanding, they remain limited i

Cited by 0SourceScholar
2026

Order within Chaos: Capturing Intrinsic Energy Anomalies for AI-Manipulated Image Forgery Localization

ICML 2026poster

Recent advancements in generative AI have led to image editing models capable of producing realistic forgeries that evade traditional image forgery localization methods, as these approaches depend on physical noise absent in synthetic data. To address this challenge, we theoretically demonstrate tha…

Cited by 0SourceScholar
2026

Perturbation Matters in Time Series Forecasting: A Wave-attention-aware Transformer Method

IJCAI 2026

Time series forecasting (TSF) refers to a fundamental task of predicting future sequential data based on historical observations. One representative category of TSF methods is transformer-based approaches, which translate time series into token sequences (i.e., as raw texts) before applying well-est

Cited by 0Scholar
2026

RGMP: Recurrent Geometric-prior Multimodal Policy for Generalizable Humanoid Robot Manipulation

AAAI 2026technical

Humanoid robots exhibit significant potential in executing diverse human-level skills. However, current research predominantly relies on data-driven approaches that necessitate extensive training datasets to achieve robust multimodal decision-making capabilities and generalizable visuomotor control.

Cited by 0SourcePDFScholar
2026

Specificity-aware reinforcement learning for fine-grained open-world classification

CVPR 2026

Classifying fine-grained visual concepts under open-world settings, i.e., without a predefined label set, demands models to be both accurate and specific. Recent reasoning Large Multimodal Models (LMMs) exhibit strong visual understanding capability but tend to produce overly generic predictions whe

Cited by 0SourcecodeScholar
2026

Spherical Geometry Diffusion: Generating High-quality 3D Face Geometry via Sphere-anchored Representations

AAAI 2026technical

A fundamental challenge in text-to-3D face generation is achieving high-quality geometry. The core difficulty lies in the arbitrary and intricate distribution of vertices in 3D space, making it challenging for existing models to establish clean connectivity and resulting in suboptimal geometry. To a

Cited by 0SourcePDFScholar
2026

StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues

CVPR 2026

Edge-based representations are fundamental cues for visual understanding, a principle rooted in early vision research and still central today. We extend this principle to vision-language alignment, showing that isolating and aligning structural cues across modalities can greatly benefit fine-tuning

Cited by 0SourcecodeScholar
2026

VKG-QA: Visual Knowledge Graph-based Question Answer for Large Multimodal Models

CVPR 2026

Understanding and reasoning over structured knowledge is a fundamental capability for intelligent systems. While Large Language Models (LLMs) have leveraged textual knowledge graphs for relational reasoning, linearizing graph structures into text often leads to token inefficiency and loss of higher-

Cited by 0SourcecodeScholar
2025

AirExo-2: Scaling up Generalizable Robotic Imitation Learning with Low-Cost Exoskeletons

CoRL 2025oral

Scaling up robotic imitation learning for real-world applications requires efficient and scalable demonstration collection methods. While teleoperation is effective, it depends on costly and inflexible robot platforms. In-the-wild demonstrations offer a promising alternative, but existing collection…

Cited by 0SourceScholar
2025

BILE: An Effective Behavior-based Latent Exploration Scheme for Deep Reinforcement Learning

IJCAI 2025

Efficient exploration of state spaces is critical for the success of deep reinforcement learning (RL). While many methods leverage exploration bonuses to encourage exploration instead of relying solely on extrinsic rewards, these bonus-based approaches often face challenges with learning efficiency

Cited by 0SourcePDFScholar
2025

Can Text-to-Video Generation help Video-Language Alignment?

CVPR 2025poster

Recent video-language alignment models are trained on sets of videos, each with an associated positive caption and a negative caption generated by large language models. A problem with this procedure is that negative captions may introduce linguistic biases, i.e., concepts are seen only as negatives…

Cited by 0SourcePDFScholar
2025

Collaborative Instance Object Navigation: Leveraging Uncertainty-Awareness to Minimize Human-Agent Dialogues

ICCV 2025poster

Language-driven instance object navigation assumes that a human initiates the task by providing a detailed description of the target to the embodied agent. While this description is crucial for distinguishing the target from other visually similar instances, providing it prior to navigation can be d…

Cited by 0SourcePDFScholar
2025

ConViS-Bench: Estimating Video Similarity Through Semantic Concepts

NeurIPS 2025poster

What does it mean for two videos to be similar? Videos may appear similar when judged by the actions they depict, yet entirely different if evaluated based on the locations where they were filmed. While humans naturally compare videos by taking different aspects into account, this ability has not be…

Cited by 0SourceScholar
2025

Do Large Language Models Truly Understand Geometric Structures?

ICLR 2025poster

Geometric ability is a significant challenge for large language models (LLMs) due to the need for advanced spatial comprehension and abstract thinking. Existing datasets primarily evaluate LLMs on their final answers, but they cannot truly measure their true understanding of geometric structures, as…

2025

Efficient Diversity-based Experience Replay for Deep Reinforcement Learning

IJCAI 2025

Experience replay is widely used to improve learning efficiency in reinforcement learning by leveraging past experiences. However, existing experience replay methods, whether based on uniform or prioritized sampling, often suffer from low efficiency, particularly in real-world scenarios with high-di

Cited by 0SourcePDFScholar
2025

ForceMimic: Force-Centric Imitation Learning with Force-Motion Capture System for Contact-Rich Manipulation

ICRA 2025

In most contact-rich manipulation tasks, humans apply time-varying forces to the target object, compensating for inaccuracies in the vision-guided hand trajectory. However, current robot learning algorithms primarily focus on trajectory-based policy, with limited attention given to learning force-re

Cited by 66SourcecodeScholar
2025

Free-form language-based robotic reasoning and grasping

IROS 2025

Performing robotic grasping from a cluttered bin based on human instructions is a challenging task, as it requires understanding both the nuances of free-form language and the spatial relationships between objects. Vision-Language Models (VLMs) trained on web-scale data, such as GPT-4o, have demonst

Cited by 8SourcecodeScholar
2025

GL-GAN: Perceiving and Integrating Global and Local Styles for Handwritten Text Generation with Mamba

COLING 2025main

Handwritten text generation (HTG) aims to synthesize handwritten samples by imitating a specific writer, which has a wide range of applications and thus has significant research value. However, current studies on HTG are confronted with a main bottleneck: dominant models lack the ability to perceive…

2025

HeMoRa: Unsupervised Heuristic Consensus Sampling for Robust Point Cloud Registration

CVPR 2025poster

Heuristic information for consensus set sampling is essential for correspondence-based point cloud registration, but existing approaches typically rely on supervised learning or expert-driven parameter tuning. In this work, we propose HeMoRa, a new unsupervised framework that trains a Heuristic info…

2025

LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text Pairing

ICCV 2025poster

Fashion design is a complex creative process that blends visual and textual expressions. Designers convey ideas through sketches, which define spatial structure and design elements, and textual descriptions, capturing material, texture, and stylistic details. In this paper, we present LOcalized Text…

Cited by 0SourcePDFScholar
2025

Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation

ICLR 2025poster

LLM self-evaluation relies on the LLM's own ability to estimate response correctness, which can greatly improve its deployment reliability. In this research track, we propose the Chain-of-Embedding (CoE) in the latent space to enable LLMs to perform output-free self-evaluation. CoE consists of all…

2025

Learned Video Compression With Refined Adaptive Flow Pyramid And Coordinate-Aware Attention

ICASSP 2025accepted

In video compression, motion estimation and motion compensation are critical for achieving efficient encoding. Although the commonly used SpyNet and bilinear interpolation have contributed in improving the compression efficiency, they still have limitations. SpyNet often loses details and fails to f…

Cited by 0SourceScholar
2025

Learning Efficient Fuse-and-Refine for Feed-Forward 3D Gaussian Splatting

NeurIPS 2025poster

Recent advances in feed-forward 3D Gaussian Splatting have led to rapid improvements in efficient scene reconstruction from sparse views. However, most existing approaches construct Gaussian primitives directly aligned with the pixels in one or more of the input images. This leads to redundancies in…

Cited by 0SourceScholar
2025

Learning from Disjoint Views: A Contrastive Prototype Matching Network for Fully Incomplete Multi-View Clustering

NeurIPS 2025poster

Multi-view clustering aims to enhance clustering performance by leveraging information from diverse sources. However, its practical application is often hindered by a barrier: the lack of correspondences across views. This paper focuses on the understudied problem of fully incomplete multi-view clus…

Cited by 0SourceScholar
2025

MIND: Towards Immersive Psychological Healing with Multi-Agent Inner Dialogue

EMNLP 2025

Mental health issues are worsening in today’s competitive society, such as depression and anxiety. Traditional healings like counseling and chatbots fail to engage effectively, they often provide generic responses lacking emotional depth. Although large language models (LLMs) have the potential to c

Cited by 0SourcePDFScholar
2025

On Large Multimodal Models as Open-World Image Classifiers

ICCV 2025poster

Traditional image classification requires a predefined list of semantic categories. In contrast, Large Multimodal Models (LMMs) can sidestep this requirement by classifying images directly using natural language (e.g., answering the prompt "What is the main object in the image?"). Despite this remar…

2025

PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts

NeurIPS 2025poster

In this paper, we introduce **PolyMath**, a multilingual mathematical reasoning benchmark covering 18 languages and 4 easy-to-hard difficulty levels. Our benchmark ensures difficulty comprehensiveness, language diversity, and high-quality translation, making it a highly discriminative multilingual m…

Cited by 0SourceScholar
2025

RetroDiff: Retrosynthesis as Multi-stage Distribution Interpolation

AISTATS 2025poster

Retrosynthesis poses a key challenge in biopharmaceuticals, aiding chemists in finding appropriate reactant molecules for given product molecules. With reactants and products represented as 2D graphs, retrosynthesis constitutes a conditional graph-to-graph (G2G) generative task. Inspired by advancem…

Cited by 0SourceScholar
2025

STAMImputer: Spatio-Temporal Attention MoE for Traffic Data Imputation

IJCAI 2025

Traffic data imputation is fundamentally important to support various applications in intelligent transportation systems such as traffic flow prediction. However, existing time-to-space sequential methods often fail to effectively extract features in block-wise missing data scenarios. Meanwhile, the

2025

Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early Decoding

NeurIPS 2025spotlight

Test-time scaling enhances large language model performance by allocating additional compute resources during decoding. Best-of-$N$ (BoN) sampling serves as a common sampling-based scaling technique, broadening the search space in parallel to find better solutions from the model distribution. Howeve…

Cited by 0SourceScholar
2025

Seeing the Abstract: Translating the Abstract Language for Vision Language Models

CVPR 2025poster

Natural language goes beyond dryly describing visual content. It contains rich abstract concepts to express feeling, creativity and properties that cannot be directly perceived. Yet, current research in Vision Language Models (VLMs) has not shed light on abstract-oriented language.Our research break…

2025

SmartExp: An Adaptive Data Expansion Strategy for Improving Handwritten Text Recognition

ICASSP 2025accepted

Constructing a highly accurate handwritten OCR system requires large amounts of high-quality training data, yet data collection is labor-intensive and costly. With the advance of generative models, high-quality synthetic images have been applied to enhance handwritten text recognition (HTR) models,…

Cited by 0SourceScholar
2025

SplatFormer: Point Transformer for Robust 3D Gaussian Splatting

ICLR 2025spotlight

3D Gaussian Splatting (3DGS) has recently transformed photorealistic reconstruction, achieving high visual fidelity and real-time performance. However, rendering quality significantly deteriorates when test views deviate from the camera angles used during training, posing a major challenge for appli…

2025

The USTC System for EEG-Music Emotion Recognition Challenge

ICASSP 2025accepted

This paper presents the Neural Harmony team’s submission to Task 1 (Person Identification) of the ICASSP 2025 EEG-Music Emotion Recognition Challenge, which aims to identify the subject from a given EEG segment. To enhance performance, we propose a novel architecture incorporating the Multiscale Con…

Cited by 0SourceScholar
2025

Training-Free Personalization via Retrieval and Reasoning on Fingerprints

ICCV 2025poster

Vision Language Models (VLMs) have lead to major improvements in multimodal reasoning, yet they still struggle to understand user-specific concepts. Existing personalization methods address this limitation butheavily rely on training procedures, that can be either costly or unpleasant to individual…

Cited by 0SourcePDFScholar
2025

Training-free Online Video Step Grounding

NeurIPS 2025poster

Given a task and a set of steps composing it, Video Step Grounding (VSG) aims to detect which steps are performed in a video. Standard approaches for this task require a labeled training set (e.g., with step-level annotations or narrations), which may be costly to collect. Moreover, they process the…

Cited by 0SourceScholar
2025

Transformer-based Speech Model Learns Well as Infants and Encodes Abstractions through Exemplars in the Poverty of the Stimulus Environment

COLING 2025main

Infants are capable of learning language, predominantly through speech and associations, in impoverished environments—a phenomenon known as the Poverty of the Stimulus (POS). Is this ability uniquely human, as an innate linguistic predisposition, or can it be empirically learned through potential li…

2025

Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of Translationese

EMNLP 2025

Translationese refers to linguistic properties that usually occur in translated texts. Previous works study translationese by framing it as a binary classification between original texts and translated texts. In this paper, we argue that translationese should be graded instead of binary and propose

Cited by 0SourcePDFScholar
2024

AirExo: Low-Cost Exoskeletons for Learning Whole-Arm Manipulation in the Wild

ICRA 2024poster

While humans can use parts of their arms other than the hands for manipulations like gathering and supporting, whether robots can effectively learn and perform the same type of operations remains relatively unexplored. As these manipulations require joint-level control to regulate the complete poses…

Cited by 40SourcecodeScholar
2024

AlignSum: Data Pyramid Hierarchical Fine-tuning for Aligning with Human Summarization Preference

EMNLP 2024finding

Text summarization tasks commonly employ Pre-trained Language Models (PLMs) to fit diverse standard datasets. While these PLMs excel in automatic evaluations, they frequently underperform in human evaluations, indicating a deviation between their generated summaries and human summarization preferenc…

2024

Automated Tone Transcription and Clustering with Tone2Vec

EMNLP 2024finding

Lexical tones play a crucial role in Sino-Tibetan languages. However, current phonetic fieldwork relies on manual effort, resulting in substantial time and financial costs. This is especially challenging for the numerous endangered languages that are rapidly disappearing, often compounded by limited…

2024

Boosting Single Positive Multi-label Classification with Generalized Robust Loss

IJCAI 2024poster

Multi-label learning (MLL) requires comprehensive multi-semantic annotations that is hard to fully obtain, thus often resulting in missing labels scenarios. In this paper, we investigate Single Positive Multi-label Learning (SPML), where each image is associated with merely one positive label. Exist…

2024

CSLM: A Framework for Question Answering Dataset Generation through Collaborative Small Language Models

EMNLP 2024finding

Collecting high-quality question-answer (QA) pairs is vital for the training of large language models (LLMs), yet this process is traditionally laborious and time-intensive. With the rapid evolution of LLMs, the potential for leveraging these models to autonomously generate QA pairs has become appar…

Cited by 0SourcePDFScholar
2024

DMMR: Cross-Subject Domain Generalization for EEG-Based Emotion Recognition via Denoising Mixed Mutual Reconstruction

AAAI 2024technical

Electroencephalography (EEG) has proven to be effective in emotion analysis. However, current methods struggle with individual variations, complicating the generalization of models trained on data from source subjects to unseen target subjects. To tackle this issue, we propose the Denoising Mixed Mu…

2024

Embedding Trajectory for Out-of-Distribution Detection in Mathematical Reasoning

NeurIPS 2024poster

Real-world data deviating from the independent and identically distributed (\textit{i.i.d.}) assumption of in-distribution training data poses security threats to deep networks, thus advancing out-of-distribution (OOD) detection algorithms. Detection methods in generative language models (GLMs) main…

2024

Geometrically-driven Aggregation for Zero-shot 3D Point Cloud Understanding

CVPR 2024highlight

Zero-shot 3D point cloud understanding can be achieved via 2D Vision-Language Models (VLMs). Existing strategies directly map VLM representations from 2D pixels of rendered or captured views to 3D points overlooking the inherent and expressible point cloud geometric structure. Geometrically similar…

2024

Harnessing Large Language Models for Training-free Video Anomaly Detection

CVPR 2024poster

Video anomaly detection (VAD) aims to temporally locate abnormal events in a video. Existing works mostly rely on training deep models to learn the distribution of normality with either video-level supervision one-class supervision or in an unsupervised setting. Training-based methods are prone to b…

Cited by 40SourcePDFScholar
2024

Learned Video Compression with Spatial-Temporal Optimization

ICASSP 2024accepted

Previous optical flow based video compression is gradually replaced by unsupervised deformable convolution (DCN) based method. This is mainly due to the fact that the motion vector (MV) estimated by the existing optical flow network is not accurate and may introduce extra artifacts. However, DCN bas…

Cited by 0SourceScholar
2024

Loose Inertial Poser: Motion Capture with IMU-attached Loose-Wear Jacket

CVPR 2024poster

Existing wearable motion capture methods typically demand tight on-body fixation (often using straps) for reliable sensing limiting their application in everyday life. In this paper we introduce Loose Inertial Poser a novel motion capture solution with high wearing comfortableness by integrating fou…

2024

Meta-Reasoning: Semantics-Symbol Deconstruction for Large Language Models

ACL 2024findings

Neural-symbolic methods have demonstrated efficiency in enhancing the reasoning abilities of large language models (LLMs). However, existing methods mainly rely on syntactically mapping natural languages to complete formal languages like Python and SQL. Those methods require that reasoning tasks be…

2024

Mind the Error! Detection and Localization of Instruction Errors in Vision-and-Language Navigation

IROS 2024

Vision-and-Language Navigation in Continuous Environments (VLN-CE) is one of the most intuitive yet challenging embodied AI tasks. Agents are tasked to navigate towards a target goal by executing a set of low-level actions, following a series of natural language instructions. All VLN-CE methods in t

Cited by 13SourceScholar
2024

R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

EMNLP 2024finding

Large language models (LLMs) have exhibited great potential in autonomously completing tasks across real-world applications. Despite this, these LLM agents introduce unexpected safety risks when operating in interactive environments. Instead of centering on the harmlessness of LLM-generated content…

2024

Rethinking Exploration in Reinforcement Learning with Effective Metric-Based Exploration Bonus

NeurIPS 2024spotlight

Enhancing exploration in reinforcement learning (RL) through the incorporation of intrinsic rewards, specifically by leveraging *state discrepancy* measures within various metric spaces as exploration bonuses, has emerged as a prevalent strategy to encourage agents to visit novel states. The critica…

Cited by 0SourcePDFScholar
2024

Retrieval-enriched zero-shot image classification in low-resource domains

EMNLP 2024main

Low-resource domains, characterized by scarce data and annotations, present significant challenges for language and visual understanding tasks, with the latter much under-explored in the literature. Recent advancements in Vision-Language Models (VLM) have shown promising results in high-resource dom…

2024

Test-Time Zero-Shot Temporal Action Localization

CVPR 2024poster

Zero-Shot Temporal Action Localization (ZS-TAL) seeks to identify and locate actions in untrimmed videos unseen during training. Existing ZS-TAL methods involve fine-tuning a model on a large amount of annotated training data. While effective training-based ZS-TAL approaches assume the availability…

2023

3DSGrasp: 3D Shape-Completion for Robotic Grasp

ICRA 2023poster

Real-world robotic grasping can be done robustly if a complete 3D Point Cloud Data (PCD) of an object is available. However, in practice, PCDs are often incomplete when objects are viewed from few and sparse viewpoints before the grasping action, leading to the generation of wrong or inaccurate gras…

Cited by 26SourcecodeScholar
2023

DATA2VEC-SG: Improving Self-Supervised Learning Representations for Speech Generation Tasks

ICASSP 2023accepted

Self-supervised learning has been successfully applied to various speech recognition and understanding tasks. However, for generative tasks such as speech enhancement and speech separation, most self-supervised speech representations did not show substantial improvements. To deal with this problem,…

Cited by 0SourceScholar
2023

DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation

ICML 2023poster

We introduce DS-1000, a code generation benchmark with a thousand data science problems spanning seven Python libraries, such as Numpy and Pandas. Compared to prior works, DS-1000 incorporates three core features. First, our problems reflect diverse, realistic, and practical use cases since we colle…

2023

Efficient Potential-based Exploration in Reinforcement Learning using Inverse Dynamic Bisimulation Metric

NeurIPS 2023poster

Reward shaping is an effective technique for integrating domain knowledge into reinforcement learning (RL). However, traditional approaches like potential-based reward shaping totally rely on manually designing shaping reward functions, which significantly restricts exploration efficiency and introd…

Cited by 9SourcePDFScholar
2023

Element-aware Summarization with Large Language Models: Expert-aligned Evaluation and Chain-of-Thought Method

ACL 2023long

Automatic summarization generates concise summaries that contain key ideas of source documents. As the most mainstream datasets for the news sub-domain, CNN/DailyMail and BBC XSum have been widely used for performance benchmarking. However, the reference summaries of those datasets turn out to be no…

2023

Label-Guided Knowledge Distillation for Continual Semantic Segmentation on 2D Images and 3D Point Clouds

ICCV 2023poster

Continual semantic segmentation (CSS) aims to extend an existing model to tackle unseen tasks while retaining its old knowledge. Naively fine-tuning the old model on new data leads to catastrophic forgetting. A common solution is knowledge distillation (KD), where the output distribution of the new…

Cited by 17PDFcodeScholar
2023

NeuS2: Fast Learning of Neural Implicit Surfaces for Multi-view Reconstruction

ICCV 2023poster

Recent methods for neural surface representation and rendering, for example NeuS, have demonstrated the remarkably high-quality reconstruction of static scenes. However, the training of NeuS takes an extremely long time (8 hours), which makes it almost impossible to apply them to dynamic scenes with…

Cited by 276PDFcodeScholar
2023

PI-Trans: Parallel-Convmlp and Implicit-Transformation Based Gan for Cross-View Image Translation

ICASSP 2023accepted

For semantic-guided cross-view image translation, it is crucial to learn where to sample pixels from the source view image and where to reallocate them guided by the target view semantic map, especially when there is little overlap or drastic view difference between the source and target images. Hen…

Cited by 0SourceScholar
2023

Query Your Model with Definitions in FrameNet: An Effective Method for Frame Semantic Role Labeling

AAAI 2023technical

Frame Semantic Role Labeling (FSRL) identifies arguments and labels them with frame semantic roles defined in FrameNet. Previous researches tend to divide FSRL into argument identification and role classification. Such methods usually model role classification as naive multi-class classification and…

2023

Self-Supervised Learning with Bi-Label Masked Speech Prediction for Streaming Multi-Talker Speech Recognition

ICASSP 2023accepted

Self-supervised learning (SSL), which utilizes the input data itself for representation learning, has achieved state-of-the-art results for various downstream speech tasks. However, most of the previous studies focused on offline single-talker applications, with limited investigations in multi-talke…

Cited by 0SourceScholar
2023

Vocabulary-free Image Classification

NeurIPS 2023poster

Recent advances in large vision-language models have revolutionized the image classification paradigm. Despite showing impressive zero-shot capabilities, a pre-defined set of categories, a.k.a. the vocabulary, is assumed at test time for composing the textual prompts. However, such assumption can be…

2023

When Visual Prompt Tuning Meets Source-Free Domain Adaptive Semantic Segmentation

NeurIPS 2023poster

Source-free domain adaptive semantic segmentation aims to adapt a pre-trained source model to the unlabeled target domain without accessing the private source data. Previous methods usually fine-tune the entire network, which suffers from expensive parameter tuning. To avoid this problem, we propos…

2022

Improving Noise Robustness of Contrastive Speech Representation Learning with Speech Reconstruction

ICASSP 2022accepted

Noise robustness is essential for deploying automatic speech recognition (ASR) systems in real-world environments. One way to reduce the effect of noise interference is to employ a preprocessing module that conducts speech enhancement, and then feed the enhanced speech to an ASR backend. In this wor…

Cited by 0SourceScholar
2022

Noise-injected Consistency Training and Entropy-constrained Pseudo Labeling for Semi-supervised Extractive Summarization

COLING 2022main

Labeling large amounts of extractive summarization data is often prohibitive expensive due to time, financial, and expertise constraints, which poses great challenges to incorporating summarization system in practical applications. This limitation can be overcome by semi-supervised approaches: consi…

2022

Spatial Commonsense Graph for Object Localisation in Partial Scenes

CVPR 2022poster

We solve object localisation in partial scenes, a new problem of estimating the unknown position of an object (e.g. where is the bag?) given a partial 3D scan of a scene. The proposed solution is based on a novel scene graph model, the Spatial Commonsense Graph (SCG), where objects are the nodes and…

Cited by 23PDFcodeScholar
2022

Wav2vec-Switch: Contrastive Learning from Original-Noisy Speech Pairs for Robust Speech Recognition

ICASSP 2022accepted

The goal of self-supervised learning (SSL) for automatic speech recognition (ASR) is to learn good speech representations from a large amount of unlabeled speech for the downstream ASR task. However, most SSL frameworks do not consider noise robustness which is crucial for real-world applications. I…

Cited by 66SourceScholar
2022

Weakly-supervised Text Classification with Wasserstein Barycenters Regularization

IJCAI 2022poster

Weakly-supervised text classification aims to train predictive models with unlabeled texts and a few representative words of classes, referred to as category words, rather than labeled texts. These weak supervisions are much more cheaper and easy to collect in real-world scenarios. To resolve this t…

2021

Double Low-Rank Representation With Projection Distance Penalty for Clustering

CVPR 2021poster

This paper presents a novel, simple yet robust self-representation method, i.e., Double Low-Rank Representation with Projection Distance penalty (DLRRPD) for clustering. With the learned optimal projected representations, DLRRPD is capable of obtaining an effective similarity graph to capture the mu…

Cited by 35PDFScholar
2021

Extracting Topics with Simultaneous Word Co-occurrence and Semantic Correlation Graphs: Neural Topic Modeling for Short Texts

EMNLP 2021finding

Short text nowadays has become a more fashionable form of text data, e.g., Twitter posts, news titles, and product reviews. Extracting semantic topics from short texts plays a significant role in a wide spectrum of NLP applications, and neural topic modeling is now a major tool to achieve it. Motiva…

2021

GraphMSE: Efficient Meta-path Selection in Semantically Aligned Feature Space for Graph Neural Networks

AAAI 2021technical

Heterogeneous information networks (HINs) are ideal for describing real-world data with different types of entities and relationships. To carry out machine learning on HINs, meta-paths are widely utilized to extract semantics with pre-defined patterns, and models such as graph convolutional networks…

2021

POMP++: Pomcp-based Active Visual Search in unknown indoor environments

IROS 2021poster

In this paper, we focus on the problem of learning online an optimal policy for Active Visual Search (AVS) of objects in unknown indoor environments. We propose POMP++, a planning strategy that introduces a novel formulation on top of the classic Partially Observable Monte Carlo Planning (POMCP) fra…

Cited by 17SourceScholar
2020

Where to Explore Next? ExHistCNN for History-aware Autonomous 3D Exploration

ECCV 2020poster

In this work we address the problem of autonomous 3D exploration of an unknown indoor environment using a depth camera. We cast the problem as the estimation of the Next Best View (NBV) that maximises the coverage of the unknown area. We do this by re-formulating NBV estimation as a classification p…

2019

Autonomous 3-D Reconstruction, Mapping, and Exploration of Indoor Environments With a Robotic Arm

RA-L 2019

We propose a novel information gain metric that combines hand-crafted and data-driven metrics to address the next best view problem for autonomous 3-D mapping of unknown indoor environments. For the hand-crafted metric, we propose an entropy-based information gain that accounts for the previous view

Cited by 48SourceScholar
2019

End-to-end Anchored Speech Recognition

ICASSP 2019accepted

Voice-controlled house-hold devices, like Amazon Echo or Google Home, face the problem of performing speech recognition of device-directed speech in the presence of interfering background speech, i.e., background noise and interfering speech from another person or media device in proximity need to b…

Cited by 20SourceScholar
2018

A Pruned Rnnlm Lattice-Rescoring Algorithm for Automatic Speech Recognition

ICASSP 2018accepted

Lattice-rescoring is a common approach to take advantage of recurrent neural language models in ASR, where a word-lattice is generated from 1st-pass decoding and the lattice is then rescored with a neural model, and an <i xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/…

Cited by 0SourceScholar
2018

Neural Network Language Modeling with Letter-Based Features and Importance Sampling

ICASSP 2018accepted

In this paper we describe an extension of the Kaldi software toolkit to support neural-based language modeling, intended for use in automatic speech recognition (ASR) and related tasks. We combine the use of subword features (letter n-grams) and one-hot encoding of frequent words so that the models…

Cited by 0SourceScholar