← Search

Bo Yang

108 accepted papers

2026

A Dual-Mode Hydraulic Actuator for a Quasi-Passive Load-Carrying Exoskeleton in Multiple Conditions

RA-L 2026

Lower limb exoskeleton robots have been widely researched for load-carrying assistance. Recently, quasi-passive exoskeletons using low-power elements to modulate mechanical characteristics have emerged. However, achieving effective damping and stiffness across varying tasks and loads remains challen

Cited by 0SourceScholar
2026

AVGGT: Rethinking Global Attention for Accelerating VGGT

CVPR 2026

Models such as VGGT and \pi^3 have shown strong multi-view 3D performance, but their heavy reliance on global self-attention results in high computational cost. Existing sparse-attention variants offer partial speedups, yet lack a systematic analysis of how global attention contributes to multi-view

Cited by 0SourceScholar
2026

Beyond Single Embedding: Modeling User Preferences as Distribution in Federated Recommendation

ICML 2026poster

Most federated recommender systems represent each user with a single embedding learned from local interaction data, implicitly assuming that user preferences are fixed and precisely identifiable. In federated settings, however, each client observes only a limited and fragmentary view of user behavio…

Cited by 0SourceScholar
2026

Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling

ICML 2026poster

Preference optimization for diffusion models relies on reward functions that are both discriminative and computationally efficient. Vision-Language Models (VLMs) have emerged as powerful reward providers. However, their computation and memory cost can be substantial, and optimizing a latent diffusio…

Cited by 0SourceScholar
2026

CCFQA: A Benchmark for Cross-Lingual and Cross-Modal Speech and Text Factuality Evaluation

AAAI 2026technical

As Large Language Models (LLMs) are increasingly popularized in the multilingual world, ensuring hallucination-free factuality becomes markedly crucial. However, existing benchmarks for evaluating the reliability of Multimodal Large Language Models (MLLMs) predominantly focus on textual or visual m

Cited by 0SourcePDFScholar
2026

CT2BSE: 3D BSE Microstructural Image Cross-Device Generation from µCT for Cement Hydration via Voxel Swin Transformer

IJCAI 2026

Acquiring three-dimensional(3D) microstructural images of cement hydration reveals critical microscale features essential for understanding hydration mechanisms and advancing material development. Micro-computed tomography(µCT) is widely used to capture these images due to its non-destructive, repea

Cited by 0Scholar
2026

DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers

CVPR 2026

Diffusion models have recently motivated great success in many generation tasks like object removal. Nevertheless, existing image decomposition methods struggle to disentangle semi-transparent or transparent layer occlusions due to mask prior dependencies, static object assumptions, and the lack of

Cited by 0SourcecodeScholar
2026

DualPrim: Compact 3D Reconstruction with Positive and Negative Primitives

CVPR 2026

We present Compact 3D Reconstruction with Positive and Negative Primitives (DualPrim), a novel approach for reconstructing compact and topologically regular 3D meshes from multi-view images. Unlike traditional methods that rely on implicit representations such as signed distance functions, or explic

Cited by 0SourceScholar
2026

DynFusion: Rethinking Condition Fusion for Adaptive Multi-Conditional Text-to-Image Generation

CVPR 2026

Text-to-image diffusion models have achieved remarkable progress, generating visually realistic and semantically coherent images from textual prompts. However, natural language alone lacks the precision required for design-centric applications that demand strict spatial and structural fidelity--part

Cited by 0SourceScholar
2026

EvObj: Learning Evolving Object-centric Representations for 3D Instance Segmentation without Scene Supervision

CVPR 2026

We introduce EvObj for unsupervised 3D instance segmentation that bridges the geometric domain gap between synthetic pretraining data and real-world point clouds. Current methods suffer from structural discrepancies when transferring object priors from synthetic datasets (e.g., ShapeNet) to real sca

Cited by 0SourcecodeScholar
2026

FairJudge : An Adaptive, Debiased, and Consistent LLM-as-a-Judge

ICML 2026poster

Existing LLM-as-a-Judge systems suffer from three fundamental limitations: \textbf{limited adaptivity} to task and domain-specific evaluation criteria, \textbf{systematic biases} driven by non-semantic cues such as position, length, format, and model provenance, and \textbf{evaluation inconsistency}…

Cited by 0SourceScholar
2026

FoundObj: Self-supervised Foundation Models as Rewards for Label-free 3D Object Segmentation

ICML 2026poster

We address the challenging task of 3D object segmentation in complex scene point clouds without relying on any scene-level human annotations during training. Existing methods are typically constrained to identifying simple objects, primarily due to insufficient object priors in the learning process.…

Cited by 0SourceScholar
2026

Game-KFS: Game-Theory-Inspired Keyframe Selection for Hybrid Representation Visual SLAM

ICRA 2026poster

Hybrid representation Visual Simultaneous Localization and Mapping (VSLAM) systems combine the inherent strengths of both discrete and field representations. They promise high-precision tracking and photo-realistic dense mapping. However, current keyframe selection methods in hybrid representation V…

Cited by 0SourceScholar
2026

GeoDiT: A Diffusion-based Vision-Language Model for Geospatial Understanding

CVPR 2026

Autoregressive models are structurally misaligned with the inherently parallel nature of geospatial understanding, forcing a rigid sequential narrative onto scenes and fundamentally hindering the generation of structured and coherent outputs. We challenge this paradigm by reframing geospatial genera

Cited by 0SourceScholar
2026

Innovative Design of Multi-Functional Supernumerary Robotic Limbs with Ellipsoid Workspace Optimization

ICRA 2026poster

Supernumerary robotic limbs (SRL) offer substantial potential in both the rehabilitation of hemiplegic patients and the enhancement of functional capabilities for healthy individuals. Designing a general-purpose SRL device is inherently challenging, particularly when developing a unified theoretical…

2026

LEGO-FL: Learning Heterogeneous Federated Models as a LEGO Assembly Games

ICML 2026poster

Just as LEGO pieces can be assembled into an unlimited variety of structures, heterogeneous federated learning (HFL) can be viewed as the assembly of diverse model components. Inspired by this analogy, we reformulate HFL as a LEGO-like assembly game. The central challenge in HFL lies in learning acr…

Cited by 0SourceScholar
2026

Learning to Drift with Individual Wheel Drive: Maneuvering Autonomous Vehicle at the Handling Limits

ICRA 2026poster

Drifting, characterized by controlled vehicle motion at high sideslip angles, is crucial for safely handling emergency scenarios at the friction limits. While recent reinforcement learning approaches show promise for drifting control, they struggle with the significant simulation-to-reality gap, as …

2026

MISF: MLLM Guided Iterative Sample Filtering for Data Fault Detection

AAAI 2026technical

High quality datasets are critical for training reliable machine learning models, yet data faults caused by insufficient annotation expertise or malicious poisoning attacks remain prevalent. Traditional classifier based methods rely on manually curated subsets for fault detection, but their limited

Cited by 0SourcePDFScholar
2026

PhysInOne: Visual Physics Learning and Reasoning in One Suite

CVPR 2026

We present PhysInOne, a large-scale synthetic dataset addressing the critical scarcity of physically-grounded training data for AI systems. Unlike existing datasets limited to merely hundreds or thousands of examples, PhysInOne provides 2 million videos across 153,810 dynamic 3D scenes, covering 71

Cited by 0SourcecodeScholar
2026

Scalable Multilingual Multimodal Machine Translation with Speech-Text Fusion

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have achieved notable success in enhancing translation performance by integrating multimodal information. However, existing research primarily focuses on image-guided methods, whose applicability is constrained by the scarcity of multilingual image-text pairs…

Cited by 0SourceScholar
2026

Single-Stage fMRI-to-3D Reconstruction via Viewpoint-Aware Embedding and Hierarchical Guidance

AAAI 2026technical

Understanding the neural basis of three-dimensional (3D) perception is a fundamental objective in cognitive neuroscience. Despite advances in decoding 2D visual stimuli from neural data, reconstructing high-fidelity 3D objects with detailed texture and geometry remains largely unexplored. In this wo

Cited by 0SourcePDFScholar
2026

SkyMoE: A Vision-Language Foundation Model for Enhancing Geospatial Interpretation with Mixture of Experts

AAAI 2026technical

The emergence of large vision-language models (VLMs) has significantly enhanced the efficiency and flexibility of geospatial interpretation. However, general-purpose VLMs remain suboptimal for remote sensing (RS) tasks. Existing geospatial VLMs typically adopt a unified modeling strategy and struggl

Cited by 0SourcePDFScholar
2026

Steering Large Language Models through the DMTA Cycle: Structure-Based Drug Design via Knowledge-Driven Bi-Level Thompson Sampling

ICML 2026poster

Structure-based drug design (SBDD) can be effectively realized through an iterative refinement via the Design-Make-Test-Analyze (DMTA) cycle, which is a common workflow used by human experts. However, most LLMs function as one-shot generators that lack feedback mechanisms, leaving the DMTA loop disc…

Cited by 0SourceScholar
2026

Towards Faithful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language Models

ICLR 2026poster

Vision-Language Models (VLMs) in remote sensing often fail at complex analytical tasks, a limitation stemming from their end-to-end training paradigm that bypasses crucial reasoning steps and leads to unverifiable outputs. To address this limitation, we introduce the Perceptually-Grounded Geospatial…

Cited by 0SourcecodeScholar
2026

Universal 3D Shape Matching via Coarse-to-Fine Language Guidance

CVPR 2026

Establishing dense correspondences between shapes is a crucial task in computer vision and graphics, while prior approaches depend on near-isometric assumptions and homogeneous subject types (i.e., only operate for human shapes). However, building semantic correspondences for cross-category objects

Cited by 0SourceScholar
2025

Brain-Inspired Spatial Continuous State Encoding for Efficient Spiking-Based Navigation

ICRA 2025

Spiking neural networks (SNNs) show great potential in mapless navigation tasks due to their low power consumption, but the continuous representation of spatial information poses a challenge to SNN training. Neuroscience findings reveal that spatial cognition cells encode spatial information through

Cited by 1SourceScholar
2025

Code-Generated Graph Representations Using Multiple LLM Agents for Material Properties Prediction

ICML 2025poster

Graph neural networks have recently demonstrated remarkable performance in predicting material properties. Crystalline material data is manually encoded into graph representations. Existing methods incorporate different attributes into constructing representations to satisfy the constraints arising…

Cited by 0SourcePDFScholar
2025

Conditional Prediction ROC Bands for Graph Classification

AISTATS 2025poster

Graph classification in medical imaging and drug discovery requires accuracy and robust uncertainty quantification. To address this need, we introduce Conditional Prediction ROC (CP-ROC) bands, offering uncertainty quantification for ROC curves and robustness to distributional shifts in test data. A…

Cited by 0SourcecodeScholar
2025

Differential-Flatness-Based Tracking Control for Tractor-Trailers in Reversing Maneuvers

IROS 2025

In this paper, we propose a differential-flatness-based controller (DFBC) for precise trajectory tracking of tractor-trailers, particularly during reversing maneuvers, which are challenging due to unstable equilibrium points. The proposed controller leverages the differential flatness property of tr

Cited by 0SourceScholar
2025

Distilling A Universal Expert from Clustered Federated Learning

IJCAI 2025

Clustered Federated Learning (CFL) addresses the challenges posed by non-IID data by training multiple group- or cluster-specific expert models. However, existing methods often overlook the shared information across clusters, which represents the generalizable knowledge valuable to all participants

Cited by 0SourcePDFScholar
2025

Energy-Efficient Omnidirectional Locomotion for Wheeled Quadrupeds via Predictive Energy-Aware Nominal Gait Selection

IROS 2025

Wheeled-legged robots combine the efficiency of wheels with the versatility of legs, but face significant energy optimization challenges when navigating diverse environments. In this work, we present a hierarchical control framework that integrates predictive power modeling with residual reinforceme

Cited by 0SourceScholar
2025

Enhancing Continual Learning for Medical Imaging: Efficient Knowledge Transfer and Multi-Disease Prediction

ICASSP 2025accepted

Deep learning models for medical disease detection require extensive labeled data, which is often scarce and expensive. Transfer learning can help by leveraging knowledge from large source domains, but directly fine-tuning these models can lead to catastrophic forgetting, making it impossible to reu…

Cited by 0SourceScholar
2025

FreeGave: 3D Physics Learning from Dynamic Videos by Gaussian Velocity

CVPR 2025poster

In this paper, we aim to model 3D scene geometry, appearance, and the underlying physics purely from multi-view videos. By applying various governing PDEs as PINN losses or incorporating physics simulation into neural networks, existing works often fail to learn complex physical motions at boundarie…

2025

GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement

ACL 2025long

The evolution of speech technology has been spurred by the rapid increase in dataset sizes. Traditional speech models generally depend on a large amount of labeled training data, which is scarce for low-resource languages. This paper presents GigaSpeech 2, a large-scale, multi-domain, multilingual s…

2025

GrabS: Generative Embodied Agent for 3D Object Segmentation without Scene Supervision

ICLR 2025spotlight

We study the hard problem of 3D object segmentation in complex point clouds without requiring human labels of 3D scenes for supervision. By relying on the similarity of pretrained 2D features or external signals such as motion to group 3D points as objects, existing unsupervised methods are usually…

2025

HSRL: A Hierarchical Control System Based on Spiking Deep Reinforcement Learning for Robot Navigation

ICRA 2025

Reinforcement Learning (RL) has shown promise in robotic navigation tasks, yet applying it to real-world environments remains challenging due to dynamic complexities and the need for dynamically feasible actions. We propose a hierarchical control framework based on Spiking Deep Reinforcement Learnin

Cited by 3SourceScholar
2025

Heterogeneous Graph Network-Based UWB Localization for Complex Indoor Environments

IROS 2025

Accurate indoor location-based services are important for mobile robots, especially in complex indoor environments. In this paper, we propose a heterogeneous graph network-based ultra-wide band (UWB) localization method to provide accurate and robust localization results for mobile robots in complex

Cited by 0SourceScholar
2025

Improving Adversarial Transferability on Vision Transformers via Forward Propagation Refinement

CVPR 2025poster

Vision Transformers (ViTs) have been widely applied in various computer vision and vision-language tasks. To gain insights into their robustness in practical scenarios, transferable adversarial examples on ViTs have been extensively studied. A typical approach to improving adversarial transferabilit…

2025

Improving Integrated Gradient-based Transferable Adversarial Examples by Refining the Integration Path

AAAI 2025technical

Transferable adversarial examples are known to cause threats in practical, black-box attack scenarios. A notable approach to improving transferability is using integrated gradients (IG), originally developed for model interpretability. In this paper, we find that existing IG-based attacks have limit…

2025

Learning to Drift With Individual Wheel Drive: Maneuvering Autonomous Vehicle at the Handling Limits

RA-L 2025

Drifting, characterized by controlled vehicle motion at high sideslip angles, is crucial for safely handling emergency scenarios at the friction limits. While recent reinforcement learning approaches show promise for drifting control, they struggle with the significant simulation-to-reality gap, as

Cited by 0SourceScholar
2025

LogoSP: Local-global Grouping of Superpoints for Unsupervised Semantic Segmentation of 3D Point Clouds

CVPR 2025poster

We study the problem of unsupervised 3D semantic segmentation on raw point clouds without needing human labels in training. Existing methods usually formulate this problem into learning per-point local features followed by a simple grouping strategy, lacking the ability to discover additional and po…

2025

LossControl: Defending Membership Inference Attacks by Controlling the Loss

ICASSP 2025accepted

Machine learning models are vulnerable to membership inference attacks (MIAs), where adversaries attempt to predict whether specific samples are part of the model’s training set. Previous studies have demonstrated a strong correlation between the distinguishability of training and testing loss distr…

Cited by 0SourceScholar
2025

Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning

ACL 2025long

Multimodal Large Language Models (MLLMs) have achieved significant success in Speech-to-Text Translation (S2TT) tasks. While most existing research has focused on English-centric translation directions, the exploration of many-to-many translation is still limited by the scarcity of parallel data. To…

2025

Multifaceted User Modeling in Recommendation: A Federated Foundation Models Approach

AAAI 2025technical

Multifaceted user modeling aims to uncover fine-grained patterns and learn representations from user data, revealing their diverse interests and characteristics, such as profile, preference, and personality. Recent studies on foundation model-based recommendation have emphasized the Transformer arch…

2025

RayletDF: Raylet Distance Fields for Generalizable 3D Surface Reconstruction from Point Clouds or Gaussians

ICCV 2025poster

In this paper, we present a generalizable method for 3D surface reconstruction from raw point clouds or pre-estimated 3D Gaussians by 3DGS from RGB images. Unlike existing coordinate-based methods which are often computationally intensive when rendering explicit surfaces, our proposed method, named…

Cited by 0SourcePDFScholar
2025

State Feedback Enhanced Graph Differential Equations for Multivariate Time Series Forecasting

IJCAI 2025

Multivariate time series forecasting holds significant theoretical and practical importance in various fields, including web analytics and transportation. Recently, graph neural networks and graph differential equations have shown exceptional capabilities in modeling spatio-temporal features. Howeve

2025

TimeEmb: A Lightweight Static-Dynamic Disentanglement Framework for Time Series Forecasting

NeurIPS 2025poster

Temporal non-stationarity, the phenomenon that time series distributions change over time, poses fundamental challenges to reliable time series forecasting. Intuitively, the complex time series can be decomposed into two factors, i.e., time-invariant and time-varying components, which indicate stati…

Cited by 0SourcecodeScholar
2025

Towards Generalizable Neural Simulators: Addressing Distribution Shifts Induced by Environmental and Temporal Variations

IJCAI 2025

With advancements in deep learning, neural simulators have become increasingly important for improving the efficiency and effectiveness of simulating complex dynamical systems in various scientific and technological fields. This paper presents a novel neural simulator called Context-informed Polymor

2025

unMORE: Unsupervised Multi-Object Segmentation via Center-Boundary Reasoning

ICML 2025poster

We study the challenging problem of unsupervised multi-object segmentation on single images. Existing methods, which rely on image reconstruction objectives to learn objectness or leverage pretrained image features to group similar pixels, often succeed only in segmenting simple synthetic objects or…

2024

CASRL: Collision Avoidance with Spiking Reinforcement Learning Among Dynamic, Decision-Making Agents

IROS 2024poster

Developing an efficient collision avoidance policy with Spiking Reinforcement Learning for dynamic, decision-making agents remains challenging. Moreover, the implementation of energy-efficient collision avoidance is important for mobile robots that operate with limited on-board computing resources.…

Cited by 0SourceScholar
2024

Decision Mamba: Reinforcement Learning via Hybrid Selective Sequence Modeling

NeurIPS 2024poster

Recent works have shown the remarkable superiority of transformer models in reinforcement learning (RL), where the decision-making problem is formulated as sequential generation. Transformer-based agents could emerge with self-improvement in online environments by providing task contexts, such as mu…

Cited by 7SourcePDFScholar
2024

Federated Adaptation for Foundation Model-based Recommendations

IJCAI 2024poster

With the recent success of large language models, particularly foundation models with generalization abilities, applying foundation models for recommendations becomes a new paradigm to improve existing recommendation systems. It becomes a new open challenge to enable the foundation model to capture…

2024

Fine-Grained Prototypes Distillation for Few-Shot Object Detection

AAAI 2024technical

Few-shot object detection (FSOD) aims at extending a generic detector for novel object detection with only a few training examples. It attracts great concerns recently due to the practical meanings. Meta-learning has been demonstrated to be an effective paradigm for this task. In general, methods ba…

2024

In-Context Decision Transformer: Reinforcement Learning via Hierarchical Chain-of-Thought

ICML 2024poster

In-context learning is a promising approach for offline reinforcement learning (RL) to handle online tasks, which can be achieved by providing task prompts. Recent works demonstrated that in-context RL could emerge with self-improvement in a trial-and-error manner when treating RL tasks as an across…

2024

Learning to Catch Reactive Objects with a Behavior Predictor

ICRA 2024poster

Tracking and catching moving objects is an important ability for robots in a dynamic world. Whilst some objects have highly predictable state evolution e.g., the ballistic trajectory of a tennis ball, reactive targets alter their behavior in response to motion of the manipulator. Reactive applicatio…

Cited by 2SourcecodeScholar
2024

Stochastic Neural Simulator for Generalizing Dynamical Systems across Environments

IJCAI 2024poster

Neural simulators for modeling complex dynamical systems have been extensively studied for various real-world applications, such as weather forecasting, ocean current prediction, and computational fluid dynamics simulation. Although they have demonstrated powerful fitting and predicting, most existi…

2023

Consecutive Inertia Drift of Autonomous RC Car via Primitive-Based Planning and Data-Driven Control

IROS 2023poster

Inertia drift is an aggressive transitional driving maneuver, which is challenging due to the high nonlinearity of the system and the stringent requirement on control and planning performance. This paper presents a solution for the consecutive inertia drift of an autonomous RC car based on primitive…

Cited by 4SourceScholar
2023

Decoupling Skill Learning from Robotic Control for Generalizable Object Manipulation

ICRA 2023poster

Recent works in robotic manipulation through reinforcement learning (RL) or imitation learning (IL) have shown potential for tackling a range of tasks e.g., opening a drawer or a cupboard. However, these techniques generalize poorly to unseen objects. We conjecture that this is due to the high-dimen…

Cited by 5SourcecodeScholar
2023

Do We Need an Encoder-Decoder to Model Dynamical Systems on Networks?

IJCAI 2023poster

As deep learning gains popularity in modelling dynamical systems, we expose an underappreciated misunderstanding relevant to modelling dynamics on networks. Strongly influenced by graph neural networks, latent vertex embeddings are naturally adopted in many neural dynamical network models. However,…

2023

DriveIRL: Drive in Real Life with Inverse Reinforcement Learning

ICRA 2023poster

In this paper, we introduce the first published planner to drive a car in dense, urban traffic using Inverse Reinforcement Learning (IRL). Our planner, DriveIRL, generates a diverse set of trajectory proposals and scores them with a learned model. The best trajectory is tracked by our self-driving v…

Cited by 32SourceScholar
2023

Dual Personalization on Federated Recommendation

IJCAI 2023poster

Federated recommendation is a new Internet service architecture that aims to provide privacy-preserving recommendation services in federated settings. Existing solutions are used to combine distributed recommendation algorithms and privacy-preserving mechanisms. Thus it inherently takes the form of…

2023

Learning Generalizable Agents via Saliency-guided Features Decorrelation

NeurIPS 2023spotlight

In visual-based Reinforcement Learning (RL), agents often struggle to generalize well to environmental variations in the state space that were not observed during training. The variations can arise in both task-irrelevant features, such as background noise, and task-relevant features, such as robot…

Cited by 9SourcePDFScholar
2023

NVFi: Neural Velocity Fields for 3D Physics Learning from Dynamic Videos

NeurIPS 2023poster

In this paper, we aim to model 3D scene dynamics from multi-view videos. Unlike the majority of existing works which usually focus on the common task of novel view synthesis within the training time period, we propose to simultaneously learn the geometry, appearance, and physical velocity of 3D scen…

2023

NeAT: Learning Neural Implicit Surfaces With Arbitrary Topologies From Multi-View Images

CVPR 2023poster

Recent progress in neural implicit functions has set new state-of-the-art in reconstructing high-fidelity 3D shapes from a collection of images. However, these approaches are limited to closed surfaces as they require the surface to be represented by a signed distance field. In this paper, we propos…

2023

NeUDF: Leaning Neural Unsigned Distance Fields With Volume Rendering

CVPR 2023poster

Multi-view shape reconstruction has achieved impressive progresses thanks to the latest advances in neural implicit surface rendering. However, existing methods based on signed distance function (SDF) are limited to closed surfaces, failing to reconstruct a wide range of real-world objects that cont…

Cited by 58SourcePDFScholar
2023

OPT-GAN: A Broad-Spectrum Global Optimizer for Black-Box Problems by Learning Distribution

AAAI 2023technical

Black-box optimization (BBO) algorithms are concerned with finding the best solutions for problems with missing analytical details. Most classical methods for such problems are based on strong and fixed a priori assumptions, such as Gaussianity. However, the complex real-world problems, especially w…

2023

RayDF: Neural Ray-surface Distance Fields with Multi-view Consistency

NeurIPS 2023poster

In this paper, we study the problem of continuous 3D shape representations. The majority of existing successful methods are coordinate-based implicit neural representations. However, they are inefficient to render novel views or recover explicit surface points. A few works start to formulate 3D shap…

2023

STS-GAN: Can We Synthesize Solid Texture with High Fidelity from Arbitrary 2D Exemplar?

IJCAI 2023poster

Solid texture synthesis (STS), an effective way to extend a 2D exemplar to a 3D solid volume, exhibits advantages in computational photography. However, existing methods generally fail to accurately learn arbitrary textures, which may result in the failure to synthesize solid textures with high fide…

Cited by 12SourcePDFScholar
2023

Spiking Reinforcement Learning with Memory Ability for Mapless Navigation

IROS 2023poster

Our study focuses on mapless navigation in robotics, which involves navigating without an established obstacle map of the environment. Spiking Neural Networks (SNNs) have recently been applied to this task using Deep Reinforcement Learning (DRL), but face challenges in dynamic and partially observab…

Cited by 3SourceScholar
2022

3PSDF: Three-Pole Signed Distance Function for Learning Surfaces With Arbitrary Topologies

CVPR 2022poster

Recent advances in learning 3D shapes using neural implicit functions have achieved impressive results by breaking the previous barrier of resolution and diversity for varying topologies. However, most of such approaches are limited to closed surfaces as they require the space to be divided into ins…

Cited by 36PDFScholar
2022

A Probabilistic Graphical Model Based on Neural-Symbolic Reasoning for Visual Relationship Detection

CVPR 2022poster

This paper aims to leverage symbolic knowledge to improve the performance and interpretability of the Visual Relationship Detection (VRD) models. Existing VRD methods based on deep learning suffer from the problems of poor performance on insufficient labeled examples and lack of interpretability. To…

Cited by 28PDFScholar
2022

HSDF: Hybrid Sign and Distance Field for Modeling Surfaces with Arbitrary Topologies

NeurIPS 2022accept

Neural implicit function based on signed distance field (SDF) has achieved impressive progress in reconstructing 3D models with high fidelity. However, such approaches can only represent closed shapes. Recent works based on unsigned distance function (UDF) are proposed to handle both watertight and…

Cited by 21SourcePDFScholar
2022

Promising or Elusive? Unsupervised Object Segmentation from Real-world Single Images

NeurIPS 2022accept

In this paper, we study the problem of unsupervised object segmentation from single images. We do not introduce a new algorithm, but systematically investigate the effectiveness of existing unsupervised models on challenging real-world images. We firstly introduce four complexity factors to quantita…

2022

SQN: Weakly-Supervised Semantic Segmentation of Large-Scale 3D Point Clouds

ECCV 2022poster

"Labelling point clouds fully is highly time-consuming and costly. As larger point cloud datasets containing billions of points become more common, we ask whether the full annotation is even necessary, demonstrating that existing baselines designed under a fully annotated assumption only degrade sli…

2022

Sentiment-Aware Automatic Speech Recognition Pre-Training for Enhanced Speech Emotion Recognition

ICASSP 2022accepted

We propose a novel multi-task pre-training method for Speech Emotion Recognition (SER). We pre-train SER model simultaneously on Automatic Speech Recognition (ASR) and sentiment classification tasks to make the acoustic ASR model more "emotion aware". We generate targets for the sentiment classifica…

Cited by 0SourceScholar
2021

Context-Sensitive Temporal Feature Learning for Gait Recognition

ICCV 2021poster

Although gait recognition has drawn increasing research attention recently, it remains challenging to learn discriminative temporal representation since the silhouette differences are quite subtle in spatial domain. Inspired by the observation that humans can distinguish gaits of different subjects…

Cited by 169PDFcodeScholar
2021

Contrastive Unsupervised Learning for Speech Emotion Recognition

ICASSP 2021accepted

Speech emotion recognition (SER) is a key technology to enable more natural human-machine communication. However, SER has long suffered from a lack of public large-scale labeled datasets. To circumvent this problem, we investigate how unsupervised representation learning on unlabeled datasets can be…

Cited by 0SourceScholar
2021

Cost-aware Graph Generation: A Deep Bayesian Optimization Approach

AAAI 2021technical

Graph-structured data is ubiquitous throughout the natural and social sciences, ranging from complex drug molecules to artificial neural networks. Evaluating their functional properties, e.g., drug effectiveness and prediction accuracy, is usually costly in terms of time, money, energy, or environme…

2021

Graph and Temporal Convolutional Networks for 3D Multi-person Pose Estimation in Monocular Videos

AAAI 2021technical

Despite the recent progress, 3D multi-person pose estimation from monocular videos is still challenging due to the commonly encountered problem of missing information caused by occlusion, partially out-of-frame target persons, and inaccurate person detection. To tackle this problem, we propose a nov…

2021

Learning Interpretable Decision Rule Sets: A Submodular Optimization Approach

NeurIPS 2021spotlight

Rule sets are highly interpretable logical models in which the predicates for decision are expressed in disjunctive normal form (DNF, OR-of-ANDs), or, equivalently, the overall model comprises an unordered collection of if-then decision rules. In this paper, we consider a submodular optimization bas…

Cited by 35SourcePDFScholar
2021

Monocular 3D Multi-Person Pose Estimation by Integrating Top-Down and Bottom-Up Networks

CVPR 2021poster

In monocular video 3D multi-person pose estimation, inter-person occlusion and close interactions can cause human detection to be erroneous and human-joints grouping to be unreliable. Existing top-down methods rely on human detection and thus suffer from these problems. Existing bottom-up methods do…

Cited by 59PDFcodeScholar
2021

OctField: Hierarchical Implicit Functions for 3D Modeling

NeurIPS 2021poster

Recent advances in localized implicit functions have enabled neural implicit representation to be scalable to large scenes. However, the regular subdivision of 3D space employed by these approaches fails to take into account the sparsity of the surface occupancy and the varying granularities of geom…

Cited by 38SourcePDFScholar
2021

SpinNet: Learning a General Surface Descriptor for 3D Point Cloud Registration

CVPR 2021poster

Extracting robust and general 3D local features is key to downstream tasks such as point cloud registration and reconstruction. Existing learning-based local descriptors are either sensitive to rotation transformations, or rely on classical handcrafted features which are neither general nor represen…

Cited by 380PDFcodeScholar
2021

Towards Semantic Segmentation of Urban-Scale 3D Point Clouds: A Dataset, Benchmarks and Challenges

CVPR 2021poster

An essential prerequisite for unleashing the potential of supervised deep learning algorithms in the area of 3D scene understanding is the availability of large-scale and richly annotated datasets. However, publicly available datasets are either in relatively small spatial scales or have limited sem…

Cited by 236PDFcodeScholar
2020

Curvature Regularization to Prevent Distortion in Graph Embedding

NeurIPS 2020spotlight

Recent research on graph embedding has achieved success in various applications. Most graph embedding methods preserve the proximity in a graph into a manifold in an embedding space. We argue an important but neglected problem about this proximity-preserving strategy: Graph topology patterns, while…

Cited by 15SourcePDFScholar
2020

Finding Second-Order Stationary Points Efficiently in Smooth Nonconvex Linearly Constrained Optimization Problems

NeurIPS 2020spotlight

This paper proposes two efficient algorithms for computing approximate second-order stationary points (SOSPs) of problems with generic smooth non-convex objective functions and generic linear constraints. While finding (approximate) SOSPs for the class of smooth non-convex linearly constrained probl…

Cited by 27SourcePDFScholar
2020

Geom-GCN: Geometric Graph Convolutional Networks

ICLR 2020spotlight

Message-passing neural networks (MPNNs) have been successfully applied in a wide variety of applications in the real world. However, two fundamental weaknesses of MPNNs' aggregators limit their ability to represent graph-structured data: losing the structural information of nodes in neighborhoods an…

Cited by 1460SourcecodeScholar
2020

RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds

CVPR 2020oral

We study the problem of efficient semantic segmentation for large-scale 3D point clouds. By relying on expensive sampling techniques or computationally heavy pre/post-processing steps, most existing approaches are only able to be trained and operate over small-scale point clouds. In this paper, we i…

Cited by 2143PDFcodeScholar
2019

DeepPCO: End-to-End Point Cloud Odometry through Deep Parallel Neural Network

IROS 2019poster

Odometry is of key importance for localization in the absence of a map. There is considerable work in the area of visual odometry (VO), and recent advances in deep learning have brought novel approaches to VO, which directly learn salient features from raw images. These learning-based approaches hav…

Cited by 61SourceScholar
2019

Learning Object Bounding Boxes for 3D Instance Segmentation on Point Clouds

NeurIPS 2019spotlight

We propose a novel, conceptually simple and general framework for instance segmentation on 3D point clouds. Our method, called 3D-BoNet, follows the simple design philosophy of per-point multilayer perceptrons (MLPs). The framework directly regresses 3D bounding boxes for all instances in a point cl…

2018

DEFO-NET: Learning Body Deformation Using Generative Adversarial Networks

ICRA 2018poster

Modelling the physical properties of everyday objects is a fundamental prerequisite for autonomous robots. We present a novel generative adversarial network (DEFO-NET), able to predict body deformations under external forces from a single RGB-D image. The network is based on an invertible conditiona…

Cited by 10SourceScholar
2017

Towards K-means-friendly Spaces: Simultaneous Deep Learning and Clustering

ICML 2017poster

Most learning approaches treat dimensionality reduction (DR) and clustering separately (i.e., sequentially), but recent research has shown that optimizing the two tasks jointly can substantially improve the performance of both. The premise behind the latter genre is that the data samples are obtaine…