← Search

Jiaxin LI

32 accepted papers

2026

DensiCrafter: Physically-Constrained Generation and Fabrication of Self-Supporting Hollow Structures

AAAI 2026technical

The rise of 3D generative models has enabled automatic 3D geometry and texture synthesis from multimodal inputs (e.g., text or images). However, these methods often ignore physical constraints and manufacturability considerations. In this work, we address the challenge of producing 3D designs that a

Cited by 0SourcePDFScholar
2026

Drugging the Undruggable: Benchmarking and Modeling Fragment-Based Screening

ICLR 2026poster

A significant portion of disease-relevant proteins remain undruggable due to shallow, flexible, or otherwise ill-defined binding pockets that hinder conventional molecule screening. Fragment-based drug discovery (FBDD) offers a promising alternative, as small, low-complexity fragments can flexibly e…

Cited by 0SourceScholar
2026

HOPS: Hierarchical Open-vocabulary Part Segmentation with Attention-Aware Filtering and Affinity-Guided Enhancement

CVPR 2026

Open-vocabulary part segmentation (OVPS) aims to segment objects into fine-grained parts while generalizing to unseen categories. Existing VLM-based methods face two challenges: (1) object over-segmentation, caused by overly broad semantic activations, and (2) part under-segmentation, resulting from

Cited by 0SourcecodeScholar
2026

InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields

CVPR 2026

Existing depth estimation methods are fundamentally limited to predicting depth on discrete image grids. Such representations restrict their scalability to arbitrary output resolutions and hinder the geometric detail recovery. This paper introduces InfiniDepth, which represents depth as neural impli

Cited by 0SourcecodeScholar
2026

OmniXtreme: Breaking the Generality Barrier in High-Dynamic Humanoid Control

RSS 2026poster

High-fidelity motion tracking serves as the ultimate litmus test for generalizable, human-level motor skills. However, current policies often hit a “generality barrier”: as motion libraries scale in diversity, tracking fidelity inevitably collapses—especially for real-world deployment of high-dynami…

Cited by 0SourceScholar
2026

Pansharpening for Thin-Cloud Contaminated Remote Sensing Images: A Unified Framework and Benchmark Dataset

AAAI 2026technical

Pansharpening under thin cloudy conditions is a practically significant yet rarely addressed task, challenged by simultaneous spatial resolution degradation and cloud-induced spectral distortions. Existing methods often address cloud removal and pansharpening sequentially, leading to cumulative erro

Cited by 0SourcePDFScholar
2025

Anchor-Aware Similarity Cohesion in Target Frames Enables Predicting Temporal Moment Boundaries in 2D

CVPR 2025poster

Video moment retrieval aims to locate specific moments from a video according to the query text. This task presents two main challenges: i) aligning the query and video frames at the feature level, and ii) projecting the query-aligned frame features to the start and end boundaries of the matching in…

2025

Extending LLM Context Window with Adaptive Grouped Positional Encoding: A Training-Free Method

ACL 2025long

Processing long input remains a significant challenge for large language models (LLMs) due to the scarcity of large-scale long-context training data and the high computational cost of training models for extended context windows. In this paper, we propose **Ada**ptive **Gro**uped **P**ositional **E*…

Cited by 0SourcePDFScholar
2025

FloNa: Floor Plan Guided Embodied Visual Navigation

AAAI 2025technical

Humans naturally rely on floor plans to navigate in unfamiliar environments, as they are readily available, reliable, and provide rich geometrical guidance. However, existing visual navigation settings overlook this valuable prior knowledge, leading to limited efficiency and accuracy. To eliminate t…

Cited by 0SourcePDFScholar
2025

HetSSNet: Spatial-Spectral Heterogeneous Graph Learning Network for Panchromatic and Multispectral Images Fusion

ICML 2025poster

Remote sensing pansharpening aims to reconstruct spatial-spectral properties during the fusion of panchromatic (PAN) images and low- resolution multi-spectral (LR-MS) images, finally generating the high-resolution multi-spectral (HR- MS) images. In the mainstream modeling strategies, i.e., CNN and T…

Cited by 0SourcePDFScholar
2025

M^3-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation

CVPR 2025poster

Intelligent robots need to interact with diverse objects across various environments. The appearance and state of objects frequently undergo complex transformations depending on the object properties, e.g., phase transitions. However, in the vision community, segmenting dynamic objects with phase tr…

2025

Modality-Aware Shot Relating and Comparing for Video Scene Detection

AAAI 2025technical

Video scene detection involves assessing whether each shot and its surroundings belong to the same scene. Achieving this requires meticulously correlating multi-modal cues, e.g., visual entity and place modalities, among shots and comparing semantic changes around each shot. However, most methods tr…

2025

Multilingual Federated Low-Rank Adaptation for Collaborative Content Anomaly Detection across Multilingual Social Media Participants

EMNLP 2025

Recently, the rapid development of multilingual social media platforms (SNS) exacerbates new challenges in SNS content anomaly detection due to data islands and linguistic imbalance. While federated learning (FL) and parameter-efficient fine-tuning (PEFT) offer potential solutions in most cases, whe

2025

Open-Vocabulary Part Segmentation via Progressive and Boundary-Aware Strategy

NeurIPS 2025poster

Open-vocabulary part segmentation (OVPS) struggles with structurally connected boundaries due to the inherent conflict between continuous image features and discrete classification mechanism. To address this, we propose PBAPS, a novel training-free framework specifically designed for OVPS. PBAPS lev…

Cited by 0SourcecodeScholar
2025

StreamForest: Efficient Online Video Understanding with Persistent Event Memory

NeurIPS 2025spotlight

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in video understanding. However, their effectiveness in real-time streaming scenarios remains limited due to storage constraints of historical visual features and insufficient real-time spatiotemporal reasoning. To a…

Cited by 0SourceScholar
2025

TSTAI: A Time-varying Brain Effective Connectivity Network Construction Method Combining with Brain Active Information

IJCAI 2025

More accurate construction of brain effective conncetivity networks remains a great challenge to achieve accurate auxiliary diagnosis of brain diseases and in-depth exploration of brain function. However, existing methods only consider higher-order or non-stationary assumptions, rather than simultan

Cited by 0SourcePDFScholar
2024

BENO: Boundary-embedded Neural Operators for Elliptic PDEs

ICLR 2024poster

Elliptic partial differential equations (PDEs) are a major class of time-independent PDEs that play a key role in many scientific and engineering domains such as fluid dynamics, plasma physics, and solid mechanics. Recently, neural operators have emerged as a promising technique to solve elliptic PD…

2024

NeRI: Implicit Neural Representation of LiDAR Point Cloud Using Range Image Sequence

ICASSP 2024accepted

This paper proposes the NeRI, an implicit neural representation (INR) based LiDAR point cloud compressor. In NeRI, we first transform a sequence of 3D LiDAR frames into a 2D range image sequence through range image projection over time. Then, we employ a neural network conditioned on the temporal fr…

Cited by 0SourceScholar
2024

Neighbor Relations Matter in Video Scene Detection

CVPR 2024poster

Video scene detection aims to temporally link shots for obtaining semantically compact scenes. It is essential for this task to capture scene-distinguishable affinity among shots by similarity assessment. However most methods relies on ordinary shot-to-shot similarities which may inveigle similar sh…

2024

Visual Loop Closure Detection with Thorough Temporal and Spatial Context Exploitation

IROS 2024poster

Despite advancements in visual Simultaneous Localization and Mapping (SLAM), prevailing visual Loop Closure Detection (LCD) methods primarily rely on computationally intensive image similarity comparisons, neglecting temporal-spatial context during long-term exploration. To address this issue, we pr…

Cited by 0SourceScholar
2023

Streaming Voice Conversion via Intermediate Bottleneck Features and Non-Streaming Teacher Guidance

ICASSP 2023accepted

Streaming voice conversion (VC) is the task of converting the voice of one person to another in real-time. Previous streaming VC methods use phonetic posteriorgrams (PPGs) extracted from automatic speech recognition (ASR) systems to represent speaker-independent information. However, PPGs lack the p…

Cited by 0SourceScholar
2022

A Deep-Learning-based System for Indoor Active Cleaning

IROS 2022poster

Cleaning public areas like commercial complexes is challenging due to their sophisticated surroundings and the vast kinds of real-life dirt. Robots are required to distinguish dirts and apply corresponding cleaning strategies. In this work, we proposed an active-cleaning framework by utilizing deep-…

Cited by 2SourcecodeScholar
2022

LiDAR Distillation: Bridging the Beam-Induced Domain Gap for 3D Object Detection

ECCV 2022poster

"In this paper, we propose the LiDAR Distillation to bridge the domain gap induced by different LiDAR beams for 3D object detection. In many real-world applications, the LiDAR points used by mass-produced robots and vehicles usually have fewer beams than that in large-scale public datasets. Moreover…

2021

Interaction via Bi-Directional Graph of Semantic Region Affinity for Scene Parsing

ICCV 2021poster

In this work, we devote to address the challenging problem of scene parsing. Previous methods, though capture context to exploit global clues, handle scene parsing as a pixel-independent task. However, it is well known that pixels in an image are highly correlated with each other, especially those f…

Cited by 19PDFScholar
2021

MINE: Towards Continuous Depth MPI With NeRF for Novel View Synthesis

ICCV 2021poster

In this paper, we propose MINE to perform novel view synthesis and depth estimation via dense 3D reconstruction from a single image. Our approach is a continuous depth generalization of the Multiplane Images (MPI) by introducing the NEural radiance fields (NeRF). Given a single image as input, MINE…

Cited by 171PDFcodeScholar
2021

MT-ORL: Multi-Task Occlusion Relationship Learning

ICCV 2021poster

Retrieving occlusion relation among objects in a single image is challenging due to sparsity of boundaries in image. We observe two key issues in existing works: firstly, lack of an architecture which can exploit the limited amount of coupling in the decoder stage between the two subtasks, namely oc…

Cited by 8PDFcodeScholar