← Search

Zhenyu Li

39 accepted papers

2026

Any Resolution Any Geometry: From Multi-View To Multi-Patch

CVPR 2026

Joint estimation of surface normals and depth is essential for holistic 3D scene understanding, yet high-resolution prediction remains difficult due to the trade-off between preserving fine local detail and maintaining global consistency. To address this challenge, we propose the Ultra Resolution Ge

Cited by 0SourcecodeScholar
2026

Depth Anything 3: Recovering the Visual Space from Any Views

ICLR 2026oral

We present Depth Anything 3 (DA3), a model that predicts spatially consistent geometry from an arbitrary number of visual inputs, with or without known camera poses. In pursuit of minimal modeling, DA3 yields two key insights: a single plain transformer (e.g., vanilla DINOv2 encoder) is sufficient…

Cited by 0SourcecodeScholar
2026

EdgeMTSC: A Lightweight Large-Kernel ConvNet for Multivariate Time Series Classification

AAAI 2026technical

In large-scale sensor networks, Multivariate Time Series Classification (MTSC) is a pivotal task for identifying events dependent on longitudinal data at the edge. However, existing methods focus on neither the inherent ability of convolutional networks to perceive subsequence features, nor the prol

Cited by 0SourcePDFScholar
2026

EgoPoseFormer v2: Accurate Egocentric Human Motion Estimation for AR/VR

CVPR 2026

Egocentric 3D human motion estimation is essential for AR/VR experiences, yet remains challenging due to limited body coverage from the egocentric viewpoint, frequent occlusions, and scarce labeled data. We present EgoPoseFormer v2, a method that addresses these challenges through two key contributi

Cited by 0SourceScholar
2026

FourierPlace: A Vision-Language Localization Framework Based on Frequency Domain Representations

RA-L 2026

Language-guided localization within 3D environments continues to pose a significant challenge for autonomous systems, primarily due to the need for precise alignment between sparse point cloud data and inherently ambiguous natural language descriptions. To address this, we present a novel vision-lan

Cited by 0SourcecodeScholar
2026

LaRI: Layered Ray Intersections for Single-view 3D Geometric Reasoning

ICML 2026poster

We present Layered Ray Intersections (LaRI), a fully supervised method for occluded geometry reasoning from a single image. Unlike conventional depth estimation, which is limited to visible surfaces, LaRI predicts multiple surfaces intersected by the camera rays using layered point maps. Compared to…

Cited by 0SourceScholar
2026

PatchRefiner V2: Fast and Lightweight Real-Domain High-Resolution Metric Depth Estimation

ICLR 2026poster

While current high-resolution depth estimation methods achieve strong results, they often suffer from computational inefficiencies due to reliance on heavyweight models and multiple inference steps, increasing inference time. To address this, we introduce PatchRefiner V2 (PRV2), which replaces heavy…

Cited by 0SourceScholar
2025

Amodal Depth Anything: Amodal Depth Estimation in the Wild

ICCV 2025poster

Amodal depth estimation aims to predict the depth of occluded (invisible) parts of objects in a scene. This task addresses the question of whether models can effectively perceive the geometry of occluded regions based on visible cues. Prior methods primarily rely on synthetic datasets and focus on m…

Cited by 0SourcePDFScholar
2025

Bridging Text and Vision: A Multi-View Text-Vision Registration Approach for Cross-Modal Place Recognition

IROS 2025

Mobile robots necessitate advanced natural language understanding capabilities to accurately identify locations and perform tasks such as package delivery. However, traditional visual place recognition (VPR) methods rely solely on single-view visual information and cannot interpret human language de

Cited by 6SourcecodeScholar
2025

COMM: Concentrated Margin Maximization for Robust Document-Level Relation Extraction

AAAI 2025technical

Document-level relation extraction (DocRE) is the process of identifying and extracting relations between entities that span multiple sentences within a document. Due to its realistic settings, DocRE has garnered increasing research attention in recent years. Previous research has mostly focused on…

Cited by 0SourcePDFScholar
2025

CoPRA: Bridging Cross-domain Pretrained Sequence Models with Complex Structures for Protein-RNA Binding Affinity Prediction

AAAI 2025technical

Accurately measuring protein-RNA binding affinity is crucial in many biological processes and drug design. Previous computational methods for protein-RNA binding affinity prediction rely on either sequence or structure features, unable to capture the binding mechanisms comprehensively. The recent em…

2025

FocusLLM: Precise Understanding of Long Context by Dynamic Condensing

ACL 2025long

Empowering LLMs with the ability to precisely understand long contexts is crucial for many downstream applications. However, handling long contexts with conventional transformer architecture requires substantial training and inference resources. Existing context condensing methods cannot accurately…

2025

M2PA: A Multi-Memory Planning Agent for Open Worlds Inspired by Cognitive Theory

ACL 2025finding

Open-world planning poses a significant challenge for general artificial intelligence due to environmental complexity and task diversity, especially in long-term tasks and lifelong learning. Inspired by cognitive theories, we propose M2PA, an open-world multi-memory planning agent. M2PA innovates by…

Cited by 0SourcePDFScholar
2025

MambaPlace: Text-to-Point-Cloud Cross-Modal Place Recognition with Attention Mamba Mechanisms

IROS 2025

Vision-Language Place Recognition (VLPR) enhances robot localization performance by incorporating natural language descriptions from images. By utilizing language information, VLPR directs robot place matching, overcoming the constraint of solely depending on vision. However, general multimodal info

Cited by 8SourcecodeScholar
2025

Maximum Score Routing For Mixture-of-Experts

ACL 2025finding

Routing networks in sparsely activated mixture-of-experts (MoE) dynamically allocate input tokens to top-k experts through differentiable sparse transformations, enabling scalable model capacity while preserving computational efficiency. Traditional MoE networks impose an expert capacity constraint…

2025

Metagent-P: A Neuro-Symbolic Planning Agent with Metacognition for Open Worlds

ACL 2025finding

The challenge of developing agents capable of open-world planning remains fundamental to artificial general intelligence (AGI). While large language models (LLMs) have made progress with their vast world knowledge, their limitations in perception, memory, and reliable reasoning still hinder LLM-base…

Cited by 0SourcePDFScholar
2025

Negative Matters: Multi-Granularity Hard-Negative Synthesis and Anchor-Token-Aware Pooling for Enhanced Text Embeddings

ACL 2025long

Text embedding models are essential for various natural language processing tasks, enabling the effective encoding of semantic information into dense vector representations. These models are typically optimized using triplets of (query, positive, negative) data pairs for contrastive learning, where…

Cited by 0SourcePDFScholar
2025

PseDet: Revisiting the Power of Pseudo Label in Incremental Object Detection

ICLR 2025poster

Incremental Objection Detection (IOD) facilitates the expansion of the usage scope of object detectors without forgetting previously acquired knowledge. Current approaches mostly adopt response-level knowledge distillation to overcome forgetting issues, by conducting implicit memory replay from the…

Cited by 0SourcePDFScholar
2025

TROI: Cross-Subject Pretraining with Sparse Voxel Selection for Enhanced fMRI Visual Decoding

ICASSP 2025accepted

fMRI (functional Magnetic Resonance Imaging) visual decoding involves decoding the original image from brain signals elicited by visual stimuli. This often relies on manually labeled ROIs (Regions of Interest) to select brain voxels. However, these ROIs can contain redundant information and noise, r…

Cited by 0SourceScholar
2025

The Devil is in the Quality: Exploring Informative Samples for Semi-Supervised Monocular 3D Object Detection

ICRA 2025

This paper tackles the challenging problem of semi-supervised monocular 3D object detection with a general framework. In specific, having observed that the bottleneck of this task lies in lacking reliable and informative samples from unlabeled data for detector learning, we introduce a novel simple

Cited by 0SourceScholar
2025

Try Before You Buy: Solving Multi-Model Complex Tasks by Model Competitions

ICASSP 2025accepted

Multi-modal large language models (MLLMs) are expanded from large language models (LLMs) with additional capabilities to infer multi-modal data. Current MLLM workflows, when dealing with complex tasks, typically begin by using an LLM to decompose the task into multiple subtasks, then heuristically s…

Cited by 0SourceScholar
2025

XIFBench: Evaluating Large Language Models on Multilingual Instruction Following

NeurIPS 2025poster

Large Language Models (LLMs) have demonstrated remarkable instruction-following capabilities across various applications. However, their performance in multilingual settings lacks systematic investigation, with existing evaluations lacking fine-grained constraint analysis across diverse linguistic c…

Cited by 0SourcecodeScholar
2024

AlphaFin: Benchmarking Financial Analysis with Retrieval-Augmented Stock-Chain Framework

COLING 2024main

The task of financial analysis primarily encompasses two key areas: stock trend prediction and the corresponding financial question answering. Currently, machine learning and deep learning algorithms (ML&DL) have been widely applied for stock trend predictions, leading to significant progress. Howev…

2024

Bio-RFX: Refining Biomedical Extraction via Advanced Relation Classification and Structural Constraints

EMNLP 2024main

The ever-growing biomedical publications magnify the challenge of extracting structured data from unstructured texts. This task involves two components: biomedical entity identification (Named Entity Recognition, NER) and their interrelation determination (Relation Extraction, RE). However, existing…

2024

DanceCamera3D: 3D Camera Movement Synthesis with Music and Dance

CVPR 2024poster

Choreographers determine what the dances look like while cameramen determine the final presentation of dances. Recently various methods and datasets have showcased the feasibility of dance synthesis. However camera movement synthesis with music and dance remains an unsolved challenging problem due t…

2024

FlexKBQA: A Flexible LLM-Powered Framework for Few-Shot Knowledge Base Question Answering

AAAI 2024technical

Knowledge base question answering (KBQA) is a critical yet challenging task due to the vast number of entities within knowledge bases and the diversity of natural language questions posed by users. Unfortunately, the performance of most KBQA models tends to decline significantly in real-world scenar…

2024

Multiple Knowledge-Enhanced Interactive Graph Network for Multimodal Conversational Emotion Recognition

EMNLP 2024finding

Multimodal Emotion Recognition in Conversations (ERC) aims to identify emotions in conversational videos. Current efforts focus on modeling both context-sensitive and speaker-sensitive dependencies and multimodal fusion. Despite the progress, models in Multimodal ERC (MERC) still struggle due to a l…

Cited by 1SourcePDFScholar
2024

PatchFusion: An End-to-End Tile-Based Framework for High-Resolution Monocular Metric Depth Estimation

CVPR 2024poster

Single image depth estimation is a foundational task in computer vision and generative modeling. However prevailing depth estimation models grapple with accommodating the increasing resolutions commonplace in today's consumer cameras and devices. Existing high-resolution strategies show promise but…

2023

BEVDistill: Cross-Modal BEV Distillation for Multi-View 3D Object Detection

ICLR 2023poster

3D object detection from multiple image views is a fundamental and challenging task for visual scene understanding. Owing to its low cost and high efficiency, multi-view 3D object detection has demonstrated promising application prospects. However, accurately detecting objects through perspective vi…

2023

Learning from Noisy Data for Semi-Supervised 3D Object Detection

ICCV 2023poster

Pseudo-Labeling (PL) is a critical approach in semi-supervised 3D object detection (SSOD). In PL, delicately selected pseudo-labels, generated by the teacher model, are provided for the student model to supervise the semi-supervised detection framework. However, such a paradigm may introduce misclas…

Cited by 15PDFcodeScholar
2022

AutoAlign: Pixel-Instance Feature Aggregation for Multi-Modal 3D Object Detection

IJCAI 2022poster

Object detection through either RGB images or the LiDAR point clouds has been extensively explored in autonomous driving. However, it remains challenging to make these two data sources complementary and beneficial to each other. In this paper, we propose AutoAlign, an automatic feature fusion strat…

Cited by 140SourcePDFScholar
2022

Deformable Feature Aggregation for Dynamic Multi-modal 3D Object Detection

ECCV 2022poster

"Point clouds and RGB images are two general perceptional sources in autonomous driving. The former can provide accurate localization of objects, and the latter is denser and richer in semantic information. Recently, AutoAlign presents a learnable paradigm in combining these two modalities for 3D ob…

2022

Not Just Plain Text! Fuel Document-Level Relation Extraction with Explicit Syntax Refinement and Subsentence Modeling

EMNLP 2022finding

Document-level relation extraction (DocRE) aims to identify semantic labels among entities within a single document. One major challenge of DocRE is to dig decisive details regarding a specific entity pair from long text. However, in many cases, only a fraction of text carries required information,…

Cited by 7SourcePDFScholar
2022

SimIPU: Simple 2D Image and 3D Point Cloud Unsupervised Pre-training for Spatial-Aware Visual Representations

AAAI 2022technical

Pre-training has become a standard paradigm in many computer vision tasks. However, most of the methods are generally designed on the RGB image domain. Due to the discrepancy between the two-dimensional image plane and the three-dimensional space, such pre-trained models fail to perceive spatial inf…

2022

Unsupervised Domain Adaptation for Monocular 3D Object Detection via Self-Training

ECCV 2022poster

"Monocular 3D object detection (Mono3D) has achieved unprecedented success with the advent of deep learning techniques and emerging large-scale autonomous driving datasets. However, drastic performance degradation remains an unwell-studied challenge for practical cross-domain deployment as the lack…

2021

Motion-Aware Robotic 3D Ultrasound

ICRA 2021poster

Robotic three-dimensional (3D) ultrasound (US) imaging has been employed to overcome the drawbacks of traditional US examinations, such as high inter-operator variability and lack of repeatability. However, object movement remains a challenge as unexpected motion decreases the quality of the 3D comp…

Cited by 29SourceScholar