← Search

Zhen Li

133 accepted papers

2026

Accelerating Benchmarking of Functional Connectivity Modeling via Structure-aware Core-set Selection

ICLR 2026poster

Benchmarking the hundreds of functional connectivity (FC) modeling methods on large-scale fMRI datasets is critical for reproducible neuroscience. However, the combinatorial explosion of model–data pairings makes exhaustive evaluation computationally prohibitive, preventing such assessments from bec…

Cited by 0SourcecodeScholar
2026

BPDQ: Bit-Plane Decomposition Quantization on a Variable Grid for Large Language Models

ICML 2026poster

Large language model (LLM) inference is often bounded by memory footprint and memory bandwidth in resource-constrained deployments, making quantization a fundamental technique for efficient serving. While post-training quantization (PTQ) maintains high fidelity at 4-bit, it deteriorates at 2–3 bits.…

Cited by 0SourceScholar
2026

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary

ICLR 2026poster

Reward models trained on human preference data have demonstrated strong effectiveness in aligning Large Language Models (LLMs) with human intent under the framework of Reinforcement Learning from Human Feedback (RLHF). However, RLHF remains vulnerable to reward hacking, where the policy exploits imp…

Cited by 0SourceScholar
2026

Cancer Survival Prediction by Cyclic Generation and Multi-grained Alignment

AAAI 2026technical

Cancer survival analysis with multimodal data is crucial for precise treatments and patient benefits. However, the following challenges prohibit integrating histopathology and genomics: (i) multimodal data is not always complete, especially for the more costly genomics data; (ii) intricate interacti

Cited by 0SourcePDFScholar
2026

ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation

ICLR 2026poster

Infographic charts are a powerful medium for communicating abstract data by combining visual elements (e.g., charts, images) with textual information. However, their visual and structural richness poses challenges for large vision-language models (LVLMs), which are typically trained on plain charts.…

Cited by 0SourcecodeScholar
2026

Composition-Incremental Learning for Compositional Generalization

AAAI 2026technical

Compositional generalization has achieved substantial progress in computer vision on pre-collected training data. Nonetheless, real-world data continually emerges, with possible compositions being nearly infinite, long-tailed, and not entirely visible. Thus, an ideal model is supposed to gradually i

Cited by 0SourcePDFScholar
2026

DLWM: Dual Latent World Models enable Holistic Gaussian-centric Pre-training in Autonomous Driving

CVPR 2026

Vision-based autonomous driving has gained much attention due to its low costs and excellent performance. Compared with dense BEV (Bird's Eye View) or sparse query models, Gaussian-centric method is a comprehensive yet sparse representation by describing scene with 3D semantic Gaussians. In this pap

Cited by 0SourceScholar
2026

DSSA: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation

ICLR 2026poster

Long-sequence processing is a critical capability for modern large language models. However, the self-attention mechanism in the standard Transformer architecture faces severe computational and memory bottlenecks when processing long sequences. While trainable sparse attention methods offer a promis…

Cited by 0SourceScholar
2026

Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield

ICLR 2026poster

Diffusion model distillation has emerged as a powerful technique for creating efficient few-step and single-step generators. Among these, Distribution Matching Distillation (DMD) and its variants stand out for their impressive performance, which is widely attributed to their core mechanism of matchi…

Cited by 0SourceScholar
2026

DriveFlow: Rectified Flow Adaptation for Robust 3D Object Detection in Autonomous Driving

AAAI 2026technical

In autonomous driving, vision-centric 3D object detection recognizes and localizes 3D objects from RGB images. However, due to high annotation costs and diverse outdoor scenes, training data often fails to cover all possible test scenarios, known as the out-of-distribution (OOD) issue. Training-free

Cited by 0SourcePDFScholar
2026

Flexibility-Aware Geometric Latent Diffusion for Full-Atom Peptide Design

ICML 2026poster

Although peptides are well suited for flexible and shallow binding interfaces, their intrinsic flexibility induces a strongly coupled sequence–structure relationship that current fixed-geometry latent models cannot simultaneously model with conformational diversity and physical feasibility, ultimate…

Cited by 0SourceScholar
2026

Hybrid Model-Learning Decoupled Control for Tendon-Driven Multi-Segment Continuum Robotic Bronchoscope

ICRA 2026poster

Flexible tendon-driven multi-segment robotic bronchoscopes can reach peripheral lung regions for minimally invasive diagnosis and therapy. However, long tendon transmissions introduce friction, elasticity, and backlash, which couple the motion of adjacent segments and reduce operational accuracy and…

Cited by 0Scholar
2026

Learning Adaptive Topology with FiLM-Guided Distillation for Tertiary Structure-Based RNA Design

ICML 2026poster

Tertiary structure-based RNA design aims to generate RNA sequences that can fold into desired 3D structures, but remains a challenging problem due to the scarcity of annotated data, structural noise, and the intrinsic complexity of RNA topology. Existing structure-to-sequence frameworks largely rely…

Cited by 0SourceScholar
2026

MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models

AAAI 2026technical

Multimodal large language models (MLLMs), which integrate language and visual cues for problem-solving, are crucial for advancing artificial general intelligence (AGI). However, current benchmarks for measuring the intelligence of MLLMs suffer from limited scale, narrow coverage, and unstructured kn

Cited by 0SourcePDFScholar
2026

Unified Map Prior Encoder for Mapping and Planning

ICRA 2026poster

Online mapping and end-to-end (E2E) planning in autonomous driving are still largely sensor-centric, leaving rich map priors—HD/SD vector maps, rasterized SD maps, and satellite imagery—underused due to heterogeneity, pose drift, and inconsistent availability at test time. We present emph{UMPE}, a U…

2026

Yume1.5: A Text-Controlled Interactive World Generation Model

CVPR 2026

Recent approaches have demonstrated the promise of using diffusion models to generate interactive and explorable worlds. However, most of these methods face critical challenges such as excessively large parameter sizes, reliance on lengthy inference steps, and rapidly growing historical context, whi

Cited by 0SourcecodeScholar
2025

3D Gaussian Flats: Hybrid 2D/3D Photometric Scene Reconstruction

NeurIPS 2025poster

Recent advances in radiance fields and novel view synthesis enable creation of realistic digital twins from photographs. However, current methods struggle with flat, texture-less surfaces, creating uneven and semi-transparent reconstructions, due to an ill-conditioned photometric reconstruction obje…

Cited by 0SourceScholar
2025

A General Framework to Enhance Fine-tuning-based LLM Unlearning

ACL 2025finding

Unlearning has been proposed to remove copyrighted and privacy-sensitive data from Large Language Models (LLMs). Existing approaches primarily rely on fine-tuning-based methods, which can be categorized into gradient ascent-based (GA-based) and suppression-based methods. However, they often degrade…

2025

ANASETC: Automatic Neural Architecture Search for Encrypted Traffic Classification

ICASSP 2025accepted

The widespread adoption of encrypted network protocols has made traffic encryption ubiquitous, creating substantial challenges for network management and security. This paper introduces a novel encrypted traffic classification system, ANASETC, which combines traffic burst features with Neural Archit…

Cited by 0SourceScholar
2025

AR-1-to-3: Single Image to Consistent 3D Object via Next-View Prediction

ICCV 2025poster

Novel view synthesis (NVS) is a cornerstone for image-to-3d creation. However, existing works still struggle to maintain consistency between the generated views and the input views, especially when there is a significant camera pose difference, leading to poor-quality 3D geometries and textures. We…

2025

Advancing Dense Endoscopic Reconstruction with Gaussian Splatting-Driven Surface Normal-Aware Tracking and Mapping

ICRA 2025

Simultaneous Localization and Mapping (SLAM) is essential for precise surgical interventions and robotic tasks in minimally invasive procedures. While recent advancements in 3D Gaussian Splatting (3DGS) have improved SLAM with high-quality novel view synthesis and fast rendering, these systems strug

Cited by 18SourcecodeScholar
2025

AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language Models

NeurIPS 2025poster

Effective human-agent collaboration in physical environments requires understanding not only what to act upon, but also where the actionable elements are and how to interact with them. Existing approaches often operate at the object level or disjointedly handle fine-grained affordance reasoning, lac…

Cited by 0SourceScholar
2025

AgentTTS: Large Language Model Agent for Test-time Compute-optimal Scaling Strategy in Complex Tasks

NeurIPS 2025poster

Test-time scaling (TTS) enhances the performance of large language models (LLMs) by allocating additional compute resources during inference. However, existing research primarily investigates TTS in single-stage tasks; while many real-world problems are multi-stage complex tasks, composed of a seque…

Cited by 0SourceScholar
2025

CLEA: Closed-Loop Embodied Agent for Enhancing Task Execution in Dynamic Environments

IROS 2025

Large Language Models (LLMs) exhibit remarkable capabilities in the hierarchical decomposition of complex tasks through semantic reasoning. However, their application in embodied systems faces challenges in ensuring reliable execution of subtask sequences and achieving one-shot success in long-term

Cited by 5SourcecodeScholar
2025

Consistency of Compositional Generalization Across Multiple Levels

AAAI 2025technical

Compositional generalization is the capability of a model to understand novel compositions composed of seen concepts. There are multiple levels of novel compositions including phrase-phrase level, phrase-word level, and word-word level. Existing methods achieve promising compositional generalization…

2025

DCP: Dual-Cue Pruning for Efficient Large Vision-Language Models

EMNLP 2025

Large Vision-Language Models (LVLMs) achieve remarkable performance in multimodal tasks but suffer from high computational costs due to the large number of visual tokens. Existing pruning methods either apply after visual tokens enter the LLM or perform pre-pruning based solely on visual attention.

Cited by 0SourcePDFScholar
2025

DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering

CVPR 2025poster

3D Question Answering (3D QA) requires the model to comprehensively understand its situated 3D scene described by the text, then reason about its surrounding environment and answer a question under that situation. However, existing methods usually rely on global scene perception from pure 3D point c…

2025

DriveGEN: Generalized and Robust 3D Detection in Driving via Controllable Text-to-Image Diffusion Generation

CVPR 2025poster

In autonomous driving, vision-centric 3D detection aims to identify 3D objects from images. However, high data collection costs and diverse real-world scenarios limit the scale of training data. Once distribution shifts occur between training and test data, existing methods often suffer from perform…

2025

Dynamic Action Localization and Recognition for Intelligent Perception of Surgical Robots

IROS 2025

Robot-assisted surgery has significantly advanced surgical precision, yet the development of autonomous surgical robots remains hindered by their limited understanding of complex surgical actions. Current systems lack the ability to effectively perceive and interpret intricate surgical relationships

Cited by 0SourceScholar
2025

ETSM: Automating Dissection Trajectory Suggestion and Confidence Map-Based Safety Margin Prediction for Robot-Assisted Endoscopic Submucosal Dissection

ICRA 2025

Robot-assisted Endoscopic Submucosal Dissection (ESD) improves the surgical procedure by providing a more comprehensive view through advanced robotic instruments and bimanual operation, thereby enhancing dissection efficiency and accuracy. Accurate prediction of dissection trajectories is crucial fo

Cited by 3SourcecodeScholar
2025

Empowering Large Language Models with 3D Situation Awareness

CVPR 2025poster

Driven by the great success of Large Language Models (LLMs) in the 2D image domain, their applications in 3D scene understanding has emerged as a new trend. A key difference between 3D and 2D is that the situation of an egocentric observer in 3D scenes can change, resulting in different descriptions…

Cited by 0SourcePDFScholar
2025

GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding

ICLR 2025poster

Recently, Multimodal Large Language Models (MLLMs) have been used as agents to control keyboard and mouse inputs by directly perceiving the Graphical User Interface (GUI) and generating corresponding commands. However, current agents primarily demonstrate strong understanding capabilities in static…

2025

High-Precision Tracking of Time-Varying Trajectories for Microsurgical Robots in Constrained Environments

IROS 2025

This research addresses the challenge of achieving high-precision tracking of time-varying trajectories under nonlinear disturbances and motion constraints in microsurgical robots. A hybrid control framework integrating fuzzy adaptive sliding mode control with radial basis function neural networks i

Cited by 0SourceScholar
2025

Implicit Disparity-Blur Alignment for Fast and Precise Autofocus in Robotic Microsurgical Imaging

IROS 2025

Creating an intelligent surgical environment requires not only advanced robotic systems but also optimized microscopic imaging. However, autofocus remains a fundamental challenge, with current methods suffering from slow iterative processes or directional ambiguity, which compromises real-time perfo

Cited by 0SourceScholar
2025

Lumina-Image 2.0: A Unified and Efficient Image Generative Framework

ICCV 2025poster

We introduce Lumina-Image 2.0, an advanced text-to-image (T2I) model that surpasses previous state-of-the-art methods across multiple benchmarks. Lumina-Image 2.0 is characterized by two key features: (1) Unification - it adopts a unified architecture (Unified Next-DiT) that treats text and image to…

2025

Multi-Sourced Compositional Generalization in Visual Question Answering

IJCAI 2025

Compositional generalization is the ability of generalizing novel compositions from seen primitives, and has received much attention in vision-and-language (V&L) recently. Due to the multi-modal nature of V&L tasks, the primitives composing compositions source from different modalities, resulting in

2025

Reduced-Dimensional Whole-Body Control Based on Model Simplification for Bipedal Robots With Parallel Mechanisms

RA-L 2025

The presence of parallel mechanisms in bipedal robots increases the complexity of modeling and control, making it crucial to manage the trade-off between model accuracy and real-time control. In this letter, we propose a reduced-dimensional whole-body controller for series-parallel bipedal robots, u

Cited by 7SourceScholar
2025

SQS: Enhancing Sparse Perception Models via Query-based Splatting in Autonomous Driving

NeurIPS 2025spotlight

Sparse Perception Models (SPMs) adopt a query-driven paradigm that forgoes explicit dense BEV or volumetric construction, enabling highly efficient computation and accelerated inference. In this paper, we introduce SQS, a novel query-based splatting pre-training specifically designed to advance SPMs…

Cited by 0SourceScholar
2025

Sekai: A Video Dataset towards World Exploration

NeurIPS 2025poster

Video generation techniques have made remarkable progress, promising to be the foundation of interactive world exploration. However, existing video generation datasets are not well-suited for world exploration training as they suffer from some limitations: limited locations, short duration, static s…

Cited by 0SourceScholar
2025

Sign2Vis: Automated Data Visualization from Sign Language

ACL 2025finding

Data visualizations, such as bar charts and histograms, are essential for analyzing and exploring data, enabling the effective communication of insights. While existing methods have been proposed to translate natural language descriptions into visualization queries, they focus solely on spoken langu…

2025

SkyVLN: Vision-and-Language Navigation and NMPC Control for UAVs in Urban Environments

IROS 2025

Unmanned Aerial Vehicles (UAVs) have emerged as versatile tools across various sectors, driven by their mobility and adaptability. This paper introduces SkyVLN, a novel framework integrating vision-and-language navigation (VLN) with Nonlinear Model Predictive Control (NMPC) to enhance UAV autonomy i

Cited by 12SourceScholar
2025

Spatiotemporal Motion Prediction of Intraocular Microsurgical Robot in Non-Visible Regions

IROS 2025

In intraocular microsurgery with minute operational scales, instruments pass through non-visible regions of the anterior segment, where robot-assisted surgery, which heavily relies on visual perception, fails to determine the instrument’s attitude relative to the eyeball. This compromises surgical f

Cited by 0SourceScholar
2025

Stepwise Perplexity-Guided Refinement for Efficient Chain-of-Thought Reasoning in Large Language Models

ACL 2025finding

Chain-of-Thought (CoT) reasoning, which breaks down complex tasks into intermediate reasoning steps, has significantly enhanced the performance of large language models (LLMs) on challenging tasks. However, the detailed reasoning process in CoT often incurs long generation times and high computation…

Cited by 0SourcePDFScholar
2025

Topo2Seq: Enhanced Topology Reasoning via Topology Sequence Learning

AAAI 2025technical

Extracting lane topology from perspective views (PV) is crucial for planning and control in autonomous driving. This approach extracts potential drivable trajectories for self-driving vehicles without relying on high-definition (HD) maps. However, the unordered nature and weak long-range perception…

Cited by 0SourcePDFScholar
2025

Unifying Within and Across: Intra-Modality Multi-View Fusion and Inter-Modality Alignment for Knowledge Graph Completion

ICASSP 2025accepted

Multi-modal knowledge graph completion (MMKGC) enhances the structural and semantic richness of knowledge graphs by integrating diverse information across modalities. However, existing methods often either overlook the diversity within a single modality or fail to ensure effective cross-modality ali…

Cited by 0SourceScholar
2025

VQA4CIR: Boosting Composed Image Retrieval with Visual Question Answering

AAAI 2025technical

Albeit progress has been made in Composed Image Retrieval (CIR), we empirically find that a certain percentage of failure retrieval results are not consistent with their relative captions. To address this issue, this work provides a Visual Question Answering (VQA) perspective to boost the performanc…

2025

VisionPAD: A Vision-Centric Pre-training Paradigm for Autonomous Driving

CVPR 2025poster

This paper introduces VisionPAD, a novel self-supervised pre-training paradigm designed for vision-centric algorithms in autonomous driving. In contrast to previous approaches that employ neural rendering with explicit depth supervision, VisionPAD utilizes more efficient 3D Gaussian Splatting to rec…

Cited by 2SourcePDFScholar
2025

VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning

ICCV 2025poster

Recent advances in diffusion models have significantly advanced image generation; however, existing models remain task-specific, limiting their efficiency and generalizability. While universal models attempt to address these limitations, they face critical challenges, including generalizable instruc…

2024

A Hybrid Admittance Control Algorithm for Automatic Robotic Cranium-Milling

ICRA 2024poster

Prior robot-assisted cranium-milling studies only considered controlling the force in the skull’s vertical direction and neglected the milling cutter’s feed force. Additionally, achieving stable force control in multiple directions is challenging for robots due to the uneven skull surface. Here a hy…

Cited by 0SourceScholar
2024

CReStyler: Text-Guided Single Image Style Transfer Method Based on CNN and Restormer

ICASSP 2024accepted

Text-guided image style transfer methods have gradually become a research hotspot. However, existing text-guided style transfer method suffers from content information missing and artifacts in the generated stylized images. Therefore, we propose CReStyler, a text-guided image style method based on t…

Cited by 0SourceScholar
2024

Chained Flexible Capsule Endoscope: Unraveling the Conundrum of Size Limitations and Functional Integration for Gastrointestinal Transitivity

ICRA 2024poster

Capsule endoscopes, predominantly serving diagnostic functions, provide lucid internal imagery but are devoid of surgical or therapeutic capabilities. Consequently, despite lesion detection, physicians frequently resort to traditional endoscopic or open surgical procedures for treatment, resulting i…

Cited by 1SourceScholar
2024

Compositional Substitutivity of Visual Reasoning for Visual Question Answering

ECCV 2024poster

"Compositional generalization has received much attention in vision-and-language and visual reasoning recently. Substitutivity, the capability to generalize to novel compositions with synonymous primitives such as words and visual entities, is an essential factor in evaluating the compositional gene…

2024

CrossBind: Collaborative Cross-Modal Identification of Protein Nucleic-Acid-Binding Residues

AAAI 2024technical

Accurate identification of protein nucleic acid binding residues poses a significant challenge with important implications for various biological processes and drug design. Many typical computational methods for protein analysis rely on a single model that could ignore either the semantic context of…

2024

DV-3DLane: End-to-end Multi-modal 3D Lane Detection with Dual-view Representation

ICLR 2024poster

Accurate 3D lane estimation is crucial for ensuring safety in autonomous driving. However, prevailing monocular techniques suffer from depth loss and lighting variations, hampering accurate 3D lane detection. In contrast, LiDAR points offer geometric cues and enable precise localization. In this pap…

2024

Design and Modeling of a Thin-walled Multi-segment Continuum Robotic Bronchoscope

IROS 2024poster

Cable-driven continuum robots in bronchoscopic procedures hold immense potential to revolutionize the diagnosis and treatment of lung cancer. However, robotic bronchoscopes in current studies are typically large in size and inflexible. Therefore, this article introduces a novel cable-driven continuu…

Cited by 0SourceScholar
2024

In-Context Compositional Generalization for Large Vision-Language Models

EMNLP 2024main

Recent work has revealed that in-context learning for large language models exhibits compositional generalization capacity, which can be enhanced by selecting in-context demonstrations similar to test cases to provide contextual information. However, how to exhibit in-context compositional generaliz…

Cited by 2SourcePDFScholar
2024

Leveraging Large Language Models for NLG Evaluation: Advances and Challenges

EMNLP 2024main

In the rapidly evolving domain of Natural Language Generation (NLG) evaluation, introducing Large Language Models (LLMs) has opened new avenues for assessing generated content quality, e.g., coherence, creativity, and context relevance. This paper aims to provide a thorough overview of leveraging LL…

2024

Magnetic-Guided Flexible Origami Robot toward Long-Term Phototherapy of H. pylori in the Stomach

ICRA 2024poster

Helicobacter pylori, a pervasive bacterial infection associated with gastrointestinal disorders such as gastritis, peptic ulcer disease, and gastric cancer, impacts approximately 50% of the global population. The efficacy of standard clinical eradication therapies is diminishing due to the rise of a…

Cited by 0SourceScholar
2024

PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding

CVPR 2024poster

Recent advances in text-to-image generation have made remarkable progress in synthesizing realistic human photos conditioned on given text prompts. However existing personalized generation methods cannot simultaneously satisfy the requirements of high efficiency promising identity (ID) fidelity and…

2024

Procedure Recognition by Knowledge-Driven Segmentation in Robotic-Assisted Vitreoretinal Surgery

ICRA 2024poster

Internal limiting membrane (ILM) peeling is a vital vitreoretinal surgery procedure. However, due to the thickness of just 1-2 micrometers and the intricacies associated with its varying density and adhesion, the difficulty of manipulation exceeds the physiological limits of human perception and ope…

Cited by 0SourceScholar
2024

RadOcc: Learning Cross-Modality Occupancy Knowledge through Rendering Assisted Distillation

AAAI 2024technical

3D occupancy prediction is an emerging task that aims to estimate the occupancy states and semantics of 3D scenes using multi-view images. However, image-based scene perception encounters significant challenges in achieving accurate prediction due to the absence of geometric priors. In this paper, w…

Cited by 21SourcePDFScholar
2024

SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge

NeurIPS 2024poster

Large vision-language models (LVLMs) are ignorant of the up-to-date knowledge, such as LLaVA series, because they cannot be updated frequently due to the large amount of resources required, and therefore fail in many cases. For example, if a LVLM was released on January 2024, and it wouldn't know th…

Cited by 2SourcePDFScholar
2024

Towards Flexible 3D Perception: Object-Centric Occupancy Completion Augments 3D Object Detection

NeurIPS 2024poster

While 3D object bounding box (bbox) representation has been widely used in autonomous driving perception, it lacks the ability to capture the precise details of an object's intrinsic geometry. Recently, occupancy has emerged as a promising alternative for 3D scene perception. However, constructing a…

2024

Unified Generation, Reconstruction, and Representation: Generalized Diffusion with Adaptive Latent Encoding-Decoding

ICML 2024poster

The vast applications of deep generative models are anchored in three core capabilities---*generating* new instances, *reconstructing* inputs, and learning compact *representations*---across various data types, such as discrete text/protein sequences and continuous images. Existing model families, l…

2024

Visual Programming for Zero-shot Open-Vocabulary 3D Visual Grounding

CVPR 2024poster

3D Visual Grounding (3DVG) aims at localizing 3D object based on textual descriptions. Conventional supervised methods for 3DVG often necessitate extensive annotations and a predefined vocabulary which can be restrictive. To address this issue we propose a novel visual programming approach for zero-…

2024

WeakPCSOD: Overcoming the Bias of Box Annotations for Weakly Supervised Point Cloud Salient Object Detection

AAAI 2024technical

Point cloud salient object detection (PCSOD) is a newly proposed task in 3D dense segmentation. However, the acquisition of accurate 3D dense annotations comes at a high cost, severely limiting the progress of PCSOD. To address this issue, we propose the first weakly supervised PCSOD (named WeakPCSO…

Cited by 3SourcePDFScholar
2024

X4D-SceneFormer: Enhanced Scene Understanding on 4D Point Cloud Videos through Cross-Modal Knowledge Transfer

AAAI 2024technical

The field of 4D point cloud understanding is rapidly developing with the goal of analyzing dynamic 3D point cloud sequences. However, it remains a challenging task due to the sparsity and lack of texture in point clouds. Moreover, the irregularity of point cloud poses a difficulty in aligning tempo…

2023

AMT: All-Pairs Multi-Field Transforms for Efficient Frame Interpolation

CVPR 2023poster

We present All-Pairs Multi-Field Transforms (AMT), a new network architecture for video frame interpolation. It is based on two essential designs. First, we build bidirectional correlation volumes for all pairs of pixels and use the predicted bilateral flows to retrieve correlations for updating bot…

2023

Amazon-M2: A Multilingual Multi-locale Shopping Session Dataset for Recommendation and Text Generation

NeurIPS 2023poster

Modeling customer shopping intentions is a crucial task for e-commerce, as it directly impacts user experience and engagement. Thus, accurately understanding customer preferences is essential for providing personalized recommendations. Session-based recommendation, which utilizes customer session d…

2023

Automated Key Action Detection for Closed Reduction of Pelvic Fractures by Expert Surgeons in Robot-Assisted Surgery

IROS 2023poster

Pelvic fractures are one of the most serious traumas in orthopedics, and the technical proficiency and expertise of the surgical team strongly influence the quality of reduction results. With the advancement of information technology and robotics, robot-assisted pelvic fracture reduction surgery is…

Cited by 0SourceScholar
2023

BEV@DC: Bird's-Eye View Assisted Training for Depth Completion

CVPR 2023poster

Depth completion plays a crucial role in autonomous driving, in which cameras and LiDARs are two complementary sensors. Recent approaches attempt to exploit spatial geometric constraints hidden in LiDARs to enhance image-guided depth completion. However, only low efficiency and poor generalization c…

Cited by 30SourcePDFScholar
2023

CowClip: Reducing CTR Prediction Model Training Time from 12 Hours to 10 Minutes on 1 GPU

AAAI 2023technical

The click-through rate (CTR) prediction task is to predict whether a user will click on the recommended item. As mind-boggling amounts of data are produced online daily, accelerating CTR prediction model training is critical to ensuring an up-to-date model and reducing the training cost. One approac…

2023

DNF: Decouple and Feedback Network for Seeing in the Dark

CVPR 2023highlight

The exclusive properties of RAW data have shown great potential for low-light image enhancement. Nevertheless, the performance is bottlenecked by the inherent limitations of existing architectures in both single-stage and multi-stage methods. Mixed mapping across two different domains, noise-to-clea…

2023

Decoupled Non-Parametric Knowledge Distillation for end-to-End Speech Translation

ICASSP 2023accepted

Existing techniques often attempt to make knowledge transfer from a powerful machine translation (MT) to speech translation (ST) model with some elaborate techniques, which often requires transcription as extra input during training. However, transcriptions are not always available, and how to impro…

Cited by 0SourceScholar
2023

EasyGaze3D: Towards Effective and Flexible 3D Gaze Estimation from a Single RGB Camera

IROS 2023poster

Eye gaze can convey rich information of human intentions, which enables the social robots to comprehend the cognition and behavior of human targets. However, the existing 3D gaze estimation methods generally have high requirements either on the dedicated hardware or the quantity and quality of train…

Cited by 3SourceScholar
2023

Exploring the Effect of Primitives for Compositional Generalization in Vision-and-Language

CVPR 2023poster

Compositionality is one of the fundamental properties of human cognition (Fodor & Pylyshyn, 1988). Compositional generalization is critical to simulate the compositional capability of humans, and has received much attention in the vision-and-language (V&L) community. It is essential to understand th…

2023

FAA: Fine-grained Attention Alignment for Cascade Document Ranking

ACL 2023long

Document ranking aims at sorting a collection of documents with their relevance to a query. Contemporary methods explore more efficient transformers or divide long documents into passages to handle the long input. However, intensive query-irrelevant content may lead to harmful distraction and high q…

Cited by 4SourcePDFScholar
2023

Fair-CDA: Continuous and Directional Augmentation for Group Fairness

AAAI 2023technical

In this work, we propose Fair-CDA, a fine-grained data augmentation strategy for imposing fairness constraints. We use a feature disentanglement method to extract the features highly related to the sensitive attributes. Then we show that group fairness can be achieved by regularizing the models on t…

Cited by 3SourcePDFScholar
2023

Geometry-Aware Network for Domain Adaptive Semantic Segmentation

AAAI 2023technical

Measuring and alleviating the discrepancies between the synthetic (source) and real scene (target) data is the core issue for domain adaptive semantic segmentation. Though recent works have introduced depth information in the source domain to reinforce the geometric and semantic knowledge transfer,…

Cited by 6SourcePDFScholar
2023

LATR: 3D Lane Detection from Monocular Images with Transformer

ICCV 2023oral

3D lane detection from monocular images is a fundamental yet challenging task in autonomous driving. Recent advances primarily rely on structural 3D surrogates (e.g., bird's eye view) built from front-view image features and camera parameters. However, the depth ambiguity in monocular images inevita…

Cited by 42PDFcodeScholar
2023

Learning Transformation-Predictive Representations for Detection and Description of Local Features

CVPR 2023poster

The task of key-points detection and description is to estimate the stable location and discriminative representation of local features, which is essential for image matching. However, either the rough hard positive or negative labels generated from one-to-one correspondences among images bring indi…

Cited by 11SourcePDFScholar
2023

MMTN: Multi-Modal Memory Transformer Network for Image-Report Consistent Medical Report Generation

AAAI 2023technical

Automatic medical report generation is an essential task in applying artificial intelligence to the medical domain, which can lighten the workloads of doctors and promote clinical automation. The state-of-the-art approaches employ Transformer-based encoder-decoder architectures to generate reports f…

2023

Multi-View Inverse Rendering for Large-Scale Real-World Indoor Scenes

CVPR 2023poster

We present a efficient multi-view inverse rendering method for large-scale real-world indoor scenes that reconstructs global illumination and physically-reasonable SVBRDFs. Unlike previous representations, where the global illumination of large scenes is simplified as multiple environment maps, we p…

2023

RankMatch: Fostering Confidence and Consistency in Learning with Noisy Labels

ICCV 2023poster

Learning with noisy labels (LNL) is one of the most important and challenging problems in weakly-supervised learning. Recent advances adopt the sample selection strategy to mitigate the interference of noisy labels and use small-loss criteria to select clean samples. However, the one-dimensional los…

Cited by 14PDFScholar
2023

SRFormer: Permuted Self-Attention for Single Image Super-Resolution

ICCV 2023poster

Previous works have shown that increasing the window size for Transformer-based image super-resolution models (e.g., SwinIR) can significantly improve the model performance but the computation overhead is also considerable. In this paper, we present SRFormer, a simple but novel method that can enjoy…

Cited by 216PDFcodeScholar
2023

Semantic Human Parsing via Scalable Semantic Transfer Over Multiple Label Domains

CVPR 2023poster

This paper presents Scalable Semantic Transfer (SST), a novel training paradigm, to explore how to leverage the mutual benefits of the data from different label domains (i.e. various levels of label granularity) to train a powerful human parsing network. In practice, two common application scenarios…

2023

SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-training

ICCV 2023poster

Skeleton sequence representation learning has shown great advantages for action recognition due to its promising ability to model human joints and topology. However, the current methods usually require sufficient labeled data for training computationally expensive models. Moreover, these methods ign…

Cited by 59PDFcodeScholar
2023

Small Total-Cost Constraints in Contextual Bandits with Knapsacks, with Application to Fairness

NeurIPS 2023poster

We consider contextual bandit problems with knapsacks [CBwK], a problem where at each round, a scalar reward is obtained and vector-valued costs are suffered. The learner aims to maximize the cumulative rewards while ensuring that the cumulative costs are lower than some predetermined cost constrain…

Cited by 2SourcePDFScholar
2023

SupFusion: Supervised LiDAR-Camera Fusion for 3D Object Detection

ICCV 2023poster

LiDAR-Camera fusion-based 3D detection is a critical task for automatic driving. In recent years, many LiDAR-Camera fusion approaches sprung up and gained promising performances compared with single-modal detectors, but always lack carefully designed and effective supervision for the fusion process.…

Cited by 19PDFcodeScholar
2022

2DPASS: 2D Priors Assisted Semantic Segmentation on LiDAR Point Clouds

ECCV 2022poster

"As camera and LiDAR sensors capture complementary information used in autonomous driving, great efforts have been made to develop semantic segmentation algorithms through multi-modality data fusion. However, fusion-based approaches require paired data, i.e., LiDAR point clouds and camera images wit…

2022

AMOS: A Large-Scale Abdominal Multi-Organ Benchmark for Versatile Medical Image Segmentation

NeurIPS 2022accept

Despite the considerable progress in automatic abdominal multi-organ segmentation from CT/MRI scans in recent years, a comprehensive evaluation of the models' capabilities is hampered by the lack of a large-scale benchmark from diverse clinical scenarios. Constraint by the high cost of collecting an…

2022

Beyond 3D Siamese Tracking: A Motion-Centric Paradigm for 3D Single Object Tracking in Point Clouds

CVPR 2022oral

3D single object tracking (3D SOT) in LiDAR point clouds plays a crucial role in autonomous driving. Current approaches all follow the Siamese paradigm based on appearance matching. However, LiDAR point clouds are usually textureless and incomplete, which hinders effective appearance matching. Besid…

Cited by 109PDFcodeScholar
2022

CLMLF:A Contrastive Learning and Multi-Layer Fusion Method for Multimodal Sentiment Detection

NAACL 2022findings

Compared with unimodal data, multimodal data can provide more features to help the model analyze the sentiment of data. Previous research works rarely consider token-level feature fusion, and few works explore learning the common features related to sentiment in multimodal data to help the model fus…

2022

Contact-Distil: Boosting Low Homologous Protein Contact Map Prediction by Self-Supervised Distillation

AAAI 2022technical

Accurate protein contact map prediction (PCMP) is essential for precise protein structure estimation and further biological studies. Recent works achieve significant performance on this task with high quality multiple sequence alignment (MSA). However, the PCMP accuracy drops dramatically while only…

2022

Divide and Contrast: Source-free Domain Adaptation via Adaptive Contrastive Learning

NeurIPS 2022accept

We investigate a practical domain adaptation task, called source-free domain adaptation (SFUDA), where the source pretrained model is adapted to the target domain without access to the source data. Existing techniques mainly leverage self-supervised pseudo-labeling to achieve class-wise global align…

2022

Don’t Take It Literally: An Edit-Invariant Sequence Loss for Text Generation

NAACL 2022long

Neural text generation models are typically trained by maximizing log-likelihood with the sequence cross entropy (CE) loss, which encourages an exact token-by-token match between a target sequence with a generated sequence. Such training objective is sub-optimal when the target sequence is not perfe…

2022

FCGCL: Fine- and Coarse-Granularity Contrastive Learning for Speech Translation

EMNLP 2022finding

It is notoriously difficult to implement end-to-end speech translation (E2E-ST) model because of the task complexity and data scarcity. Existing techniques often attempt to carry out implicit knowledge transfer from machine translation (MT) to ST model by imposing various constraints. However, in th…

2022

Graph Enhanced Contrastive Learning for Radiology Findings Summarization

ACL 2022long

The impression section of a radiology report summarizes the most prominent observation from the findings section and is the most important section for radiologists to communicate to physicians. Summarizing findings is time-consuming and can be prone to error for inexperienced radiologists, and thus…

2022

Let Images Give You More: Point Cloud Cross-Modal Training for Shape Analysis

NeurIPS 2022accept

Although recent point cloud analysis achieves impressive progress, the paradigm of representation learning from single modality gradually meets its bottleneck. In this work, we take a step towards more discriminative 3D point cloud representation using 2D images, which inherently contain richer appe…

2022

PhyIR: Physics-Based Inverse Rendering for Panoramic Indoor Images

CVPR 2022poster

Inverse rendering of complex material such as glossy, metal and mirror material is a long-standing ill-posed problem in this area, which has not been well solved. Previous approaches cannot tackle them well due to simplified BRDF and unsuitable illumination representations. In this paper, we present…

Cited by 26PDFcodeScholar
2022

Reciprocal Learning of Knowledge Retriever and Response Ranker for Knowledge-Grounded Conversations

COLING 2022main

Grounding dialogue agents with knowledge documents has sparked increased attention in both academia and industry. Recently, a growing body of work is trying to build retrieval-based knowledge-grounded dialogue systems. While promising, these approaches require collecting pairs of dialogue context an…

Cited by 4SourcePDFScholar
2022

Towards an End-to-End Framework for Flow-Guided Video Inpainting

CVPR 2022poster

Optical flow, which captures motion information across frames, is exploited in recent video inpainting methods through propagating pixels along its trajectories. However, the hand-crafted flow-based processes in these methods are applied separately to form the whole inpainting pipeline. Thus, they a…

Cited by 186PDFcodeScholar
2022

Weakly Supervised Object Localization through Inter-class Feature Similarity and Intra-Class Appearance Consistency

ECCV 2022poster

"Weakly supervised object localization (WSOL) aims at detecting objects through only image-level labels. Class activation maps (CAMs) are the commonly used features for WSOL. However, existing CAM-based methods tend to excessively pursue discriminative features for object recognition and hence ignor…

Cited by 14SourcePDFScholar
2022

X-Trans2Cap: Cross-Modal Knowledge Transfer Using Transformer for 3D Dense Captioning

CVPR 2022poster

3D dense captioning aims to describe individual objects by natural language in 3D scenes, where 3D scenes are usually represented as RGB-D scans or point clouds. However, only exploiting single modal information, e.g., point cloud, previous approaches fail to produce faithful descriptions. Though ag…

Cited by 96PDFcodeScholar
2021

Adaptive Residue-wise Profile Fusion for Low Homologous Protein Secondary Structure Prediction Using External Knowledge

IJCAI 2021poster

Protein secondary structure prediction (PSSP) is essential for protein function analysis. However, for low homologous proteins, the PSSP suffers from insufficient input features. In this paper, we explicitly import external self-supervised knowledge for low homologous PSSP under the guidance of resi…

2021

Box-Aware Feature Enhancement for Single Object Tracking on Point Clouds

ICCV 2021poster

Current 3D single object tracking approaches track the target based on a feature comparison between the target template and the search area. However, due to the common occlusion in LiDAR scans, it is non-trivial to conduct accurate feature comparisons on severe sparse and incomplete shapes. In this…

Cited by 122PDFcodeScholar
2021

Free-Form Description Guided 3D Visual Graph Network for Object Grounding in Point Cloud

ICCV 2021poster

3D object grounding aims to locate the most relevant target object in a raw point cloud scene based on a free-form language description. Understanding complex and diverse descriptions, and lifting them directly to a point cloud is a new and challenging topic due to the irregular and sparse nature of…

Cited by 100PDFcodeScholar
2021

InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds Through Instance Multi-Level Contextual Referring

ICCV 2021poster

Compared with the visual grounding on 2D images, the natural-language-guided 3D object localization on point clouds is more challenging. In this paper, we propose a new model, named InstanceRefer, to achieve a superior 3D visual grounding through the grounding-by-matching strategy. In practice, our…

Cited by 147PDFcodeScholar
2021

Local Representation is Not Enough: Soft Point-Wise Transformer for Descriptor and Detector of Local Features

IJCAI 2021poster

Significant progress has been witnessed for the descriptor and detector of local features, but there still exist several challenging and intractable limitations, such as insufficient localization accuracy and non-discriminative description, especially in repetitive- or blank-texture regions, which h…

Cited by 12SourcePDFScholar
2021

PSSM-Distil: Protein Secondary Structure Prediction (PSSP) on Low-Quality PSSM by Knowledge Distillation with Contrastive Learning

AAAI 2021technical

Protein secondary structure prediction (PSSP) is an essential task in computational biology. To achieve the accurate PSSP, the general and vital feature engineering is to use multiple sequence alignment (MSA) for Position-Specific Scoring Matrix (PSSM) extraction. However, when only low-quality PSSM…

2021

PointLIE: Locally Invertible Embedding for Point Cloud Sampling and Recovery

IJCAI 2021poster

Point Cloud Sampling and Recovery (PCSR) is critical for massive real-time point cloud collection and processing since raw data usually requires large storage and computation. This paper addresses a fundamental problem in PCSR: How to downsample the dense point cloud with arbitrary scales while pres…

2021

Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

ICML 2021oral

Pre-trained representations are becoming crucial for many NLP and perception tasks. While representation learning in NLP has transitioned to training on raw text without human annotations, visual and vision-language representations still rely heavily on curated training datasets that are expensive o…

Cited by 4467SourcePDFScholar
2021

Shallow Feature Matters for Weakly Supervised Object Localization

CVPR 2021poster

Weakly supervised object localization (WSOL) aims to localize objects by only utilizing image-level labels. Class activation maps (CAMs) are the commonly used features to achieve WSOL. However, previous CAM-based methods did not take full advantage of the shallow features, despite their importance f…

Cited by 117PDFcodeScholar
2021

Sparse Single Sweep LiDAR Point Cloud Segmentation via Learning Contextual Shape Priors from Scene Completion

AAAI 2021technical

LiDAR point cloud analysis is a core task for 3D computer vision, especially for autonomous driving. However, due to the severe sparsity and noise interference in the single sweep LiDAR point cloud, the accurate semantic segmentation is non-trivial to achieve. In this paper, we propose a novel spars…

2021

Temporal Modulation Network for Controllable Space-Time Video Super-Resolution

CVPR 2021poster

Space-time video super-resolution (STVSR) aims to increase the spatial and temporal resolutions of low-resolution and low-frame-rate videos. Recently, deformable convolution based methods have achieved promising STVSR performance, but they could only infer the intermediate frame pre-defined in the t…

Cited by 113PDFcodeScholar
2020

Attention-Guided Lightweight Network for Real-Time Segmentation of Robotic Surgical Instruments

ICRA 2020poster

The real-time segmentation of surgical instruments plays a crucial role in robot-assisted surgery. However, it is still a challenging task to implement deep learning models to do real-time segmentation for surgical instruments due to their high computational costs and slow inference speed. In this p…

Cited by 61SourcecodeScholar
2020

BARNet: Bilinear Attention Network with Adaptive Receptive Fields for Surgical Instrument Segmentation

IJCAI 2020poster

Surgical instrument segmentation is crucial for computer-assisted surgery. Different from common object segmentation, it is more challenging due to the large illumination variation and scale variation in the surgical scenes. In this paper, we propose a bilinear attention network with adaptive recept…

Cited by 0SourcePDFScholar
2020

Hierarchical Chinese Legal event extraction via Pedal Attention Mechanism

COLING 2020main

Event extraction plays an important role in legal applications, including case push and auxiliary judgment. However, traditional event structure cannot express the connections between arguments, which are extremely important in legal events. Therefore, this paper defines a dynamic event structure fo…

2020

PointASNL: Robust Point Clouds Processing Using Nonlocal Neural Networks With Adaptive Sampling

CVPR 2020poster

Raw point clouds data inevitably contains outliers or noise through acquisition from 3D sensors or reconstruction algorithms. In this paper, we present a novel end-to-end network for robust point clouds processing, named PointASNL, which can deal with point clouds with noise effectively. The key com…

Cited by 764PDFcodeScholar
2020

Towards Content-Independent Multi-Reference Super-Resolution: Adaptive Pattern Matching and Feature Aggregation

ECCV 2020poster

Recovering realistic textures from a largely down-sampled low resolution (LR) image with complicated patterns is a challenging problem in image super-resolution. This work investigates a novel multi-reference based super-resolution problem by proposing a Content Independent Multi-Reference Super-Res…

Cited by 34SourcePDFScholar
2019

Feedback Network for Image Super-Resolution

CVPR 2019poster

Recent advances in image super-resolution (SR) explored the power of deep learning to achieve a better reconstruction performance. However, the feedback mechanism, which commonly exists in human visual system, has not been fully exploited in existing deep learning based image SR methods. In this pap…

Cited by 1053PDFcodeScholar
2019

Path Planning for Surgery Robot with Bidirectional Continuous Tree Search and Neural Network

IROS 2019poster

Solving a thorny issue of real-time path planning for surgery robot in uncertain environments, a novel algorithm named bidirectional continuous tree search (BCTS) is proposed. Most partially observable markov decision process (POMDP) planners address challenges of unknown environments with discrete…

Cited by 11SourceScholar
2019

Semi-Supervised Video Salient Object Detection Using Pseudo-Labels

ICCV 2019poster

Deep learning-based video salient object detection has recently achieved great success with its performance significantly outperforming any other unsupervised methods. However, existing data-driven approaches heavily rely on a large quantity of pixel-wise annotated video frames to deliver such promi…

Cited by 154PDFScholar
2018

Deep Neural Nets with Interpolating Function as Output Activation

NeurIPS 2018poster

We replace the output layer of deep neural nets, typically the softmax function, by a novel interpolating function. And we propose end-to-end training and testing algorithms for this new architecture. Compared to classical neural nets with softmax function as output activation, the surrogate with in…

Cited by 40SourcePDFScholar
2017

High-Resolution Shape Completion Using Deep Neural Networks for Global Structure and Local Geometry Inference

ICCV 2017spotlight

We propose a data-driven method for recovering missing parts of 3D shapes. Our method is based on a new deep learning architecture consisting of two sub-networks: a global structure inference network and a local geometry refinement network. The global structure inference network incorporates a long…

Cited by 367PDFScholar
2015

Learning Semantic Relationships for Better Action Retrieval in Images

CVPR 2015poster

Human actions capture a wide variety of interactions between people and objects. As a result, the set of possible actions is extremely large and it is difficult to obtain sufficient training examples for all actions. However, we could compensate for this sparsity in supervision by leveraging the ric…

Cited by 150SourcePDFScholar