← Search

Qi Chen

97 accepted papers

2026

Beyond the Global Scores: Fine-Grained Token Grounding as a Robust Detector of LVLM Hallucinations

CVPR 2026

Large vision-language models (LVLMs) achieve strong performance on visual reasoning tasks but remain highly susceptible to hallucination. Existing detection methods predominantly rely on coarse, whole-image measures of how an object token relates to the input image. This global strategy is limited:

Cited by 0SourceScholar
2026

Contrastive Symbolic Regression: Aligned Representations, Adaptive Prediction, and Diverse Ensembles

ICML 2026poster

Existing symbolic regression approaches primarily focus on learning explicit input-output mappings, often neglecting relational structures among data instances. This paper introduces Contrastive Symbolic Regression (CSR), a feature-construction-based symbolic regression approach that integrates evol…

Cited by 0SourceScholar
2026

CooperDrive: Enhancing Driving Decisions through Cooperative Perception

ICRA 2026poster

Autonomous vehicles equipped with robust onboard perception, localization, and planning still face limitations in occlusion and non-line-of-sight (NLOS) scenarios, where delayed reactions can increase collision risk. We propose CooperDrive, a cooperative perception framework that augments situationa…

2026

Discretized Density-Guided Source-Free Adaptation for Continuous Targets

ICML 2026spotlight

Source-Free Domain Adaptation (SFDA) enables model adaptation under distribution shifts without access to source data, providing a practical solution for privacy-sensitive applications and having shown substantial progress in classification. In contrast, regression involves ordered and continuous ta…

Cited by 0SourceScholar
2026

Foundation VAE for CT Reconstruction, Augmentation, and Generation

ICML 2026poster

Variational autoencoders (VAEs) compress high resolution CT volumes into compact latents while preserving clinically relevant structure. However, training CT-specific VAEs from scratch or heavily fine-tuning them incurs substantial computational and engineering cost, and often degrades under heterog…

Cited by 0SourceScholar
2026

Future-Gain Guided Test-Time Learning for Large Language Models

ICML 2026poster

Large language models (LLMs) inevitably encounter distribution shifts during real-world deployment, leading to performance degradation. Although test-time learning (TTL) adapts LLMs from unlabeled test streams, applying entropy minimization to autoregressive generation faces two challenges: (i) earl…

Cited by 0SourceScholar
2026

Geoint-R1: Formalizing Multimodal Geometric Reasoning with Dynamic Auxiliary Constructions

CVPR 2026

Mathematical geometric reasoning is essential for scientific discovery and educational development, requiring precise logic and rigorous formal verification. While recent advances in Multimodal Large Language Models (MLLMs) have improved reasoning tasks, existing models typically struggle with forma

Cited by 0SourcecodeScholar
2026

ISSE: AN INSTRUCTION-GUIDED SPEECH STYLE EDITING DATASET AND BENCHMARK

ICASSP 2026poster

Speech style editing refers to modifying the stylistic properties of speech while preserving its linguistic content and speaker identity. However, most existing approaches depend on explicit labels or reference audio, which limits both flexibility and scalability. More recent attempts to use natural…

Cited by 0SourcePDFScholar
2026

Learning Patient-Specific Disease Dynamics With Latent Flow Matching For Longitudinal Imaging Generation

ICLR 2026poster

Understanding disease progression is a central clinical challenge with direct implications for early diagnosis and personalized treatment. While recent generative approaches have attempted to model progression, key mismatches remain: disease dynamics are inherently continuous and monotonic, yet late…

Cited by 0SourceScholar
2026

M3CAD: Towards Generic Cooperative Autonomous Driving Benchmark

ICRA 2026poster

We introduce M3CAD, a comprehensive benchmark designed to advance research in generic cooperative autonomous driving. M3CAD comprises 204 sequences with 30,000 frames. Each sequence includes data from multiple vehicles and different types of sensors, e.g., LiDAR point clouds, RGB images, and GPS/IMU…

2026

MapReduce LoRA: Advancing the Pareto Front in Multi-Preference Optimization for Generative Models

CVPR 2026

Reinforcement learning from human feedback (RLHF) with reward models has advanced alignment of generative models to human aesthetic and perceptual preferences. However, jointly optimizing multiple rewards often incurs an alignment tax--improving one dimension while degrading others. To address this,

Cited by 0SourcecodeScholar
2026

PAPL-SLAM: Principal Axis-Anchored Monocular Point-Line SLAM

ICRA 2026poster

In point-line Simultaneous Localization and Mapping (SLAM) systems, the utilization of line structural information and the optimization of lines are two significant problems. The former is usually addressed through structural regularities, while the latter typically involves using minimal parameter …

2026

PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object Detection

CVPR 2026

Detection Transformer (DETR) has redefined object detection by casting it as a set prediction task within an end-to-end framework. Despite its elegance, DETR and its variants still rely on fixed learnable queries and suffer from severe query utilization imbalance, which limits adaptability and leave

Cited by 0SourceScholar
2026

RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation

ICLR 2026poster

Large language models excel at generating individual functions or single files of code, yet generating complete repositories from scratch remains a fundamental challenge. This capability is key to building coherent software systems from high-level specifications and realizing the full potential of a…

Cited by 0SourcecodeScholar
2026

Tavatar: Topology-Aware Gaussian Attribute Derivation for Animatable Human Avatars

CVPR 2026

Reconstructing high-fidelity, animatable human avatars from monocular videos remains a critical challenge. Existing 3DGS-based human animation methods constrain Gaussian parameters but exclude scale, which we argue is crucial for adapting human poses to challenging out-of-distribution poses. To achi

Cited by 0SourceScholar
2026

Tracking the Unstable: Appearance-Guided Motion Modeling for Robust Multi-Object Tracking in UAV-Captured Videos

AAAI 2026technical

Multi-object tracking (MOT) aims to track multiple objects while maintaining consistent identities across frames of a given video. In unmanned aerial vehicle (UAV) recorded videos, frequent viewpoint changes and complex UAV-ground relative motion dynamics pose significant challenges, which often lea

Cited by 0SourcePDFScholar
2026

UniABG: Unified Adversarial View Bridging and Graph Correspondence for Unsupervised Cross-View Geo-Localization

AAAI 2026technical

Cross-view geo-localization (CVGL) matches query images (e.g., drone) to geographically corresponding opposite-view imagery (e.g., satellite). While supervised methods achieve strong performance, their reliance on extensive pairwise annotations limits scalability. Unsupervised alternatives avoid ann

Cited by 0SourcePDFScholar
2026

Zero-source LLM Hallucination Detection with Human-like Criteria Probing

ICML 2026poster

Large language models (LLMs) often hallucinate by generating factually incorrect or unfaithful content, posing significant risks to their safe use. Detecting such hallucinations is particularly challenging under the zero-source constraint, where no model internals or external references are availabl…

Cited by 0SourceScholar
2025

A2DO: Adaptive Anti-Degradation Odometry with Deep Multi-Sensor Fusion for Autonomous Navigation

ICRA 2025

Accurate localization is essential for the safe and effective navigation of autonomous vehicles, and Simultaneous Localization and Mapping (SLAM) is a cornerstone technology in this context. However, The performance of the SLAM system can deteriorate under challenging conditions such as low light, a

Cited by 1SourceScholar
2025

Alleviating Performance Degradation Caused by Out-of-Distribution Issues in Embedding-Based Retrieval

EMNLP 2025

In Embedding Based Retrieval (EBR), Approximate Nearest Neighbor (ANN) algorithms are widely adopted for efficient large-scale search. However, recent studies reveal a query out-of-distribution (OOD) issue, where query and base embeddings follow mismatched distributions, significantly degrading ANN

Cited by 0SourcePDFScholar
2025

Are Pixel-Wise Metrics Reliable for Computerized Tomography Reconstruction?

NeurIPS 2025poster

Widely adopted evaluation metrics for sparse-view CT reconstruction, such as Structural Similarity Index Measure and Peak Signal-to-Noise Ratio, prioritize pixel-wise fidelity but often fail to capture the completeness of critical anatomical structures, particularly small or thin regions that are ea…

Cited by 0SourceScholar
2025

Attention-Driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models Without Fine-Tuning

AAAI 2025technical

Recent advancements in Multimodal Large Language Models (MLLMs) have generated significant interest in their ability to autonomously interact with and interpret Graphical User Interfaces (GUIs). A major challenge in these systems is grounding—accurately identifying critical GUI components such as te…

2025

Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video

ICCV 2025poster

We propose a novel and general framework to disentangle video data into its dynamic motion and static content components. Our proposed method is a self-supervised pipeline with less assumptions and inductive biases than previous works: it utilizes a transformer-based architecture to jointly generate…

Cited by 0SourcePDFScholar
2025

ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question Answering

EMNLP 2025

Chart question answering (CQA) has become a critical multimodal task for evaluating the reasoning capabilities of vision-language models. While early approaches have shown promising performance by focusing on visual features or leveraging large-scale pre-training, most existing evaluations rely on r

Cited by 0SourcePDFScholar
2025

Dual Energy-Based Model with Open-World Uncertainty Estimation for Out-of-distribution Detection

CVPR 2025poster

Out-of-distribution (OOD) detection is crucial for machine learning models deployed in open-world environments. However, existing methods often struggle with model over-confidence or rely heavily on empirical energy value estimation, limiting their scalability and generalizability. This paper introd…

2025

Efficient Scale-Uniform 3D Visual Coverage Algorithm for UAV Based on Elastic Photogrammetric Constraints

ICRA 2025

Unmanned aerial vehicles equipped with modern vision algorithms are crucial for missions such as reconstruction and target acquisition. However, when deployed in the field, undulating terrain can cause significant fluctuations in image scale and degrade the performance of vision algorithms. Instead

Cited by 0SourceScholar
2025

Efficiently Selecting Response Generation Strategies for Synthetic Data Construction by Self-Aligned Perplexity

EMNLP 2025

Fine-tuning large language models (LLMs) typically relies on producing large sets of input-output pairs. Yet for a given question, there can be many valid outputs. In practice, these outputs are often derived by distilling knowledge from teacher models, and they can vary depending on the specific te

2025

Enhancing Large Language Model Performance with Gradient-Based Parameter Selection

AAAI 2025technical

Large language models (LLMs) have revolutionized numerous fields of research, driving significant advancements in natural language processing, machine translation, and beyond. Although the extensive number of parameters contributes a lot to the great success, existing studies indicate that not all m…

Cited by 0SourcePDFScholar
2025

EpiCoder: Encompassing Diversity and Complexity in Code Generation

ICML 2025poster

Existing methods for code generation use code snippets as seed data, restricting the complexity and diversity of the synthesized data. In this paper, we introduce a novel feature tree-based synthesis framework, which revolves around hierarchical code features derived from high-level abstractions of…

Cited by 4SourcePDFScholar
2025

From Words to Structured Visuals: A Benchmark and Framework for Text-to-Diagram Generation and Editing

CVPR 2025highlight

We introduce the task of text-to-diagram generation, which focuses on creating structured visual representations directly from textual descriptions. Existing approaches in text-to-image and text-to-code generation lack the logical organization and flexibility needed to produce accurate, editable dia…

Cited by 2SourcePDFScholar
2025

Generalization in VAE and Diffusion Models: A Unified Information-Theoretic Analysis

ICLR 2025poster

Despite the empirical success of Diffusion Models (DMs) and Variational Autoencoders (VAEs), their generalization performance remains theoretically underexplored, especially lacking a full consideration of the shared encoder-generator structure. Leveraging recent information-theoretic tools, we prop…

Cited by 0SourcePDFScholar
2025

IPDN: Image-enhanced Prompt Decoding Network for 3D Referring Expression Segmentation

AAAI 2025technical

3D Referring Expression Segmentation (3D-RES) aims to segment point cloud scenes based on a given expression. However, existing 3D-RES approaches face two major challenges: feature ambiguity and intent ambiguity. Feature ambiguity arises from information loss or distortion during point cloud acquisi…

2025

InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles

EMNLP 2025

LLMs have shown strong performance on human-centric reasoning tasks. While previous evaluations have explored whether LLMs can infer intentions or detect deception, they often overlook the individualized reasoning styles that influence how people interpret and act in social contexts. Social deductio

Cited by 0SourcePDFScholar
2025

Integrative Decoding: Improving Factuality via Implicit Self-consistency

ICLR 2025poster

Self-consistency-based approaches, which involve repeatedly sampling multiple outputs and selecting the most consistent one as the final response, prove to be remarkably effective in improving the factual accuracy of large language models. Nonetheless, existing methods usually have strict constraint…

Cited by 4SourcePDFScholar
2025

Localizing Before Answering: A Benchmark for Grounded Medical Visual Question Answering

IJCAI 2025

Medical Large Multi-modal Models (LMMs) have demonstrated remarkable capabilities in medical data interpretation. However, these models frequently generate hallucinations contradicting source evidence, particularly due to inadequate localization reasoning. This work reveals a critical limitation in

Cited by 0SourcePDFScholar
2025

Multi-Hierarchical Fine-Grained Feature Mapping Driven by Feature Contributions for Molecular Odor Prediction

IJCAI 2025

Molecular odor prediction involves using a molecule's structure to estimate its odor. While accurate prediction remains challenging, AI models can suggest potential odors. Existing methods, however, often rely on basic descriptors or handcrafted fingerprints, which lack expressive power and hinder e

Cited by 0SourcePDFScholar
2025

OVG-HQ: Online Video Grounding with Hybrid-modal Queries

ICCV 2025poster

Video grounding (VG) task focuses on locating specific moments in a video based on a query, usually in text form. However, traditional VG struggles with some scenarios like streaming video or queries using visual cues. To fill this gap, we present a new task named Online Video Grounding with Hybrid-…

Cited by 0SourcePDFScholar
2025

PanTS: The Pancreatic Tumor Segmentation Dataset

NeurIPS 2025poster

PanTS is a large-scale, multi-institutional dataset curated to advance research in pancreatic CT analysis. It contains 36,390 CT scans from 145 medical centers, with expert-validated, voxel-wise annotations of over 993,000 anatomical structures, covering pancreatic tumors, pancreas head, body, and t…

Cited by 0SourceScholar
2025

RAG-SR: Retrieval-Augmented Generation for Neural Symbolic Regression

ICLR 2025spotlight

Symbolic regression is a key task in machine learning, aiming to discover mathematical expressions that best describe a dataset. While deep learning has increased interest in using neural networks for symbolic regression, many existing approaches rely on pre-trained models. These models require sign…

Cited by 0SourcePDFScholar
2025

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

NeurIPS 2025poster

Transformer-based Large Language Models (LLMs) have become increasingly important. However, scaling LLMs to longer contexts incurs slow inference speed and high GPU memory consumption for caching key-value (KV) vectors. This paper presents RetrievalAttention, a training-free approach to both acceler…

Cited by 0SourcecodeScholar
2025

Scaling Tumor Segmentation: Best Lessons from Real and Synthetic Data

ICCV 2025poster

AI for tumor segmentation is limited by the lack of large, voxel-wise annotated datasets, which are hard to create and require medical experts. In our proprietary JHH dataset of 3,000 annotated pancreatic tumor scans, we found that AI performance stopped improving after 1,500 scans. With synthetic d…

2025

Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding

ICCV 2025poster

Visual grounding (VG) is the capability to identify the specific regions in an image associated with a particular text description. In medical imaging, VG enhances interpretability by highlighting relevant pathological features corresponding to textual descriptions, improving model transparency and…

Cited by 0SourcePDFScholar
2025

Separation of Powers: On Segregating Knowledge from Observation in LLM-enabled Knowledge-based Visual Question Answering

CVPR 2025poster

Knowledge-based visual question answering (KBVQA) separates image interpretation and knowledge retrieval into separate processes, motivated in part by the fact that they are very different tasks. In this paper, we transform the KBVQA into linguistic question-answering tasks so that we can leverage t…

Cited by 0SourcePDFScholar
2025

SketchAgent: Generating Structured Diagrams from Hand-Drawn Sketches

IJCAI 2025

Hand-drawn sketches are a natural and efficient medium for capturing and conveying ideas. Despite significant advancements in controllable natural image generation, translating freehand sketches into structured, machine-readable diagrams remains a labor-intensive and predominantly manual task. The p

Cited by 0SourcePDFScholar
2025

TSTAI: A Time-varying Brain Effective Connectivity Network Construction Method Combining with Brain Active Information

IJCAI 2025

More accurate construction of brain effective conncetivity networks remains a great challenge to achieve accurate auxiliary diagnosis of brain diseases and in-depth exploration of brain function. However, existing methods only consider higher-order or non-stationary assumptions, rather than simultan

Cited by 0SourcePDFScholar
2025

Training-Free Class Purification for Open-Vocabulary Semantic Segmentation

ICCV 2025poster

Fine-tuning pre-trained vision-language models has emerged as a powerful approach for enhancing open-vocabulary semantic segmentation (OVSS). However, the substantial computational and resource demands associated with training on large datasets have prompted interest in training-free methods for OVS…

2025

Uncertainty-Guided Robotic Manipulation Through Variational Information Bottleneck in Imitation Learning

RA-L 2025

Traditional imitation learning methods typically rely on high-quality expert demonstrations and exhibit poor generalization when deployed in unfamiliar environments. A key limitation is their inability to effectively quantify and utilize epistemic uncertainty in the decision-making process. To addre

Cited by 0SourcecodeScholar
2025

VQTalker: Towards Multilingual Talking Avatars Through Facial Motion Tokenization

AAAI 2025technical

We present VQTalker, a Vector Quantization-based framework for multilingual talking head generation that addresses the challenges of lip synchronization and natural motion across diverse languages. Our approach is grounded in the phonetic principle that human speech comprises a finite set of distinc…

Cited by 0SourcePDFScholar
2025

Velocitune: A Velocity-based Dynamic Domain Reweighting Method for Continual Pre-training

ACL 2025long

It is well-known that a diverse corpus is critical for training large language models, which are typically constructed from a mixture of various domains. In general, previous efforts resort to either sampling training data from different domains with static proportions or dynamically adjusting these…

2024

3D Point Cloud Semantic Segmentation Based on Diffusion Model

ICASSP 2024accepted

Point cloud segmentation plays a crucial role in extracting unique attributes and separating various objects, thereby enabling semantic comprehension and analysis. In this paper, we introduce a novel point cloud segmentation approach based on Diffusion Probabilistic Network (DDPM). The proposed mode…

Cited by 0SourceScholar
2024

3D-STMN: Dependency-Driven Superpoint-Text Matching Network for End-to-End 3D Referring Expression Segmentation

AAAI 2024technical

In 3D Referring Expression Segmentation (3D-RES), the earlier approach adopts a two-stage paradigm, extracting segmentation proposals and then matching them with referring expressions. However, this conventional paradigm encounters significant challenges, most notably in terms of the generation of l…

2024

CREAD: A Classification-Restoration Framework with Error Adaptive Discretization for Watch Time Prediction in Video Recommender Systems

AAAI 2024technical

The watch time is a significant indicator of user satisfaction in video recommender systems. However, the prediction of watch time as a target variable is often hindered by its highly imbalanced distribution with a scarcity of observations for larger target values and over-populated samples for smal…

Cited by 6SourcePDFScholar
2024

G-NeRF: Geometry-enhanced Novel View Synthesis from Single-View Images

CVPR 2024poster

Novel view synthesis aims to generate new view images of a given view image collection. Recent attempts address this problem relying on 3D geometry priors (e.g. shapes sizes and positions) learned from multi-view images. However such methods encounter the following limitations: 1) they require a set…

2024

IRGen: Generative Modeling for Image Retrieval

ECCV 2024poster

"While generative modeling has become prevalent across numerous research fields, its integration into the realm of image retrieval remains largely unexplored and underjustified. In this paper, we present a novel methodology, reframing image retrieval as a variant of generative modeling and employing…

2024

Learning Multiscale Consistency for Self-Supervised Electron Microscopy Instance Segmentation

ICASSP 2024accepted

Electron microscopy (EM) images are notoriously challenging to segment due to their complex structures and lack of effective annotations. Fortunately, large-scale self-supervised pretraining offers a promising solution by allowing us to acquire prior knowledge of cell and subcellular tissue structur…

Cited by 0SourceScholar
2024

Neighborhood-Enhanced Multimodal Collaborative Filtering for Item Cold Start Recommendation

ICASSP 2024accepted

The lack of interaction data of new items in recommendation systems leads to the problem of cold-start item recommendations. Current methods usually approximate content features of the items to interaction embeddings and then use content features for prediction. However, these methods typically lear…

Cited by 0SourceScholar
2024

PairAug: What Can Augmented Image-Text Pairs Do for Radiology?

CVPR 2024poster

Current vision-language pre-training (VLP) methodologies predominantly depend on paired image-text datasets a resource that is challenging to acquire in radiology due to privacy considerations and labelling complexities. Data augmentation provides a practical solution to overcome the issue of data s…

2024

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation

NeurIPS 2024oral

3D Referring Expression Segmentation (3D-RES) aims to segment 3D objects by correlating referring expressions with point clouds. However, traditional approaches frequently encounter issues like over-segmentation or mis-segmentation, due to insufficient emphasis on spatial information of instances. I…

2024

SIMFALL: A Data Generator for RF-Based Fall Detection

ICASSP 2024accepted

Fall detection using Radio Frequency (RF) signals with deep learning has exhibited significant promise in recent years. However, the costly collection of RF data with falls has hampered the performance of existing methods. While there has been approaches which can generate RF signals using various s…

Cited by 0SourceScholar
2024

STT: Stateful Tracking with Transformers for Autonomous Driving

ICRA 2024poster

Tracking objects in three-dimensional space is critical for autonomous driving. To ensure safety while driving, the tracker must be able to reliably track objects across frames and accurately estimate their states such as velocity and acceleration in the present. Existing works frequently focus on t…

Cited by 0SourceScholar
2024

SiCP: Simultaneous Individual and Cooperative Perception for 3D Object Detection in Connected and Automated Vehicles

IROS 2024poster

Cooperative perception for connected and automated vehicles is traditionally achieved through the fusion of feature maps from two or more vehicles. However, the absence of feature maps shared from other vehicles can lead to a significant decline in 3D object detection performance for cooperative per…

Cited by 6SourcecodeScholar
2024

Towards Generalizable Tumor Synthesis

CVPR 2024poster

Tumor synthesis enables the creation of artificial tumors in medical images facilitating the training of AI models for tumor detection and segmentation. However success in tumor synthesis hinges on creating visually realistic tumors that are generalizable across multiple organs and furthermore the r…

2024

Towards Understanding Evolving Patterns in Sequential Data

NeurIPS 2024spotlight

In many machine learning tasks, data is inherently sequential. Most existing algorithms learn from sequential data in an auto-regressive manner, which predicts the next unseen data point based on the observed sequence, implicitly assuming the presence of an \emph{evolving pattern} embedded in the da…

Cited by 1SourcePDFScholar
2024

Weak-eval-Strong: Evaluating and Eliciting Lateral Thinking of LLMs with Situation Puzzles

NeurIPS 2024poster

While advancements in NLP have significantly improved the performance of Large Language Models (LLMs) on tasks requiring vertical thinking, their lateral thinking capabilities remain under-explored and challenging to measure due to the complexity of assessing creative thought processes and the scarc…

2024

WebVLN: Vision-and-Language Navigation on Websites

AAAI 2024technical

Vision-and-Language Navigation (VLN) task aims to enable AI agents to accurately understand and follow natural language instructions to navigate through real-world environments, ultimately reaching specific target locations. We recognise a promising opportunity to extend VLN to a comparable navigati…

2023

Algorithm-Dependent Bounds for Representation Learning of Multi-Source Domain Adaptation

AISTATS 2023poster

We use information-theoretic tools to derive a novel analysis of Multi-source Domain Adaptation (MDA) from the representation learning perspective. Concretely, we study joint distribution alignment for supervised MDA with few target labels and unsupervised MDA with pseudo labels, where the latter is…

2023

Improving Few-Shot Learning for Talking Face System with TTS Data Augmentation

ICASSP 2023accepted

Audio-driven talking face has attracted broad interest from academia and industry recently. However, data acquisition and labeling in audio-driven talking face are labor-intensive and costly. The lack of data resource results in poor synthesis effect. To alleviate this issue, we propose to use TTS (…

Cited by 0SourceScholar
2023

Model-enhanced Vector Index

NeurIPS 2023poster

Embedding-based retrieval methods construct vector indices to search for document representations that are most similar to the query representations. They are widely used in document retrieval due to low latency and decent recall performance. Recent research indicates that deep retrieval solutions o…

2023

On the Stability-Plasticity Dilemma in Continual Meta-Learning: Theory and Algorithm

NeurIPS 2023poster

We focus on Continual Meta-Learning (CML), which targets accumulating and exploiting meta-knowledge on a sequence of non-i.i.d. tasks. The primary challenge is to strike a balance between stability and plasticity, where a model should be stable to avoid catastrophic forgetting in previous tasks and…

2023

Prompt Switch: Efficient CLIP Adaptation for Text-Video Retrieval

ICCV 2023poster

In text-video retrieval, recent works have benefited from the powerful learning capabilities of pre-trained text-image foundation models (e.g., CLIP) by adapting them to the video domain. A critical problem for them is how to effectively capture the rich semantics inside the video using the image en…

Cited by 40PDFcodeScholar
2023

Self-Supervised Neuron Segmentation with Multi-Agent Reinforcement Learning

IJCAI 2023poster

The performance of existing supervised neuron segmentation methods is highly dependent on the number of accurate annotations, especially when applied to large scale electron microscopy (EM) data. By extracting semantic information from unlabeled data, self-supervised methods can improve the performa…

2022

A Neural Corpus Indexer for Document Retrieval

NeurIPS 2022accept

Current state-of-the-art document retrieval solutions mainly follow an index-retrieve paradigm, where the index is hard to be directly optimized for the final retrieval target. In this paper, we aim to show that an end-to-end deep neural network unifying training and indexing stages can significantl…

Cited by 148SourcePDFScholar
2022

Fair Representation Learning through Implicit Path Alignment

ICML 2022spotlight

We consider a fair representation learning perspective, where optimal predictors, on top of the data representation, are ensured to be invariant with respect to different sub-groups. Specifically, we formulate this intuition as a bi-level optimization, where the representation is learned in the oute…

Cited by 30SourcePDFScholar
2022

Global Evolution Neural Network for Segmentation of Remote Sensing Images

ICASSP 2022accepted

The popular convolutional neural networks (CNNs) have been successfully used in very high-resolution remote sensing image semantic segmentation. However, these networks often suffer from performance limitations. First, although deeper networks usually provide better feature representation, they may…

Cited by 0SourceScholar
2022

Human-Robot Variable Impedance Skills Transfer Learning Based on Dynamic Movement Primitives

RA-L 2022

Endowing robots with human-like abilities to perform motor skills smoothly and naturally is one of the important goals of robotics. Learning from demonstration (LfD) has been successfully applied for learning tasks on robots, for which the human tutor can demonstrate a successful execution. Learning

Cited by 35SourceScholar
2022

On Learning Fairness and Accuracy on Multiple Subgroups

NeurIPS 2022accept

We propose an analysis in fair learning that preserves the utility of the data while reducing prediction disparities under the criteria of group sufficiency. We focus on the scenario where the data contains multiple or even many subgroups, each with limited number of samples. As a result, we present…

2022

Optimization-Induced Graph Implicit Nonlinear Diffusion

ICML 2022spotlight

Due to the over-smoothing issue, most existing graph neural networks can only capture limited dependencies with their inherently finite aggregation layers. To overcome this limitation, we propose a new kind of graph convolution, called Graph Implicit Nonlinear Diffusion (GIND), which implicitly has…

2022

Self-Supervised Image-Specific Prototype Exploration for Weakly Supervised Semantic Segmentation

CVPR 2022poster

Weakly Supervised Semantic Segmentation (WSSS) based on image-level labels has attracted much attention due to low annotation costs. Existing methods often rely on Class Activation Mapping (CAM) that measures the correlation between image pixels and classifier weight. However, the classifier focuses…

Cited by 192PDFcodeScholar
2021

Contrastive Neural Architecture Search With Neural Architecture Comparators

CVPR 2021poster

One of the key steps in Neural Architecture Search (NAS) is to estimate the performance of candidate architectures. Existing methods either directly use the validation performance or learn a predictor to estimate the performance. However, these methods can be either computationally expensive or very…

Cited by 87PDFcodeScholar
2021

Generalization Bounds For Meta-Learning: An Information-Theoretic Analysis

NeurIPS 2021spotlight

We derive a novel information-theoretic analysis of the generalization property of meta-learning algorithms. Concretely, our analysis proposes a generic understanding in both the conventional learning-to-learn framework \citep{amit2018meta} and the modern model-agnostic meta-learning (MAML) algorith…

2021

PolarStream: Streaming Object Detection and Segmentation with Polar Pillars

NeurIPS 2021poster

Recent works recognized lidars as an inherently streaming data source and showed that the end-to-end latency of lidar perception models can be reduced significantly by operating on wedge-shaped point cloud sectors rather then the full point cloud. However, due to use of cartesian coordinate systems…

Cited by 58SourcePDFScholar
2021

SPANN: Highly-efficient Billion-scale Approximate Nearest Neighborhood Search

NeurIPS 2021spotlight

The in-memory algorithms for approximate nearest neighbor search (ANNS) have achieved great success for fast high-recall search, but are extremely expensive when handling very large scale database. Thus, there is an increasing request for the hybrid ANNS solutions with small memory and inexpensive s…

2021

StrokeGAN: Reducing Mode Collapse in Chinese Font Generation via Stroke Encoding

AAAI 2021technical

The generation of stylish Chinese fonts is an important problem involved in many applications. Most of existing generation methods are based on the deep generative models, particularly, the generative adversarial networks (GAN) based models. However, these deep generative models may suffer from the…

2020

Closed-Loop Matters: Dual Regression Networks for Single Image Super-Resolution

CVPR 2020poster

Deep neural networks have exhibited promising performance in image super-resolution (SR) by learning a nonlinear mapping function from low-resolution (LR) images to high-resolution (HR) images. However, there are two underlying limitations to existing SR methods. First, learning the mapping function…

Cited by 425PDFcodeScholar
2020

Every View Counts: Cross-View Consistency in 3D Object Detection with Hybrid-Cylindrical-Spherical Voxelization

NeurIPS 2020poster

Recent voxel-based 3D object detectors for autonomous vehicles learn point cloud representations either from bird eye view (BEV) or range view (RV, a.k.a. the perspective view). However, each view has its own strengths and weaknesses. In this paper, we present a novel framework to unify and leverage…

Cited by 129SourcePDFScholar
2020

Intelligent Home 3D: Automatic 3D-House Design From Linguistic Descriptions Only

CVPR 2020poster

Home design is a complex task that normally requires architects to finish with their professional skills and tools. It will be fascinating that if one can produce a house plan intuitively without knowing much knowledge about home design and experience of using complex designing tools, for example, v…

Cited by 47PDFcodeScholar
2020

Object as Hotspots: An Anchor-Free 3D Object Detection Approach via Firing of Hotspots

ECCV 2020poster

Accurate 3D object detection in LiDAR based point clouds suffers from the challenges of data sparsity and irregularities. Existing methods strive to organize the points regularly, e.g. voxelize, pass them through a designed 2D/3D neural network, and then define object-level anchors that predict offs…

Cited by 211SourcePDFScholar
2019

NAT: Neural Architecture Transformer for Accurate and Compact Architectures

NeurIPS 2019poster

Designing effective architectures is one of the key factors behind the success of deep neural networks. Existing deep architectures are either manually designed or automatically searched by some Neural Architecture Search (NAS) methods. However, even a well-searched architecture may still contain ma…

2015

Maintaining constant towing tension between cable ship and burying system under sea waves by hybrid FUZZY P + ID controller

IROS 2015poster

In this paper, we propose a hybrid FUZZY P + ID controller to stabilize the towing cable tension between a cable ship and a burying system. First, we develop the model of a winch system driven by valve-controlled hydraulic motors and evaluate the step responses yielded by the conventional PID and th…

Cited by 5SourceScholar