← Search

Jing Zhang

253 accepted papers

2026

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention

CVPR 2026

Vision-Language-Action (VLA) models have shown remarkable progress in embodied tasks recently, but most methods process visual observations independently at each timestep. This history-agnostic design treats robot manipulation as a Markov Decision Process, even though real-world robotic control is i

Cited by 0SourceScholar
2026

All Roads Lead to Rome: Incentivizing Divergent Thinking in Vision-Language Models

CVPR 2026

Recent studies have demonstrated that Reinforcement Learning (RL), notably Group Relative Policy Optimization (GRPO), can intrinsically elicit and enhance the reasoning capabilities of Vision-Language Models (VLMs). However, despite the promise, the underlying mechanisms that drive the effectiveness

Cited by 0SourcecodeScholar
2026

AnesSuite: A Comprehensive Benchmark and Dataset Suite for Anesthesiology Reasoning in LLMs

ICLR 2026poster

The application of large language models (LLMs) in the medical field has garnered significant attention, yet their reasoning capabilities in more specialized domains like anesthesiology remain underexplored. To bridge this gap, we introduce AnesSuite, the first comprehensive dataset suite specifical…

Cited by 0SourcecodeScholar
2026

Any2Any: Unified Arbitrary Modality Translation for Remote Sensing

ICML 2026poster

Multi-modal remote sensing imagery provides complementary observations of the same geographic scene, yet such observations are frequently incomplete in practice. Existing cross-modal translation methods treat each modality pair as an independent task, resulting in quadratic complexity and limited ge…

Cited by 0SourceScholar
2026

Attribution Analysis-based Concept Alignment: A Human-in-the-loop Data Debugging Framework

AAAI 2026technical

Ensuring consistently high-quality training data is essential for developing reliable machine learning systems. Recent research demonstrates that incorporating human supervision into training set debugging effectively improves model performance, especially for text classification tasks. However, suc

Cited by 0SourcePDFScholar
2026

Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing

CVPR 2026

Document parsing is a fine-grained task where image resolution significantly impacts performance. While advanced research leveraging vision-language models benefits from high-resolution input to boost model performance, this often leads to a quadratic increase in the number of vision tokens and sign

Cited by 2SourcecodeScholar
2026

CRAG: Can 3D Generative Models Help 3D Assembly?

ICML 2026poster

Most existing 3D assembly methods treat the problem as pure pose estimation, rearranging observed parts via rigid transformations. In contrast, human assembly naturally couples structural reasoning with holistic shape inference. Inspired by this intuition, we reformulate 3D assembly as a joint probl…

Cited by 0SourceScholar
2026

DCMM-Transformer: Degree-Corrected Mixed-Membership Attention for Medical Imaging

AAAI 2026technical

Medical images exhibit latent anatomical groupings, such as organs, tissues, and pathological regions, that standard Vision Transformers (ViTs) fail to exploit. While recent work like SBM-Transformer attempts to incorporate such structures through stochastic binary masking, they suffer from non-diff

Cited by 0SourcePDFScholar
2026

Degradation-Aware Metric Prompting for Hyperspectral Image Restoration

ICML 2026poster

Unified hyperspectral image (HSI) restoration aims to recover diverse degradations within a single model. However, current methods often rely on impractical explicit priors or opaque black-box representations that overfit to training distributions, hampering generalization to unseen scenarios. To br…

Cited by 0SourceScholar
2026

Efficient-SAM2: Accelerating SAM2 with Object-Aware Visual Encoding and Memory Retrieval

ICLR 2026poster

Segment Anything Model 2 (SAM2) shows excellent performance in video object segmentation tasks; however, the heavy computational burden hinders its application in real-time video processing. Although there have been efforts to improve the efficiency of SAM2, most of them focus on retraining a lightw…

Cited by 0SourceScholar
2026

FedFINFO: A General Full-Informativeness Federated Graph Learning from Open Cross-Domain Data

IJCAI 2026

Open cross-domain federated graph learning facilitates collaborative learning among clients from distinct graph domains while preserving privacy. However, severe structure and feature heterogeneity in open scenarios exacerbates the multiplicative amplification of structural and feature noises within

Cited by 0Scholar
2026

From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models

ICML 2026oral

Latent actions serve as an intermediate representation that enables consistent modeling of vision-language-action (VLA) models across heterogeneous datasets. However, approaches to supervising VLAs with latent actions are fragmented and lack a systematic comparison. This work structures the study of…

Cited by 0SourceScholar
2026

GPFlow: Gaussian Prototype Probability Flow for Unsupervised Multi-Modal Anomaly Detection

CVPR 2026

In this paper, we study unsupervised multi-modal anomaly detection under challenging few-shot conditions, where only a few normal training samples are available for each class. To prevent the trivial reconstruction of anomalies, recent methods often rely on discrete prototypes to establish an inform

Cited by 0SourceScholar
2026

GeoBridge: A Semantic-Anchored Multi-View Foundation Model Bridging Images and Text for Geo-Localization

CVPR 2026

Cross-view geo-localization infers a location by retrieving geo-tagged reference images that visually correspond to a query image. However, the traditional satellite-centric paradigm limits robustness when high-resolution or up-to-date satellite imagery is unavailable. It further underexploits compl

Cited by 0SourcecodeScholar
2026

Heuristic-inspired Reasoning Priors Facilitate Data-Efficient Referring Object Detection

CVPR 2026

Most referring object detection (ROD) models, especially the modern grounding detectors, are designed for data-rich conditions, yet many practical deployments, such as robotics, augmented reality, and other specialized domains, would face severe label scarcity. In such regimes, end-to-end grounding

Cited by 0SourcecodeScholar
2026

Learning Spatial Decay for Vision Transformers

AAAI 2026technical

Vision Transformers (ViTs) have revolutionized computer vision, yet their self-attention mechanism lacks explicit spatial inductive biases, leading to suboptimal performance on spatially-structured tasks. Existing approaches introduce data-independent spatial decay based on fixed distance metrics, a

Cited by 0SourcePDFScholar
2026

MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs

ICLR 2026poster

The Mixture-of-Experts (MoE) architecture has become a predominant paradigm for scaling large language models (LLMs). Despite offering strong performance and computational efficiency, large MoE-based LLMs like DeepSeek-V3-0324 and Kimi-K2-Instruct present serious challenges due to substantial memory…

Cited by 0SourceScholar
2026

More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models

ICLR 2026poster

Reasoning has emerged as a pivotal capability in Large Language Models (LLMs). Through Reinforcement Learning (RL), typically Group Relative Policy Optimization (GRPO), these models are able to solve complex tasks such as mathematics and code generation. Building on these advances, recent research h…

Cited by 0SourcecodeScholar
2026

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization

ICML 2026poster

Large Language Models (LLMs) have demonstrated remarkable capabilities in understanding and generation tasks. However, their massive parameter scale leads to significant resource consumption and latency during inference. Post-training weight-only quantization offers a promising solution by reducing …

Cited by 0SourceScholar
2026

PIRN: Prototypical-based Intra-modal Reconstruction with Normality Communication for Multi-modal Anomaly Detection.

ICLR 2026poster

Unsupervised Multimodal anomaly detection (MAD) — identifying defects by jointly analyzing RGB images and 3D data — is crucial for quality control in manufacturing. However, existing MAD methods struggle when only a few normal samples are available. Cross-modal alignment models fail to learn stable…

Cited by 0SourceScholar
2026

PP-OCRv5: A Specialized 5M-Parameter Model Rivaling Billion-Parameter Vision-Language Models on OCR Tasks

CVPR 2026

The advent of "OCR 2.0" and large-scale vision-language models (VLMs) has set new benchmarks in text recogni- tion. However, these unified architectures often come with significant computational demands, challenges in precise text localization within complex layouts, and a propen- sity for textual h

Cited by 0SourcecodeScholar
2026

PTQ4ARVG: Post-Training Quantization for AutoRegressive Visual Generation Models

ICLR 2026poster

AutoRegressive Visual Generation (ARVG) models retain an architecture compatible with language models, while achieving performance comparable to diffusion-based models. Quantization is commonly employed in neural networks to reduce model size and computational latency. However, applying quantization…

Cited by 0SourcecodeScholar
2026

Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning

CVPR 2026

Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced the reasoning capabilities of Large Language Models (LLMs) and is now being applied to Vision-Language Models (VLMs). However, vanilla RLVR for VLMs verifies only the final textual output, critically neglecting the foun

Cited by 0SourcecodeScholar
2026

Probing and Bridging Geometry-Interaction Cues for Affordance Reasoning in Vision Foundation Models

CVPR 2026

What does it mean for a visual system to truly understand affordance? We argue that this understanding hinges on two complementary capacities: geometric perception, which identifies the structural parts of objects that enable interaction, and interaction perception, which models how an agent's actio

Cited by 0SourceScholar
2026

Prune4Web: DOM Tree Pruning Programming for Web Agent

AAAI 2026technical

Web automation uses intelligent agents to perform high-level tasks by mimicking human interactions with webpages. Despite recent advances in LLM-based web agents, efficiently navigating complex, real-world webpages remains challenging due to massive DOM structures (10,000 ~ 100,000 tokens). Current

Cited by 0SourcePDFScholar
2026

Residual Diffusion Bridge Model for Image Restoration

CVPR 2026

Diffusion bridge models establish probabilistic paths between arbitrary paired distributions and exhibit great potential for universal image restoration. Most existing methods merely treat them as simple variants of stochastic interpolants, lacking a unified analytical perspective. Besides, they ind

Cited by 0SourcecodeScholar
2026

RnG: A Unified Transformer for Complete 3D Modeling from Partial Observations

CVPR 2026

Humans perceive the 3D world from limited 2D observations. While recent feed-forward generalizable 3D reconstruction models can recover structures from sparse images, they typically represent only observed regions, leaving unseen geometry unmodeled. This raises a fundamental question: Can we infer c

Cited by 0SourceScholar
2026

SABER: Spatially Consistent 3D Universal Adversarial Objects for BEV Detectors

CVPR 2026

Adversarial robustness of BEV 3D object detectors is critical for autonomous driving (AD). Existing invasive attacks require altering the target vehicle itself (e.g. attaching patches), making them unrealistic and impractical for real-world evaluation. While non-invasive attacks that place adversari

Cited by 0SourceScholar
2026

SAQ-SAM: Semantically-Aligned Quantization for Segment Anything Model

AAAI 2026technical

Segment Anything Model (SAM) exhibits remarkable zero-shot segmentation capability; however, its prohibitive computational costs make edge deployment challenging. Although post-training quantization (PTQ) offers a promising compression solution, existing methods yield unsatisfactory results when app

Cited by 0SourcePDFScholar
2026

SARMAE: Masked Autoencoder for SAR Representation Learning

CVPR 2026

Synthetic Aperture Radar (SAR) imagery plays a critical role in all-weather, day-and-night remote sensing applications. However, existing SAR-oriented deep learning is constrained by data scarcity, while the physically grounded speckle noise in SAR imagery further hampers fine-grained semantic repre

Cited by 0SourcecodeScholar
2026

Scaling Agentic Verifier for Competitive Coding

ICML 2026poster

Large language models (LLMs) have demonstrated strong coding capabilities but still struggle to solve competitive programming problems correctly in a single attempt. Execution-based re-ranking offers a promising test-time scaling strategy, yet existing methods are constrained by either difficult tes…

Cited by 0SourceScholar
2026

SteerMusic: Enhanced Musical Consistency for Zero-shot Text-Guided and Personalized Music Editing

AAAI 2026technical

Music editing is an important step in music production, which has broad applications, including game development and film production. Most existing zero-shot text-guided editing methods rely on pretrained diffusion models by involving forward-backward diffusion processes. However, these methods ofte

Cited by 0SourcePDFScholar
2026

Text Before Vision: Staged Knowledge Injection Matters for Agentic RLVR in Ultra-High-Resolution Remote Sensing Understanding

ICML 2026poster

Multimodal reasoning for ultra-high-resolution (UHR) remote sensing (RS) is usually bottlenecked by visual evidence acquisition: the model necessities localizing tiny task-relevant regions in massive pixel spaces. While Agentic Reinforcement Learning with Verifiable Rewards (RLVR) using zoom-in tool…

Cited by 0SourceScholar
2026

Thinking in 360deg: Humanoid Visual Search in the Wild

CVPR 2026

Humans rely on the synergistic control of head (cephalomotor) and eye (oculomotor) to efficiently search for visual information in 360deg. However, prior approaches to visual search are limited to a static image, neglecting the physical embodiment and its interaction with the 3D world. How can we de

Cited by 0SourcecodeScholar
2026

UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes

CVPR 2026

Instruction-driven segmentation in remote sensing generates masks from guidance, offering great potential for accessible and generalizable applications. However, existing methods suffer from fragmented task formulations and limited instruction data, hindering effective understanding and generalizati

Cited by 0SourcecodeScholar
2026

VAnim: Rendering-Aware Sparse State Modeling for Structure-Preserving Vector Animation

ICML 2026poster

Scalable Vector Graphics (SVG) animation generation is pivotal for professional design due to their structural editability and resolution independence. However, this task remains challenging as it requires bridging discrete code representations with continuous visual dynamics. Existing optimization-…

Cited by 0SourceScholar
2026

Wanderland: Geometrically Grounded Simulation for Open-World Embodied AI

CVPR 2026

Reproducible closed-loop evaluation remains a major bottleneck in Embodied AI such as visual navigation. A promising path forward is high-fidelity simulation that combines photorealistic sensor rendering with geometrically grounded interaction in complex, open-world urban environments. Although rece

Cited by 0SourcecodeScholar
2025

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking

ICCV 2025poster

Vision-language tracking aims to locate the target object in the video sequence using a template patch and a language description provided in the initial frame. To achieve robust tracking, especially in complex long-term scenarios that reflect real-world conditions as recently highlighted by MGIT, i…

2025

Adversarial Exploitation of Data Diversity Improves Visual Localization

ICCV 2025poster

Visual localization, which estimates a camera's pose within a known scene, is a fundamental capability for autonomous systems. While absolute pose regression (APR) methods have shown promise for efficient inference, they often struggle with generalization. Recent approaches attempt to address this t…

Cited by 0SourcePDFScholar
2025

BEVTrack: A Simple and Strong Baseline for 3D Single Object Tracking in Bird's-Eye View

IJCAI 2025

3D Single Object Tracking (SOT) is a fundamental task in computer vision and plays a critical role in applications like autonomous driving. However, existing algorithms often involve complex designs and multiple loss functions, making model training and deployment challenging. Furthermore, their rel

2025

Black Sheep in the Herd: Playing with Spuriously Correlated Attributes for Vision-Language Recognition

ICLR 2025poster

Few-shot adaptation for Vision-Language Models (VLMs) presents a dilemma: balancing in-distribution accuracy with out-of-distribution generalization. Recent research has utilized low-level concepts such as visual attributes to enhance generalization. However, this study reveals that VLMs overly rely…

Cited by 0SourcePDFScholar
2025

Brain-Inspired Spiking Neural Networks for Energy-Efficient Object Detection

CVPR 2025poster

Brain-inspired spiking neural networks (SNNs) have the capability of energy-efficient processing of temporal information. However, leveraging the rich dynamic characteristics of SNNs and prior works in artificial neural networks (ANNs) to construct an effective object detection model for visual task…

Cited by 0SourcePDFScholar
2025

CAD-Coder: Text-to-CAD Generation with Chain-of-Thought and Geometric Reward

NeurIPS 2025poster

In this work, we introduce CAD-Coder, a novel framework that reformulates text-to-CAD as the generation of CadQuery scripts—a Python-based, parametric CAD language. This representation enables direct geometric validation, a richer modeling vocabulary, and seamless integration with existing LLMs. To…

Cited by 0SourceScholar
2025

CARE Transformer: Mobile-Friendly Linear Visual Transformer via Decoupled Dual Interaction

CVPR 2025highlight

Recently, large efforts have been made to design efficient linear-complexity visual Transformers. However, current linear attention models are generally unsuitable to be deployed in resource-constrained mobile devices, due to suffering from either few efficiency gains or significant accuracy drops.…

Cited by 0SourcePDFScholar
2025

CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal Features

ICML 2025poster

Effectively modeling and utilizing spatiotemporal features from RGB and other modalities (e.g., depth, thermal, and event data, denoted as X) is the core of RGB-X tracker design. Existing methods often employ two parallel branches to separately process the RGB and X input streams, requiring the mod…

2025

CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos

CVPR 2025poster

Navigating dynamic urban environments presents significant challenges for embodied agents, requiring advanced spatial reasoning and adherence to common-sense norms. Despite progress, existing visual navigation methods struggle in map-free or off-street settings, limiting the deployment of autonomous…

2025

CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis

ACL 2025long

Current inference scaling methods, such as Self-consistency and Best-of-N, have proven effective in improving the accuracy of LLMs on complex reasoning tasks. However, these methods rely heavily on the quality of candidate responses and are unable to produce correct answers when all candidates are i…

2025

Consistency Rating of Semantic Transparency: an Evaluation Method for Metaphor Competence in Idiom Understanding Tasks

COLING 2025main

Idioms condense complex semantics into fixed phrases, and their meaning is often not directly connected to the literal meaning of their constituent words, making idiom comprehension a test of metaphor competence. Metaphor, as a cognitive process in human beings, has not yet found an effective evalua…

Cited by 0SourcePDFScholar
2025

DDPA-3DVG: Vision-Language Dual-Decoupling and Progressive Alignment for 3D Visual Grounding

IJCAI 2025

3D visual grounding aims to localize target objects in point clouds based on free-form natural language, which often describes both target and reference objects. Effective alignment between visual and text features is crucial for this task. However, existing two-stage methods that rely solely on obj

2025

DGSolver: Diffusion Generalist Solver with Universal Posterior Sampling for Image Restoration

NeurIPS 2025poster

Diffusion models have achieved remarkable progress in universal image restoration. However, existing methods perform naive inference in the reverse process, which leads to cumulative errors under limited sampling steps and large step intervals. Moreover, they struggle to balance the commonality of d…

Cited by 0SourcecodeScholar
2025

Diffusion Actor-Critic: Formulating Constrained Policy Iteration as Diffusion Noise Regression for Offline Reinforcement Learning

ICLR 2025poster

In offline reinforcement learning, it is necessary to manage out-of-distribution actions to prevent overestimation of value functions. One class of methods, the policy-regularized method, addresses this problem by constraining the target policy to stay close to the behavior policy. Although several…

2025

Dynamic Parallel Tree Search for Efficient LLM Reasoning

ACL 2025long

Tree of Thoughts (ToT) enhances Large Language Model (LLM) reasoning by structuring problem-solving as a spanning tree. However, recent methods focus on search accuracy while overlooking computational efficiency. The challenges of accelerating the ToT lie in the frequent switching of reasoning focus…

2025

Dynamic Scaling of Unit Tests for Code Reward Modeling

ACL 2025long

Current large language models (LLMs) often struggle to produce accurate responses on the first attempt for complex reasoning tasks like code generation. Prior research tackles this challenge by generating multiple candidate solutions and validating them with LLM-generated unit tests. The execution r…

Cited by 0SourcePDFScholar
2025

Empowering LLMs to Understand and Generate Complex Vector Graphics

CVPR 2025poster

The unprecedented advancements in Large Language Models (LLMs) have profoundly impacted natural language processing but have yet to fully embrace the realm of scalable vector graphics (SVG) generation. While LLMs encode partial knowledge of SVG data from web pages during training, recent findings su…

2025

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues

ICASSP 2025accepted

Vision-Language Tracking (VLT) aims to localize a target in video sequences using a visual template and language description. While textual cues enhance tracking potential, current datasets typically contain much more image data than text, limiting the ability of VLT methods to align the two modalit…

Cited by 0SourceScholar
2025

FacLens: Transferable Probe for Foreseeing Non-Factuality in Fact-Seeking Question Answering of Large Language Models

EMNLP 2025

Despite advancements in large language models (LLMs), non-factual responses still persist in fact-seeking question answering. Unlike extensive studies on post-hoc detection of these responses, this work studies non-factuality prediction (NFP), predicting whether an LLM will generate a non-factual re

2025

FlightPatchNet: Multi-Scale Patch Network with Differential Coding for Short-Term Flight Trajectory Prediction

UAI 2025

Accurate multi-step flight trajectory prediction plays an important role in Air Traffic Control, which can ensure the safety of air transportation. Two main issues limit the flight trajectory prediction performance of existing works. The first issue is the negative impact on prediction accuracy caus

Cited by 0SourcePDFScholar
2025

Fusionsense: Bridging Common Sense, Vision, and Touch for Robust Sparse-View Reconstruction

ICRA 2025

Humans effortlessly integrate common-sense knowledge with sensory input from vision and touch to understand their surroundings. Emulating this capability, we introduce FusionSense, a novel 3D reconstruction framework that enables robots to fuse priors from foundation models with highly sparse observ

Cited by 6SourceScholar
2025

GARF: Learning Generalizable 3D Reassembly for Real-World Fractures

ICCV 2025poster

3D reassembly is a challenging spatial intelligence task with broad applications across scientific domains. While large-scale synthetic datasets have fueled promising learning-based approaches, their generalizability to different domains is limited. Critically, it remains uncertain whether models tr…

Cited by 0SourcePDFScholar
2025

GazeScope: A Framework of Gaze Attention-Based Automatic Field-of-View Adjustment for Laparoscopic Robots

RA-L 2025

The procedure of laparoscopic minimally invasive surgery (MIS) heavily relies on the effective and efficient adjustment of the laparoscopic field-of-view (FoV). However, most existing robot-assisted laparoscopic FoV adjustment methods either require additional surgeon interactions or neglect surgeon

Cited by 3SourceScholar
2025

GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution

NeurIPS 2025spotlight

Ultra-high-resolution (UHR) remote sensing (RS) imagery offers valuable data for Earth observation but pose challenges for existing multimodal foundation models due to two key bottlenecks: (1) limited availability of UHR training data, and (2) token explosion caused by the large image size. To addre…

Cited by 0SourcecodeScholar
2025

Harnessing Massive Satellite Imagery with Efficient Masked Image Modeling

ICCV 2025poster

Masked Image Modeling (MIM) has become an essential method for building foundational visual models in remote sensing (RS). However, the limitations in size and diversity of existing RS datasets restrict the ability of MIM methods to learn generalizable representations. Additionally, conventional MIM…

2025

Human-Imperceptible, Machine-Recognizable Images

IJCAI 2025

Massive human-related data is collected to train neural networks for computer vision tasks. A major conflict is exposed relating to software engineers between better developing AI systems and distancing from the sensitive training data. To reconcile this conflict, the paper proposes an efficient pri

2025

Identifying and Mitigating Position Bias of Multi-image Vision-Language Models

CVPR 2025poster

The evolution of Large Vision-Language Models (LVLMs) has progressed from single-image understanding to multi-image reasoning. Despite this advancement, our findings indicate that LVLMs struggle to robustly utilize information across multiple images, with predictions significantly affected by the al…

2025

L-Diffusion: Laplace Diffusion for Efficient Pathology Image Segmentation

ICML 2025poster

Pathology image segmentation plays a pivotal role in artificial digital pathology diagnosis and treatment. Existing approaches to pathology image segmentation are hindered by labor-intensive annotation processes and limited accuracy in tail-class identification, primarily due to the long-tail distri…

2025

MOCID: Motion Context and Displacement Information Learning for Moving Infrared Small Target Detection

AAAI 2025technical

In the field of Moving Infrared Small Target Detection (MIRSTD), current methods typically use sequential modeling with two individual modules for spatial and temporal processing. However, such a modeling strategy lacks clear guidance on the motion and displacement difference between moving targets…

2025

MOL-Mamba: Enhancing Molecular Representation with Structural & Electronic Insights

AAAI 2025technical

Molecular representation learning plays a crucial role in various downstream tasks, such as molecular property prediction and drug design. To accurately represent molecules, Graph Neural Networks (GNNs) and Graph Transformers (GTs) have shown potential in the realm of self-supervised pretraining. Ho…

2025

MapNav: A Novel Memory Representation via Annotated Semantic Maps for VLM-based Vision-and-Language Navigation

ACL 2025long

Vision-language navigation (VLN) is a key task in Embodied AI, requiring agents to navigate diverse and unseen environments while following natural language instructions. Traditional approaches rely heavily on historical observations as spatio-temporal contexts for decision making, leading to signif…

2025

Multi-axis Prompt and Multi-dimension Fusion Network for All-in-one Weather-degraded Image Restoration

AAAI 2025technical

Existing approaches aiming to remove adverse weather degradations compromise the image quality and incur the long processing time. To this end, we introduce a multi-axis prompt and multi-dimension fusion network (MPMF-Net). Specifically, we develop a multi-axis prompts learning block (MPLB), which l…

2025

Open-Vocabulary Fine-Grained Hand Action Detection

IJCAI 2025

In this work, we address the new challenge of open-vocabulary fine-grained hand action detection, which aims to recognize hand actions from both known and novel categories using textual descriptions. Traditional hand action detection methods are limited to closed-set detection, making it difficult f

Cited by 0SourcePDFScholar
2025

P2 Law: Scaling Law for Post-Training After Model Pruning

ACL 2025long

Pruning has become a widely adopted technique for reducing the hardware requirements of large language models (LLMs). To recover model performance after pruning, post-training is commonly employed to mitigate the resulting performance degradation. While post-training benefits from larger datasets, o…

Cited by 0SourcePDFScholar
2025

ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization

NeurIPS 2025poster

The optimal bit-width for achieving the best trade-off between quantized model size and accuracy has been a subject of ongoing debate. While some advocate for 4-bit quantization, others propose that 1.58-bit offers superior results. However, the lack of a cohesive framework for different bits has le…

Cited by 0SourceScholar
2025

Patch-level Sounding Object Tracking for Audio-Visual Question Answering

AAAI 2025technical

Answering questions related to audio-visual scenes, i.e., the AVQA task, is becoming increasingly popular. A critical challenge is accurately identifying and tracking sounding objects related to the question along the timeline. In this paper, we present a new Patch-level Sounding Object Tracking (PS…

Cited by 6SourcePDFScholar
2025

Probability Density Geodesics in Image Diffusion Latent Space

CVPR 2025poster

Diffusion models indirectly estimate the probability density over a data space, which can be used to study its structure. In this work, we show that geodesics can be computed in diffusion latent space, where the norm induced by the spatially-varying inner product is inversely proportional to the pro…

2025

Rethink Sparse Signals for Pose-guided Text-to-image Generation

ICCV 2025poster

Recent works favored dense signals (e.g., depth, DensePose), as an alternative to sparse signals (e.g., OpenPose), to provide detailed spatial guidance for pose-guided text-to-image generation. However, dense representations raised new challenges including editing difficulties and potential inconsis…

2025

RoMA: Scaling up Mamba-based Foundation Models for Remote Sensing

NeurIPS 2025poster

Recent advances in self-supervised learning for Vision Transformers (ViTs) have fueled breakthroughs in remote sensing (RS) foundation models. However, the quadratic complexity of self-attention poses a significant barrier to scalability, particularly for large models and high-resolution images. Whi…

Cited by 0SourcecodeScholar
2025

SAIST: Segment Any Infrared Small Target Model Guided by Contrastive Language-Image Pretraining

CVPR 2025poster

Infrared Small Target Detection (IRSTD) aims to identify low signal-to-noise ratio small targets in infrared images with complex backgrounds, which is crucial for various applications. However, existing IRSTD methods typically rely solely on image modalities for processing, which fail to fully captu…

Cited by 0SourcePDFScholar
2025

SAM Decoding: Speculative Decoding via Suffix Automaton

ACL 2025long

Speculative decoding (SD) has been demonstrated as an effective technique for lossless LLM inference acceleration.Retrieval-based SD methods, one kind of model-free method, have yielded promising speedup, but they often rely on single retrieval resources, inefficient retrieval methods, and are const…

2025

SafeMap: Robust HD Map Construction from Incomplete Observations

ICML 2025poster

Robust high-definition (HD) map construction is vital for autonomous driving, yet existing methods often struggle with incomplete multi-view camera data. This paper presents SafeMap, a novel framework specifically designed to ensure accuracy even when certain camera views are missing. SafeMap integr…

Cited by 0SourcePDFScholar
2025

Self-calibration Enhanced Whole Slide Pathology Image Analysis

IJCAI 2025

Pathology images are considered the ``gold standard" for cancer diagnosis and treatment, with gigapixel images providing extensive tissue and cellular information. Existing methods fail to simultaneously extract global structural and local detail features for comprehensive pathology image analysis e

Cited by 0SourcePDFScholar
2025

Semi-supervised Infrared Small Target Detection with Thermodynamic-Inspired Uneven Perturbation and Confidence Adaptation

AAAI 2025technical

Single-frame Infrared Small Target (SIRST) detection has made significant advancements, but it still faces challenges due to limited labeled data and the foreground-background class imbalance. To address these issues, we introduce a novel Semi-Supervised SIRST Detection (S^3D) pipeline in this paper…

Cited by 0SourcePDFScholar
2025

Streamlining Redundant Layers to Compress Large Language Models

ICLR 2025spotlight

This paper introduces LLM-Streamline, a pioneer work on layer pruning for large language models (LLMs). It is based on the observation that different layers have varying impacts on hidden states, enabling the identification of less important layers to be pruned. LLM-Streamline comprises two parts:…

2025

Structural-Aware Disentangled Learning with CLIP for Hyperbolic Zero-Shot Sketch-Based Image Retrieval

ICASSP 2025accepted

The zero-shot sketch-based image retrieval task faces two key challenges: domain gap and knowledge transfer. Our innovation is recognizing that directly aligning cross-domain features weakens the discriminative ability of the model, as it overlooks the asymmetry between sketches and images. Addition…

Cited by 0SourceScholar
2025

Synergistic Prompting for Robust Visual Recognition with Missing Modalities

ICCV 2025poster

Large-scale multi-modal models have demonstrated remarkable performance across various visual recognition tasks by leveraging extensive paired multi-modal training data. However, in real-world applications, the presence of missing or incomplete modality inputs often leads to significant performance…

Cited by 0SourcePDFScholar
2025

TableLLM: Enabling Tabular Data Manipulation by LLMs in Real Office Usage Scenarios

ACL 2025finding

We introduce TableLLM, a robust large language model (LLM) with 8 billion parameters, purpose-built for proficiently handling tabular data manipulation tasks, whether they are embedded within documents or spreadsheets, catering to real-world office scenarios. We propose a distant supervision method…

2025

UAWTrack: Universal 3D Single Object Tracking in Adverse Weather

AAAI 2025technical

3D single object tracking (3D SOT) in LiDAR point clouds is essential for autonomous driving. Most existing 3D SOT methods focus on clear weather, where point clouds are more defined. However, adverse weather conditions lead to sparser and noisier point clouds, significantly degrading tracking perfo…

2025

UICOMPASS: UI Map Guided Mobile Task Automation via Adaptive Action Generation

EMNLP 2025

Mobile task automation is an emerging technology that leverages AI to automatically execute routine tasks by users’ commands on mobile devices like Android, thus enhancing efficiency and productivity. While large language models (LLMs) excel at general mobile tasks through training on massive datase

2025

Uncovering the Impact of Chain-of-Thought Reasoning for Direct Preference Optimization: Lessons from Text-to-SQL

ACL 2025long

Direct Preference Optimization (DPO) has proven effective in complex reasoning tasks like math word problems and code generation. However, when applied to Text-to-SQL datasets, it often fails to improve performance and can even degrade it. Our investigation reveals the root cause: unlike math and co…

2025

What Makes for Text to 360-degree Panorama Generation with Stable Diffusion?

ICCV 2025poster

Recent prosperity of text-to-image diffusion models, e.g. Stable Diffusion, has stimulated research to adapt them to 360-degree panorama generation. Prior work has demonstrated the feasibility of using conventional low-rank adaptation techniques on pre-trained diffusion models to generate panoramic…

2025

When CLIP Meets PHOC: A Dual-Branch Network for Historical Document Image Retrieval

ICASSP 2025accepted

In this paper, we leverage Contrastive Language-Image Pre-training (CLIP) for Historical Document Image Retrieval (HDIR). We are largely inspired by recent advances on CLIP and its exceptional generalization capabilities, but for the first time, we tailor it to benefit HDIR. We put forward a dual-br…

Cited by 0SourceScholar
2025

XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?

CVPR 2025highlight

The astonishing breakthrough of multimodal large language models (MLLMs) has necessitated new benchmarks to quantitatively assess their capabilities, reveal their limitations, and indicate future research directions. However, this is challenging in the context of remote sensing (RS), since the image…

2024

A Cause-Effect Look at Alleviating Hallucination of Knowledge-grounded Dialogue Generation

COLING 2024main

Empowered by the large-scale pretrained language models, existing dialogue systems have demonstrated impressive performance conducting fluent and natural-sounding conversations. However, they are still plagued by the <b>hallucination</b> problem, causing unpredictable factual errors in the generated…

2024

Adversarial Purification with the Manifold Hypothesis

AAAI 2024technical

In this work, we formulate a novel framework for adversarial robustness using the manifold hypothesis. This framework provides sufficient conditions for defending against adversarial examples. We develop an adversarial purification method with this framework. Our method combines manifold learning wi…

2024

AlignBench: Benchmarking Chinese Alignment of Large Language Models

ACL 2024long

Alignment has become a critical step for instruction-tuned Large Language Models (LLMs) to become helpful assistants. However, effective evaluation of alignment for emerging Chinese LLMs is still significantly lacking, calling for real-scenario grounded, open-ended, challenging and automatic evaluat…

2024

Automatic Field of View Adjustment of an RCM Constraint-Free Continuum Laparoscopic Robot

IROS 2024poster

Automatic laparoscopic field of view (FOV) adjustment can effectively assist surgeons in minimally invasive surgery (MIS). However, existing work based on rod-shaped laparoscopes is inevitably constrained by the remote center of motion (RCM) during the process of FOV adjustment. The RCM limits lapar…

Cited by 2SourceScholar
2024

BEVNav: Robot Autonomous Navigation via Spatial-Temporal Contrastive Learning in Bird's-Eye View

RA-L 2024

Goal-driven mobile robot navigation in map-less environments requires effective state representations for reliable decision-making. Inspired by the favorable properties of Bird's-Eye View (BEV) in point clouds for visual perception, this paper introduces a novel navigation approach named BEVNav. It

Cited by 10SourceScholar
2024

Beyond Accuracy: Tracking more like Human via Visual Search

NeurIPS 2024poster

Human visual search ability enables efficient and accurate tracking of an arbitrary moving target, which is a significant research interest in cognitive neuroscience. The recently proposed Central-Peripheral Dichotomy (CPD) theory sheds light on how humans effectively process visual information and…

2024

Cross-Modal Feature Distribution Calibration for Few-Shot Visual Question Answering

AAAI 2024technical

Few-shot Visual Question Answering (VQA) realizes few-shot cross-modal learning, which is an emerging and challenging task in computer vision. Currently, most of the few-shot VQA methods are confined to simply extending few-shot classification methods to cross-modal tasks while ignoring the spatial…

Cited by 3SourcePDFScholar
2024

Data-Free Generalized Zero-Shot Learning

AAAI 2024technical

Deep learning models have the ability to extract rich knowledge from large-scale datasets. However, the sharing of data has become increasingly challenging due to concerns regarding data copyright and privacy. Consequently, this hampers the effective transfer of knowledge from existing data to novel…

2024

Deciphering Rumors: A Multi-Task Learning Approach with Intent-aware Hierarchical Contrastive Learning

EMNLP 2024main

Social networks are rife with noise and misleading information, presenting multifaceted challenges for rumor detection. In this paper, from the perspective of human cognitive subjectivity, we introduce the mining of individual latent intentions and propose a novel multi-task learning framework, the…

Cited by 1SourcePDFScholar
2024

Decomposing Semantic Shifts for Composed Image Retrieval

AAAI 2024technical

Composed image retrieval is a type of image retrieval task where the user provides a reference image as a starting point and specifies a text on how to shift from the starting point to the desired target image. However, most existing methods focus on the composition learning of text and reference im…

2024

Disentangling Domain and General Representations for Time Series Classification

IJCAI 2024poster

Modeling time series data has become a very at tractive research topic due to its wide application, such as human activity recognition, financial forecasting and sensor-based automatic system monitoring. Recently deep learning models have shown great advances in modeling the time series data but the…

2024

Distilling Causal Effect of Data in Continual Few-shot Relation Learning

COLING 2024main

Continual Few-Shot Relation Learning (CFRL) aims to learn an increasing number of new relational patterns from a data stream. However, due to the limited number of samples and the continual training mode, this method frequently encounters the catastrophic forgetting issues. The research on causal in…

2024

Diversifying Question Generation over Knowledge Base via External Natural Questions

COLING 2024main

Previous methods on knowledge base question generation (KBQG) primarily focus on refining the quality of a single generated question. However, considering the remarkable paraphrasing ability of humans, we believe that diverse texts can express identical semantics through varied expressions. The abov…

2024

DreamSteerer: Enhancing Source Image Conditioned Editability using Personalized Diffusion Models

NeurIPS 2024poster

Recent text-to-image (T2I) personalization methods have shown great premise in teaching a diffusion model user-specified concepts given a few images for reusing the acquired concepts in a novel context. With massive efforts being dedicated to personalized generation, a promising extension is persona…

2024

Encoder-Minimal and Decoder-Minimal Framework for Remote Sensing Image Dehazing

ICASSP 2024accepted

Haze obscures remote sensing images, hindering valuable information extraction. To this end, we propose RSHazeNet, an encoder-minimal and decoder-minimal framework for efficient remote sensing image dehazing. Specifically, regarding the process of merging features within the same level, we develop a…

Cited by 0SourceScholar
2024

GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching

NeurIPS 2024poster

Beyond the text detection and recognition tasks in image text spotting, video text spotting presents an augmented challenge with the inclusion of tracking. While advanced end-to-end trainable methods have shown commendable performance, the pursuit of multi-task optimization may pose the risk of prod…

2024

HENet: Hyperbolic-Based Encoder-Decoder Network for Word Spotting in Historical Mongolian Documents

ICASSP 2024accepted

In the domain of historical Mongolian document image retrieval (HMDIR), word spotting poses a inherent challenge due to the frequent appearance of out-of-vocabulary (OOV) words. Existing methods have mainly focused on query-by-example (QBE), neglecting the query-by-string (QBS) approach. Meanwhile,…

Cited by 0SourceScholar
2024

IMPUS: Image Morphing with Perceptually-Uniform Sampling Using Diffusion Models

ICLR 2024poster

We present a diffusion-based image morphing approach with perceptually-uniform sampling (IMPUS) that produces smooth, direct and realistic interpolations given an image pair. The embeddings of two images may lie on distinct conditioned distributions of a latent diffusion model, especially when they…

2024

IRPruneDet: Efficient Infrared Small Target Detection via Wavelet Structure-Regularized Soft Channel Pruning

AAAI 2024technical

Infrared Small Target Detection (IRSTD) refers to detecting faint targets in infrared images, which has achieved notable progress with the advent of deep learning. However, the drive for improved detection accuracy has led to larger, intricate models with redundant parameters, causing storage and co…

2024

IRSAM: Advancing Segment Anything Model for Infrared Small Target Detection

ECCV 2024poster

"The recent Segment Anything Model (SAM) is a significant advancement in natural image segmentation, exhibiting potent zero-shot performance suitable for various downstream image segmentation tasks. However, directly utilizing the pretrained SAM for Infrared Small Target Detection (IRSTD) task falls…

2024

Is Your HD Map Constructor Reliable under Sensor Corruptions?

NeurIPS 2024poster

Driving systems often rely on high-definition (HD) maps for precise environmental information, which is crucial for planning and navigation. While current HD map constructors perform well under ideal conditions, their resilience to real-world challenges, \eg, adverse weather and sensor failures, is…

Cited by 17SourcePDFScholar
2024

LA-UCL: LLM-Augmented Unsupervised Contrastive Learning Framework for Few-Shot Text Classification

COLING 2024main

The few-shot tasks require the model to have the ability to generalize from a few samples. However, due to the lack of cognitive ability, the current works cannot fully utilize limited samples to expand the sample space and still suffer from overfitting issues. To address the problems, we propose a…

Cited by 11SourcePDFScholar
2024

LUWA Dataset: Learning Lithic Use-Wear Analysis on Microscopic Images

CVPR 2024highlight

Lithic Use-Wear Analysis (LUWA) using microscopic images is an underexplored vision-for-science research area. It seeks to distinguish the worked material which is critical for understanding archaeological artifacts material interactions tool functionalities and dental records. However this challeng…

Cited by 4SourcePDFScholar
2024

Latent Optimal Paths by Gumbel Propagation for Variational Bayesian Dynamic Programming

ICML 2024poster

We propose the stochastic optimal path which solves the classical optimal path problem by a probability-softening solution. This unified approach transforms a wide range of DP problems into directed acyclic graphs in which all paths follow a Gibbs distribution. We show the equivalence of the Gibbs d…

2024

LeMeViT: Efficient Vision Transformer with Learnable Meta Tokens for Remote Sensing Image Interpretation

IJCAI 2024poster

Due to spatial redundancy in remote sensing images, sparse tokens containing rich information are usually involved in self-attention (SA) to reduce the overall token numbers within the calculation, avoiding the high computational cost issue in Vision Transformers. However, such methods usually obtai…

2024

MapDistill: Boosting Efficient Camera-based HD Map Construction via Camera-LiDAR Fusion Model Distillation

ECCV 2024poster

"Online high-definition (HD) map construction is an important and challenging task in autonomous driving. Recently, there has been a growing interest in cost-effective multi-view camera-based methods without relying on other sensors like LiDAR. However, these methods suffer from a lack of explicit d…

Cited by 14SourcePDFScholar
2024

MemVLT: Vision-Language Tracking with Adaptive Memory-based Prompts

NeurIPS 2024poster

Vision-language tracking (VLT) enhances traditional visual object tracking by integrating language descriptions, requiring the tracker to flexibly understand complex and diverse text in addition to visual information. However, most existing vision-language trackers still overly rely on initial fixed…

Cited by 5SourcePDFScholar
2024

Multi-Dimension Queried and Interacting Network for Stereo Image Deraining

ICASSP 2024accepted

Eliminating the rain degradation in stereo images poses a formidable challenge, which necessitates the efficient exploitation of mutual information present between the dual views. To this end, we devise MQINet, which employs multi-dimension queries and interactions for stereo image deraining. More s…

Cited by 0SourceScholar
2024

Multi-Modality Affinity Inference for Weakly Supervised 3D Semantic Segmentation

AAAI 2024technical

3D point cloud semantic segmentation has a wide range of applications. Recently, weakly supervised point cloud segmentation methods have been proposed, aiming to alleviate the expensive and laborious manual annotation process by leveraging scene-level labels. However, these methods have not effectiv…

2024

Object-Aware Adaptive-Positivity Learning for Audio-Visual Question Answering

AAAI 2024technical

This paper focuses on the Audio-Visual Question Answering (AVQA) task that aims to answer questions derived from untrimmed audible videos. To generate accurate answers, an AVQA model is expected to find the most informative audio-visual clues relevant to the given questions. In this paper, we propos…

2024

OxyGenerator: Reconstructing Global Ocean Deoxygenation Over a Century with Deep Learning

ICML 2024poster

Accurately reconstructing the global ocean deoxygenation over a century is crucial for assessing and protecting marine ecosystem. Existing expert-dominated numerical simulations fail to catch up with the dynamic variation caused by global warming and human activities. Besides, due to the high-cost d…

Cited by 5SourcePDFScholar
2024

PCQPR: Proactive Conversational Question Planning with Reflection

EMNLP 2024main

Conversational Question Generation (CQG) enhances the interactivity of conversational question-answering systems in fields such as education, customer service, and entertainment. However, traditional CQG, focusing primarily on the immediate context, lacks the conversational foresight necessary to gu…

Cited by 2SourcePDFScholar
2024

PowerPM: Foundation Model for Power Systems

NeurIPS 2024poster

The proliferation of abundant electricity time series (ETS) data presents numerous opportunities for various applications within power systems, including demand-side management, grid stability, and consumer behavior analysis. Deep learning models have advanced ETS modeling by effectively capturing s…

2024

Q-Distribution guided Q-learning for offline reinforcement learning: Uncertainty penalized Q-value via consistency model

NeurIPS 2024poster

``Distribution shift'' is the primary obstacle to the success of offline reinforcement learning. As a learning policy may take actions beyond the knowledge of the behavior policy (referred to as Out-of-Distribution (OOD) actions), the Q-values of these OOD actions can be easily overestimated. Conseq…

2024

Quantum-Inspired Neural Network with Runge-Kutta Method

AAAI 2024technical

In recent years, researchers have developed novel Quantum-Inspired Neural Network (QINN) frameworks for the Natural Language Processing (NLP) tasks, inspired by the theoretical investigations of quantum cognition. However, we have found that the training efficiency of QINNs is significantly lower th…

Cited by 2SourcePDFScholar
2024

Question Calibration and Multi-Hop Modeling for Temporal Question Answering

AAAI 2024technical

Many models that leverage knowledge graphs (KGs) have recently demonstrated remarkable success in question answering (QA) tasks. In the real world, many facts contained in KGs are time-constrained thus temporal KGQA has received increasing attention. Despite the fruitful efforts of previous models i…

Cited by 6SourcePDFScholar
2024

SGSH: Stimulate Large Language Models with Skeleton Heuristics for Knowledge Base Question Generation

NAACL 2024findings

Knowledge base question generation (KBQG) aims to generate natural language questions from a set of triplet facts extracted from KB. Existing methods have significantly boosted the performance of KBQG via pre-trained language models (PLMs) thanks to the richly endowed semantic knowledge. With the ad…

2024

SVGDreamer: Text Guided SVG Generation with Diffusion Model

CVPR 2024poster

Recently text-guided scalable vector graphics (SVGs) synthesis has shown promise in domains such as iconography and sketch. However existing text-to-SVG generation methods lack editability and struggle with visual quality and result diversity. To address these limitations we propose a novel text-gui…

2024

SimDistill: Simulated Multi-Modal Distillation for BEV 3D Object Detection

AAAI 2024technical

Multi-view camera-based 3D object detection has become popular due to its low cost, but accurately inferring 3D geometry solely from camera data remains challenging and may lead to inferior performance. Although distilling precise 3D geometry knowledge from LiDAR data could help tackle this challeng…

2024

SoundLoCD: An Efficient Conditional Discrete Contrastive Latent Diffusion Model for Text-to-Sound Generation

ICASSP 2024accepted

We present SoundLoCD, a novel text-to-sound generation framework, which incorporates a LoRA-based conditional discrete contrastive latent diffusion model. Unlike recent large-scale sound generation models, our model can be efficiently trained under limited computational resources. The integration of…

Cited by 0SourceScholar
2024

SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation

NeurIPS 2024spotlight

We introduce SpreadsheetBench, a challenging spreadsheet manipulation benchmark exclusively derived from real-world scenarios, designed to immerse current large language models (LLMs) in the actual workflow of spreadsheet users. Unlike existing benchmarks that rely on synthesized queries and simpli…

Cited by 5SourcePDFScholar
2024

SurgicalSAM: Efficient Class Promptable Surgical Instrument Segmentation

AAAI 2024technical

The Segment Anything Model (SAM) is a powerful foundation model that has revolutionised image segmentation. To apply SAM to surgical instrument segmentation, a common approach is to locate precise points or boxes of instruments and then use them as prompts for SAM in a zero-shot manner. However, we…

2024

Training A Small Emotional Vision Language Model for Visual Art Comprehension

ECCV 2024poster

"This paper develops small vision language models to understand visual art, which, given an art work, aims to identify its emotion category and explain this prediction with natural language. While small models are computationally efficient, their capacity is much limited compared with large models.…

2024

Transferable and Efficient Non-Factual Content Detection via Probe Training with Offline Consistency Checking

ACL 2024long

This paper proposes PiNose, which trains a probing model on offline self-consistency checking results, thereby circumventing the need for human-annotated data and achieving transferability across diverse data distributions. As the consistency check process is offline, PiNose reduces the computationa…

2024

UniMix: Towards Domain Adaptive and Generalizable LiDAR Semantic Segmentation in Adverse Weather

CVPR 2024poster

LiDAR semantic segmentation (LSS) is a critical task in autonomous driving and has achieved promising progress. However prior LSS methods are conventionally investigated and evaluated on datasets within the same domain in clear weather. The robustness of LSS models in unseen scenes and all weather c…

Cited by 46SourcePDFScholar
2023

A Generation-based Deductive Method for Math Word Problems

EMNLP 2023long main

Math word problems (MWP) involving advanced operators such as linear equation solver cannot be easily tackled by earlier MWP methods, because the existing generation methods suffer from repeated sub-expression generation and deductive methods are restricted to dealing with binary operations. This pa…

Cited by 0SourcecodeScholar
2023

Constrained Policy Optimization with Explicit Behavior Density For Offline Reinforcement Learning

NeurIPS 2023poster

Due to the inability to interact with the environment, offline reinforcement learning (RL) methods face the challenge of estimating the Out-of-Distribution (OOD) points. Existing methods for addressing this issue either control policy to exclude the OOD action or make the $Q$ function pessimistic. H…

2023

DPText-DETR: Towards Better Scene Text Detection with Dynamic Points in Transformer

AAAI 2023technical

Recently, Transformer-based methods, which predict polygon points or Bezier curve control points for localizing texts, are popular in scene text detection. However, these methods built upon detection transformer framework might achieve sub-optimal training efficiency and performance due to coarse po…

2023

Decoupling Learning and Remembering: A Bilevel Memory Framework With Knowledge Projection for Task-Incremental Learning

CVPR 2023poster

The dilemma between plasticity and stability arises as a common challenge for incremental learning. In contrast, the human memory system is able to remedy this dilemma owing to its multi-level memory structure, which motivates us to propose a Bilevel Memory system with Knowledge Projection (BMKP) fo…

2023

DeepSolo: Let Transformer Decoder With Explicit Points Solo for Text Spotting

CVPR 2023poster

End-to-end text spotting aims to integrate scene text detection and recognition into a unified framework. Dealing with the relationship between the two sub-tasks plays a pivotal role in designing effective spotters. Although Transformer-based methods eliminate the heuristic post-processing, they sti…

2023

DiffSketcher: Text Guided Vector Sketch Synthesis through Latent Diffusion Models

NeurIPS 2023poster

Even though trained mainly on images, we discover that pretrained diffusion models show impressive power in guiding sketch synthesis. In this paper, we present DiffSketcher, an innovative algorithm that creates \textit{vectorized} free-hand sketches using natural language input. DiffSketcher is deve…

2023

Domain Specified Optimization for Deployment Authorization

ICCV 2023poster

This paper explores Deployment Authorization (DPA) as a means of restricting the generalization capabilities of vision models on certain domains to protect intellectual property. Nevertheless, the current advancements in DPA are predominantly confined to fully supervised settings. Such settings requ…

Cited by 8PDFScholar
2023

Dual Path Modeling for Semantic Matching by Perceiving Subtle Conflicts

ICASSP 2023accepted

Transformer-based pre-trained models have achieved great improvements in semantic matching. However, existing models still suffer from insufficient ability to capture subtle differences. The modification, addition and deletion of words in sentence pairs may make it difficult for the model to predict…

Cited by 0SourceScholar
2023

Dynamic Focus-Aware Positional Queries for Semantic Segmentation

CVPR 2023poster

The DETR-like segmentors have underpinned the most recent breakthroughs in semantic segmentation, which end-to-end train a set of queries representing the class prototypes or target segments. Recently, masked attention is proposed to restrict each query to only attend to the foreground regions predi…

2023

ESSAformer: Efficient Transformer for Hyperspectral Image Super-resolution

ICCV 2023poster

Single hyperspectral image super-resolution (single-HSI-SR) aims to restore a high-resolution hyperspectral image from a low-resolution observation. However, the prevailing CNN-based approaches have shown limitations in building long-range dependencies and capturing interaction information between s…

Cited by 83PDFcodeScholar
2023

Explicit Boundary Guided Semi-Push-Pull Contrastive Learning for Supervised Anomaly Detection

CVPR 2023poster

Most anomaly detection (AD) models are learned using only normal samples in an unsupervised way, which may result in ambiguous decision boundary and insufficient discriminability. In fact, a few anomaly samples are often available in real-world applications, the valuable knowledge of known anomalies…

2023

FC-KBQA: A Fine-to-Coarse Composition Framework for Knowledge Base Question Answering

ACL 2023long

The generalization problem on KBQA has drawn considerable attention. Existing research suffers from the generalization issue brought by the entanglement in the coarse-grained modeling of the logical expression, or inexecutability issues due to the fine-grained modeling of disconnected classes and re…

2023

FFAEval: Evaluating Dialogue System via Free-For-All Ranking

EMNLP 2023long findings

Evaluating open-domain dialogue systems is currently an open question. Automatic evaluation metrics have shown poor correlation with human assessment in dialogue generation tasks. Human evaluation, which involves annotators for multi-dimension scoring, is trustworthy but time-consuming. In this wor…

Cited by 0SourceScholar
2023

Feature Decomposition for Reducing Negative Transfer: A Novel Multi-Task Learning Method for Recommender System (Student Abstract)

AAAI 2023technical

We propose a novel multi-task learning method termed Feature Decomposition Network (FDN). The key idea of the proposed FDN is to reduce the phenomenon of feature redundancy by explicitly decomposing features into task-specific features and task-shared features with carefully designed constraints. Ex…

Cited by 13SourcePDFScholar
2023

GLT-T: Global-Local Transformer Voting for 3D Single Object Tracking in Point Clouds

AAAI 2023technical

Current 3D single object tracking methods are typically based on VoteNet, a 3D region proposal network. Despite the success, using a single seed point feature as the cue for offset learning in VoteNet prevents high-quality 3D proposals from being generated. Moreover, seed points with different impor…

2023

Leverage Interactive Affinity for Affordance Learning

CVPR 2023poster

Perceiving potential "action possibilities" (i.e., affordance) regions of images and learning interactive functionalities of objects from human demonstration is a challenging task due to the diversity of human-object interactions. Prevailing affordance learning algorithms often adopt the label assig…

2023

MPMQA: Multimodal Question Answering on Product Manuals

AAAI 2023technical

Visual contents, such as illustrations and images, play a big role in product manual understanding. Existing Product Manual Question Answering (PMQA) datasets tend to ignore visual contents and only retain textual parts. In this work, to emphasize the importance of multimodal contents, we propose a…

2023

Model Calibration in Dense Classification with Adaptive Label Perturbation

ICCV 2023poster

For safety-related applications, it is crucial to produce trustworthy deep neural networks whose prediction is associated with confidence that can represent the likelihood of correctness for subsequent decision-making. Existing dense binary classification models are prone to being over-confident. To…

Cited by 4PDFcodeScholar
2023

Modeling the Distributional Uncertainty for Salient Object Detection Models

CVPR 2023poster

Most of the existing salient object detection (SOD) models focus on improving the overall model performance, without explicitly explaining the discrepancy between the training and testing distributions. In this paper, we investigate a particular type of epistemic uncertainty, namely distributional u…

Cited by 23SourcePDFScholar
2023

Multimodal Variational Auto-encoder based Audio-Visual Segmentation

ICCV 2023poster

We propose an Explicit Conditional Multimodal Variational Auto-Encoder (ECMVAE) for audio-visual segmentation (AVS), aiming to segment sound sources in the video sequence. Existing AVS methods focus on implicit feature fusion strategies, where models are trained to fit the discrete samples in the da…

Cited by 42PDFcodeScholar
2023

OSP2B: One-Stage Point-to-Box Network for 3D Siamese Tracking

IJCAI 2023poster

Two-stage point-to-box network acts as a critical role in the recent popular 3D Siamese tracking paradigm, which first generates proposals and then predicts corresponding proposal-wise scores. However, such a network suffers from tedious hyper-parameter tuning and task misalignment, limiting the tra…

2023

P2C: Self-Supervised Point Cloud Completion from Single Partial Clouds

ICCV 2023poster

Point cloud completion aims to recover the complete shape based on a partial observation. Existing methods require either complete point clouds or multiple partial observations of the same object for learning. In contrast to previous approaches, we present Partial2Complete (P2C), the first self-supe…

Cited by 29PDFcodeScholar
2023

RESDSQL: Decoupling Schema Linking and Skeleton Parsing for Text-to-SQL

AAAI 2023technical

One of the recent best attempts at Text-to-SQL is the pre-trained language model. Due to the structural property of the SQL queries, the seq2seq model takes the responsibility of parsing both the schema items (i.e., tables and columns) and the skeleton (i.e., SQL keywords). Such coupled targets incr…

2023

RPEFlow: Multimodal Fusion of RGB-PointCloud-Event for Joint Optical Flow and Scene Flow Estimation

ICCV 2023poster

Recently, the RGB images and point clouds fusion methods have been proposed to jointly estimate 2D optical flow and 3D scene flow. However, as both conventional RGB cameras and LiDAR sensors adopt a frame-based data acquisition mechanism, their performance is limited by the fixed low sampling rates,…

Cited by 24PDFcodeScholar
2023

SAMRS: Scaling-up Remote Sensing Segmentation Dataset with Segment Anything Model

NeurIPS 2023poster

The success of the Segment Anything Model (SAM) demonstrates the significance of data-centric machine learning. However, due to the difficulties and high costs associated with annotating Remote Sensing (RS) images, a large amount of valuable RS data remains unlabeled, particularly at the pixel level…

2023

ST${2}$: Spatial-Temporal State Transformer for Crowd-Aware Autonomous Navigation

RA-L 2023

Empowering an intelligent agent with the ability of autonomous navigation in complex and dynamic environments is an important and active research topic in embodied artificial intelligence. In this letter, we address this challenging task from the view of exploiting both the spatial and temporal stat

Cited by 34SourceScholar
2023

Sensitivity-Aware Visual Parameter-Efficient Fine-Tuning

ICCV 2023oral

Visual Parameter-Efficient Fine-Tuning (PEFT) has become a powerful alternative for full fine-tuning so as to adapt pre-trained vision models to downstream tasks, which only tunes a small number of parameters while freezing the vast majority ones to ease storage burden and optimization difficulty. H…

Cited by 63PDFcodeScholar
2022

"JPerceiver: Joint Perception Network for Depth, Pose and Layout Estimation in Driving Scenes"

ECCV 2022poster

"Depth estimation, visual odometry (VO), and bird’s-eye-view (BEV) scene layout estimation present three critical tasks for driving scene perception, which is fundamental for motion planning and navigation in autonomous driving. Though they are complementary to each other, prior works usually focus…

2022

"Towards Scale-Aware, Robust, and Generalizable Unsupervised Monocular Depth Estimation by Integrating IMU Motion Dynamics"

ECCV 2022poster

"Unsupervised monocular depth and ego-motion estimation has drawn extensive research attention in recent years. Although current methods have reached a high up-to-scale accuracy, they usually fail to learn the true scale metric due to the inherent scale ambiguity from training with monocular sequenc…

2022

3DJCG: A Unified Framework for Joint Dense Captioning and Visual Grounding on 3D Point Clouds

CVPR 2022oral

Observing that the 3D captioning task and the 3D grounding task contain both shared and complementary information in nature, in this work, we propose a unified framework to jointly solve these two distinct but closely related tasks in a synergistic fashion, which consists of both shared task-agnosti…

Cited by 120PDFScholar
2022

APT-36K: A Large-scale Benchmark for Animal Pose Estimation and Tracking

NeurIPS 2022accept

Animal pose estimation and tracking (APT) is a fundamental task for detecting and tracking animal keypoints from a sequence of video frames. Previous animal-related datasets focus either on animal tracking or single-frame animal pose estimation, and never on both aspects. The lack of APT datasets hi…

2022

Audio—Visual Segmentation

ECCV 2022poster

"We propose to explore a new problem called audio-visual segmentation (AVS), in which the goal is to output a pixel-level map of the object(s) that produce sound at the time of the image frame. To facilitate this research, we construct the first audio-visual segmentation benchmark (AVSBench), provid…

2022

BMD: A General Class-Balanced Multicentric Dynamic Prototype Strategy for Source-Free Domain Adaptation

ECCV 2022poster

"Source-free Domain Adaptation (SFDA) aims to adapt a pre-trained source model to the unlabeled target domain without accessing the well-labeled source data, which is a much more practical setting due to the data privacy, security, and transmission issues. To make up for the absence of source data,…

2022

DSM: Question Generation over Knowledge Base via Modeling Diverse Subgraphs with Meta-learner

EMNLP 2022main

Existing methods on knowledge base question generation (KBQG) learn a one-size-fits-all model by training together all subgraphs without distinguishing the diverse semantics of subgraphs. In this work, we show that making use of the past experience on semantically similar subgraphs can reduce the le…

2022

DearKD: Data-Efficient Early Knowledge Distillation for Vision Transformers

CVPR 2022poster

Transformers have been successfully applied to computer vision due to its powerful modelling capacity with self-attention. However, the good performance of transformers heavily depends on enormous training images. Thus, a data-efficient transformer solution is urgently needed. In this work, we propo…

Cited by 99PDFScholar
2022

Energy-Based Generative Cooperative Saliency Prediction

AAAI 2022technical

Conventional saliency prediction models typically learn a deterministic mapping from an image to its saliency map, and thus fail to explain the subjective nature of human attention. In this paper, to model the uncertainty of visual saliency, we study the saliency prediction problem from the perspec…

2022

Exploring Figure-Ground Assignment Mechanism in Perceptual Organization

NeurIPS 2022accept

Perceptual organization is a challenging visual task that aims to perceive and group the individual visual element so that it is easy to understand the meaning of the scene as a whole. Most recent methods building upon advanced Convolutional Neural Network (CNN) come from learning discriminative rep…

Cited by 20SourcePDFScholar
2022

FIBA: Frequency-Injection Based Backdoor Attack in Medical Image Analysis

CVPR 2022poster

In recent years, the security of AI systems has drawn increasing research attention, especially in the medical imaging realm. To develop a secure medical image analysis (MIA) system, it is a must to study possible backdoor attacks (BAs), which can embed hidden malicious behaviors into the system. Ho…

Cited by 123PDFcodeScholar
2022

FP-DETR: Detection Transformer Advanced by Fully Pre-training

ICLR 2022poster

Large-scale pre-training has proven to be effective for visual representation learning on downstream tasks, especially for improving robustness and generalization. However, the recently developed detection transformers only employ pre-training on its backbone while leaving the key component, i.e., a…

2022

FakeCLR: Exploring Contrastive Learning for Solving Latent Discontinuity in Data-Efficient GANs

ECCV 2022poster

"Data-Efficient GANs (DE-GANs), which aim to learn generative models with a limited amount of training data, encounter several challenges for generating high-quality samples. Since data augmentation strategies have largely alleviated the training instability, how to further improve the generative pe…

2022

GMFlow: Learning Optical Flow via Global Matching

CVPR 2022oral

Learning-based optical flow estimation has been dominated with the pipeline of cost volume with convolutions for flow regression, which is inherently limited to local correlations and thus is hard to address the long-standing challenge of large displacements. To alleviate this, the state-of-the-art…

Cited by 459PDFcodeScholar
2022

ISNet: Shape Matters for Infrared Small Target Detection

CVPR 2022poster

Infrared small target detection (IRSTD) refers to extracting small and dim targets from blurred backgrounds, which has a wide range of applications such as traffic management and marine rescue. Due to the low signal-to-noise ratio and low contrast, infrared targets are easily submerged in the backgr…

Cited by 353PDFcodeScholar
2022

Improving RGB-D Point Cloud Registration by Learning Multi-Scale Local Linear Transformation

ECCV 2022poster

"Point cloud registration aims at estimating the geometric transformation between two point cloud scans, in which accurate correspondence estimation is the key to its success. In addition to previous methods that seek correspondences by hand-crafted or learnt geometric features, recent point cloud r…

2022

Knowledge-augmented Self-training of A Question Rewriter for Conversational Knowledge Base Question Answering

EMNLP 2022finding

The recent rise of conversational applications such as online customer service systems and intelligent personal assistants has promoted the development of conversational knowledge base question answering (ConvKBQA). Different from the traditional single-turn KBQA, ConvKBQA usually explores multi-tur…

2022

MeshMAE: Masked Autoencoders for 3D Mesh Data Analysis

ECCV 2022poster

"Recently, self-supervised pre-training has advanced Vision Transformers on various tasks w.r.t. different data modalities, e.g., image and 3D point cloud data. In this paper, we explore this learning paradigm for 3D mesh data analysis based on Transformers. Since applying Transformer architectures…

Cited by 59SourcePDFScholar
2022

Multiview Long-Short Spatial Contrastive Learning For 3D Medical Image Analysis

ICASSP 2022accepted

The success of supervised deep learning heavily depends on large labeled datasets whose construction is often challenging in medical image analysis. Contrastive learning, a variant of self-supervised learning, is a potential solution to alleviate the strong demand for data annotation. In this work,…

Cited by 0SourceScholar
2022

PolyphonicFormer: Unified Query Learning for Depth-Aware Video Panoptic Segmentation

ECCV 2022poster

"The Depth-aware Video Panoptic Segmentation (DVPS) is a new challenging vision problem that aims to predict panoptic segmentation and depth in a video simultaneously. The previous work solves this task by extending the existing panoptic segmentation method with an extra dense depth prediction and i…

2022

RU-Net: Regularized Unrolling Network for Scene Graph Generation

CVPR 2022poster

Scene graph generation (SGG) aims to detect objects and predict the relationships between each pair of objects. Existing SGG methods usually suffer from several issues, including 1) ambiguous object representations, as graph neural network-based message passing (GMP) modules are typically sensitive…

Cited by 52PDFcodeScholar