← Search

Bo Zhang

183 accepted papers

2026

A Phase-Change-Material-Based Variable Stiffness Sheath Inspired by a Multi-Layer Wave Spring Structure for Flexible Upper Gastrointestinal Endoscopic Robots

ICRA 2026poster

Continuum robots employed in flexible gastrointestinal endoscopy require the capability of transitioning between the flexible and the rigid states. Phase-change-material-based variable stiffness (VS) methods exhibit a significant stiffness change ratio but are typically time-consuming. Besides, thes…

Cited by 0SourceScholar
2026

BAMI: Training-Free Bias Mitigation in GUI Grounding

CVPR 2026

GUI grounding is a critical capability for enabling GUI agents to execute tasks such as clicking and dragging. However, in complex scenarios like the ScreenSpot-Pro benchmark, existing models often suffer from suboptimal performance. Utilizing the proposed Masked Prediction Distribution (MPD) attrib

Cited by 0SourcecodeScholar
2026

Beyond Gemini-3-Pro: Revisiting LLM Routing and Aggregation at Scale

ICML 2026poster

Large Language Models (LLMs) have rapidly advanced, with Gemini-3-Pro setting a new performance milestone. In this work, we explore collective intelligence as an alternative to monolithic scaling, and demonstrate that open-source LLMs' collaboration can surpass Gemini-3-Pro. We first revisit LLM rou…

Cited by 0SourceScholar
2026

Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token Interactions

ICML 2026poster

Multimodal Large Language Models (MLLMs) demonstrate impressive cross-modal capabilities, yet their substantial size poses significant deployment challenges. Knowledge distillation (KD) is a promising solution for compressing these models, but existing methods primarily rely on static next-token ali…

Cited by 0SourceScholar
2026

CareCom: Generative Image Composition with Calibrated Reference Features

AAAI 2026technical

Image composition aims to seamlessly insert foreground object into background. Despite the huge progress in generative image composition, the existing methods are still struggling with simultaneous detail preservation and foreground pose/view adjustment. To address this issue, we extend the existin

Cited by 0SourcePDFScholar
2026

Contrastive Cross-Bag Augmentation for Multiple Instance Learning-based Whole Slide Image Classification

CVPR 2026

Recent pseudo-bag augmentation methods for Multiple Instance Learning (MIL)-based Whole Slide Image (WSI) classification sample instances from a limited number of bags, resulting in constrained diversity. To address this issue, we propose Contrastive Cross-Bag Augmentation (C2Aug) to sample instance

Cited by 0SourcecodeScholar
2026

Dual-Level Hypergraph Generation for Addressing Feature Scarcity in Whole-Slide Image Classification

CVPR 2026

Lymph node metastasis diagnosis in pathological images is a highly challenging four-class classification task, comprising macrometastasis, micrometastasis, isolated tumor cells (ITC), and negative lesions.Unlike conventional classification settings, this four-class scenario simultaneously suffers fr

Cited by 0SourcecodeScholar
2026

Force-Centric Bidirectional Cross-Domain Interpretation Between Soft-Body Simulation and Piezoresistive E-Skin

RA-L 2026

Cutaneous tactile feedback is essential for robotic manipulation, particularly when interacting with deformable objects with limited visual sensing. Learning-based approaches are increasingly adopted but rely on large volumes of tactile data. To meet this demand, Sim-to-Real strategies are widely us

Cited by 0SourceScholar
2026

GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation

CVPR 2026

Driving World Models (DWMs) have been developing rapidly with the advances of generative models. However, existing DWMs lack 3D scene understanding capabilities and can only generate content conditioned on input data, without the ability to interpret or reason about the driving environment. Moreover

Cited by 0SourcecodeScholar
2026

GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation

ICLR 2026poster

Geometric problem solving constitutes a critical branch of mathematical reasoning, requiring precise analysis of shapes and spatial relationships. Current evaluations of geometric reasoning in vision-language models (VLMs) face limitations, including the risk of test data contamination from textbook…

Cited by 0SourcecodeScholar
2026

Global-Local Confidence Fusion for Hallucination Detection in Mathematical Reasoning Task

AAAI 2026technical

Large Reasoning Models (LRMs) achieve promising results on complex reasoning tasks but remain susceptible to hallucinations. Existing hallucination detection methods based on Large Language Models (LLMs) often focus solely on final answers, overlooking inconsistencies between the answer and reasonin

Cited by 0SourcePDFScholar
2026

Grounding LLMs in Scientific Discovery via Embodied Actions

ICML 2026poster

Large Language Models (LLMs) have shown significant potential in scientific discovery but struggle to bridge the gap between theoretical reasoning and verifiable physical simulation. Existing solutions operate in a passive "execute-then-response" loop and thus lack runtime perception, obscuring agen…

Cited by 0SourceScholar
2026

LBA: Textual Hard-Label Adversarial Attack Under Low Query Budgets

IJCAI 2026

Generating high-quality adversarial texts with low query budgets remains a challenging problem in the hard-label scenario. Most existing approaches rely on greedy algorithms, where one position in the text is selected for substitution, followed by the substitutions of other positions. This local sea

Cited by 0Scholar
2026

LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have achieved significant advances in integrating visual and linguistic information, yet their ability to reason about complex and real-world scenarios remains limited. Existing benchmarks are usually constructed in a task-oriented manner, without a guarantee…

Cited by 0SourceScholar
2026

LLM-MatLogic: Executable Exchange Contracts for Knowledge-Graph Query Answering with Scoped Negation

ICML 2026poster

LLM-to-KG systems frequently fail on exclusion-rich questions because natural-language negation is both scope-sensitive and evidence-dependent: it may constrain only one subgoal/branch and only certain supporting paths, yet such attachment is rarely explicit in text. We propose the Executable Exchan…

Cited by 0SourceScholar
2026

Learning to LEAP: Efficient Dense Point Tracking by Focusing Where It Matters

AAAI 2026technical

Tracking Any Point (TAP) is a foundational task in computer vision with broad applicability. The state-of-the-art self-supervised TAP method leverages a global matching transformer and contrastive random walks to learn point correspondences. However, its dense all-pairs attention and correlation vol

Cited by 0SourcePDFScholar
2026

ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering

ICML 2026poster

The emergence of large language model (LLM)-based agents has significantly advanced the development of autonomous machine learning (ML) engineering. However, the dominant prompt-based paradigm exhibits limitations: smaller models lack the capacity to learn from execution trajectories for generalizat…

Cited by 0SourcecodeScholar
2026

MME-Reasoning: A Broad-Spectrum Benchmark for Evaluating Logical Reasoning in MLLMs

ICML 2026poster

Logical reasoning is a fundamental aspect of human intelligence and an essential capability for multimodal large language models (MLLMs). Despite the significant advancement in multimodal reasoning, existing benchmarks fail to comprehensively evaluate their reasoning abilities due to the lack of exp…

Cited by 0SourceScholar
2026

MapDream: Task-Driven Map Learning for Vision-Language Navigation

ICML 2026poster

Vision-Language Navigation (VLN) requires agents to follow natural language instructions in partially observed 3D environments, motivating map representations that aggregate spatial context beyond local perception. However, most existing approaches rely on hand-crafted maps constructed independently…

Cited by 0SourceScholar
2026

Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports from Scratch with Agentic Framework

AAAI 2026technical

Visualizations play a crucial part in effective communication of concepts and information. Recent advances in reasoning and retrieval augmented generation have enabled Large Language Models (LLMs) to perform deep research and generate comprehensive reports. Despite its progress, existing deep resear

Cited by 0SourcePDFScholar
2026

OmniGen2: Towards Instruction-Aligned Multimodal Generation

CVPR 2026

Multimodal generative models can process instructions in various modalities and demonstrate outstanding performance across a wide range of image generation tasks. However, their robustness in complex real-world scenarios remains limited due to insufficient generalized instruction alignment. We intro

Cited by 0SourcecodeScholar
2026

Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation

CVPR 2026

Vision-Language Navigation requires agents to act coherently over long horizons by understanding not only local visual context but also how far they have advanced within a multi-step instruction.However, recent Vision-Language-Action models focus on direct action prediction and earlier progress meth

Cited by 0SourceScholar
2026

Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis

ICML 2026poster

The advancement of Medical Vision-Language Models (VLMs) for 3D Computed Tomography (CT) analysis is hindered by a misalignment between optimization objectives and clinical rigor. Current Reinforcement Learning (RL) paradigms rely on lexical proxy signals that induce ``\textbf{evaluation hallucinati…

Cited by 0SourceScholar
2026

SAIL: Self-Amplified Iterative Learning for Diffusion Model Alignment with Minimal Human Feedback

ICLR 2026poster

Aligning diffusion models with human preferences remains challenging, particularly when reward models are unavailable or impractical to obtain, and collecting large-scale preference datasets is prohibitively expensive. This raises a fundamental question: can we achieve effective alignment using only…

Cited by 0SourceScholar
2026

Semore: VLM-guided Enhanced Semantic Motion Representations for Visual Reinforcement Learning

AAAI 2026technical

The growing exploration of Large Language Models (LLM) and Vision-Language Models (VLM) has opened avenues for enhancing the effectiveness of reinforcement learning (RL). However, existing LLM-based RL methods often focus on the guidance of control policy and encounter the challenge of limited repre

Cited by 0SourcePDFScholar
2026

TEMPFLOW-GRPO: WHEN TIMING MATTERS FOR GRPO IN FLOW MODELS

ICLR 2026poster

Recent flow matching models for text-to-image generation have achieved remarkable quality, yet their integration with reinforcement learning for human preference alignment remains suboptimal, hindering fine-grained reward-based optimization. We observe that the key impediment to effective GRPO train…

Cited by 0SourcecodeScholar
2026

THEMIS: Towards Holistic Evaluation of MLLMs for Scientific Paper Fraud Forensics

ICLR 2026poster

We present **THEMIS**, a novel multi-task benchmark designed to comprehensively evaluate Multimodal Large Language Models (MLLMs) on visual fraud reasoning within real-world academic scenarios. Compared to existing benchmarks, THEMIS introduces three major advancements. (1) **Real-world Scenarios &…

Cited by 0SourcecodeScholar
2026

UniMERNet: A Universal Network for Real-World Mathematical Expression Recognition

CVPR 2026

This paper introduces UniMERNet, a high-accuracy, computation-efficient algorithm for Mathematical Expression Recognition (MER) across diverse real-world scenarios. To facilitate UniMERNet's training, we constructed UniMER-1M, a million-scale dataset whose unprecedented diversity endows the model wi

Cited by 0SourcecodeScholar
2026

UniSplat: Unified Spatio-Temporal Fusion via 3D Latent Scaffolds for Dynamic Driving Scene Reconstruction

ICLR 2026poster

Feed-forward 3D reconstruction for autonomous driving has advanced rapidly, yet existing methods struggle with the joint challenges of sparse, non-overlapping camera views and complex scene dynamics. We present UniSplat, a general feed-forward framework that learns robust dynamic scene reconstructi…

Cited by 0SourceScholar
2026

Unveiling the Surprising Efficacy of Navigation Understanding in End-To-End Autonomous Driving

ICRA 2026poster

Global navigation information and local scene understanding are two crucial components of autonomous driving systems. However, our experimental results indicate that many end-to-end autonomous driving systems tend to over-rely on local scene understanding while failing to utilize global navigation i…

2026

VKG-QA: Visual Knowledge Graph-based Question Answer for Large Multimodal Models

CVPR 2026

Understanding and reasoning over structured knowledge is a fundamental capability for intelligent systems. While Large Language Models (LLMs) have leveraged textual knowledge graphs for relational reasoning, linearizing graph structures into text often leads to token inefficiency and loss of higher-

Cited by 0SourcecodeScholar
2026

ViG-RAG: Video-aware Graph Retrieval-Augmented Generation via Temporal and Semantic Hybrid Reasoning

AAAI 2026technical

Retrieval-augmented generation (RAG) has greatly improved Large Language Models (LLMs) by adding external knowledge. However, current RAG-based methods face difficulties with long-context video understanding due to two main challenges. First, Current RAG-based methods for long-context video understa

Cited by 0SourcePDFScholar
2026

ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking

CVPR 2026

CoT has significantly enhanced the reasoning ability of LLMs while it faces challenges when extended to multimodal domains, particularly in mathematical tasks.Existing MLLMs typically perform textual reasoning solely from a single static mathematical image, overlooking dynamic visual acquisition dur

Cited by 0SourcecodeScholar
2026

VisualScore: Learning Holistic Visual Quality Scores via Multi-Task Reasoning

ICML 2026poster

Image quality assessment (IQA) is inherently multi-mage quality assessment (IQA) is inherently multi-dimensional, yet existing reward models are typically limited to a single task and become unstable when extended to multi-task settings. In particular, heterogeneous reward scales and variances acros…

Cited by 0SourceScholar
2026

WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research

ICLR 2026poster

This paper tackles \textbf{open-ended deep research (OEDR)}, a complex challenge where AI agents must synthesize vast web-scale information into insightful reports. Current approaches are plagued by dual-fold limitations: static research pipelines that decouple planning from evidence acquisition and…

Cited by 0SourcecodeScholar
2025

A Phase-Change-Material-Based Variable Stiffness Sheath Inspired by a Multi-Layer Wave Spring Structure for Flexible Upper Gastrointestinal Endoscopic Robots

RA-L 2025

Continuum robots in flexible gastrointestinal endoscopy require transitioning between flexible and rigid states. Phase-change-material-based variable stiffness (VS) methods exhibit a significant stiffness change ratio but are typically time-consuming. These materials are commonly fabricated as simpl

Cited by 2SourceScholar
2025

A Semantic Knowledge Complementarity based Decoupling Framework for Semi-supervised Class-imbalanced Medical Image Segmentation

CVPR 2025poster

The limited data annotations have made semi-supervised learning (SSL) increasingly popular in medical image analysis. However, the use of pseudo labels in SSL degrades the performance of decoders that heavily rely on high-accuracy annotations. This issue is particularly pronounced in class-imbalance…

2025

A Training-free LLM-based Approach to General Chinese Character Error Correction

ACL 2025long

Chinese spelling correction (CSC) is a crucial task that aims to correct character errors in Chinese text. While conventional CSC focuses on character substitution errors caused by mistyping, two other common types of character errors, missing and redundant characters, have received less attention.…

2025

An End-to-End Learning-Based Multi-Sensor Fusion for Autonomous Vehicle Localization

ICRA 2025

Multi-sensor fusion is essential for autonomous vehicle localization, as it is capable of integrating data from various sources for enhanced accuracy and reliability. The accuracy of the integrated location and orientation depends on the precision of the uncertainty modeling. Traditional methods of

Cited by 2SourceScholar
2025

Bench4Merge: A Comprehensive Benchmark for Merging in Realistic Dense Traffic with Micro-Interactive Vehicles

IROS 2025

While the capabilities of autonomous driving have advanced rapidly, merging into dense traffic remains a significant challenge, many motion planning methods for this scenario have been proposed but it is hard to evaluate them. Most existing closed-loop simulators rely on rule-based controls for othe

Cited by 0SourcecodeScholar
2025

Bi-directional Cable-driven Ankle Exoskeleton Coupled with Series Elastic Actuator for Compliant Gait Assisting*

IROS 2025

This paper presents a lightweight bidirectional cable-driven ankle exoskeleton system (total mass: 2.6 kg) based on a series elastic actuation architecture (actuator module mass: 1.05 kg). The system utilizes a waist-mounted drive unit and Bowden cables to deliver bidirectional assistance to the ank

Cited by 0SourceScholar
2025

Breaking the Compression Ceiling: Data-Free Pipeline for Ultra-Efficient Delta Compression

NeurIPS 2025poster

With the rise of the fine-tuned–pretrained paradigm, storing numerous fine-tuned models for multi-tasking creates significant storage overhead. Delta compression alleviates this by storing only the pretrained model and the highly compressed delta weights (the differences between fine-tuned and pretr…

Cited by 0SourcecodeScholar
2025

Chimera: Improving Generalist Model with Domain-Specific Experts

ICCV 2025poster

Large Multi-modal Models (LMMs), trained on web-scale datasets predominantly composed of natural images, have demonstrated remarkable performance on general tasks. However, these models often exhibit limited specialized capabilities for domain-specific tasks that require extensive domain prior knowl…

Cited by 0SourcePDFScholar
2025

DGVO: A Dynamically Constrained Gradient Velocity Obstacle Approach for Mobile Robots in Dynamic Environments

IROS 2025

In this paper, we propose a framework based on velocity obstacles to address dynamic obstacle avoidance problem for constrained mobile robots. The framework establishes a nonlinear mapping from the control domain to the velocity space based on the robot’s kinematic model and input constraints. This

Cited by 0SourceScholar
2025

DISC: Plug-and-Play Decoding Intervention with Similarity of Characters for Chinese Spelling Check

ACL 2025long

One key characteristic of the Chinese spelling check (CSC) task is that incorrect characters are usually similar to the correct ones in either phonetics or glyph. To accommodate this, previous works usually leverage confusion sets, which suffer from two problems, i.e., difficulty in determining whic…

2025

DRARL: Disengagement-Reason-Augmented Reinforcement Learning for Efficient Improvement of Autonomous Driving Policy

IROS 2025

With the increasing presence of automated vehicles on open roads under driver supervision, disengagement cases are becoming more prevalent. While some data-driven planning systems attempt to directly utilize these disengagement cases for policy improvement, the inherent scarcity of disengagement dat

Cited by 3SourceScholar
2025

DiffCalib: Reformulating Monocular Camera Calibration as Diffusion-Based Dense Incident Map Generation

AAAI 2025technical

Monocular camera calibration is a key precondition for numerous 3D vision applications. Despite considerable advancements, existing methods often hinge on specific assumptions and struggle to generalize across varied real-world scenarios, and the performance is limited by insufficient training data.…

Cited by 4SourcePDFScholar
2025

Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback

ACL 2025long

The scientific research paradigm is undergoing a profound transformation owing to the development of Artificial Intelligence (AI). Recent works demonstrate that various AI-assisted research methods can largely improve research efficiency by improving data analysis, accelerating computation, and fost…

2025

DriveX: Omni Scene Modeling for Learning Generalizable World Knowledge in Autonomous Driving

ICCV 2025poster

Data-driven learning has advanced autonomous driving, yet task-specific models struggle with out-of-distribution scenarios due to their narrow optimization objectives and reliance on costly annotated data. We present DriveX, a self-supervised world model that learns generalizable scene dynamics and…

Cited by 0SourcePDFScholar
2025

Dynamic Evil Score-Guided Decoding: An Efficient Decoding Framework For Red-Team Model

ACL 2025finding

Large language models (LLMs) have achieved significant advances but can potentially generate harmful content such as social biases, extremism, and misinformation. Red teaming is a promising approach to enhance model safety by creating adversarial prompts to test and improve model robustness. However…

Cited by 0SourcePDFScholar
2025

GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training

ICLR 2025poster

Despite their proficiency in general tasks, Multi-modal Large Language Models (MLLMs) struggle with automatic Geometry Problem Solving (GPS), which demands understanding diagrams, interpreting symbols, and performing complex reasoning. This limitation arises from their pre-training on natural images…

Cited by 8SourcePDFScholar
2025

Image Over Text: Transforming Formula Recognition Evaluation with Character Detection Matching

CVPR 2025poster

Formula recognition presents significant challenges due to the complicated structure and varied notation of mathematical expressions. Despite continuous advancements in formula recognition models, the evaluation metrics employed by these models, such as BLEU and Edit Distance, still exhibit notable…

2025

JiSAM: Alleviate Labeling Burden and Corner Case Problems in Autonomous Driving via Minimal Real-World Data

CVPR 2025poster

Deep-learning-based autonomous driving (AD) perception introduces a promising picture for safe and environment-friendly transportation. However, the over-reliance on real labeled data in LiDAR perception limits the scale of on-road attempts. 3D real world data is notoriously time-and-energy-consumin…

2025

LaTeXNet: A Specialized Model for Converting Visual Tables and Equations to LaTeX Code

ICASSP 2025accepted

LaTeX provides precise representation of complex elements (i.e., tables and equations) in scientific documents. However, the automated transcription of visual representations into LaTeX code is challenging and prone to errors. This paper introduces LaTeXNet, a specialized model designed to automate…

Cited by 0SourceScholar
2025

LiON: Learning Point-Wise Abstaining Penalty for LiDAR Outlier DetectioN Using Diverse Synthetic Data

AAAI 2025technical

LiDAR-based semantic scene understanding is an important module in the modern autonomous driving perception stack. However, identifying outlier points in a LiDAR point cloud is challenging as LiDAR point clouds lack semantically-rich information. While former SOTA methods adopt heuristic architectur…

2025

Lumina-Image 2.0: A Unified and Efficient Image Generative Framework

ICCV 2025poster

We introduce Lumina-Image 2.0, an advanced text-to-image (T2I) model that surpasses previous state-of-the-art methods across multiple benchmarks. Lumina-Image 2.0 is characterized by two key features: (1) Unification - it adopts a unified architecture (Unified Next-DiT) that treats text and image to…

2025

MLVU: Benchmarking Multi-task Long Video Understanding

CVPR 2025poster

The evaluation of Long Video Understanding (LVU) performance poses an important but challenging research problem. Despite previous efforts, the existing video understanding benchmarks are severely constrained by several issues, especially the insufficient lengths of videos, a lack of diversity in vi…

2025

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

ICML 2025poster

Answering questions with Chain-of-Thought (CoT) has significantly enhanced the reasoning capabilities of Large Language Models (LLMs), yet its impact on Large Multimodal Models (LMMs) still lacks a systematic assessment and in-depth investigation. In this paper, we introduce **MME-CoT**, a specializ…

Cited by 0SourcePDFScholar
2025

MovieDreamer: Hierarchical Generation for Coherent Long Visual Sequences

ICLR 2025poster

Recent advancements in video generation have primarily leveraged diffusion models for short-duration content. However, these approaches often fall short in modeling complex narratives and maintaining character consistency over extended periods, which is essential for long-form video production like…

Cited by 24SourcePDFScholar
2025

OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

ICLR 2025spotlight

Image-text interleaved data, consisting of multiple images and texts arranged in a natural document format, aligns with the presentation paradigm of internet data and closely resembles human reading habits. Recent studies have shown that such data aids multimodal in-context learning and maintains th…

2025

OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations

CVPR 2025poster

Document content extraction is a critical task in computer vision, underpinning the data needs of large language models (LLMs) and retrieval-augmented generation (RAG) systems. Despite recent progress, current document parsing methods have not been fairly and comprehensively evaluated due to the nar…

2025

PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training

ICLR 2025spotlight

This paper aims to address the challenge of hallucinations in Multimodal Large Language Models (MLLMs) particularly for dense image captioning tasks. To tackle the challenge, we identify the current lack of a metric that finely measures the caption quality in concept level. We hereby introduce HalF…

Cited by 0SourcePDFScholar
2025

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

ICCV 2025poster

Large Language Models have demonstrated remarkable reasoning capability in complex textual tasks. However, multimodal reasoning, which requires integrating visual and textual information, remains a significant challenge. Existing visual-language models often struggle to effectively analyze and reaso…

2025

Risk Euclidean Distance-based Model Predictive Path Integral to Safety-Critical Obstacle Avoidance

IROS 2025

Sampling-based Model Predictive Control (MPC) algorithms such as Model Predictive Path Integral (MPPI) excel in managing nonlinear constraints and complex systems. However, their conventional sampling strategies often result in suboptimal local solutions. To address this problem, we propose RESM-MPP

Cited by 0SourceScholar
2025

SURVEYFORGE : On the Outline Heuristics, Memory-Driven Generation, and Multi-dimensional Evaluation for Automated Survey Writing

ACL 2025long

Survey paper plays a crucial role in scientific research, especially given the rapid growth of research publications. Recently, researchers have begun using LLMs to automate survey generation for better efficiency. However, the quality gap between LLM-generated surveys and those written by human rem…

2025

SafeConf: A Confidence-Calibrated Safety Self-Evaluation Method for Large Language Models

EMNLP 2025

Large language models (LLMs) have achieved groundbreaking progress in Natural Language Processing (NLP). Despite the numerous advantages of LLMs, they also pose significant safety risks. Self-evaluation mechanisms have gained increasing attention as a key safeguard to ensure safe and controllable co

Cited by 0SourcePDFScholar
2025

Self-adaptive Dataset Construction for Real-World Multimodal Safety Scenarios

EMNLP 2025

Multimodal large language models (MLLMs) are rapidly evolving, presenting increasingly complex safety challenges. However, current dataset construction methods, which are risk-oriented, fail to cover the growing complexity of real-world multimodal safety scenarios (RMS). And due to the lack of a uni

Cited by 0SourcePDFScholar
2025

TELL ME: Tackle Electrocardiogram with Large Language Model Effectively

ICASSP 2025accepted

Electrocardiogram (ECG) signals are crucial indicators of various human physiological states. A thorough analysis of these signals is indispensable for applications such as disease prediction, mental stress assessment, and other medical diagnostics. Despite the rapid progress in large language model…

Cited by 0SourceScholar
2025

TREND: Unsupervised 3D Representation Learning via Temporal Forecasting for LiDAR Perception

NeurIPS 2025spotlight

Labeling LiDAR point clouds is notoriously time-and-energy-consuming, which spurs recent unsupervised 3D representation learning methods to alleviate the labeling burden in LiDAR perception via pretrained weights. Existing work focus on either masked auto encoding or contrastive learning on LiDAR po…

Cited by 0SourceScholar
2025

Temporal Overlapping Prediction: A Self-supervised Pre-training Method for LiDAR Moving Object Segmentation

ICCV 2025poster

Moving object segmentation (MOS) on LiDAR point clouds is crucial for autonomous systems such as self-driving vehicles. While previous supervised approaches rely on costly manual annotations, LiDAR sequences naturally capture temporal motion cues that can be leveraged for self-supervised learning. I…

2025

VDTF-ACT: ACT-based Multimodal Space Fine Manipulation Method with Visual Depth Tactile Fusion

IROS 2025

Autonomous fine manipulation in space for orbital assembly continues to present a critical challenge in the field of aerospace engineering. Under low-gravity conditions, during satellite manipulator operations on free-floating objects, the absence of significant gravitational forces and friction con

Cited by 0SourcecodeScholar
2025

What Is a Good Question? Assessing Question Quality via Meta-Fact Checking

AAAI 2025technical

Knowledge-based questions are typically employed to evaluate LLM's knowledge boundaries; meanwhile, numerous studies focus on question generation as a means to enhance the capabilities of both models and individuals. However, there is a lack of in-depth exploration about what constitutes a good ques…

2024

A Simple yet Effective Training-free Prompt-free Approach to Chinese Spelling Correction Based on Large Language Models

EMNLP 2024main

This work proposes a simple training-free prompt-free approach to leverage large language models (LLMs) for the Chinese spelling correction (CSC) task, which is totally different from all previous CSC approaches. The key idea is to use an LLM as a pure language model in a conventional manner. The LL…

2024

Better Regression Makes Better Test-time Adaptive 3D Object Detection

ECCV 2024poster

"Domain Adaptation (DA) has been widely explored and made significant progress on cross-domain 3D tasks recently. Despite being effective, existing works fail to deal with rapidly changing domains due to the unpredictable test time scenarios and meanwhile fast response time requirement. Thus, we exp…

2024

ComFusion: Enhancing Personalized Generation by Instance-Scene Compositing and Fusion

ECCV 2024poster

"Recent progress in personalizing text-to-image (T2I) diffusion models has demonstrated their capability to generate images based on personalized visual concepts using only a few user-provided examples. However, these models often struggle with maintaining high visual fidelity, particularly when mod…

Cited by 1SourcePDFScholar
2024

Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving

NeurIPS 2024poster

Autonomous driving has advanced significantly due to sensors, machine learning, and artificial intelligence improvements. However, prevailing methods struggle with intricate scenarios and causal relationships, hindering adaptability and interpretability in varied environments. To address the above p…

2024

DAGCN: Distance-based and Aspect-oriented Graph Convolutional Network for Aspect-based Sentiment Analysis

NAACL 2024findings

Aspect-based sentiment analysis (ABSA) is a task that aims to determine the sentiment polarity of aspects by identifying opinion words. Recent advancements have predominantly been rooted either in semantic or syntactic methods. However, both of them tend to interference from local factors such as ir…

2024

DiffMap: Enhancing Map Segmentation With Map Prior Using Diffusion Model

RA-L 2024

Constructing high-definition (HD) maps is a crucial requirement for enabling autonomous driving. In recent years, several map segmentation algorithms have been developed to address this need, leveraging advancements in Bird's-Eye View (BEV) perception. However, existing models still encounter challe

Cited by 17SourceScholar
2024

DomainGallery: Few-shot Domain-driven Image Generation by Attribute-centric Finetuning

NeurIPS 2024poster

The recent progress in text-to-image models pretrained on large-scale datasets has enabled us to generate various images as long as we provide a text prompt describing what we want. Nevertheless, the availability of these models is still limited when we expect to generate images that fall into a spe…

2024

DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion Prior

ICLR 2024poster

We present DreamCraft3D, a hierarchical 3D content generation method that produces high-fidelity and coherent 3D objects. We tackle the problem by leveraging a 2D reference image to guide the stages of geometry sculpting and texture boosting. A central focus of this work is to address the consistenc…

2024

Embodied Intelligence: Bionic Robot Controller Integrating Environment Perception, Autonomous Planning, and Motion Control

RA-L 2024

This letter proposes a bionic robot controller equipped with intelligent perception and autonomous planning modules to address the manufacturing industry's requirements for small-batch, customized, and autonomous task. Three crucial components: motion control module, vision perception module, and au

Cited by 18SourceScholar
2024

Exploring and Exploiting the Asymmetric Valley of Deep Neural Networks

NeurIPS 2024poster

Exploring the loss landscape offers insights into the inherent principles of deep neural networks (DNNs). Recent work suggests an additional asymmetry of the valley beyond the flat and sharp ones, yet without thoroughly examining its causes or implications. Our study methodically explores the factor…

Cited by 3SourcePDFScholar
2024

Language-Driven Anchors for Zero-Shot Adversarial Robustness

CVPR 2024poster

Deep Neural Networks (DNNs) are known to be susceptible to adversarial attacks. Previous researches mainly focus on improving adversarial robustness in the fully supervised setting leaving the challenging domain of zero-shot adversarial robustness an open question. In this work we investigate this d…

2024

LiDAR-PTQ: Post-Training Quantization for Point Cloud 3D Object Detection

ICLR 2024poster

Due to highly constrained computing power and memory, deploying 3D lidar-based detectors on edge devices equipped in autonomous vehicles and robots poses a crucial challenge. Being a convenient and straightforward model compression approach, Post-Training Quantization (PTQ) has been widely adopted i…

2024

LogFormer: A Pre-train and Tuning Pipeline for Log Anomaly Detection

AAAI 2024technical

Log anomaly detection is a key component in the field of artificial intelligence for IT operations (AIOps). Considering log data of variant domains, retraining the whole network for unknown domains is inefficient in real industrial scenarios. However, previous deep models merely focused on extractin…

2024

Multimodal Clickbait Detection by De-confounding Biases Using Causal Representation Inference

EMNLP 2024main

This paper focuses on detecting clickbait posts on the Web. These posts often use eye-catching disinformation in mixed modalities to mislead users to click for profit. That affects the user experience and thus would be blocked by content provider. To escape detection, malicious creators use tricks t…

Cited by 0SourcePDFScholar
2024

Norm Tweaking: High-Performance Low-Bit Quantization of Large Language Models

AAAI 2024technical

As the size of large language models (LLMs) continues to grow, model compression without sacrificing accuracy has become a crucial challenge for deployment. While some quantization methods, such as GPTQ, have made progress in achieving acceptable 4-bit weight-only quantization, attempts at lower-bi…

Cited by 34SourcePDFScholar
2024

Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

ACL 2024long

A pivotal advancement in the progress of large language models (LLMs) is the emergence of the Mixture-of-Experts (MoE) LLMs. Compared to traditional LLMs, MoE LLMs can achieve higher performance with fewer active parameters, but it is still hard to deploy them due to their immense parameter sizes. D…

2024

OWL: A Large Language Model for IT Operations

ICLR 2024poster

With the rapid advancement of IT operations, managing and analyzing large data volumes efficiently for practical applications has become increasingly critical. Natural Language Processing (NLP) techniques have demonstrated remarkable capabilities in various tasks, including named entity recognition,…

2024

On the Emergence of Cross-Task Linearity in Pretraining-Finetuning Paradigm

ICML 2024poster

The pretraining-finetuning paradigm has become the prevailing trend in modern deep learning. In this work, we discover an intriguing linear phenomenon in models that are initialized from a common pretrained checkpoint and finetuned on different tasks, termed as Cross-Task Linearity (CTL). Specifical…

Cited by 5SourcePDFScholar
2024

Once for Both: Single Stage of Importance and Sparsity Search for Vision Transformer Compression

CVPR 2024poster

Recent Vision Transformer Compression (VTC) works mainly follow a two-stage scheme where the importance score of each model unit is first evaluated or preset in each submodule followed by the sparsity score evaluation according to the target sparsity constraint. Such a separate evaluation process in…

2024

ReSimAD: Zero-Shot 3D Domain Transfer for Autonomous Driving with Source Reconstruction and Target Simulation

ICLR 2024poster

Domain shifts such as sensor type changes and geographical situation variations are prevalent in Autonomous Driving (AD), which poses a challenge since AD model relying on the previous domain knowledge can be hardly directly deployed to a new domain without additional costs. In this paper, we provid…

2024

Realistic Rainy Weather Simulation for LiDARs in CARLA Simulator

IROS 2024poster

Data augmentation methods to enhance perception performance in adverse weather have recently attracted considerable attention. Most of the LiDAR data augmentation methods post-process the existing dataset by physics-based models or machine-learning methods. However, due to the limited environmental…

Cited by 4SourcecodeScholar
2024

Shadow Generation for Composite Image Using Diffusion Model

CVPR 2024poster

In the realm of image composition generating realistic shadow for the inserted foreground remains a formidable challenge. Previous works have developed image-to-image translation models which are trained on paired training data. However they are struggling to generate shadows with accurate shapes an…

2024

Towards Better Utilization of Multi-Reference Training Data for Chinese Grammatical Error Correction

ACL 2024findings

For the grammatical error correction (GEC) task, there usually exist multiple correction ways for an erroneous input sentence, leading to multiple references. Observing the high proportion of multi-reference instances in Chinese GEC training data, we target a systematic study on how to better utiliz…

2024

Training-Free Adaptive Diffusion with Bounded Difference Approximation Strategy

NeurIPS 2024poster

Diffusion models have recently achieved great success in the synthesis of high-quality images and videos. However, the existing denoising techniques in diffusion models are commonly based on step-by-step noise predictions, which suffers from high computation cost, resulting in a prohibitive latency…

2024

ZOPP: A Framework of Zero-shot Offboard Panoptic Perception for Autonomous Driving

NeurIPS 2024poster

Offboard perception aims to automatically generate high-quality 3D labels for autonomous driving (AD) scenes. Existing offboard methods focus on 3D object detection with closed-set taxonomy and fail to match human-level recognition capability on the rapidly evolving perception tasks. Due to heavy re…

2024

mABC: Multi-Agent Blockchain-inspired Collaboration for Root Cause Analysis in Micro-Services Architecture

EMNLP 2024finding

Root cause analysis (RCA) in Micro-services architecture (MSA) with escalating complexity encounters complex challenges in maintaining system stability and efficiency due to fault propagation and circular dependencies among nodes. Diverse root cause analysis faults require multi-agents with diverse…

2024

mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

EMNLP 2024finding

Structure information is critical for understanding the semantics of text-rich images, such as documents, tables, and charts. Existing Multimodal Large Language Models (MLLMs) for Visual Document Understanding are equipped with text recognition ability but lack general structure understanding abilit…

2023

AD-PT: Autonomous Driving Pre-Training with Large-scale Point Cloud Dataset

NeurIPS 2023poster

It is a long-term vision for Autonomous Driving (AD) community that the perception models can learn from a large-scale point cloud dataset, to obtain unified representations that can achieve promising results on different tasks or benchmarks. Previous works mainly focus on the self-supervised pre-tr…

2023

Bi3D: Bi-Domain Active Learning for Cross-Domain 3D Object Detection

CVPR 2023poster

Unsupervised Domain Adaptation (UDA) technique has been explored in 3D cross-domain tasks recently. Though preliminary progress has been made, the performance gap between the UDA-based 3D model and the supervised one trained with fully annotated target domain is still large. This motivates us to con…

2023

Conditional Positional Encodings for Vision Transformers

ICLR 2023poster

We propose a conditional positional encoding (CPE) scheme for vision Transformers. Unlike previous fixed or learnable positional encodings that are predefined and independent of input tokens, CPE is dynamically generated and conditioned on the local neighborhood of the input tokens. As a result, CPE…

2023

Delving Into Shape-Aware Zero-Shot Semantic Segmentation

CVPR 2023poster

Thanks to the impressive progress of large-scale vision-language pretraining, recent recognition models can classify arbitrary objects in a zero-shot and open-set manner, with a surprisingly high accuracy. However, translating this success to semantic segmentation is not trivial, because this dense…

2023

Generative Diffusion Prior for Unified Image Restoration and Enhancement

CVPR 2023poster

Existing image restoration methods mostly leverage the posterior distribution of natural images. However, they often assume known degradation and also require supervised training, which restricts their adaptation to complex real applications. In this work, we propose the Generative Diffusion Prior (…

Cited by 240SourcePDFScholar
2023

Improving Seq2Seq Grammatical Error Correction via Decoding Interventions

EMNLP 2023long findings

The sequence-to-sequence (Seq2Seq) approach has recently been widely used in grammatical error correction (GEC) and shows promising performance. However, the Seq2Seq GEC approach still suffers from two issues. First, a Seq2Seq GEC model can only be trained on parallel data, which, in GEC task, is of…

Cited by 0SourcecodeScholar
2023

LipsNet: A Smooth and Robust Neural Network with Adaptive Lipschitz Constant for High Accuracy Optimal Control

ICML 2023poster

Deep reinforcement learning (RL) is a powerful approach for solving optimal control problems. However, RL-trained policies often suffer from the action fluctuation problem, where the consecutive actions significantly differ despite only slight state variations. This problem results in mechanical com…

Cited by 17SourcePDFScholar
2023

Make-It-3D: High-fidelity 3D Creation from A Single Image with Diffusion Prior

ICCV 2023poster

In this work, we investigate the problem of creating high-fidelity 3D content from only a single image. This is inherently challenging: it essentially involves estimating the underlying 3D geometry while hallucinating unseen textures. To address this challenge, we leverage prior knowledge in a well-…

Cited by 301PDFcodeScholar
2023

MetaPortrait: Identity-Preserving Talking Head Generation With Fast Personalized Adaptation

CVPR 2023poster

In this work, we propose an ID-preserving talking head generation framework, which advances previous methods in two aspects. First, as opposed to interpolating from sparse flow, we claim that dense landmarks are crucial to achieving accurate geometry-aware flow fields. Second, inspired by face-swapp…

2023

NaSGEC: a Multi-Domain Chinese Grammatical Error Correction Dataset from Native Speaker Texts

ACL 2023findings

We introduce NaSGEC, a new dataset to facilitate research on Chinese grammatical error correction (CGEC) for native speaker texts from multiple domains. Previous CGEC research primarily focuses on correcting texts from a single domain, especially learner essays. To broaden the target domain, we anno…

2023

OPT-GAN: A Broad-Spectrum Global Optimizer for Black-Box Problems by Learning Distribution

AAAI 2023technical

Black-box optimization (BBO) algorithms are concerned with finding the best solutions for problems with missing analytical details. Most classical methods for such problems are based on strong and fixed a priori assumptions, such as Gaussianity. However, the complex real-world problems, especially w…

2023

Paint by Example: Exemplar-Based Image Editing With Diffusion Models

CVPR 2023poster

Language-guided image editing has achieved great success recently. In this paper, we investigate exemplar-guided image editing for more precise control. We achieve this goal by leveraging self-supervised training to disentangle and re-organize the source image and the exemplar. However, the naive ap…

2023

Physically Realizable Natural-Looking Clothing Textures Evade Person Detectors via 3D Modeling

CVPR 2023poster

Recent works have proposed to craft adversarial clothes for evading person detectors, while they are either only effective at limited viewing angles or very conspicuous to humans. We aim to craft adversarial texture for clothes based on 3D modeling, an idea that has been used to craft rigid adversar…

2023

RODIN: A Generative Model for Sculpting 3D Digital Avatars Using Diffusion

CVPR 2023highlight

This paper presents a 3D diffusion model that automatically generates 3D digital avatars represented as neural radiance fields (NeRFs). A significant challenge for 3D diffusion is that the memory and processing costs are prohibitive for producing high-quality results with rich details. To tackle thi…

Cited by 369SourcePDFScholar
2023

ROME: Robustifying Memory-Efficient NAS via Topology Disentanglement and Gradient Accumulation

ICCV 2023poster

Albeit being a prevalent architecture searching approach, differentiable architecture search (DARTS) is largely hindered by its substantial memory cost since the entire supernet resides in the memory. This is where the single-path DARTS comes in, which only chooses a single-path submodel at each ste…

Cited by 12PDFScholar
2023

UMC: A Unified Bandwidth-efficient and Multi-resolution based Collaborative Perception Framework

ICCV 2023poster

Multi-agent collaborative perception (MCP) has recently attracted much attention. It includes three key processes: communication for sharing, collaboration for integration, and reconstruction for different downstream tasks. Existing methods pursue designing the collaboration process alone, ignoring…

Cited by 40PDFcodeScholar
2023

Uni3D: A Unified Baseline for Multi-Dataset 3D Object Detection

CVPR 2023poster

Current 3D object detection models follow a single dataset-specific training and testing paradigm, which often faces a serious detection accuracy drop when they are directly deployed in another dataset. In this paper, we study the task of training a unified 3D detector from multiple datasets. We obs…

2022

A Joint Acceleration Estimation Method Based on a High-Order Disturbance Observer

RA-L 2022

Joint acceleration feedback is widely used in the design of controllers and observers since joint accelerations reflect the joint dynamics of robots, especially in physical human-robot interaction. However, joint acceleration acquisition is a technical difficulty for robots. The dynamics-based metho

Cited by 9SourceScholar
2022

Adversarial Texture for Fooling Person Detectors in the Physical World

CVPR 2022oral

Nowadays, cameras equipped with AI systems can capture and analyze images to detect people automatically. However, the AI system can make mistakes when receiving deliberately designed patterns in the real world, i.e., physical adversarial examples. Prior works have shown that it is possible to print…

Cited by 147PDFcodeScholar
2022

Analytic-DPM: an Analytic Estimate of the Optimal Reverse Variance in Diffusion Probabilistic Models

ICLR 2022oral

Diffusion probabilistic models (DPMs) represent a class of powerful generative models. Despite their success, the inference of DPMs is expensive since it generally needs to iterate over thousands of timesteps. A key problem in the inference is to estimate the variance in each timestep of the reverse…

2022

CamMap: Extrinsic Calibration of Non-Overlapping Cameras Based on SLAM Map Alignment

RA-L 2022

Multiple cameras have emerged as a promising technology for robots and vehicles due to their broad fields of view (FoV) and high resolution. However, there are often limited or no overlapping FoVs among cameras, bringing challenges to estimating extrinsic camera parameters. To overcome this problem,

Cited by 12SourceScholar
2022

Enhanced Accuracy and Robustness via Multi-Teacher Adversarial Distillation

ECCV 2022poster

"Adversarial training is an effective approach for improving the robustness of deep neural networks against adversarial attacks. Although bringing reliable robustness, adversarial training (AT) will reduce the performance of identifying clean examples. Meanwhile, Adversarial training can bring more…

2022

Estimating the Optimal Covariance with Imperfect Mean in Diffusion Probabilistic Models

ICML 2022spotlight

Diffusion probabilistic models (DPMs) are a class of powerful deep generative models (DGMs). Despite their success, the iterative generation process over the full timesteps is much less efficient than other DGMs such as GANs. Thus, the generation performance on a subset of timesteps is crucial, whic…

2022

Fast Lossless Neural Compression with Integer-Only Discrete Flows

ICML 2022spotlight

By applying entropy codecs with learned data distributions, neural compressors have significantly outperformed traditional codecs in terms of compression ratio. However, the high inference latency of neural networks hinders the deployment of neural compressors in practical applications. In this work…

2022

Human-Centric Image Cropping with Partition-Aware and Content-Preserving Features

ECCV 2022poster

"Image cropping aims to find visually appealing crops in an image, which is an important yet challenging task. In this paper, we consider a specific and practical application: human-centric image cropping, which focuses on the depiction of a person. To this end, we propose a human-centric image crop…

2022

MuCGEC: a Multi-Reference Multi-Source Evaluation Dataset for Chinese Grammatical Error Correction

NAACL 2022long

This paper presents MuCGEC, a multi-reference multi-source evaluation dataset for Chinese Grammatical Error Correction (CGEC), consisting of 7,063 sentences collected from three Chinese-as-a-Second-Language (CSL) learner sources. Each sentence is corrected by three annotators, and their corrections…

2022

Real-Time Neural Character Rendering with Pose-Guided Multiplane Images

ECCV 2022poster

"We propose pose-guided multiplane image (MPI) synthesis which can render an animatable character in real scenes with photorealistic quality. We use a portable camera rig to capture the multi-view images along with the driving signal for the moving subject. Our method generalizes the image-to-image…

2022

StyleSwin: Transformer-Based GAN for High-Resolution Image Generation

CVPR 2022poster

Despite the tantalizing success in a broad of vision tasks, transformers have not yet demonstrated on-par ability as ConvNets in high-resolution image generative modeling. In this paper, we seek to explore using pure transformers to build a generative adversarial network for high-resolution image sy…

Cited by 318PDFcodeScholar
2022

SynGEC: Syntax-Enhanced Grammatical Error Correction with a Tailored GEC-Oriented Parser

EMNLP 2022main

This work proposes a syntax-enhanced grammatical error correction (GEC) approach named SynGEC that effectively incorporates dependency syntactic information into the encoder part of GEC models. The key challenge for this idea is that off-the-shelf parsers are unreliable when processing ungrammatical…

2022

Vector Quantized Diffusion Model for Text-to-Image Synthesis

CVPR 2022oral

We present the vector quantized diffusion (VQ-Diffusion) model for text-to-image generation. This method is based on a vector quantized variational autoencoder (VQ-VAE) whose latent space is modeled by a conditional variant of the recently developed Denoising Diffusion Probabilistic Model (DDPM). We…

Cited by 959PDFcodeScholar
2021

A Unified Span-Based Approach for Opinion Mining with Syntactic Constituents

NAACL 2021long

Fine-grained opinion mining (OM) has achieved increasing attraction in the natural language processing (NLP) community, which aims to find the opinion structures of “Who expressed what opinions towards what” in one sentence. In this work, motivated by its span-based representations of opinion expres…

2021

AutoKWS: Keyword Spotting with Differentiable Architecture Search

ICASSP 2021accepted

Smart audio devices are gated by an always-on lightweight keyword spotting program to reduce power consumption. It is however challenging to design models that have both high accuracy and low latency for accurate and fast responsiveness. Many efforts have been made to develop end-to-end neural netwo…

Cited by 0SourceScholar
2021

CoCosNet v2: Full-Resolution Correspondence Learning for Image Translation

CVPR 2021poster

We present the full-resolution correspondence learning for cross-domain images, which aids image translation. We adopt a hierarchical strategy that uses the correspondence from coarse level to guide the fine levels. At each hierarchy, the correspondence can be efficiently computed via PatchMatch tha…

Cited by 371PDFcodeScholar
2021

DARTS-: Robustly Stepping out of Performance Collapse Without Indicators

ICLR 2021poster

Despite the fast development of differentiable architecture search (DARTS), it suffers from a standing instability issue regarding searching performance, which extremely limits its application. Existing robustifying methods draw clues from the outcome instead of finding out the causing factor. Vario…

2021

FairNAS: Rethinking Evaluation Fairness of Weight Sharing Neural Architecture Search

ICCV 2021poster

One of the most critical problems in weight-sharing neural architecture search is the evaluation of candidate models within a predefined search space. In practice, a one-shot supernet is trained to serve as an evaluator. A faithful ranking certainly leads to more accurate searching results. However,…

Cited by 432PDFcodeScholar
2021

Lywal: a Leg-Wheel Transformable Quadruped Robot with Picking up and Transport Functions

ICRA 2021poster

This paper introduces a leg-wheel transformable quadruped robot named Lywal which can switch to the leg-mode and the wheel-mode for locomotion, and the claw-mode for picking up and transport functions. First, the mechanical structure of Lywal is designed by using an innovative 2-DoF transformable me…

Cited by 11SourceScholar
2021

MagDR: Mask-Guided Detection and Reconstruction for Defending Deepfakes

CVPR 2021poster

Deepfakes raised serious concerns on the authenticity of visual contents. Prior works revealed the possibility to disrupt deepfakes by adding adversarial perturbations to the source data, but we argue that the threat has not been eliminated yet. This paper presents MagDR, a mask-guided detection and…

Cited by 44PDFScholar
2021

Matching Distributions between Model and Data: Cross-domain Knowledge Distillation for Unsupervised Domain Adaptation

ACL 2021long

Unsupervised Domain Adaptation (UDA) aims to transfer the knowledge of source domain to the unlabeled target domain. Existing methods typically require to learn to adapt the target model by exploiting the source data and sharing the network architecture across domains. However, this pipeline makes t…

Cited by 22SourcePDFScholar
2021

Prototypical Pseudo Label Denoising and Target Structure Learning for Domain Adaptive Semantic Segmentation

CVPR 2021poster

Self-training is a competitive approach in domain adaptive segmentation, which trains the network with the pseudo labels on the target domain. However inevitably, the pseudo labels are noisy and the target features are dispersed due to the discrepancy between source and target domains. In this paper…

Cited by 631PDFcodeScholar
2021

Stability and Generalization of Bilevel Programming in Hyperparameter Optimization

NeurIPS 2021poster

The (gradient-based) bilevel programming framework is widely used in hyperparameter optimization and has achieved excellent performance empirically. Previous theoretical work mainly focuses on its optimization properties, while leaving the analysis on generalization largely open. This paper attempts…

2021

Style-Based Point Generator With Adversarial Rendering for Point Cloud Completion

CVPR 2021poster

In this paper, we proposed a novel Style-based Point Generator with Adversarial Rendering (SpareNet) for point cloud completion. Firstly, we present the channel-attentive EdgeConv to fully exploit the local structures as well as the global shape in point features. Secondly, we observe that the conca…

Cited by 106PDFcodeScholar
2021

Twins: Revisiting the Design of Spatial Attention in Vision Transformers

NeurIPS 2021poster

Very recently, a variety of vision transformer architectures for dense prediction tasks have been proposed and they show that the design of spatial attention is critical to their success in these tasks. In this work, we revisit the design of the spatial attention and demonstrate that a carefully dev…

2021

Variational (Gradient) Estimate of the Score Function in Energy-based Latent Variable Models

ICML 2021spotlight

This paper presents new estimates of the score function and its gradient with respect to the model parameters in a general energy-based latent variable model (EBLVM). The score function and its gradient can be expressed as combinations of expectation and covariance terms over the (generally intracta…

2020

A Wasserstein Minimum Velocity Approach to Learning Unnormalized Models

AISTATS 2020poster

Score matching provides an effective approach to learning flexible unnormalized models, but its scalability is limited by the need to evaluate a second-order derivative. In this paper, we present a scalable approximation to a general family of learning objectives including score matching, by observi…

2020

An Attention-based Model for Conversion Rate Prediction with Delayed Feedback via Post-click Calibration

IJCAI 2020poster

Conversion rate (CVR) prediction is becoming increasingly important in the multi-billion dollar online display advertising industry. It has two major challenges: firstly, the scarce user history data is very complicated and non-linear; secondly, the time delay between the clicks and the correspondin…

Cited by 0SourcePDFScholar
2020

Bi-level Score Matching for Learning Energy-based Latent Variable Models

NeurIPS 2020poster

Score matching (SM) provides a compelling approach to learn energy-based models (EBMs) by avoiding the calculation of partition function. However, it remains largely open to learn energy-based latent variable models (EBLVMs), except some special cases. This paper presents a bi-level score matching (…

2020

Cross-Domain Correspondence Learning for Exemplar-Based Image Translation

CVPR 2020oral

We present a general framework for exemplar-based image translation, which synthesizes a photo-realistic image from the input in a distinct domain (e.g., semantic segmentation mask, or edge map, or pose keypoints), given an exemplar image. The output has the style (e.g., color, texture) in consisten…

Cited by 515PDFScholar
2020

Fair DARTS: Eliminating Unfair Advantages in Differentiable Architecture Search

ECCV 2020poster

Differentiable Architecture Search (DARTS) is now a widely disseminated weight-sharing neural architecture search method. However, it suffers from well-known performance collapse due to an inevitable aggregation of skip connections. In this paper, we first disclose that its root cause lies in an unf…

2020

Internal and Contextual Attention Network for Cold-start Multi-channel Matching in Recommendation

IJCAI 2020poster

Real-world integrated personalized recommendation systems usually deal with millions of heterogeneous items. It is extremely challenging to conduct full corpus retrieval with complicated models due to the tremendous computation costs. Hence, most large-scale recommendation systems consist of two mod…

2020

Training Interpretable Convolutional Neural Networks by Differentiating Class-specific Filters

ECCV 2020poster

Convolutional neural networks (CNNs) have been successfully used in a range of tasks. However, CNNs are often viewed as ""black-box"" and lack of interpretability. One main reason is due to the filter-class entanglement -- an intricate many-to-many correspondence between filters and classes. Most ex…

2020

Understanding and Stabilizing GANs’ Training Dynamics Using Control Theory

ICML 2020poster

Generative adversarial networks (GANs) are effective in generating realistic images but the training is often unstable. There are existing efforts that model the training dynamics of GANs in the parameter space but the analysis cannot directly motivate practically effective stabilizing methods. To t…

Cited by 36SourcePDFScholar
2019

Deep Exemplar-Based Video Colorization

CVPR 2019poster

This paper presents the first end-to-end network for exemplar-based video colorization. The main challenge is to achieve temporal consistency while remaining faithful to the reference style. To address this issue, we introduce a recurrent framework that unifies the semantic correspondence and color…

Cited by 261PDFcodeScholar
2019

Function Space Particle Optimization for Bayesian Neural Networks

ICLR 2019poster

While Bayesian neural networks (BNNs) have drawn increasing attention, their posterior inference remains challenging, due to the high-dimensional and over-parameterized nature. To address this issue, several highly flexible and scalable variational inference procedures based on the idea of particle…

2019

Multi-objects Generation with Amortized Structural Regularization

NeurIPS 2019poster

Deep generative models (DGMs) have shown promise in image generation. However, most of the existing methods learn a model by simply optimizing a divergence between the marginal distributions of the model and the data, and often fail to capture rich structures, such as attributes of objects and their…

2018

DeepExposure: Learning to Expose Photos with Asynchronously Reinforced Adversarial Learning

NeurIPS 2018poster

The accurate exposure is the key of capturing high-quality photos in computational photography, especially for mobile phones that are limited by sizes of camera modules. Inspired by luminosity masks usually applied by professional photographers, in this paper, we develop a novel algorithm for learni…

Cited by 116SourcePDFScholar
2018

Interpret Neural Networks by Identifying Critical Data Routing Paths

CVPR 2018poster

Interpretability of a deep neural network aims to explain the rationale behind its decisions and enable the users to understand the intelligent agents, which has become an important issue due to its importance in practical applications. To address this issue, we develop a Distillation Guided Routing…

Cited by 122SourcePDFScholar
2018

Message Passing Stein Variational Gradient Descent

ICML 2018oral

Stein variational gradient descent (SVGD) is a recently proposed particle-based Bayesian inference method, which has attracted a lot of interest due to its remarkable approximation ability and particle efficiency compared to traditional variational inference and Markov Chain Monte Carlo methods. How…

Cited by 104SourcePDFScholar
2018

Semi-crowdsourced Clustering with Deep Generative Models

NeurIPS 2018poster

We consider the semi-supervised clustering problem where crowdsourcing provides noisy information about the pairwise comparisons on a small subset of data, i.e., whether a sample pair is in the same cluster. We propose a new approach that includes a deep generative model (DGM) to characterize low-le…

2018

Smooth Neighbors on Teacher Graphs for Semi-Supervised Learning

CVPR 2018poster

The recently proposed self-ensembling methods have achieved promising results in deep semi-supervised learning, which penalize inconsistent predictions of unlabeled data under different perturbations. However, they only consider adding perturbations to each single data point, while ignoring the conn…

2018

Textbook Question Answering Under Instructor Guidance With Memory Networks

CVPR 2018poster

Textbook Question Answering (TQA) is a task to choose the most proper answers by reading a multi-modal context of abundant essays and images. TQA serves as a favorable test bed for visual and textual reasoning. However, most of the current methods are incapable of reasoning over the long contexts an…

2017

Gradient magnitude similarity deviation on multiple scales for color image quality assessment

ICASSP 2017accepted

Recently, various image quality assessment (IQA) metrics based on gradient similarity have been developed. In this paper, we extend the work of gradient magnitude similarity deviation (GMSD) and propose a more efficient metric. First, a novel similarity index is proposed, which gives the flexibility…

Cited by 0SourceScholar
2015

Convolutional Neural Networks with Intra-Layer Recurrent Connections for Scene Labeling

NeurIPS 2015poster

Scene labeling is a challenging computer vision task. It requires the use of both local discriminative features and global context information. We adopt a deep recurrent convolutional neural network (RCNN) for this task, which is originally proposed for object recognition. Different from traditional…

Cited by 80SourcePDFScholar
2015

Robust widely linear beamformer based on a projection constraint

ICASSP 2015accepted

For noncircular signals, optimal widely linear (WL) minimum variance distortionless response (MVDR) beamformer has a powerful performance by exploiting the noncircularity of the received signals. Though, the noncircularity rate can be estimated by the steering vector (SV) of the signal of interest (…

Cited by 0SourceScholar