← Search

Weidong Zhang

22 accepted papers

2026

Expert Divergence Learning for MoE-based Language Models

ICLR 2026poster

The Mixture-of-Experts (MoE) architecture is a powerful technique for scaling language models, yet it often suffers from expert homogenization, where experts learn redundant functionalities, thereby limiting MoE's full potential. To address this, we introduce Expert Divergence Learning, a novel pre-…

Cited by 0SourceScholar
2026

RHCNet: Residual-Guided Hierarchical Calibration Network for Robust Underwater Object Detection

CVPR 2026

Underwater images commonly suffer from foreground-background ambiguity, loss of structural details, and severely reduced contrast, which collectively make underwater object detection (UOD) an inherently challenging task. To handle this issue, we present a residual-guided hierarchical calibration net

Cited by 0SourcecodeScholar
2026

S2C2Seg: Semantic-Spatial Consistency and Category Optimization for Open-Vocabulary Segmentation

CVPR 2026

Open-vocabulary semantic segmentation extends pixel-level recognition to arbitrary text-described categories. Despite strong global semantic understanding, vision-language models such as CLIP exhibit limited spatial precision and semantic ambiguity across large vocabularies, constraining their effec

Cited by 0SourceScholar
2025

Adaptive Noise Rejection Strategy for Cooperative Motion Control of Dual-Arm Robots

RA-L 2025

Dual-arm robots possess exceptional collaborative capabilities and versatility, demonstrating broad application prospects across various fields. As a significant research area for dual-arm robots, the requirements for coordinated motion control are gradually increasing. In practical applications, ro

Cited by 3SourceScholar
2025

CVLN-Think: Causal Inference with Counterfactual Style Adaptation for Continuous Vision-and-Language Navigation

IROS 2025

Vision-and-Language Navigation in Continuous Environments (VLN-CE) presents challenges due to environmental variations and domain shifts, making it difficult for agents to generalize beyond seen environments. Most existing methods rely on learning correlations between observations and actions from t

Cited by 0SourceScholar
2025

Investigating Numerical Translation with Large Language Models

ICASSP 2025accepted

The inaccurate translation of numbers can lead to significant security issues, ranging from financial setbacks to medical inaccuracies. While large language models (LLMs) have made significant advancements in machine translation, their capacity for translating numbers has not been thoroughly explore…

Cited by 0SourceScholar
2025

PhysGCN-DL: Physics-Informed Graph Convolutional Networks with Diversity-Aware Loss Optimization for Multimodal Pedestrian Trajectory Prediction

IROS 2025

Pedestrian trajectory prediction ensures safe navigation in autonomous driving and intelligent robots. Existing methods have shown promising results but still face challenges in handling dynamic environments, social interactions, and high-dimensional data. In this paper, we propose a novel PhysGCN-D

Cited by 0SourceScholar
2024

Capturing Closely Interacted Two-Person Motions with Reaction Priors

CVPR 2024poster

In this paper we focus on capturing closely interacted two-person motions from monocular videos an important yet understudied topic. Unlike less-interacted motions closely interacted motions contain frequently occurring inter-human occlusions which pose significant challenges to existing capturing a…

Cited by 1SourcePDFScholar
2024

LatEval: An Interactive LLMs Evaluation Benchmark with Incomplete Information from Lateral Thinking Puzzles

COLING 2024main

With the evolution of LLMs, they are endowed with impressive logical reasoning, or vertical thinking capabilities. But can they think out of the box? Do they possess proficient lateral thinking abilities? Following the setup of Lateral Thinking Puzzles, we propose a novel evaluation benchmark, LatEv…

2024

Online-Learning-Based Distributionally Robust Motion Control with Collision Avoidance for Mobile Robots

ICRA 2024poster

Collision-free navigation is a critical issue in robotic systems as the environment is often dynamic and uncertain. This paper investigates a data-stream-driven motion control problem for mobile robots to avoid randomly moving obstacles when the probability distribution of the obstacle’s movement is…

Cited by 0SourceScholar
2023

CLIPVG: Text-Guided Image Manipulation Using Differentiable Vector Graphics

AAAI 2023technical

Considerable progress has recently been made in leveraging CLIP (Contrastive Language-Image Pre-Training) models for text-guided image manipulation. However, all existing works rely on additional generative models to ensure the quality of results, because CLIP alone cannot provide enough guidance in…

2023

Learning Analytical Posterior Probability for Human Mesh Recovery

CVPR 2023poster

Despite various probabilistic methods for modeling the uncertainty and ambiguity in human mesh recovery, their overall precision is limited because existing formulations for joint rotations are either not constrained to SO(3) or difficult to learn for neural networks. To address such an issue, we de…

2023

Robust Target Interception Strategy for a USV With Experimental Validation

RA-L 2023

This letter addresses the problem of designing an interception strategy for an underactuated uncrewed surface vessel (USV) in the presence of uncertain external disturbances and unknown internal parameters i.e., linear and nonlinear damping coefficients, vehicle mass, etc. The interception strategy

Cited by 11SourceScholar
2023

Underwater Ranker: Learn Which Is Better and How to Be Better

AAAI 2023technical

In this paper, we present a ranking-based underwater image quality assessment (UIQA) method, abbreviated as URanker. The URanker is built on the efficient conv-attentional image Transformer. In terms of underwater images, we specially devise (1) the histogram prior that embeds the color distribution…

2022

Face2Faceρ: Real-Time High-Resolution One-Shot Face Reenactment

ECCV 2022poster

"Existing one-shot face reenactment methods either present obvious artifacts in large pose transformations, or cannot well-preserve the identity information in the source images, or fail to meet the requirements of real-time applications due to the intensive amount of computation involved. In this p…

Cited by 37SourcePDFScholar
2022

Flexible Collision-free Platooning Method for Unmanned Surface Vehicle with Experimental Validations

IROS 2022poster

This paper addresses the flexible formation problem for unmanned surface vehicles in the presence of obstacles. Building upon the leader-follower formation scheme, a hybrid line-of-sight based flexible platooning method is proposed for follower vehicle to keep tracking the leader ship. A fusion arti…

Cited by 3SourceScholar
2021

SARG: A Novel Semi Autoregressive Generator for Multi-turn Incomplete Utterance Restoration

AAAI 2021technical

Dialogue systems in open domain have achieved great success due to the easily obtained single-turn corpus and the development of deep learning, but the multi-turn scenario is still a challenge because of the frequent coreference and information omission. In this paper, we investigate the incomplete…

2021

SPatchGAN: A Statistical Feature Based Discriminator for Unsupervised Image-to-Image Translation

ICCV 2021poster

For unsupervised image-to-image translation, we propose a discriminator architecture which focuses on the statistical features instead of individual patches. The network is stabilized by distribution matching of key statistical features at multiple scales. Unlike the existing methods which impose mo…

Cited by 35PDFcodeScholar
2020

GeoLayout: Geometry Driven Room Layout Estimation Based on Depth Maps of Planes

ECCV 2020poster

The task of room layout estimation is to locate the wall-floor, wall-ceiling, and wall-wall boundaries. Most recent methods solve this problem based on edge/keypoint detection or semantic segmentation. However, these approaches have shown limited attention on the geometry of the dominant planes and…

Cited by 34SourcePDFScholar