← Search

Jenq-Neng Hwang

40 accepted papers

2026

CLEP: Contrastive Language-Pose Pretraining

CVPR 2026

Aligning natural language descriptions with precise 3D human poses remains a big challenge due to the scarcity of effective pose representation mechanisms and large-scale, semantically rich datasets. To overcome these limitations, we first introduce **CLEP-2M**, the largest 3D pose-language dataset

Cited by 0SourceScholar
2026

Intrinsic Entropy of Context Length Scaling in LLMs

ICLR 2026oral

There has been work discussing the impact of long context on Language Model performance: some find that long irrelevant context could harm performance, while some experimentally summarize loss reduction by relevant long context as Scaling Laws. This calls for a more thorough understanding on how lon…

Cited by 0SourcecodeScholar
2026

Learning to Learn Weight Generation via Local Consistency Diffusion

CVPR 2026

Diffusion-based algorithms have emerged as promising techniques for weight generation. However, existing solutions are limited by two challenges: generalizability and missing local supervision targets. The first challenge stems from the inherent lack of cross-task transferability in existing single-

Cited by 0SourceScholar
2025

AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

ICLR 2025poster

Video detailed captioning is a key task which aims to generate comprehensive and coherent textual descriptions of video content, benefiting both video understanding and generation. In this paper, we propose AuroraCap, a video captioner based on a large multimodal model. We follow the simplest archit…

Cited by 6SourcePDFScholar
2025

Bayesian Optimization for Controlled Image Editing via LLMs

ACL 2025finding

In the rapidly evolving field of image generation, achieving precise control over generated content and maintaining semantic consistency remain significant limitations, particularly concerning grounding techniques and the necessity for model fine-tuning. To address these challenges, we propose Bayes…

Cited by 0SourcePDFScholar
2025

Exploring Probabilistic Modeling Beyond Domain Generalization for Semantic Segmentation

ICCV 2025poster

Domain Generalized Semantic Segmentation (DGSS) is a critical yet challenging task, as domain shifts in unseen environments can severely compromise model performance. While recent studies enhance feature alignment by projecting features into the source domain, they often neglect intrinsic latent dom…

Cited by 0SourcePDFScholar
2025

Is Meta-Learning Out? Rethinking Unsupervised Few-Shot Classification with Limited Entropy

ICCV 2025poster

Meta-learning is a powerful paradigm for tackling few-shot tasks. However, recent studies indicate that models trained with the whole-class training strategy can achieve comparable performance to those trained with meta-learning in few-shot classification tasks. To demonstrate the value of meta-lear…

Cited by 0SourcePDFScholar
2025

MambaMOT: State-Space Model as Motion Predictor for Multi-Object Tracking

ICASSP 2025accepted

In the field of multi-object tracking (MOT), traditional methods often rely on the Kalman filter for motion prediction, leveraging its strengths in linear motion scenarios. However, the inherent limitations of these methods become evident when confronted with complex, nonlinear motions and occlusion…

Cited by 0SourceScholar
2025

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression

NeurIPS 2025poster

Visual Autoregressive (VAR) modeling has garnered significant attention for its innovative next-scale prediction approach, which yields substantial improvements in efficiency, scalability, and zero-shot generalization. Nevertheless, the coarse-to-fine methodology inherent in VAR results in exponenti…

Cited by 0SourcecodeScholar
2025

MonoTAKD: Teaching Assistant Knowledge Distillation for Monocular 3D Object Detection

CVPR 2025poster

Monocular 3D object detection (Mono3D) holds noteworthy promise for autonomous driving applications owing to the cost-effectiveness and rich visual context of monocular camera sensors. However, depth ambiguity poses a significant challenge, as it requires extracting precise 3D scene geometry from a…

2025

PAMN: Multi-phase Correlation Modeling for Contrast-Enhanced 3D Medical Image Retrieval

EMNLP 2025

Contrast-enhanced 3D Medical imaging (e.g., CT, MRI) leverages phase sequences to uncover temporal dynamics vital for diagnosing tumors, lesions, and vascular issues. However, current retrieval models primarily focus on spatial features, neglecting phase-specific progression detailed in clinical rep

Cited by 0SourcePDFScholar
2025

The Role of Deductive and Inductive Reasoning in Large Language Models

ACL 2025long

Large Language Models (LLMs) have demonstrated impressive capabilities in reasoning tasks, yet their reliance on static prompt structures and limited adaptability to complex scenarios remains a major challenge. In this paper, we propose the **Deductive and Inductive (DID)** method, a novel framework…

Cited by 0SourcePDFScholar
2025

ToSA: Token Merging with Spatial Awareness

IROS 2025

Token merging has emerged as an effective strategy to accelerate Vision Transformers (ViT) by reducing computational costs. However, existing methods primarily rely on the visual token’s feature similarity for token merging, overlooking the potential of integrating spatial information, which can ser

Cited by 5SourcecodeScholar
2025

Zero-shot 3D Question Answering via Voxel-based Dynamic Token Compression

CVPR 2025poster

Recent advancements in 3D Large Multi-modal Models (3D-LMMs) have driven significant progress in 3D question answering. However, recent multi-frame Vision-Language Models (VLMs) demonstrate superior performance compared to 3D-LMMs on 3D question answering tasks, largely due to the greater scale and…

Cited by 0SourcePDFScholar
2024

2D Human Pose Estimation Calibration and Keypoint Visibility Classification

ICASSP 2024accepted

The confidence scores of 2D pose estimation are widely utilized in various fields, including multi-view 3D human pose estimation, skeleton-based human tracking, human action recognition, human re-identification, etc. Despite widespread use, confidence scores from 2D pose estimation methods are unrel…

Cited by 0SourceScholar
2024

A Density-Guided Temporal Attention Transformer for Indiscernible Object Counting in Underwater Videos

ICASSP 2024accepted

Dense object counting or crowd counting has come a long way thanks to the recent development in the vision community. However, indiscernible object counting, which aims to count the number of targets that are blended with respect to their surroundings, has been a challenge. Image-based object counti…

Cited by 0SourceScholar
2024

Contrastive Pre-Training with Multi-View Fusion for No-Reference Point Cloud Quality Assessment

CVPR 2024poster

No-reference point cloud quality assessment (NR-PCQA) aims to automatically evaluate the perceptual quality of distorted point clouds without available reference which have achieved tremendous improvements due to the utilization of deep neural networks. However learning-based NR-PCQA methods suffer…

Cited by 17SourcePDFScholar
2024

MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

CVPR 2024poster

Recently integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific pre-defined vision tasks. Yet existing systems can only handle videos with very few frames. For long videos the computation complexity memory cost and…

2024

UniAP: Towards Universal Animal Perception in Vision via Few-Shot Learning

AAAI 2024technical

Animal visual perception is an important technique for automatically monitoring animal health, understanding animal behaviors, and assisting animal-related research. However, it is challenging to design a deep learning-based perception model that can freely adapt to different animals across various…

Cited by 9SourcePDFScholar
2023

Global Adaptation Meets Local Generalization: Unsupervised Domain Adaptation for 3D Human Pose Estimation

ICCV 2023poster

When applying a pre-trained 2D-to-3D Human Pose lifting model to a target unseen dataset, a large performance degradation is commonly encountered due to domain shift issues. We observe that the degradation is caused by two factors: 1) the large distribution gap over global positions of poses between…

Cited by 30PDFcodeScholar
2022

GLIPv2: Unifying Localization and Vision-Language Understanding

NeurIPS 2022accept

We present GLIPv2, a grounded VL understanding model, that serves both localization tasks (e.g., object detection, instance segmentation) and Vision-Language (VL) understanding tasks (e.g., VQA, image captioning). GLIPv2 elegantly unifies localization pre-training and Vision-Language Pre-training (V…

2022

Grounded Language-Image Pre-Training

CVPR 2022oral

This paper presents a grounded language-image pre-training (GLIP) model for learning object-level, language-aware, and semantic-rich visual representations. GLIP unifies object detection and phrase grounding for pre-training. The unification brings two benefits: 1) it allows GLIP to learn from both…

Cited by 1294PDFcodeScholar
2022

LUNA: Localizing Unfamiliarity Near Acquaintance for Open-Set Long-Tailed Recognition

AAAI 2022technical

The predefined artificially-balanced training classes in object recognition have limited capability in modeling real-world scenarios where objects are imbalanced-distributed with unknown classes. In this paper, we discuss a promising solution to the Open-set Long-Tailed Recognition (OLTR) task utili…

Cited by 14SourcePDFScholar
2021

ACE: Ally Complementary Experts for Solving Long-Tailed Recognition in One-Shot

ICCV 2021poster

One-stage long-tailed recognition methods improve the overall performance in a "seesaw" manner, i.e., either sacrifice the head's accuracy for better tail classification or elevate the head's accuracy even higher but ignore the tail. Existing algorithms bypass such trade-off by a multi-stage trainin…

Cited by 185PDFcodeScholar
2021

Absolute 3d Pose Estimation and Length Measurement of Severely Deformed Fish from Monocular Videos in Longline Fishing

ICASSP 2021accepted

Monocular absolute 3D fish pose estimation allows for efficient fish length measurement in the longline fisheries, where fishes are under severe deformation during the catching process. This task is challenging since it requires locating absolute 3D fish keypoints based on a short monocular video cl…

Cited by 0SourceScholar
2021

Hierarchical Pose Classification for Infant Action Analysis and Mental Development Assessment

ICASSP 2021accepted

Based on Alberta Infant Motor Scale (AIMS), a questionnaire that tracks an infant’s motor function, an infant’s mental development can be evaluated by recording poses a baby can achieve. Therefore, it is meaningful to propose a systematic image-based pose classifier to classify infant actions based…

Cited by 0SourceScholar
2021

Track Without Appearance: Learn Box and Tracklet Embedding With Local and Global Motion Patterns for Vehicle Tracking

ICCV 2021poster

Vehicle tracking is an essential task in the multi-object tracking (MOT) field. A distinct characteristic in vehicle tracking is that the trajectories of vehicles are fairly smooth in both the world coordinate and the image coordinate. Hence, models that capture motion consistencies are of high nece…

Cited by 77PDFcodeScholar
2021

Vehicle 3d Localization in Road Scenes VIA a Monocular Moving Camera

ICASSP 2021accepted

Knowing the 3D locations of the surrounding vehicles is of vital importance in autonomous driving scenarios. It can be pretty challenging to make an accurate estimation from a monocular moving camera. In this paper, we present an effective vehicle 3D localization method, that utilizes 2D key-points…

Cited by 0SourceScholar
2019

CityFlow: A City-Scale Benchmark for Multi-Target Multi-Camera Vehicle Tracking and Re-Identification

CVPR 2019oral

Urban traffic optimization using traffic cameras as sensors is driving the need to advance state-of-the-art multi-target multi-camera (MTMC) tracking. This work introduces CityFlow, a city-scale traffic camera dataset consisting of more than 3 hours of synchronized HD videos from 40 cameras across 1…

Cited by 511PDFcodeScholar
2016

Chute based automated fish length measurement and water drop detection

ICASSP 2016accepted

Image processing and analysis techniques have drawn increasing attention since they enable a non-extractive and non-lethal approach to fisheries survey, such as fish size measurement, abundance prediction, catch estimation and compliance, species recognition and population counting. In this work, we…

Cited by 0SourceScholar
2016

Multiple-kernel adaptive segmentation and tracking (MAST) for robust object tracking

ICASSP 2016accepted

In a video surveillance system with static cameras, object segmentation often fails when part of the object has similar color with the background, resulting in poor performance of the subsequent object tracking. Multiple kernels have been utilized in object tracking to deal with occlusion, but the p…

Cited by 0SourceScholar
2015

Combined estimation of camera link models for human tracking across nonoverlapping cameras

ICASSP 2015accepted

Human tracking across multiple cameras is highly demanded for large scale video surveillance. To successfully track human across multiple uncalibrated cameras that have no overlapping field of views, a system to train more reliable camera link models is proposed in this paper. We employ a novel appr…

Cited by 0SourceScholar
2015

Deformable multiple-kernel based human tracking using a moving camera

ICASSP 2015accepted

In this paper, we propose an innovative human tracking algorithm, which efficiently integrates the deformable part model (DPM) into the multiple-kernel based tracking using a moving camera. By representing each part model of a DPM detected human as a kernel, the proposed algorithm iteratively mean-s…

Cited by 0SourceScholar