← Search

Hui Huang

66 accepted papers

2026

A Reasoning Paradigm for Named Entity Recognition

AAAI 2026technical

Generative LLMs typically improve Named Entity Recognition (NER) performance through instruction tuning. They excel at generating entities by semantic pattern matching but lack an explicit, verifiable reasoning mechanism. This "cognitive shortcutting" leads to suboptimal performance and weak general

Cited by 0SourcePDFScholar
2026

Augmented Radiance Field: A General Framework for Enhanced Gaussian Splatting

ICLR 2026poster

Due to the real-time rendering performance, 3D Gaussian Splatting (3DGS) has emerged as the leading method for radiance field reconstruction. However, its reliance on spherical harmonics for color encoding inherently limits its ability to separate diffuse and specular components, making it challengi…

Cited by 0SourceScholar
2026

CARE: Class-Adaptive Expert Consensus for Reliable Learning with Long-Tailed Noisy Labels

ICML 2026poster

Learning from real-world data is frequently hindered by the compound challenge of long-tailed class distributions and noisy annotations. Existing methods partially address these issues but typically ignore the non-uniform impact of label noise across classes, resulting in ineffective correction for …

Cited by 0SourceScholar
2026

Cycle-Consistent Tuning for Layered Image Decomposition

CVPR 2026

Disentangling visual layers in real-world images is a persistent challenge in vision and graphics, as such layers often involve non-linear and globally coupled interactions, including shading, reflection, and perspective distortion. In this work, we present an in-context image decomposition framewor

Cited by 0SourcecodeScholar
2026

EnvSocial-Diff: A Diffusion-Based Crowd Simulation Model with Environmental Conditioning and Individual-Group Interaction

ICLR 2026poster

Modeling realistic pedestrian trajectories requires accounting for both social interactions and environmental context, yet most existing approaches largely emphasize social dynamics. We propose EnvSocial-Diff: a diffusion-based crowd simulation model informed by social physics and augmented with env…

Cited by 0SourceScholar
2026

Long-form RewardBench: Evaluating Reward Models for Long-form Generation

AAAI 2026technical

The widespread adoption of reinforcement learning-based alignment highlights the growing importance of reward models. Various benchmarks have been built to evaluate reward models in various domains and scenarios. However, a significant gap remains in assessing reward models for long-form generation,

Cited by 0SourcePDFScholar
2026

Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory

AAAI 2026technical

The evaluation of large language models (LLMs) via benchmarks is widespread, yet inconsistencies between different leaderboards and poor separability among top models raise concerns about their ability to accurately reflect authentic model capabilities. This paper provides a critical analysis of ben

Cited by 0SourcePDFScholar
2026

Multi-view Crowd Tracking Transformer with View-Ground Interactions Under Large Real-World Scenes

CVPR 2026

Multi-view crowd tracking estimates each person's tracking trajectories on the ground of the scene. Recent research works mainly rely on CNNs-based multi-view crowd tracking architectures, and most of them are evaluated and compared on relatively small datasets, such as Wildtrack and MultiviewX. Sin

Cited by 0SourcecodeScholar
2026

RM-Distiller: Exploiting Generative LLM for Reward Model Distillation

IJCAI 2026

Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. Due to the difficulty of obtaining high-quality human preference annotations, distilling preferences from generative LLMs has emerged as a standard practice. However, existing approaches predomin

Cited by 0Scholar
2026

StrokeFusion: Vector Sketch Generation via Joint Stroke-UDF Encoding and Latent Sequence Diffusion

AAAI 2026technical

In the field of sketch generation, raster-format trained models often produce non-stroke artifacts, while vector-format trained models typically lack a holistic understanding of sketches, leading to compromised recognizability. Moreover, existing methods struggle to extract common features from simi

Cited by 0SourcePDFScholar
2026

TG-Field: Geometry-Aware Radiative Gaussian Fields for Tomographic Reconstruction

AAAI 2026technical

3D Gaussian Splatting (3DGS) has revolutionized 3D scene representation with superior efficiency and quality. While recent adaptations for computed tomography (CT) show promise, they struggle with severe artifacts under highly sparse-view projections and dynamic motions. To address these challenges,

Cited by 0SourcePDFScholar
2026

Think-J: Learning to Think for Generative LLM-as-a-Judge

AAAI 2026technical

LLM-as-a-Judge refers to the automatic modeling of preferences for responses generated by Large Language Models (LLMs), which is of significant importance for both LLM evaluation and reward modeling. Although generative LLMs have made substantial progress in various tasks, their performance as LLM-J

Cited by 0SourcePDFScholar
2025

2D-DPO: Scaling Direct Preference Optimization with 2-Dimensional Supervision

NAACL 2025findings

Recent advancements in Direct Preference Optimization (DPO) have significantly enhanced the alignment of Large Language Models (LLMs) with human preferences, owing to its simplicity and effectiveness. However, existing methods typically optimize a scalar score or ranking reward, thereby overlooking…

Cited by 2SourcePDFScholar
2025

AIR: Complex Instruction Generation via Automatic Iterative Refinement

EMNLP 2025

With the development of large language models, their ability to follow simple instructions has significantly improved. However, adhering to complex instructions remains a major challenge. Current approaches to generating complex instructions are often irrelevant to the current instruction requiremen

2025

An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

ACL 2025finding

Recently, there has been a growing trend of utilizing Large Language Model (LLM) to evaluate the quality of other LLMs. Many studies have fine-tuned judge models based on open-source LLMs for evaluation. While the fine-tuned judge models are claimed to achieve comparable evaluation capability with G…

2025

ArcPro: Architectural Programs for Structured 3D Abstraction of Sparse Points

CVPR 2025highlight

We introduce ArcPro, a novel learning framework built on architectural programs to recover structured 3D abstractions from highly sparse and low-quality point clouds. Specifically, we design a domain-specific language (DSL) to hierarchically represent building structures as a program, which can be e…

Cited by 0SourcePDFScholar
2025

Attention Distillation: A Unified Approach to Visual Characteristics Transfer

CVPR 2025poster

Recent advances in generative diffusion models have shown a notable inherent understanding of image style and semantics. In this paper, we leverage the self-attention features from pretrained diffusion networks to transfer the visual characteristics from a reference to generated images. Unlike previ…

2025

Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models

ACL 2025long

New LLM benchmarks are important to align with the rapid development of Large Language Models (LLMs). In this work, we present Chinese SimpleQA, the first comprehensive Chinese benchmark to evaluate the factuality ability of LLMs to answer short questions, and Chinese SimpleQA mainly has five proper…

2025

DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language Models

NAACL 2025long

Multimodal Large Language Models (MLLMs) pose unique safety challenges due to their integration of visual and textual data, thereby introducing new dimensions of potential attacks and complex risk combinations. In this paper, we begin with a detailed analysis aimed at disentangling risks through ste…

2025

EmoEdit: Evoking Emotions through Image Manipulation

CVPR 2025poster

Affective Image Manipulation (AIM) seeks to modify user-provided images to evoke specific emotions. This task is inherently complex due to its twofold objective: evoking the intended emotion while preserving image composition. Existing AIM methods primarily adjust color and style, often failing to e…

2025

LDIR: Low-Dimensional Dense and Interpretable Text Embeddings with Relative Representations

ACL 2025finding

Semantic text representation is a fundamental task in the field of natural language processing. Existing text embedding (e.g., SimCSE and LLM2Vec) have demonstrated excellent performance, but the values of each dimension are difficult to trace and interpret. Bag-of-words, as classic sparse interpret…

2025

Legal Fact Prediction: The Missing Piece in Legal Judgment Prediction

EMNLP 2025

Legal judgment prediction (LJP), which enables litigants and their lawyers to forecast judgment outcomes and refine litigation strategies, has emerged as a crucial legal NLP task. Existing studies typically utilize legal facts, i.e., facts that have been established by evidence and determined by the

2025

MuSC: Improving Complex Instruction Following with Multi-granularity Self-Contrastive Training

ACL 2025long

Complex instruction-following with elaborate constraints is imperative for Large Language Models (LLMs). While existing methods have constructed data for complex instruction alignment, they all rely on a more advanced model, especially GPT-4, limiting their application. In this paper, we propose a M…

2025

Out-of-Distribution Detection with Prototypical Outlier Proxy

AAAI 2025technical

Out-of-distribution (OOD) detection is a crucial task for deploying deep learning models in the wild. One of the major challenges is that well-trained deep models tend to perform over-confidence on unseen test data. Recent research attempts to leverage real or synthetic outliers to mitigate the issu…

2025

Towards Multi-Document Question Answering in Scientific Literature: Pipeline, Dataset, and Evaluation

EMNLP 2025

Question-Answering (QA) systems are vital for rapidly accessing and comprehending information in academic literature.However, some academic questions require synthesizing information across multiple documents. While several prior resources consider multi-document QA, they often do not strictly enfor

2025

View Transformation Robustness for Multi-View 3D Object Reconstruction with Reconstruction Error-Guided View Selection

AAAI 2025technical

View transformation robustness (VTR) is critical for deep-learning-based multi-view 3D object reconstruction models, which indicates the methods' stability under inputs with various view transformations. However, existing research seldom focused on view transformation robustness in multi-view 3D obj…

2024

CRAYM: Neural Field Optimization via Camera RAY Matching

NeurIPS 2024poster

We introduce camera ray matching (CRAYM) into the joint optimization of camera poses and neural fields from multi-view images. The optimized field, referred to as a feature volume, can be “probed” by the camera rays for novel view synthesis (NVS) and 3D geometry reconstruction. One key reason for ma…

Cited by 1SourcePDFScholar
2024

EmoGen: Emotional Image Content Generation with Text-to-Image Diffusion Models

CVPR 2024poster

Recent years have witnessed remarkable progress in image generation task where users can create visually astonishing images with high-quality. However exsiting text-to-image diffusion models are proficient in generating concrete concepts (dogs) but encounter challenges with more abstract ones (emoti…

Cited by 21SourcePDFScholar
2024

FRI-Net: Floorplan Reconstruction via Room-wise Implicit Representation

ECCV 2024poster

"In this paper, we introduce a novel method called FRI-Net for 2D floorplan reconstruction from 3D point cloud. Existing methods typically rely on corner regression or box regression, which lack consideration for the global shapes of rooms. To address these issues, we propose a novel approach using…

Cited by 0SourcePDFScholar
2024

Feature Fusion from Head to Tail for Long-Tailed Visual Recognition

AAAI 2024technical

The imbalanced distribution of long-tailed data presents a considerable challenge for deep learning models, as it causes them to prioritize the accurate classification of head classes but largely disregard tail classes. The biased decision boundary caused by inadequate semantic information in tail c…

2024

Generating Non-Stationary Textures using Self-Rectification

CVPR 2024poster

This paper addresses the challenge of example-based non-stationary texture synthesis. We introduce a novel two-step approach wherein users first modify a reference texture using standard image editing tools yielding an initial rough target for the synthesis. Subsequently our proposed method termed "…

2024

Improving Visual Prompt Tuning by Gaussian Neighborhood Minimization for Long-Tailed Visual Recognition

NeurIPS 2024poster

Long-tailed visual recognition has received increasing attention recently. Despite fine-tuning techniques represented by visual prompt tuning (VPT) achieving substantial performance improvement by leveraging pre-trained knowledge, models still exhibit unsatisfactory generalization performance on tai…

2024

InterFusion: Text-Driven Generation of 3D Human-Object Interaction

ECCV 2024poster

"In this study, we tackle the complex task of generating 3D human-object interactions (HOI) from textual descriptions in a zero-shot text-to-3D manner. We identify and address two key challenges: the unsatisfactory outcomes of direct text-to-3D methods in HOI, largely due to the lack of paired text-…

2024

Multi-View People Detection in Large Scenes via Supervised View-Wise Contribution Weighting

AAAI 2024technical

Recent deep learning-based multi-view people detection (MVD) methods have shown promising results on existing datasets. However, current methods are mainly trained and evaluated on small, single scenes with a limited number of multi-view frames and fixed camera views. As a result, these methods may…

2024

Reconstruct and Match: Out-of-Distribution Robustness via Topological Homogeneity

NeurIPS 2024spotlight

Since deep learning models are usually deployed in non-stationary environments, it is imperative to improve their robustness to out-of-distribution (OOD) data. A common approach to mitigate distribution shift is to regularize internal representations or predictors learned from in-distribution (ID) d…

Cited by 0SourcePDFScholar
2024

Self-Evaluation of Large Language Model based on Glass-box Features

EMNLP 2024finding

The proliferation of open-source Large Language Models (LLMs) underscores the pressing need for evaluation methods. Existing works primarily rely on external evaluators, focusing on training and prompting strategies. However, a crucial aspect – model-aware glass-box features – is overlooked. In this…

2024

Synchronized Dual-arm Rearrangement via Cooperative mTSP

ICRA 2024poster

Synchronized dual-arm rearrangement is widely studied as a common scenario in industrial applications. It often faces scalability challenges due to the computational complexity of robotic arm rearrangement and the high-dimensional nature of dual-arm planning. To address these challenges, we formulat…

Cited by 0SourceScholar
2024

TERD: A Unified Framework for Safeguarding Diffusion Models Against Backdoors

ICML 2024poster

Diffusion models have achieved notable success in image generation, but they remain highly vulnerable to backdoor attacks, which compromise their integrity by producing specific undesirable outputs when presented with a pre-defined trigger. In this paper, we investigate how to protect diffusion mode…

2023

ARO-Net: Learning Implicit Fields From Anchored Radial Observations

CVPR 2023poster

We introduce anchored radial observations (ARO), a novel shape encoding for learning implicit field representation of 3D shapes that is category-agnostic and generalizable amid significant shape variations. The main idea behind our work is to reason about shapes through partial observations from a s…

2023

EmoSet: A Large-scale Visual Emotion Dataset with Rich Attributes

ICCV 2023poster

Visual Emotion Analysis (VEA) aims at predicting people's emotional responses to visual stimuli. This is a promising, yet challenging, task in affective computing, which has drawn increasing attention in recent years. Most of the existing work in this area focuses on feature design, while little att…

Cited by 52PDFScholar
2023

Improving Translation Quality Estimation with Bias Mitigation

ACL 2023long

State-of-the-art translation Quality Estimation (QE) models are proven to be biased. More specifically, they over-rely on monolingual features while ignoring the bilingual semantic alignment. In this work, we propose a novel method to mitigate the bias of the QE model and improve estimation performa…

Cited by 5SourcePDFScholar
2023

Iterative Nearest Neighbour Machine Translation for Unsupervised Domain Adaptation

ACL 2023findings

Unsupervised domain adaptation of machine translation, which adapts a pre-trained translation model to a specific domain without in-domain parallel data, has drawn extensive attention in recent years. However, most existing methods focus on the fine-tuning based techniques, which is non-extensible.…

2023

NIFT: Neural Interaction Field and Template for Object Manipulation

ICRA 2023poster

We introduce NIFT, Neural Interaction Field and Template, a descriptive and robust interaction representation of object manipulations to facilitate imitation learning. Given a few object manipulation demos, NIFT guides the generation of the interaction imitation for a new object instance by matching…

Cited by 9SourceScholar
2023

Semi-Weakly Supervised Object Kinematic Motion Prediction

CVPR 2023poster

Given a 3D object, kinematic motion prediction aims to identify the mobile parts as well as the corresponding motion parameters. Due to the large variations in both topological structure and geometric details of 3D objects, this remains a challenging task and the lack of large scale labeled data als…

Cited by 11SourcePDFScholar
2022

"Capturing, Reconstructing, and Simulating: The UrbanScene3D Dataset"

ECCV 2022poster

"We present UrbanScene3D, a large-scale data platform for research of urban scene perception and reconstruction. UrbanScene3D contains over 128k high-resolution images covering 16 scenes including large-scale real urban regions and synthetic cities with 136 km2 area in total. The dataset also contai…

2022

Iterative Constrained Back-Translation for Unsupervised Domain Adaptation of Machine Translation

COLING 2022main

Back-translation has been proven to be effective in unsupervised domain adaptation of neural machine translation (NMT). However, the existing back-translation methods mainly improve domain adaptability by generating in-domain pseudo-parallel data that contains sentence-structural knowledge, paying l…

2022

ShapeFormer: Transformer-Based Shape Completion via Sparse Representation

CVPR 2022poster

We present ShapeFormer, a transformer-based network that produces a distribution of object completions, conditioned on incomplete, and possibly noisy, point clouds. The resultant distribution can then be sampled to generate likely completions, each of which exhibits plausible shape details, while be…

Cited by 160PDFScholar
2022

Sim2real Learning of Obstacle Avoidance for Robotic Manipulators in Uncertain Environments

RA-L 2022

Obstacle avoidance for robotic manipulators can be challenging when they operate in unstructured environments. This problem is probed with the sim-to-real (sim2real) deep reinforcement learning, such that a moving policy of the robotic arm is learnt in a simulator and then adapted to the real world.

Cited by 38SourceScholar
2021

Improving Stylized Neural Machine Translation with Iterative Dual Knowledge Transfer

IJCAI 2021poster

Stylized neural machine translation (NMT) aims to translate sentences of one style into sentences of another style, which is essential for the application of machine translation in a real-world scenario. However, a major challenge in this task is the scarcity of high-quality parallel data which is s…

2021

Origami-Inspired Snap-through Bistability in Parallel and Curved Mechanisms Through the Inflection of Degree Four Vertexes

ICRA 2021poster

Origami, the art of folding paper, can impart useful design inspirations to the creation of mechanical structures and mechanisms. Bistability is a useful property for origami designs, which can help compartmentalize different actuations and stiffness tuning regimes. Given the benefits of bistability…

Cited by 10SourceScholar
2021

Saliency-based Multi-View Mixed Language Training for Zero-shot Cross-lingual Classification

EMNLP 2021finding

Recent multilingual pre-trained models, like XLM-RoBERTa (XLM-R), have been demonstrated effective in many cross-lingual tasks. However, there are still gaps between the contextualized representations of similar words in different languages. To solve this problem, we propose a novel framework named…

Cited by 8SourcePDFScholar
2019

ETNet: Error Transition Network for Arbitrary Style Transfer

NeurIPS 2019poster

Numerous valuable efforts have been devoted to achieving arbitrary style transfer since the seminal work of Gatys et al. However, existing state-of-the-art approaches often generate insufficiently stylized results under challenging cases. We believe a fundamental reason is that these approaches try…

2019

Patch-Based Progressive 3D Point Set Upsampling

CVPR 2019poster

We present a detail-driven deep neural network for point set upsampling. A high-resolution point set is essential for point-based rendering and surface reconstruction. Inspired by the recent success of neural image super-resolution techniques, we progressively train a cascade of patch-based upsampli…

Cited by 353PDFcodeScholar
2019

ZigZagNet: Fusing Top-Down and Bottom-Up Context for Object Segmentation

CVPR 2019poster

Multi-scale context information has proven to be essential for object segmentation tasks. Recent works construct the multi-scale context by aggregating convolutional feature maps extracted by different levels of a deep neural network. This is typically done by propagating and fusing features in a on…

Cited by 86PDFcodeScholar
2018

Multi-Scale Context Intertwining for Semantic Segmentation

ECCV 2018poster

Accurate semantic image segmentation requires the joint consideration of local appearance, semantic information, and global scene context. In today’s age of pre-trained deep networks and their powerful convolutional features, state-of-the-art semantic segmentation approaches differ mostly in how the…

Cited by 211SourcePDFScholar
2018

Specular-to-Diffuse Translation for Multi-View Reconstruction

ECCV 2018poster

Most multi-view 3D reconstruction algorithms, especially when shape-from-shading cues are used, assume that object appearance is predominantly diffuse. To alleviate this restriction, we introduce S2Dnet, a generative adversarial network for transferring multiple views of objects with specular reflec…

Cited by 28SourcePDFScholar
2017

Cascaded Feature Network for Semantic Segmentation of RGB-D Images

ICCV 2017poster

Fully convolutional network (FCN) has been successfully applied in semantic segmentation of scenes represented with RGB images. Images augmented with depth channel provide more understanding of the geometric information of the scene in the image. The question is how to best exploit this additional i…

Cited by 175PDFScholar
2017

Learning to Aggregate Ordinal Labels by Maximizing Separating Width

ICML 2017poster

While crowdsourcing has been a cost and time efficient method to label massive samples, one critical issue is quality control, for which the key challenge is to infer the ground truth from noisy or even adversarial data by various users. A large class of crowdsourcing problems, such as those involvi…

Cited by 9SourcePDFScholar