← Search

Yongjae Lee

10 accepted papers

2026

Decision-focused Sparse Tangent Portfolio Optimization

ICML 2026poster

Sparse tangent portfolio optimization aims to learn an interpretable, low-cardinality portfolio in the tangency direction of the mean–variance frontier, yet the associated cardinality-constrained formulation is NP-hard and standard predict-then-optimize pipelines often misalign forecasting accuracy …

Cited by 0SourceScholar
2026

Position: Evaluating LLMs in Finance Requires Explicit Bias Consideration

ICML 2026poster

Large Language Models (LLMs) are increasingly integrated into financial workflows, but evaluation practice has not kept up. Finance-specific biases can inflate performance, contaminate backtests, and make reported results useless for any deployment claim. We identify five recurring biases in financi…

Cited by 0SourceScholar
2025

BitAbuse: A Dataset of Visually Perturbed Texts for Defending Phishing Attacks

NAACL 2025findings

Phishing often targets victims through visually perturbed texts to bypass security systems. The noise contained in these texts functions as an adversarial attack, designed to deceive language models and hinder their ability to accurately interpret the content. However, since it is difficult to obtai…

2025

Geodesic Flow Kernels for Semi-Supervised Learning on Mixed-Variable Tabular Dataset

AAAI 2025technical

Tabular data poses unique challenges due to its heterogeneous nature, combining both continuous and categorical variables. Existing approaches often struggle to effectively capture the underlying structure and relationships within such data. We propose GFTab (Geodesic Flow Kernels for Semi-Supervise…

2024

Ignore Me But Don’t Replace Me: Utilizing Non-Linguistic Elements for Pretraining on the Cybersecurity Domain

NAACL 2024findings

Cybersecurity information is often technically complex and relayed through unstructured text, making automation of cyber threat intelligence highly challenging. For such text domains that involve high levels of expertise, pretraining on in-domain corpora has been a popular method for language models…

Cited by 3SourcePDFScholar
2024

LP-3DGS: Learning to Prune 3D Gaussian Splatting

NeurIPS 2024poster

Recently, 3D Gaussian Splatting (3DGS) has become one of the mainstream methodologies for novel view synthesis (NVS) due to its high quality and fast rendering speed. However, as a point-based scene representation, 3DGS potentially generates a large number of Gaussians to fit the scene, leading to h…

Cited by 6SourcePDFScholar
2023

DarkBERT: A Language Model for the Dark Side of the Internet

ACL 2023long

Recent research has suggested that there are clear differences in the language used in the Dark Web compared to that of the Surface Web. As studies on the Dark Web commonly require textual analysis of the domain, language models specific to the Dark Web may provide valuable insights to researchers.…

2023

Deep Value Function Networks for Large-Scale Multistage Stochastic Programs

AISTATS 2023poster

A neural networks-based stagewise decomposition algorithm called Deep Value Function Networks (DVFN) is proposed for large-scale multistage stochastic programming (MSP) problems. Traditional approaches such as nested Benders decomposition and its stochastic variant, stochastic dual dynamic programmi…

2022

Shedding New Light on the Language of the Dark Web

NAACL 2022long

The hidden nature and the limited accessibility of the Dark Web, combined with the lack of public datasets in this domain, make it difficult to study its inherent characteristics such as linguistic properties. Previous works on text classification of Dark Web domain have suggested that the use of de…

Cited by 17SourcePDFScholar