Joonhun Lee

Physical AI Engineer from 🇰🇷 for now, working on production-level dual-arm robots. I translate research-stage ideas into deployable systems that hold up under real-world constraints — including RL execution policies that ran live in APAC equity and futures markets.

  • Reinforcement Learning
  • Dual-Arm Manipulation
  • Uncertainty & Distribution Shift
  • Live Decision Systems

Experience

PLAiF

Physical AI Engineer

The open question is not whether a dual-arm policy can move in a simulator, but whether the reward and the evaluation still mean something under the constraints of a real line. Building those policies and rewards for production robots, with transferable manipulation skills across manufacturing processes as the longer aim rather than a result already shown.

Aug '26 - Present

Qraft Technologies

AI Research Team Lead

Live execution can fail in the gap between a decision and its fill, and between what training can see and what serving can see. A validation basis-point figure is an upper bound on live edge, not the edge. Formulated the objective, state, action, and reward, then held candidates to champion-challenger evaluation against the production baseline across KRX equities and index futures, TWSE/TPEx equities, and HKEX equities. The same evaluation discipline covered trading-research agents and asset allocation.

AI Researcher

The earlier systems had to learn in noisy, partially observable markets where an order and its fill do not land in the same step. Built the actor-learner stack and the release gate for that delay. Uncertainty-aware TD-error modeling and feature alignment stayed separate research, rather than production claims.

Dec '23 - Aug '26

Wavebridge

Quantitative Developer

A centralized-exchange fill and an on-chain fill are not the same event, so a shared backtest would have scored the wrong thing. Built venue-aware simulators and low-latency on-chain collection for liquidity-provision research, across 5+ global and 3+ Korean chains.

Sep '23 - Nov '23

DoctorNow

Chief of Staff

The slow step was pharmacy matching: distance-based allocation waited on stock the nearest pharmacy did not have. OCR-assisted extraction supported matching by inventory rather than by distance alone. Separately supported the close of a ₩40B Series B and led the launch of triage-room waiting-time prediction during the Omicron wave.

Oct '21 - Feb '22

Research

Generalized Gaussian Temporal Difference Error for Uncertainty-aware Reinforcement Learning

A Gaussian variance head assumes TD error is thin-tailed. Real transition noise is often heavy-tailed, so the shape, not only the variance, is the uncertainty. Replaced that head with a state-conditioned generalized-Gaussian shape and kurtosis-aware regularization, improving over Gaussian variants in several continuous- and discrete-control settings with task- and regime-dependent gains. Official code.

Under review · Equal contribution

Feature-aligned N-BEATS with Sinkhorn divergence

Aligning every residual would have disturbed N-BEATS residual stacking, which is how the model forecasts. Aligned stack-wise feature measures with Sinkhorn divergence instead, so domain shift could be reduced without flattening that basis. Out-of-domain accuracy improved on macroeconomic and weather benchmarks under severe distribution shift. Official code.

ICLR '24 Spotlight · Equal contribution

MINR: Implicit Neural Representations with Masked Image Modelling

A representation that only reconstructs the masks it trained on has not learned the image. Predicted implicit neural representation weights from the visible pixels, so reconstruction could hold for unseen masks and out-of-domain images, at roughly 7× fewer parameters than MAE-Large. Official code.

ICCV '23 Workshop · Equal contribution

Education

Seoul National University

M.S. in Computational Science and Technology
Mar '22 - Feb '24

CFA Institute

Passed all three levels of the CFA Program

Seoul National University

B.S. in Physics Education
Mar '17 - Feb '22

Skills

AI / ML
  • Reinforcement Learning
  • Distributional RL
  • Uncertainty-Aware RL
  • Deep Learning
  • Time-Series Modeling
  • Representation Learning
  • Domain Generalization
Evaluation
  • Offline-to-Live Evaluation
  • Candidate Validation
  • Baseline Comparison
  • Failure-Mode Analysis
  • Release Readiness
Programming
  • Python
  • PyTorch
  • SQL
  • Rust
  • Go (familiar)
  • C++ (familiar)

Interests

Personal Agents · Mechanical Watches · Golf & Tennis