Field note / ML systems

LightGBM missing values are a deployment contract.

A generated tree can preserve every threshold and still disagree with LightGBM if it sends missing features down the wrong branch.

Compiling a boosted-tree model to source code looks like a threshold-printing problem. The edge cases reveal the real job: reproducing the inference contract of the training runtime.

A split contains policy, not just a number

At a numerical node, ordinary values are compared with a threshold. A missing value does not participate in that comparison in the usual way. LightGBM stores a default direction for the node, and the prediction follows that branch when the feature is missing.

If generated code treats every missing value as zero, always routes it left, or relies on a target language’s accidental comparison behavior, it changes the model. The error may remain invisible on clean benchmark rows and appear only in production, where incomplete inputs are common.

Make the branch explicit

In the model-to-code compiler I built, each target receives an explicit missing-value expression and the emitted conditional changes with the node’s stored direction. Python, C++17, and JavaScript express the check differently, but they must choose the same leaf.

COMPILER PRINCIPLEDo not translate syntax. Translate observable behavior.

Parity tests need hostile rows

A useful parity suite cannot be a random sample of complete feature vectors. It should deliberately place missing values in features used by splits, exercise values on both sides of thresholds, and include multiple outputs when the model supports them.

The project’s verification path generates predictions from the original LightGBM model, restores missing values as NaN, runs every row through the generated module, and compares the outputs. The same cases then compile or execute in the target runtime. This catches translation errors at the boundary where they actually matter.

Reject what you cannot preserve

A compiler earns trust partly through refusal. Unknown missing-value modes, unsupported decision types, and malformed tree structures should fail generation with a useful message. Silent approximation creates a valid-looking artifact with invalid semantics.

Treat deployment as a model change

Moving inference out of the training library changes the execution environment. That deserves the same discipline as changing features or retraining: pinned model artifacts, adversarial parity cases, per-target execution, tolerances appropriate to floating-point behavior, and a record of unsupported constructs.

The broader LightGBM model-to-code case study covers the compiler architecture and runnable verification path. Missing-value routing is the small detail that makes the larger point concrete: production parity lives in edge-case semantics.