llama.cpp Release b11532 Adds Exact GELU Support for ModernBERT
llama.cpp released version b11532, adding exact GELU activation function support for ModernBERT encoder models. The update also updates function mappings, including mapping gelu_python to ggml_geglu_erf. Proper activation alignment is essential for inference fidelity when porting transformer models to llama.cpp's lightweight backend. This ensures ModernBERT models perform accurately without precision degradation during encoding tasks. The release specifically maps `gelu_python` to `ggml_geglu_erf` while maintaining `tanh GELU` aliases on `ggml_geglu`. This fix addresses architectural nuances between PyTorch implementation variants and ggml internal operators.
## BACKGROUND
GELU (Gaussian Error Linear Unit) is a widely used smooth activation function in modern transformer architectures such as BERT and GPT. ModernBERT updates the classic BERT encoder structure with modernized techniques, requiring accurate operator compatibility in inference runtime engines like llama.cpp.