llama.cpp Release b11495 Adds Classifier Activation Support for Reranker Models
llama.cpp release b11495 introduces support for classifier activation functions in reranker models. The update maps classifier GELU activations to `gelu_erf`, accepts `tanh`, sets `tanh` as the default classifier activation, and configures ModernBERT models to fall back to `gelu_erf`. Reranker models are essential components in Retrieval-Augmented Generation (RAG) pipelines for accurately scoring and re-ordering retrieved documents. Proper support for modern classifier activation functions ensures that modern architectures, such as those built on ModernBERT, can run accurately during local inference. Implemented via PR #29692, the update sets `act_cls` to default to `tanh` for rerankers while routing ModernBERT to `gelu_erf`. Pre-built binary releases were made available across Linux, Windows, macOS, Android, and iOS platforms with multi-backend GPU/NPU acceleration options.
## BACKGROUND
llama.cpp is a widely used C/C++ inference library that allows running Large Language Models and embedding/reranker models locally with optimized hardware support. Rerankers are specialized cross-encoder models that evaluate query-passage relevance in search systems using a classification layer on top of transformer encoders. ModernBERT is a modernized BERT architecture trained on 2 trillion tokens that incorporates modern design choices like GeGLU activations and rotary positional embeddings.