~/LLMS/abliterlitics-analysis-of-23-gemma-4-e4b-models-reveals-popular-fine-tunes

Abliterlitics Analysis of 23 Gemma 4 E4B Models Reveals Popular Fine-Tunes Are Broken

An evaluation of 23 Gemma 4 E4B models using the abliterlitics framework revealed that the most downloaded fine-tune, "OBLITERATUS", is severely broken and performs worse than the base model. Additionally, reasoning distill fine-tunes from Claude and Gemini failed to improve reasoning, instead degrading the model's capabilities. This highlights a critical issue in the open-source AI community where highly popular, unverified fine-tunes and "uncensored" models can be severely degraded compared to their base versions. It underscores the necessity of systematic evaluation tools like abliterlitics to verify model claims before deployment. The "OBLITERATUS" model, despite having nearly 800,000 downloads, showed the lowest Attack Success Rate (ASR) on HarmBench and a high KL divergence of 1.1. In contrast, the "heretic" variants performed best, achieving around 95% ASR on HarmBench while preserving most of the base model's capabilities.

## BACKGROUND

Ablitteration is a technique used to uncensor Large Language Models (LLMs) by identifying and removing the mathematical directions in the model's weights that cause it to refuse prompts. Abliterlitics is an open-source forensic toolkit designed to analyze and compare these safety-removed variants against their base models to measure structural and behavioral changes.

## REFERENCES

## KEYWORDS

#LLMs#Model Evaluation#Gemma#Open Source AI

$ subscribe --daily

Abliterlitics Analysis of 23 Gemma 4 E4B Models Reveals Popular Fine-Tunes Are Broken | Daily News