~/LLMS/uncensored-llms-are-measurably-more-optimistic-and-confident-than-their-base-models

"Uncensored" LLMs are measurably more optimistic and confident than their base models

An empirical study involving 21,600 decisions revealed that removing refusal mechanisms from LLMs like Gemma and Qwen alters their confidence and optimism levels. While the uncensored models exhibited longer, more confident reasoning and more optimistic predictions, their actual accuracy remained unchanged. This finding demonstrates that "abliteration" does not just bypass safety filters but fundamentally shifts a model's disposition and reasoning style. Understanding these behavioral drifts is crucial for developers deploying uncensored models in decision-making tasks like financial forecasting. The study observed model-dependent effects: after abliteration, confidence increased in Qwen but decreased in Gemma. Additionally, the uncensored models produced fewer hedging words like "maybe" or "uncertain" when predicting stock market movements.

## BACKGROUND

Uncensored LLMs are models modified to remove safety alignments or refusal mechanisms, allowing them to respond to sensitive prompts. "Abliteration" (refusal vector ablation) is a popular technique that achieves this without retraining by identifying and neutralizing the specific mathematical directions in the model's activation space that trigger refusals.

## REFERENCES

## KEYWORDS

#LLMs#AI Alignment#Abliteration#Model Evaluation

$ subscribe --daily

"Uncensored" LLMs are measurably more optimistic and confident than their base models | Daily News