~/LOCAL LLM/quantized-local-llm-injects-unrelated-knowledge-into-chain-of-thought-reasoning

Quantized Local LLM Injects Unrelated Knowledge into Chain-of-Thought Reasoning

A user reported that an ultra-low precision local language model (`strata-swift-iq3_xxs`) randomly injected a biography of Singapore's founding father, Lee Kuan Yew, into its internal chain-of-thought reasoning during a coding task. Despite this unexpected hallucination in its thought process, the model seamlessly returned to analyzing file line endings and completing the coding task. This incident demonstrates how aggressive quantization (such as IQ3_XXS, which uses roughly 3-bit precision) can degrade internal model routing and cause memory leakage or hallucinations during reasoning. It highlights the trade-offs between speed or VRAM efficiency and output stability when running large open-weight models on consumer GPUs. The model was running via the Strata inference engine on an Nvidia RTX 5070 Ti to double execution speed over 4-bit variants. While generating a step-by-step plan to detect end-of-line characters using a node script, the model printed several detailed sentences about Lee Kuan Yew's education and political career right between two coding instructions.

## BACKGROUND

Quantization reduces the numerical precision of a language model's weights to fit larger models into limited GPU VRAM. Modern reasoning models use chain-of-thought (CoT) prompting to write out an internal scratchpad before producing their final response, which can expose raw activation glitches caused by weight quantization.

## REFERENCES

## KEYWORDS

#Local LLM#Quantization#LLM Hallucinations#AI Engineering

$ subscribe --daily

Quantized Local LLM Injects Unrelated Knowledge into Chain-of-Thought Reasoning | Daily News