~/OPEN SOURCE /release-of-victoria-and-maple-open-weights-fine-tuned-models

Release of Victoria and Maple Open-Weights Fine-Tuned Models

Researchers have released two fine-tuned open-weights models: Victoria, an agentic coding model derived from Qwen3.8-Flash-Next with 44% of its experts pruned, and Maple, a fine-tune designed for Canadian regulatory and civic queries. Victoria was retrained natively in 4-bit NVFP4 format, achieving a 70.0% score on Terminal-Bench 2.1 along with high inference throughput. The release demonstrates how expert pruning paired with quantization-aware training can drastically cut LLM memory overhead while preserving execution accuracy. It also highlights how region-specific fine-tuning like Maple can reduce standard US-centric biases in foundation models. Victoria utilized REAP to reduce experts from 512 down to 288 per layer and achieved 280 tokens/second single-stream performance on a Dell B300 setup using speculative decoding draft heads. Maple improved full answer correctness on Canadian evaluation questions from 6.6% to 21.8% while increasing official government source citations to 62.9%.

## BACKGROUND

Mixture-of-Experts (MoE) architectures route input tokens to specialized sub-networks, enabling high parameter capacity with lower per-token compute, though they require massive memory. Router-weighted Expert Activation Pruning (REAP) compresses MoE models by pruning redundant experts based on gate values and activation norms. NVFP4 is NVIDIA's 4-bit floating-point format designed to maximize inference bandwidth and efficiency on modern hardware.

## REFERENCES

## KEYWORDS

#Open Source AI#LLM Pruning#Quantization#Model Fine-Tuning#AI Benchmarks

$ subscribe --daily

Release of Victoria and Maple Open-Weights Fine-Tuned Models | Daily News