webAI Releases TwIL LM3 Pro: A 3.66B Logic-Specialized Model
webAI released TwIL LM3 Pro, a 3.66B parameter open model built on IBM Granite 4.2 3B and fine-tuned specifically for formal logic tasks using reinforcement learning against programmatic verifiers. On formal logic benchmarks, it achieves a composite score of 55.4 and 95.4% on BBH logic, performing on par with or ahead of larger models like Qwen3-8B. This release demonstrates that small language models can match the formal reasoning capabilities of models twice their size when trained with programmatic verifier reinforcement learning (RLVR). It offers a compact model—down to 2.09 GiB in Q4 GGUF format—capable of running on local consumer hardware to act as a logic checker for contract rules, policy verification, or agent step validation. The model was trained using LoRA SFT, checkpoint merging, and RL against code verifiers, though it trades off some performance on general benchmarks (averaging 79.0). It generates long chain-of-thought responses averaging around 1,900 tokens per prompt, with roughly 25% of responses hitting context length limits, and is released under a non-commercial license.
## BACKGROUND
Programmatic verifier reinforcement learning (RLVR) trains language models by providing exact, checkable feedback from code execution or automated theorem provers like Lean, eliminating reliance on human feedback or LLM evaluators. Benchmarks like BIG-Bench Hard (BBH) evaluate complex multi-step deductive logic and reasoning tasks that general-purpose language models traditionally struggle to solve.