Pentagon Reportedly Requested OpenAI Model with 'Minimum Refusal Rate' for Military Commands
Leaked FOIA documents obtained by The Intercept reveal that the US Department of Defense requested a custom OpenAI model for national security with a 'minimum refusal rate' when executing military commands. Both OpenAI and the Pentagon denied that the clause made it into the final executed agreement, claiming the released document was an early draft. The leak highlights intense ethical debates over deploying commercial frontier AI models within military and classified operations. It underscores growing concerns about whether standard safety guardrails preventing hazardous or lethal AI behavior are being removed for defense applications. The 'minimum refusal rate' phrasing appeared in documents for contract 'P00003,' an expansion of a prototype agreement worth up to $200 million. Although Department of Justice lawyers initially confirmed the document was an executed contract, officials later backtracked, and OpenAI recently signed a classified network agreement claiming strict red lines against autonomous killing and domestic surveillance.
## BACKGROUND
Commercial large language models are aligned using safety guardrails to automatically decline prompts involving harm, such as targeting drone strikes or synthesizing hazardous materials. In AI safety engineering, model refusals occur when safety classifiers or alignment mechanisms flag an input prompt as violating safety policies.