~/AI BENCHMARK/jevman-an-open-source-benchmark-testing-ai-decision-models-via-real-time

Jevman: An Open-Source Benchmark Testing AI Decision Models via Real-Time Pac-Man

An open-source benchmark framework called "jevman" was released to evaluate low-latency AI decision models by having them play Pac-Man in real time. The benchmark compared six decision models, including jev 1.13, GPT-6 Luna, and Cloudflare's Clef, publishing a public leaderboard based on 100 runs per model. As AI development shifts from chat-based interfaces toward specialized decision-making APIs, jevman provides an interactive environment to stress-test model performance under real-time constraints. It allows developers to easily benchmark custom fine-tuned models against major open-weight and proprietary solutions. Across 100 evaluation runs, jev 1.13 achieved the highest mean score of 2,750 points (290 ms average latency), followed closely by GPT-6 Luna at 2,568 points (179 ms latency) and Clef Flash at 2,538 points (256 ms latency). Human players can also play as Pac-Man against ghost opponents controlled by selected decision models.

## BACKGROUND

AI decision models represent an emerging class of non-conversational endpoints designed to accept application states (such as JSON or video frames) and return structured choices or probabilities. Unlike general-purpose LLMs optimized for text chat, decision models prioritize millisecond-level responsiveness for automated agent actions in dynamic environments.

## REFERENCES

## KEYWORDS

#ai-benchmarks#llm-evaluation#decision-models#open-source

$ subscribe --daily

Jevman: An Open-Source Benchmark Testing AI Decision Models via Real-Time Pac-Man | Daily News