~/AI AGENTS/androidlife-benchmark-tests-qwen3-8-27b-ai-agent-on-real-android-phone

AndroidLife Benchmark Tests Qwen3.8-27B AI Agent on Real Android Phone

A real-device benchmark using AndroidLife tested Alibaba's Qwen3.8-27b text model across 60 continuous Android phone tasks, achieving a 56.7% success rate. The test revealed heavy hardware strain, draining 69% of the battery and pushing chip temperatures up to 98.2°C. This evaluation highlights the major gap between theoretical LLM performance and practical deployment on mobile devices, demonstrating severe thermal and power penalties. It shows that current AI agents still struggle with multi-app workflows and proactive user interaction in real-world mobile operating environments. Qwen3.8-27b averaged 29.25 steps (~6 minutes) and $0.118 per task, performing well on easy single-app tasks (80.8%) but failing sharply on hard multi-app workflows (23.5%). The agent frequently hallucinated completed outcomes on calendar tasks and failed to ask clarifying questions when encountering ambiguous user inputs.

## BACKGROUND

AndroidLife is a real-phone evaluation benchmark that tests AI agents on physical mobile devices via ADB across dozens of everyday applications, grading success based on actual device end-states. Qwen is a family of open-weight large language models developed by Alibaba Cloud, with models like Qwen-27B often used for complex agentic reasoning tasks.

## REFERENCES

## KEYWORDS

#ai-agents#llm-benchmarks#mobile-ai#on-device-ai#qwen

$ subscribe --daily

AndroidLife Benchmark Tests Qwen3.8-27B AI Agent on Real Android Phone | Daily News