~/AI HARDWARE/south-korean-startup-hyperaccel-mass-produces-4nm-bertha-ai-inference-chip

South Korean Startup HyperAccel Mass-Produces 4nm 'Bertha' AI Inference Chip

South Korean AI startup HyperAccel, in partnership with design solution provider SEMIFIVE, has entered mass production of its Bertha 500 LLM inference chip on Samsung Foundry's 4nm node. The chip features a large die area exceeding 500mm² designed specifically for data center LLM inference workloads. This milestone highlights the growing momentum behind custom, domain-specific AI accelerators challenging general-purpose GPUs in cost and power efficiency for inference. It also demonstrates Samsung Foundry and SEMIFIVE's capabilities in delivering turnkey, large-die custom silicon solutions for emerging hardware startups. The Bertha 500 delivers 768 TFLOPS of FP8 compute performance within a 250W TDP envelope, featuring 256MB of on-chip SRAM and up to 256GB of LPDDR5X memory (546GB/s bandwidth). HyperAccel claims the dual-slot PCIe accelerator provides up to 2x the throughput, 19x the cost-effectiveness, and 12x the energy efficiency of an NVIDIA H100 GPU for LLM inference.

## BACKGROUND

Large language model (LLM) inference requires high memory bandwidth and low latency to generate output tokens efficiently at scale. While general-purpose GPUs like NVIDIA's H100 dominate AI workloads, custom application-specific integrated circuits (ASICs) optimized for token generation can offer significant cost and power savings. SEMIFIVE functions as a design platform provider that enables startups to bring complex, custom SoCs to volume production on Samsung's advanced semiconductor processes.

## REFERENCES

## KEYWORDS

#AI Hardware#Semiconductors#AI Inference#Data Center#Samsung Foundry

$ subscribe --daily

South Korean Startup HyperAccel Mass-Produces 4nm 'Bertha' AI Inference Chip | Daily News