AI is more important than everything else; the iPhone 18 Pro’s NPU is more powerful than the GPU

The new flagship devices from Apple, including iPhone 18 Pro, iPhone 18 Pro Max and the novel iPhone Duo, are set to debut on the market, driven by a processor in which artificial intelligence takes an absolute priority.

Rumors and early data reveal that the new chip A20 Pro boasts a Neural Engine capable of surpassing the peak power of the same integrated graphics chip.

Apple A20 Pro, the new NPU for AI is incredibly powerful

iPhone 18 Pro
Credits: Apple

The new SoC from Apple presents itself as a high-end processor, establishing new records in Geekbench 6 for single-core performance. Its power is such that it registers a 27% faster speed compared to the previous A19 Pro and M3 in multi-core tests, able even to outperform the mighty M5 Max.

To reveal a particularly interesting detail about this component and the M6 processor, Max Weinbach shared it via the platform X. Weinbach pointed out that the performances of the dual 16-core Neural Engine can surpass those of the 7-core GPU, provided the system must handle tasks specially programmed to exploit such architecture.

The resource management inside the A20 Pro follows an extremely precise logic, capable of freeing a huge workload from both the CPU and the GPU. The new Neural Engine handles with great ease on-device AI models of small size, with 3 billion parameters, and the timely processing of prompts.

The pre-compilation phase for quantized models aligns directly with the NPU pipelines, optimizing response times.

In addition are advanced features such as on-device voice recognition, real-time audio enhancements and numerous camera-related operations, including depth estimation of the surrounding space, application of background blur, text recognition and complex image segmentation.

The limits of the Neural Engine and the role of the GPU

Despite the incredible peak power of the NPU, several operations continue to require the intervention of the more classical processors. High-intensity graphical workloads and particular animations would prove extremely slow if entrusted to neural cores.

Similarly, non-quantized 32- or 64-bit floating-point operations, as well as handling tensors that change shape dynamically, remain a prerogative of CPU and GPU.

Even the generation of large language models with a single token can suffer dispatch delays if entrusted solely to the NPU architecture, making traditional processors much more suitable for the purpose.