All About Circuits

Apple Targets AI Compute With New M6 and M5 Ultra Processors

M6 is Apple's first 2-nm chip and features a dual 16-core Neural Engine, while the M5 Ultra joins four dies to achieve 1.2 TB/s of unified memory bandwidth.


News September 11, 2026 by Jake Hertz

Continuing its run of in-house silicon, Apple recently announced the M6 and M5 Ultra desktop system-on-chips (SoCs). As AI becomes more pervasive in everyday workflows, Apple continues to raise the silicon bar to deliver local inference in new MacBook generations.

 

The M6 and M5 Ultra

The M6 and M5 Ultra. Image used courtesy of Apple
 

With features like a 2-nanometer (nm) process and a quad-die architecture, Apple turned to new architectural advancements to deliver greater performance and efficiency.

 

The M6

Today’s computers impose diverse workloads on the underlying processors, some of which thrive on one fast core, while others benefit from being spread out over many weaker cores. To improve performance across workloads, Apple designed M6 with three kinds of cores. Specifically, M6’s 12-core CPU includes two super cores, four performance cores, and six efficiency cores. By implementing this multi-core architecture on the company’s first 2-nm process, Apple claims M6 delivers the world's fastest single-threaded performance and 1.2 times faster multithreaded performance compared to M5.

 

M6

With a dual 16-core Neural Engine, a 12-core GPU with Neural Accelerators, and 170 GB/s of unified memory bandwidth, M6 is designed to boost on-device AI performance. Image used courtesy of Apple
​

In addition to improving CPU performance, Apple also enhanced on-device AI on M6 with the new Neural Engine, GPU Neural Accelerators, and improved unified memory bandwidth. The company claims M6’s Dual 16-core Neural Engine doubles the SoC’s peak compute against previous generations because its frameworks can schedule work onto both engines at once.

Apple also added a Neural Accelerator to each core of M6’s 12-core GPUs to deliver 30 percent higher peak AI compute over M5 and faster prompt processing. To unlock this level of AI performance, the company designed M6 to offer up to 170 GB/s of unified memory bandwidth.

 

The M5 Ultra

To scale all of M6’s performance to workstation levels, Apple designed M5 Ultra with a company-first quad-die architecture. The company created the SoC by connecting two dual-die M5 Max chips with UltraFusion technology, which unlocks more than 4.4 TB/s of inter-die bandwidth and over six times the previous generation's connection density. According to Apple, M5 Ultra’s high-performance die-to-die links let programmers ignore signal latency and treat the four dies as one processor.

​

M5 Ultra

With up to 512 GB of unified memory and 1.2 TB/s of memory bandwidth, the M5 Ultra is designed to let developers run large AI models locally. Image used courtesy of Apple
​

M5 Ultra’s quad-die architecture lets Apple include a 36-core CPU in the SoC. With 12 super cores and 24 performance cores, the chip is designed to offer 1.25 times the single-threaded and 1.3 times the multithreaded performance of M3 Ultra. Alongside the CPU, Apple fitted a 32-core Neural Engine and an 80-core GPU with a Neural Accelerator in every core into M5 Ultra, which the company credits with delivering up to 4.5 times the peak AI compute of M3 Ultra.

To improve performance further, Apple gave M5 Ultra up to 512 GB of unified memory at 1.2 TB/s, which is 50 percent higher than M3 Ultra’s memory bandwidth. AI developers using M5 Ultra can potentially keep an entire model in memory, which the company says will boost tokens-per-second speed.

 

Multi-Die Architectures

Chip designers may split a large processor into multiple dies for two main reasons. 

The first is the reticle limit, the largest area a lithography scanner can pattern in one exposure. No monolithic design can exceed that ceiling, which sits near 850 mm2, so chip designers who want more processor area must create multiple dies. 

The second reason is yield, or the share of chips on a wafer that come out without defects. The larger each chip, the higher the defect rate, so manufacturers have an incentive to split one large processor in two to boost yield.

​

UCIe’s SoC architecture

UCIe’s SoC architecture, a blueprint for multi-die designs. Image used courtesy of UCIe Consortium
 

However, once chip designers split one processor design into multiple dies, they face new problems, namely in inter-die communication. Signals that once stayed within a single die now must travel through a package substrate or interposer, adding latency and energy costs.

Thus, multi-die chip designers strive to maximize bandwidth per millimeter of die edge and minimize energy per bit for every die-to-die link. If they succeed, system programmers can treat the multiple dies as one cohesive processor. Otherwise, programmers will notice the costs from inter-die communication and be forced to adjust how the multiple dies share work to maintain overall system performance. Even with their best efforts, however, programmers cannot fully eliminate the costs from deficient die-to-die links, especially in high-compute workloads.

 

Pricing and Availability

The M6 and M5 Ultra are only available as part of Apple’s Mac desktops. The Mac mini hosts the M6, and the Mac Studio features the M5 Ultra, with 256 GB of unified memory and a 1 TB SSD configurable to 16 TB. Both systems are currently available for purchase, and Apple plans to ship them to customers on September 22.