Apple has introduced two new silicon architectures aimed at local artificial intelligence workloads: the M6, manufactured on a 2-nanometer process, and the M5 Ultra, a quad-die processor offering up to 512GB of unified memory.
The chips debut across updated desktop lines. The M6 powers an entry Mac mini starting at $899, while the M5 Ultra configures into the Mac Studio, where M5 Max base configurations start at $2,499. Both product lines are scheduled to begin customer deliveries on September 22, 2026.
Quad-Die Interconnect and Unified Memory Scaling
The primary architectural shift in the M5 Ultra is its multi-die packaging. Rather than building a monolithic die, Apple fused two dual-die M5 Max packages using an updated generation of its UltraFusion interconnect. The resulting four-die package presents to the operating system as a single unified system-on-chip.
Apple reports that the UltraFusion interconnect delivers more than 4.4TB/s of inter-die bandwidth, representing a sixfold increase in interconnect density compared to prior iterations. This low-latency interconnect allows the processor to address up to 512GB of shared unified memory at 1.2TB/s of bandwidth without partitioning memory across distinct CPU and GPU address spaces.

At maximum configuration, the M5 Ultra contains a 36-core CPU and an 80-core GPU. Apple states that every GPU core incorporates a dedicated Neural Accelerator alongside a 32-core Neural Engine and upgraded media engines with hardware-accelerated AV1 decoding. The company claims the 512GB memory pool provides sufficient headroom to run large language models exceeding several hundred billion parameters entirely on local hardware.
M6 Architecture and 2nm Transition
At the lower end of the product stack, the M6 serves as Apple's first production chip built on a 2nm node. The base configuration features a 12-core CPU, a 12-core GPU with per-core Neural Accelerators, and a Dual 16-core Neural Engine.
The M6 supports up to 32GB of unified memory delivering 170GB/s of bandwidth, a 10 percent increase over the base M5. Apple's internal testing indicates a 1.2x improvement in multithreaded CPU performance over the M5 and a near 30 percent boost in peak GPU compute for AI operations. For workloads demanding higher memory throughput in a small form factor, Apple also made the Mac mini available with an M5 Pro configuration delivering up to 307GB/s of bandwidth and 64GB of unified memory starting at $1,699.
Local Inference Economics
The release comes as hardware makers increasingly orient desktop silicon toward local model execution and developer workstations. High unified memory capacities reduce reliance on remote API endpoints and quantized weights for developer inference, local agent loops, and fine-tuning.
Apple named local inference tools including Ollama, LM Studio Bionic, and Draw Things among the target environments optimized for the hardware. While the company claims up to 4.5x faster peak GPU compute for AI compared to the M3 Ultra, all performance metrics represent internal vendor benchmarks measured on preproduction hardware.



