Architectural overview, decision framework, and evidence contracts for shared-memory inference platforms and multi-device stacks -- no invented numbers.
Illustrative launch-preview data. Live price tracking is not connected yet.
Back to gpu-memory.com
Illustrative price signals / launch preview
Explore
Apple Silicon, AMD Ryzen AI Max, Intel, and multi-device decisions
Planned capability
Unified memory means the GPU and CPU share the same physical RAM, so a 128GB Mac Studio can host models that otherwise need multi-GPU setups. The two mature inference runtimes are ollama and MLX. Spec-level facts are sourced from Apple developer documentation; per-tier tokens/sec values will be added only after measurement against the evidence contract methodology.
Planned capability
AMD's shared-memory APU ships up to 128GB of unified memory on a single socket. The inference path is llama.cpp with HIP backend on Windows and ROCm on Linux. Spec facts are sourced from AMD's product brief and the llama.cpp HIP/ROCm build documentation. The evidence contract names the specific Ryzen AI Max SKUs queued for the first measurement batch.
Planned capability
Intel's current consumer shared-memory story is bandwidth-limited compared to Apple and AMD. The platform runs smaller quantized models well but is not yet a primary path for 70B-class local inference. Sources: Intel ARK product pages and the official product briefs. Re-evaluation cadence: each new Intel consumer platform generation. The evidence contract records the bandwidth and memory tiers this site plans to track.
Planned capability
A single higher-memory device almost always beats two networked lower-memory devices at the same total memory, because interconnect is the bottleneck. The framework covers when to stack (NVLink dual NVIDIA), when to network (Exo, llama.cpp --rpc), and when to wait. Per-config tokens/sec values will be added only after measurement, not now. Architecture-level trade-offs are sourced from the official project documentation cited in the evidence contract.
Planned capability
If your target model fits in one device, buy one device. If it does not, stacking is acceptable but tokens/sec will be a fraction of a hypothetical single device at that memory size. This is the llama.cpp documentation position, restated for the site. The full source-backlinked reasoning lives in docs/memorysites/gpu-memory-strategy.md under Multi-Device Architectures.
Planned capability
This site does not yet publish tokens/sec, price-per-GB, or power-draw numbers for the platforms covered here; those values will be added only after measurement against the public tooling. The full evidence-contract list (vendor, tooling, model, quantization, hardware skew, methodology) is in docs/memorysites/dom-3hi.4.4-evidence-contracts.md. Readers can verify any future number against that contract before treating it as authoritative.
Planned capability
This page covers shared-memory platforms (Apple Silicon, AMD Ryzen AI Max, Intel) and multi-device architectures for inference. Discrete-GPU compatibility (NVIDIA RTX, datacenter cards) is documented on the parent gpu-memory.com home and in the Model Compatibility Matrix. Boundaries are deliberate: stacking a Mac and an NVIDIA box is a llama.cpp --rpc or Exo configuration, not a unified-memory system.
Planned capability
Apple silicon and AMD Strix Halo move on the same generation cadence as their vendor releases. Intel shared-memory is the slowest-moving story in the set; the editorial policy is to re-evaluate at every new consumer platform generation, not quarterly. Source-backlinks to the vendor newsroom on each page revision are recorded in the evidence contract.
Planned capability
Specific tokens/sec at any model size, specific price-per-GB, specific power draw, specific watt-hours-per-million-tokens -- none of these are claimed on this page. The page is the architectural and decision framework; the measured numbers arrive in a later phase, each tied to the methodology in the evidence contract. This is the editorial policy and the reason the parent site carries the Illustrative launch-preview data disclosure.
Planned capability
Start with the Decision Framework card if you are choosing hardware. Read the Apple Silicon and AMD cards for per-platform architecture (specs only, with the public source). Read the Multi-Device card before considering a stack. Read the Evidence Contract last, to know which numbers on the site are measured and which are pending measurement. Nothing on this page substitutes for the actual evidence contract when comparing hardware.
Reading the deep-dive is a starting point