What Actually Matters in an AI PC: RAM Before TOPS, and Other Rules
Every laptop sticker this year has a number with “TOPS” after it. Microsoft’s own bar for a Copilot+ PC is an NPU that can hit more than 40 trillion operations per second, and chipmakers have treated that number like horsepower ever since.
It mostly doesn’t decide what you can do. If you want to run AI on the machine itself, rather than renting it from a data center through a chat window, the spec that matters is memory: how much you have and how fast the chip can read it. Here’s what each part of an “AI PC” does, and where your money should go.
What an NPU is, and what TOPS counts
A neural processing unit is a block on the processor built for the repetitive matrix math that neural networks run on. Microsoft describes it as a chip for “AI-intensive processes like real-time translations and image generation.” Its advantage is efficiency. It can run a background blur on your webcam or transcribe a meeting all day without draining the battery the way the GPU would.
TOPS, in Microsoft’s plain definition, “measures trillions of operations a system can perform per second.” That’s a peak figure, usually measured on low-precision integer math, and vendors aren’t always careful to say which part of the chip they’re counting. AMD’s Ryzen AI Max+ 395, for example, is listed at up to 50 TOPS for the NPU and up to 126 TOPS for the whole chip. Both numbers are real. Only one is relevant to Copilot+ certification.
A bigger TOPS figure gets you a faster small model. It doesn’t let you run a bigger one.
Which Windows features actually need the NPU
The Copilot+ label currently requires a 40+ TOPS NPU, 16GB of RAM and 256GB of storage. In return you get a specific set of Windows features: Recall, Click to Do, improved Windows Search, live caption translation, Cocreator and Restyle for images, and Windows Studio Effects for your camera. Microsoft adds that experiences “vary by device and market.”
Plenty of AI isn’t on that list. ChatGPT, Claude, Gemini and the Copilot chatbot all run on someone else’s servers, so a 2019 laptop with a decent browser handles them fine. If that’s your whole use of AI, the NPU will sit idle and you’ve paid for it anyway.
Want to see whether yours is doing anything? On machines that have one, Task Manager shows NPU usage on the Performance tab, next to CPU, GPU and memory.
Why RAM decides which models you can run
Running a large language model locally means holding its weights in memory. All of them, the whole time. Popular local tools such as Ollama and LM Studio mostly run those models on the GPU or CPU, not the NPU. So the questions that matter are whether the model fits and how quickly the chip can stream it.
Download sizes from Ollama’s model library make this concrete. These are 4-bit compressed versions, which is how most people run models at home:
- 16GB: Qwen3 8B is a 5.2GB download and runs comfortably. OpenAI’s gpt-oss-20b, a 14GB download, “can run on systems with as little as 16GB memory”, but on a 16GB Windows laptop that leaves very little room for your browser. Fitting and fitting comfortably are different things.
- 32GB: The useful middle tier. Qwen3 30B and 32B (19GB and 20GB) fit with room left over for the operating system and a long conversation.
- 128GB: This is where 70B-class models live. Llama 3.3 70B is a 43GB download, and gpt-oss-120b is 65GB. Qwen3’s 235B flagship, at 142GB, still won’t fit.
Memory bandwidth determines speed. Generating each word requires reading roughly the whole active model out of memory, so a rough ceiling on output speed is bandwidth divided by model size. That rule of thumb is ours, and mixture-of-experts models, which read only part of their weights per token, beat it. Apple publishes its number: the base M5 has 153GB/s of unified memory bandwidth and tops out at 32GB. By that arithmetic, a 20GB dense model would top out around seven or eight tokens (word fragments) a second. You’d be able to use it, but you’d be waiting on it.
Also check whether the RAM is soldered. Most thin AI laptops use soldered LPDDR5x, so the amount you choose at checkout is permanent. Spend here before you spend on a higher-TOPS chip.
How the chip families differ
Apple M-series
Apple Silicon popularized unified memory: one pool shared by the CPU, GPU and Neural Engine, so the GPU can use most of your RAM for a model. With M5, Apple added a Neural Accelerator to each GPU core and claims over four times M4’s peak GPU compute for AI. “With the introduction of Neural Accelerators in the GPU, M5 delivers a huge boost to AI workloads,” Apple hardware chief Johny Srouji said in the launch release. If you’re shopping Apple for local AI, how much memory you configure matters more than which chip tier you pick.
Nvidia RTX Spark (N1X)
The new entrant. Nvidia’s Arm-based chip, which it said launches in October, pairs a Grace CPU with a Blackwell GPU and comes in two versions, according to Tom’s Hardware. One has 20 CPU cores, 6,144 CUDA cores and 24GB to 128GB of unified memory, for laptops and mini PCs. The other has 18 cores, 5,120 CUDA cores and 24GB to 32GB, for laptops only. Lenovo, Acer, Asus, Dell, MSI, HP and Microsoft have signed on. Nvidia hasn’t announced pricing, TOPS or bandwidth figures, so wait for independent reviews before you buy. Our read on why the 128GB version matters: it puts a 70B-class model in reach of an Nvidia GPU, and CUDA is what most local AI software is built for first.
Qualcomm Snapdragon X2
Qualcomm had Windows on Arm to itself until N1X arrived. Its X2 Plus now powers Microsoft’s 12-inch Surface Pro and 13-inch Surface Laptop, and the X2 line is getting Debian support by the end of 2026, according to The Next Web’s Snapdragon Summit coverage. Snapdragon machines have cleared the Copilot+ NPU bar since the first Snapdragon X Elite wave. Their strength is battery life. For local LLMs, check memory size first.
Intel and AMD
On x86, Intel’s Core Ultra and AMD’s Ryzen AI chips have qualified for Copilot+ since the Core Ultra 200V and Ryzen AI 300 series, and they run every x86 Windows app natively. AMD’s Strix Halo, sold as Ryzen AI Max, takes the Apple approach: up to 128GB of 256-bit LPDDR5x shared with Radeon graphics with 40 graphics cores. It’s the x86 option for people who want big models without buying a desktop graphics card.
A note for everyone on Arm, whether Qualcomm or Nvidia: check that your must-have apps, VPN, printer drivers and any games with anti-cheat run there before you commit.
Who should skip the AI premium
If you use AI only through a browser or chat app, buy for the screen, keyboard, battery and 16GB of RAM, and ignore TOPS. Any recent Copilot+ machine covers the Windows features, and none of them need the most expensive tier.
If you want to run models locally, rank your options by memory: 32GB as a floor, 64GB or more if you want 70B-class models, and bandwidth as the tiebreaker. Then think about the GPU, with NPU TOPS last.
Microsoft technical fellow Steven Bathiche, in the company’s NPU explainer, called agents “the north star” of Windows. If that’s where Windows is headed, those agents will need room to hold their models, and on most of these laptops the RAM you pick at checkout is what you’re stuck with.
Sources
James Whitfield covers hardware for prompt/power: chips, semiconductors, laptops, components and the benchmarks behind the launch-day claims. He thinks the most important number on any spec sheet is usually the one in the footnote.
Latest from prompt/power
- Thomson Reuters Won the First AI Training Appeal. Footnote 7 Is the CatchOct 7
- The Family Safe Word: How to Beat AI Voice-Clone Emergency ScamsOct 7
- Reflection AI’s Beam, Explained: The 501B Open-Weight Model Aimed at ChinaOct 7
- Your SSN or SIN Leaked in a Breach? Do These 8 ThingsOct 7
- Apple’s Oct. 13 Event Rumour, Plus iPhone Duo Pre-Order Dates for CanadaOct 7
Leave a Reply