AI, decoded

Why is it so hard to move AI workloads off NVIDIA?

Because much of the switching cost lives in the software, not the silicon. Memory and math formats are catchable on a roadmap; what is hard to match is an ecosystem where new frontier models run on day one, which is what AMD has spent the last few years working to match and what Google sidestepped by owning its stack from the compiler up.

· Updated · Chain of Thought

Level 5: Production agents · 5.4 Infrastructure choices

AI HardwareAI InfrastructureOpen Source AI

An iceberg with a small tip labeled specs above the waterline and a much larger mass labeled software below, beside the line: the moat is software.

1. The moat is day-one model support

AMD’s Anush Elangovan describes what the alternative platform used to feel like: ROCm support was a “port to platform.” Someone would launch a model, and then someone else would go in and make it work. The change he points to is that the recent frontier releases (DeepSeek, Llama, Qwen) run natively from day zero, as well as they do on the competing platform. That confidence is what a buyer is paying NVIDIA for: a model announced this morning runs on their cluster this afternoon.

Update, October 3, 2026: At 9:13 in episode 72, Anush Elangovan describes AMD’s expanded testing across frameworks and hardware generations. Elangovan says AMD can block upstream pull requests and test them on AMD hardware to help prevent regressions from changes tested only on NVIDIA. For a team evaluating the day-one support discussed above, upstream testing is an ongoing compatibility check to ask about.

2. The specs are the catchable part

Hardware differences show up on a roadmap and get closed on a roadmap. AMD’s CDNA4 generation added FP6 as a stepping stone toward FP4, so a team can move parts of a model down in precision in stages. The 350 series ships 288 GB of memory per GPU, which on AMD’s own framing puts a 500-billion-parameter model on a single card. NVIDIA’s Blackwell Ultra now matches that 288 GB, and AMD’s newer MI455X, launched in July 2026, raises it to 432 GB. Memory capacity is where AMD keeps pressing its case. None of that, by itself, moves a workload.

3. The other exit is owning the whole stack

Google took the path almost nobody else can afford. TPUs sit under XLA as the target compiler, and JAX sits on top of XLA, the framework Google DeepMind uses in research and in production. Paige Bailey’s point is that JAX is loved partly because it uses TPUs so well, and scales not just across many TPUs but across multiple data centers, which is what training Gemini requires. Vertical integration from silicon to framework is why the optimizations available to Google are not available to a team renting someone else’s stack.

4. Open is the counter-strategy, and it is a software strategy

AMD’s bet is that an open ecosystem out-innovates a closed one, and the tell is where they applied it: the software review process is moving outside the company, so any developer working in an area reviews the changes in it. Partners build switches, NICs, and interconnects around them rather than at the periphery of a proprietary fabric. The pitch is about the stack rather than the accelerator: you should not have to buy the whole thing from the company that sells you the chip.

Why it matters

If you are evaluating accelerators on FLOPS and memory, you are comparing the part that converges. Ask instead how long after a model release your team can run it, who fixes the kernel when it breaks, and what you would have to rewrite to leave. That answer is the switching cost.

Why moving AI workloads off NVIDIA is hard: the specs are the tip, the software is the iceberg An iceberg. Above the waterline, the part a roadmap can catch: math formats, such as AMD's FP6 as a stepping stone to FP4, and memory, 288 GB per GPU on AMD's 350 series, which on AMD's framing fits a 500-billion-parameter model on one card. Below the waterline, the much larger part that is the switching cost: whether a new model runs on day zero or has to be ported to the platform, the compiler and framework, who fixes the kernel when it breaks, and what you would have to rewrite to leave. On the right, two ways out. Own the whole stack, as Google does: TPUs under the XLA compiler under the JAX framework, which scales across multiple data centers to train Gemini. Or open it, as AMD is betting: source repositories that are all external, outside developers contributing and reviewing changes that AMD merges, and partners building switches, NICs and interconnects. The specs are the tip. The software is the iceberg. Why AI workloads are hard to move off NVIDIA, from AMD's Anush Elangovan and Google DeepMind's Paige Bailey. WATERLINE CATCHABLE ON A ROADMAP FP6 as a step to FP4 288 GB per GPU (AMD 350 series) memory capacity, bandwidth: where AMD claims a lead THE SWITCHING COST Does a new model run on day zero, or get ported? the compiler and the framework who fixes the kernel when it breaks what you'd rewrite to leave TWO WAYS OUT Own the whole stack Google's path, which few can afford JAX XLA compiler TPUs scales across data centers to train Gemini Open the stack AMD's bet, a software strategy chip outside devs NICs switches repos all external; review moving outside
Specs converge; the software layer is the switching cost. The AMD details are Anush Elangovan’s, episode 28; the TPU, XLA and JAX stack is Paige Bailey’s, episode 47. Download the image

Hear it from the guest

“The internal source repositories and external source repositories are exactly the same. It is all external, right? And so what that means is, as an external developer, you can contribute any code changes you want and we actually take that seriously and merge it in.”
“Why people love JAX so much other than JAX being just like a friendly NumPy like interface to do work is because it can use TPUs so optimally.”
“We can block PRs that go in and are always tested on AMD”

Quotes lightly edited to remove filler words.

From the conversation

This explainer is drawn from these episodes — each carries its full transcript.

Concepts in this explainer

TPU (Tensor Processing Unit)Frontier ModelModel ParametersPrecision and Recall