AI, decoded

Can switching to an open model actually cut your AI costs?

On the narrow, high-volume tasks, yes. Intercom was spending roughly $250,000 a month, by its Chief AI Officer's own estimate, running one summarization job through a closed-model API and replaced it with a fine-tuned 14-billion-parameter Qwen model. The decision is made per task rather than per platform: the hardest prompt in the same product still runs on a frontier model.

· Chain of Thought

Level 3: Leading with AI · 3.2 Build, buy or rent

Open Source AIEnterprise AIAI Infrastructure

A row of pipeline blocks where only one, the query-rewrite step, has been swapped for an open model, beside the words: Move one task, not the stack.

1. Pick the task, not the stack

Intercom’s line item, which Fergal Reid estimates at roughly $250,000 a month, was one job: query canonicalization, a summarization task running at volume. That is the profile where a smaller fine-tuned model wins: narrow, repetitive, well-specified, and enormous in aggregate. The hardest prompt inside the same product, the one that answers the end user’s question, still runs on a frontier model in production. Reid, Intercom’s Chief AI Officer, moved one task off the API and left the hardest prompt where it was.

2. The open weights had to get good enough first

Reid dates the shift to around DeepSeek’s release, when open-weight models got close enough to closed ones to be viable for core tasks. That threshold moved, which means the answer to this question carries a date. The Qwen 3 models were the ones that back-tested well for Intercom’s workload, and Reid is explicit that a large training expense was absorbed by whoever released the weights.

Update, October 2026: Reid wrote on LinkedIn: “The important piece is a post training pipeline that works well, not any one base model. We have since switched to Nvidia and Google base models.” The base model turned out to be the dated part of this answer; the post-training pipeline is what carried over.

3. The savings are partly a headcount trade

Intercom’s AI team went from 10 people to 55 in about two and a half years, with plans to double again. Their standard training setup is a node of H200s, and the approach moved from LoRA and other parameter-efficient methods toward distributed full fine-tuning and reinforcement learning. If you count only the inference bill you will get the decision wrong: you are buying an in-house post-training capability and the people who operate it.

4. Do not let a benchmark make the call

Reid’s warning is Goodhart’s law in its native habitat. It is easy for a researcher optimizing a benchmark to teach a model the benchmark’s task and then watch it fail to generalize. Intercom’s discipline is their own back tests plus A/B testing in production, run by people hired as scientists rather than only as engineers. A public leaderboard is not evidence about your workload.

Why it matters

“Open source is cheaper” is not a strategy, and neither is “frontier models are better.” The teams saving real money are profiling their own traffic, finding the one or two tasks that are high-volume and narrow, and moving those. Everything else stays where it works.

Intercom moved one task to an open model, not the whole stack Two rows show the same pipeline inside Intercom's Fin agent, from episode 49. A user's question goes through a step that summarizes and canonicalizes the query, then retrieval, then the hardest prompt, which answers the end user's question. Before: the query step ran on GPT-4.1 over an API, a bill Fergal Reid estimates at roughly 250,000 dollars a month, and the answer step ran on Claude Sonnet. After: the query step moved to a post-trained 14-billion-parameter Qwen 3 model, which Reid says saved almost all of that spend; the answer step still runs on a frontier model. A threshold marker notes that open weights only got close enough around DeepSeek's release. A receipt lists the post-training program behind the switch: an AI team that grew from 10 to 55 people in about two and a half years, training on a node of H200s, and their own back tests plus A/B tests in production instead of a public leaderboard. Move one task, not the stack Inside Intercom’s Fin agent, episode 49: one narrow, high-volume step changed models. The hardest prompt did not. QUESTION REWRITE THE QUERY RETRIEVAL ANSWER IT (HARDEST PROMPT) BEFORE “How do I…” GPT-4.1 via API ~$250K a month* retrieval Claude Sonnet a frontier model the model swap unchanged AFTER “How do I…” Qwen 3 14B, post-trained saved almost all of it retrieval Still a frontier model for the hardest prompt FIRST, A THRESHOLD Around DeepSeek, open weights got close enough for some core tasks. THE REST OF THE RECEIPT · THE PROGRAM BEHIND THE SWAP A team to run it Intercom’s AI group: 10 → 55 people in ~2.5 years Hardware to train on A node of H200s; full fine-tuning plus RL Proof on your own traffic Their own back tests and A/B tests, not a leaderboard *Fergal Reid’s own estimate: “I don’t remember exactly, but it could have been like a quarter million dollars a month.”
Move one narrow, high-volume task, not the stack. The roughly $250K a month is Fergal Reid’s own estimate from episode 49; the team size comes from the episode’s introduction. Download the image

Hear it from the guest

“And we were able to go and take that model and train that model, post train that model to do that summarization task. and able to replace that and saved almost all of those … several hundred thousand a month on inference just by doing that one thing.”
“It's so easy for a researcher or a scientist trying to optimize a benchmark to get good at teaching the model how to do the kind of task that's in the benchmark and then suffer generalization failure.”

Quotes lightly edited to remove filler words.

From the conversation

This explainer is drawn from these episodes — each carries its full transcript.

Concepts in this explainer

Small Language ModelFine-TuningOpen WeightsAI BenchmarkFrontier ModelFoundation ModelPost-trainingInference