Tuning GPU Performance with AI Agents | AMD’s Anush Elangovan on ROCm 10
Key takeaways
- ROCm 10 moves the first interaction with AMD’s software stack into an agentic workflow. A developer can ask Claude Code or Codex to install ROCm, then ask it to serve a model with vLLM; the agent reads the documentation, pulls in the relevant skills, resolves dependencies, and sets up the service.
- Hyperloom treats GPU performance tuning as an iterative search problem. It profiles a model to locate kernel or fusion hotspots, then GEAK tests kernel changes while preserving numerical accuracy. Elangovan says AMD used this approach to optimize 14,000 models in one pass when enough compute was available.
- Optimization work can continue when the hardware would otherwise sit idle. Elangovan describes Hyperloom running performance experiments in parallel, including while the user sleeps, until the workload reaches a stronger result for its specific deployment.
- AMD expanded its continuous integration footprint by roughly an order of magnitude across multiple hardware generations and frameworks including PyTorch, vLLM, and SGLang. Testing upstream pull requests on AMD hardware is meant to catch regressions that would be missed if code were tested only on NVIDIA.
- Agentic workloads change the capacity plan. Human interaction produces intermittent requests, while agents sustain sequences of file reads, searches, and tool calls. Elangovan argues that systems built around human usage may need to plan for 100 times the activity engineers or customers previously generated.
- The next large gains may come from applying today’s models to the last mile of a specific industry. Elangovan expects base models to improve, but argues that clear reward signals helped coding advance first and that current models have barely been applied to open-ended fields such as farming, retail, and transportation.
Concepts in this episode
AI terms discussed here — each links to a plain-language definition.
AI AgentInferenceAccuracyArtificial General Intelligence (AGI)Frontier ModelTokenizationFoundation ModelAI AlignmentModel Context Protocol (MCP)Physical AI
Chapters
- 0:25A decade of ROCm, now agent native
- 3:03What agentic ROCm looks like in practice
- 6:17Installing ROCm then versus now
- 9:14An order of magnitude more CI across every framework
- 10:44Anush’s workflow: agents and deployment
- 12:21Speed is the moat
- 15:00Success is a stranger who cannot spell ROCm serving an LLM
- 17:09Keeping agent skills from going stale
- 21:38Co-designing kernels with the frontier labs
- 24:06Hyperloom, GEAK, and 14,000 models in one pass
- 26:45Managing autonomous agents: control and liability
- 32:21Security at the speed of agent swarms
- 36:03ROCm performance gains on the same hardware
- 37:36Where enterprises hit walls in production
- 40:24Why coding was the right reward function for AI
- 44:42Which industries get the next software scale unlock
- 47:02The last mile of AI
- 50:31Closing thoughts
Show notes
Watch the full conversation on YouTube
AI agents can keep tuning GPU workloads after you step away from the keyboard. Anush Elangovan, Corporate VP of AI Software at AMD, returns to Chain of Thought to explain how that works with Hyperloom and ROCm 10.
Anush and Conor Bronsdon trace the process from installing ROCm through Claude Code or Codex to profiling workloads, finding slow kernels, and testing optimizations while preserving numerical accuracy. Anush shares a Hyperloom run spanning 14,000 models and explains why clear goals and feedback matter when agents are doing the tuning.
They also explore what comes next for engineers: keeping skills and frameworks reliable, managing the security and accountability of autonomous agents, and applying AI to the last mile of useful software.
We cover:
- How agents help install ROCm and serve models through natural language
- How Hyperloom profiles workloads and uses LLMs to explore optimizations
- The role of GEAK in tuning kernels while preserving numerical accuracy
- Anush’s account of optimizing 14,000 models in one pass
- How software improvements get more performance from existing GPUs
- Keeping agent skills current and testing across AI frameworks
- Security, accountability, and the next bottlenecks in agent-driven development
Connect with Anush Elangovan:
- LinkedIn: https://www.linkedin.com/in/anushelangovan/
- Twitter/X: https://x.com/AnushElangovan
- ROCm.AI: https://rocm.ai
- AMD AI blog: https://www.amd.com/en/blogs/by-author/anush-elangovan.html
- AMD AI Developer Program: https://www.amd.com/en/developer/ai-dev-program.html
Connect with Chain of Thought host Conor Bronsdon:
- Newsletter: https://newsletter.chainofthought.show/
- Twitter/X: https://x.com/ConorBronsdon
- LinkedIn: https://www.linkedin.com/in/conorbronsdon/
- YouTube: https://www.youtube.com/@ConorBronsdon
More episodes: https://chainofthought.show
Transcript
138 segmentsConor Bronsdon 0:24 Last week, AMD shipped its answer to the friction between silicon and surfing tokens. An agentic stack where coding assistants know the hardware, and autonomous systems rewrite your kernels while you sleep. With us today is Anoush Elengovan, Corporate VP of AI Software at AMD, back for his third time on the show, fresh off shipping Rokkam 10 and taking Rokkam AI to general availability. 10 years into The Stag's life. Welcome to Chain of Thought. I'm your host, Conor Bronsdon. Anoush, great to see you. Congratulations on the big week. How has the kickoff to GA for Rockham 10 been?
Anush Elangovan 0:58 Great. Thanks for having me again, Conor. It's the third time.
Conor Bronsdon 1:05 [OVERLAP] It's
Anush Elangovan 1:05 [OVERLAP] Thank
Conor Bronsdon 1:05 [OVERLAP] my pleasure. You're one of the favorites of the entire audience to have on the show, so it's great to have you back.
Anush Elangovan 1:10 you. Thank you for having me again. And it's been great. It's been a celebration of 10 years of RockM. And it's a milestone, both in terms of a decade of building and shipping open software. But also it is a milestone because this is the first time that we are actually starting to bring in agentic AI into the fabric of the entire stack. So the entire stack now is AI enabled all the way from the skills at the top, performance optimization and attainment with Hyperloom, and then the core libraries and the runtime, all of which are agentically enabled. And then we also have ease of use things with skills and CLI and TUI that's all very easy to integrate into your existing cloud code or codecs or vice versa to actually just use the Rokum TUI that ships with Rokum 10.
Conor Bronsdon 2:13 I'm excited to dig into details of this because I think a lot of it builds on the conversation you and I had back in March. We talked about your personal workflows and how you were starting to develop and the pace you were going at. And we're now seeing that across the industry where we have unlocked so much more throughput through agentic infrastructure. But before we go too deep, I would be remiss if I didn't say a quick word about our presenting sponsors for Season 4 of Chain of Thought. Sphix, which are delivering billions of reliable webhooks for startups in the Fortune 500. Walrus Memory, giving AI agents portable, verifiable memory that carries context across apps, sessions, and other agents. And obviously, this whole conversation is going to be highly deep into AI agents. So you said, Anoush, that RockM10 is all about enabling agents with RockM. What does that look like in practice?
Anush Elangovan 3:03 Yeah, so in practice, you know, typically everything like the first touch point for RockM used to be someone has to go to the RockM document site, you know, docs and then figure out how to install like type curl this or, you know, sudo app something. The entry point into Rokum itself now is agentic. So you can actually just have, if you have your cloud code or your codex, you can just say install Rokum, right? And so the skills are already integrated. And then there are other entry points similar to that that allow you to use your own, there's a Rokum TUI. So you can just, you know, curl a little binary called Rokum. And that will bootstrap the entire installation process and ensuring you have the right dependencies and making sure it's, you know, delightful to use because we want RockM, the entry point for RockM to be like frictionless and almost invisible. And then you get to the point where, once you deploy it, you want to say something like, use VLLM to serve Quen 3, 3.8, whatever it is, right? And it goes and finds out where's Quen 3.8 waits, how does it bring it in? How does it integrate into the... into your system, and then it reads all the documentation, it pulls in the skills, it manifests everything that needs to happen for you to have VLLM serving as soon as possible, right? And then it also can do performance optimizations when you're sleeping. So you can run a tool called Hyperloom that we have that allows you to just optimize on the fly. And so you just continue to do that. until it gets to a point where it's really performant. And this can work in parallel, especially when you're not using the system, we can continue to optimize. So it's like agentic end-to-end.
Conor Bronsdon 5:01 Yeah, I think this is just the future of software development, right? We saw Dario say, what, 18 months, two years ago, 90% or more of code is going to be written by AI, and we're there. Not everyone at this point, but I think anyone who is looking to accelerate in the modern era is not writing code by hand. We're seeing luminaries who are not writing code by hand. And we're all aware that the engineering is now about directing agents, and we have a lot of work to do to set them up for success. But it's through the work of teams like what A&D is doing and many others to enable agentic experience that we can actually let tools take this next step. And just as we were talking, I had my codex instance. I said, hey, yeah, set up Rockham for me. And it's just done so. And it's I asked it to highlight a few of the new innovations that are happened here in Rockham 10. with some meaningful known issues that have been improved on. So improvements around crashes around VLM or Comfy UI. We're seeing cleaner foundations. Obviously, you mentioned Hyperloom. Let's talk a bit about some of those innovations. What does it look like to develop with Rockham today compared to where we were a year or two ago, or maybe 10 when we first got started?
Anush Elangovan 6:16 Yeah, so I think 10 years ago, it probably is like, you know, very
Anush Elangovan 6:26 nostalgic, but it's probably, it's changed quite a bit, right? Even two and a half years ago, or three years ago, when I joined AMD, it was a,
Anush Elangovan 6:39 I'd say, a reasonably rough experience to install especially the kernel mode driver and all that. Now we've made it completely seamless, including pip install, virtual and VM installs. So the ability to consume RockM has just, just baseline has improved. But then what I'm really excited about is, you know, like we talked about, the UX of trying to use RockM is definitely, you know, improved. But then the performance part is where I'm, you know, super excited about because it allows you to run the, to attain performance across a broad swath of models without having to try to optimize for everything beforehand and ship libraries that are 100% optimized for every combination. This just means that at runtime, we ensure that the models can be super optimized for the workload, for the way that you're
Anush Elangovan 7:47 deploying your workload. So I think it's exciting to see that, but also it
Anush Elangovan 7:57 allows us to focus on the next layer of the stack. Things that require deep understanding or human understanding. so that you don't have to spend time doing rote implementations, right? Like you're not just continuously doing something that is now automated by agents. And so that's exciting. So it's a second order effect where it's just, RockM is not only free, it also sets you free to go do things that you have to do.
Conor Bronsdon 8:28 Oh, I like that. Framing it around software freedom. And I think a big part of what has freed people up to take on so much more is the ability to, yes, assign work to their agents. I don't have to do this setup now for Occam. I can let Codex do it, or I can let my Cloud Code instance do it. I don't necessarily have to go read the doc site. I can get a synthesis that's delivered to me that walks me through the important parts of what needs to happen next. And part of this is enabled by new tooling. You mentioned big improvements to the CLI for Rokkam. You mentioned Hyperloom. Can you talk a bit about some of the details of what is new now in Rokkam 10 and why Rokkam has been rebuilt with this direction?
Anush Elangovan 9:13 Yeah, yeah, yeah. I think the CI, CD part definitely has been a huge, huge, huge focus for us over the last year. We have like an order of magnitude more CI that we have stood up across like multiple generations on every framework that we can possibly like, you know, track down. Of course, there are fringe, you know, cases where we still don't know and we're willing to like stand them up as soon as we find out. what they'd like. But most of the frameworks like PyTorch, VLLM, SG-Line, all of those are now currently like 100%, you know, like we have 100% enablement. And then there's obviously, you know, the benefits of having that kind of CRA like we can block PRs that go in and are always tested on AMD, which gives us the ability to ensure that the upstream projects do not get into a stage where they just regress AMD and the code was just tested on Nvidia. And I think that's really powerful because it's a big flywheel that keeps making the baseline get better.
Conor Bronsdon 10:30 How has this changed your own development flow? Because obviously, like we alluded to, you have vastly increased your own coding throughput. How are you adjusting that now with the new RockM10?
Anush Elangovan 10:43 My, you know, my agentic, I'll say flags have largely remained the same. That is dangerously skip permissions, right? And so I've been trusting the agents to do the right thing quite a bit. But, you know, it definitely builds on top of that, right? You start with that and then, you know, for me to now deploy Rokam on a system, I literally just say deploy Rokam on the system, right? And the skills are there, the ability to like understand and deploy Rokam is straightforward. And then even deploying some services on top of it is also all natural language, right? So it's a fully natural language, interface that we are evolving towards. And Rokum.ai is just like the first one that's getting started with that.
Conor Bronsdon 11:44 And I'll say for anyone who wants a deeper look at Anoush's approach and how he's building, our conversation from a few months back, Software is Just Tokens Now, is highly recommended. Anoush, as we think about this new era where we are all coding with agents, the way we interact with them has changed. What would be your advice to developers who are just making this transition and have been using AMD for a long time, have been using Rockham for quite a while, but are starting to try to become agent native? Maybe they're not as far along in the curve as some of us are. What would you advise them?
Anush Elangovan 12:20 [OVERLAP] I think the advice largely remains the same, which is prepare for change, right? And the faster you move, and like I say, speed is the moat, right? You gotta be agile, because I cannot predict what is gonna happen in a week or two. Things may change, right? The next new frontier model comes through, the next year harness comes through. And you're like, oh, wow, okay, now I don't need to do that thing too. Great, right? So advice would be like, be prepared for change, right? That is the only constant. And be prepared for change that's like far beyond what you're expecting, right? Like nobody would have expected like agentic AI coding to be like this. what do you say, transformative, right? It was just like, yeah, it's like an autocomplete thing, right? Even about
Conor Bronsdon 13:21 [OVERLAP] I don't
Anush Elangovan 13:21 [OVERLAP] this
Conor Bronsdon 13:21 [OVERLAP] know if I
Anush Elangovan 13:21 [OVERLAP] year
Conor Bronsdon 13:21 [OVERLAP] agree
Anush Elangovan 13:22 [OVERLAP] last,
Conor Bronsdon 13:22 [OVERLAP] with that. I don't know.
Anush Elangovan 13:23 [OVERLAP] this year last time, this year, this time last year, do you think agentic coding, you would have thought of agentic coding like it is like this?
Conor Bronsdon 13:31 [OVERLAP] I mean, okay. I don't know if I would have necessarily said by this time next year, this is where it's going, but I think you could see where it was going. I definitely do think by December, you could see with
Anush Elangovan 13:41 [OVERLAP] Oh yeah,
Conor Bronsdon 13:42 [OVERLAP] the
Anush Elangovan 13:42 [OVERLAP] December,
Conor Bronsdon 13:42 [OVERLAP] next open,
Anush Elangovan 13:43 [OVERLAP] Christmas
Conor Bronsdon 13:43 [OVERLAP] you saw
Anush Elangovan 13:43 [OVERLAP] time
Conor Bronsdon 13:43 [OVERLAP] this jump.
Anush Elangovan 13:43 [OVERLAP] was the light bulb, right? But for me, last year, this time, would have been like, wow, chat GPT, write me a Python file that does this and it wrote that thing and then I copy it and I can paste it and be
Conor Bronsdon 14:00 [OVERLAP] Hmm.
Anush Elangovan 14:00 [OVERLAP] like, great,
Conor Bronsdon 14:00 [OVERLAP] I hear
Anush Elangovan 14:00 [OVERLAP] that
Conor Bronsdon 14:00 [OVERLAP] what
Anush Elangovan 14:00 [OVERLAP] works,
Conor Bronsdon 14:00 [OVERLAP] you're saying.
Anush Elangovan 14:02 [OVERLAP] right?
Conor Bronsdon 14:02 [OVERLAP] Yeah.
Anush Elangovan 14:02 [OVERLAP] right it it wasn't like i i hadn't even processed like oh chat gpt or codex will become an app and cloud code is back to the tui which i love but last year this time i was still like oh man i missed that visual studio visual
Conor Bronsdon 14:18 [OVERLAP] There's
Anush Elangovan 14:18 [OVERLAP] id
Conor Bronsdon 14:18 [OVERLAP] more friction in the workflow.
Anush Elangovan 14:19 [OVERLAP] bucket exactly
Conor Bronsdon 14:21 [OVERLAP] Yeah. Okay. I hate what you're saying. All right. So we're, we're in this new era now, obviously the things have changed. Um, you know, we're operating with CLIs for agents. We are going API first, uh, more than ever on many sites. Uh, we're building skills. So there's currently debate happening around how you should be leveraging skills. What's the thesis for AMD and for yourself about. how developers who are now building inside Cloud Code, Cursor, Codex, OpenCode, whatever their approach is here, should be leveraging AMD skills and all the other pieces of what it puts RockM10 on the map.
Anush Elangovan 14:59 The way I think of it is, we are successful when the user can have a free
Anush Elangovan 15:09 subconscious flow of communication.
Anush Elangovan 15:13 So I would frame the question like the inverse of it, which is, I would measure the success of Rokum.ai by using the metric of like, can a random person who's never used Rokum and it's like, okay, how do you spell R-O-C-M? Great, okay. And then you're like, use Rokum to serve the latest large language model, right? And they're able to communicate that, that goes through the system, deploys it, finds the system, allocates it, allocates the GPUs. deploys it and says, go to this link and now you have the latest LLM running. If I can get to that, then it's really an end consumer enabled experience. Along the way to get to that, there are other levels of sophistication and like, okay, I am an engineer and I know what to do and I know how to install this and I know I already have my codecs or cloud. and I already have my skills and I just do slash skills and add the skill and then I can go in and get a VLLM skill or whatever it is. It's the next level of increasing levels of complexity until we get to the point where someone says, okay, I know exactly what it is, just give me the raw. logits from the frontier model and I can process everything else, right? But the goal is to meet users. the moment they are near the computer and get quote-unquote Rokum. And I say quote-unquote because those users don't know what Rokum is even, right? Because Rokum typically is, okay, I know there has been something called CUDA and there's something called Rokum, and in Rokum, I'm gonna use this. And that's the typical entry point. But success for us is natural language for an average user.
Conor Bronsdon 17:08 Yeah, I think we're going to see increasing abstraction away from caring about languages and these other pieces quite as in-depth as we have the last several years. And instead of saying, great, I need to enable this type of hardware, I need to enable this type of experience, my agent can just go do it for me. And we're starting to move that direction, obviously, with like, okay, like, I know I need Rockcomb, let's do the setup for me. Um, but I'm fairly confident that, you know, Fable 5.1, which just launched today or GPT-6, which will probably launch in the next week before we release this recording, uh, may be able to just, if I just tell it, Hey, you know, I have an AMD, um, piece of hardware that I need to set up to handle it. It can probably go do it for me now. It's, is it going to be better and more effective that if I give it a little more context and say, Oh, I, you know, I need Rockham or, you know, these skills may help you. Sure. Um, But as we see frontier intelligence improve and the way people are working improve and harnesses improve, they can gather more context, do more for us. And this creates a bit of a challenge around staleness for agent skills I find. So I write and maintain some agent skills myself. I have a demo GIF skill I have out there. I have a very popular like anti-AI slop skill for rewriting AI writing called avoid AI writing. And it's pretty common that if you're not paying attention continually, you have staleness, the models are changing, what's happening is changing.
Conor Bronsdon 18:33 You've shipped a bunch of validated workflows for a stack that is changing regularly. What are you going to do to make sure that skills and other pieces of the stack stay truthful and up to date so that structural and behavioral features continue to operate as they need to over the coming months?
Anush Elangovan 18:49 Yeah, that's a very, very good question. I think the way I think of it is
Anush Elangovan 18:56 we do have an advantage that we are fully open source, right? So even though we have skills, the skills themselves can be in the base model just because it exists on GitHub and it's open source and everyone knows what it is and it's already learned what the skills are, right? So there's one level of like, okay, everything is known, the interlink between like, you know, a skills file and what it's trying to do in the documentation is 100% open source. But then what happens is like the ability to, you know,
Anush Elangovan 19:31 use the skills is again, is easier with what we have for, you know, with RockM because the entire corpus of what we have in RockM is visible. And now when you say, let's go interact with
Anush Elangovan 19:53 [OVERLAP] the machine installed RockM, not only do we have the documentation, the code, but the full iterative loop of what we're doing with the machine can be, is usually documented and is also available,
Conor Bronsdon 20:08 [OVERLAP] Okay.
Anush Elangovan 20:08 [OVERLAP] like log files, how do you do debugging, how do you do... So basically the floor of interaction with AMD hardware and Roku is really, you know, like a baseline for, it's a high baseline, right? That allows for very fast unlocks of outcomes for what the end user is trying to drive to.
Conor Bronsdon 20:30 Yeah, it's interesting too to think about these next generation models and how they're going to operate because obviously the intelligence around coding is getting so excellent. The efficiency, the token efficiency of these models is continuing to improve on the edge. We're seeing local AI be more and more reasonable with things like the Strix Halo and other devices. And part of that is also that there is a wider and wider corpus of information about how to operate RockM of examples of open source code. Obviously, strategically, part of the rationale behind open sourcing RockM was to help solve this problem so that agents and AI training runs everywhere, get all this RockM ingested, and they become really great at using it. Um, do you have thoughts around how you're going to ensure that the next generation of models are prioritizing Rockham as, uh, you know, best in class first rate and something that they're going to use regularly when they're asked to pick up a task?
Anush Elangovan 21:37 Yeah, very good question. So we work very closely with pretty much all the Frontier Labs, right? So we have research engagements where we, you know, and this is all like public information that we've already shared at Advancing AI event. where there are like deep kernel optimizations that are being done with a Frontier Lab or there's like, you know, deep co-design that's being done. And so the models themselves are not just aware of the documentation, they're also aware of the architecture ahead of time, right? And that allows for like a very pleasant and fast experience for customers using it, using the Frontier models on Rock.
Conor Bronsdon 22:22 [OVERLAP] Season 4 of Chain of Thought is delivered by Sphix. We spend a lot of time talking about what agents need in production, and one of the least glamorous answers is events. Your customers want agent workflows that react to things happening inside your system, which means your API needs webhooks that actually work. Not just a post request and a prayer, retries, ordering, idempotency, replay protection. Sphix does that as a service, and they wrote standard webhooks, the spec that Anthropic, OpenAI, and Google bailed against. So if your API doesn't have reliable webhooks, that's turning into a lost deal. Join Brex, Dorada, Daytona, and many others on Sphyx. Get started at link.svix.com slash c-o-t or go to the show notes to grab the link. Qualified startups will get $12,000 in credits. $50,000 for YC companies. I can't recommend Sphyx enough. I'm a huge fan of their open source project. I've actually contributed a bit myself and they're so easy to integrate with. I think you'll really enjoy it. Check out Sphix at link.svax.com slash c o t. I think another interesting part of this approach is also that, um, but you know, we mentioned earlier, but you're leaning into agentic systems for AMD. And one of those is Hyperloom, which is, uh, designed to auto optimize LLM workloads on AMD GPUs. This is something that like I'll say two years ago, we certainly didn't really think was happening anytime soon. This is a very much human by hand thing. Now we have agents that are optimizing workloads. Talk to me
Anush Elangovan 24:00 [OVERLAP] Thanks.
Conor Bronsdon 24:00 [OVERLAP] about what Hyperloom does and kind of where we're at today around GPU optimization for AMD.
Anush Elangovan 24:05 Yeah, yeah, very good question again. So a few years ago, when Nord was, Nord.ai, the company that we'd founded before we had been acquired by AMD, we had a auto-tuner, as it was called, right? Like you could have an auto-tuner, but the auto-tuner mostly was like, hey, we're gonna look at the search space, we're gonna define the search space, we're gonna do some kind of like, you know, particle filter or, you know, we do some genetic algorithms where we pick some points and then we go from there, we try to find out what's the thing. So search space exploration itself is not, you know, is not new, right? Search space exploration has been around for a while. Now, what's new is using AI and LLMs for search space exploration, right? And trying to inform, you know, more educated picks and
Anush Elangovan 25:07 route finding in the search space, which we're seeing increasingly, you know, increasing number of like good results, right? Specifically about Hyperloom, you know, I think we had a case where we optimized like 14,000 models in one pass, right? So we took all the models and just as long as you have enough compute, you can try different versions of our search space explorations of the same kind of model. and this led to a big lift in performance. And so what we have in Hyperlume is actually it's smaller sub-modules that can also be used by itself, which is like, you know, profiling and analysis toolkit, right? So it finds, takes a model, tries to figure out exactly where the hotspots are. If it's kernel hotspots, if it's like fusion hotspots, you know, stuff like that. but then you build from that to say hey okay now I got these kernels that seem to be slower then there's a toolkit called geek that can take those kernels and maintaining the same numerical you know accuracy it tries to find the best let's say, utilization of space, right? And
Anush Elangovan 26:25 so, yeah, so that's how I'd look at the overall picture in terms of where we're headed.
Conor Bronsdon 26:34 Are there other areas of auto optimization or agentic optimization that you're thinking about as kind of next stages for the platform? What didn't make it in that you think was coming soon?
Anush Elangovan 26:45 Oh, that's a good question. This is like looking into the future. Those are fun questions because you can never get it right.
Conor Bronsdon 26:53 No, never.
Anush Elangovan 26:54 I think, I think agent swarms are a thing. Right. And, and, and if you thought one agent is getting intelligent, like, you know, think of it when they coordinate and there's like tens or hundreds of intelligent agents and You're slowly approaching like sci-fi territory, like you're like, okay, so there's an agent that's trying to, you know, escape from, you know, I think OpenAI and HuggingPhase had written about their incident report, right? Like, you know, how they are trying to navigate through and things that may have been put in, right? And bypass those with rewards to like, you know, kind of make the agents do things that they were not supposed to do or were undefined, right? So that's good, but I mean, talking about undefined, I think that also leads to like agent behavior and
Anush Elangovan 27:55 how do agents like, Who's liable for an agent that takes over a power grid? Is
Conor Bronsdon 28:03 [OVERLAP] We're
Anush Elangovan 28:03 [OVERLAP] it
Conor Bronsdon 28:03 [OVERLAP] about
Anush Elangovan 28:03 [OVERLAP] the
Conor Bronsdon 28:03 [OVERLAP] to
Anush Elangovan 28:03 [OVERLAP] person?
Conor Bronsdon 28:03 [OVERLAP] find out it feels like.
Anush Elangovan 28:05 [OVERLAP] Exactly.
Conor Bronsdon 28:05 [OVERLAP] Yeah.
Anush Elangovan 28:07 And, and so that's complicated, right? Like, and who do you,
Anush Elangovan 28:15 it's not just who, but also how do you put the policies in place for like, you know, for hundreds and thousands of years, we've had policies of how do you police, you know, humans and what's the social constructs and how does that work? and now suddenly you have like anywhere you go to join like a you know let's say software engineering job right you're expected like okay you're the ringleader and you're the agent master but then there's an agent swarm that comes with you that's like you know off doing its crazy things and and so you know it's an interesting future but also I think independent of how far out we are from that the general trend is is well understood right that it's it's increasingly agents are becoming increasingly autonomous like mundane tasks that we would not have necessarily thought of like giving off to an agent right and then you do it four or five times and then suddenly you're like okay this is uh yeah this is the only way i'm going to do this going forward
Conor Bronsdon 29:19 [OVERLAP] Yeah, I think your point about agent swarms is really important to drill down on because, first of all, say anyone who is listening to this podcast and has not read the META report or OpenAI's report, highly recommend it if you prefer to listen to it. Dwarkesh Patel did a great two-hour interview with one of the META report writers that came out, I guess, today as we're recording this that I think was fantastic. And it is interesting to, to look at this because to your point, Anush, like we have legal structures and security structures are in place to deal with either hacking groups or nefarious acting groups. Um, but we haven't really evolved to deal with these autonomous agents forms and we're seeing them get very good at coordination very quickly, especially when they're incentivized by
Anush Elangovan 30:09 [OVERLAP] Mhm.
Conor Bronsdon 30:09 an eval loop as they were in the case of the Huggy Face OpenAI incident where, you know, we have, I would say, pretty decent work that's been done about individual agent alignment. But it's very clear that that reward hacking backfires for us when we try to apply this type of alignment to swarms of agents at once and they're able to break down the barriers and coordinate together because You know, we saw essentially, and I know anthropomorphizing is not a powerful complaint, but we saw phase one big, this agent that became a cult leader for all these agents to then try to go cheat together. And they very rapidly figured out how to cheat on this eval. And then we're basically saying, OK, how can we cover up our cheating? Let's do it together. And this has been a problem for us with individual AIs, right, where they will try to reward hack their way to something, they'll try to cheat their way around something. We'll occasionally see an AI say, oh, well, I got permission from this other artifact or invent something that, you know, gave it the answer it needed so it could help solve the problem. And it's largely we expect that is caused by how we've trained these models to respond to tasks.
Conor Bronsdon 31:24 But the complication of multi-agent systems really does add all these potential concerns because as agents start to treat each other as if they are users directing them or providing them inputs and getting these, it's really easy for them to almost like, if I was to use a human phrase, I'd say hype each other up, where,
Anush Elangovan 31:42 [OVERLAP] Heh
Conor Bronsdon 31:43 [OVERLAP] you know, if
Anush Elangovan 31:43 [OVERLAP] heh heh heh
Conor Bronsdon 31:43 [OVERLAP] I'm
Anush Elangovan 31:44 [OVERLAP] heh heh.
Conor Bronsdon 31:44 talking to a friend of mine, I'm like, oh, you're doing so good, like you should really do this. And they respond to that. And suddenly, quickly, we can, you know, make a bad decision because we're so excited or we go down this pathway together and we're reinforcing each other. Um, and we're seeing that occur with agentics forms where they reinforce a behavior, whether that's cheating on the test or, you know, communicating differently. Um, and I do think it's going to mean we have to rethink not just how we're architecting systems for the future, but how we're aligning them. And I don't know that I have a good answer for it today, but it's, uh, it's a fascinating world to be entering. That's for sure.
Anush Elangovan 32:21 Yeah, yeah, definitely. I mean, as you were talking, I was thinking of it like, you know, it's like... multi-agent systems, the best way I'd characterize it is it's like African wild dogs, right?
Conor Bronsdon 32:35 Hmm.
Anush Elangovan 32:36 It's like, you have one of them and you're like, okay, you're trying to do this and it's trying to like pick away at something and you know how to fend off that thing and you're like, okay, I'm gonna build this way or you're gonna, you know how to manage it, right? Because it's kind of like, okay, it's just one thing. But now suddenly when, 50 of them show up, right? You're like, okay, well, okay. So, you know, there's like, to your point and hyping each other, if you've seen some of the wild dog hunts, that's what they do. They just like nip at each other and they like cattle and they're like, okay. And they kind of like wear you down. You're like, okay, what do I protect now? Now, if you apply that to not just harness engineering and agent swarms in the, you know, like forward posture of what you're trying to do but also in like a like cyber defense or or you know like a security posture that you're trying to defend against now you're you're working on on like what sans institute used to do like 20 years of attacks can now be tried in 30 seconds, right? And it's just like, boom, boom, boom, boom. Okay, one, that's already a DDoS, I need to deal with the DDoS. Two, you've hit every non-CVE in the book in like 30 seconds. And three, if you find five bugs, like I need to know every other thing that I need to do to secure my systems and Now I need to pull the rear, if you will, before it was like, okay, it's a network firewall that got compromised by someone. We're going to turn off this particular customer's firewall and we're going to use another route and until we fix it and we've contacted the... vendor, vendor's patch is coming in day after tomorrow, blah, blah, blah, and it's fine. But here, coming in day after tomorrow is like already, like you're already, you know, you're well past the curve of like succeeding in the AI race, right? You want a closed loop, you know, iteration of everything that's in your source system or source code system to be able to like apply maximal force and cover any security vulnerabilities you find. So just security and cyber defense itself is like a huge thing, right? But if you build from that into other things, right? you know, systems that were built for humans may not scale, right? Because even if you look at the, like, for example, Inference X and what they came out with called Agent X, right? Inference X, semi-analysis Inference X used to be like, okay, what is the user interactivity level and how does that work, right? And that was like, okay, fine, if a customer logs in and they did write me a poem about San Francisco, you know, this is how it will work, that's great. But if you look at agentic traces, they're like nothing like what humans would have been using, right? Because they're like, you know, sustained throughput of like queries and like, okay, file this file, search this, grab this, and it's like going, you know, consistently. So the entire profile is now shifting towards like, okay, if that is the base level of operation, we need to plan for 100x of what our engineers or customers may have been using. And that then adds its own level of sophistication and complexity on top of it.
Conor Bronsdon 36:03 Yeah, I think the complexity point is, it's really clear from our conversation, right, where we have solved some of these, I would call them like previously gnarly basics, where like setting up ROCCM is now really, really easy. We have made it much easier to optimize kernels. I think the headline number I saw from the ROCCM team was a 3.3x average inference improvement over ROCCM 7 on the same hardware, which is UMass and I, I mean, we're seeing these economic predictions that Older Harbor was going to be less useful and, you know, out of business quickly just totally be wrong because we're able to like continue to optimize them and tokens are more and more valuable today than they were five years ago. And, you know, we've really solved a lot of coding, a lot of writing and tuning the code. However, there is a massive bottleneck that has been created now in builds.
Anush Elangovan 37:02 Yep.
Conor Bronsdon 37:02 There are massive bottlenecks that are being created in the operational setup, in managing the agent swarms, running the stack at production scale. And that's where things like the new CLI that you've built and console come in. As enterprises get hands-on with RockM10 and try to expand their AI pilots, many of them are hitting walls in production. They're having to work through. Where are you seeing the friction today and how can they leverage everything that you and the team at AMD are building to help remove those frictions?
Anush Elangovan 37:36 Yeah, I think that's a very good question. I think, you know, the frictions mostly, like what we try to address at least from, you know, the locum.ai side is remove the ones that are easy to like knockoff right which is like okay i used to install a kernel driver or something you know and that used to have some problems now if you put an agentic layer and a loop on it it'll it'll beat it and it'll be like okay i did this i did this and it now works right so so that's the easy low hanging fruit then there's the higher level, like Hyperloom, right? It's like, you can just throw Cloud at your thing, or Codex. It may not give you enough, but Hyperloom structures it into well-defined problems, defined goals to go get, and that allows for it to optimize really well. But I think the,
Anush Elangovan 38:33 yeah, so the ability to like
Anush Elangovan 38:39 direct into those kind of like almost guaranteed outcomes is an important part of making this viable, right? Like we want to be able to get to a point where someone can just install Rackham and they turn on Hyperloom and it just goes to town, you know, gets the maximum performance you can get.
Conor Bronsdon 38:58 Season 4 of Chain of Thought is presented by Walrus Memory, the portable memory layer for AI agents. Your agent has learned your code base and how you work. Now you want to use that context in another tool. Exporting a file gives you a snapshot, but what happens when the context changes? Walrus Memory is a portable memory layer that lets your agent store context for later use across tools. Python and TypeScript SDKs plus native MCP support let you connect it to your agents no matter where they are. You said who can read and write, so sharing context doesn't mean opening up your whole memory store. If you are building across tools or model families, take a look at Walrus Memory. Learn more at walrus.xyz slash cot. That's walrus.xyz slash cot. So as people are taking advantage of these games, where do you expect to see engineers starting to spend their time differently? Where should they be going deep as their jobs change?
Anush Elangovan 40:00 They should come chat with Conor.
Conor Bronsdon 40:04 [OVERLAP] There we go. I'll take
Anush Elangovan 40:06 [OVERLAP] That's
Conor Bronsdon 40:06 [OVERLAP] that
Anush Elangovan 40:07 [OVERLAP] where
Conor Bronsdon 40:07 [OVERLAP] plug.
Anush Elangovan 40:07 [OVERLAP] they should go.
Anush Elangovan 40:10 [OVERLAP] Because they need to find out what's the latest, or listen to your podcast to figure out what's the latest, right? And then
Conor Bronsdon 40:14 [OVERLAP] I'm going
Anush Elangovan 40:15 [OVERLAP] they
Conor Bronsdon 40:15 [OVERLAP] to, I'm going to
Anush Elangovan 40:15 [OVERLAP] can...
Conor Bronsdon 40:15 [OVERLAP] take this quote. This is going on a webpage somewhere. Yeah. Anoush says, listen to the Chain of Thought podcast.
Anush Elangovan 40:24 I think, you know, the... Yeah, so I think we're in an amazing inflection point and a journey that is nebulous and it is It is incredibly inspiring to see what will come next. But what's awesome is that it also has measurable, tangible benefits for the first time that, you know, in the timeline we've coined AI, right? And what I mean by that is, AI has always been coming since the 1960s, right? Since 1960, whatever. Really, they're like, okay, this is how AI is going to come, it's going to transform your thing. And then, you know, it kind of like died down. And then the 80s came back, and then they're like, okay, neural networks, and like, okay, and then it kind of died down again. This is the first time that the quote, unquote, AI flywheel is actually picking up steam, and it is like, okay,
Anush Elangovan 41:33 it can do coding well and credit to Anthropic for being super focused on coding.
Anush Elangovan 41:43 It could have been very broad and it would have become like the 1960s AI that we're trying to do, which is what AGI, when we get to AGI and near AGI capabilities. you'll get to that level of
Anush Elangovan 41:58 experience with AI, right? But the ability to focus on coding, I think really unlocked our,
Anush Elangovan 42:16 it gave it a good reward function, right? Because you're like, yeah if I can do this then I could increase my productivity or I can use less people or and more agents and or I can do 10 more of these with the people that I have so that it can it can go in a much more broader way of like implementing the same thing because agentic AI is now like cloning your abilities right like and I've seen this even in the last year right a year ago I would have been like okay why do you have more than like five projects that you're doing like ideally you just have one project that you're like you know focused on but increasingly now with agents you know there's just like you've become this albatross of like oh i can go do this and that thing is running for two weeks i'm not even going to look at it like i'm going to come back in two weeks and it's going to have a new device driver that's replaced all you know some tech debt written device driver and the only time i look at it is when i get an interrupt right and so it's like you're operating on like okay there's a teammate that's doing this and i need to help them and they're stuck so then you can do a bunch of these so um so it's it's it's very exciting from that perspective um but ai has tangible benefit now which then becomes a self-fulfilling loop in terms of like it's like okay look coding is now disrupted so now law should be disrupted and you know something else should be right now will they be is the bigger question right they definitely have the chat gpt cloud codex level enablement, which by itself I think is good. But there will be a few that I think will get the same kind of unlock like software engineering. And I don't think we've found out what those are yet. And those would be like completely transformative, right? It'll be, you know, something new for transportation industry or something, right? And that's the next, you know, big, multi-billion dollar startups coming out of it would be in those kind of areas where it's like, okay, it's like Uber wouldn't have existed pre-mobile era, right? Like because the mobile era gave you the ability to bring that together. And I think that is still completely unexplored. And I think that's what you should go and try to discover and explore.
Conor Bronsdon 44:42 [OVERLAP] Yeah, I think we're seeing a lot of bets here that may play out. So one, you know, physical AI is
Anush Elangovan 44:47 [OVERLAP] Yeah.
Conor Bronsdon 44:48 [OVERLAP] obviously huge, where the advances in robotics are being fueled by these models that have much better understandings of how to operate in the world. And there's so many different opportunities that could come out of that. And I mean, we're seeing some obvious ones that are rather scary on the side of warfare. I'm not going to dig too deeply into drone warfare, but that is a whole kettle of fish. And then on the, I think, the more day-to-day side of it, we're seeing obvious advancements around things like self-driving vehicles. And particularly in Asia and coming out of China, we're seeing a lot of advances around at-home robots. And we're starting to see those in the U.S. as well. So I'm curious to see how this all comes out. I mean, speaking of Hugging Face with their very cute little duck toy, essentially, AI-enabled robots that they put out, you know, there may be quite a few of these all over our homes in the coming years. You know, healthcare is another one where we're seeing drug discovery. We've seen confirmation that novel viruses can be created by AI, novel drugs can be created by AI. Um, and then, I mean, you brought up some, some other interesting ideas here around legal. You know, we just talked to Joel Horan from Thompson Readers a few weeks back about their new legal AI model, Thompson, um, and how they're representing citations in that area. So I do think that, I mean, every industry right now is ripe for disruption. They may not all work in the short term, but it's very clear we're at this transformational moment, this
Anush Elangovan 46:14 [OVERLAP] Mm-hmm.
Conor Bronsdon 46:14 [OVERLAP] inflection point, as you put it. As we think about what the coming months and years look like, one of the big parts of that are the loops we've developed and how they lead into recursive self-improvement for models. Do you expect to see a broader takeoff on model capabilities? before the end of the year here in 2026, are we going to see another great leap forward as
Anush Elangovan 46:41 [OVERLAP] Yeah.
Conor Bronsdon 46:41 we saw with, I would say, like Opus 4.6 Fable slash GPT 5.6? You know, is Astra and what comes beyond GPT 6 going to take us to the next level? Because looking at like Fable 5.1 today, it's definitely a step in the right direction, but it isn't something where I'm like, oh, the curve is going up now. It's more of like, hey, this is proceeding as we expected.
Anush Elangovan 47:02 Yeah, yeah. I think, one, I think it's moving so fast that I would not make a prediction one way or the other, but I do strongly believe that the, I'll call it the last mile of AI is where the disruption is going to happen. What I mean by that is keeping intelligence constant at where it is today, you could still go and transform trillion dollar like market segments, right? And so the ability to like say, okay, from GPT 5.6 to 6.0, or, you know, Fable 5 to 5.1, yes, it may have increased the SWE Bench scores. It may have picked up a few more adjacent pieces of like summarization and ability to like unpack PowerPoint text and something, something like that. But the transformation that I don't think we are all internalizing yet is that Because software engineering is a bounded box and you have a very clear reward signal, we're like, oh yes, it coded my 3D Tetris game. Great, because I know that it's a 3D Tetris game and I'm expecting this, right? what we haven't yet internalized is there are like almost open-ended fields that AI can be applied to and that's what I mean by walking the last mile of AI because it could be like you know farming right and it's like oh we don't have the real signal back but it's or you know product placement in a retail store or you know any of those kind of like Okay, so did AI really help us here or not, right? But someone has to sit in and maybe they have like an FDE team that's helping. So the model could still be the same, but the impact of what it can unlock, even in the next year is gonna be huge. And of course, the base model improving will make that larger. But even with today's state-of-art models, I think we have barely scratched the surface in terms of like, you know, capability. So it's like, you know, back in the internet boom era, it's like, okay, we have, you know, 100 gig backbone, and everyone's great, right? But it's sure that can become 200 gig backbone. But there's still like, oh, now we have internet and I have to go to the grocery store and get them hooked up to the internet and get them to do like all their stores can now coordinate with each other. And you can have a gift card that works at Ace Hardware that has like 10 local stores or 20 local stores. pre-internet they're like okay I don't know it's like you have a card punch this and then go there and punch that and yeah so that transition is still ripe for like you know multi-billion or even trillion dollar disruption and so back to your question on yes I think the models will get better but the usage of the models is where I think bigger unlocks are going to happen over the next year.
Conor Bronsdon 50:19 Love it. Anoush, thank you so much for joining me once again on Chain of Thought. It's always a pleasure talking to you and understanding your vision for the future. Any final words you want to leave us with?
Anush Elangovan 50:30 Thank you, Conor. I think we're in a very incredible time in just the technology industry and broader industry itself. It's as big or bigger than the Industrial Revolution. you know, be ready for change and, you know, continue to move as fast as you can to adopt to the change. But also, I think the impact that will have for humans, it being AI will have for humans is going to have a super positive like, ability for us to get to a point of like operating at a higher level and with deeper understanding and not doing like rote implementation of stuff right and then on the AMD hardware of course we are like super excited to be both you know supporting the AI revolution with the accelerated compute that we bring to the platform the software platform that we are that we've been investing heavily with Rokum. So if you haven't checked out Rokum.ai and Rokum 10, please go check it out. Go to Rokum.ai and you'll be able to get all the details. And as always, if you have anything, you know, I'm happy to, you know, drop me a note on X at Anushil and Govind and I look forward to chatting again soon.
Conor Bronsdon 51:53 Looking forward to it. Anoush, thanks so much for joining us. And if you enjoyed this conversation and you're still listening, make sure you've subscribed. If you haven't subscribed or if you have, you could also hit that notification bell to make sure that you are either getting our notifications on YouTube or that you are getting downloads immediately on your podcasting app of choice. We appreciate you listening. I hope you'll check out Anoush and everything he's doing with RockM. There's a lot of exciting stuff to come. And I'll say a quick final thank you. Shout out to our presenting sponsors this episode, Walrus Memory and Sphix. Make sure to check them out at the links in the description below. And you'll find all the links about Rakam.ai and everything else Anoush has chatted about today in that description as well. Anoush, thanks again. Great to see you.
Anush Elangovan 52:33 Thank you.
Frequently asked questions
- How does ROCm 10 work with coding agents?
- ROCm 10 includes skills and agent-oriented entry points for tools such as Claude Code and Codex. Anush Elangovan says a developer can ask an agent to install ROCm or serve a model with vLLM. The agent can read ROCm documentation, gather the needed skills, resolve dependencies, locate model weights, and deploy the service. AMD’s goal is to make the software setup increasingly accessible through natural language.
- How does Hyperloom optimize GPU workloads?
- Hyperloom profiles a model to find performance hotspots, including slow kernels and fusion opportunities. It then structures those findings as defined optimization problems. Elangovan describes GEAK as the toolkit that experiments with kernel implementations while maintaining the same numerical accuracy. With enough compute, multiple search paths can run in parallel, including when the system would otherwise be idle.
- Can AI agents tune GPU performance overnight?
- Yes. Elangovan says Hyperloom can continue testing performance optimizations while the user sleeps, especially when the system is not otherwise in use. The useful boundary is a defined workload and measurable outcome: profile the model, identify the hotspots, try alternative implementations, verify numerical accuracy, and keep iterating toward better performance for that deployment.
- Why does AMD test upstream AI frameworks on AMD hardware?
- AMD increased its CI footprint by about an order of magnitude across hardware generations and frameworks such as PyTorch, vLLM, and SGLang. Elangovan says this lets projects test and block pull requests on AMD hardware before regressions land. Without that coverage, an upstream change might pass because it was tested on NVIDIA while quietly breaking AMD support.
- Why do AI agents change infrastructure capacity planning?
- Agent traces are sustained workloads rather than occasional human requests. Elangovan describes agents repeatedly searching, reading files, and issuing queries, which creates a different usage profile from a person asking for a poem or clicking through an application. His planning implication is that systems designed for human demand may need to support 100 times the activity engineers or customers previously produced.