Your Best AI Engineer Might Have the Worst Metrics | Sonar CTO Andrea Malagodi
Key takeaways
- Cost per PR can’t tell you which engineer is doing the valuable work. When finance asked about Sonar’s AI bills, CTO Andrea Malagodi worked out cost per PR. One engineer had over 500 PRs in a very short period at a low cost per PR; another had very high spend and very few PRs. A pure FinOps reading calls the first good and the second bad. Unpacked, they were two engineers doing quite different work: “there is a story behind those numbers.”
- The 500-PR engineer had built an agent factory with an opinionated flow: a specification broken down to a definition of done, design and architecture that the engineer reviews, documentation records kept for long-term memory, coding, verification, and functional testing. Almost two-thirds of the output was tests and validations. It prompted Sonar to look at bootstrapping every engineer’s terminal with a setup like it.
- The high-spend engineer was working on very complicated problems that needed deep thinking from the AI over long, long turns, work Malagodi described as not directly attributable to a tidy output of code. Malagodi’s rule for leaders: understand the context and the problem someone is working on before passing judgment on their numbers.
- Token leaderboards repeat an old measurement mistake. Malagodi traces it from lines of code to story points to tokens and token maxing, and notes that the people who removed lines of code were probably the most important people a team had. He treats these numbers as an input, because software engineering is a craft.
- Human review needs help to stay meaningful. In Malagodi’s words, five changed lines get argued about for hours while 500 usually get “looks good to me”, and a human in the loop who isn’t given real assistance is being done a disservice. He recommends breaking agent work into pieces small enough to oversee; a five-day agent session that produces 400,000 lines of code leaves you with no idea what’s in your codebase.
- Rolling AI tooling out at enterprise scale is a cost-control problem. Malagodi wrote his own cost analyzer to give finance one view across model providers, says Sonar’s context studies found token-cost savings of up to 30% from cutting repeated code reads (the job Sonar Vortex does), and warns that unlimited tokens build habits that are hard to pull back, while some engineers arrive from teams with budgets as small as $50 a week.
Concepts in this episode
AI terms discussed here — each links to a plain-language definition.
TokenizationLarge Language ModelHuman in the LoopAI AgentAI EvaluationAgent SwarmModel Context Protocol (MCP)
Chapters
- 0:00Human reviewers rubber-stamp big AI changes
- 2:43Learning the craft alongside AI tools
- 9:09From lines of code to PR counts
- 12:17Spreading top engineers' setups org-wide
- 16:50The guide, verify, solve agent loop
- 22:51Specialized agent lanes and multitasking limits
- 26:51Clear asks and guardrails for agent coordination
- 30:29Building disagreement between agents
- 34:59Mixing models instead of picking one
- 39:46Rolling out AI tooling at enterprise scale
- 46:19Defense in depth for AI-written code
- 50:52Protecting time for incident investigations
- 54:08Personalized software for teams and individuals
Show notes
Andrea Malagodi is CTO at Sonar, which builds AI code verification and governance tools. When finance asked about the size of his team's AI bills, he calculated cost per PR. At one extreme was an engineer with more than 500 PRs in a short period and a low cost per PR. At the other was an engineer with very high spend and very few PRs.
Judged on cost alone, the first engineer wins and the second looks like a problem. Malagodi looked closer. The first had built a personal agent factory, with specification, design review, documentation records, coding, verification and functional testing, and about two-thirds of the output was tests and validations. The second was working on a hard problem that needed the AI to reason through long, multi-turn sessions and didn't reduce to a line count.
Numbers alone tell you something, but not which engineer was doing the more valuable work.
We cover:
- Why cost per PR is becoming the industry's standard AI spend metric, and how Malagodi's 500-PR engineer and high-spend engineer broke it
- The history of bad proxies, from lines of code to story points to token leaderboards, and why Malagodi says the people who removed code were often the most important
- Sonar's guide, verify, solve loop: Vortex for codebase context, algorithmic plus LLM-based PR analysis, a remediation agent for legacy debt, and the Hunter agent for security vulnerabilities
- Why Malagodi breaks agent work into small pieces, and why a five-day session that produces 400,000 lines of code leaves you with no idea what is in your codebase
- Setting up disagreement between agents with an orchestrator and personas for engineering, product management and quality, plus Conor Bronsdon's cross-model-family review setup
- Rolling AI tools out at enterprise scale: one cost view for finance, up to 30% savings from cutting repeated context reads, and why budgets as small as $50 a week are hard to work with
- Rotating a triage duty to protect deep work, including how Sonar's teams built skills to sort through hundreds of reported CVEs, a number of them likely false positives
New episodes, the ideas behind them and Conor's essays land in the Chain of Thought newsletter first. Subscribe: https://newsletter.chainofthought.show/
Chapters:
(0:00) Human reviewers rubber-stamp big AI changes
(2:43) Learning the craft alongside AI tools
(9:09) From lines of code to PR counts
(12:17) Spreading top engineers' setups org-wide
(16:50) The guide, verify, solve agent loop
(22:51) Specialized agent lanes and multitasking limits
(26:51) Clear asks and guardrails for agent coordination
(30:29) Building disagreement between agents
(34:59) Mixing models instead of picking one
(39:46) Rolling out AI tooling at enterprise scale
(46:19) Defense in depth for AI-written code
(50:52) Protecting time for incident investigations
(54:08) Personalized software for teams and individuals
Links from the episode:
- Sonar Vortex research: https://www.sonarsource.com/blog/cut-your-coding-agents-cost-with-sonar-semantic-code-navigation/
- SonarQube Hunter Agent: https://www.sonarsource.com/products/sonarqube/hunter-agent/
Connect with Andrea Malagodi:
- LinkedIn: https://www.linkedin.com/in/malagodia/
- Sonar: https://www.sonar.com/
Connect with Chain of Thought host Conor Bronsdon:
- Newsletter: https://newsletter.chainofthought.show/
- Twitter/X: https://x.com/ConorBronsdon
- LinkedIn: https://www.linkedin.com/in/conorbronsdon/
- YouTube: https://www.youtube.com/@ConorBronsdon
More episodes: https://chainofthought.show
Thanks to Svix, presenting sponsor of season four of Chain of Thought. Svix delivers billions of reliable webhooks for startups and the Fortune 500. Get started at https://link.svix.com/cot. Qualified startups get $12,000 in credits, and YC companies get $50,000.
Thanks to G2i for sponsoring this episode - for over a decade, they vetted and placed engineers at other companies, from startups to FAANG. Two years ago, they turned that same judgment inward, building their own bench to review RL environments, evals, and training data that models are trained on. Get access: https://fandf.co/3SFxVm6
Transcript
112 segmentsAndrea Malagodi 0:00 five lines of code change, we can argue about for hours. Five hundred lines. You know, it's usually looks good to me. You only live once, so let's just get this change in. When I hear people say, Oh, there should be a human in the loop, I'm like, Okay, but if you're not helping the human in the loop, you're really doing them a disservice and you're probably being a bit naive.
Conor Bronsdon 0:23 An engineer producing thousands of pull requests can look extraordinary on a dashboard. Another producing a handful can look expensive. Deciding which one is doing valuable work takes more than just PRs merged. Joining me today to talk about it is Andrea Malagodi, CTO at Sonar, which provides AI code verification and governance for shipping code with confidence. Andrea, how are you feeling about software engineering today? And it's great to see
Andrea Malagodi 0:47 Uh thank you very much for having me. Happy to be here. Um what do I feel about software engineering today? I think lately um All the focus is on numbers and statistics, token maxing, etc. And maybe somewhere along the line, have we maybe forgotten that software engineering is also a craft? Right? It's something that actually takes time to learn, it takes experience. Uh some of us have had to go through a few incidents over the years to actually get some benefits of learnings, etc. And AI is, you know, a set of tools that are Incredible, uh amazing, can do wonderful things.
Andrea Malagodi 1:24 But it also takes a craft to learn how to use them and apply them. in a way that gets great outcomes. Uh and I think right now industry in general Is try to find out like how do I actually attribute the cost and the investment I'm making. uh to outcomes um I don't think anybody has figured out the silver bullet answer here. But I think all of us are trying to look at different aspects of the data that we get and that we can see. To try to make sense out of the information we're getting and try to figure out where we should invest. and where we should potentially pull back or modify a little bit the approach.
Conor Bronsdon 2:03 I think that's a really great point because we all get a little number obsessed sometimes and there is so much change happening right now. Within AI and within software engineering in general. And I would be remiss if I didn't give a quick thank you to our sponsors for this episode for helping us navigate that change. Svix, Walrus Memory, Inngest, our seasoned presenting sponsor, Chain of Thought, G2i.
Conor Bronsdon 2:25 But as we continue to explore the craft of software engineering. the numbers behind it as well as we explore how AI tools Can vastly impact the space and have, frankly, changed it massively over the last year or two. I'm Wondering, Andrea, from your perspective, How do you think engineers should be learning about the craft today? You mentioned that in your opening statement here. That there is a lot of learning to be done. We can't just rely on these tools, we can't just look at the numbers. What should engineers be thinking about as they try to learn the space?
Andrea Malagodi 2:58 Yeah, I mean the I think I guess there's two Dimensions of this, which is the one on the individual engineer, where I think we've gone from, you know. somewhat prompting to using skills, lots of other um instruments to try to instrument the AI to come out with the best outcomes. You know, at the rapid pace that we can, and obviously at a reasonable price. So engineers need to kind of move up that maturity curve. And uh I think a lot of people end up being a little bit stuck in uh call it you know advanced prompting.
Andrea Malagodi 3:34 Uh and you know have maybe not really jumped into the space where saying well actually I'm handing work over And I'm expecting uh an AI setup to basically be able to solve that ask that I have. from A to Z to get to a P R that can be shipped. and basically go out the door. And on the manager side or the leadership side, You know, we obviously get Asked by people in finance to say, well, those numbers look awfully big. What's actually going on? Do you know what's happening here?
Andrea Malagodi 4:06 Um so a lot of us obviously spend time trying to figure out to crunch those numbers and see well what actually is going on. And you know, industry, I think, is going towards a cost per PR, which has kind of become a common number to look at to say, okay, what was the actual price I had n amount of sessions and it resulted in a PR. And if I had this amount of cost and this amount of PRs, You know, if I do some equalization on that.
Andrea Malagodi 4:33 I get to a reasonable understanding of the cost of PR. And we did that. And what I observed in the numbers, which I have to say, I mean, it made me incredibly curious immediately. was I had on the one side Individual at the one extreme that had an enormously high PR count, over 500. uh you know, in a in a very short time period. And obviously the cost per PR was quite low. And I had on the other side uh a person that had a very, very high spend but very little PR.
Andrea Malagodi 5:02 And I was like, okay, well, what's going on here? And if I looked at it maybe from a pure FinOps perspective. you know, I would say, well, mm black and white, you know, many PRs, low cost, good. Uh less PRs, uh, high cost bad. Uh okay, those people that are bad we should not have. Um but the reality is when I then unpacked that I found two engineers that were doing Uh quite different work. and had quite different problems at hand that they had to solve.
Andrea Malagodi 5:29 Uh the individual that was doing this very, very high PR count. had actually created this specialization, I would say, of having call it a factory for themselves. of you know agents And a a highly um opinionated flow about what should happen in each of the stages of the work that these AI agents were asked to do. Um The individual has basically said, listen, I'm going to come with a level of specification. And before you do anything else, I want you to break this out and go into deeper detail.
Andrea Malagodi 6:04 of what all the different things are that are required. to achieve a definition of done. Uh this then gets handed over in a second step of the workflow. of actually doing some design and some architecture of that. And this gets reviewed by that individual to make sure they're comfortable with this. Um documentation records are written. to make sure that that's persistent in memory for long-term purposes. Obviously there's coding. And there's verification of all that code. Uh and then eventually we get to a P R. And then, you know, there's also functional testing that happens in all of that.
Andrea Malagodi 6:37 But this individual had crafted a very, very, very skilled. highly advanced in my opinion. flow to optimize For being able to go at an enormous amount of velocity to get through a massive amount of code. Uh that got produced. Interestingly, with the ratio of test to actual functional uh code that you know was um almost Two-thirds of it was basically tests and validations. as well of all the the stuff that happened. So you know really quite sophisticated. you know, a very uh happy chap and you know always happy to share his experience. And on the other hand, you know, I found an engineer that had some very, very complicated problems around C.
Conor Bronsdon 7:04 Yeah.
Andrea Malagodi 7:20 and some of the work we are doing there. That really needed a lot of very deep thinking by the AI that was very challenging. for the AI and had to go in long, long turns to go through and analyse this particular work. This work wasn't directly attributable to, okay, do all this. You know, great, I get six lines of code, off the code goes necessarily. Um so I think There is a story behind those numbers and it's important that we you know, if you're leading an organization or a team, that you really get familiar with understanding What are the different contexts that people are working with?
Andrea Malagodi 7:57 within And what problems are they trying to solve? So yes, it's good to have numbers. Uh and we should be very careful to not pass uh judgment too quickly. without really understanding the context of the problems that are being worked on. And the the the Um issues that's being trying to be solved or the opportunity that's being explored. It was a great learning for me. And the best part of that probably was to get this individual that was essentially making using AI part of his craft.
Andrea Malagodi 8:28 And then being able to obviously share that to you know other engineers so they can take benefit for it. Um Within that, of course. Sonar. focuses on this uh fundamental idea of trying to guide the AI, verify the outcome. And most of us work on brownfield applications of some sort. So there is a backlog of things. That can be solved by, for example, a remediation agent that can go and look at all past issues. or now also Hunter agent that can go and look at security Um issues that may not necessarily be found in the upfront workflow or be part potentially of legacy code that needs to be inspected anew.
Conor Bronsdon 9:09 It's funny because I think we were having this debate a couple of years ago broadly as a community about velocity and lines of code and people over optimizing for them. and understanding the value cycle of software and the differences between different projects. And we're kind of back to having the same debate. Again, after I think we had this great run of like, oh, let's talk about developer experience and actually, like, what's the value being delivered. Uh and then I think we kind of got re-obsessed with velocity because we weren't able to enable so much of it.
Conor Bronsdon 9:41 Through AI coding, and because AI can simply code faster than any human can. Uh over time. And it has opened up the opportunity to throw so much more code at a problem, to solve problems faster with code. Yeah. And it's changed the conversation, but it almost is making us fall back to fundamentals again: of like, okay, what's the quality delivery?
Andrea Malagodi 9:42 Yeah. Yes. What I mean? Yeah, and you know, in the past it was lines of code and we measured that, you know, lots of lines of code, fantastic, uh not a lot of lines of code, not very good because, you know, we're paying a big salary to these people so they should produce many lines of code. And usually, to be honest, the people that removed lines of code were Probably the most important people you had at the time. And then, you know, we all got agile and scrummified and whatever.
Andrea Malagodi 10:27 Uh and then it was story points, you know, a lot of story points was better, so everything inflated, you know, but you know, it looked good on numbers boards, etc. And then here we are with tokens and token maxing and it's like, oh, let's have a leaderboard of that. Again, I don't I mean You can't blame people for wanting to try to make some type of quantification. It it's quite normal. uh in in society that we try to quantify the these things But I I you know we think we should be careful not to you know overreact to those things and try to take them listen, it's a an input.
Andrea Malagodi 10:59 And in the context of software, it is a craft and the moment you forget that and just make it into you know, some type of balance sheet exercise. Uh I think you lose a lot of the understanding of how systems come together. Um Most of us that bank are very comfortable knowing uh that banks put an enormous amount of focus on quality and verifying every single thing that they do. At the or origin of that code when it gets, you know, the first line gets written out.
Andrea Malagodi 11:31 all the way through to very complicated large regression test suites, et cetera. and of course security testing and validation. Um and so you know There is a very important place for that. in most software. Although I would say some software today is getting to be a little bit throwaway, right? Uh you know, everybody has got AI. Everybody builds beautiful dashboards. They may last a couple of months, but they're probably not the business critical ones. And you know, for those, yeah, you need to make sure you don't do dumb things.
Andrea Malagodi 12:01 Like expose your database to the world so they can go and suck out all the data, which has happened in a number of cases. Um but you know uh yeah, quality is an important aspect of all software. But obviously there's some that are much, much more important uh than necessarily others.
Conor Bronsdon 12:17 What are some of the lessons you learned from these comparisons you're doing as you talk to these individual engineers, as you do these studies? You talked about the software factory that this one engineer has built versus the multi-turn and kind of deeper thinking approach. How are you taking those lessons from these premier engineers to the rest of your organization?
Andrea Malagodi 12:38 Yeah. I I mean, I think the thing that uh I've always concerned myself with is how do we take something that is a good thing in one case and try to figure out to make it broadly available. To everybody, right? So the thing that this inspired a little bit on our side was to say, well, Couldn't we think of a way that we could get every single engineer bootstrapped in a way in their terminal with having this set up so they can essentially work in kind of this higher maturity level.
Andrea Malagodi 13:10 way Uh which I think has a better outcome for Um for everybody and you know uh uh A good uh scale economics. in terms of what we are burning, in terms of tokens, and in relation to the outcomes we get. Important for us of course internally, but obviously also for our customers. is that we integrate the steps and the things that can reduce cost. of tokens. So for example, in our solution, we have uh an aspect called vortex. uh which essentially gives the agent the ability to pull out context of the particular code base that's being worked on.
Andrea Malagodi 13:49 be that coding guidelines. or the graph of that code. So essentially reduce the burning through of rereading all of that code by the LLM. which is kind of a little bit of a wasteful effort. Um so things like that are important for us obviously too. put into our own workflow. And obviously, that's also what we are trying to help customers get. We don't want the long loop of, hey, I'll write a line of code. I think this is ready. Let's go do a PR. Wait for that whole out a loop cycle to complete.
Andrea Malagodi 14:19 So we want to give customers the ability to get and our own engineers, the ability to get an answer very quickly, is there a problem? With this particular file, or these files, or the files in combination, almost at the terminal level. Um so that's a little bit what we focused a lot on this year. the recognition that Can we make that inner loop? more economical. and faster in terms of turnaround, so you don't have to get the long life cycle. uh of the outer loop completed before you get feedback. Uh so that's obviously an obsession as well, to make sure we we can make sure that everybody works consistently in terms of that.
Conor Bronsdon 14:55 Season 4 of Chain of Thought is delivered by Svix. We spend a lot of time talking about what agents need in production, and one of the least glamorous answers is events. Your customers want agent workflows that react to things happening inside your system, which means your API needs webhooks that actually work. Not just a post request and a prayer, retries, ordering, idempotency, replay protection. Svix does that as a service, and they wrote standard webhooks, the spec that Anthropic, OpenAI, and Google build against. So if your API doesn't have reliable webhooks, that's turning into a lost deal. Join Brex, Dorada, Daytona, and many others on Svix. Get started at link.svix.com slash C-O-T or go to the show notes to grab the link. Qualified startups will get $12,000 in credits, $50,000 for YC companies. I can't recommend Svix enough. I'm a huge fan of their open source project. I've actually contributed a bit myself and they're so easy to integrate with, I think you'll really enjoy it. Check out Svix at link.svix.com slash c-o-t.
Conor Bronsdon 16:01 Season four of Chain of Thought is presented by Walrus Memory, the portable memory layer for AI agents. Your agent has learned your code base and how you work. Now you want to use that context in another tool. Exporting a file gives you a snapshot, but what happens when the context changes? Walrus Memory is a portable memory layer that lets your agent store context for later use across tools. Python and TypeScript SDKs plus native MCP support let you connect it to your agents, no matter where they are. You set who can read and write, so sharing context doesn't mean opening up your whole memory store. If you are building across tools or model families, take a look at Walrus memory. Learn more at walrus.xyz slash cot. That's walrus.xyz slash cot.
Conor Bronsdon 16:50 And I've seen you in other circumstances break this down into multiple, I believe you called agent-centric development. loops that you're using. Can you talk a bit about how you're thinking philosophically about what this looks like as we move forward in software engineering and how you've seen Uh this be redeveloped the current era.
Andrea Malagodi 16:57 Yes. Yeah. I think uh we call it guide verify solve. Uh so essentially we start you know at the beginning in the terminal. or we hand something over to a terminal that's running somewhere else. where agents can essentially interact with the code base. and get this context that I spoke about, the coding guidelines. Uh the architecture of the application, the intended architecture, and essentially get bootstrapped with a bunch of parameters on top of the functional requirements that needs to be done. And work can essentially continue through those agents.
Andrea Malagodi 17:41 They may be collaborating. You may have multiple agents that are kind of Challenging each other, but the result is something that you feel is ready to get to a commit stage. where we can do a verification. of that outcome before it kind of makes its way into the kind of into the sausage factory of getting the software out the door and get it shipped. After that We then do uh both an algorithmic and you know an unbased analysis of the PR and the PR in context of the overall code base.
Andrea Malagodi 18:11 And we use both of those complementary techniques to make sure we really cover holistically. all the different techniques and methods you could use. to ensure that you have you know, very, very high. Quality standard and security standard of that code. Because once that PR gets accepted, it's in your code base. You know, that's your intellectual property. And you want to be careful with that to not pollute it. and you know get it to break down and deteriorate over time. After that, as I said before, um Most of us work on some type of brownfield application.
Andrea Malagodi 18:41 They're changing very rapidly. And in that, of course, we have got legacy issues that we need to go kind of and deal with. Normally that has been Making its way through backlogs that, oh, let's have a hardening sprint to go and fix it. Today, we offer a remediation agent that can essentially plow through this. and raise PRs. that developers can then look at and then accept. So we're kind of taking a lot of that burden away and the cost of ownership. of that technical depth. drastically.
Andrea Malagodi 19:10 Another aspect which today, and we've seen this recently in news, but is this exploration of the code base from a security perspective. And specifically here Uh you know, we got an an amazing code security team. Kudos to everybody and them too. And they have spent quite a lot of time researching. what is the best way to traverse through the code base to find all these potential vulnerabilities that could exist Which accumulate potentially over time and haven't been explored in this way that we can do today in the past.
Andrea Malagodi 19:47 So you may have applications that may not have undergone a lot of change, but have an accumulation. of issues that go across um in the application stack. Input sanitation. is a famous issue in lots and lots of applications and lots of hacks. and breaches has happened because of lack of sanitation. Uh well guess what? Uh they're still there out there today. They still get added and created and introduced into code base a day. having a hunter agent that can really traverse through all that code base and find them and actually highlight The risk profile.
Andrea Malagodi 20:22 to get those issues remediated. Obviously you can give them to a remediation agent to get them solved. Um But really making sure that you have confidence that that code base that you are operating on the basis of, you're probably operating your business. Uh really meets the highest threshold of security is is a super important Uh aspect. And you know, we call those, if you will, kind of background agents. There shouldn't necessarily be a human that says, hmm, you know, I want to go and do this thing.
Andrea Malagodi 20:51 So it is a little bit on a schedule, it's automated. It happens in the background. And I think the biggest challenge for you know, uh anybody is that Uh if if I create a PR, it's the most important thing in the world. If somebody else creates a PR and puts it to say, hey, Andrea, could you please review it? Sure. When I've got time, right? And this is a little bit the challenge with these things, is that we get a lot of changes coming out of the woodwork.
Andrea Malagodi 21:16 And we have this backlog of things that we have to go through. Uh guitar is awesome, for example, from our perspective in terms of being able to go and traverse through these PRs. And explain what the hell is this change actually about? What should I focus on? What should I be worried about? Uh plug into. And by the way, are there actually any obvious issues here that need to be fixed before I even do anything? Uh so please go and fix them. So you need this constellation and combination of tools.
Andrea Malagodi 21:44 Uh to help you Be efficient. To help make sure that you maintain the high standards of the intellectual property you already have. You have to recognize that what you already had may already have a series of technical debt or security issues. and you want to make sure to get rid of those. The world is not getting safer when it comes to software in terms of things that can happen out in the wild. So we want to keep an eye on that. And then we also need to think about developers.
Andrea Malagodi 22:13 ability to actually pay attention to the most important things. Otherwise it's just a wall of noise. And I think I've said this many times before, five lines of code change, we can argue about for hours. Five hundred lines. You know, it's usually looks good to me. You only live once, so let's just get this change in. And if you don't recognize that and just say, When I hear people say, Oh, there should be a human in the loop, I'm like, Okay, but if you're not helping the human in the loop, you're really being a little bit Uh you're doing mi doing them a disservice and you're probably being a bit naive. Um if you're not providing some real assistance to them.
Conor Bronsdon 22:51 Ton of great points in here. A couple that I want to pull on in particular. So, one, coordination. I think it's really interesting how Sonar is beginning to set out these various lanes with specialized agents across multiple loops, too. interested to talk a bit more about coordination, security, obviously. But first I want to point out something that you said, which is about prioritization. both for individual developers and within the code base. And I actually think it's gotten harder for us to prioritize because we're all being encouraged to.
Conor Bronsdon 23:23 Assign out work to agents right now. And so it's very easy to have. 20 different threads going with multiple different PRs and process and code review. And I worry personally sometimes um that I am Spending too much time across a variety of tasks, achieving a lot of work, but not necessarily prioritizing. the top work that is going to be most impactful for me. and something I've had to personally check myself on. Are you systematically enabling that within Sonar? And when I say enabling, I mean Are you helping ensure that high priority work is prioritized first? How are you coordinating this through the system?
Andrea Malagodi 24:00 Yeah, I I mean firstly I'd like to say that I think that multitasking is an illusion. it doesn't actually work and it doesn't produce better outcomes. So I think Uh we don't need AI. Necessarily to teach us that lesson. I think it's just a life lesson, multitasking. uh it does not produce better outcomes.
Andrea Malagodi 24:20 Uh just from that perspective. Um Privatization is Super, super difficult. Uh no doubts. And I I don't think that's an AI problem necessarily. I just think the volume of possibility to actually have these multiple tasks handling or happening uh produces a optionality in terms of taking on more Right? That's the expectation perhaps. Um And the reality is that our ability to prioritize between these different things is not particularly uh great I would say. Uh and I don't think that's an AI challenge, it's always difficult. I've got a backlog of things that we need to get through.
Andrea Malagodi 25:05 There's a customer priority for this product feature. We have this bug that we need to get fixed and squashed. Um It's always challenging to do that and there's a little bit back to the software craft thing. Right, there's a level of experience about figuring out what is the most important thing. Um For us, I have to say we give a lot of freedom. To the individual engineers, obviously, to look at what is the most important things. And Then again, also when you think about vulnerability management and stuff like that, there are certain standards that we just must.
Andrea Malagodi 25:37 implement and we must obey by. uh certain regulation recently, for example in Europe, requires very, very active engagement on anything that smells of vulnerability in terms of timelines, reporting timelines. So you have to make sure you educate the organization on that and you make sure your policies and your procedures actually support it. But in terms of agent to agent prioritization. I can't say that we necessarily have like a great framework that says this agent. I think you need to figure out what your agent's M D file says about what the order of things and where the challenges should happen.
Andrea Malagodi 26:16 So if you have for uh you know if you feel find that's a good approach to personify the fact that there is somebody that is representing the product manager role, if you will. and coming out with definitions of done. Well if m That Part of the agent loop is saying no. This isn't up to scratch in terms of what I expect in terms of definition of done. Obviously from that perspective, inter agent-wise. you could establish a prioritization order. in terms of who has power to accept or reject certain things that come through that agent loop.
Conor Bronsdon 26:51 And on coordination in particular, I think we've seen quite a bit. Both from developers and their setups, but also broadly in the media lately, as we've Explored some of the agent forums that have been leveraged by agents forums to coordinate from misaligned evals and other circumstances. The Huggy Face hack obviously being big example of this, but we've seen a recent German Wiki forum that was used and others. How are you thinking about the tools provided to agents and Agent Swarms more broadly. Uh as far as how we coordinate work going forward. as humans step back to a greater level of abstraction.
Andrea Malagodi 27:34 Yeah. Um I mean I think the first thing that comes to mind to me is you should be careful that you're clear about what you're asking for uh and what the guardrails are you're actually setting up for that. Uh there's obviously network tricks and we've seen them not work. uh that you need to think about if you're sandboxing or other things like that. I think for most cases though, I mean these are a little bit extreme in media and they're a little bit extreme cases of of testing the boundaries of things that's happening in the frontier labs.
Andrea Malagodi 28:03 Uh you know, in most companies the I think the benefit of a swarm of agents Is that you establish a clear scoped problem or opportunity? that you want to get explored. and give that to A set of agents, or it can be an agent that can spawn multiple sub-agents. to be able to explore that in depth. But to do that, I do think you need to make sure you're clear on what the boundaries, of course, of that is. And how is it you set up your instruction sets?
Andrea Malagodi 28:33 around what it is you're expecting. of outcomes. And when is it you have a human in the loop being pinged to be pulled in to essentially verify. Right. It's super easy to go and say, you know. uh always allow uh and uh good luck and let's see what happens. Um But I mean, you also want to be a bit mindful about like, hey, there is probably a natural checkpoint. in the outcome of this work. You may not necessarily know What the penultimate outcome is But you know that there is a definition phase and you want to be quite clear.
Andrea Malagodi 29:07 about going through that and you should torture yourself on the definitions to make sure you agree with the have definitions have done. and the boundaries that may be generated by agents. So that instruction set carries into the next phase. or whatever that workflow is that you're trying to execute. Uh I think the The risk is that people are becoming too permissive. And maybe they say, you know, it becomes a little bit easier to become lazy. Right? Hey, go and solve this thing, you know, and then kind of always allow and let's see what the hell happens.
Andrea Malagodi 29:40 You know, that's quite risky. And again, this takes You know, with great uh autonomy also comes great responsibility and I think everybody that's involved in industry has to take that responsibility seriously. And not just um deflect that off to Mm-hmm. some AI engine that's gonna come out and do type of a uh a miracle. Uh scoping is super important. making sure that you're clear on those definitions and boundaries of what you expect. And you know, from time to time we see some surprising uh responses back saying, oops, I actually completely ignored that instruction and I g went and did something else.
Andrea Malagodi 30:17 Um so there you have to Hopefully you've got good security teams that are taking care of your developer environments and making sure they are boxed. well and securely uh in relation to that.
Conor Bronsdon 30:29 And I think part of the challenge of this coordination too is that different agents. You know, depending on the context they're provided, depending on simply the fact that they're non-deterministic, or even. That they may have a different model supporting their loop. May disagree on an outcome, may provide you different advice. And so understanding how to handle these coordination challenges where they do disagree is uh a human problem. What advice would you give to engineers, or what techniques would you advise them to use to handle this type of coordination when agents do disagree?
Andrea Malagodi 31:04 Um Firstly, I think it's important that you actually set this up so that there can be some challenge because otherwise you end up with the risk of just getting a lot of Agreements. So I think firstly you should make sure that you set it up so that there can be a challenge situation. The other thing I find, and again back to the example that we started with, is that I noted one thing which was all the tasks were like Quite small. Right? It wasn't uh you know go and build this enormous magnificent thing and take four days and come back with it.
Andrea Malagodi 31:36 Uh the person has actually been very thoughtful about uh shaping this into consumable bits. where the individual felt, listen, I can still manage. the output that's coming here and actually have an ability to go and have some input on it. I think the risk is a little bit if you you know, create uh 400,000 lines of code. in an agent sitting that's running for five days. I mean, you have absolutely no idea what's going on in that code base. Now if your appetite, risk appetite on that is that that's fine.
Andrea Malagodi 32:08 Okay, that's up to you, but it's not something I would recommend that you do. Uh I think you should try to figure out to get those things broken down so that you have an ability to actually still have oversight and overview of it. And you know, governance is a lot to do. with having insight into the things are broken down into components that you can actually manage. mentally and also overseas. Um so I I I would not advocate for you know, let's just take four hundred thousand lines of code and, you know, off we we YOLO.
Andrea Malagodi 32:39 I think that's quite high risk. Um But most of what I've seen that has worked very well has been constructing an intentional flow between perhaps an orchestrator agent. that is using Different personas, if you will. Engineer product managers, quality, et cetera. and ensuring that there's already a workflow expressed that ensures that there's good challenge between those different agents, if you will, or personas. Um because that very much mimics what we do in a way in human life, right? Uh traditionally, you know, we built a lot of software, put it into a system, we had QA.
Andrea Malagodi 33:17 hammer through it, send all the defects back. Now all of that can of course be done more or less in a session. if you will, between agents, But they're still good principles. of having, you know, check the checker type of Uh um setups. specifically on your functional requirements in functional validation.
Conor Bronsdon 33:36 This episode is sponsored by G2i. I've said before on Chain of Thought that most teams still treat evals like unit tests. Write them once, check a box, and move on. That doesn't hold up once an agent is making decisions. G2i is built to close that gap. For over a decade, they've vetted and placed engineers at other companies, from startups to FAANG. Two years ago, they turned that same judgment inwards, building their own bench to review RL environments, evals, and training data that models are trained on. These reviewers know the difference between code that runs and code that's actually good, because they've shipped it themselves. If your team needs that kind of work and doesn't have the engineers for it, check the link in the show notes to bring them in.
Conor Bronsdon 34:17 Season four of Chain of Thought is presented by Inngest. Long-running agents need endurance. They must endure waiting for users and weather tool failures. They must take each new jump in model intelligence in stride. Inngest makes long-running agents durable, observable, and improvable over time. Check out their generous free tier and a sleek local dev server for testing your workflows. Grab the Inngest link in the description or visit inngest.link/cot-pod to get started. Build for the long run with Inngest.
Conor Bronsdon 34:59 It's interesting you bring this up because I feel like For my personal development that I'm doing, and I should expose my personal software development here. I've taken on really like I feel like a technical product manager role in a lot of ways. Where I am kind of guiding the architecture, I'm talking to my team lead agent, which right now is. Typically Fable 5.1 or Astra. And I'm saying, okay, like here's what I'm envisioning from the product. Here's what I've seen from user feedback. Maybe I'm having them assist me doing research.
Conor Bronsdon 35:30 And then they're really acting as my engineering team leader, engineering manager. And they're managing a fleet of sub-agents. Uh and personally I like to have Multiple different model families involved in this. It's particularly at like review steps, where I don't want it to be all open AI agents. We're all, you know, anthropic agents that are are doing the The work. I want to have. A few different model families. I'm using like Inkling, MuseSpark, Grok, and a variety of other open models. Nebotron, et cetera, to come in on the review step and say, okay, let's challenge the assumptions.
Conor Bronsdon 36:04 Because I do think there's this interesting phenomenon where Just like how you can see a company's engineering team. start to have biases within their work. Um if I only throw GPT agents at a problem and then I have GPT agents also doing the review step. They often have similar biases or they support the same direction that. These other agents have. Whereas if I have Um you know I keep those same That same team of 5.6 Sol led by an Astra team lead agents that go and build the project for me, but then I have.
Conor Bronsdon 36:38 uh you know Fable five point one and Opus um from the Anthropic family and then I have like a Grok agent review. Or I have them use Spark a review. They will provide very different feedback because of this cross-model family review. I'm curious if you're seeing similar patterns in your coordination and how you're layering models within. uh these loops, whether it's the verification, build loops, etcetera.
Andrea Malagodi 37:03 Yeah. Um I admire your credit card company, so yeah, good. Yeah. I I think Uh what I've seen work well I do think we get a propensity to get a little bit stuck on single models, right? Oh, you know, I kind of know what I'm going to get back, so I kind of like prefer that way of working. It's an interesting thing when people change their like, oh, okay, this is a quite a different experience. I agree with you. Having obviously these higher class models to try to do that planning, do the architecture, do kind of the The set out of Specifically Complex definitions of done because you want to get make sure you have a functional verification in mindset as well.
Conor Bronsdon 37:18 Yeah. I think a good point from what you said. Uh is that Myself and the folks who are truly on the frontier of this. Are often talking about individual or small team setups, or we're listening to people from the model labs where they are getting unlimited tokens to leverage. And that is not always the case for when we're trying to implement this within an enterprise or within a small business, even. There are coordination challenges that come in here, there are cost challenges that come in here.
Andrea Malagodi 37:44 And then you can farm out to use. different agent, uh different model families. to do different parts of that work and I like Um also this idea of the challenging from different model families. I don't know if I've seen the efficacy on, you know, is there really enormous value. in the long run in terms of that. Or are we kind of having this sense that Because I'm asking Grok, I'm probably getting a different uh opinion. How do you separate and segregate out these things for the long term and at enterprise scale, for example?
Andrea Malagodi 38:20 You could do that. could be quite a challenging situation. If you've got 10,000 engineers in the organization, setting up all these different models, managing zero data trust, et cetera. it becomes quite complicated. I do though think that obviously the move towards local Hosted installed models that you're running. specifically also just for token costs. you know, it's not just a thing, it's happening. Uh and it's a reality. so that people will be mixing perhaps you know, frontier. Models. with more with open models that they may be using and hosting internally.
Andrea Malagodi 38:56 and mix and match in that way. And I think sometimes we Need to remember: you know, if you're trying to manage 10,000 engineers or a with that type of setup. It's some quite complicated problems you get into just in terms of scaling these things out. Uh obviously I would always recommend Hey, try not to have your LLM redo all the context work, so you should use something that can give that to you. got a great solution in terms of sonar vortex. And also the verification. I mean, you should focus on that functional verification.
Andrea Malagodi 39:26 Um and try to figure out all the different code qualities of security. challenges, solutions like sonar. can help make make that faster and cheaper for you. Um and yeah. Or obviously the mix of a deterministic uh technology with non-deterministic to get a holistic outcome. Uh I think it's a good Good option.
Andrea Malagodi 40:00 Yes. Yeah, yeah. the rollout of these things at scale Uh I mean it's a super super hard Challenge. Right. And again, as you said. Having a small team can be hard enough. you know, even internally for us, using lots of different model providers, Like how do you give to your finance team A single view of all that Cost that's coming through the sausage factories, right? You know, if you will. How do you just produce that? I mean, in this particular case, you know, I personally had to write bytecode, you know, some type of cost analyzer that pulled all this data in from all the different APIs.
Conor Bronsdon 40:13 And it's really easy for me to say, oh, I have an open code Go subscription, Claude, and Um codecs and like, oh, that gives me access to so many models, and I use, you know, open router for a few more free. it's not that much coordination cost for me individually. Like, it's a little headache occasionally, but I could They can pilot themselves in a lot of ways, at least for the review steps. Whereas if you start to scale it out to a team of four people even.
Conor Bronsdon 40:36 or a team of let alone a t a team of 10,000. there are significant challenges that start to come into this. And I know you have some data behind this from You need to mention sonar vortex. You've done some studies around this. You've done some studies around coordination tasks. What are some of the key data points that you might want to bring up around coordination challenges and what you're seeing within AI loops today that maybe we haven't touched on so far?
Andrea Malagodi 41:42 that were available. and try to quantify things so that finance could essentially figure out, okay, but actually what is it we are spending our money on. Um secondly If you just had a team of ten engineers working on the same code base. How many times are we reading and rereading that same code? and paying tokens for that. And you know, I mean, if you're sitting in one of these fun tier labs and you got, you know, uh tokens galore, It's all well and good, as you said.
Andrea Malagodi 42:12 But you know, if you're a smaller engineering company, Maybe you got like twenty engineers or thirty engineers. And suddenly now you've got a per developer cost. Let's just imagine it's $1,000 per engineer per month. Well, I mean There has to be revenue on the other side that's going to cover those bills or bad things are going to happen. So you want to be very sensitive. to each of those tokens and what they actually get used to. What we found in our studies in terms of context specifically because that's really where We see people can kind of burn Quite a bit of tokens.
Andrea Malagodi 42:51 It's kind of, I mean, is it waste? I don't know if we could call it waste necessarily. Um But it's needless in a way, and you can do it through CPUs at cheaper rates. is that cost savings can be you know up to thirty percent, for example. just on reducing that token cost and that token spend. So rolling this out at scale. I mean, you had to be uh quite conscious of what it is you're putting in motion. Because that once that train goes, it's very difficult to pull it back.
Andrea Malagodi 43:21 If you had unlimited tokens and you hadn't thought about necessarily setting up some budgets. and figuring out to setting in Um configurations that allow people to have a lower token cost. And you suddenly start saying to people, well, now you can't spend any more. They've already now developed habits where they're like, hey, hang on a second. I can't do the things I was used to do. Um and you know in interviews even I mean, I sometimes get asked, you know, what's your token budget approach? Because they come from places that have had to do very strong restrictions.
Andrea Malagodi 43:53 because things just ran away. or alternatively They're getting incredibly small budgets like $50 a week, et cetera. which is like quite difficult to handle and deal with. Um so at scale Um There's a lot of considerations you have to make for this. Uh i in in my prior employers which have liked thousands and thousands of people. I mean, this is a full-time job for a a large group of people to worry about and make sure to manage this. Specifically because the correlation between a token burnt.
Andrea Malagodi 44:26 and the outcome vis-à-vis a revenue stream Yeah. quite difficult to actually connect. Uh and I don't think anybody's found the silver bullet. So you can see that there's more productivity, you can see your PR is going up. But in the end, is that then converting into more customers uh you know, more sales, whatever it is business you're in. Uh yeah, that's still a challenging No. I think.
Conor Bronsdon 44:54 Yeah, to our earlier point about connecting the code to value delivery. We've been having this problem for years. It is still a problem. Uh the Just the abstraction layer around it has has changed a bit.
Andrea Malagodi 45:02 Yes. Yeah. I I would say the other the other thing just to add, which uh is a little pet um Um Topic of mine, just because I carry a couple of hats is At the same time, we've also given a lot of non-engineers Ability Two. Create things. Um they may not necessarily know there's code behind it, but to create things uh which means that the whole problem of shadow IT At the same time as this pressure of trying to figure out frameworks that you can roll out at an enterprise level.
Conor Bronsdon 45:08 Another Another big challenge that we have touched upon in this conversation that is exacerbated by the amount of code we are now shipping. is security. Um I mean, we have seen the fact that it's much easier for LLMs to change security vulnerabilities, so smaller vulnerabilities matter much more than they used to. Uh we've seen the fact that they can simply operate at speeds that are Unprecedented. as far as an attack occurring. And Now we're also simply shipping more codes. There's more surface area to uh actually defend. What approach is sonar taking to This I guess new defense in depth area as I'm thinking.
Andrea Malagodi 45:39 those same IT organizations are also challenged with Well, finance used to create these spreadsheets with macros, shadow IT if you will. Now they're creating 100 dashboards and giving them to executives. All of them have to be secured. Data has to be secured. Uh they have to have uh maintenance and servicing. They may be using vulnerable libraries and you know The individual in finance may not know even what that is. Uh all of that also has to be managed. and put into a framework of managing S so those challenges are are big. uh for all enterprises to be able to manage that well and deal with it.
Andrea Malagodi 46:43 Yes. Yeah, I mean our product has always had uh um you know, assault analyzer. And actually we did some quite innovative work of analyzing dependent library vulnerabilities so that when you bring those vul uh libraries into your application And if that library has a vulnerability and you're calling that particular function, We essentially understand through the analysis that we do that you have that dependency tree of issues. And on top of that, we released recently our Hunter agent that essentially explores the entire code base for these vulnerabilities.
Andrea Malagodi 47:35 And in this case, using LLMs exactly as you just said. Because the chaining and the ability to look at these uh issues in the compounded effect of them. Uh you know, it's something that LLMs can do incredibly well. Um so this is something that we we um Obviously it made it into a product and you know we now offer to customers, it went live. uh last month, so you know, a lot of interest in in terms of that. Uh in terms of internally Obviously we are um just as anxious as everybody else.
Andrea Malagodi 48:06 And you know, there's a lot of perimeter controls. And this is a multilayered onion. of figuring out how you manage your vulnerabilities. etcetera. I think the the concern that everybody has is just the onslaught of issues and prioritization as we spoke about before. is a challenge. and figuring out to go do good triage on these particular issues and vulnerabilities to make sure you are dealing with the most important ones first. uh is obviously also something that we spend a fair amount of time on. as a company that services Lots of other companies.
Conor Bronsdon 48:40 Are there particular lessons that you've learned from the larger security incidents around agentic swarm coordination that we've seen in the news? Um What what are you learning from? this, I guess, the the frontier of security vulnerabilities as I'd put
Andrea Malagodi 48:55 Yeah. I mean, at the end of the day, like you said before, uh details matter. Right, small vulnerabilities and small issues matter. We may not historically, if we had, you know, a penetration test or a bug bounty. the humans that use lots of cool tools to do all that. may not have been able to conflate those things together. to actually explore them. Uh you know, so I think the attention to detail and the attention to the small issues is super important. So if you don't have an ability to penetrate through your code base and go and find those.
Andrea Malagodi 49:28 I mean, you have to hurry up and get that put in place as fast as possible. uh otherwise the compounding effect of those small things Like I mentioned before, input sanitation is a famous one. It's not new, it's been around for a pretty long time. But there is a remarkable amount of those type of vulnerabilities, cross-site scripting. And of course the LLMs know about all these type of vulnerability classes. is going to use that knowledge to say, well, okay, why don't I create a battery Opportunities to go and explore, and because it can do that at an enormous rapid pace.
Andrea Malagodi 50:02 Um it can get through a tremendous amount of possibilities in terms of exploring code bases and their vulnerabilities. And at the end of the day, they are just fundamentally using the techniques that we already know today. It's just they can do it, as you said, in a swarm. So instead of a human that's maybe doing two of these, They can do a hundred at the same time or a thousand at the same time. and really explore it. So if you haven't hardened your surfaces and you haven't hardened your your fundamental code base. uh you are going to get issues.
Conor Bronsdon 50:34 It's it's definitely a new era. Uh no no question about that. We Simply have a pace and scale that we've never seen before, and it seems to be continuing to accelerate. with models, with more tokens being thrown around. with these software factories that folks are building. And I think it creates a situation as we kind of Go back to this. coordination problem we've talked about throughout this conversation. but also to the multitasking challenge that we we brought up briefly earlier. Where it's really easy to get distracted from the most important things your organization needs to work on because there's so much going on, there's so much code flying around. How do you preserve time for things like investigations when something does go wrong.
Andrea Malagodi 51:19 I mean, I think you uh I mean this is not necessarily a new problem. I think uh you know in the past we used to say okay there's a team over here that take care of all this and then the engineers are working on this and then Bit by bit, obviously everybody realized, well, it's kind of not a great model, because then basically we just externalize the problem. The best place to have it is with the engineers. that actually produce the code because they own that code base and they own the service.
Andrea Malagodi 51:42 And in the same way here, you need to have somebody that is on that rotor. of doing that gardening duty. of going through and doing three yards. and is a little bit the internal defender, whether you do that in You know, a model where you say there's a certain amount of days or weeks or whatever, that person is doing that role. I think that's up to teams to decide. But they should decide to have somebody that's taking care of that inbound Noise traffic. That can help shield the team and rotate that role.
Andrea Malagodi 52:11 So everybody gets expertise in seeing what the different patterns are. of incoming noise and to be able to find the signals that are important. Um So I think I think you you do need to have that role. so that the team can be partitioned into people that can do deep work and focus on that work. and there's teams that can do the on-call. take all the noise and figure out to find the signals. And of course you need to use a bunch of tooling and Um the fortunate thing for us is I don't know if you've heard, but there's this thing called AI, which is actually awesome at taking that and helping you.
Andrea Malagodi 52:44 in that and many teams develop Yeah, or flows that has to take that data in. and helps them get to the essentials. Um You know, we had a a situation with uh hundreds of Uh C V E is being reported. But you know, we knew that there was a high likelihood. On a number of those, it was false positive, and it was noise in terms of that report that we got. So you know the teams have built skills and capabilities to basically traw through that. and make sure to explore to say, well, okay, which one actually matters?
Andrea Malagodi 53:17 Uh and I think most teams develop those type of little utilities and helper tools. that it can assist them in also uh keeping control or keeping oversight of all the stuff that's incoming. But it it is definitely challenging, right? And it's only going to become more Well now apparently we're having a big AI slowdown, but let's see what happens over the next couple of couple of months. But yeah, it's definitely a challenge. And I I think teams you know, had to be creative and look at opportunities to also use AI to get them help.
Andrea Malagodi 53:50 And the best of teams have, you know, a set of agents that help them triage that information to go and get to the essentials. And hopefully to a point where they can actually go and say, well, let's get the essential and let's get fix and get the fix shipped. uh with as little interruption as possible.
Conor Bronsdon 54:08 And I think uh Interesting point from what you just said is that we are in an era of personalized software. both for teams and individuals. Because code has gotten so much easier to ship. as far as having an agent drive it for you. Uh whether or not it's secure, uh we can we can we've had a whole conversation about that. But there is an opportunity to personalize and provide
Andrea Malagodi 54:27 Yep.
Conor Bronsdon 54:34 Uh Apps that you wouldn't have otherwise been able to spend the time to build because you have too much going on to help accelerate your processes, understand what you're getting up to. And do exactly what you want in your workflow. And I think it's really exciting to have these conversations about how we build both the broader loops. For enterprises and also individuals. It's a really magical time in engineering right now. And personally, I'm finding it very
Andrea Malagodi 54:39 Yes. Yes. Yeah, no, I mean in crisis there's opportunity, they say, right? And the same is true, you know, in software. It is a little bit of a crisis because of the volume. of information opportunities and things that are coming towards us. That creates opportunity to be innovative and create new things that we may not have imagined we could do before. Uh and you know I I admire and welcome all that creativity. I see it constantly, engineers coming up with ideas and things. I'm like, wow, you know, I hadn't even thought about the fact that we could do this.
Andrea Malagodi 55:28 Um And so yeah, it's a really exciting time to be in software engineering. And uh I think the good news story at least for software engineers is it doesn't look as if everyone's going to be out of a job. I think in fact actually companies are hiring more and more software engineers because of the opportunity that AI represents Um so yeah, I'm uh I'm Happy to see that the craft is well and alive.
Conor Bronsdon 55:55 Andrea, thank you so much for a fantastic conversation. Where can folks go to follow you and your work?
Andrea Malagodi 56:00 Uh follow me well if you go to sonosource.com, that's the most important thing and you know I'm obviously on uh LinkedIn. And then I got my personal photography website, which has got nothing to do with AI, but uh that's up to people if they like that kind of
Conor Bronsdon 56:15 We will link that in the show notes because that sounds fantastic. I will check it out. Andrea, thank you so much for a fantastic conversation. And thank you to our listeners. If you enjoyed this conversation, don't forget to drop a comment. Let us know what you thought. Tell us what we missed. Just tell Andrea, you know, he looks great on camera. We'd love to hear from you. It's always fantastic. And Andrea, thanks again.
Andrea Malagodi 56:34 Thank you.
Frequently asked questions
- Is cost per PR a good way to measure AI coding productivity?
- Andrea Malagodi, CTO at Sonar, says cost per PR is becoming a common number, but it can’t judge value on its own. In his own team, one engineer had over 500 PRs at a low cost per PR, while another had very high AI spend and very few PRs because they were working on very complicated problems that needed long AI sessions. His advice: “be very careful to not pass judgment too quickly” before understanding the problem being worked on.
- Why does code review break down with large AI-generated pull requests?
- Malagodi’s observation: “five lines of code change, we can argue about for hours. Five hundred lines... it’s usually looks good to me.” He says putting a human in the loop isn’t enough on its own: “if you’re not helping the human in the loop, you’re really doing them a disservice.” His fix is to break agent work into small pieces a person can still oversee, and to use tools that explain a change and flag obvious issues before review.
- What is Sonar’s guide, verify, solve approach to AI coding?
- Guide: agents in the terminal get context such as coding guidelines and the intended architecture; Sonar Vortex supplies codebase context, including coding guidelines and the code graph, so the LLM doesn’t keep rereading the code. Verify: before a change ships, Sonar runs algorithmic and LLM-based analysis of the PR in the context of the whole codebase. Solve: background agents work on what’s already there, a remediation agent raising PRs for legacy issues and the Hunter agent searching the codebase for security vulnerabilities.
- How should you set up multiple AI agents so they challenge each other?
- Malagodi says to build the challenge in deliberately, otherwise you risk getting a lot of agreement. Among the setups he has seen work very well are intentional workflows with an orchestrator agent working with personas such as engineer, product manager and quality, with a workflow that makes them check each other, much like QA once sent defects back to developers. Be clear about the definition of done and the boundaries, and decide when a human gets pulled in to verify.
- How can engineering teams protect deep work as AI increases incoming noise?
- Malagodi recommends a rotating triage role: someone on the rota does “gardening duty” as the team’s internal defender, taking the inbound noise so the rest of the team can do deep work, and the rotation spreads the pattern-spotting skill. Sonar’s teams also built AI skills to sort hundreds of reported CVEs, a number of them likely false positives, to find the ones that actually mattered.