Token price is one input. An agent can make several model calls, query paid tools, retry failures, and require human review. Our recommendation is to choose an explicit accounting boundary and compare the same costs across candidate systems.
Worked example: 100 attempted tasks cost $20 in model and tool charges, and 80 succeed. Cost per attempt is $0.20; cost per success is $0.25. Both are valid measurements, but they answer different questions. Human review would need an additional cost estimate if it is part of the comparison.
In episode 64, Jiaona Zhang argues for measuring returned time and useful outcomes instead of rewarding token consumption. Conor’s Intercom write-up offers a case study of specializing one pipeline step. Neither gives a universal cost target: the task’s value and acceptable failure rate still matter.
Sources
- Episode 64: Jiaona Zhang on useful AI work — Primary interview motivating outcome-based accounting; the arithmetic example is illustrative.