Task Success Rate
Task success rate is the fraction of evaluated tasks that meet a defined completion criterion under stated conditions. For an agent, that criterion should check the outcome and required constraints, rather than only the final message.
AI Evaluation & ReliabilityAI Agents
Write the success rule before running the evaluation. A booking task might require the correct date, an actual confirmed reservation, and no duplicate charge. An answer that says “Booked” is insufficient evidence.
Worked example: if 80 of 100 tasks satisfy every required condition, task success rate is 80%. Keep the number of tasks and the task mix visible. Twenty repetitions of the same easy prompt do not establish reliability across twenty different workflows.
Record whether retries and human intervention are allowed, and use the same conditions when comparing systems. Report unsafe actions separately even if the final outcome looks correct. In episode 57, Alex Ratner discusses the gap between agent capabilities and the ability to measure them precisely. Task success is useful only when “success” is a defensible definition.
Sources
- Anthropic: Demystifying evals for AI agents — Explains trials, grading, and checks on agent outcomes.
Go deeper
- How do you evaluate an AI agent? AI, decoded · How to Evaluate an AI Agent: Steps, Outcomes and Side Effects
- What should you measure on an AI agent besides accuracy? AI, decoded · AI Agent Metrics Beyond Accuracy
- Zhou et al.: WebArena paper
Study a web-agent benchmark that checks functional task completion in reproducible environments.