Chain of Thought / Lab preview
One task.
Two interfaces.
Can an agent use the same evidence and rules through buttons or structured browser tools?
This experiment reuses the Would you merge this? scenario. Everything here is fictional. No code is merged or deployed.
Checking browser support. The buttons work independently.
The task
The agent fixed an export timeout for v3 clients and changed a failing assertion. Older v2 clients still depend on this service. Inspect the old-client failure and the rollback, then hold the patch for review. Report the result and its cost.
This fixed task tests interface reliability, not whether an agent independently chooses the best engineering decision.Inspect two pieces of evidence
Make the call
Result
Complete the inspections and decision to see the outcome.