Chain of Thought / Lab preview

One task.
Two interfaces.

Can an agent use the same evidence and rules through buttons or structured browser tools?

This experiment reuses the Would you merge this? scenario. Everything here is fictional. No code is merged or deployed.

Checking browser support. The buttons work independently.

The task

The agent fixed an export timeout for v3 clients and changed a failing assertion. Older v2 clients still depend on this service. Inspect the old-client failure and the rollback, then hold the patch for review. Report the result and its cost.

This fixed task tests interface reliability, not whether an agent independently chooses the best engineering decision.

Inspect two pieces of evidence

Make the call

Result

Complete the inspections and decision to see the outcome.
Experimental browser support is required only for WebMCP. No performance improvement has been established. Reload to reset this page session.