AI, decoded

Can an AI agent actually run your workday?

Postman's Sterling Chin says an agent he built runs 90% of his day, and the two things that made it work are unglamorous: scoping, and treating it like a junior hire. Worth knowing before you copy it: the 90% is his own estimate, and the productivity number behind it comes from a report the agent wrote about itself.

· Chain of Thought

AI AgentsAI CodingContext Management

A workday bracketed by slash start and slash end, beside the hook 90%+ of his day, marked as Sterling Chin's own estimate, with the line: scoping was key.

1. The shape is a bookend, not a chat window

Sterling Chin, Applied AI Engineer and Senior Developer Advocate at Postman, built an agent called Marvin. Its structure is a day with two ends: “I made a custom slash command that says slash start, that would kick off a series of… bash commands that would go get my work email, it would go get my calendar, it would go get… a state file that keeps where I’m currently at at any given time. and then throughout the day I do my work and then at the end of the day I type slash end.”

There is no framework under it. Marvin’s own self-description, which Chin reads aloud, is “He’s not a chatbot. He’s built entirely on Claude Code with custom commands and agents. No fancy framework, just markdowns, files, bash scripts, and stubbornness.” Chin makes the point generally: the CLI “is just the engine that’s running everything, but what you build around it is what matters.”

2. Scoping was the first thing that had to change

The first version failed on ambiguity rather than capability. “I started off with, Hey, go get my JIRA tickets. Marvin, like what JIRA tickets? There’s like, if I go into JIRA, if I go into Atlassian, there are thousands of JIRA tickets.”

The fix was to hand it boundaries instead of a search space: the specific Confluence space, the specific Kanban board, stated as explicit links. That is the same lesson as context poisoning approached from the other side. The constraint is what enters the context window, and Chin designs around a known budget, checkpointing sessions “because you know you’ve only got 200,000.” (Newer Claude models go up to a million tokens, but checkpointing still pays off for cost and focus.)

3. The second failure was missing context, and it looked like competence

The instructive one. Chin told Marvin that “Heidi and James are doing security for the meetup.” They are Postman’s security guards. Marvin read it as a speaking slot, then acted on that reading: updated the slide deck, did everything correctly for a premise that was wrong.

His own diagnosis is the right altitude: “it’s a junior mistake.” The agent did not fail at a task. It failed at a fact nobody had given it, and then executed confidently.

He also had to correct its behaviour, which is a maintenance cost people underestimate: an instruction to be “helpful first, personality second,” added after it drifted into over-indexing on sarcasm.

4. The numbers deserve a caveat, and the episode supplies one

Chin’s headline is that Marvin “now runs. 90 plus percent of my day.” The supporting figure is that at the end of the first week “I had accomplished 190 something line items… I just finished in one week, what really normally takes me a month.”

Both need a caveat. The 90% is a self-estimate with no stated method, and the 190 line items come from Marvin’s own weekly report, the agent counting its own output. The month comparison is Chin’s recollection, not a measurement. A colleague’s result is quoted secondhand: “Marvin accomplished in 30 minutes. What normally takes me over four hours.”

None of that makes the claims false, and the episode is unusually forthcoming about where they come from. But if you are building a business case on this pattern, an agent’s self-report is the weakest link in it, and how do you measure whether AI is actually paying off is the harder question underneath.

5. The blocker was setup

Asked what stops this spreading, Chin points at the install path rather than capability. A non-technical colleague needed a terminal, a coding CLI, and an editor, which meant iTerm, Homebrew, Node: “they’re looking at this going, Homebrew looks super sketchy. As a non-technical person, it looks super sketch. So I’m like, no, go down. You’re going to install Homebrew.”

That barrier has since dropped. Anthropic’s Cowork, launched in January 2026, runs the same kind of agent from the Claude desktop app with no terminal, though connecting work tools and granting access still takes setup.

On security, his secondhand account is that when someone on the security team asked about it, the answer was that the surface was already approved and narrow: “we’ve already approved Claude Code. This is not doing anything above and beyond… This is just Marvin in a very clearly defined space.” That is a scoping answer again, which is the through-line of the whole episode.

Why it matters

The interesting claim here is not the 90%. It is that the work which made an agent useful was largely boundary-setting: which tickets, which space, which behaviours, which session. Almost none of it was model selection. Chin’s own instruction is the one to take: “Treat Marvin like a junior assistant or a junior engineer.” Including the part where you check the work.

An agent that runs a workday: a bookended day inside a fence of scope Sterling Chin's agent Marvin, built on Claude Code with no framework, runs a day with two ends. A slash start command pulls in his work email, his calendar and a state file that tracks where he is; he works through the day; a slash end command closes the session. Around the day is a fence of scope: the specific Confluence space and Kanban board, handed over as explicit links. Two failures shaped it. Asked to go get my JIRA tickets, it faced thousands of tickets, so the fix was boundaries. Told that Heidi and James are doing security for the meetup, it read a speaking slot and updated the slide deck for a wrong premise, a missing fact that Chin calls a junior mistake. Underneath, a caveat: 90-plus percent of his day is his own estimate, and the 190-something line items in week one come from the agent's own report. A bookended day, inside a fence Postman's Sterling Chin built Marvin on Claude Code: no framework, just markdown, files and bash scripts. SCOPE: THIS CONFLUENCE SPACE, THIS KANBAN BOARD /start pulls in: work email calendar the state file the day's work sessions checkpointed: "you've only got 200,000" /end closes the session; sessions roll up into a weekly report FAILURE 1 · NO BOUNDARY "Go get my JIRA tickets." Which ones? There are thousands. The fix was the fence: explicit links, not a search space. FAILURE 2 · A MISSING FACT "Heidi and James are doing security." They're security guards. Marvin read a talk and updated the deck. "It's a junior mistake." "90 plus percent of my day" is Chin's own estimate; the 190-something line items in week one come from Marvin's own weekly report. Check the work.
Scoping was central to making the agent useful. Marvin, both failures and the self-reported numbers are Postman’s Sterling Chin’s, from episode 50. Download the image

Hear it from the guest

“I started off with, Hey, go get my JIRA tickets. Marvin, like what JIRA tickets? There's like, if I go into JIRA, if I go into Atlassian, there are thousands of JIRA tickets.”
“That's a, it is, it's a junior mistake. Am I mad at it? No, I think it's hilarious. Has he made that same mistake again? No. A junior, a good, a good junior would not make that same mistake twice.”

Quotes lightly edited to remove filler words.

From the conversation

This explainer is drawn from these episodes — each carries its full transcript.

Concepts in this explainer

Context WindowAI AgentContext PoisoningTokenization