Watch the video
Overview
Summary
In this technical discussion, host Matt Pocock interviews Lauren Tan (creator of pstack) to unpack the systems and philosophies enabling engineers to ship thousands of pull requests monthly using autonomous AI agents. Tan begins by recounting the origins of pstack following burnout at Meta, illustrating how acting as a manual ‘meat proxy’ between single agents and developer tools motivated a progression up the ‘trust ladder’. Pocock and Tan establish that as frontier models advance, the engineering bottleneck shifts from raw implementation to precise natural-language intent articulation and the systematic construction of constrained execution environments.
To reframe the popular ‘software factory’ narrative, Tan introduces the ‘Michelin Kitchen’ metaphor, conceptualizing the modern engineer as an executive chef who orchestrates tools, ingredients, and sous-chefs rather than cooking every dish manually. The dialogue dives deeply into the mechanics of agent reliability, highlighting that autonomous loops depend entirely on verification skills and deterministic CLI wrappers that execute mechanical tasks while preserving model judgment for high-level reasoning. Architectural guardrails—such as TypeScript narrowing, strict linting rules, and rigid directory structures like the internal ‘Dune’ framework—are presented as foundational mechanisms that prevent error sprawl before code is ever written.
The conversation culminates in an examination of multi-agent scaling topologies and continuous integration. Tan explains how marrying an outer context-gathering loop (Grokbot) with an inner execution loop (Cursor Projects) facilitates parallel task delegation and multi-agent fuzzing, allowing autonomous systems to safely test and merge their own PRs. Rather than performing synchronous manual code reviews on thousands of pull requests, engineering oversight shifts toward post-merge statistical sampling, telemetry auditing, and mining past chat transcripts to distill operational workflows into reusable Markdown skills.
Topics
Introduction and the Trust Ladder Concept
- Pocock introduces Tan following a viral presentation on shipping 2,500 monthly PRs to production.
- Tan recounts discovering the friction of micromanaging single agents manually after leaving Meta.
- The concept of climbing the ‘trust ladder’ emerges as engineers transition from manual proxies to supervisors of autonomous agent workflows.
Domain Expertise and Intent Formulation
- Domain expertise remains critical because clearly expressing intent and boundaries is the primary engineering bottleneck.
- High-density natural language and concepts (such as eliminating tautological tests) allow agents to reason more effectively.
- Subject-matter experts with baseline technical literacy gain leverage by directing agentic execution.
The ‘Michelin Kitchen’ Architectural Metaphor
- Tan prefers ‘Michelin Kitchen’ over ‘software factory’ to emphasize craft, software quality, and end-user experience.
- The engineer transitions from an individual line cook to an executive chef who designs environments, tools, and processes.
- Codebases, tools, and environment configurations serve as the modern ingredients for high-velocity software delivery.
Verification Infrastructure and Deterministic Tooling
- Verification skills provide agents with ‘eyes and hands’ to run, test, and debug code autonomously.
- Deterministic scripts and CLIs (Playwright/Chrome DevTools) handle mechanical operations, conserving context and LLM judgment for non-deterministic tasks.
- Automated verification loops enable iterative hill-climbing optimization without active human intervention.
Codebase Constraints and Guardrails
- Engineers must invest significant effort into ‘sharpening knives’ by configuring the agent’s operating environment.
- Strict lint rules and TypeScript type narrowing eliminate invalid coding paths and prevent agent regressions.
- The ‘Dune’ framework enforces feature-directory conventions to eliminate massive god files and streamline code discovery.
Scaling PR Velocity: Outer and Inner Loops
- High PR volume is sustained by separating concerns into an outer context loop (Grokbot) and an inner execution loop (Cursor Projects).
- Cloud-based coordinator agents act as chiefs of staff, receiving issue payloads and delegating sub-tasks across worker topologies.
- Buffer queues for code gardening enable holistic pattern recognition rather than isolated bug patching.
Post-Merge Quality Control and Autopilot Merging
- Quality assurance shifts from blocking pull request reviews to statistical sampling and auditing commit histories.
- Full autopilot spawns multi-agent verifiers that actively fuzz applications and patch regressions before merging.
- Formal programmatic verification and mathematical proofs (e.g., Lean, Bend) convert irreversible one-way architectural doors into safe two-way doors.
Authoring Custom Skills from Execution Transcripts
- Skills represent process definitions written in plain Markdown rather than rigid code scripts.
- Engineers are encouraged to mine past agent chat transcripts and intervention patterns to author specialized skills like pstack’s ‘recall’.
- Skills libraries like pstack and Pocock’s collections are complementary building blocks for custom workflows.
Topics
[00:00:00] Introduction and Context
Matt Pocock introduces guest Lauren Tan (poteto) to discuss using AI agents at scale, software factory workflows, and Tan’s viral talk about shipping thousands of PRs monthly.
- Matt introduces poteto following a viral talk on shipping 2,500 PRs in a single month.
- The conversation is set up as a deep-dive Q&A exploring agent trust, velocity, and skill creation.
“So what is your story of how you climbed the trust ladder and how did that work when like you got SpaceX and started climbing more and more?”
[00:01:58] Origins and Climbing the Trust Ladder
Lauren Tan shares the background of starting side projects after Meta, discovering the inefficiency of manually micromanaging single agents, and developing skills to transfer personal expertise into agents.
- Tan experienced burnout after Meta and started side projects using AI for coding.
- Observing the excessive time spent acting as a ‘meat proxy’ for a single agent sparked the idea for pstack and personal agent skills.
- Joining Cursor to work on agent performance revealed that frontier models frequently take shortcuts, necessitating rigorous verification.
“I was spending like so many hours just micromanaging one agent, right?” “I was sort of the meat proxy between my agent and Chrome DevTools.”
[00:06:23] Domain Expertise and Precision of Language
Pocock and Tan debate the role of domain expertise in an AI-assisted world, agreeing that deep domain knowledge and precise linguistic articulation are the true bottlenecks of engineering.
- Tan argues domain expertise is more important than ever because clear intent formulation is the main bottleneck.
- Matt and Lauren emphasize the precise value of language and concepts like eliminating tautological tests to anchor agent reasoning.
“it almost becomes like the bottleneck is no longer the agent, right? it becomes your ability to express your intent and your goals in a clear way”
[00:10:33] The Michelin Kitchen Metaphor
Tan contrasts the utilitarian concept of a ‘software factory’ with a ‘Michelin Kitchen,’ framing the engineer’s role as an executive chef organizing environments, tools, and processes.
- Tan prefers ‘Michelin Kitchen’ over ‘software factory’ to highlight craft, quality, and user experience.
- Engineers transition from manual solo cooks to executive chefs who lead, prepare environments, and orchestrate sous-chefs.
“how you set up your kitchen right and how you set up your skills your environment your codebase I think are ultimately the new ingredients that go into um building product”
[00:15:16] Verification Skills and Deterministic CLIs
Tan explains why verification is the cornerstone of autonomous agent loops and describes embedding deterministic CLI tools into skills to save context and enforce reliable execution.
- Verification skills give agents ‘eyes and hands’ to run, test, and debug code autonomously.
- Verification enables iterative optimization (‘hill climbing’) without human intervention.
- Deterministic scripts and CLIs (Playwright/Chrome DevTools) handle mechanical tasks while preserving LLM judgment for non-deterministic tasks.
“the single most important skill that should be in your toolkit is verification because without verification… there’s no way it can actually iterate”
[00:24:52] Codebase Constraints and Architectural Guardrails
The speakers discuss shaping codebases and architecture with strict lint rules, TypeScript narrowing, and rigid folder structures like the ‘Dune’ framework to eliminate bad code paths.
- Engineers must invest time ‘sharpening knives’ by engineering the environment itself.
- TypeScript type narrowing and category theory principles inspire constraining codebase boundaries.
- The internal ‘Dune’ framework utilizes conventional feature directories and strict linting to prevent god-file sprawl and agent errors.
“every time you see a mistake every time you see something that could be done better you think you step back and think how do I turn this into a lint rule?”
[00:33:01] Scaling PR Velocity: Connecting Inner and Outer Loops
Tan details the mechanics of scaling to 2,500 PRs per month by bridging an outer context loop (Grokbot) with an inner execution loop (Cursor Projects).
- High PR volume is achieved by opening ‘chain restaurants’ rather than manually typing prompts.
- The outer loop (Grokbot) ingests issues from Slack, linear, and email to supply real context.
- The inner loop (Cursor Projects) coordinates cloud-based chief of staff agents that assign sub-tasks to worker agents.
- Buffers and queues for code gardening allow humans or coordinator agents to recognize wider structural patterns before applying fixes.
“Cursor projects are my inner loop and grabbot is my outer loop.” “the big unlock for me for getting to 2,000 PRs is starting from the question and working backwards of how do I get to the point where my agent can merge its own code?”
[00:49:17] Post-Merge Review and Autopilot Verification
Tan explains post-merge verification, sampling techniques for code quality, and how automated fuzzing enables agents to safely merge code autonomously.
- Instead of blocking every PR with manual review, quality control relies on statistical sampling and auditing commit histories post-merge.
- Full autopilot spawns verifier agents to fuzz the app, detect regressions, and patch bugs autonomously before merging.
- Domains with high programmatic verifiability or formal mathematical proofs (e.g., Bend, Lean) allow one-way door changes to become safe two-way doors.
“it becomes more about sampling right and thinking about the processes” “Full autopilot is is something that you can do in in PAC and that will trigger off this very intense rigorous verification loop”
[00:58:40] Customizing Skills and Mining Past Transcripts
Pocock and Tan discuss how developers should approach skill libraries, advising users to mine their own chat transcripts to author custom, workflow-oriented skills.
- Skills from different creators (pstack and Pocock’s skills) are complementary process definitions written in Markdown.
- Engineers should mine their own prompt intervention histories and error transcripts to build personal skills (like pstack’s ‘recall’).
- Modern agent skills are trending toward concise workflow descriptions rather than rigid script implementations.
“at the end of the day a skill is just English or or language. It’s just markdown.” “The the past chats I I often say is like a a treasure trove of context because that, you know, it’s it’s like the process materialized”