Transcripts
← The collection

Building Codex with Tibo Sottiaux

Watch the video

Watch on The Pragmatic Engineer on YouTube Video not playing? Open it on YouTube.
01

Overview

Summary

In this extensive conversation, Tibo Sottiaux, lead of Core Products & Platform at OpenAI, traces the architectural and conceptual evolution of Codex alongside host Gergely Orosz. The narrative opens with Sottiaux’s background in applied mathematics, early startups, and his tenure at Google and DeepMind—where he co-developed an early internal LLM chat interface—contextualizing his transition to OpenAI to tightly couple research with product execution. Sottiaux recounts how Codex originated as an internal tool to accelerate Python research workflows before leadership directed it into a public product, merging with the AS3 (Autonomous Software Engineer) initiative. He explains key foundational decisions, including choosing Rust to enforce modular boundaries and static correctness, as well as making Codex open source and multi-model compatible to foster community trust and compete strictly on product merit rather than proprietary lock-in.

The dialogue pivots to the symbiotic co-design loop between models and the agent harness, exploring how the development paradigm itself is being reshuffled. Sottiaux conceptualizes the harness as temporary scaffolding and operational guardrails—pragmatic crutches that progressively shrink as frontier models acquire stronger autonomous reasoning and self-reflection capabilities. This evolution transforms the software development lifecycle across OpenAI: routine maintenance and sweeping architectural rewrites become low-cost operations, while automated AI verification achieves superhuman performance in spotting logic flaws and gating pull requests on security. As a result, human engineering reviews are transitioning away from syntactic checking toward alignment on high-level intent, system contracts, and domain invariants.

The discussion culminates in an examination of ‘The Merge’—the engineering effort to translate Codex’s local CLI capabilities into ChatGPT Work’s scalable cloud infrastructure using managed Kata containers. Sottiaux reflects on his own mobile-centric and voice-dictated productivity workflows, illustrating how frictionless agentic execution enables rapid weekend prototyping and broad contextual inquiry. Concluding the interview, Sottiaux advises aspiring AI engineers that deep curiosity, the ability to grok complex systems quickly, and acute product taste and user empathy remain the ultimate durable skills in an era of automated code generation.

Topics

Origins, Background, and the Genesis of Codex

  • Sottiaux’s early career spanned applied mathematics consulting, pharmaceutical supply chain optimization, and research tooling at Google and DeepMind.
  • At DeepMind, Sottiaux co-developed an internal LLM chat interface a year before ChatGPT, which went viral across the organization.
  • Sottiaux joined OpenAI due to its lean team culture and tightly coupled research and product cycles.
  • Codex originally started as an internal agent trained to accelerate OpenAI’s Python research codebase before merging with the AS3 initiative to become a public tool.

Architectural Choices: Rust and Open-Source Philosophy

  • Rust was selected for the Codex core to enforce robust security boundaries, static verification, scale, and separation between the agent and product UI.
  • Open sourcing the CLI and SDK enabled community contributions and dogfooding, though it introduced trade-offs like PR noise and competitors copying features in development.
  • Supporting non-OpenAI model providers prevents artificial vendor lock-in, gives users optionality, and forces OpenAI to compete strictly on merit.

Harness Mechanics and Harness-Model Co-Design

  • Codex defaults to local sandboxed execution while expanding into cloud Kata containers via ChatGPT Work for compute-heavy workloads.
  • The harness acts as temporary scaffolding ahead of model capabilities; developer prompt instructions shrink as reasoning models improve.
  • Engineering and research teams collaborate continuously to determine whether capabilities belong in harness logic or model pre-training.

Transforming the SDLC, Code Reviews, and Maintenance

  • New OpenAI engineers leverage Codex directly to query cross-organizational context across public Slack channels, Notion, and codebases.
  • Automated reasoning models provide superhuman code and security verification, serving as strict merge gates for pull requests.
  • Human code reviews are pivoting from line-by-line syntax verification to high-level alignment on intent, interfaces, and system invariants.
  • Routine maintenance, dependency upgrades, and full system re-architectures have become vastly cheaper and largely automated.

The ChatGPT Merge, Executive Workflows, and Career Advice

  • The Merge integrated local Codex agent workflows into ChatGPT’s scalable cloud infrastructure, using Codex itself as an internal documentarian.
  • The temporary ‘Work toggle’ is an interim step toward complete intelligence unification across ChatGPT.
  • Sottiaux conducts mobile-first executive workflows via voice dictation, custom skills, and rapid weekend prototyping.
  • Advice for engineers entering AI emphasizes cultivating deep curiosity, rapid system comprehension with agent tools, and strong product empathy and user clarity.
02

Topics

[00:00:00] Introduction to Codex and Tibo Sottiaux

The host introduces the episode, previewing topics like the origins of Codex, engineering decisions around Rust and open source, changes in code review, and the merger into ChatGPT.

  • Tibo Sottiaux leads Core Products & Platform at OpenAI, encompassing Codex.
  • Key themes of the episode include Rust architecture, open source dynamics, changing code review practices, and the integration of Codex into ChatGPT.

“Today we cover how Codex started and why it was built in Rust and made open source.”

[00:03:07] Early Career and Applied Mathematics Startup

Tibo discusses how growing up in rural Belgium sparked his interest in computers, leading to applied mathematics studies and an early startup focusing on supply chain optimization.

  • Moved to a remote village in Belgium as a child, finding learning and community through early computing and the internet.
  • Studied applied mathematics and graduated early while consulting on supply chain and complex mathematical problems.
  • Founded a startup utilizing classical optimization and Monte Carlo simulations for pharmaceutical supply chains and clinical trials.

“I was obsessed with applied mathematics is just really this idea of you have theoretical mathematics… and then there was like the real world”

[00:07:20] Engineering at Google and Early LLMs at DeepMind

Tibo recounts his work at Google London, learning from canceled web acceleration projects, working on Google Maps, and experimenting with internal LLM chat systems at DeepMind.

  • Worked on a mobile web acceleration initiative at Google Ads that was eventually cancelled for lack of Google-scale product-market fit.
  • Transitioned to DeepMind to build research tooling and infrastructure.
  • Co-developed an early internal LLM chat interface at DeepMind a year before ChatGPT launched, which went viral internally.

“It felt more than like a research project or like a research a project for researchers.”

[00:12:41] Transitioning to OpenAI

Tibo explains why he moved from DeepMind to OpenAI, drawn by its tight coupling of research and product, lean teams, and rapid execution culture.

  • Sought an environment where research and product were tightly integrated to deliver direct real-world impact.
  • Astonished to learn ChatGPT was operated by a nimble team of only around 20 engineers.
  • Joined OpenAI to work on research infrastructure and early reasoning model launches like o1-preview.

“I wanted to join a group where you know like all the parameters were sort of like considered together where you know research and product were like really co-designing.”

[00:15:19] The Genesis of Codex from Internal Tooling

Tibo traces the evolution of Codex from internal models trained to accelerate OpenAI’s Python codebase into a general product merging with the AS3 initiative.

  • Recognized the imperative to use reasoning models to accelerate internal OpenAI research and tooling.
  • Trained internal Python-focused coding agents to build infrastructure and assist researchers.
  • Leadership encouraged expanding internal tooling into a public-facing product, merging internal efforts with the AS3 (Autonomous Software Engineer) initiative.

“Greg was very adamant that you know we would we would not just focus on ourselves but we would also focus on benefiting uh the world”

[00:18:20] Architectural Rationale: Choosing Rust for Codex

Tibo explains why Codex chose Rust for its agent core despite models historically being less on-distribution for Rust than Python or TypeScript.

  • Prioritized a strict architectural separation between the core agent and product interfaces.
  • Rust provided correctness, static verification, performance at scale, and robust sandboxing boundaries.
  • A strict language boundary prevented sloppy code intertwining that could impede long-term innovation.

“Turns out it was quite clear that Rust as a language would actually be quite good for agents fairly quickly if we decided to put some effort into it.”

[00:21:15] Open Sourcing the CLI and SDK: Benefits and Trade-offs

Tibo explores the philosophy of releasing Codex open source, examining the benefits of community collaboration alongside the downsides of copied features and PR spam.

  • Believed a coding agent should be open source so the community can point the tool at itself to improve it.
  • Benefits include streamlined onboarding for new hires and high community energy.
  • Downsides include competitors copying unreleased features developed in the open and handling high volumes of low-quality contributions.

“There was something really cool about the idea of having the code open source because fundamentally what you’re building is you’re building a coding agent.”

[00:25:50] Supporting Multi-Model Providers

Tibo details why Codex supports non-OpenAI models, arguing that locking users in hinders product merit and community trust.

  • Coupling the open source harness to OpenAI models would have inevitably forced community forks.
  • Multi-model support provides user optionality and gives OpenAI direct feedback across competitive models.
  • Commits to winning users on product and model merit rather than proprietary lock-in.

“I want us to win users by having, you know, the best models, the most efficient models, the best product”

[00:32:09] Harness Execution: Local Sandboxing vs. Cloud VMs

The discussion covers how Codex executes commands securely inside local sandboxes and how execution is evolving into cloud-managed Kata containers.

  • Default execution runs locally within a secure sandbox, prompting users for escalated permissions.
  • Codex can execute inside managed cloud VMs using Kata containers via ChatGPT Work.
  • Cloud execution frees developers from local compute constraints as agent workloads scale.
  • Agents can automatically configure and sync cloud development environments, eliminating traditional setup friction.

“As models just get better and uh more capable… they can leverage so much more compute and many more resources than are available on your local machine”

[00:36:44] Harness and Model Co-Design

Tibo describes the iterative loop between the harness and the underlying models, showing how harness ‘crutches’ are phased out as model capabilities advance.

  • The harness acts as temporary scaffolding and guardrails ahead of model capabilities.
  • As models improve at reflection and autonomous verification, system developer instructions shrink.
  • Engineering and research teams collaborate to decide whether a capability belongs in harness logic or model training.

“The harness in a sense is always a little bit ahead of the model… you set it up with like a couple of crutches so that it can actually do the thing”

[00:41:19] Software Development Life Cycle Inside the Codex Team

Tibo describes the internal engineering workflows at OpenAI, highlighting ubiquitous Codex integration, radical documentation transparency, and high developer autonomy.

  • New engineers are encouraged to ask Codex first, as it indexes Slack, Notion, docs, and codebases.
  • Open document permissions and public Slack channels ground agents in organizational context.
  • Individual engineers are empowered to ship large changes directly to hundreds of millions of users with automated guardrails.
  • The long-term vision is a unified, proactive personal AGI controlled through voice and natural language.

“The thing that they hear the most about when they have a question is like, ‘Have you asked Codex?’”

[00:46:39] The Evolution of Code Reviews and Automated Verification

Tibo discusses how reasoning models have achieved superhuman performance in logic verification and security review, shifting human review toward validating high-level intent.

  • Early research produced code review models that perform multi-layer dependency analysis to uncover deep logic bugs.
  • Automated AI security reviews now strictly gate PR merges across OpenAI.
  • Human code reviews are transitioning from syntactic verification to alignment on intent and design invariants.
  • Clear interface boundaries allow the code inside the box to change dynamically without constant human oversight.

“When we benchmark them it’s like they’re like super human in code review and this is not just true for correctness. This is also true for security”

[00:52:09] Automating Maintenance and Lowering the Cost of Rearchitecture

Tibo analyzes how AI drastically cuts the tax of software maintenance, dependency upgrades, and full system re-architectures.

  • Routine maintenance like library version upgrades and patch application is becoming fully automated.
  • The cost and risk of major architectural rewrites are drastically reduced by agentic refactoring.
  • Solid software engineering principles like clean abstractions and invariants become even more essential with autonomous agents.

“So maintenance is really sort of like a tax that you pay over time just to keep things running. But where I think it changes is like a lot of it is just going to be automated.”

[00:56:43] Expanding Engineering Capabilities and Mindsets

Tibo and the host reflect on how AI accelerates iteration cycles, shifting the engineering craft from manual coding to rapid problem-solving.

  • Software goes through development lifecycles in days or weekends rather than months or years.
  • Viewing code as a problem-solving tool rather than an end in itself unlocks unprecedented velocity.
  • Eliminates painful coding roadblocks like dead-end refactors and compilation errors.
  • Long-running tasks no longer require artificial prompt scaffolding as frontier models sustain multi-day reasoning.

“It makes you it should make you a better engineer if you just really care about, you know, the outcome and the system working well.”

[01:02:30] The Merge: Integrating Codex into ChatGPT

Tibo outlines the technical and architectural challenges of bringing local Codex capabilities into ChatGPT’s scalable cloud infrastructure.

  • Required reconciling two distinct architectures: fully local CLI agent vs. managed cloud stack.
  • Implemented secure, powerful cloud VMs capable of arbitrary workloads (e.g., training models, 3D rendering) within ChatGPT Work.
  • Codex served as an internal ‘journalist’ documenting design debates and architectural decisions across the company.
  • The ‘Work toggle’ is an interim step toward complete unification of intelligence across ChatGPT.

“What we’re trying to build is like one unified product that gives you access uh to the same intelligence but in the way that you want to use it.”

[01:07:16] Tibo’s Personal Workflows with Codex and Mobile Work

Tibo shares his personal productivity system, relying heavily on voice dictation, ChatGPT Work on mobile, and weekend prototyping.

  • Conducts extensive executive and engineering workflows via voice dictation on mobile.
  • Configures custom instructions to generate structured reports, slide decks, and code explorations.
  • Queries production logs, team statuses, and feature sentiment in under 30 minutes.
  • Builds complete weekend prototypes to test and share product concepts with teams.

“It’s like any question I have I can get an answer to like you know within 30 minutes. And so that’s how I use it. I use it for everything.”

[01:10:44] Advice for Engineers and Closing Summary

Tibo offers advice for engineers entering AI, emphasizing curiosity, system comprehension, and user empathy, followed by the host’s concluding thoughts.

  • Cultivate deep curiosity and the ability to grok complex systems quickly using agents.
  • Develop clarity of thought, strong product taste, and deep alignment with user needs.
  • Host reflects on the demotivating versus empowering aspects of harness code obsolescence and AI observability.

“If you can’t explain what you’re trying to achieve, if you can’t explain your intent, if you don’t have a tie to a community… it’s going to be much harder to do great work.”

Source & file details

https://www.youtube.com/watch?v=sLSTM9znQNs

Visibility
public
Result status
completed
Source type
youtube
Caption type
automatic
Language
en-orig
Source duration
4462 seconds

Published files

  • Overview.md 5,275 bytes
    SHA-256 da93e12af4b4c9393469f1c0ec975651dbb0558514e76fdb4249e03e58740f84
  • Topics.md 12,037 bytes
    SHA-256 f0d60da503314dbb0056016e062d486a7383c6264e6258cce5014198c9195e20
  • Source.srt 123,020 bytes
    SHA-256 ff8120d1962e85552d26b5e42172ffb5f628eb5d0f83bdc215bed71468351165
  • transcript.srt 123,020 bytes
    SHA-256 ff8120d1962e85552d26b5e42172ffb5f628eb5d0f83bdc215bed71468351165

These notes are generated from the source and may contain errors. Refer to the original for context.