Speed of thought

LLMs—with the right guardrails and tools—can safely automate the design, implementation, and verification of production-ready software. What’s left for humans is to understand the problem and steer the machine toward a solution. This is production-ready software development at the speed of thought.

Background

After six years of research, four years working with LLMs, almost a year of agentic coding in a production-ready environment, I found that:

LLMs do the heavy lifting. They are indispensable when creating production-ready software.

Research

In autumn 2022, when ChatGPT came out, I started a journey to find a method for creating likely-correct software through rapid iterations.

Last year, I presented a study on how to use a mix of formal and semi-formal methods to achieve likely-correct understanding and design—the first two phases of the software development life cycle (SDLC).

This year I cover the next two phases: implementation and verification—again—the likely-correct way. The remaining phases, deployment and maintenance, are not relevant in this context.

The goal of these studies is to find out what is possible in terms of correctness, and at what cost, in software development. To find a pragmatic approach to LLMs, to discover where they are truly useful and what we can do to make them truly useful. To eliminate the unpleasant surprises they often throw at us.

The final goal is to create better software faster.

Production-ready code, faster

There is no exact definition of what production-ready software is or how to create it. There is broad agreement that such software is:

You can see an example of production-ready code in A likely-correct list, a prequel to this current study.

The study presents:

The study concludes:

And in a future stage, it should provide AI/ML practitioners with insights to combine generative AI’s speed with the rigor of these new techniques to produce better software, faster.

Now let’s pick up from here.

In January 2026, after the Claude Code hype settled down, I decided to start the implementation of the production-ready version of the Olog editor—a project where the MVP received positive feedback and in which I saw potential for further investment. In my workflow, ologs ensure a likely-correct understanding of a problem domain.

The idea was to vibecode the new web app with strong guardrails.

I had on hand a strong set of React, Next.js and Typescript templates created in Silicon Valley environments and battle-tested in production, ensuring that the written code was of high quality.

I had some strong opinions about an ideal code architecture. Organizing and maintaining a codebase is a recurring top pain point for front-end developers and I had spent endless effort figuring it out.

If it’s a pain point for humans it will be a pain point for LLMs. I asked the LLMs for help, and with three American and two Chinese models, we came up with such a solid production-ready architecture that they said they had never seen something similar before. Lol 😀

In the end, we’ve put together a generic multi-dimensional guardrail system in which:

With this toolset—my best knowledge ever—I started vibecoding the features one by one.

A detailed workflow document would instruct the LLM how to implement a feature step by step across the architectural layers. My job was easy:

We—the junior dev team provided by the LLM and I—were flying high. The speed was right. The codebase looked promising. The investment in the code architecture and the templates paid off. Even Simon Willison praised the results.

On the other hand, the product experience was strange. Subtle user interface errors, messy business logic, once-fixed problems resurfacing, code smells, inconsistency, duplicated code, ghost code. I had a bad feeling—We were not ready for production.

Also the Claude experience turned out to be awful. As the problems within the codebase grew that intellectual, tortured genius became more stubborn, over-refusing and always lecturing.

What’s next?

The problem was that while the primary focus—implementing the specs—went almost flawlessly, the secondary focus—following the architectural layers and coding rules—drifted at a subtle level while it was looking good on the surface.

To fix this we’ve started a second iteration now focusing on the rows of the guardrail matrix, on the non-user-facing coding and organizational rules that make a product robust.

We also chose a much better harness (Opencode) and a coding agent with a dedicated engineering mindset (Deepseek Flash) at a fraction of the price and hassle.

Also changed the rules of the game with LLMs: Instruct — Never ask — Always verify.

The result is more than promising.

After two perpendicular iterations and an overarching, integrating error-management implementation I’ve got production-ready code: it follows the specs, it follows the rules, smells good, looks good, it is fully tested and easy to iterate on.

I can affirm that the implementation and verification SDLC phases are safely doable with LLMs when the right guardrails, coding agents, and human supervisors are in place.

Why LLMs for I+V?

Better production-ready code, faster

While I was figuring out the magic formula ...

SDLC + LLM = U + LikelyCorrect(D) + SafelyAutomated(I + V)

... others reached the same or even better conclusions and created useful tools.

The age-old practice of Contract programming (Design by contract) took off in a new form—Vericode—and enhanced BDD to produce formally verified code.

A React/Typescript implementation, which I’ll use in my next projects, comes with this thesis:

  1. Humans define what should be built
  2. AI handles implementation
  3. Machines guarantee correctness

Meanwhile, Shopify went even further.

They’ve managed to get the agents write specs (although from an existing production-ready app and its codebase) and implementation plans; then execute and verify these plans.

Before you get too excited I should fill in the details about the invisible and hard work behind the scenes that makes that possible. According to Shopify:

Production-ready software at the speed of thought

In March 2024, the CIA presented a report on the promise and peril of artificial intelligence concluding “Generative AI is neither quite so wondrous nor quite so bleak”.

In autumn 2026, I can confirm this. You cannot change the world with a single prompt, but you can safely automate most of the software development process.

Today in the SDLC+LLM process D+I+V are—formally and semi-formally—solved.

Now agentic software development reduces to these steps:

  1. You start by understanding the problem domain
  2. (AI helps you) by sketching out the ontology and taxonomy in an Olog editor
  3. In the Stately FSM editor, you model (together) how these parts behave and interact
  4. When the solution to the problem is clear you design the user interface and experience by using the concept design DSL
  5. Finally you let it loose: Using the guardrails and the vericoding techniques the AI creates the production-ready code

This is software development at the speed of thought.

Resources

  1. Likely correct software — Osequi, 2023
  2. Rapid iteration in software development — Osequi, 2023
  3. A likely-correct list — Osequi, 2025
  4. Future of Intelligence: “The Incalculable Element”: The Promise and Peril of Artificial Intelligence — CIA, 2024
  5. JavaScript Pain Points — Stack Overflow, State Of Javascript, 2023
  6. To React with best practices — Metamn, 2018-2021
  7. A Hacker News comment thread on Software Architecture Guide — Simon Willison, 2026
  8. A Hacker News comment thread on: Ask HN: Why is the HN crowd so anti-AI? — The author, 2026
  9. Design by contract — Wikipedia
  10. From Vibecoding to Vericoding: A Gradient, Not a Jump — Scidonia, 2026
  11. Midspiral — 2026
  12. Migrating Shop app from React Native to native — Shopify, September 2026
  13. Rails World 2026 Opening Keynote — DHH, September 2026
  14. Helix: The internal tool powering our Shopify app's native migration — Shopify, September 2026
  15. Stately — 2026

About the author

The author holds a degree in mathematics and computer science. He is a self-taught UX/UI designer with works featured in online galleries.

Recently, he has been running a research and development studio specializing in software correctness and rapid iteration, providing consulting services to companies.

Credits

This document was created using Notion and published using a modified version of Tufte CSS.

It was checked with the Google Docs spell checker and the write-good naive linter for English prose.

Then LLMs corrected the non-native English words and constructions—I’m inherently prone to using—to your delight.