Thomas Howard

Why I'm Building an Engineering OS

I'm trying to find out how much of a production engineering organization can be turned into a system, and how far one engineer can take it.

September 2026 · Thomas Howard

I'm building a software company largely by myself. That creates an obvious problem: building the application is only a small part of operating a real software product. Someone has to design the architecture, implement features, review changes, test them, deploy them, monitor production, respond to incidents, patch infrastructure, manage security, investigate failures, control costs, maintain documentation, and decide what gets built next.

Historically, those responsibilities get distributed across an engineering organization. I don't have an engineering organization, so I've started building one. Not by pretending an AI coding agent is an engineering department, but by asking which responsibilities of an engineering organization can be encoded into software, agents, automation, policy, and platform infrastructure.

I call it the Engineering OS.

The problem isn't writing enough code

AI has made producing software dramatically easier. I can describe a feature, investigate an architecture, generate an implementation, create tests, modify Kubernetes resources, analyze failures, and work through unfamiliar code much faster than I could a few years ago. That's incredibly powerful, but it also exposes a misconception about software engineering: writing code was never the entire job.

Suppose an AI agent produces 5,000 lines of perfectly valid code while I'm asleep. Should those changes ship? I still need to know whether they followed the architecture, whether the right tests ran, whether the implementation introduced a security problem, whether infrastructure changed, what evidence proves the change works, who is allowed to deploy it, and what happens if production behaves differently than the test environment.

Making implementation faster doesn't eliminate those questions. It makes answering them more important. That's why I don't think of Engineering OS as a giant prompt that says "run my engineering department." I'm trying to build the operating system around software delivery.

Intent → Planning → Implementation → Review → Evidence → Deployment → Observation → Learning

Different agents and platform components can participate at different points in that lifecycle. An implementation agent might write code. Another system can independently review the result. Automated tests can produce evidence. Policy can determine whether that evidence is sufficient. GitOps can control how an approved artifact reaches Kubernetes. Observability can determine what happened afterward, and incident workflows can react when reality doesn't match expectations.

No single agent is the engineering organization. The system is.

Autonomy isn't the goal. Trust is.

It's tempting to measure AI engineering by how autonomous an agent is. Can it implement a ticket without help? Open a pull request? Deploy? Modify production? Those are interesting capabilities, but they're not the metric I care about most. I care about whether I can trust the engineering system to operate when I'm not watching it.

An agent capable of modifying production infrastructure at 3:00 AM isn't necessarily an advanced engineering system. It might just be an extremely efficient way to create an incident. The more autonomy I add, the more important constraints become: what an agent is authorized to change, what evidence must exist before work progresses, which actions are reversible, which decisions require human approval, and where final authority lives. These aren't theoretical questions anymore. They're architecture.

The goal isn't to remove myself from engineering. It's to remove myself from work that doesn't require me. I don't need to personally watch a successful deployment. I shouldn't need to notice that a pod is crash-looping, repeatedly check whether a dependency is healthy, or remind the system that a required test suite hasn't run. Machines are very good at that kind of work.

What I do want to own are consequential decisions: what gets built, which risks are acceptable, what architecture we're committing to, when the system should stop, when automation should escalate, and when the evidence is insufficient. The Engineering OS should make those decisions easier for me, not quietly make all of them on my behalf.

Desired state applies to more than containers

Kubernetes is part of this experiment for a reason. Its operating model fits the problem remarkably well: declare what you want, continuously observe reality, and reconcile the difference. That idea extends beyond containers. A release can have a desired state. So can a reliability requirement, an engineering policy, a synthetic check, or the evidence required before production.

Declare what you want, continuously observe reality, and reconcile the difference.

That's why I'm increasingly interested in controllers, platform APIs, policies, and automated evidence. They provide ways to move operational knowledge out of documentation and into systems that continuously enforce it. Documentation asks people to remember. Platforms can make the desired behavior the default.

Complexity doesn't disappear when you do this, and I don't think it should. Distributed systems are complicated. Production infrastructure, security, and reliability are complicated. Pretending otherwise often just pushes the complexity onto developers or operators. I'd rather concentrate necessary complexity inside the platform while keeping the interface to that complexity simple.

Here's what I want.
Here's what the system did.
Here's the evidence.
Here's what requires my decision.

That's increasingly what I want the Engineering OS to feel like from the outside, even if sophisticated controllers, agents, policies, deployment workflows, and observability systems exist underneath it.

Production is the experiment

I'm using BusyNow to explore much of this because I don't want Engineering OS to become an elaborate demo. Demos are easy to automate. Production is where the assumptions get exposed. Dependencies become unavailable. Networks get slow. Deployments fail. Credentials expire. Infrastructure drifts. APIs throttle requests. Agents misunderstand instructions. Tests pass while users still encounter problems.

An engineering system has to deal with that reality or I've just built an impressive CI pipeline. BusyNow forces the abstractions to have an actual job and gives me consequences when I get them wrong.

The long-term question is much more interesting to me than whether an AI agent can implement a feature.

Can one engineer build and operate a serious software product serving thousands, or eventually tens of thousands, of users by surrounding themselves with sufficiently capable agents, automation, and platform infrastructure?

I don't know. There are obvious reasons the answer might be no. Customer support could become the bottleneck. Product decisions could become the bottleneck. Security could require specialization. Operational complexity could eventually overwhelm the system. I don't want to assume the thesis is correct. I want to test it, and if I eventually reach a point where another engineer is required, I want to understand why.

How much organization can become a system?

That's what I'm going to document here. Not every commit or every new tool, but the engineering questions that emerge as I build it. How do you establish trust between autonomous agents? What should an AI agent actually be allowed to do in production? How do you prove automated work was completed correctly? How should architecture constrain autonomous implementation? Where should humans remain in the control loop? How do you design automation that fails safely?

I expect some of my assumptions to be wrong. I expect parts of the architecture to change, and I expect to build things that turn out to be unnecessary. That's part of the experiment.

The goal isn't to prove that one person can replace an engineering organization. The goal is to find out how much of an engineering organization can be made into a system, and then see what becomes possible when it is.