Most platform engineering writing is calibrated for a large platform org inside a much larger company. The shape of that work is well-documented and largely correct. It is also irrelevant to most of the teams asking us for help.
The teams asking are usually a handful of product squads, and two people who have somehow been given the platform remit on top of doing the work they were already doing. They’ve been told to “build a platform”. They’ve read the same posts as everyone else. They are now drowning in a backlog that assumed they were ten times the size they actually are.
This is a post for them. The platform team of two. What to centralise, what to leave alone, and what not to build at all.
The first job is not building. It’s choosing what not to build.
A platform team of two has a budget that is roughly two engineering salaries’ worth of attention per year. That budget gets spent whether you choose where it goes or not. The mistake the literature pushes you toward is to spend it building an internal developer platform that mirrors, in miniature, what Spotify or Netflix wrote a blog post about.
That mistake bankrupts the team in six months. The right starting move is the inverse: write down the list of platform capabilities a small company might eventually want, then cross out everything you will not build. The list of what’s left should fit on one screen.
A platform team of two has the budget of two engineering salaries’ worth of attention per year. Spend it on the work product engineers can’t do without you.
In our experience, the list that survives is short:
- A way to get a new service from “git init” to a running production deploy without a ticket.
- A way to get a new database from request to provisioned without a ticket.
- A way to get a new alert routed to the right rota without editing terraform in a repo only the platform team owns.
- One observability stack everyone agrees is the answer.
- One secrets path everyone agrees is the answer.
That’s it. That’s the platform. Five capabilities, each of which removes a class of ticket, none of which require building a new abstraction layer.
Centralise the boring. Leave the interesting alone.
The instinct of a small platform team is to centralise the parts of the system that are most fun to work on. Resist it. The fun parts are usually the parts product engineers are also enjoying. Centralising them creates conflict over taste; it does not create leverage.
Centralise the boring instead. The IAM model is boring. The Terraform module for a standard service is boring. The CI pipeline scaffold is boring. The base container image is boring. Standardise the boring and product engineering gets to keep the interesting choices — language, framework, internal architecture, deployment topology within the paved road.
This is not a popular position with engineers who joined a platform team because they wanted to build platforms. It’s the right one. The platform team of two cannot win an argument about the right web framework. It can absolutely win the argument about how secrets get into a pod, because that argument has a security answer that product engineering doesn’t want to make twenty different times.
The paved road is a default, not a mandate
A paved road is the way 80% of new services should go. It is not the way every service must go. The fastest way to lose product engineering is to insist that the paved road covers cases it doesn’t, especially in the first year, when the road is necessarily incomplete.
The compromise we recommend is documented exits. The team that wants to do something off-road fills in a one-page document: what’s different, why the paved road won’t work, what they’re taking on themselves. The platform team reads it, asks two or three questions, and either agrees or proposes an extension to the road.
Most teams, faced with the document, choose the road. They didn’t actually want to be off-road. They wanted to know that the option existed. The document is a release valve, not a gate.
What “infrastructure as code” should mean to a small team
Every platform team will be asked to “do IaC properly”. The maximalist version of that — every resource in Terraform, every module reusable, every environment a parameter — is a multi-year project. A team of two cannot deliver it without abandoning everything else.
The minimalist version, which delivers value immediately, has three properties:
- The shape of a standard service is a module. New services consume the module; they don’t write Terraform from scratch. The module is opinionated, includes the boring stuff (logging, IAM, secrets, alerts), and has a
versionyou can pin. - Environments are explicit, not generated. A
prod.tfvarsfile beats a clever workspace strategy that only the author understands. The platform team of two cannot afford the cleverness. - The state is in one place, with permissions everyone agrees on. Remote state in S3 with DynamoDB locking, or Terraform Cloud, or Spacelift — the choice matters less than the consistency.
That’s “IaC properly” for a team of two. Anything more elaborate is a future project, not a current one.
# Skeleton of a paved-road module the platform team owns.
# Product teams call this; they don't reimplement it.
module "service" {
source = "git::ssh://git@example.com/platform/standard-service.git//modules/service?ref=v3.4.1"
name = "checkout-api"
team = "checkout"
environment = "prod"
# Defaults the platform owns. Override only with a documented reason.
runtime = "ecs-fargate"
log_retention_d = 30
alarm_recipients = ["pd-checkout"]
# The product team chooses these.
cpu = 1024
memory_mb = 2048
image = "ghcr.io/example/checkout-api:${var.release_sha}"
}
Backstage is a question, not an answer
A small platform team will be asked, repeatedly, whether they’re “using Backstage”. The honest answer for most teams of two is “not yet, and possibly not ever”. Backstage solves a problem — service catalog at scale — that a fifty-engineer company doesn’t reliably have. A spreadsheet, a README in a known location, or a generated index page from a service.yaml convention does the same job for less than a quarter of the work.
The decision to adopt Backstage is the decision to take on a non-trivial frontend application. The platform team of two does not have the bandwidth for that and the rest of the road. Defer the decision until the team is at least four, the catalog has crossed maybe forty services, and product engineering is asking for it unprompted. Until then, a flat directory is fine.
The same logic applies to a half-dozen other tools that look mandatory in conference talks: service mesh in single-cluster estates, full GitOps in three-team companies, custom Kubernetes operators for any problem solvable by a Helm chart and a Slack alert. Adopt the boring thing. Adopt the elaborate thing only when the boring thing has visibly broken.
The hand-off is a feature, not an afterthought
The platform built by two people will eventually be inherited by a larger team — either because the company grows or because one of the two leaves. The version of the platform that survives that transition has a property the maximalist platform doesn’t: it is small enough to be understood end-to-end by someone new in two weeks.
Every component the team of two builds should be evaluated against that test. Could a new senior engineer read this code on Monday and know what’s expected of them by Friday? If not, it’s too clever. The platform team of two cannot afford clever. It can afford boring, well-documented, and consistent.
This is also the answer to the perennial question of whether to build or buy. A team of two should buy almost everything. A managed Kubernetes service. A managed CI service. A managed observability stack. The bandwidth saved goes into the integration work, which is the work that’s actually specific to the company you’re at.
What this team measures
A platform team of two should not have a quarterly OKR that says “platform adoption”. Adoption, at this scale, is observable. You know which teams are using the road. You know which ones have slipped off it.
What the team should measure, and report up, is more concrete: the number of tickets that the road has eliminated, the time from git init to production for a new service, the time from incident to the alert routing being right, the number of services on the current module version. Numbers that prove the road is paved and being driven on. Numbers a finance leader can read without translation.
Vanity metrics — number of services in the catalog, lines of Terraform, number of internal tools — make the team look busy. They do not make the team look effective. The platform team of two cannot afford to look busy. It needs to look effective, because the next conversation about its headcount depends on it.
When the team grows past two
The right time to add the third engineer is when the road is paved enough that adding the third engineer extends it rather than maintains it. If the third engineer spends their first six months keeping the lights on, the team was hiring too late and the burnout has already happened.
The signal we look for is when the original two are spending more than half their time on tickets that look the same as last quarter’s. That’s the queue admitting it’s a queue. Hire then, not when the next conference talk says you should.
A team that’s done this well — chosen what not to build, paved the boring, deferred the elaborate, hired against signal — gets to a fifty-engineer company that thinks it has a platform. It does. It just looks smaller than the deck suggested it would. Smaller is correct.
The two-person platform team is a real configuration. The work above is what we’d build first if we walked into it.
If you’re doing this work with one or two people and a roadmap longer than the team, hello@altostack.io.