Every consultancy has a sales deck. Fewer have a refusal list. The refusal list is more useful, because it tells you what the team actually believes about production — not what it’s willing to say in front of a buyer.

This is ours. It’s short on purpose. The point of writing it down is to make the answer the same on day one and day ninety, regardless of who’s asking and how much pressure is in the room. If we wouldn’t ship it for a friend’s startup, we wouldn’t ship it for the largest enterprise we’d ever quote.

The point of writing the refusal list down is to make the answer the same on day one and day ninety.

The single-region production database

There is always a reason. The replica is “coming next quarter”. The team is “small enough that we’ll just restore from backup if it goes”. The region “has never gone down”.

Single-region databases for a workload anyone is paying for are not a cost-saving measure. They are a deferred outage with interest. The cost of a quiet weekend running a cross-region read replica is a fraction of the cost of a Tuesday morning where a control plane in eu-west-1 is unreachable and the recovery plan is a Slack message that says “trying RDS console again”.

If a workload is genuinely throwaway — internal tooling, a marketing microsite, a batch job that can be re-run — that’s a different conversation, and we’ll happily run it cheaply. But for anything customers depend on, multi-AZ is a default and multi-region is a deliberate decision the architecture document has to make explicitly. “We didn’t get to it” is not a decision.

Manual production access without a paved alternative

Every production estate has people who can SSH or use the console with admin rights. Sometimes that’s appropriate — break-glass exists for a reason. What we won’t ship is a system that requires that path for routine work.

If the way you deploy a config change in production is “Sarah pulls up CloudShell on her laptop”, Sarah is the platform. Sarah will leave, or get sick, or be on the train when the alert fires. The fix is not a runbook called “what to do when Sarah is on the train”. The fix is a paved road: a CI pipeline, a terraform apply behind review, a parameter store with proper IAM. Boring, traceable, repeatable.

We’ve seen estates where the IAM model is so complex that the paved road is slower than the manual one. That’s not an argument for the manual one. That’s a separate piece of work to fix the IAM, and we’ll usually quote for it before we quote for anything else.

Observability that only the platform team understands

Dashboards built by the people who built the platform, for the people who built the platform, are a category. The category is “vanity observability”. It looks great in a steering committee. It is useless at 02:00 when the on-call engineer is a backend developer who joined eight weeks ago.

Observability that earns its place answers questions a tired engineer would ask in plain English. Is the checkout flow degraded? For which tenant? Since when? What changed? If the dashboard requires you to already know the answer in order to phrase the query, it is not observability. It is a museum.

We’d refuse to hand over a platform whose alerts only one person can triage. The handover criterion is simple: the second engineer on the rota, the one who didn’t write any of it, can take an alert at 02:00 and reach a sensible action within ten minutes. If they can’t, the alert is wrong, the dashboard is wrong, or the runbook is wrong. Usually all three.

”We’ll harden it later”

Security work that is scheduled after the launch never happens at the resolution it needed. The window where it would have been cheap closes the moment the system has users. After that, every change is a migration, every migration is a risk, and the risk argument always wins against the security argument because the security argument is statistical and the risk argument is concrete.

Hardening goes in during build. Not “security review at the end”. Not “pen test in Q3”. The IAM least-privilege, the secret management, the network segmentation, the key rotation, the audit logging — these are non-negotiable line items in the build sprint, sized and demoed live like everything else.

This is unglamorous. It also means the security team in the client estate isn’t the last to find out that the new system exists. They’ve reviewed the design. They’ve seen the Terraform. The penetration test confirms what the design said it would, instead of discovering it.

A migration plan with no rollback

If you cannot answer “what do we do at 03:00 if this is wrong” with something more specific than “open a ticket”, the cutover is not ready. We’ve seen migration plans that depend on a forward-only DNS change because the team didn’t want to spend two days building a route reversal. Two days is the trade. Take the trade.

Rollback isn’t a backup plan. It’s the plan you actually use when the smoke test goes red. Every wave we sequence has a rollback path that has been tested in pre-production with the same tooling and the same people who’ll execute it on the night. If it hasn’t been rehearsed, it isn’t a rollback. It’s a wish.

The corollary is that some migrations can’t have a clean rollback — usually because data has moved and the source has been catching up against a writable target. Those are not the migrations to do at 03:00 on a Saturday. Those are the migrations to plan with a freeze window, a written go/no-go, and a board-level sign-off on the freeze.

The “AI feature” with no data layer

We’ve been called in more than once in the last while to add a chat-shaped feature on top of a system whose own data engineers can’t reliably answer questions about. The model layer is a distraction. The actual problem is that the source data is duplicated across four systems, none of which agree, and the only person who knows which one is the source of truth left in 2022.

The refusal isn’t “no AI”. It’s “no AI on top of a data layer we wouldn’t trust for a basic operational dashboard”. The plumbing comes first: ingest, lineage, governance, retrieval. Once the system can answer the simple questions deterministically, the model can answer the harder ones probabilistically — and you can tell which is which when something goes wrong. The other order produces a chatbot whose hallucinations are indistinguishable from the underlying data being a mess.

”Fast on a fresh account”

Demos on greenfield AWS or Azure tenants are not engineering. They are theatre. The system you have to ship into is the one with seven-year-old service control policies, a landing zone that was correct in 2019, and a network team that owns the transit gateway and isn’t on this Slack.

What we’d refuse to ship is a reference architecture that only works in the demo account. The diagnose phase exists specifically to read the estate as it is — not as the marketing diagram suggests. The design is sized against what is actually deployable, with explicit notes on what the existing controls require us to change, what we’d recommend changing later, and what we’d advise leaving exactly as-is.

Handovers without a rota

The work isn’t finished when the system is in production. It’s finished when someone other than us is on call and the alert routes to a human who can act on it. Until then, we’re the ones holding it, even if the contract says otherwise.

We refuse the kind of handover that’s a Confluence page and a calendar invite. The handover we ship is a rota with named people, runbooks they’ve actually executed under supervision, an incident review they’ve led, and a thirty-day shadow period where they hold the pager and we shadow. We leave when the rota tells us to.

Why write any of this down

Refusal lists are not for clients. They’re for the team. Every project compresses the values of the people who built it, and a project under deadline pressure compresses them hard. Writing the refusals down — and revisiting them when the room is calm — is what stops the compression from quietly producing a system we wouldn’t recognise as ours.

It’s also a useful filter at the start of the engagement. If the refusal list is incompatible with what the buyer wants, the time to find that out is week one, not month three. The work we want is the work where the refusal list isn’t controversial. There’s plenty of it.

The work we want is the work where the refusal list isn’t controversial. hello@altostack.io.