Windsor Harlow Start a conversation

Services / Cloud, DevOps & Infrastructure

Cloud infrastructure your own team can still operate.

Cost and operability designed in as requirements, not cleaned up after the invoice arrives.

Typical first engagement
6–14 weeks, fixed scope
Starts with
Architecture and cost review against your actual usage
Ends with
Terraform in your repo, runbooks, and on-call handover

The problem

What this practice is actually for.

The expensive cloud mistakes are architectural and made early. Default instance families, egress nobody modelled, a Kubernetes cluster running four services.

We design infrastructure for the team that inherits it: fewer moving parts, everything in Terraform, and a documented answer to what each alert means at 3am.

  • Everything reproducible from code — no hand-built resources in production
  • Cost modelled at design time and tagged for attribution from day one
  • Failure modes written down, with the recovery path tested rather than assumed
  • Migrations sequenced so there is always a way back

In the work

What Cloud code looks like when we write it.

   
Running 0.0s
pipeline main ● live run

Everything reproducible from code.

No hand-built resources in production, cost tags applied at creation, and a policy scan that fails the plan before anything reaches an account.

Scope a Cloud engagement

Capabilities

Cloud, DevOps & Infrastructure — the full stack.

13 capability groups

AWS
EC2, Lambda, ECS/EKS, RDS, S3, CloudFront, API Gateway, IAM, VPC design, CloudFormation & CDK
Google Cloud
Compute Engine, GKE, Cloud Functions, Cloud Run, BigQuery, Cloud Storage
Azure
Azure VMs, AKS, Azure Functions, Azure DevOps, Azure Active Directory
Infrastructure as code
Terraform, CloudFormation, AWS CDK, module design and state strategy
Containers & orchestration
Docker, Kubernetes on EKS / GKE / AKS, Helm, Istio, Linkerd, GitOps with ArgoCD and Flux, Operators and CRDs, HPA / VPA / KEDA autoscaling
CI/CD
GitHub Actions, GitLab CI, Jenkins
Serverless
Step Functions, EventBridge, Azure Logic Apps, GCP Workflows, Lambda@Edge, CloudFront Functions
Data infrastructure
Kafka / Amazon MSK, Redis / ElastiCache, Aurora, DynamoDB, data lakes on S3 + Glue + Athena, Snowflake, Redshift, BigQuery
Migrations
Lift-and-shift, replatforming and full re-architecture — from on-premise or between providers
Resilience
Multi-region architecture, active-active and active-passive failover, backup and recovery strategy, tested restores
Security
Vault, AWS Secrets Manager, Azure Key Vault, OPA, Checkov, tfsec, zero-trust network design, SSO, IAM policy design, SSL/TLS
Observability
Dynatrace, CloudWatch, APM and log aggregation, alerting that maps to runbooks
FinOps
Rightsizing, reserved instance and savings plan strategy, idle-resource audits, tagging, Kubecost showback and chargeback

Typical engagements

How this work usually starts.

Indicative scope and duration

Fixed scope

On-premise to cloud migration

Assessment, landing zone, and a sequenced cutover with rollback at every stage. Replatforming where it pays for itself, lift-and-shift where it does not.

10–16 weeks · Terraform · staged cutover

Advisory

Cloud cost and architecture audit

Where the spend actually goes, what is idle, what is over-provisioned, and which architectural decisions are generating recurring cost. Findings ranked by saving against effort.

2–3 weeks · written findings · savings model

Retainer

Platform operation and on-call

Ongoing operation of a live platform: patching, capacity, incident response and continuous cost control, with a named senior engineer.

Monthly · defined response times

What you get

Deliverables, every time.

01

Infrastructure as code

Your entire environment in Terraform, in your repository, reviewable and reproducible from a clean account.

02

Architecture decision records

Every significant choice with the alternatives considered, the trade-off accepted, and the conditions under which it should be revisited.

03

Runbooks and alerting

Alerts that map to documented procedures, so on-call means following a runbook rather than reverse-engineering a system.

04

Cost baseline

Tagged, attributable spend with a projection at your growth rate and the specific levers available to change it.

Questions

What clients ask before they commit.

Usually the one your team already knows, unless a specific requirement overrides that — a managed service you depend on, a data residency rule, or committed spend you have already negotiated. Migration between providers rarely pays for itself on cost grounds alone, and the retraining bill is routinely underestimated. We will say so even when it means a smaller engagement.

Often not. Kubernetes earns its operational cost when you have many services, several teams deploying independently, and a real need for portability. For four services and one team, ECS, Cloud Run or plain container hosting will cost less to build and far less to run. We will recommend the smaller thing when the smaller thing is correct.

That is the preferred arrangement. We embed in your sprint process, work in your repositories to your review standards, and document as we go so that capability stays with your team. Handover is a requirement of the engagement rather than a phase we run out of time for.

Have a Cloud problem worth a senior pair of eyes?

Tell us the system, the constraint, and what happens if it is not solved. A senior engineer replies within one business day.

Scope an engagement