Skip to content

AWS · DevOps · Production operations

Operating a Multi-Service Production Platform on AWS

Anonymised case note. The platform pre-dated my involvement; client and infrastructure identifiers are omitted.

Deployment and troubleshooting work across ECS Fargate, CloudFront, private service connectivity and two release pipelines—inside an existing AWS platform.

Context

The platform was already live. My role was to operate within it.

A media and digital-content platform served external users through a multi-service AWS production environment. My work crossed release pipelines, container runtime, traffic routing, service discovery and private database connectivity.

The architecture below shows the environment; the following sections state only the work I personally handled.

Environment I worked within

A simplified production landscape.

Conceptual production environment operated within; simplified and anonymised.

Components are grouped by operating surface. This is an environment inventory, not a verified network topology or dependency graph.

Frontend surface

Environment components

  • Users
  • CloudFront
  • S3 frontend / microfrontends

Release components

  • GitHub Actions CI/CD
  • Amazon S3
  • CloudFront

Application services

Environment components

  • Application Load Balancer
  • ECS Fargate services
  • Service-to-service communication
  • AWS Cloud Map / service discovery
  • Private Amazon DocumentDB

Release components

  • Azure DevOps CI/CD
  • Amazon ECR
  • ECS deployment workflow

Work I personally handled

The operational boundaries I worked across.

Release workflows
Handled container releases through Azure DevOps, ECR and ECS, and frontend releases through GitHub Actions, S3 and CloudFront.
Container runtime
Worked with the deployment and runtime behaviour of ECS Fargate services inside the existing environment.
Service connectivity
Investigated communication and reachability across services, dependencies and network boundaries.
Traffic and target health
Traced application reachability and target-health behaviour across the load-balanced request path.
Service discovery
Investigated AWS Cloud Map and service-discovery behaviour used by the platform.
Private data access
Troubleshot connectivity between application services and a private Amazon DocumentDB environment.

Diagnostic vignette

Tracing a service through the production path.

A focused example of how I investigated reachability and target-health behaviour across more than one layer.

Two Node.js services were associated with ports 8090 and 9191, with service URLs supplied through environment variables. When the 9191 service showed reachability symptoms, the investigation followed the full request path.

  1. ALB
  2. Target group
  3. ECS task
  4. Container
  5. Application

Evidence boundary: the investigation is confirmed; the root cause, configuration change and measured outcome are not, so they are not claimed here.

Diagnostic surface

What the investigation checked.

The investigation checked the boundaries below. They are areas examined, not individual causes asserted by this case note.

  • ECS task and service state
  • Expected application or service port
  • Load-balancer and target behaviour
  • Health-check behaviour
  • Network and security reachability
  • Application and service logs
  • Deployment state

Observe the symptom trace the request path verify each boundary isolate the mismatch scope the next change

What this demonstrates

Production incidents rarely respect team or service boundaries.

This case demonstrates practical work across application expectations, release workflows, container runtime, traffic routing, service discovery and private dependencies—without claiming authorship of the platform or an outcome the evidence cannot prove.

Next

Need a second set of eyes on an existing AWS environment?