AWS · DevOps · Production operations
Operating a Multi-Service Production Platform on AWS
Anonymised case note. The platform pre-dated my involvement; client and infrastructure identifiers are omitted.
Deployment and troubleshooting work across ECS Fargate, CloudFront, private service connectivity and two release pipelines—inside an existing AWS platform.
Context
The platform was already live. My role was to operate within it.
A media and digital-content platform served external users through a multi-service AWS production environment. My work crossed release pipelines, container runtime, traffic routing, service discovery and private database connectivity.
The architecture below shows the environment; the following sections state only the work I personally handled.
Environment I worked within
A simplified production landscape.
Components are grouped by operating surface. This is an environment inventory, not a verified network topology or dependency graph.
Frontend surface
Environment components
- Users
- CloudFront
- S3 frontend / microfrontends
Release components
- GitHub Actions CI/CD
- Amazon S3
- CloudFront
Application services
Environment components
- Application Load Balancer
- ECS Fargate services
- Service-to-service communication
- AWS Cloud Map / service discovery
- Private Amazon DocumentDB
Release components
- Azure DevOps CI/CD
- Amazon ECR
- ECS deployment workflow
Work I personally handled
The operational boundaries I worked across.
- Release workflows
- Handled container releases through Azure DevOps, ECR and ECS, and frontend releases through GitHub Actions, S3 and CloudFront.
- Container runtime
- Worked with the deployment and runtime behaviour of ECS Fargate services inside the existing environment.
- Service connectivity
- Investigated communication and reachability across services, dependencies and network boundaries.
- Traffic and target health
- Traced application reachability and target-health behaviour across the load-balanced request path.
- Service discovery
- Investigated AWS Cloud Map and service-discovery behaviour used by the platform.
- Private data access
- Troubleshot connectivity between application services and a private Amazon DocumentDB environment.
Diagnostic vignette
Tracing a service through the production path.
A focused example of how I investigated reachability and target-health behaviour across more than one layer.
Two Node.js services were associated with ports 8090 and 9191, with service URLs supplied through environment variables. When the 9191 service showed reachability symptoms, the investigation followed the full request path.
- ALB
- Target group
- ECS task
- Container
- Application
Evidence boundary: the investigation is confirmed; the root cause, configuration change and measured outcome are not, so they are not claimed here.
Diagnostic surface
What the investigation checked.
The investigation checked the boundaries below. They are areas examined, not individual causes asserted by this case note.
- ECS task and service state
- Expected application or service port
- Load-balancer and target behaviour
- Health-check behaviour
- Network and security reachability
- Application and service logs
- Deployment state
Observe the symptom trace the request path verify each boundary isolate the mismatch scope the next change
What this demonstrates
Production incidents rarely respect team or service boundaries.
This case demonstrates practical work across application expectations, release workflows, container runtime, traffic routing, service discovery and private dependencies—without claiming authorship of the platform or an outcome the evidence cannot prove.
Next