aws

The 5 ECS Decisions That Waste 6 Weeks (And What to Pick Instead)

Opinionated picks for deploying Python apps on ECS: Fargate vs EC2, service discovery, CI/CD, secrets, and monitoring, from real consulting work.

I’ve been helping Python teams deploy to AWS for the past 2 years. The pattern repeats: a team has a working FastAPI or Django app running fine on a laptop with docker-compose up, and then someone says “let’s put this in ECS.” Six weeks later, they’re still arguing about Fargate versus EC2.

The problem isn’t that ECS is hard. The problem is that teams treat infrastructure decisions like they’re permanent. They’re not.

Last year I worked with a startup that spent 5 weeks evaluating container orchestration options. Five weeks. They had a working app and paying customers waiting, but the engineering team was stuck in a loop of “what if we need to scale?” and “shouldn’t we future-proof this?”

They launched on Fargate. It took 3 days once they stopped debating.

The wrong choice costs a few hundred dollars a month. The debate over the right one costs weeks of engineering time.

Here are the 5 decisions that waste the most time, and what I tell every client to pick.

Fargate vs EC2: pick Fargate

This decision wastes more time than all the others combined.

EC2 looks cheaper on paper. Run the numbers and you can show that at 50 containers you’ll save $400/month over Fargate. The finance person gets excited, someone mentions spot instances, and now you’re three meetings deep into capacity planning for traffic you don’t have yet.

Here’s what EC2 costs you in practice: a week on instance types, another week on auto-scaling groups, then a placement failure because the bin-packing algorithm can’t find space for your containers. Your “cheaper” option has eaten two sprints of engineering time.

Fargate works without any of that. You tell it how much CPU and memory you need, and it runs your container. No instances to manage, no patching, no capacity planning.

“But it’s more expensive.”

Sure, maybe 20-30% more at scale. But you’re not at scale, you’re trying to ship. An extra $200/month on your AWS bill is nothing next to the $30k+ in engineering salaries you’re burning while you debate this.

Python apps benefit from Fargate more than most. A Django app with Celery workers is memory-heavy and I/O bound, not CPU-intensive, so Fargate lets you right-size memory without playing Tetris with EC2 instance types.

Pick Fargate. Revisit it when you’re running 200 containers 24/7 and have real cost data. Until then, move on.

ECS service discovery: use an internal ALB

When your services need to talk to each other, AWS gives you three options: Cloud Map, internal ALB, or Service Connect. I’ve seen teams spend weeks evaluating all three: proof-of-concepts, whitepapers, the works.

Use an internal ALB.

I know, it’s not exciting. It’s a load balancer that’s been around forever. That’s exactly why it works:

  • It gives you a stable DNS name your services can call
  • Health checks work out of the box
  • You get access logs for debugging
  • Every developer on your team already understands HTTP

Your FastAPI service calls http://api-internal.yourdomain.local/users and it works. No service mesh, no Envoy sidecars, no DNS caching gotchas.

Cloud Map is fine, but I’ve debugged too many issues where services couldn’t find each other because of DNS TTL problems. Service Connect is powerful, but then you’re operating a service mesh: debugging Envoy proxy configuration when your actual problem is a database query.

The internal ALB is boring. Boring is good. Boring means you’re debugging your application code instead of your infrastructure.

CI/CD for ECS: use GitHub Actions

If your code is on GitHub, use GitHub Actions. Don’t overthink this.

“But CodePipeline is AWS-native.”

Yes, and it requires a pipeline with Source, Build, and Deploy stages, IAM roles for each stage, buildspec files, and wiring it all together. It’s more YAML for the same result.

“But Jenkins gives us more control.”

It’s 2025. Don’t set up a Jenkins server. You’ll spend more time maintaining Jenkins than deploying your app.

GitHub Actions has an official AWS action that handles ECS deployments:

yaml
- name: Deploy to ECS
  uses: aws-actions/amazon-ecs-deploy-task-definition@v1
  with:
    task-definition: task-definition.json
    service: my-service
    cluster: my-cluster
    wait-for-service-stability: true

That’s the whole thing. It registers your task definition, updates the service, and waits for the deployment to stabilize. AWS maintains it. It works.

Your deployment workflow lives in your repo, your team already knows GitHub Actions from running tests, and you’re not managing another piece of infrastructure. If you want to see how these pipeline runners work internally, I wrote a deep dive on building a CI/CD pipeline runner from scratch in Python.

If you’re on GitLab, use GitLab CI. If you’re on Bitbucket, use Bitbucket Pipelines. Use whatever’s already integrated with your code, and don’t add complexity.

ECS secrets management: use SSM Parameter Store

Where do you store your database passwords and API keys?

Not in your task definition. I’ve seen that. Don’t.

The two real options are SSM Parameter Store and Secrets Manager. Teams debate this endlessly because Secrets Manager has automatic rotation and sounds more “enterprise.”

SSM Parameter Store is free, integrates natively with ECS, and handles 99% of use cases.

json
{
  "secrets": [
    {
      "name": "DATABASE_URL",
      "valueFrom": "arn:aws:ssm:us-east-1:123456789:parameter/myapp/database_url"
    }
  ]
}

Your Python app reads os.environ['DATABASE_URL'] like it does locally. No SDK, no code changes.

Secrets Manager costs $0.40 per secret per month and is worth it if you need automatic rotation for RDS credentials, but you probably don’t need that on day one. Start with SSM, then migrate specific secrets to Secrets Manager later if you need rotation.

Don’t set up HashiCorp Vault unless a compliance requirement specifically mandates it. You’d be operating a distributed system just to store passwords. That’s not simplifying your life.

ECS logging and monitoring: use CloudWatch

Every team wants to evaluate Datadog, New Relic, Honeycomb, and then maybe self-host Prometheus and Grafana “for cost savings.”

Stop. Use CloudWatch.

Add this to your task definition:

json
{
  "logConfiguration": {
    "logDriver": "awslogs",
    "options": {
      "awslogs-group": "/ecs/my-service",
      "awslogs-region": "us-east-1",
      "awslogs-stream-prefix": "ecs"
    }
  }
}

Done. Your container logs go to CloudWatch, queryable with Log Insights. Enable Container Insights and you get CPU/memory metrics. Set up a few alarms and you now have better observability than 80% of startups.

Datadog is genuinely good software, I like it, but it costs $15+ per host per month from day one, plus another vendor relationship to manage. Add it later when you need distributed tracing or APM.

Self-hosted observability is a trap. I’ve seen teams spend months building ELK stacks and Prometheus clusters. That’s infrastructure work that doesn’t ship features. Unless you have a dedicated platform team, don’t volunteer for this.

The pattern behind these picks

Here’s what I recommended:

  • Fargate over EC2
  • Internal ALB over Cloud Map or Service Connect
  • GitHub Actions over CodePipeline or Jenkins
  • SSM over Secrets Manager or Vault
  • CloudWatch over Datadog or self-hosted

Every choice optimizes for the same thing: less stuff to manage.

Some of these cost slightly more money. Some are less flexible. But they all share one property: they let you ship faster and debug easier.

And here’s what nobody puts in their architecture decision records: every one of these choices is reversible.

  • Fargate to EC2? Task definitions work on both.
  • ALB to Service Connect? A DNS change.
  • SSM to Secrets Manager? Same integration pattern.
  • CloudWatch to Datadog? Add the agent, keep CloudWatch as backup.

The “wrong” choice costs you maybe a few hundred dollars a month in inefficiency. The debate over the “right” choice costs you weeks of engineering time.

What happens when teams ship this way

Teams that follow this advice ship in about a week:

  • Day 1-2: Fargate cluster up, first service running
  • Day 3: ALB routing traffic, services talking to each other
  • Day 4: GitHub Actions deploying on push to main
  • Day 5: Secrets in SSM, logs in CloudWatch, basic alarms set up

Week 2: building features.

Teams that “do it right” are still having meetings about networking topology in week 6.

I’ve watched startups run out of runway while their infrastructure was still “almost ready.” I’ve seen senior engineers burn out on DevOps work instead of building the product that got them excited about in the first place.

Your Python app on Fargate with CloudWatch logs isn’t going to fall over at 1,000 users. Probably not at 10,000 either. By the time scale is a real problem, you’ll have the traffic data and revenue to solve it properly.

Ship first. Optimize later.


If you found this helpful, share it on X and tag me @muhammad_o7 - I’d love to hear about your ECS deployment experiences. You can also connect with me on LinkedIn.

Need Help? I’m available for AWS and DevOps consulting. If you’re stuck in ECS decision paralysis or need help getting to production faster, reach out via email or DM me on X/Twitter.