Running a model in a Jupyter notebook is trivial. Running a model that serves 500 predictions per second with 99.9% uptime, auto-scales with traffic, recovers from node failures, and costs less than $2,000 per month is an infrastructure problem. The tool you choose to orchestrate your inference containers determines whether that infrastructure problem is a one-time setup or an ongoing operational burden.
Kubernetes, AWS ECS, and Fly.io represent three points on the complexity spectrum. Kubernetes gives you maximum control at maximum complexity. ECS gives you AWS-integrated orchestration at moderate complexity. Fly.io gives you minimal-configuration deployment at minimum complexity. The right choice depends on your team size, your scale, and your tolerance for infrastructure work.
Kubernetes: the default that costs more than you think
Kubernetes is the industry standard for container orchestration. It handles scheduling, scaling, networking, storage, and service discovery across a cluster of machines. For ML inference workloads, Kubernetes provides GPU scheduling, horizontal pod autoscaling based on request latency or queue depth, and rolling deployments that maintain availability during model updates.
The capability is real. Kubernetes can run any workload at any scale. The cost is operational complexity that most ML teams underestimate. A production Kubernetes cluster requires cluster provisioning and maintenance, networking configuration (ingress controllers, service mesh, DNS), storage provisioning for model artifacts, GPU driver management, monitoring and alerting for cluster health, and security configuration (RBAC, network policies, secrets management).
Managed Kubernetes offerings (EKS, GKE, AKS) reduce the operational burden by handling the control plane, but the application-layer configuration remains your responsibility. A team deploying their first ML inference service on EKS should budget two to four weeks for initial setup and one to two hours per week for ongoing maintenance.
Kubernetes makes sense when you are running multiple ML models with different resource requirements, different scaling profiles, and different deployment schedules. It also makes sense when your organisation already runs Kubernetes for other workloads and adding ML inference is incremental rather than net-new. If Kubernetes is not already in your stack, do not adopt it solely for ML inference.
ECS: the pragmatic middle ground
AWS Elastic Container Service is AWS’s managed container orchestration service. It runs containers on EC2 instances or on Fargate (serverless compute). For ML inference, ECS with Fargate handles simple CPU-based models. ECS with EC2-backed clusters handles GPU-based inference workloads.
ECS’s advantage over Kubernetes is simplicity. The service model is straightforward: define a task definition (container image, CPU/memory requirements, environment variables), define a service (how many copies, scaling rules), and ECS handles the rest. Networking integrates with VPC natively. Load balancing integrates with ALB natively. Logging integrates with CloudWatch natively. There are no CRDs to manage, no Helm charts to maintain, and no YAML files to debug.
For GPU inference on ECS, you configure a task definition with GPU resource requirements, launch it on a GPU-enabled EC2 instance type (p3, p4, g4, g5), and ECS schedules it. The setup is simpler than Kubernetes GPU scheduling because AWS manages the GPU driver installation on the ECS-optimised AMI.
The limitation is AWS lock-in. ECS is an AWS-only service. If you need multi-cloud deployment or on-premises inference, ECS is not an option. Within AWS, this limitation is irrelevant for most teams because the integration with other AWS services (CloudWatch, IAM, Secrets Manager, SageMaker for model artifacts) reduces operational overhead that multi-cloud architectures create.
ECS also lacks some of Kubernetes’ advanced features: no built-in service mesh, no native support for canary deployments (though you can achieve this with ALB weighted target groups), and no equivalent to Kubernetes’ extensive ecosystem of operators and controllers. For ML inference workloads that do not need these features, the absence is not a problem.
Fly.io: the minimal option
Fly.io takes a radically different approach: deploy containers close to users with minimal configuration. You write a Dockerfile, run fly launch, and your container is running on Fly.io’s edge infrastructure. Auto-scaling, load balancing, and health checks are configured with sensible defaults.
For ML inference, Fly.io works well for lightweight models: small transformer models, classification models, feature engineering services, and API gateways. The deployment experience is the fastest of the three: a working inference endpoint can be running in under ten minutes.
Fly.io supports GPU inference through its GPU-enabled machines. You configure GPU type and count in your fly.toml file and Fly provisions a GPU machine for your workload. The GPU offering is newer and less mature than ECS or Kubernetes GPU support, with fewer instance type options and less fine-grained control over GPU memory allocation.
The limitation is scale and maturity. Fly.io is designed for application workloads with moderate resource requirements. If your inference workload requires dozens of GPUs, complex multi-model serving with resource isolation, or integration with enterprise monitoring and governance tools, Fly.io is not the right platform. It is the right platform for small ML teams that want to deploy models quickly without managing infrastructure.
Cost comparison
At small scale (one to five inference endpoints, moderate traffic), Fly.io is the cheapest option because you pay only for the compute you use with no cluster management overhead. ECS with Fargate is comparable for CPU workloads. Kubernetes is the most expensive because even managed control planes have a base cost and you typically over-provision cluster capacity.
At medium scale (ten to fifty endpoints, variable traffic), ECS and Kubernetes become cost-competitive because their auto-scaling capabilities reduce idle resource waste. Fly.io’s per-machine pricing can become expensive if you need dedicated GPU instances that run continuously.
At large scale (hundreds of endpoints, thousands of requests per second), Kubernetes has the best cost profile because you can bin-pack multiple models onto shared GPU nodes, use spot instances for non-critical workloads, and fine-tune resource allocation per model. ECS is comparable if you use EC2-backed clusters rather than Fargate. Fly.io is not designed for this scale.
Model update and deployment patterns
ML models update more frequently than traditional software. A/B testing new model versions, canary deployments, and blue-green deployments are standard patterns. The three platforms handle these differently.
Kubernetes supports all deployment patterns natively through its rolling update, canary (via service mesh or Argo Rollouts), and blue-green strategies. The ecosystem around Kubernetes for ML deployments (KServe, Seldon Core, BentoML) provides inference-specific features like traffic splitting, model versioning, and request routing.
ECS supports rolling deployments natively and blue-green deployments through CodeDeploy integration. Canary deployments require ALB weighted target groups, which work but are less flexible than Kubernetes-based canary routing. For ML workloads where canary deployment is important, ECS requires more manual configuration.
Fly.io supports rolling deployments and blue-green deployments through its release command and health check mechanisms. Canary deployment is possible through its traffic routing features but is not as granular as Kubernetes-based canary. For most ML inference workloads at Fly.io’s target scale, rolling deployment with health checks is sufficient.
Decision framework
Use Kubernetes when you are running a platform with multiple ML models, when your organisation already uses Kubernetes, when you need advanced deployment patterns (canary, A/B testing, multi-model serving), or when you need multi-cloud or hybrid-cloud deployment. Budget for an infrastructure team or platform engineering function.
Use ECS when you are in AWS and want a simpler alternative to Kubernetes for ML inference. ECS is the right choice for teams with two to twenty inference endpoints that need auto-scaling, GPU support, and AWS-native integration without Kubernetes complexity.
Use Fly.io for small ML teams that want to deploy models quickly, for lightweight inference workloads, for edge deployment where latency matters, or for prototyping and staging environments. It is the wrong choice for GPU-heavy workloads at scale.
The honest heuristic: if you do not have a dedicated platform engineering team, start with ECS or Fly.io. Kubernetes is a two-year commitment, not a two-week project.