Writing

Cloud Run or EC2: how I chose for a model that listens

In short: if the service takes a finished recording, scores it and answers, Cloud Run is usually the simpler and cheaper home. If it holds a live audio stream open, needs a GPU, or has to stay warm…

published
read time
5 min
words
994
lang
en
filed under
Engineering

In short: if the service takes a finished recording, scores it and answers, Cloud Run is usually the simpler and cheaper home. If it holds a live audio stream open, needs a GPU, or has to stay warm all day, a plain EC2 instance is easier to reason about. I've shipped services on both, and the deciding questions are always the same three.

A model that listens is an awkward thing to deploy. The inputs are large compared to a JSON request. The model is often slow to load. And the traffic is lumpy: nothing for an hour, then a clinic uploads a batch of recordings at once. Those three facts point at the three questions I ask before picking a platform.

  1. Cold starts. How long does it take from "no instance running" to "first answer", and does anyone notice?
  2. Cost when idle. How many hours a day does the service do nothing?
  3. Streaming. Does the audio arrive as a file, or as a live stream that has to stay open?

The two options in one look

Cloud Run starts containers as requests arrive and scales them back to zero when nobody is calling. EC2 gives you a virtual machine that runs until you stop it. Everything else follows from that difference.

Side by side

Cloud Run

  • Container in, HTTPS endpoint out
  • Scales to zero, pays per use
  • Cold start includes loading the model
  • Request size limits push you to upload audio to storage first
  • Little to patch or babysit

EC2

  • A machine you own and maintain
  • Pays by the hour, busy or idle
  • Model stays loaded, no cold start
  • Long-lived connections and streams are easy
  • GPU types are a dropdown away

Cold starts are mostly model loading

People blame the platform for cold starts. In my experience, for an audio model the container start is the small part. The big part is the process importing its libraries and loading weights into memory. A small model loads quickly. A large one takes long enough that the first request after a quiet period feels broken.

Three things help on Cloud Run:

  • Load the model once, at startup, in a global, not inside the request handler. Obvious, and still the most common mistake I see in notebook code turned into a service.
  • Bake the weights into the image instead of downloading them from a bucket on every start.
  • Keep one instance warm with a minimum instance count, if a slow first answer is a real problem. This trades a little idle cost for predictable latency.

On EC2 the question disappears. The model loads once when the service starts and stays there. You pay for that with the machine running around the clock.

Cost and latency, roughly

I'm not going to put prices here. They change, they depend on region, and the right comparison is your own traffic. What doesn't change is the shape.

SituationCloud RunEC2
Idle most of the dayClose to nothing while idleFull hourly price while idle
Bursty uploadsScales out on its ownQueue up, or set up autoscaling yourself
Steady load all dayPer-use pricing adds upOften cheaper once it is always busy
First request after quietSlow unless one instance is kept warmSame as any other request
Live audio streamPossible, but bounded by request timeoutsSimple, the connection lives as long as you like
Needs a GPUNewer and more limitedPick a GPU instance type

The line that usually decides it is the first one. Medical audio tools are often used during clinic hours and nothing else. A machine that runs all night to score nothing is money spent on a warm room.

Streaming changes the answer

If the recording is finished before it reaches you, treat it as a file. The pattern I like: the client uploads the audio straight to a storage bucket with a short-lived signed URL, then calls the scoring service with the object path. The service never receives the large body, so request size limits stop mattering, and the raw audio lands in one place where you can audit and delete it.

client bucket scorer 1 audio file 2 object path only 3 read 4 score
The large body never passes through the service. Only a path does.

If the audio is live, for example a voice agent or a device that streams while the patient speaks, the shape is different. You hold a connection open for minutes, keep state per session, and care about the delay on every chunk. That is where a long-running process on a machine you control is simply less friction. Cloud Run can serve WebSockets, but you'll spend your time working around timeouts and instance lifetimes instead of on the model.

A deploy that respects the model

When Cloud Run is the right answer, most of the work is in a few flags. Concurrency is the one people miss. An audio model usually uses a full CPU per request, so letting one instance accept dozens of requests at once just makes all of them slow.

gcloud run deploy voice-scorer \
  --image REGION-docker.pkg.dev/PROJECT/models/voice-scorer:1.4.0 \
  --region REGION \
  --cpu 2 --memory 4Gi \
  --concurrency 2 \
  --min-instances 1 \
  --max-instances 20 \
  --timeout 300 \
  --cpu-boost \
  --no-allow-unauthenticated

Pin the image to a version, not latest, so you always know which model answered. Set --min-instances to 0 if a slow first request is fine, and to 1 if it isn't.

How to choose for your own service

Write down three numbers before you pick anything: how long your model takes to load, how many hours a day it is busy, and whether the audio arrives as a file or a stream. Load the model in a container on your laptop and time the first request. If it's busy most of the day or streams, start on EC2. If it's quiet most of the day and takes files, start on Cloud Run with the upload going straight to storage. Either way, you can move later. The container is the same.

related

Keep reading