Google CloudSeptember 22, 20264 min read

How to Find Idle Google Cloud Resources Still Billing You

Google already built the tool, runs it every day, and tells you what it found. It just will not tell you the numbers it used to decide.

Google has been quietly doing this work for you already. There is a system called Active Assist that watches your project, decides what is not being used, and writes it down. It is free, it is already on, and most people have never opened it.

In the console it shows up as Recommendations, on the home dashboard and on each service's page. Underneath, it is a set of separate recommenders, each looking at one kind of resource.

The ones that find dead money

There are six worth knowing by name, because their names tell you exactly what they do.

The idle VM recommender, google.compute.instance.IdleResourceRecommender, whose whole description is "Remove unused VMs". This is the expensive one, because a machine nobody uses is still a machine you rent by the hour.

The idle disk recommender, google.compute.disk.IdleResourceRecommender. Google's own wording for this one is "Backup and remove unused disks", and I like that they put backup first, because that is the right order and it is the step people skip.

The idle IP recommender, google.compute.address.IdleResourceRecommender, "Remove unused IPs". Worth doing early, because on Google Cloud a reserved address that is not attached to anything is charged at a higher rate than the same address in use. The idle one costs you more than the working one.

The idle Cloud SQL recommender, google.cloudsql.instance.IdleRecommender. A database nobody queries bills exactly the same as one under load.

The GKE recommender, google.container.DiagnosisRecommender, which finds unused clusters. An empty Kubernetes cluster is one of the most expensive ways to run nothing.

And the machine type recommender, google.compute.instance.MachineTypeRecommender, which does not remove anything but tells you where a machine is bigger than its workload.

How Google decides something is idle

Read this before you act on anything it says, because it changes what the recommendations mean.

The observation window is stated plainly: "By default, the historical observation period is 14 days, or, for new VMs, starting one day after VM creation." And "New recommendations are generated once per day."

So a machine you spun up on Monday will not be judged yet, and something you fixed yesterday may still be listed until tomorrow. Neither is a bug.

The test itself is this: "If CPU and network usage are below predefined thresholds, the Recommender classifies the VM as idle."

Now the part worth knowing. Google does not publish what those thresholds are, and says they "might change in the future". You are being handed a conclusion without the working.

They do give you one sentence that explains the intent better than a number would: "Google designs the default thresholds so that if monitoring agents generate the majority of CPU and network usage, then your VM is classified as idle."

That is a genuinely good definition. A machine whose only real activity is the agent reporting that the machine is alive is not doing anything. It is a server whose entire job is telling you it exists.

What it will not catch

Recommenders look at usage. Things that cost money without usage patterns slip past.

Disks attached to a stopped VM are the clearest case. Stopping an instance ends the charge for its vCPUs and memory, and the disks keep billing at full price the whole time. The machine is off, the storage is not, and nothing is idle in a way the recommender measures.

Snapshots are another. They accumulate quietly, they are individually cheap, and nobody has ever been thanked for deleting one.

And the biggest gap is not a resource at all. If a scheduled query or a dashboard is scanning large tables in BigQuery on a timer, you are paying per query, and no idle-resource recommender will ever mention it because there is nothing sitting there to call idle.

Before you delete anything

Take Google's own advice from the disk recommender and back it up first. A snapshot of a disk costs a fraction of the disk, and it buys you the right to be wrong.

Then wait. A fortnight between the snapshot and the deletion costs almost nothing and catches the quarterly job nobody remembered. If a resource has been idle for 14 days and nobody claims it in another 14, it is genuinely unused.

The one thing I would not do is act on a list of recommendations in bulk without reading it. Idle for fourteen days is not the same as unnecessary. Disaster recovery standby machines are idle by design, and so is anything seasonal.

The part that does not stick

Run this once and you will find things. Run it again in four months and you will find different things, because waste is not an event, it is what ordinary work leaves behind.

Somebody tests an idea and moves on. A project ends and its database stays. A machine gets stopped rather than deleted because stopping felt safer, and its disks bill quietly for a year.

Where we are with this

Being straight with you: Liberra is an AI that connects to your cloud, indexes it, and answers questions like this without you opening six screens. It runs on AWS and Azure today. Google Cloud is next and I am not going to put a date on it here, because a date given before something ships is a guess wearing a suit.

So use the list above, it is free and it is the same list I would be running. Open Recommendations in your console, start with the idle IP addresses because that one is pure profit, then the disks, then the VMs.

One thing that will be true when Google Cloud does land, because it is true on both clouds we run on now: Liberra cannot delete anything. The delete commands are blocked in the code itself. It will find the idle disk, tell you what is on it and what it costs, and then the decision stays yours, in your own console. For a list where the whole risk is removing something that turned out to matter, that is where it belongs.

Founder, Liberra AI