
Last Update: October 9, 2026
BY
eric
Keywords
Vibe coding is the trend of the moment, and for good reason. AI has knocked down the barrier that used to separate "has an idea" from "can ship software," and people who have never written a line of code are deploying working services. And it isn't only beginners: plenty of experienced developers have largely stopped hand-writing code altogether, because the frontier models are simply good enough that describing the change is faster than typing it.
The barrier that came down, though, wasn't only a barrier to writing code. It was also the thing that used to force you to understand what you were writing. Vibe coding collapses the distance between "I have an idea" and "it's running," and the part it quietly removes is where you'd normally stop and ask how often does this run, and what does each run cost? An AI model writing your polling loop or your alarm handler has no idea what pricing plan you're on. It optimizes for "this satisfies the prompt," not "this is cheap to operate at scale."
What follows is three real stories of AI-written code that shipped with a correctness bug which was also, quietly, a cost bug, and nobody noticed until the invoice did the noticing for them. Our companion post covers what Cloudflare's D1, Durable Objects, Workers, Email Sending, and Containers actually cost; this one is about what happens when the code driving those meters is wrong.
Lesson 1: the alarm that never stopped alarming
A Cloudflare user, @shmily7, described it bluntly in a post that's since been seen nearly half a million times: "遭遇Cloudflare帐单刺客,爆了一万刀的帐单,原因是一个项目 Durable Object alarm 死循环,读写了6万亿次,Vibe Coding误我". Translated: hit by a Cloudflare "billing assassin," a $10,000 bill, caused by a Durable Object alarm stuck in a dead loop that read and wrote 6 trillion times. Their own verdict: vibe coding did them wrong.

(X's automatic translation renders 6万亿 as "60 trillion"; the Chinese figure is 6 trillion. Either way, the point stands.)
Durable Object alarms are a single callback, alarm(), that Cloudflare invokes once at a time you set with setAlarm(). The pattern for a recurring job is to call setAlarm() again at the end of your own handler to schedule the next run. That's also exactly where the bug lives: if the rescheduling math is wrong (a delay of 0, a timestamp computed in the past, an error path that retries immediately instead of backing off) the alarm fires again the instant it finishes, forever, and each firing bills as a request plus duration plus whatever it reads or writes. At six trillion operations, even the included free quota's rounding error is a five-figure invoice. The logic a careful engineer would write defensively (minimum delay, a sanity check on the computed next-run time, an exponential backoff on error) is exactly the kind of unglamorous edge-case handling that a quick vibe-coded prompt tends to skip, because the happy path, "alarm fires, does work, schedules next one," looks completely correct in a two-minute test.
Lesson 2: the poller that never waited
@keking99 built a DeepSeek voting station with GPT-6 Astra, OpenAI's frontier model, released barely a month earlier, "in just a few sentences," and got a $1,000 Cloudflare bill for the month. The description of the bug is the most useful part, so here it is in X's own translation:
I wanted you to display the vote count in real time, and you give me this thing that polls api/result every 3 seconds as long as the webpage is open, infinite querying, right? The request volume alone hit 2.6 billion times, you can make such a basic mistake? Now you see the consequences of blindly vibing without looking at the code, huh. Boomerang to the back of the head.

It's worth sitting with the fact that this was not some budget model cutting corners. Astra was the most capable thing OpenAI had shipped, posting near-perfect scores on reasoning benchmarks, and it still produced this. That's the uncomfortable part: capability and cost-awareness are different properties, and no amount of the former delivers the latter. The model wasn't asked what the code would cost to run, so it didn't consider it. Nor will it, however good it gets, because "correct" and "cheap to operate at scale" are separate questions and only one of them was in the prompt.
Note too that what it produced isn't stupid. It's reasonable-looking. Asked for a live-updating vote count, it reached for the most obvious implementation: poll the result endpoint every three seconds. On one open tab that's 1,200 requests an hour, which is fine. The part nobody costed is that it runs per open page, for as long as that page is open, across every visitor, and a voting page is precisely the kind of thing people leave open. Multiply a reasonable-looking interval by an audience and a duration and you arrive at 2.6 billion requests without any single decision ever looking wrong.
That's the second, more common shape of this failure: not an infinite loop, just an interval nobody multiplied out. A setInterval, a while (true) with a fetch() and an undersized sleep, a retry with no backoff. Each looks like working code, and none look dangerous in a two-minute test, because a human tester doesn't leave the tab open for a week. The AI that wrote it never ran it at all.
Lesson 3: when the "fully managed" database costs more than the server you didn't want to run
This one's ours, and it's not a Cloudflare story, which is the point. We used Google App Engine with Datastore to log machine/operation records: every command execution, every status check, written as a document. It felt like the right call: fully managed, no server to patch, scales automatically. The bill for database operations alone came to roughly $300, comfortably more than it would have cost to run our own MySQL or Postgres instance and handle the same write volume ourselves.
The mechanism is the same shape as the other two, just slower-burning: a pay-per-operation managed store (Datastore, Firestore, DynamoDB, D1, Durable Objects storage; the specific product doesn't matter) charges per read and per write, while a self-hosted relational database on a flat-rate VM charges the same whether you run a thousand queries a day or a hundred million. Below some volume, managed-and-metered wins easily. Above it, the per-operation meter quietly overtakes a box you could have rented for $20/month. Nobody sat down and did that napkin math before wiring up the logging. It just seemed like the "serverless-native" choice, which is exactly the trap: convenience at prototype time doesn't come with a label telling you where the crossover point is.
The pattern underneath all three
Different clouds, different products, same shape of mistake:
- Code that is functionally correct, in that it does the thing it's supposed to do, while being economically unbounded, because nothing in the code or the review caps how many times, how fast, or how expensively it can run.
- A happy-path test that looks fine, because happy-path tests don't run for a month unattended or process a table that's grown 100x since launch.
- A review step that used to happen almost for free, as a side effect of how slow writing code by hand used to be, that vibe coding genuinely removes. You can now ship in the time it used to take to plan, which is great, except planning was also when "how often does this run?" used to get asked.
None of this is an argument against AI-assisted coding. It's an argument for putting the review back in deliberately, since it no longer happens automatically as a side effect of friction.
Where it bites, service by service
Staying with Cloudflare for a moment, since it's the platform in two of the three stories: each of its metered services has a characteristic way of turning a small coding mistake into a large number. The companion post has the full price sheet; this is what goes wrong against it.
- D1 bills per row, not per query, and writes cost roughly 1,000x what reads cost. An
UPDATEthat touches every row in a growing table, run on every request, scales its cost with table size. It's cheap in testing against 50 rows and ruinous at a million. This is the one that cost the thread's author over a thousand dollars. - Durable Objects run three meters at once: requests, duration in GB-seconds, and D1-rate storage for anything persisted. An alarm handler that both reschedules itself and writes state on every tick is spinning all three, which is why a bad reschedule compounds so fast.
- Workers bill CPU time separately from requests. A tight loop, an unbounded recursive parse, or synchronous crypto on a large payload can burn seconds of CPU per call instead of milliseconds. The guard here is cheap and almost always skipped: set
limits.cpu_msinwrangler.tomland a runaway request gets killed rather than billed to completion. - Email Sending isn't a runaway-loop risk so much as an abuse surface. A magic-link or signup endpoint with no rate limiting is a free way for bots to make you pay per attempt. The thread's author was seeing several dollars a day from bots hammering a signup form, which is a rounding error right up until it isn't.
- Containers meter on wall-clock time whether or not they're doing anything. The trap is assuming the $5/month Workers Paid plan covers it; a small always-on container runs to roughly $30/month having served no traffic at all. Several people got caught by this after a Cloudflare demo made it look like a $5 add-on.
What to check before you ship vibe-coded infrastructure code
- Every loop, retry, reschedule, or alarm: what's the exit condition, and what happens on the path where that condition is never met? Walk the failure branch, not just the success branch.
- Every poller: is there a minimum interval, a backoff on failure, and a cache so a quiet period doesn't mean a constant stream of identical calls?
- Every write inside a request handler or a loop: does its cost scale with table size, request volume, or iteration count? If yes, what's the worst case, not the typical case?
- Usage alerts configured before the first deploy, not after the first scary invoice. On Cloudflare that's the Usage Based Billing Alert (Pro plan and above, pay-as-you-go accounts); most providers have an equivalent. Turn it on before you need it.
limits.cpu_msset inwrangler.tomlif you're on Workers. Cloudflare's own docs pitch it exactly this way, to prevent "accidental runaway bills or denial-of-wallet attacks."- A back-of-envelope cost estimate for anything that talks to a pay-per-operation service: expected operations per day times the unit price, compared honestly against what a flat-rate self-hosted alternative would cost at the same volume.
The other half: platforms could meet us halfway
Everything above is the developer's job, and none of these were exotic bugs. An alarm that reschedules itself too eagerly and a poller with no throttle are rookie mistakes, and owning them is part of shipping software. But "don't write bugs" has never been a workable safety strategy, because all code has bugs. That's precisely why we have type checkers, linters, circuit breakers, and rate limits: we assume the mistake will happen and build something that catches it.
Cloud billing is conspicuously missing that layer. The controls that exist are weak in a specific way:
- Alerts notify; they don't stop anything. A usage alert tells you the meter is spinning. If it fires at 2am on a Saturday, the meter keeps spinning until you wake up and act. The one mechanism that actually stops work,
limits.cpu_ms, caps a single invocation, and every incident described above was unbounded invocation count, not one expensive call. Nothing capped that. - There's no hard spend ceiling. You cannot tell Cloudflare "whatever happens, never bill me more than $100 this month, and shut it down if we get there." For a hobby project or a small business site, that's the single control that would matter most, and the inability to set it is why a weekend project can generate a mortgage payment.
And this is plainly solvable, because other providers already solve it. Vercel has Spend Management: you set an on-demand budget, and with Pause Production Deployments enabled it automatically pauses every project on the team when you hit it, so visitors get a 503 rather than you getting a surprise invoice. Google Cloud documents a budget-triggered function that removes the billing account from a project outright, shutting the whole thing down. Neither is elegant, and both ship with honest caveats: Vercel checks spend "every few minutes," so pausing isn't instantaneous, and Google states flatly that its approach "doesn't guarantee that you won't spend more than your budget."
Those caveats are usually where the argument against spend limits starts: metering is distributed, usage data arrives late, an exact cap is genuinely hard. All true, and all beside the point. A cap that overshoots by a few minutes turns a $10,000 incident into a $20 one. Nobody asking for a spend limit expects it to be accurate to the cent; they want the runaway to stop while it's still a rounding error. "We can't make it exact" is not a reason to offer nothing, and the existence of imperfect-but-real implementations elsewhere makes it a product decision rather than a technical impossibility.
And the anomaly here was not subtle. A site with a handful of users does not legitimately generate six trillion database operations, or 2.6 billion API calls. Against its own baseline, that traffic is self-evidently broken, and the platform is in a far better position to see it than the developer is, because the platform is the thing counting. It has the traffic history, it knows the account's normal range, and it is already doing real-time analysis on every request for security purposes. A provider that can detect and mitigate a DDoS in seconds can notice that a free-tier-sized project has started writing a billion rows an hour.
None of this excuses shipping the bug. But there's a reasonable expectation that the party holding the meter, watching the traffic, and writing the invoice should be the one to say "this looks wrong, we've paused it," rather than quietly counting to six trillion and sending a bill. Until that exists, the burden sits entirely on the developer, which is exactly the wrong place for it when the tooling has made it so easy for non-specialists to deploy code they don't fully understand.
References
- @shmily7 on X, the $10,000 Durable Objects alarm bill
- @keking99 on X, the $1,000 bill from a three-second polling loop
- GPT-6 Astra, OpenAI, the model that wrote the polling loop
- @mcwangcn on X, a thread cataloguing Cloudflare billing traps, including a passing mention of @keking99's requests-driven bill
- Durable Objects alarms API, Cloudflare Developer Docs
- Workers pricing, Cloudflare Developer Docs, source of the
limits.cpu_msguidance on runaway bills - Available notifications, Cloudflare Developer Docs, covering the Usage Based Billing Alert and its plan restrictions
- Spend Management, Vercel Docs, on budget-triggered pausing of production deployments
- Disable billing with notifications, Google Cloud Docs, on budget-triggered shutdown and its limitations
Want the actual price sheet behind why these bugs got so expensive? See the companion post on what Cloudflare's developer platform costs, service by service.





Comments (0)
Leave a Comment