Free observability for a personal service with no budget left over
As I wrote in the previous post, I am building an anonymous chat service as a side project. It is not finished yet, but I have an ADR from partway through the design — the record of how I decided on observability — so here it is, as I wrote it at the time.
Short version: I split the work across CloudWatch, RDS Performance Insights, and Grafana Cloud Free, and the whole thing costs nothing extra.
Here is what this post covers:
- Assumptions (as of August 2026)
- What the problem actually was
- “Observability” turns out to mean three different things
- What I decided
- What I turned down, and why
- The trade-offs I accepted for going free
- The status stays “Proposed,” on purpose
- Wrapping up
- References
Assumptions (as of August 2026)
Since this involves dollar figures, the assumptions first.
- Timing: this is an August 2026 estimate. Cloud pricing and free-tier terms change often, so don’t take it at face value — check the current official page yourself.
- Exchange rate: figures are in USD; I am not converting to yen.
- Estimate, not measurement. The service is not in production yet, so this is “probably around this much.” I expect it to be off once it is actually running.
- Setup: AWS. 2× EC2 + RDS PostgreSQL (Multi-AZ) + ElastiCache Redis. Fixed instance count, no autoscaling.
- Scale: designed for up to 1,000 concurrent connections. One person builds and runs it.
What the problem actually was
The infrastructure estimate looked like this.
| Item | Monthly estimate |
|---|---|
| WebSocket servers (2 small instances) | ~$40 |
| Redis | ~$9–20 |
| PostgreSQL (Multi-AZ) | ~$70 |
| DB storage (90 days, ~240GB) | ~$28 |
| Backups | ~$23 |
| Data transfer (egress) | ~$109 |
| Total | ~$280 |
The target was “around $280 a month,” so I had already spent the whole thing.
Which means “well, $30 a month for monitoring is fine” was never on the table. That’s the starting point for this post.
Incidentally, the single biggest line is egress (traffic leaving the server) at $109. A chat service keeps streaming data over WebSocket the whole time, so bandwidth ends up as the largest cost — a little counterintuitive.
“Observability” turns out to mean three different things
Once I broke down what I actually wanted, three things of different character had gotten mixed together.
1. Is the infrastructure alive
CPU usage, memory, connection count, and how much I’ve spent this month. This is the layer that watches whether the server is up and whether the wallet is on fire.
2. Which DB query is slow
When something feels sluggish, I want to name the query responsible. The infrastructure graphs only tell you “CPU is high,” so I need a different tool.
3. Numbers specific to the app
- How many people are connected right now
- How many seconds people wait before a match is found
- How often the profanity filter is triggering
Whether the service is actually usable only shows up here.
Trying to cover all three with one tool means something ends up expensive or something ends up weak. So I split them by layer instead.
What I decided
| Layer | Tool | Cost |
|---|---|---|
| Infrastructure metrics and alarms | CloudWatch + AWS Budgets | Within free tier |
| DB query analysis | RDS Performance Insights (7-day retention) | Free |
| App metrics and dashboards | Grafana Cloud Free | Free |
| Load testing | k6 (500 VU-hours/month on Grafana Cloud Free) | Free |
On the app side, I expose a /metrics endpoint in Prometheus format. Grafana Alloy (a collector agent) polls it and ships the data to Grafana Cloud.
A quick glossary
Terms that might be new, spelled out.
- CloudWatch: AWS’s built-in monitoring service. Basic numbers for EC2 and RDS show up automatically, with zero setup.
- AWS Budgets: lets you set “notify me once I’ve spent $140 this month.” I set alerts at 50% / 80% / 100% of budget.
- Performance Insights: an RDS feature that lists slow queries. Free if you keep 7 days of retention.
- Prometheus format: the de facto standard text format for exposing metrics — plain lines like
chat_active_connections 342. - Grafana Cloud Free: hosted dashboards. The free tier is “10,000 active series, 14-day retention, up to 3 users.”
- Series: one metric × one combination of label values = one series. This is the number that caps the free tier, and it matters later.
- k6 / VU: a load-testing tool. VU = Virtual User. “500 VU-hours” means, for example, 100 simulated users for 5 hours.
What I decided not to send out
Chat message bodies, session tokens, and raw IP addresses never go to an external monitoring service.
Since this is anonymous chat, conversation logs get deleted after 90 days and only reported ones are kept. Sender information also has to be handled inside a legal-compliance framework. The most common way this kind of thing goes wrong is logging a message body “just for debugging” and having it end up in a third-party SaaS, so I wrote the rule down explicitly.
What I turned down, and why
This is the section I personally find most valuable in an ADR. What you cut matters more than what you picked, once you’re reading it back later.
Option A: Datadog
The default choice for monitoring SaaS. When I actually priced it out, it came to $30/month for Infrastructure Pro on 2 hosts, $92/month with APM added. RDS and ElastiCache aren’t billed by host, so it was cheaper than I expected.
Still turned it down, simply because the budget is already spent — there’s no room for even $30. The free plan only keeps 1 day of history, which defeats the purpose of “how does this compare to last week.”
That said, I kept one option open for later: spin it up on-demand ($18/host) just for the duration of a load test, then tear it down. That’s fine as long as it isn’t a standing cost.
Option B: Push everything into CloudWatch
Staying entirely inside AWS is a real advantage — everything lives in one place.
The problem is custom metric pricing: $0.30 per metric per month. Looks cheap, but it multiplies.
8 topics × 2 room types × 4 indicators = 64 metrics
64 × $0.30 = $19/month
What bothered me wasn’t the dollar amount so much as the structure where measuring more granularly costs more.
Match-wait time is exactly the kind of thing I’m still figuring out from real data — how many seconds is too long to make someone wait. If “measuring this costs money” is pulling at that decision, it warps the judgment toward measuring less, and that’s what I wanted to avoid.
Option C: Run Prometheus + Grafana myself on EC2
Cost would be just the instance. Technically a normal setup too.
Two reasons I turned it down.
- It fails alongside the thing it’s watching. If AWS itself has a problem, the monitoring that’s supposed to tell me about it goes down with it — useless.
- Solo development doesn’t leave time to babysit a monitoring stack. Spending build time running a monitoring server instead is backwards.
The trade-offs I accepted for going free
Getting it for free means the cost shows up somewhere else. Here it is, plainly.
- One more container. Grafana Alloy runs on top of the two small instances. Resource usage isn’t zero.
- I have to stay aware of the series cap. A Prometheus histogram (the mechanism behind “how many requests finished within X seconds”) consumes one series per bucket — roughly
bucket count + 2per histogram. A histogram per topic would eat through the 10,000-series limit fast, so the label dimensions have to be decided up front. You could say the “measure less” pressure I disliked in Option B shows up again here, just in a different shape. - Only 3 people can log in. Not a problem solo. If the team grows, I move to the $19/month-and-up plan.
The status stays “Proposed,” on purpose
This ADR is not marked Accepted yet. It’s staying at Proposed.
The reason is that implementation is a later milestone. Free-tier terms change. RDS Performance Insights, in fact, is in the middle of being folded into a broader framework called Database Insights, and how much retention stays free is something I’ll only know for sure once I implement it.
Writing a decision as if it’s final, when the premises might still shift, means whoever reads it later (my future self, six months from now) will believe it without checking.
I wrote down what would trigger a re-think, too.
- If ad revenue or something else widens the budget
- If I hit Grafana Cloud Free’s series limit
Either of those, and I reconsider — Datadog included.
Wrapping up
- Even a personal service with a spent-out budget can get free-tier observability, if you split it by layer
- Judge pricing models by what behavior they push you toward, not just the dollar figure. A scheme that gets more expensive the more granularly you measure quietly pushes you to measure less
- Writing down what you turned down, and why, saves you from starting from zero when conditions change later
- It’s fine to leave something at Proposed when you can’t fully decide. “Haven’t verified this yet” is a legitimate thing to record
I expect the real numbers to miss the estimate once this is actually running. I’ll write about it again when they do.
References
The pricing and free-tier figures in this post reflect what each page said as of August 2026. Terms change — always check the current official page.
Pricing and free tiers (where the numbers came from)
- Amazon CloudWatch pricing — the $0.30/metric/month for custom metrics
- Publishing custom metrics - CloudWatch docs — how metrics and dimensions are counted
- Amazon EC2 On-Demand pricing — instance cost and egress rates
- Amazon RDS Performance Insights pricing — the conditions for free retention
- Grafana Cloud pricing — the Free plan’s series count, retention, and user limits
- Understanding metrics billing - Grafana Cloud — what “active series” actually means
- Datadog pricing — used for the Option A estimate
Tool documentation
- AWS Budgets — setting up budget alerts
- Amazon RDS Performance Insights
- Amazon RDS Database Insights overview — where Performance Insights is being folded into. Worth checking at implementation time
- Grafana Alloy docs — the metrics collector agent
- Grafana k6 docs — load testing
How Prometheus thinks about metrics
- Metric types - Prometheus — Counter vs Gauge vs Histogram
- Histograms and summaries - Prometheus — how buckets consume series
- Metric and label naming - Prometheus — guidance on not over-expanding label dimensions
ADRs (the format this post is based on)
- Documenting Architecture Decisions - Michael Nygard — the original source
- Architectural Decision Records — formats and templates
- architecture-decision-record - joelparkerhenderson — a template collection