We Stopped Reading Logs - But We Still Fix Issues Before Users Notice

Annotation

When you're maintaining a long-term project, you hopefully have logs. We do - we have a lot of logs - thousands and thousands of logs a day from different parts of the system.

And we filtered these logs and reviewed them daily - a human skimming them and escalating whatever felt off... And it worked well, we were able to proactively fix issues before users reported them...

However, as the lazy developer I am, I never had the intention to do this manually forever - so I spent a couple of hours every week and vibe-coded the perfect solution to replace me!

Now we have an automated LLM-enabled pipeline that ingests logs from the frontend, the API, all of the failed calls to 3rd party integrations, normalizes them, matches front-end errors to the backend errors while aware of both codebases, maintains history of issues and correctly escalates new issues and supresses known noise - and delivers a better report than I would every morning into our Slack channel.

And the kicker? It's much easier to implement than you'd think, more deterministic than you'd think, and much cheaper to run than you would expect.