A few weeks ago I built a fairly complex automation with Next.js on the frontend and Python handling most of the backend logic. Everything looked fine. Tests were passing. Logs looked normal. The automation was doing what it was supposed to do. Then one night around 2 AM, the LLM returned malformed JSON. Normally you'd expect something to fail loudly. This one didn't. Part of the workflow continued, part of it stopped, and the failure wasn't obvious until the client noticed something was wrong. That was the uncomfortable part. The system hadn't really failed. It had failed silently. The client found it before I did. I was honestly lucky they did, because this was a high-value client and it could have turned into a very different conversation. I started looking at how I was monitoring the automation and realised something I hadn't really thought about before: Most of my tooling was very good at telling me what happened. But I needed something that could answer: Did the automation actually produce the outcome it was supposed to produce? So I started building a layer around the automation that doesn't just watch errors. It checks the output against the expected structure, looks for suspicious states, catches things like malformed responses, retries and rate-limit behaviour, records the evidence, and tries to determine whether the workflow should actually be considered successful. The interesting part is that I've found quite a few cases where: HTTP 200 workflow completed LLM said "success" …still didn't mean the business outcome actually happened. I'm about 60% through rebuilding this properly, and I'm beginning to think the difficult part of AI automation isn't getting the workflow to run. It's proving that it ran correctly. Now I'm going down a bit of a rabbit hole with failure reproduction and regression testing too, because fixing a failure once doesn't tell you whether the same thing will happen again three weeks later. Curious how other people are handling this.