How we built a weekly automated flow health monitor with Claude
Generated by Klaviyo AI
Particle's VP of retention explains how the team built a custom dashboard that pulls data from the Claude x Klaviyo integration to monitor 140 active flows and 1,000+ live emails weekly. Instead of manually spot-checking top flows, the system scores every flow's health and flags specific underperforming emails against thresholds the team set, then routes fixes straight into their Monday.com board.
- Every email gets its own alert thresholds: The system tracks drops in open rate, click rate, revenue per recipient, and recipients, plus increases in spam complaints, unsubscribes, and bounce rate — separating performance alerts from deliverability alerts since they need different responses.
- Claude acts as a context-aware analyst, not just a dashboard: Each email has a Claude chat scoped to it, already aware of the metrics history; the team can ask it to diagnose a metric drop or suggest A/B tests, and Claude reads the email's actual HTML to ground its suggestions in the real subject lines and content.
- The payoff: 400% more A/B tests and faster fixes: The team now runs more than 30 new tests per week, a 400% increase in testing capacity, and catches silent technical failures (like a broken trigger or filter) within a week instead of letting them run undetected for months.
Every retention team has the same weakness: automated flows break, but you don’t necessarily notice where they’re breaking, let alone why.
A single email inside an 8-email welcome series tanks. Open rates slip, spam complaints creep up, revenue drops. But the flow-level numbers still look fine, so nobody notices for weeks.
I’m VP of retention at Particle, a leading DTC men’s personal care brand with an email list of over a million subscribers. My team and I have over 1,000 live emails and text messages across 140 active flows, with 5–30 emails in each flow.
At that kind of volume, monitoring email performance and making sure each flows is working to its full potential would take a team an entire week.
So, we built a fix with Claude.
The problem with how most teams manage automated flows
Most retention teams do the same thing: they go to the flows generating the most revenue, pick one or two emails, run a test, and watch the results. And then they repeat.
It’s not a terrible strategy, but it’s not true optimization, as it only addresses the top 3% of your email program.
The other 97% just runs, often with nobody checking for the deeper issues like whether:
- The fifth email in the abandoned cart flow is still performing as well as it was 18 months ago
- The design is up to date
- The offer is stale
- Something in the trigger logic stopped firing
- Spam complaints or bounce rates have increased
Klaviyo gives you flow-level numbers, and if the flow overall looks fine, it’s easy to assume everything inside it is fine, too.
For us, with 140 flows and so many active messages, that assumption was leaving a lot of money on the table. But we couldn’t see exactly where these deeper issues were happening without investing a lot of time.
Building a flow health score with Claude
We built a weekly flow monitor: a dashboard that pulls data from the Claude x Klaviyo integration and our entire Klaviyo account automatically every week.
Here’s how it works:
1. Every flow gets a health score.
2. Every email inside every flow gets monitored against thresholds we set ourselves:
- Drop in open rate
- Drop in click rate
- Drop in revenue per recipient
- Drop in recipients
- Increase in spam complaints
- Increase in unsubscribes
- Increase in bounce rate.
3. When something crosses a threshold, it triggers an alert. We separate performance alerts from deliverability alerts because they require different responses.
4. The result is a prioritized view of exactly which flows need attention and which specific emails inside those flows are the problem. It tells you the specific email, with a direct link to that email in Klaviyo.
Flow health pro tip: We set minimum recipient thresholds, too. If an email is too new or has too little volume to generate meaningful data, it doesn’t trigger alerts. The system only surfaces real signals.
Getting root-cause answers from Claude in seconds, not days
Here’s where our weekly flow health monitor gets even more useful.
Inside every email view in the dashboard, there’s a Claude chat window scoped to that specific email. When an alert fires, we get a notification, and Claude acts as a conversation partner that already knows the context: what the email is, what the metrics looked like last week, what changed, and what the possible causes are.
So, we can ask it directly:
- What caused the drop in click rate?
- Is the decline in recipients the likely driver of the revenue drop, or is something else happening?
- What A/B tests would you run on this email given what you’re seeing?
Claude can read the HTML of that specific email and run a more accurate analysis based on subject lines, preview lines, and the actual content, which makes the suggestions less generic. It generates the test suggestions in seconds. We review them, create the task, and move on.
Before our flow health monitor existed, surfacing that same set of recommendations manually would have required pulling data from Klaviyo, building a report, analyzing it, and then writing a brief from scratch. For one email. For us, with more than 1,000 live messages at a given time, this just wasn’t possible.
The Monday.com integration: closing the gap between finding a problem and fixing it
When we decide to act on an alert, the task goes straight to our Monday board with all the data attached: the flow name, the email, the alert type, and the metrics. The right person gets assigned. Once it’s resolved, it clears from the board.
This matters because the analysis and the work used to live in completely different places. You would find a problem in one tool, write it up, paste it somewhere else, and assign it manually. Now, the problem identification and the task happen in the same place.
Flow health pro tip: If you set up a similar process, make sure you send tasks to your project management system.
400% more tests, tens of thousands more in revenue
With our new weekly flow health monitor, we’re running more than 30 new A/B tests per week across our flows.
At the volume of tests we’re now running across the entire program, our flow health monitor has increased our testing capacity by 400%. With a list our size, where every 1% lift in click rate generates meaningful revenue, that expanded capability potentially translates to tens of thousands of dollars in incremental monthly revenue.
The other thing our flow health monitor catches is technical failures. Flows break for reasons that have nothing to do with copy or design: an integration stops firing, a trigger condition changes, or a filter accidentally excludes too many subscribers. At our scale, that kind of silent failure used to go undetected for weeks. Now, it shows up in an alert the following Sunday.
3 things we would do differently
We’re a personal care brand, not a software company. We don’t have a big stable of engineers. Building this health monitor took real effort, and there are 3 things we would approach differently now if we were doing it over again:
- Threshold setting: We spent a lot of time calibrating what counts as a real alert vs. noise. Too sensitive, and the dashboard floods with warnings that don’t require action. Not sensitive enough, and real problems slip through. Getting that calibration right took a lot of iteration, and we’re still tuning it.
- The scoring system: We built our own health score for each flow, which is useful, but the methodology was somewhat arbitrary at the start. We’ve refined it over time, but if we were starting today, we would spend more time defining exactly what a healthy flow looks like before writing the first line of logic.
- Coverage beyond flows: We started with flows, and now we have a separate module for campaigns. If we were building again from scratch, we would design the architecture to handle both from day one rather than treating them as separate builds.
What full visibility does for your email program
Your Klaviyo program may be under-optimized because your visibility needs improvement. And, of course, you can’t fix what you can’t see. If you’re like us, managing hundreds of emails across dozens of flows, the things you can’t see vastly outnumber the things you can.
The combination of Klaviyo's data, Claude's analysis, and a purpose-built dashboard changes that ratio. But it doesn't replace the judgment calls. A human still reviews every alert, decides which tests to run, and approves every change.
What it does mean is that you can make those judgment calls on all 1,000+ of your flow emails, not just your top 3 flows.
If building something like this yourself isn't realistic for your team, this is close to what Klaviyo Composer's Flow Audit already does for you, right inside Klaviyo. Ask it how a flow is performing, and it walks the sequence step by step, points to where customers are dropping off, and explains why in plain language, grounded in your actual segments, catalog, and performance history.
What my team put together took real engineering effort, custom thresholds, and a dashboard to maintain over time. Composer gets you to a similar answer in one conversation, with no build required, so your time goes toward acting on what's underperforming instead of figuring out how you'd find it.
Composer still puts the decision in your hands: what to test, what to ship, and when. But if you want that kind of visibility without taking on the maintenance, it's already built in.
Working to scale your email program with a similar approach?




