Notes

Detection, not prevention

How I shipped a production integrity control in a day, and published its limits before anyone asked.

Case study

Client
A federal health system
Role
Lead architect (as VP, Delivery & Technology at govSlackers)
Built with
Slack audit logs, Claude Code, Slackbot AI
Timeline
Live same day; running since September 2026

The problem

A shared reporting view in the client’s Slack workspace kept getting silently overwritten — three times in under two weeks. The underlying data was never at risk. What broke was the shared view everyone reported from, and each time we found out because the customer told us, which is the worst way to find out.

The client’s executive asked the direct question: how do we prevent this?

The answer nobody wanted

This wasn’t a misconfiguration. In Slack Lists, edit access to a list is edit access to that list’s views. There is no separate view-level permission to tighten, so anyone who can legitimately edit the data can legitimately destroy the shared view. That’s the product’s design, not a setting.

Prevention wasn’t available. I told the executive so in writing and offered what was: detection. Catch it fast, route it to an owner, and stop learning about it from the customer.

Before proposing it, I measured the problem: 645 view changes in the prior 30 days, about 22 a day. That number is what made detection a proportionate answer rather than a retreat.

What I shipped

Ticket opened at 10:55. Working build by 11:32. Live in production by 9:06 that evening.

I built the first version with Claude Code. The monitor watches Slack’s audit log. When a monitored view changes, it opens an urgent ticket linked to the exact view, names the person who changed it, and posts to a dedicated channel. Changes by trusted admins get a quiet confirmation instead of an alert — a control that cries wolf gets muted within a week, and a muted control is worse than none.

Limits, disclosed up front

The audit log tells you who, which view, and when. It doesn’t tell you what changed: with no API for views, there’s no before-and-after and no diff. It’s a change notification, not a change record. Remediation stays manual, so I wrote the recovery procedure into the ticket before shipping — the gap got a workaround, not just a caveat.

The next morning the first scheduler proved unreliable: one run in about seven hours instead of twenty-eight. I moved execution, and disclosed the degraded latency and the overnight coverage gap in the same update that announced the monitor was live, naming the cadence I was working toward rather than implying we’d reached it. Detection now runs every hour on Claude Routines.

I was also clear that this is a stopgap. The durable fix belongs in the product, so I filed it with Slack as two feature requests: version history with restore, and row-scoped permissions with field-level locking.

Results

First ten days in production:

Measure Result
View changes caught 31
From non-admin users 20
From trusted admins 11
Tickets opened 10
Distinct users triggering alerts 14
Steady-state rate ~2 ticketed changes/day
Detection cadence Hourly

Three things matter more than the numbers.

The load went to the right person. Alerts were auto-assigning to me while a teammate did the actual resolving. I scoped a one-line fix, pushed back when the agent implementing it claimed the behavior was already correct, shipped it, and verified it against a real incident the next morning.

It ran without me. The following week I was at a conference. The monitor operated unattended, opened seven incident tickets, and routed every one to the right owner. A control that only works when its author is watching isn’t a control.

Management sees the pattern, not just the alerts. A Slackbot AI skill and weekly task now gather the week’s incidents and summarize them in a Slack surface for management. The conversation moved from individual tickets to trends: who is changing views, how often, and whether it’s improving. Claude Code built the detector; Slackbot AI reports on it.

What this shows

Working inside someone else’s platform, the most valuable thing you can often do is establish precisely what it cannot do — then build the best available thing on the right side of that line. The failure mode is promising prevention, shipping something that mostly works, and letting the customer find the gap during an incident.

The pattern: separate what’s controllable from what isn’t, ship the controllable half fast, publish the ceiling so nobody over-trusts it, push the structural fix to whoever owns the product, and route the operational load to whoever should carry it.