Harrier
An AI agent closes security alerts for forty client companies at once. I designed the console that tells the analyst how far to trust it, client by client.
- The problem
- Managed detection providers want AI agents to take more of the alert queue, but the provider carries the liability for every wrong close. The whole category sells “fewer cases reach the human”. Nobody shows the human the queue that remains, or how much the agent has earned the right to decide.
- The goal
- Let one analyst cover more client companies without more reversals. The working target: 40 percent more tenants per analyst at a flat reversal rate, and a verdict in under four minutes.
- What I did
- Read 13 competitors live and four trust benchmarks from outside security, then took a two-screen take-home brief through the whole method: research to handoff, a coded system and a Figma library at two widths.
- What it gave
- 14 screens and 66 state pages in eight days, a 75-component system checked by automated rules, and a handoff that two agents with no context used to add a feature on their own.
The category sells fewer cases. Nobody shows the ones left.
An analyst at a managed detection provider watches forty or more client companies, each one a separate tenant with its own normal. AI agents now triage the alerts, and every vendor makes the same promise: fewer cases reach the human. But the provider, not the vendor, carries the liability when an agent closes a real intrusion as noise. So the provider cannot safely hand the agent more work without a way to show, client by client, how much latitude it has earned.
The analyst arrives already distrustful. In the SANS 2025 SOC survey, AI and machine learning tools rank last in analyst satisfaction. A console that asks for trust will not get it. A console that shows its evidence might.
The gap narrowed twice as I read the market. One vendor already sells earned autonomy; another sells it per tenant. What was left is the thing nobody does: make the fleet of earned trust readable where the analyst works. That is a design claim, not a capability claim, which is exactly the kind a designer can own.
Autonomy is earned per client, and it is always on screen.
The agent is called Clerk, and the name is the contract: a clerk prepares the case file, the judge rules. Clerk correlates signals, investigates and files a verdict. The analyst accepts, amends or rejects it, and every rejection is a signal that tunes how much Clerk may do alone on that tenant.
The best answers to “how much do I trust the machine” were outside security, so four benchmarks came from there. Aviation's flight mode annunciator separates what is armed from what is active. A rain forecast states its claim, its area and its window. A self-driving car shows what it sees. A chess engine shows its line, not just its move. And a study of explainable AI in security, a survey of 248 people and 24 interviews, found that trust follows the quality of the evidence, not the accuracy number: more explanation is not more trust.
A confidence number means nothing until it says about what, over what, and out of how many.
Every row is a decision, not a record.
The console is a split pane. The queue sits on the left; on the right, at rest, is the fleet: every tenant with the autonomy Clerk currently holds there. Open a case and the right pane becomes the case file, with the verdict, the evidence and a provenance strip saying where each fact came from. A chat-first layout was rejected because it inverts the contract: the analyst would be asking the clerk, instead of ruling on its work.


The shift brief answers the handover problem that none of the thirteen products addresses. Research on shift handovers found that handover is signposting, not a document, so the brief is a short list of what changed and where to look, not a report. Density is the feature: Archivo replaced Space Grotesk because it sets the same text 11 percent narrower, and the theme is dark by default for night shifts, against the reading research, with a light theme beside it.






Five calls, and what each one cost.
Autonomy per tenant, not one global threshold.
One global rule would put a regional bank and a dental practice under the same latitude. Each tenant earns its own.
CostForty dials to explain instead of one.
The fleet is the resting state, not a dashboard behind a tab.
The analyst sees the state of trust every time a case closes, which is where the calibration happens.
CostLess screen for the case itself when nothing is open.
Confidence states its claim, scope, window and count.
“92 percent” is replaced by what the number is about, over which tenant, over what period, and out of how many cases.
CostLonger labels in a table that is already dense.
Dark by default, against the reading research.
Most reading studies favour dark text on light. This product is read at 3 a.m. on a 24/7 floor, so dark leads and light is a first-class second theme.
CostSlower reading in long sessions for anyone who keeps the default.
An honest failure on record.
One of the five design principles is that rejecting should be as easy as accepting. It still fails: reject costs four taps against accept's two, and the principle says so in its own file.
CostNothing yet. It is the first thing to fix.
A complete console in eight days, and the targets it is built to move.
Shipped
Built to move
The targets are hypotheses, and the research marks each one as a hypothesis with its baseline unknown. The planned first test is eight to ten analysts working real cases, which is where every one of these numbers would start to be measured.
Untested with an analyst.
[?] Does a readable fleet actually calibrate trust. The taxonomies, the confidence grammar and the autonomy levels are all argued from research and none has met a working analyst yet. The test is written: eight to ten tier-2 analysts, real cases, and a measure of how often they overrule Clerk where they should and leave it alone where they should.
Fire Serpent