B2B SaaSSecurity operationsHuman oversight of AIDesktop web app

Harrier

An AI agent closes security alerts for forty client companies at once. I designed the console that tells the analyst how far to trust it, client by client.

Harrier · console, prototype1440 × 900 @2x00:00 / 00:00
The problem
Managed detection providers want AI agents to take more of the alert queue, but the provider carries the liability for every wrong close. The whole category sells “fewer cases reach the human”. Nobody shows the human the queue that remains, or how much the agent has earned the right to decide.
The goal
Let one analyst cover more client companies without more reversals. The working target: 40 percent more tenants per analyst at a flat reversal rate, and a verdict in under four minutes.
What I did
Read 13 competitors live and four trust benchmarks from outside security, then took a two-screen take-home brief through the whole method: research to handoff, a coded system and a Figma library at two widths.
What it gave
14 screens and 66 state pages in eight days, a 75-component system checked by automated rules, and a handoff that two agents with no context used to add a feature on their own.
RoleProduct design, soloTimeline8 days, Aug 2026Built withHTML · CSS · FigmaStage11 of 12journey covered by flows
01 The problem

The category sells fewer cases. Nobody shows the ones left.

An analyst at a managed detection provider watches forty or more client companies, each one a separate tenant with its own normal. AI agents now triage the alerts, and every vendor makes the same promise: fewer cases reach the human. But the provider, not the vendor, carries the liability when an agent closes a real intrusion as noise. So the provider cannot safely hand the agent more work without a way to show, client by client, how much latitude it has earned.

The analyst arrives already distrustful. In the SANS 2025 SOC survey, AI and machine learning tools rank last in analyst satisfaction. A console that asks for trust will not get it. A console that shows its evidence might.

The gap narrowed twice as I read the market. One vendor already sells earned autonomy; another sells it per tenant. What was left is the thing nobody does: make the fleet of earned trust readable where the analyst works. That is a design claim, not a capability claim, which is exactly the kind a designer can own.

02 The bet

Autonomy is earned per client, and it is always on screen.

The agent is called Clerk, and the name is the contract: a clerk prepares the case file, the judge rules. Clerk correlates signals, investigates and files a verdict. The analyst accepts, amends or rejects it, and every rejection is a signal that tunes how much Clerk may do alone on that tenant.

The best answers to “how much do I trust the machine” were outside security, so four benchmarks came from there. Aviation's flight mode annunciator separates what is armed from what is active. A rain forecast states its claim, its area and its window. A self-driving car shows what it sees. A chess engine shows its line, not just its move. And a study of explainable AI in security, a survey of 248 people and 24 interviews, found that trust follows the quality of the evidence, not the accuracy number: more explanation is not more trust.

A confidence number means nothing until it says about what, over what, and out of how many.
03 The product

Every row is a decision, not a record.

The console is a split pane. The queue sits on the left; on the right, at rest, is the fleet: every tenant with the autonomy Clerk currently holds there. Open a case and the right pane becomes the case file, with the verdict, the evidence and a provenance strip saying where each fact came from. A chat-first layout was rejected because it inverts the contract: the analyst would be asking the clerk, instead of ruling on its work.

The Harrier case queue: seven cases with severity, tenant, Clerk's verdict and age, beside the fleet pane listing what Clerk may do alone at each tenant.
Queue and fleet. The fleet is what the right pane shows when no case is open.
The Harrier case file for Larkfield Logistics: what Clerk filed, what happened, evidence with source links, and accept, amend, reject and escalate.
Case file. The verdict, the evidence and where each fact came from.

The shift brief answers the handover problem that none of the thirteen products addresses. Research on shift handovers found that handover is signposting, not a document, so the brief is a short list of what changed and where to look, not a report. Density is the feature: Archivo replaced Space Grotesk because it sets the same text 11 percent narrower, and the theme is dark by default for night shifts, against the reading research, with a light theme beside it.

Harrier tenant detail: one client's latitude, what is normal there, and what Clerk may do at this tenant.
Tenant. How much Clerk may do alone here, and why.
The Harrier shift brief: what waits and what moved this shift, with the handover panel.
Shift brief. Signposts, not a report.
The Harrier reject dialog: a checklist of reasons over the dimmed queue.
Reject. The reason is what tunes Clerk on this tenant.
The Harrier design system overview: primitives, roles, usage rules, components and patterns.
The system, checked by automated rules.
The Harrier case file on a phone, with the tenant's latitude, the timeline, evidence and an escalate button.
Escalating a case on a phone: sent to the SOC lead with a handover note.
Away from the desk. On a phone the job narrows to one: read the case and hand it up.
04 Decisions

Five calls, and what each one cost.

D1

Autonomy per tenant, not one global threshold.

One global rule would put a regional bank and a dental practice under the same latitude. Each tenant earns its own.

CostForty dials to explain instead of one.

D2

The fleet is the resting state, not a dashboard behind a tab.

The analyst sees the state of trust every time a case closes, which is where the calibration happens.

CostLess screen for the case itself when nothing is open.

D3

Confidence states its claim, scope, window and count.

“92 percent” is replaced by what the number is about, over which tenant, over what period, and out of how many cases.

CostLonger labels in a table that is already dense.

D4

Dark by default, against the reading research.

Most reading studies favour dark text on light. This product is read at 3 a.m. on a 24/7 floor, so dark leads and light is a first-class second theme.

CostSlower reading in long sessions for anyone who keeps the default.

D5

An honest failure on record.

One of the five design principles is that rejecting should be as easy as accepting. It still fails: reject costs four taps against accept's two, and the principle says so in its own file.

CostNothing yet. It is the first thing to fix.

05 What it gave

A complete console in eight days, and the targets it is built to move.

Shipped

8 daysFrom the first research note to handoff
66State pages across 14 screens, plus 62 grey wireframes
75Components and 4 patterns, rule-checked by Playwright
2Agents with no context who built a feature from the docs alone

Built to move

+40%[?] Tenants per analyst at a flat reversal rate
< 4 min[?] From case open to verdict
60%[?] Of tenants above entry autonomy after 90 days
50%[?] Of rejections that lead to a tuning change within 14 days

The targets are hypotheses, and the research marks each one as a hypothesis with its baseline unknown. The planned first test is eight to ten analysts working real cases, which is where every one of these numbers would start to be measured.

06 Still open

Untested with an analyst.

[?] Does a readable fleet actually calibrate trust. The taxonomies, the confidence grammar and the autonomy levels are all argued from research and none has met a working analyst yet. The test is written: eight to ten tier-2 analysts, real cases, and a measure of how often they overrule Clerk where they should and leave it alone where they should.

07 Artifacts

The whole method, public.

Next case

Fire Serpent