Part 1 opened with a talk called “AIOps: Prove It!”, an open letter from Charity Majors and Fred Hebert asking vendors selling AI to SREs for data on how often their systems produce useful, actionable results.

That question sat behind every decision in Part 2 and Part 3. This part tries to answer it.

The short version: a tool that tells you what is wrong, and never checks whether it stayed wrong, is asking for trust it has not earned. So verification is not a feature bolted on at the end here. It is the output.


The bug that made verification real

The verification loop existed on paper before it worked. Findings carried a status.verificationOutcome of StillPresent, Resolved or Recurred. The problem was that nothing could ever set Resolved.

Here is why, and it is a nice example of a bug you only find by thinking about the whole path rather than the component.

Candor’s Trivy provider reads VulnerabilityReport resources and translates them into internal signals. If a report had no vulnerabilities, the translator returned “nothing to see here” and the pipeline discarded it. Reasonable in isolation. But consider what that means end to end:

  1. A workload has a CRITICAL vulnerability. Trivy reports it. Candor creates a Finding.
  2. Someone patches the image. Trivy now reports zero vulnerabilities.
  3. The translator sees a clean report, produces nothing, and the signal is discarded.
  4. The Finding sits there forever, still saying CRITICAL.

The tool would have told you something was broken, watched it get fixed, and never mentioned it. Exactly the failure the “prove it” critique is about.

The fix is small and slightly counterintuitive: a clean report still produces a signal.

// A workload-linked report with zero vulnerabilities at every severity still produces a Signal
// (ok=true), with Severity left empty: signal.AtLeast never matches an unrecognised severity
// [...] so a previously-vulnerable workload that's now clean gets its Finding marked Resolved

An empty severity can never clear any policy threshold, so the signal is always filtered out. That sounds pointless until you notice that “filtered out” and “never existed” are different states. The filtered path already knows how to look for an existing Finding and resolve it. The empty-severity signal is a carrier: it reaches that path and says “this source is clean now”, and the Finding gets marked Resolved instead of rotting.

The lesson I would take from it: when a component correctly returns “nothing”, check what the caller loses by never being told.

Outcomes get recomputed, not asserted

The other half is that verification is not a one-time judgement.

Every reconcile of a Finding’s source recomputes the outcome. The interesting case is the transition table, which is three lines and encodes the whole idea:

  • New or ongoing problem: StillPresent
  • Filtered signal and an open Finding exists: Resolved
  • Signal is back, and the previous outcome was Resolved: Recurred

That last one is the one that matters for trust. A tool that only tracks resolved-versus-open can report a 100% fix rate while the same problem comes back every week. Recurred is the state that makes the claim falsifiable, and it is published as a metric:

candor_verification_transitions_total{outcome="resolved"}
candor_verification_transitions_total{outcome="recurred"}

The resolved-to-recurred ratio over time is the number. Not a benchmark score, not a self-assessment. Just: of the things this tool said were fixed, how many stayed fixed.

Worth being precise about what that is not. It is not a calibrated accuracy score against ground truth, and I have deliberately not claimed one. Measuring whether a hypothesis was correct requires labelled data this project does not have. Measuring whether a finding stayed resolved requires only honesty about what already happened. The second one is weaker, and it is real.

Suppression without a second source of truth

Alert fatigue tools usually grow a mute list, and mute lists rot. Somebody silences an alert in March, the underlying situation changes in June, and nothing resurfaces because the mute was keyed on an alert name.

Candor mutes a fingerprint:

for i := range list.Items {
	s := &list.Items[i]
	if s.Spec.Fingerprint != fingerprint {
		continue
	}
	if s.Spec.ExpiresAt != nil && !s.Spec.ExpiresAt.After(now) {
		continue
	}
	return s, nil
}

Exact string match, nothing clever. The cleverness was already spent in Part 2: the fingerprint is a hash of the content, so if the situation genuinely changes, the fingerprint changes, the suppression no longer matches, and the Finding resurfaces by itself.

There is no “is this suppression still relevant” logic, no expiry sweep required for correctness, no second source of truth to drift. A reason is mandatory on the CRD, because a silent mute with no recorded justification is how teams end up afraid to delete anything.

The dashboard is the deliverable

For a tool whose pitch is publishing its own accuracy, the dashboard is not decoration. It is the artifact the claim is made in. It ships in the Helm chart, labelled for the kube-prometheus-stack sidecar to pick up automatically.

It also took several rounds to get right, and the failures were instructive.

The first version answered the wrong question. It led with Candor’s own cost and efficiency metrics: calls made, calls avoided, budget consumed. Which is to say it opened with “look how efficient we are” rather than what an operator actually opens a dashboard to find out, which is what needs attention right now. Findings by severity moved to the top; the self-observability moved below the fold. If your dashboard’s first row is about the tool rather than about the cluster, it is a brochure.

Grafana’s documentation and Grafana’s behaviour disagreed. A stat panel configured to show multiple values with a default range query rendered one box per timestamp sample, producing a wall of repeated numbers. The fix was instant: true on the targets. Separately, orientation: vertical did not stack panels the way the docs describe. I checked the live instance rather than assuming I had misconfigured it, and then stopped relying on that setting entirely in favour of explicit grid positions. When documented behaviour and observed behaviour disagree, the observed behaviour is the one users will get.

increase() returns fractions. Prometheus extrapolates, so a counter that went up by 3 can render as 3.4 “findings resolved”. Wrapping in floor() was the fix. Claiming a fractional number of resolved security findings undermines the entire point of a dashboard whose job is to be credible.

How do you test a dashboard?

This turned out to be the question with no good off-the-shelf answer, and it needed two separate solutions.

In CI, a job runs a real Grafana as a service container, registers a Prometheus datasource so the dashboard’s template variable resolves, imports the JSON through the real API, and asserts that every panel survived the round trip:

got=$(curl -sf -u admin:admin http://localhost:3000/api/dashboards/uid/candor-overview | jq '.dashboard.panels | length')
want=$(python3 -c "import json; print(len(json.load(open('charts/chart/files/grafana-dashboard.json'))['panels']))")
if [ "$got" != "$want" ]; then
  echo "panel count mismatch after import: got $got, want $want"
  exit 1
fi

The comment above that job in the repo is the important part, because it states what the test does not prove:

# Grafana's dashboard-save API is lenient - it does NOT reject an invalid panel type or a
# broken query, only structurally unimportable JSON (verified locally against a real
# instance before writing this job).

I found that out by trying it. A test whose limits you have not established is a test you do not know the value of.

For the human check, there is a local compose stack: a fake metrics generator, Prometheus, and Grafana with both the datasource and the dashboard auto-provisioned. No cluster, no operator, no real scan data. Open localhost:3000 and look at it.

One detail there is worth more than it seems. The first metrics generator drove “findings currently open” and “findings resolved” as two independent random walks, so the preview could show zero outstanding findings next to thirty resolved ones. Nothing the real system can produce, since both numbers come from the same lifecycle. It now simulates one small state machine per namespace and severity, and the invariant that holds in production holds in the preview:

open == still_present + recurred - resolved

A preview that can display impossible states will eventually convince you a real bug is fine, or that a fine layout is broken.

Getting it out of the cluster

The last piece is the boring one that determines whether anyone sees any of this. Findings and digests go out through a generic webhook:

// Package notify is Candor's generic webhook sink (docs/design.md: "one code path covers
// Slack/Teams/PagerDuty/anything"). It knows nothing about any particular chat or paging vendor -
// it POSTs a JSON payload to a URL and nothing more.

No Slack SDK, no PagerDuty event schema, no per-vendor integrations to maintain as five vendors independently change their APIs. One JSON POST. Formatting is the receiving end’s problem.

That is a deliberate trade: slightly more setup for the user, dramatically less maintenance surface, and no vendor gets to be a first-class citizen because they were integrated first. A periodic digest goes out the same path, summarising findings raised, resolved and recurred over the window. That digest is the artifact the “prove it” critique actually asks for, arriving on its own rather than waiting for someone to go looking.

Sending is always best-effort. A broken webhook endpoint never blocks a reconcile, it increments a failure counter. Notification is not important enough to be a dependency of correctness.

Where this leaves the question

Back to the open letter. Can Candor answer “how often does your system produce useful, actionable results”?

Not yet. Here is exactly where it has got to.

Measured today. Findings raised, resolved and recurred, by severity, over any window. LLM calls made, and calls avoided by the fingerprint gate. Budget consumed against the ceiling. Whether the operator is currently degraded. All in Prometheus, all in the shipped dashboard.

Not measured. Whether a given hypothesis was correct. That needs labelled ground truth and a fixture suite. The resolved-to-recurred ratio tells you whether a fix held. It tells you nothing about whether the diagnosis was right in the first place, and I am not going to imply otherwise by putting the two next to each other and letting the reader assume.

Untested at scale. Everything above has run against test fixtures and a local cluster. Nobody has pointed this at a busy production cluster for three months and come back with numbers. Until someone has, the cost claims are architectural arguments, not field results.

I do not have a conclusion for you, and I think reaching for one would be the wrong move in a project called Candor. The honest position is that this is work in progress with its measurements published and its gaps written down.

What I would rather have than a conclusion is scrutiny. The design doc carries the full evidence base, including citations I got wrong and corrected after re-checking them. The dashboard is in the chart. The tests that enforce the cost claims are in the repo and you can read what they do and do not prove. If the fingerprint approach breaks in a way I have not thought of, if the resolved-to- recurred ratio is a worse proxy than I think it is, if there is a whole use for this I have not considered, those are more useful to me than agreement.

Issues and disagreement welcome. That is not a closing pleasantry, it is the point of publishing the numbers in the first place.


That is the series. Four parts: the decisions before the code, the operator pattern under a cost constraint, putting a model behind a gate and a budget, and measuring what came out. The code, the design doc with its full evidence base, and the dashboard are all at teerakarna/candor.