Part 1 was about the week before any code: research, positioning, naming, governance. This part is the code.

The thing that makes this a useful way to learn the operator pattern is the constraint. Most operator tutorials reconcile something free. If your reconcile loop runs twice when it should run once, nothing happens, nobody notices. Here, an unnecessary reconcile can mean an unnecessary LLM call, and unnecessary LLM calls are exactly the failure mode documented in the incumbents: 164 findings turning into 9,300 model calls in three days.

So every design decision below is shaped by one question. When this loop runs again, what stops it spending money?


Two CRDs, and what goes in each

Candor ships three CRDs. Two matter here.

SignalPolicy is configuration: what to watch, and how strictly.

apiVersion: candor.dev/v1alpha1
kind: SignalPolicy
metadata:
  name: team-a-policy
  namespace: team-a
spec:
  providers: ["trivy"]
  minSeverity: HIGH

Finding is the output: one object per thing worth knowing about.

A SignalPolicy governs signals in its own namespace only. No namespace selector field, no label matching across the cluster. This follows the same convention as ResourceQuota and NetworkPolicy: the object’s own namespace is its scope. It is one less concept to learn, and it means RBAC works the way people already expect.

One small decision in that spec that paid off later. providers is a free-form string list, not a CRD enum:

// providers this policy enables, e.g. ["trivy"]. Deliberately a free-form string list rather
// than a closed CRD enum: each provider's own reconciler simply ignores policies that don't
// name it, so adding a new provider later never requires a schema migration here.
// +kubebuilder:validation:MinItems=1
Providers []string `json:"providers"`

An enum would have been tighter validation. It would also mean that adding a second provider later becomes a CRD schema change, which means a migration, which means a version bump on an API other people may already be running. Free-form strings push that cost to zero. Each provider’s reconciler ignores policies that do not name it.

The provider pattern: consume, do not bundle

Candor does not scan images. Trivy already does that, very well, and it has an operator that publishes results as VulnerabilityReport custom resources.

So Candor watches those resources. It does not embed Trivy, vendor Trivy’s Go module, or shell out to the Trivy binary. The entire integration is a controller watching a CRD that somebody else owns:

// SetupWithManager sets up the controller with the Manager. Watches VulnerabilityReport as an
// unstructured type - no compile-time dependency on Trivy Operator's Go module, matching the
// provider pattern's loose coupling.
func (r *Reconciler) SetupWithManager(mgr ctrl.Manager) error {
	u := &unstructured.Unstructured{}
	u.SetGroupVersionKind(GroupVersionKind)

	return ctrl.NewControllerManagedBy(mgr).
		For(u).
		Named("trivy-provider").
		Complete(r)
}

Note unstructured.Unstructured rather than an imported Go type. Importing Trivy Operator’s types would give compile-time safety, at the cost of pinning Candor’s build to Trivy’s module version forever. For a CRD Candor only reads a handful of fields from, unstructured access is the better trade. The provider’s job is small: read the report, translate it to an internal Signal, hand it off.

That internal type is the actual seam:

// Signal is one provider's deterministic reading of one piece of evidence about one Kubernetes
// resource. Never carries an LLM opinion - Summary is generated from the signal's own data, not
// inferred.
type Signal struct {
	Provider  string
	Severity  string
	Namespace string
	Kind      string
	Name      string
	Summary   string
	RefKind   string
	RefName   string
}

A provider produces a Signal. Everything after that happens once, centrally: policy filtering, fingerprinting, enrichment, budget accounting, notification. Adding Falco later means writing a translator, not reimplementing the pipeline. The value of a seam is not the abstraction, it is that the expensive logic behind it is written once.

Do not crash the manager for a CRD you do not own

Here is a failure mode worth knowing about if you build anything on the provider pattern.

If you register a controller watching a CRD that is not installed in the cluster, controller-runtime fails at startup. Not a warning. The manager does not come up, which takes down every other controller in the same binary with it.

For Candor that would be unacceptable. Trivy Operator is optional. A cluster without it is a supported configuration, not a broken one. So the CRD’s presence is checked before registration:

// CRDInstalled reports whether gvk's CRD exists in the cluster the given RESTMapper is talking
// to. Every provider watches a CRD it doesn't own (that's the point - see docs/design.md's
// provider pattern), so a cluster without that CRD installed must not crash the whole manager;
// callers use this to skip registering a provider's watch instead, and log why.
func CRDInstalled(mapper meta.RESTMapper, gvk schema.GroupVersionKind) (bool, error) {
	_, err := mapper.RESTMapping(gvk.GroupKind(), gvk.Version)
	if err == nil {
		return true, nil
	}
	if meta.IsNoMatchError(err) {
		return false, nil
	}
	return false, err
}

And at startup, the provider is skipped with an actionable log line rather than a stack trace:

VulnerabilityReport CRD not found - skipping the Trivy provider
(install Trivy Operator to enable it)

The distinction that matters in that function is meta.IsNoMatchError. “This CRD is not installed” is a normal condition and returns false, nil. Any other error, such as the API server being unreachable, is a real error and gets returned. Collapsing both into “false” would silently disable the provider during an API server blip, which is the kind of bug you find six months later while wondering why nothing was scanned on a particular Tuesday.

The fingerprint

This is the piece everything else in the project leans on.

The naive way to avoid redundant work is a timestamp or a TTL: do not re-analyse anything checked in the last hour. That breaks in both directions. Content that changed two minutes ago gets ignored, and content that has not changed in a year gets re-analysed on the hour, forever.

Candor keys on content instead:

func Fingerprint(sig Signal) string {
	h := sha256.New()
	for _, field := range []string{sig.Provider, sig.Namespace, sig.Kind, sig.Name, sig.Severity, sig.Summary} {
		h.Write([]byte(field))
		h.Write([]byte{0}) // separator, so ("ab","c") and ("a","bc") never collide
	}
	return hex.EncodeToString(h.Sum(nil))
}

That null separator is not decoration. Without it, hashing the concatenation of fields means ("ab", "c") and ("a", "bc") produce identical hashes. Rare, but it is a correctness bug in the one function the entire cost model depends on, and it costs one byte per field to eliminate.

Two fingerprints are stored on Finding.status: fingerprint is what the content hashes to now, and enrichedFingerprint is what it hashed to when the LLM last looked at it. The gate is the comparison between them:

func NeedsEnrichment(f *candorv1alpha1.Finding) bool {
	return f.Status.EnrichedFingerprint != f.Status.Fingerprint
}

Three lines, and they answer every case correctly. Never enriched, so enrichedFingerprint is empty and they differ: needs enrichment. Enriched, and nothing has changed: identical, skip. Enriched, but the content has since changed: they differ again, re-enrich. There is no separate cache to invalidate, because the Finding’s own status field is the cache. Nothing to expire, nothing to keep in sync, nothing extra to deploy.

This also quietly solves a problem that arrives two slices later. Suppression needs to resurface a muted finding when the underlying situation genuinely changes. Because suppression keys on the fingerprint, and the fingerprint changes when content changes, that behaviour falls out for free rather than needing its own mechanism.

Proving it, rather than claiming it

A README can claim “we only call the model once per distinct finding.” A test can enforce it.

// N more ingests of the identical signal. Every single one must report "no enrichment
// needed" - this is the actual cost claim: unchanged content costs nothing, N times over,
// not just once.
const n = 10
for i := range n {
	if _, err := Ingest(ctx, c, scheme, sig, nil); err != nil {
		t.Fatal(err)
	}
	if NeedsEnrichment(getFinding()) {
		t.Fatalf("iteration %d: unchanged content reported as needing enrichment - this is the exact k8sgpt failure mode (9,300 calls for 164 stable findings) this mechanism exists to prevent", i)
	}
}

That failure message is doing deliberate work. When this test breaks in two years, whoever is looking at it does not need to reconstruct why anyone cared. The reason is in the assertion.

The test ingests genuinely new content and asserts enrichment is needed. It marks the fingerprint enriched, simulating the model call that does not exist yet. Then it ingests the identical signal ten more times and asserts that not one of them reports work to do. Finally it changes the content and asserts the gate reopens.

Worth being precise about what this proves at this stage. There is no LLM in the codebase yet, so it cannot prove “one LLM call.” It proves that NeedsEnrichment makes exactly one true-to-false transition across N ingests of identical content, and reverses only on real change. When the model call lands in the next slice, it plugs into this gate unchanged, and the test that counts real calls is trivial to write because the mechanism it sits on is already proven.

I would rather ship a test that proves a smaller claim honestly than one that implies a bigger claim it cannot support.

What this bought

Three things worth carrying into any operator, not just this one.

Reconcile loops run more than you think. Design as though every loop will run thousands of times against unchanged state, because it will. The question is always what the second run costs.

Content addressing beats time addressing. A hash of what something is survives restarts, re-deploys, cache eviction and clock skew. A timestamp survives none of them.

Optional dependencies must be optional at runtime, not just in the README. If your operator watches a CRD it does not own, check it exists before you watch it.

Part 3 puts an actual model behind this gate, which is where the budget ceiling, the degraded mode and the “no API key is a supported configuration” stance come in. The code is at teerakarna/candor.