I had an abandoned repo sitting on GitHub called kube-quota-watch. One commit, from 2021, no code. The kind of thing you start on a Sunday and never touch again.

I wanted to build something real in that space: a Kubernetes operator that ingests security and reliability signals, uses an LLM to enrich them, and helps a team act on them. The obvious move was to open the editor and start scaffolding.

I did not do that. The first week went into research, positioning, naming, and repo governance, with no feature code at all. This post is about why, and what it changed.

This is part 1 of a series about building Candor in the open. Later parts get into CRDs, controllers, LLM integration and the verification loop. This one is about everything that happens before any of that.


The competitors already existed, and they were good

The first useful thing research did was kill my assumption that this was an empty field.

There are at least four serious projects doing roughly this job:

  • k8sgpt and its operator. CNCF Sandbox, around 8k stars, actively maintained. Deterministic analyzers plus LLM explanation.
  • Kubernaut. The closest analogue to what I wanted. Alert, then LLM investigation, then remediation, with approval gates, OPA policies and audit trails. It has an effectiveness monitor that scores whether a fix actually worked, and the automated half of it has shipped.
  • HolmesGPT. Read-only investigation with strong runbook integration, and a CNCF Sandbox project. An SRE team running it in production put the lesson well: without runbooks, the model just guesses.
  • kagent. A generic agent runtime rather than a remediation product.

If you are about to build something and you find four maintained projects already doing it, you have two honest options. Stop, or find a reason to be better rather than just different. “Different” is easy and worthless. I needed to know whether there was a real gap.

So instead of reading their READMEs, I read their issue trackers.

Issue trackers are where the truth is

A README tells you what a project intends to do. The issue tracker tells you what it actually does to people who run it.

Four findings changed the design:

Cost is unbounded. k8sgpt-operator issue #769, still open at the time of writing: a user configured an hourly scan, and it produced 164 results while calling Bedrock roughly 9,300 times in three days, confirmed against CloudTrail. Issue #730 found the mechanism behind that class of problem: change detection was comparing result hashes, finding them identical, and updating anyway. That one has since been fixed. Issue #419, which proposed decoupling LLM request timing from the reconcile loop, sat for over two years and was closed with “Closing as stale. This has had no traction in 6+ months and is not on the current roadmap.”

Suppression does not exist. k8sgpt issue #372, “Exclude a list of known issues”, has been open since May 2023 and is still open. A maintainer’s reply, ten days in: “We have no concensus on the design yet, do you want to propose something first?” A contributor offered a draft proposal. It never landed. Users were still asking in 2025. A tool whose pitch is reducing alert fatigue, with no way to mute a known issue, recreates the fatigue it claims to solve.

Verification of remediation is partial. k8sgpt-operator’s AUTO_REMEDIATION.md is refreshingly honest about its own limits. It checks that a Deployment rollout completed and replicas are available before treating a finding as resolved. But it lists “targeted re-analysis that proves the original finding is resolved” as future work, and says plainly that “a missing Result remains the finding-resolution signal.” Absence of a complaint is not proof of a fix.

Cluster writes break GitOps shops. This one is structural rather than anecdotal. Flux and ArgoCD are designed to revert anything that drifts from Git. That is the entire point of them. So if your remediation patches the cluster directly, the GitOps controller will faithfully undo it, and now you have two automated systems fighting over the same object.

None of these are obscure. They are all public, all documented, all in the open. They just do not show up on a landing page.

The gaps became the spec

I ended up with eight of these, written down as a table: the evidenced gap on the left, my answer on the right. That table became the actual design document. Every architectural decision in Candor traces back to a row in it.

Four examples of how that translated:

Evidenced gapWhat I built instead
9,300 LLM calls in 3 daysContent-addressed fingerprinting. Hash the resource identity, the relevant state, and the analyzer verdict. The LLM runs once per distinct fingerprint, ever. Plus a hard budget ceiling per window.
Suppression unshipped since 2023A Suppression CRD that mutes a fingerprint, with a required reason and optional expiry. Exact by construction: if the underlying state changes, the fingerprint changes, and the finding comes back on its own.
Verification of remediation is partial at bestEvery finding carries a verification outcome (resolved, still present, recurred), re-checked on every reconcile of its source and published as a metric. Worth being precise here: Kubernaut has since shipped automated effectiveness assessment too, so this is a difference of shape rather than of existence.
Direct cluster writesThe default write path is a pull request against the GitOps repo, not a cluster mutation.

There is a trap here worth naming. It is easy to read a competitor’s issue tracker and come away feeling clever. That feeling is not worth much. What matters is whether the gap is structural or just unfinished. Suppression in k8sgpt is genuinely hard because they have no content hash to key it on, which is why the design still has no consensus after three years. Fingerprinting makes suppression almost trivial. That is the difference between “they have not got to it yet” and “their architecture makes this awkward”, and only the second one is a real reason to build.

The trust problem is the actual problem

The deeper the research went, the clearer it became that the blocker in this field is not capability. It is trust.

Charity Majors and Fred Hebert gave a talk at SREcon25 called “AIOps: Prove It!”. It is effectively an open letter asking vendors for data on how often their systems produce useful, actionable results. Building a feedback loop is not the same as publishing what it found, and as far as I can tell no OSS tool in this space has published that data, though Kubernaut now has the machinery to.

The numbers back that up. A survey of 696 practitioners run by The Register with NeuBird AI in April 2026 found 60% naming lack of trust as the top adoption barrier, well ahead of ROI, security and data quality, which each landed around 12% to 13%. 59% said they need near-perfect accuracy before they would adopt. And adoption reflects that: 73% were not using AIOps at all, 19% were piloting, and 8% had it in production.

There was one comment on the Hacker News launch thread for HyperProbe, an AI production-debugging tool, that I kept coming back to. A commenter described three incidents where confident but incorrect diagnoses burned significant debugging time, and concluded: “A single confident answer that’s wrong is worse than no answer, because it sends a human down a road with the agent’s credibility behind it.”

That sentence is why this project is built the way it is. Ranked hypotheses with confidence, never one asserted cause. A verification outcome on every finding. Self-metrics published as part of the product rather than as a nice-to-have. If the pitch is “trust this”, the burden of proof sits with the tool.

Naming it

My first draft was called KubeGuard. I checked before committing to it, which turned out to matter. The name collides with an AppsCode Kubernetes auth webhook, a separate Java security scanner on GitHub, and an arXiv paper titled “KubeGuard: LLM-Assisted Kubernetes Hardening via Configuration Files and Runtime Logs Analysis”. That last one is not just a name clash, it is the same idea in the same field, published while I was sketching mine.

The rename took an afternoon and cost nothing. Finding out after publishing would have cost a lot more.

I picked Candor because the word is the posture. The tool is meant to be honest about uncertainty and to publish its own accuracy. A name you have to explain is a bad name, but a name that states the thesis is doing real work.

Two practical notes if you are naming something. Search the arXiv papers and trademark registers, not just GitHub. And check whether the GitHub org or user name is free before you fall in love with it.

Governance before features

Here is the part that looks like procrastination and is not.

Before writing a single controller, the repo got:

  • Branch protection on main, with 10 required status checks
  • Required signed commits
  • Dependabot with auto-merge for patch updates
  • OpenSSF Scorecard running on a schedule
  • GOVERNANCE.md, SECURITY.md, CONTRIBUTING.md, CODE_OF_CONDUCT.md
  • Release automation: goreleaser, cosign signing, SBOM, and build provenance on the image
  • The Helm chart published to GHCR

That is a lot of setup for a repo with no features. Three reasons it went first.

The security posture has to predate the code. Candor is a security tool that consumes untrusted input and calls an LLM. If the supply chain around it is casual, nothing else it claims matters. A tool that tells you about your vulnerabilities, shipped from an unsigned build with no SBOM, is not a serious proposition.

Retrofitting governance is worse than starting with it. Turning on required signed commits after 200 commits means either rewriting history or living with an unsigned tail forever. Turning it on at commit zero costs nothing.

It sets the standard for the code that follows. Once CI blocks on lint, vet, tests across three Go versions, e2e, govulncheck, trivy-config, gitleaks and codeql, you cannot merge a lazy slice. The bar is set by machine and it does not negotiate.

There is one thing I would flag as a cost, not just a benefit. Required signed commits and a bot-automated repo interact badly, and I lost an afternoon to a pull request that showed every check green and still refused to merge, with GitHub declining to say why. That story is its own post.

What “AI-first” actually meant here

“AI-first” usually means the AI is the product. That is not what it meant for this project.

Every meaningful decision in the first week was about constraining the model rather than showcasing it: how to avoid calling it, how to cap what it costs, how to check whether its output was right, how to make sure a bad answer cannot execute anything. The LLM is the most expensive, least predictable component in the system, and the architecture is mostly scaffolding to keep it bounded.

That is a design stance you can only take before you write code. Once the model call is sitting in the middle of your reconcile loop, adding a budget around it is surgery. Deciding on day one that nothing reaches the model without passing a fingerprint gate is free.

The research week produced no commits. It produced the constraint the rest of the project is built around. Whether that was the right trade is something the later parts, and anyone who runs this, are better placed to judge than I am.


Part 2 gets concrete: the CRDs, the provider pattern, and the fingerprinting mechanism that all of this depends on. The code is at teerakarna/candor.