SIGN IN SIGN UP

feat(fleet): print a fleet summary after a multi-context scan [LFX 2026] (#3929)

* feat(fleet): print a fleet summary after a multi-context scan

--fleet-report writes a combined report as JSON and nothing else, so
reading the result of a fleet scan means piping it through jq. Four
rounds of aggregation have landed and none of it is visible in a
terminal.

This prints a summary at the end of a fleet scan: how each cluster did,
how the fleet did, and where the clusters disagree.

It is deliberately not an IPrinter. That interface takes one cluster's
OPASessionObj, which a fleet report is not and cannot be made into, so
implementing it would mean pretending a fleet is a cluster. What is
shared instead is everything below the interface. The printer lives in
the same package as the pretty printer so it uses getColor and
severityRank directly rather than re-exporting or copying them, the
tables are built in the house style, and the text goes through the
existing cautils display helpers. The severity scales line up already:
severityRank returns the same ordinals apis.SeverityCritical and friends
use, so colouring a divergence row needs no new code at all.

The control matrix is not printed. It is one column per cluster and
stops being readable somewhere around five of them, and it is already in
the JSON for anyone who wants it.

Three things the summary refuses to round off. A cluster that could not
be scanned keeps its row, because a table that quietly omits it makes
the fleet look smaller and healthier than it is. A measurement nobody
took renders as a dash rather than 0%, for the same reason the report
leaves the field absent: zero reads as fully non-compliant when the
truth is never measured. And the fleet score never appears without the
count of clusters behind it, because a number printed alone invites more
confidence than it has earned.

Divergence rows are sorted worst severity first, so the row worth acting
on is at the top. That sort is local to the printer: the report itself
stays ordered by control ID so two runs of an unchanged fleet still
compare equal, and the printer works on a slice of pointers so it never
reorders the report it was handed.

The summary prints whether or not the file was written. A run that
scanned every cluster and then failed to save the result should still
say what it found, since the summary is the only place the operator
would otherwise see it. A plain --kube-contexts run without
--fleet-report prints nothing extra, because no aggregate is built there
at all, which is what keeps that run's memory flat.

Also corrects the CoverageGap comment. It said the flag is set when
another cluster reached a verdict, but the predicate counts a deliberate
skip too, which this package treats as a decision rather than a verdict.
The behaviour is unchanged and already pinned by the wire-format test,
only the wording was loose.

Also adds the test for the timestamp invariant this package has been
relying on in prose. The comment on fleetScan says contexts are scanned
one at a time because k8sinterface's connection state is process global
with no locking around it, but nothing checked it. The test records when
each context's scan began and ended, and which context was live while it
ran, then asserts the windows never overlap and that the process was
pointed at the right cluster throughout.

It is written to be able to fail. Dropping EnterClusterContext trips it,
and so does rewriting the loop to scan contexts concurrently, which was
checked both ways rather than assumed. The recorder takes a mutex for
the same reason: without one, a concurrent rewrite would trip the race
detector inside the test helper instead of reporting the invariant that
actually broke.

Signed-off-by: Ady0333 <adityashinde1525@gmail.com>

* fix(fleet): don't report agreement a fleet never reached

BuildDivergence drops any control fewer than two clusters reached an
outcome on, so an empty divergence means one of two different things.
With two clusters reporting it means they agreed. With fewer it means
nothing was ever compared, and the summary said they agreed anyway.

That turned the worst runs into the cleanest looking ones. A run where
every context was unreachable, or where one context succeeded and the
rest failed, printed a green "Every cluster agreed on every control it
ran" over a fleet nobody had measured. It is the same mistake the rest
of this file exists to avoid: an unmeasured result must not read as a
good one.

The count comes from the control matrix rather than from how many
clusters were scanned, because the matrix is what the divergence is
computed over. A cluster whose scan produced no control results
contributes no cells and cannot agree or disagree with anyone, so
counting it would put the vacuous message back for that case.

Signed-off-by: Ady0333 <adityashinde1525@gmail.com>

* fix(fleet): decide comparability per control, not across the matrix

The previous guard counted the union of cluster IDs across every row of
the matrix. That is the wrong shape, because BuildDivergence works one
control at a time and drops any row fewer than two clusters reached an
outcome on.

So two clusters that scanned disjoint control sets, A reporting only
C-0016 and B only C-0038, gave a union of two clusters and an empty
divergence. The summary read that as agreement and said every cluster
agreed on every control it ran, when in fact no control had been
compared at all.

Comparability is now counted per control row: how many controls more
than one cluster reported on. When none qualify the summary says so
instead, and distinguishes the three ways that happens, nobody
reporting, one cluster reporting, or several clusters sharing no
control. A fleet that does share a control and agrees on it still gets
told so, which is the case worth keeping.

The regression test builds the divergence with BuildDivergence rather
than writing it by hand, since the bug lives in the interaction between
the two rather than in either one alone.

Signed-off-by: Ady0333 <adityashinde1525@gmail.com>

---------

Signed-off-by: Ady0333 <adityashinde1525@gmail.com>
A
Aditya Shinde committed
4910049909942e0342919666c74aa5cb7ace34ac
Parent: 3fb3215
Committed by GitHub <noreply@github.com> on 9/25/2026, 12:00:10 PM