SIGN IN SIGN UP

storage: kill broker process if log_reader detects a corrupt segment

In cases where a segment file had a zeroed batch header mid-file, the
log reader would return the batches before it and then report a clean
end of stream at the broken header. The offset translator, seeing a
successful read that stopped short of the log tip, would then throw out
of consensus::start(), causing the raft group to never register on that
shard.

log_reader should fail there instead of passing the fault up as a
success. The stream passed to the continuous_batch_parser in
log_segment_batch_reader::initialize is limited by the filesize at
construction, which will match get_stable_offset(). recover_segments
sets stable_offset for each segment from the index (written only after
flushing the appender) or from recovering the segment. Thus, reads up
through the segment's stable offset must succeed -- invalid or missing
batches indicate either a redpanda bug or a fault in the underlying
storage. In either case, we don't want to propagate that state to
clients or other brokers.

consume()'s result<size_t> will report success if any data is read (and
for some errors if no data is read). We need to use error() to see the
actual error result. We also need a bit of additional information to
distinguish a correct end_of_stream from a premature one -- capture
_stable_at_stream_start at continuous_batch_parser construction time
in initialize() and use it to check whether an end_of_stream from
consume() occurred prematurely. The other errors are automatically
fatal as they indicate invalid batches in a range that must be valid.

By default, detecting a fault in a segment kills the broker process as
continuing to operate with a corrupt storage layer is dangerous. Setting
storage_abort_on_corrupt_segment to false will disable this abort and
allow the broker to tolerate the problem. It is a node property because
it must be settable on a broker that will not start.

The following read paths may hit this case:

- kafka fetch, including fetch from a follower
- raft recovery, feeding a lagging follower from the leader's log
- offset translator startup sync
- tiered storage upload
- suffix truncation's offset-to-filepos scan
- compaction, via create_segment_full_reader

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
S
Samuel Just committed
125b15a5ecff13d0f73ef474aa65d309102f4c57
Parent: 19facd6