storage: kill broker process if log_reader detects a corrupt segment
In cases where a segment file had a zeroed batch header mid-file, the log reader would return the batches before it and then report a clean end of stream at the broken header. The offset translator, seeing a successful read that stopped short of the log tip, would then throw out of consensus::start(), causing the raft group to never register on that shard. log_reader should fail there instead of passing the fault up as a success. The stream passed to the continuous_batch_parser in log_segment_batch_reader::initialize is limited by the filesize at construction, which will match get_stable_offset(). recover_segments sets stable_offset for each segment from the index (written only after flushing the appender) or from recovering the segment. Thus, reads up through the segment's stable offset must succeed -- invalid or missing batches indicate either a redpanda bug or a fault in the underlying storage. In either case, we don't want to propagate that state to clients or other brokers. consume()'s result<size_t> will report success if any data is read (and for some errors if no data is read). We need to use error() to see the actual error result. We also need a bit of additional information to distinguish a correct end_of_stream from a premature one -- capture _stable_at_stream_start at continuous_batch_parser construction time in initialize() and use it to check whether an end_of_stream from consume() occurred prematurely. The other errors are automatically fatal as they indicate invalid batches in a range that must be valid. By default, detecting a fault in a segment kills the broker process as continuing to operate with a corrupt storage layer is dangerous. Setting storage_abort_on_corrupt_segment to false will disable this abort and allow the broker to tolerate the problem. It is a node property because it must be settable on a broker that will not start. The following read paths may hit this case: - kafka fetch, including fetch from a follower - raft recovery, feeding a lagging follower from the leader's log - offset translator startup sync - tiered storage upload - suffix truncation's offset-to-filepos scan - compaction, via create_segment_full_reader Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
S
Samuel Just committed
125b15a5ecff13d0f73ef474aa65d309102f4c57
Parent: 19facd6