+ Add Minecraft MCA region creation and compression support - #412
Merged
Merged
Conversation
Hawkynt
force-pushed
the
feat/mca-region-write-support
branch
3 times, most recently
from
September 30, 2026 16:11
5fc41b2 to
8007649
Compare
MCA archives only exposed extraction despite having a concrete, independently addressable chunk layout. Their reader also discarded timestamps and treated newer compression types as stored bytes. Add clean-room region writing for all four in-region compression methods, decode raw LZ4, validate sector bounds, and carry chunk timestamps through listing and extraction. Reference: PaperMC SectorTool specification and Minecraft Wiki region layout. No implementation code was copied.
The first PR revision still described only legacy compression ids and could leave extracted chunk times at the extraction time. Document the raw LZ4 and external payload cases, preserve timestamp zero too, and guard the nullable filename API.
MCA compression id 4 uses LZ4-Java blocks rather than naked LZ4 blocks or LZ4 frames. Add the 64 KiB block framing, raw fallback, xxHash32 checksums, terminator, validation, and tests for both block forms. Reference: lz4-java LZ4BlockOutputStream and Minecraft Map Format write-up. The framing and constants were derived from format behavior; no external code was copied.
The encoder chooses a borrowed source block for raw fallback and an owned compressed array otherwise. Give that conditional an explicit read-only span target so both arms bind without relying on inference between Span and ReadOnlySpan.
Compression id 4 is framed with lz4-java blocks. Cite the Minecraft Map Format write-up beside the existing region layout references.
Each LZ4-Java data block and the end marker carries its own magic bytes. A separate stream prefix duplicated the first header and made compression id 4 unreadable by Minecraft.
… listing Symptom: CI failed RoundTripsItsOwnOutput_Mca and every Convert_*__to__Mca pair with InvalidDataException on names like PROBE.TXT. Review also found that compression id 4 chunks written by Minecraft would be rejected with a checksum mismatch, and that List() threw on any damaged region. Root cause: - An input that is not chunk_X_Z.nbt is a caller contract violation; the writer reported it as corrupt data, which neither the round-trip fixture nor the conversion matrix treats as a clean refusal. - lz4-java stores the block checksum through StreamingXXHash32.asChecksum(), whose getValue() keeps only the low 28 bits. The block stream wrote and expected the full 32-bit XXH32. - The reader threw on the first bad location entry and also refused an unpadded final sector that Minecraft itself reads. Fix: - Name, duplicate-coordinate and 255-sector refusals throw ArgumentException. - Block checksum is XXH32(seed 0x9747B28C) & 0x0FFFFFFF on write and read. - McaReader gains a lenient mode with a Problems list; List() and OpenEntry use it and surface a metadata.ini (parse_status=partial), Extract stays strict so an integrity test still reports damage. Stored chunks report their size without decoding; a chunk that fails to decode lists as -1. - A chunk only has to fit inside the file and its own allocation, not have its whole allocation present. Sourcing: rung 1. lz4-java (Apache-2.0) was read for the framing and used as the oracle: its LZ4BlockOutputStream produced the committed reference streams (raw block, two compressed 64 KiB blocks), our raw-block output is byte-identical to it, and its LZ4BlockInputStream decodes our Fast/HC/Max output for 1, 13, 65536 and 200000 bytes. anvil-parser2 (MIT) reads the zlib regions we write (chunks at 0,0 / 31,31 / 5,7) and we read the regions its writer produces.
Hawkynt
force-pushed
the
feat/mca-region-write-support
branch
from
September 30, 2026 18:08
a89ab47 to
e614db3
Compare
Hawkynt
added a commit
that referenced
this pull request
Sep 30, 2026
…and #412 were rebased Symptom: after #399 and #412 were merged, the package README still showed the old OneNote row (no note) and the old MCA row (raw LZ4, no mention of LZ4-Java block streams or oversized .mcc chunks). Root cause: while rebasing those branches onto main, the conflict in the generated API-count line of Hawkynt.FileFormats.Archives/README.md was resolved by taking main's whole file, which also discarded each branch's edit to its own support-matrix row. Fix: restore both rows exactly as the original PR commits wrote them.
Hawkynt
added a commit
that referenced
this pull request
Sep 30, 2026
…and #412 were rebased (#433) Symptom: after #399 and #412 were merged, the package README still showed the old OneNote row (no note) and the old MCA row (raw LZ4, no mention of LZ4-Java block streams or oversized .mcc chunks). Root cause: while rebasing those branches onto main, the conflict in the generated API-count line of Hawkynt.FileFormats.Archives/README.md was resolved by taking main's whole file, which also discarded each branch's edit to its own support-matrix row. Fix: restore both rows exactly as the original PR commits wrote them.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
chunk_X_Z.nbtinputs with gzip, zlib, stored and LZ4-Java block-stream (compression id 4) payloads; per-chunk timestamps viaTimestamp/Timestamp.X.Z.StreamingXXHash32.asChecksum()stores them.chunk_X_Z.nbt(X, Z in 0..31), duplicates and chunks beyond 255 sectors are refused with ArgumentException; external.mccchunks are listed but not written or extracted.metadata.ini(parse_status=partial); extraction stays strict.Verification
LZ4BlockOutputStream(raw block, two compressed 64 KiB blocks) are committed and decoded; our raw-block output is byte-identical to lz4-java's; itsLZ4BlockInputStreamdecodes our Fast/HC/Max output for 1, 13, 65536 and 200000 bytes.