Generated by make_fixtures.py with pyarrow 22.0.0. Payload: payload.csv, 1567 bytes, sha256 94a81135a89562afe3272534f928fe6383a11382801d8a506413c5463db905c2.
One compressed copy of payload.csv per codec, real compressor output. Feed each to Codec.Decompress WITHOUT an uncompressedSize argument: every one of these streams lets the size be derived, so the call doubles as a test of the per-codec size derivation. The result must equal payload.csv byte for byte.
payload.csv.snappy: 487 bytes, sha256 f1b2acb1e0ced82f27d31e7dd4e036fb86725a201756e2fc645c3652f9de2afb,Codec.Decompress(blob, Compression.Snappy)payload.csv.gz: 283 bytes, sha256 02d6d386d0e8541959e14f49d66ff12cba7614df5cf87d7c05ea44a3a6f0609f,Codec.Decompress(blob, Compression.GZip)payload.csv.br: 199 bytes, sha256 1c930b5cd12d2dfc20f2aea891989740682fee0b8b9f1a4f7df117a27496b940,Codec.Decompress(blob, Compression.Brotli)payload.csv.zst: 229 bytes, sha256 0ccd1748899fd95a60ddfc5254d63fbf8f899b4aef2cdc3e9721b277a5bacd81,Codec.Decompress(blob, Compression.Zstandard)payload.csv.lz4raw: 500 bytes, sha256 e2dcf759453cf8a8205bc35ab6c1a32ff5f7f30867fa81e7bd6e637e14eaf393,Codec.Decompress(blob, Compression.LZ4)payload.csv.lz4hadoop: 508 bytes, sha256 3420f3fa2c127620e85ba6466651a9873715137590a3e29dbf41447b6a0cf0a7,Codec.Decompress(blob, Compression.LZ4) // hadoop framing auto-detected
Control: does Parquet.Document read a real-world file with this codec at all.
A_none.parquet: codec written as UNCOMPRESSED, 1931 bytes, sha256 129c2326dbef130766b2df59624c459e8365feaaae56eec600db5f75a45832baA_snappy.parquet: codec written as SNAPPY, 853 bytes, sha256 8b926795f017e7b362cbaac111f59cb162c3c324244dd102add34fae12b608ceA_gzip.parquet: codec written as GZIP, 649 bytes, sha256 af0bc42119d0a80d46518ef7e332dc7917110c9ae7c9a7144d4d773b40aeb912A_brotli.parquet: codec written as BROTLI, 562 bytes, sha256 f4db3a49eebc9f17a60fd56c9872a6c439ebc53d5863374c68738dfe87ed3c6eA_zstd.parquet: codec written as ZSTD, 596 bytes, sha256 ef18594e16f1f1fae585627b6c77bc1261a183715876ed315ed68ed0a3bbff65A_lz4.parquet: codec written as LZ4, 864 bytes, sha256 a174afc400f48ed81592c890d9175f09ed4aedfd8e6f599e5cefbb2b87f10b44
The compressed page IS the codec stream, byte for byte. This is the oracle test: if Parquet.Document returns the payload, the engine decompressed an arbitrary external blob.
B_uncompressed.parquet: codec id 0, blob 1567 bytes, file 1712 bytes, sha256 c187f5abc9f49898756b392a9f0d39921fee085369fb561bb629776a1cc519d2B_snappy.parquet: codec id 1, blob 487 bytes, file 632 bytes, sha256 5d106de43a12ec2a8211f482ffa2ddef6a6032590161fd019a2e34210ff46ebdB_gzip.parquet: codec id 2, blob 283 bytes, file 428 bytes, sha256 e208f4b0ddc8a7d1f461512fa66efba800ce7d04ec0f2f52f643abb4025c5041B_brotli.parquet: codec id 4, blob 199 bytes, file 344 bytes, sha256 0ff024fb4f0827f2c3ecbea1bf871d81d814184eb2af7301ada0dbf789b26962B_zstd.parquet: codec id 6, blob 229 bytes, file 374 bytes, sha256 d6dfcc60cf7dc1f03fb2bd68ff1ba6148c5e00101db725eb4bcbfd4a7abc3194B_lz4raw.parquet: codec id 7, blob 500 bytes, file 645 bytes, sha256 6d768785b881d5078a7a7a69545ee3df72df2b56e43fdb743e9155c13fd7a7c2B_lz4hadoop.parquet: codec id 5, blob 508 bytes, file 653 bytes, sha256 79b372dc9ac81d6589d65bf7241f32fd2223229842a4bd03bdec9fa4d8652a9c
Mirror check passed: the hardcoded-byte construction used by Codec.Decompress.pq reproduces every B file byte-for-byte.
- A fails, B fails: codec genuinely absent from the engine.
- A works, B fails: codec present but the wrapper is malformed for Microsoft's reader (fix the wrapper; the codec claim still stands).
- A works, B works: Parquet.Document is a decompression oracle for this codec; any format compressing its blocks with it becomes readable in pure M.
- B_uncompressed and B_gzip are wrapper sanity controls: gzip is known-implemented, so if B_gzip fails the wrapper is wrong, not the codec.