Python-Blosc2 4.14.1 is a security and feature release introducing safe
persisted-object deserialization by default, Caterva2 remote stores and tables,
compressed slicing for RemoteArray, and performance optimizations for
expression gather operations.
- Safe persisted-object deserialization by default. Persisted-input APIs
(
blosc2.open(),blosc2.load(),blosc2.from_cframe(), CFrame constructors, store traversal, and CTable variable-length/batch columns) now default todeserialize="safe". Safe mode accepts passive MessagePack values but rejects embedded objects, references, remote/source descriptors, proxies, and lazy recipes before reconstruction. Trusted callers wishing to deserialize rich values must passdeserialize="full"explicitly; blocked values raise the new publicblosc2.UnsafeDeserializationError. - Caterva2 remote stores and tables. Caterva2 groups and dataset tables can
now be opened via
blosc2.open()asRemoteStoreandRemoteCTablerespectively. Caterva2 tables support on-demand reading, bounded row slices, column projection (select()), iteration, and materialization. Large row slices and materialization fetch data in bounded batches to limit peak memory usage, withmaterialize(urlpath=...)streaming batches directly to disk. Authentication tokens are bound at open time and disk caches are partitioned per token for multi-tenant isolation. - Compressed slicing for ``RemoteArray``. Added
RemoteArray.slice()to extract a slice selection directly as an independent compressedNDArraywithout materializing an uncompressed NumPy array intermediate. Custom compression parameters (cparams) can be passed directly, and fetching, construction, and cache eviction are serialized into a single atomic operation allowing slices larger than the cache quota. - Expression gather performance optimization. Avoided zero-initializing
full NumPy input blocks in
miniexprgather operations, as full blocks are overwritten byte-for-byte by raw NumPy gather (#728). Only partial edge blocks retain zero-padding. Contributed by Johnny Kao (@Johnny-Kao).
Install it with:
pip install blosc2 --upgrade # if you prefer wheels conda install -c conda-forge python-blosc2 mkl # if you prefer conda and MKL
For more info, see the release notes at:
https://github.com/Blosc/python-blosc2/releases
Python-Blosc2 is a high-performance compressor, compute engine, and format for binary data containers that are portable and open-source. It comes with a lazy expression engine allowing for complex calculations on compressed data, whether stored in memory, on disk, or over the network (e.g., via Caterva2). It is especially optimized for storing and retrieving data from N-dimensional arrays (NDArray) and columnar tables (CTable), bringing a query/indexing layer too. The main use case is fast, compressed, out-of-core numerical data — especially when data is too large to fit comfortably in RAM.
More info: https://www.blosc.org/python-blosc2/getting_started/overview.html
The sources and documentation are managed through GitHub services at:
https://github.com/Blosc/python-blosc2
Python-Blosc2 is distributed using the BSD license, see https://github.com/Blosc/python-blosc2/blob/main/LICENSE.txt for details.
Follow https://fosstodon.org/@Blosc2 to get informed about the latest developments.
Enjoy!
- Blosc Development Team Compress Better, Compute Bigger