Skip to content

Commit 11428b4

Browse files
committed
Getting ready for release 4.12.0
1 parent e776671 commit 11428b4

6 files changed

Lines changed: 49 additions & 117 deletions

File tree

‎ANNOUNCE.rst‎

Lines changed: 22 additions & 40 deletions
Original file line numberDiff line numberDiff line change
@@ -1,45 +1,27 @@
1-
Announcing Python-Blosc2 4.11.0
1+
Announcing Python-Blosc2 4.12.0
22
===============================
33

4-
Nullability in ``CTable`` is rebuilt on Arrow's own model: a nullable column
5-
keeps its nulls in a validity sidecar instead of reserving a value from its own
6-
range. That makes it lossless, and everything above it — predicates, indexes,
7-
Arrow/Parquet/CSV round-trips — follows. Wheels become a single Stable ABI
8-
build per platform.
9-
10-
- **Mask-based nullable columns, and they are the default.** A bare
11-
``nullable=True`` no longer steals a value from the dtype, so an ``int8``
12-
column can hold ``-128``, a ``utf8`` one can hold ``""``, a ``float64`` one
13-
can tell ``NaN`` from missing, and ``complex128`` is nullable at all for the
14-
first time. ``None`` is how you write a null. Note that a table with a
15-
mask column records schema version 3, which readers older than 4.11.0 refuse
16-
to open.
17-
18-
- **Predicates over nulls follow three-valued (Kleene) logic.** A comparison
19-
against a null is now *unknown* rather than ``False``, so ``~(t.price > 10)``
20-
returns the rows definitely not above 10 instead of every null row, and
21-
``~((a > 10) & (b == 999))`` stops dropping rows that qualify. Both the
22-
operator and the string query form agree with SQL. A predicate can also be
23-
asked about its unknown rows: ``p.is_null()``, ``p.null_count()``,
24-
``p.fillna(True)``.
25-
26-
- **Column indexes are null-aware.** Per-segment ``min``/``max`` are taken over
27-
the rows that carry a value, so ``Column.min``/``Column.max`` answer from the
28-
index for a nullable column instead of scanning, and ``where()`` with an ``OR``
29-
over a nullable indexed column no longer falls back to a full scan (**1.6x**).
30-
``rebuild_index()`` promotes indexes written by an earlier release.
31-
32-
- **A single Stable ABI (abi3) wheel per platform**, serving CPython 3.11 and
33-
every later version — so a new CPython is installable from a wheel without
34-
waiting for a blosc2 release. Free-threaded 3.14 and 3.15 ship alongside as
35-
version-specific wheels. No measurable performance cost.
36-
37-
- **Plus a long list of fixes** around nullable columns: CSV import/export,
38-
``extend()`` between storages, sorted-view reductions, descending sorts of
39-
the widest integers, timestamp writes through ``col[key] = value``, scalar
40-
broadcast on mask columns, nested columns surviving ``convert_nulls()`` and a
41-
save/reopen cycle, and more. Outside ``CTable``, ``asarray()`` no longer
42-
corrupts arrays whose chunks overhang the shape.
4+
This release makes remote Blosc2 arrays faster and easier to use, from object
5+
stores and plain HTTP servers to Caterva2.
6+
7+
- **Read and write through fsspec URLs.** ``blosc2.open()``, ``save_array()``
8+
and ``save_tensor()`` support URLs such as ``s3://``, ``gs://``, ``https://``
9+
and chained filesystems via the new ``blosc2[fsspec]`` extra.
10+
11+
- **Fetch only the bytes a slice needs.** Lazy remote proxies read individual
12+
compressed blocks, overlap requests and can reuse a validated on-disk cache.
13+
This cuts traffic, latency and peak memory for small remote slices.
14+
15+
- **Improved Caterva2 access.** Stored arrays and leaves inside ``.b2z``
16+
containers use HTTP byte ranges. Pre-sized remote arrays can also be filled
17+
concurrently with ``C2Array.update_chunk()`` and ``aupdate_chunk()``.
18+
19+
- **Lean UTF-8 index lookups.** FULL-index queries bisect the vocabulary on
20+
disk instead of materializing it, greatly reducing memory use for
21+
high-cardinality string columns.
22+
23+
- **Bundled C-Blosc2 3.3.3**, together with expanded remote-array documentation
24+
and benchmarks.
4325

4426
Install it with::
4527

‎RELEASE_NOTES.md‎

Lines changed: 23 additions & 73 deletions
Original file line numberDiff line numberDiff line change
@@ -1,27 +1,23 @@
11
# Release notes
22

3-
## Changes from 4.11.0 to 4.11.1
3+
## Changes from 4.11.0 to 4.12.0
44

5-
XXX version-specific blurb XXX
5+
This release focuses on efficient remote arrays. Blosc2 containers can now be
6+
read and written through fsspec URLs, while lazy proxies fetch only the required
7+
blocks, overlap requests and reuse validated caches. Caterva2 arrays gain
8+
block-range reads and concurrent chunk writers. UTF-8 FULL-index lookups also
9+
use substantially less memory by bisecting their vocabulary on disk.
610

711
### Improvements
812

913
* New `blosc2[fsspec]` extra: `blosc2.open()`, `save_array()` and `save_tensor()`
1014
accept any [fsspec](https://filesystem-spec.readthedocs.io) URL — `s3://`,
11-
`gs://`, `https://`, `zip://`, chained ones like
12-
`zip://inner.b2nd::s3://bucket/a.zip`.
13-
`open()` reads the container whole, or through a staleness-checked local copy
14-
with `cache_storage=` (which is what covers `.b2d` stores, sparse frames,
15-
`offset` and `mmap_mode`), or a piece at a time with `lazy=True`, which
16-
leaves a huge frame where it is and fetches only what a slice touches
17-
through the new `blosc2.FsspecNDSource`. The two combine: `lazy=True` with a
18-
`cache_storage=` keeps what was fetched there, so a later run starts from
19-
it. Protocol drivers (`s3fs`, `gcsfs`...) and credentials stay the caller's
20-
business. On the write side `NDArray.save()` and `blosc2.save()` upload the
21-
whole array as one object; containers cannot be *backed* by a URL while they
22-
are written (the C layer rewrites a frame's header and offsets as chunks land,
23-
which an object store has no way to serve), so constructors given a URL now say
24-
that instead of failing deep in C.
15+
`gs://`, `https://`, `zip://` and chained URLs. `open()` can download the
16+
container, keep a validated local copy with `cache_storage=`, or fetch only
17+
requested slices with `lazy=True` and the new `blosc2.FsspecNDSource`; lazy
18+
reads can also use a persistent cache. Saves upload one complete object, but
19+
URL-backed mutable containers are not supported. Protocol drivers and
20+
credentials remain the caller's responsibility.
2521

2622
* A `C2Array` can be written to a chunk at a time, which is how several
2723
processes fill one remote array at once: `update_chunk()` (and its async
@@ -70,49 +66,13 @@ XXX version-specific blurb XXX
7066
with the traffic, since a fetch in flight is now a block rather than a chunk.
7167
`bench/ndarray/fsspec-block-granularity.py` measures both on any array.
7268

73-
* A `Proxy` over a `C2Array` reads **blocks** too, straight out of the stored
74-
frame over HTTP byte ranges. Caterva2 serves a stored dataset from a file, so
75-
the `Range` header is honoured and composes with the auth cookie; no new
76-
endpoint is involved. On cat2.cloud's `kevlar-tomo.b2nd` a corner slice costs
77-
0.031 MB instead of 2.723 MB, and a slice touching ten chunks takes 0.14 s
78-
against 1.01 s. Three things add up to it: one pooled HTTP client instead of a
79-
connection per request (0.162 s → 0.046 s each), fetches that overlap by
80-
default as `afetch()` already did, and one request carrying the whole wave of
81-
ranges (`multipart/byteranges`, which no object store offers). A dataset the
82-
subscriber *computes* — a lazy expression, an HDF5 leaf, a `.b2z` member — is
83-
fetched a whole chunk at a time as before; which it is costs at most one small
84-
request to find out, and is not asked again -- unless the subscriber could not
85-
say, which a busy or unreachable one cannot: a 5xx or a connection that failed
86-
is asked again on the next fetch rather than written off, since neither
87-
downloaded anything to find out. The same holds once blocks are being read: a
88-
dataset that stops being served from a file, or a subscriber too busy to serve
89-
it, costs the granularity and not the fetch — whatever is still missing comes
90-
as whole chunks, which every dataset can be read as. `blosc2.ByteRangeNDSource` is
91-
the frame reader `FsspecNDSource` and the new `C2NDSource` share: subclass it
92-
with a `read_range(offset, size)` to give any transport the same treatment.
93-
Opening a remote frame through either of them now costs two requests instead
94-
of four (0.237 s → 0.138 s against cat2.cloud), and one for a frame small
95-
enough to arrive whole in the first read: the two reads that only measured the
96-
next one are guessed at generously instead, since over a network a few hundred
97-
bytes and a few kilobytes cost the same. Of those two, only the header is read
98-
when the frame is opened — it is what says the frame can be read this way at
99-
all — and where the chunks are waits for the first chunk anything asks about.
100-
So `blosc2.open(url, lazy=True)` reads once, and so does a whole run over a
101-
`cache_storage=` that already holds the slice wanted: it fetches nothing, and
102-
now asks for no index either. `FsspecNDSource` also asks the filesystem for
103-
the object's metadata, which is where its `stamp` comes from, so its floor is
104-
that call plus the header read; `C2Array` gets geometry and stamp together
105-
from `api/info`, and its floor is that one request. A persisted cache keeps
106-
what the source read about *where* things are — the frame's chunk offsets, and
107-
the block offsets of the chunks it holds only part of — so a later run over it
108-
starts from those instead of reading them again. A warm fetch of blocks missing
109-
from a chunk already half held goes from 4 requests to 2 against a subscriber,
110-
and drops the offsets read and one layout read per chunk touched against an
111-
object store. Only for a source that can name the bytes it read: positions in a
112-
frame are worth nothing against a frame that was replaced, so an unstamped
113-
source keeps none of this and reads as before.
114-
`bench/ndarray/cat2-block-granularity.py` measures all of it on any dataset,
115-
against a real subscriber or a stand-in it starts itself.
69+
* A `Proxy` over a `C2Array` now reads only the required compressed blocks from
70+
file-backed datasets using concurrent, batched HTTP byte ranges. On
71+
cat2.cloud's `kevlar-tomo.b2nd`, this reduced a corner slice from 2.723 MB to
72+
0.031 MB and a ten-chunk slice from 1.01 s to 0.14 s. Computed datasets fall
73+
back to whole-chunk reads. `blosc2.ByteRangeNDSource` provides the same frame
74+
reader to `FsspecNDSource`, `C2NDSource` and custom transports, with lazy,
75+
persistent caching of frame layout metadata.
11676

11777
* `DictStore.member_window(key)` says where a leaf's frame lies inside a `.b2z`,
11878
as `(offset, nbytes)`. A zip store keeps each external leaf uncompressed, so
@@ -126,20 +86,10 @@ XXX version-specific blurb XXX
12686
output it is, and said plainly when it is the store's own super-chunk.
12787

12888
* `blosc2.Proxy(src, urlpath=..., mode="a")` now adopts the cache left by an
129-
earlier run instead of failing on the existing file, so a proxy's cache can
130-
outlive the process. The cache must come from a proxy over a source of the same
131-
shape and dtype; anything else at that path raises. A cache that holds only
132-
some blocks of a chunk keeps them across runs too. Sources that can name the
133-
bytes they read are held to that as well, so a remote array *replaced* while
134-
keeping its shape is noticed rather than served stale: `FsspecNDSource` uses
135-
fsspec's token and `C2Array` the subscriber's mtime, both free with metadata
136-
they already fetch. A proxy that `blosc2.open` rebuilds over its own cache
137-
gets the same treatment without the raise -- there is no `mode="w"` to offer
138-
it -- so a cache whose stamp no longer matches starts as though nothing had
139-
been fetched and fills again from the bytes served now; opened read-only there
140-
is nothing to empty, and every read falls through to the source. Caches from
141-
earlier 4.11.1 development builds are not adopted (the stamp moved to a
142-
`proxy-stamp` entry); pass `mode="w"` once.
89+
earlier run, including partially fetched chunks. It validates the cache's
90+
shape, dtype and source stamp, refetching stale data if the remote array was
91+
replaced and raising for incompatible caches. Caches from pre-release 4.12.0
92+
builds require one fresh open with `mode="w"`.
14393

14494
* The source protocol moved to its own module, `blosc2.proxy_source`:
14595
`ProxySource`, `ProxyNDSource`, `ByteRangeNDSource`, `FsspecNDSource` and the

‎doc/guides/optimization_tips.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -402,7 +402,7 @@ t.where("title == 'some exact title'") # looks it up, no scan
402402

403403
A `utf8()` column is indexed by *alphabetical rank*: the query literal is located by bisecting the index's vocabulary, and the rows that match are a contiguous run of the sorted-positions sidecar. None of that depends on how many different values the column holds, so the index is worth having at either cardinality — a scan costs tens of milliseconds, a lookup a few. The first lookup of a session is the dearer one only because it opens the sidecars; later ones reuse them.
404404

405-
```{versionchanged} 4.11.1
405+
```{versionchanged} 4.12.0
406406
The literal is bisected out of the vocabulary sidecar. Earlier versions materialized the whole vocabulary on the first lookup, which on the near-unique column here cost ~62 ms and 739 MiB of peak memory instead of the new ~12 ms and 5.5 MiB.
407407
```
408408

‎plans/cat2-concurrent-writers.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -23,7 +23,7 @@ one (see [Non-goals](#non-goals)).
2323

2424
## What was verified
2525

26-
Measured on this machine (Apple M4 Pro, APFS, python-blosc2 4.11.1.dev0 against
26+
Measured on this machine (Apple M4 Pro, APFS, python-blosc2 4.12.0.dev0 against
2727
c-blosc2 3.3.2), on local files. Code references: c-blosc2 at
2828
`/Users/faltet/blosc/c-blosc2` (`main`), Caterva2 at
2929
`/Users/faltet/ironArray/Caterva2` (`range-honesty`).

‎pyproject.toml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -41,7 +41,7 @@ dependencies = [
4141
"rich",
4242
"threadpoolctl; platform_machine != 'wasm32'",
4343
]
44-
version = "4.11.1.dev0"
44+
version = "4.12.0"
4545
[project.entry-points."array_api"]
4646
blosc2 = "blosc2"
4747

‎src/blosc2/version.py‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,2 +1,2 @@
1-
__version__ = "4.11.1.dev0"
1+
__version__ = "4.12.0"
22
__array_api_version__ = "2024.12"

0 commit comments

Comments
 (0)