Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
20 commits
Select commit Hold shift + click to select a range
44ebca5
Update to use latest C-Blosc2 sources
FrancescAlted Jul 15, 2026
fe00423
Make paragraphs one-line long
FrancescAlted Jul 15, 2026
ee82741
Add Arrow PyCapsule interchange protocol to CTable (Gap A)
FrancescAlted Jul 15, 2026
02eaa33
Make CTable views read-only for value writes (Gap B)
FrancescAlted Jul 15, 2026
bf39bad
Add Column.fillna() / CTable.dropna() and fix timestamp null detectio…
FrancescAlted Jul 15, 2026
a18bbc8
Propagate nulls through Column arithmetic and comparisons (Gap C2b)
FrancescAlted Jul 15, 2026
3e0cbff
Add UDF aggregations, engine dispatch plumbing, and CTable.apply() (G…
FrancescAlted Jul 15, 2026
560d149
Add enhancing-ctable plan with implementation notes for Gaps A-D
FrancescAlted Jul 15, 2026
76b7616
Guard pandas import in test_groupby.py (fixes CI collection failure)
FrancescAlted Jul 15, 2026
580aef6
Address Copilot review comments on UDF aggregations (Gap D2)
FrancescAlted Jul 15, 2026
376f8e5
Treat a UDF aggregation returning None as a null result (Gap D2)
FrancescAlted Jul 16, 2026
0571ea8
Polish review fixes for Gaps A-D: apply() return type/validation, fro…
FrancescAlted Jul 16, 2026
fddd251
Extend test coverage for Gaps A-D review
FrancescAlted Jul 16, 2026
05a5a67
Fix NameError when nan/inf scalars appear in lazy expressions
FrancescAlted Jul 16, 2026
5fbbe62
Speed up string-key groupby with exact hash-based factorization (Gap …
FrancescAlted Jul 16, 2026
f39fd40
Make reductions on derived nullable expressions skip nulls (closes C2…
FrancescAlted Jul 16, 2026
8b142ce
Remove plan-task references (Gap A-E) from docstrings and comments
FrancescAlted Jul 16, 2026
c972b7b
Phase 2 plan added
FrancescAlted Jul 16, 2026
cc0d84c
Changed default from self.col_names (all columns including computed) …
FrancescAlted Jul 16, 2026
ce36aa0
ShapeInferencer no longer ignores user shapes for nan/inf
FrancescAlted Jul 16, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ project(python-blosc2)

# blosc2_ext.pyx calls blosc2_schunk_lock()/unlock(), added in c-blosc2 3.2.x
set(BLOSC2_MIN_VERSION 3.2.1)
set(BLOSC2_BUNDLED_VERSION v3.2.2)
set(BLOSC2_BUNDLED_VERSION v3.2.3)

if(WIN32 AND NOT CMAKE_C_COMPILER_ID STREQUAL "Clang")
message(FATAL_ERROR "Windows builds require clang-cl. Set CC/CXX to clang-cl or configure CMake with -T ClangCL.")
Expand Down
71 changes: 71 additions & 0 deletions bench/ctable/bench_groupby_keys.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,71 @@
#######################################################################
# Copyright (c) 2019-present, Blosc Development Team <blosc@blosc.org>
# All rights reserved.
#
# SPDX-License-Identifier: BSD-3-Clause
#######################################################################

"""Group-by aggregation speed by key type (1e7 rows, low-cardinality keys).
String keys are the case to watch: they miss every
Cython fast path and go through hash-based factorization in
``_factorize_fixed_width_str``."""

import time
from dataclasses import dataclass

import numpy as np

import blosc2
from blosc2 import CTable

N = 10_000_000
rng = np.random.default_rng(42)

int_keys = rng.integers(0, 20, N)
float_vals = rng.random(N) * 100
cities = np.array(["Paris", "Rome", "Berlin", "Madrid", "Lisbon"])
str_keys = cities[rng.integers(0, 5, N)]


@dataclass
class Row:
ikey: int = blosc2.field(blosc2.int64())
skey: str = blosc2.field(blosc2.string(max_length=8))
dkey: str = blosc2.field(blosc2.dictionary())
val: float = blosc2.field(blosc2.float64())


print(f"building table ({N:.0e} rows)...", flush=True)
t = CTable(Row)
t.extend(
{"ikey": int_keys, "skey": str_keys, "dkey": [str(s) for s in str_keys], "val": float_vals},
validate=False,
)


def bench(label, fn, reps=3):
times = []
for _ in range(reps):
t0 = time.perf_counter()
fn()
times.append(time.perf_counter() - t0)
print(f"{label:45s} {min(times) * 1000:8.1f} ms", flush=True)


bench("int key, sum", lambda: t.group_by("ikey").sum("val"))
bench("int key, mean", lambda: t.group_by("ikey").agg({"val": "mean"}))
bench("string key, sum", lambda: t.group_by("skey").sum("val"))
bench("dict key, sum", lambda: t.group_by("dkey").sum("val"))
bench("two keys (int+dict), sum", lambda: t.group_by(["ikey", "dkey"]).sum("val"))

try:
import pandas as pd
except ImportError:
pass
else:
df = pd.DataFrame({"ikey": int_keys, "skey": str_keys, "val": float_vals})
bench("pandas int key, sum", lambda: df.groupby("ikey")["val"].sum())
bench("pandas string key, sum", lambda: df.groupby("skey")["val"].sum())

# rough speed-of-light reference for the int-key case
bench("numpy bincount int key, sum", lambda: np.bincount(int_keys, weights=float_vals))
3 changes: 1 addition & 2 deletions doc/guides/optimization_tips.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,6 @@
# Optimization tips

This page collects small idioms that make a measurable difference in speed or memory (often both). Each one is backed by a small benchmark in [`bench/optim_tips/`](https://github.com/Blosc/python-blosc2/tree/main/bench/optim_tips), which you can run yourself — see the
[bench/optim_tips README](https://github.com/Blosc/python-blosc2/tree/main/bench/optim_tips/README.md).
This page collects small idioms that make a measurable difference in speed or memory (often both). Each one is backed by a small benchmark in [`bench/optim_tips/`](https://github.com/Blosc/python-blosc2/tree/main/bench/optim_tips), which you can run yourself — see the [bench/optim_tips README](https://github.com/Blosc/python-blosc2/tree/main/bench/optim_tips/README.md).

Numbers below were measured on an Apple M4 Pro Mac Mini (macOS, Python 3.14); absolute values will differ on your machine, but the direction and rough magnitude of each effect should not.

Expand Down
100 changes: 100 additions & 0 deletions doc/reference/ctable.rst
Original file line number Diff line number Diff line change
Expand Up @@ -164,6 +164,61 @@ Sentinels are resolved in this order: explicit ``null_value`` in the schema,
.. autofunction:: get_null_policy


Nulls in expressions
---------------------

Arithmetic and comparisons built from :class:`Column` objects (``t.x + 1``,
``t.x > 0``, ...) propagate nulls on nullable int/timestamp/bool columns
without touching storage or the on-disk format — a sentinel is a validity
bitmap encoded in-band, so propagation is purely a rewrite of the lazy
expression:

- **Arithmetic** (``+ - * / // % **``): the result is null (NaN) wherever
any nullable operand is null, promoting integer/timestamp results to
``float64`` — the same promotion pandas' legacy (pre-nullable-dtype)
integer-null arithmetic performs. Non-nullable columns are unaffected and
pay no overhead::

(t.price + 1) # NaN where price is null; float64 either way
(t.price + t.tax) # NaN where either operand is null

- **Comparisons** (``< <= > >= == !=``): SQL ``WHERE`` semantics — a null
never satisfies any comparison, including ``==``/``!=`` against the raw
sentinel value itself. Only :meth:`Column.is_null` /
:meth:`Column.notnull` test for nullness::

t[t.price < 0] # excludes rows where price is null
t[t.price == -1] # False for a null row even if -1 is the sentinel
t[t.price.is_null()] # the only way to select null rows

- **Boolean combinators** (``& | ~``) need no special handling: their
inputs are already-resolved comparison results with null-ness folded to
``False``. Mind the consequence for negation: since a null compares
``False``, ``~(t.price > 0)`` *selects* null rows (they are "not > 0").
To exclude them, write the complementary comparison instead
(``t.price <= 0``), which never matches nulls.

Kleene three-valued logic (where ``null > 0`` evaluates to null rather than
``False``) is intentionally out of scope — it needs a validity channel on
boolean intermediates, i.e. masks, which CTable does not use.

Reductions on derived expressions skip nulls too: arithmetic involving a
nullable column returns a ``NullableExpr`` — a thin wrapper that remembers
which rows are null — so ``sum``/``mean``/``min``/``max``/``std`` on it skip
nulls (and deleted rows) exactly like the corresponding :meth:`Column.sum`
etc. do on a real column::

t.price.sum() # skips nulls (real Column)
(t.price + 1).sum() # skips nulls too (NullableExpr)
((t.price + 1) * 2).mean() # chaining keeps the null channel

Only *nulls* are skipped: a NaN produced by the arithmetic itself (e.g.
``0/0`` on non-null values) is a value, not a null, and poisons the
reduction as usual. ``min()``/``max()`` raise ``ValueError`` when every
value is null; ``sum()`` returns ``0.0`` and ``mean()``/``std()`` return
NaN — the same conventions as the Column reductions.


Attributes
----------

Expand Down Expand Up @@ -257,6 +312,7 @@ When a NumPy structured array is needed, materialize explicitly::
.. autosummary::

CTable.where
CTable.dropna
CTable.view
CTable.take
CTable.select
Expand All @@ -268,6 +324,7 @@ When a NumPy structured array is needed, materialize explicitly::
CTable.group_by

.. automethod:: CTable.where
.. automethod:: CTable.dropna
.. automethod:: CTable.view
.. automethod:: CTable.take
.. automethod:: CTable.select
Expand Down Expand Up @@ -321,6 +378,32 @@ several ways; the list-of-pairs form additionally lets you pass ``Column``
objects (``t.sales``) instead of name strings; use explicit names when you want a
specific (and easily addressable) output column name.

**Custom UDF aggregations** are accepted only via the named form --
``output_name=(column, callable)``, optionally with an explicit output dtype as
a third element::

by_city.agg(sales_range=(t.sales, lambda a: a.max() - a.min()))
# explicit dtype instead of inferring it from the callable's results:
by_city.agg(sales_range=(t.sales, lambda a: a.max() - a.min(), blosc2.float32()))

The callable receives a 1-D NumPy array of the group's live, non-null values
(same null semantics as the built-in aggregations) and returns a scalar; it is
called once per group with a plain Python loop -- there is no JIT
acceleration for arbitrary UDFs yet. A group with no non-null values for that
column never calls the callable, producing a null result instead (the same
convention ``sum``/``min``/``max`` already use). The mapping and
list-of-pairs auto-named forms cannot derive an output column name for an
arbitrary callable, so they only accept the string/blosc2-reduction-function
ops described above.

**Execution engine.** :meth:`CTable.group_by` takes an ``engine=`` parameter
for the *built-in* aggregations (``size``/``count``/``sum``/``mean``/``min``/
``max``/``argmin``/``argmax``): ``"auto"`` (default) and ``"numpy"`` both use
the NumPy/Cython chunked implementation; ``"jit"`` is reserved for a future
miniexpr-JIT path and currently raises :class:`NotImplementedError`. UDF
aggregations always run the plain Python per-group loop regardless of
``engine``.

**Output ordering.** ``sort`` controls how groups are ordered, and is *always
by the group key(s), never by the aggregated value*:

Expand Down Expand Up @@ -414,6 +497,7 @@ in place.
CTable.drop_computed_column
CTable.drop_column
CTable.rename_column
CTable.apply

.. automethod:: CTable.delete
.. automethod:: CTable.compact
Expand All @@ -423,6 +507,7 @@ in place.
.. automethod:: CTable.drop_computed_column
.. automethod:: CTable.drop_column
.. automethod:: CTable.rename_column
.. automethod:: CTable.apply


Indexes
Expand Down Expand Up @@ -506,12 +591,24 @@ well as import/export paths for CSV, Arrow, and Parquet data.
CTable.to_cframe
CTable.to_csv
CTable.to_arrow
CTable.__arrow_c_stream__
CTable.to_parquet
CTable.from_arrow
CTable.from_parquet
CTable.from_csv
ctable_from_cframe

CTable also implements the `Arrow PyCapsule interchange protocol
<https://arrow.apache.org/docs/format/CDataInterface/PyCapsuleInterface.html>`_
via :meth:`~blosc2.CTable.__arrow_c_stream__`, so pyarrow, DuckDB, Polars, and
pandas >= 2.2 can consume a CTable directly as a stream of record batches
(``pa.table(ct)``, ``duckdb.sql("SELECT ... FROM ct")``, ``pl.DataFrame(ct)``),
and :meth:`~blosc2.CTable.from_arrow` accepts any object implementing that
protocol on ingest. Strict zero-copy is not possible — the underlying data is
compressed, so decompression is unavoidably a copy — but there is no
intermediate materialization: batches are decompressed and handed to the
consumer one at a time, so memory use stays bounded regardless of table size.

.. automethod:: CTable.load
.. automethod:: CTable.open
.. automethod:: CTable.save
Expand All @@ -520,6 +617,7 @@ well as import/export paths for CSV, Arrow, and Parquet data.
.. automethod:: CTable.to_cframe
.. automethod:: CTable.to_csv
.. automethod:: CTable.to_arrow
.. automethod:: CTable.__arrow_c_stream__
.. automethod:: CTable.to_parquet
.. automethod:: CTable.from_arrow
.. automethod:: CTable.from_parquet
Expand Down Expand Up @@ -653,10 +751,12 @@ Nullable helpers
Column.is_null
Column.notnull
Column.null_count
Column.fillna

.. automethod:: Column.is_null
.. automethod:: Column.notnull
.. automethod:: Column.null_count
.. automethod:: Column.fillna


Unique values
Expand Down
Loading
Loading