Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -104,6 +104,9 @@ contributor expanding the library, explore the documentation below:

* **[Documentation Sitemap](sitemap.md)**: Complete table of contents and
layout of the DPSynth documentation.
* **[Specifying Attribute Domains](specifying_domains.md)**: Upfront domain
requirements, supported attribute types, the DPSynth domain language, and
inferring domains from Pydantic, Protobuf, and Pandas schemas.
* **[Data Model & Terminology](data_and_terminology.md)**: Attributes
(categorical vs. numerical), schema deduction, and `domain.yaml`
specifications.
Expand Down Expand Up @@ -151,6 +154,7 @@ APIs:
:caption: Getting Started
:hidden:

specifying_domains
data_and_terminology
in_memory_api
scalable_pipeline_api
Expand Down
17 changes: 17 additions & 0 deletions docs/sitemap.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,23 @@

--------------------------------------------------------------------------------

<details>
<summary>📁 <a href="specifying_domains.md">Specifying Attribute Domains</a></summary>

* [Why Domains Must Be Specified Up Front](specifying_domains.md#why-domains-must-be-specified-up-front)
* [Supported Attribute Types](specifying_domains.md#supported-attribute-types)
* [Out-of-Domain Handling](specifying_domains.md#out-of-domain-handling)
* [Specifying Domains in the DPSynth Language](specifying_domains.md#specifying-domains-in-the-dpsynth-language)
* [Deriving Domains from Structured Containers](specifying_domains.md#deriving-domains-from-structured-containers)
* [1. From Pydantic Models](specifying_domains.md#1-from-pydantic-models)
* [2. From Protocol Buffers](specifying_domains.md#2-from-protocol-buffers)
* [3. From Pandas Dtypes](specifying_domains.md#3-from-pandas-dtypes)
* [Customizing Inferred Domains](specifying_domains.md#customizing-inferred-domains)

</details>

--------------------------------------------------------------------------------

<details>
<summary>📁 <a href="data_and_terminology.md">Data Model and Terminology</a></summary>

Expand Down
225 changes: 225 additions & 0 deletions docs/specifying_domains.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,225 @@
<!-- Copyright 2026 Google LLC

Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at

http://www.apache.org/licenses/LICENSE-2.0

Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License. -->

# Specifying Attribute Domains

<!-- disableFinding(LINK_RELATIVE_G3DOC) -->

Every synthesis mechanism in DPSynth requires an explicit **domain
specification** describing the attributes (columns) of your dataset and the set
of valid values each attribute can take. This guide explains why domains must be
defined before touching private data, how to express domains directly in the
DPSynth domain language, and how to derive them automatically from structured
schema definitions like Pydantic models, Protocol Buffers, and Pandas dtypes.

## Why Domains Must Be Specified Up Front

In differential privacy, the domain of an attribute defines the sample space of
noisy measurements and bounds the sensitivity of each record's contribution.
Crucially, **the domain must be specified before inspecting the sensitive
dataset**:

* **Data-dependent bounds leak privacy**: Computing minimum and maximum values
or extracting unique categories directly from private data is not
differentially private. Adding or removing a single outlier record can
change the observed range or reveal the presence of a rare category.
* **Calibration precedes execution**: In DPSynth's [three-step mechanism
lifecycle](mechanism_api.md#three-step-pipeline), you pass the domain to the
calibration step to allocate privacy budgets and construct per-column
initializers *before* the calibrated mechanism ever sees private data.
* **Handling unknown domains privately**: When the valid categories of a
string column are not known from public metadata, use an open-set
categorical attribute rather than scanning the raw data. DPSynth will
automatically spend a fraction of the privacy budget on differentially
private partition selection to discover common categories.

## Supported Attribute Types

DPSynth provides four attribute classes in the domain module to model tabular
columns:

| Attribute Type | When to Use | Required Public Metadata | DP Initialization |
| :--- | :--- | :--- | :--- |
| `CategoricalAttribute` | Finite set of known values, such as booleans, enum values, US states, or status codes. | List of valid categories. | None (domain is already discrete and fixed). |
| `NumericalAttribute` | Ordered integer or continuous floating-point numbers, such as age, balance, or duration. | Inclusive minimum and maximum bounds, plus integer or float data type. | Discretizes into bins via DP quantiles (or uses fixed bin edges if provided). |
| `OpenSetCategoricalAttribute` | Categorical strings whose valid values are unknown in advance, such as free-entered city names or job titles. | None (optional fallback sentinel and optional public seed categories). | Discovers frequent categories via DP Gaussian thresholding (requires positive delta). |
| `FreeFormTextAttribute` | Unstructured text fields synthesized by a language model. | Maximum token length. | Handled by text-generation mechanisms. |

### Out-of-Domain Handling

Real-world datasets frequently contain missing values or records that fall
outside expected bounds. Because raising a runtime error on unexpected records
would leak information about the private data, DPSynth handles out-of-domain
values deterministically during encoding:

* **Categorical attributes**: Any value not present in the declared list of
possible values is mapped to a designated fallback index (the first category
by default). If missing or unexpected values may occur, place a typed
sentinel (such as an unknown label for strings or a negative number for
integers) at the first position. If all records are guaranteed to be
in-domain, omit the sentinel so the synthesizer never generates it.
* **Numerical attributes**: By default, values outside the declared minimum
and maximum bounds are clipped to the nearest bound, and missing entries map
to the minimum bound. When range clipping is disabled, out-of-range and
missing entries are routed to a dedicated out-of-domain bin during
discretization and decoded back to a sentinel value.

## Specifying Domains in the DPSynth Language

You can define a domain directly in Python as a dictionary mapping column names
to attribute instances:

```python
import dpsynth
from dpsynth import domain

domains = {
"age": domain.NumericalAttribute(min_value=18, max_value=100, dtype="int"),
"tier": domain.CategoricalAttribute(possible_values=["FREE", "PRO"]),
"city": domain.OpenSetCategoricalAttribute(),
}
```

Domain dictionaries can also be saved to and loaded from YAML:

```python
yaml_str = dpsynth.to_yaml(domains)
restored_domains = dpsynth.from_yaml(yaml_str)
```

## Deriving Domains from Structured Containers

In many applications, your dataset schema is already captured in a structured
container such as a **Pydantic model**, a **Protocol Buffer descriptor**, or
**Pandas column dtypes** (for example, from a Parquet schema). Manually
rewriting tens or hundreds of fields in the DPSynth domain language creates
unnecessary boilerplate and risks schema drift.

DPSynth provides domain inference helpers under the adapters package that
automatically translate these schema definitions into a domain dictionary.
Because the output is a standard Python dictionary, you can inspect or override
individual attributes before calibrating your mechanism.

### 1. From Pydantic Models

Pydantic models can encode both categorical types (booleans, enums, and literal
unions) and numerical bounds directly on the class definition. Optional fields
automatically include a missing-value category or disable numerical range
clipping:

```python
from typing import Literal
import dpsynth.adapters.pydantic
import pydantic


class UserRecord(pydantic.BaseModel):
age: int = pydantic.Field(ge=18, le=100)
tier: Literal["FREE", "PRO"]
city: str


domains = dpsynth.adapters.pydantic.infer_domain(UserRecord)
```

If numeric fields on the Pydantic model do not carry field bound annotations,
you can supply their bounds via the numerical bounds argument. Nested models are
also supported as long as they are non-optional, and are flattened into
dot-separated attribute names.

### 2. From Protocol Buffers

Protocol Buffer descriptors define field types and enum values, but do not store
numeric minimum and maximum ranges:

```protobuf
message UserRecord {
enum Tier {
FREE = 0;
PRO = 1;
}
int32 age = 1;
Tier tier = 2;
string city = 3;
}
```

Pass your Protobuf message class or descriptor along with numerical bounds:

```python
import dpsynth.adapters.protobuf

domains = dpsynth.adapters.protobuf.infer_domain(
UserRecord,
numerical_bounds={"age": (18, 100)},
)
```

Enums and booleans map to categorical attributes, numeric fields map to
numerical attributes using the provided bounds, and strings map to open-set
categorical attributes. Nested submessages are also supported as long as they
are non-repeated (flattened into dot-separated keys), and unsupported or
unbounded fields can be skipped automatically.

### 3. From Pandas Dtypes

When working with typed tabular formats such as Parquet or Arrow, or when a
public schema is already represented as Pandas dtypes, you can infer the domain
from a Series of dtypes or a column-to-dtype mapping:

```python
import dpsynth.adapters.pandas
import pandas as pd

schema_dtypes = {
"age": "int64",
"tier": pd.CategoricalDtype(["FREE", "PRO"]),
"city": "string",
}

domains = dpsynth.adapters.pandas.infer_domain(
schema_dtypes,
numerical_bounds={"age": (18, 100)},
)
```

```{warning}
When inferring a categorical attribute from a Pandas CategoricalDtype, the list
of categories **must** come from a public schema. Never cast a sensitive column
to a categorical dtype on private data and pass the resulting dtypes to domain
inference, as doing so extracts the exact active support of the private dataset
without differential privacy. For string columns with unknown categories, leave
the dtype as a string (or object) so it maps to an open-set categorical
attribute.
```

### Customizing Inferred Domains

Because domain inference returns a plain Python dictionary, you can easily
customize specific columns after inference when a schema container does not
capture a nuance of your domain (for example, replacing an inferred open-set
attribute on a string column with a known closed-set categorical attribute, or
specifying custom bin edges on a numerical column):

```python
domains = dpsynth.adapters.protobuf.infer_domain(
UserRecord,
numerical_bounds={"age": (18, 100)},
)

# Override an open-set string field when valid values are publicly known:
domains["city"] = domain.CategoricalAttribute(
possible_values=["OTHER", "NYC", "LA", "SF"],
)
```
Loading