Skip to content

Latest commit

 

History

History
193 lines (157 loc) · 6.48 KB

File metadata and controls

193 lines (157 loc) · 6.48 KB

In-Memory DataFrame API Guide

[TOC]

The In-Memory API is the fastest way to experiment with DPSynth. Built on top of Pandas and NumPy, this interface is designed for researchers, rapid prototypers, and software engineers operating on datasets that comfortably fit within a single machine's RAM.


Python API: dpsynth.TabularConfig

The primary entry point for in-memory synthesis is dpsynth.TabularConfig. It accepts a dictionary of attribute domains and mechanism options, is calibrated with a privacy budget to produce a dpsynth.TabularMechanism, and generates a fully synthetic, differentially private DataFrame matching the exact schema and data types of your input.

Usage

import dpsynth
from dpsynth import discrete_mechanisms
import numpy as np
import pandas as pd

config = dpsynth.TabularConfig(
    domains=domains,
    discrete_mechanism=discrete_mechanisms.MSTConfig(),
)
mechanism = config.calibrate(epsilon=1.0, delta=1e-6)
result = mechanism(np.random.default_rng(), sensitive_df)
synthetic_df = result.synthetic_data

Key Configuration Arguments

When initializing dpsynth.TabularConfig:

  • domains: Mapping of column names to domain specifications (CategoricalAttribute, NumericalAttribute, or OpenSetCategoricalAttribute). Every key must exist in data.columns.
  • discrete_mechanism: Configuration object specifying which DP synthesis mechanism to run (e.g., MSTConfig(), AIMConfig(), IndependentConfig()).
  • numerical_bins: Number of equal-frequency quantile buckets used to discretize continuous numerical columns (default: 32).
  • init_budget_fraction: Fraction of total (epsilon, delta) budget allocated for per-column initialization such as bounds computation and partition selection (default: 0.1).
  • cross_attribute_constraints: Optional sequence of constraints to enforce on generated data.

When calling config.calibrate(...):

  • epsilon, delta: Total differential privacy budget parameters. Returns a runnable TabularMechanism.

Standalone End-to-End Python Example

Here is a complete, self-contained Python script demonstrating how to specify a domain, set up a TabularConfig, calibrate the mechanism with a privacy budget, load sensitive data, synthesize records, and print the first few rows.

import dpsynth
from dpsynth import discrete_mechanisms
from dpsynth import domain
import numpy as np
import pandas as pd

# 1. Domain Specification: Define the schema of the tabular dataset
attribute_domains = {
    "age": domain.NumericalAttribute(lower_bound=18, upper_bound=90),
    "workclass": domain.CategoricalAttribute(
        allowed_values=["Private", "Self-emp", "Gov", "Other"]
    ),
    "education": domain.CategoricalAttribute(
        allowed_values=["HS-grad", "Bachelors", "Masters", "PhD"]
    ),
}

# 2. Setup Config: Configure synthesizer with domain and mechanism choices
config = dpsynth.TabularConfig(
    domains=attribute_domains,
    discrete_mechanism=discrete_mechanisms.MSTConfig(),
    numerical_bins=16,
)

# 3. Calibrate Mechanism: Allocate privacy budget to get runnable mechanism
mechanism = config.calibrate(epsilon=1.0, delta=1e-5)

# 4. Load Data: Create sensitive input DataFrame matching the domain schema
sensitive_df = pd.DataFrame({
    "age": [25, 42, 30, 55, 62, 29, 38, 47, 51, 33],
    "workclass": [
        "Private",
        "Gov",
        "Private",
        "Self-emp",
        "Other",
        "Private",
        "Gov",
        "Private",
        "Self-emp",
        "Private",
    ],
    "education": [
        "Bachelors",
        "Masters",
        "HS-grad",
        "PhD",
        "HS-grad",
        "Bachelors",
        "HS-grad",
        "Masters",
        "Bachelors",
        "HS-grad",
    ],
})

# 5. Synthesize Data: Run the calibrated mechanism on the sensitive data
rng = np.random.default_rng(seed=42)
result = mechanism(rng, sensitive_df)
synthetic_df = result.synthetic_data

# 6. Print the first few rows of the generated synthetic dataset
print("Generated Synthetic Data:")
print(synthetic_df.head())

Command-Line Interface: bin/main.py

For immediate execution without writing custom Python scripts, use the standalone binary bin/main.py. It provides command-line flags for all standard configuration parameters.

CLI Execution Syntax

python3 bin/main.py \
  --dataset=/path/to/dataset.csv \
  --domain=/path/to/domain.yaml \
  --epsilon=1.0 \
  --delta=1e-8 \
  --mechanism=mst \
  --seed=12345 \
  --output_path=/tmp/synthetic_output.csv

Supported CLI Flags

  • --dataset: Path to the input CSV file. (Supports standard CSV parsing arguments via --read_csv_args).
  • --domain: Path to the YAML domain specification file.
  • --epsilon, --delta: Total DP privacy budget.
  • --mechanism: Supported options are mst, aim, and independent.
  • --seed: Integer seed for reproducible randomness across DP sampling and PGM inference.
  • --output_path: Destination filepath where the synthetic CSV will be written.

Under the Hood: The In-Memory Lifecycle

When you configure and run TabularConfig, the library performs the following single-machine pipeline:

  1. Discretization: Continuous numerical columns are bucketed into numerical_bins quantiles using pipeline_dp.LocalBackend. Open-set strings are evaluated via DP partition selection.
  2. Integer Encoding: All columns are mapped to dense integer indices [0, K-1].
  3. Domain Compression: DPSynth measures 1-way marginals with Gaussian noise and merges rare categories into an "Other" bucket, producing an un-noised discrete dataset (mbi.Dataset).
  4. Mechanism Execution: Calls the configured discrete mechanism (AIM, MST, etc.) on the discrete dataset. The mechanism fits a Markov Random Field (mbi.MarkovRandomField) via Private-PGM mirror descent.
  5. Sampling & Inversion: Samples synthetic integer records from the graphical model, unpacks "Other" categories, and inverts the integer encoding back to original Pandas dtypes (strings, integers, floating points).