Domains

Set a domain on a column

Configure a fixed, hierarchical (inline, derived, or table), or reference domain - each has its own endpoint.

A domain restricts which values a column will accept - beyond a data type ("any string") to "one of these values" - and it's enforced on every insert_rows and merge_rows, so bad data never lands. Domains are stored as metadata on the Delta table, so the rules travel with the data.

TypeRestricts a column to…Use for
Fixed (enum)a finite value liststatus codes, country codes, any closed set
Hierarchicalvalues that depend on a parent column's valuecategory → subcategory, country → region
Referencevalues present in another table's columncross-table referential integrity (like a foreign key)

Hierarchical domains come in three flavors, differing in where the parent→child mapping lives:

FlavorMapping livesUse when
Inlinein the domain itself (curated, editable per pair)you own the taxonomy and want to control every pair
Derivedrecomputed from the table's own rowsthe data is the source of truth; only block new combinations
Tablean external lookup table, read at query timethe mapping is maintained elsewhere (org chart, geo reference)

Each type has its own typed endpoint - PUT …/domain/fixed, …/domain/hierarchical/{inline,derived,table}, …/domain/reference - mirrored by CLI subcommands. Over MCP, the single set_column_domain tool covers the same ground with a type parameter. Pick the type that fits, set it, then validate.

Fixed (enum)

A finite list of allowed values:

eddytor set domain eddytor sales orders status fixed --values active,inactive,draft
PUT /v1/tables/eddytor/cfg_xxx/<uuid>_orders/columns/status/domain/fixed
Authorization: Bearer edd_live_…

{ "allowed_values": ["active", "inactive", "draft"] }
set_column_domain(table, "status", "fixed",
  values=["active", "inactive", "draft"])

Grow or shrink the list later without re-setting the domain - manage domain values.

Hierarchical - inline (curated mapping)

A child column whose allowed values depend on a parent column's value, with the mapping stored in the domain itself. Set the parent's domain first, then either seed the mapping from data or build it by hand.

Seed from existing data

If the table already holds the parent/child pairs you want, let Eddytor scan the rows and build the mapping from the distinct parent/child value pairs found:

eddytor set domain eddytor sales orders subcategory hierarchical-inline \
  --seed-from-data --parent-column category
PUT /v1/tables/…/columns/subcategory/domain/hierarchical/inline

{ "parent_column": "category", "seed_from_data": true }
set_column_domain(table, "subcategory", "hierarchical",
  parent_column="category", seed_from_data=true)

Seeding requires the parent column to already carry a fixed or inline hierarchical domain, is mutually exclusive with an explicit mapping, and rejects tables whose distinct-pair count exceeds the seeding cap. The response's skipped_parents lists parent values found in the data but absent from the parent's domain - those pairs are skipped, not invented.

Build the mapping by hand

Key mappings by the parent values' UUIDs (not the strings):

# 1. read the parent domain's value IDs
get_allowed_values(table, "category")
# -> [{ "id": "a1b2c3d4-…", "value": "Electronics" },
#     { "id": "c3d4e5f6-…", "value": "Clothing" }, …]

# 2. key the mapping by those UUIDs
PUT …/columns/subcategory/domain/hierarchical/inline
{ "parent_column": "category",
  "mappings": { "a1b2c3d4-…": ["Phones", "Laptops"],
                "c3d4e5f6-…": ["Shirts", "Pants"] } }

Heads up

mappings keys are parent-value UUIDs, not strings - a string key fails with Invalid hierarchy format. Read them with get_allowed_values(table, parent_column) first. Setting the domain replaces the entire mapping; to change individual pairs afterwards, use the granular parent/child endpoints instead of re-sending the map.

Deeper chains (grandparent → parent → child)

Chains aren't limited to two levels: an inline hierarchical column can itself be the parent of the next level (fixed → inline → inline …). Read the parent level's value IDs - for a hierarchical parent that's its child values across all of its own parents - then map or seed the next level the same way:

# category (fixed) → subcategory (inline) → product_type (inline)
eddytor set domain eddytor sales orders product_type hierarchical-inline \
  --seed-from-data --parent-column subcategory

Heads up

Chain levels resolve by string, so child values must be unique across the whole level. If the same value string appears under two different parents (two UUIDs), the first match wins when the next level resolves against it - rename one of the duplicates before chaining deeper.

Hierarchical - derived (mapping from the data)

No stored mapping: the allowed child values are derived from the table's own rows and recomputed as the data changes. Use it when the data itself is the source of truth and you only want to prevent new parent/child combinations.

eddytor set domain eddytor sales orders subcategory hierarchical-derived \
  --parent-column category
PUT /v1/tables/…/columns/subcategory/domain/hierarchical/derived

{ "parent_column": "category" }

Hierarchical - table (external lookup)

The mapping lives in another table and is resolved from it at query time - use it when a reference dataset (org chart, product taxonomy) is maintained elsewhere:

eddytor set domain eddytor sales orders region hierarchical-table \
  --source-table eddytor.cfg_xxx.<uuid>_geo --source-column region_name \
  --parent-column country
PUT /v1/tables/…/columns/region/domain/hierarchical/table

{ "parent_column": "country", "table_name": "<uuid>_geo",
  "table_path": "s3://bucket/path/to/geo",
  "parent_key": "country_name", "child_key": "region_name" }

The lookup table needs a parent-key column and a child-key column. Because the mapping is read live, changes to the lookup table apply immediately - and deleting lookup rows can orphan dependents (validate finds them).

Reference (cross-table)

Link a column to another table's domain - referential integrity across tables:

eddytor set domain eddytor sales orders customer_id reference \
  --source-table-path s3://bucket/path/customers \
  --source-table-name <uuid>_customers --source-column customer_id
PUT /v1/tables/…/columns/customer_id/domain/reference

{ "source_table": { "catalog": "eddytor", "schema": "cfg_xxx",
                    "table": "<uuid>_customers" },
  "source_column": "customer_id" }
set_column_domain(table, "customer_id", "reference",
  source_table="eddytor.cfg_xxx.<uuid>_customers", source_column="customer_id")

The source column must already have a domain configured. Reference domains validate live, so deleting a source value can orphan dependents - find them with validate_domain_values.

Order & gotchas

  • Set parent before child for hierarchies, and set domains before loading data - rejecting a bad value at write time is far cheaper than cleaning it up afterward.
  • Domain values are case-sensitive - "Active" and "active" are different values.
  • Requires the domains:write scope (Builder or Admin; workspace role governs on workspace-owned configs).

On this page