spread
Partitions the data and lays the groups out along an axis, with a gap between them. The workhorse operator for bar charts and small multiples.
from gofish import chart, spread, rect
chart(seafood).flow(spread(by="lake", dir="x")).mark(rect(h="count")).render(
w=500, h=300, axes=True
)Signature
spread(*, by=None, dir, **options) -> OperatorParameters
spread — Operator form
Arrange children along dir with spacing, aligning them on the cross axis.
| Option | Type | Default | Description |
|---|---|---|---|
by | str | Any | Field to partition rows by; also accepts a field(...) accessor carrying domain ops (sort/reverse/bin). | |
dir | str | Axis to spread along: x, y, or an axis name the enclosing coordinate space declares (polar theta/r, geo lon/lat). | |
spacing | float | 8 | Gap between children, px. |
alignment | str | "baseline" | Cross-axis alignment ("start" | "middle" | "end" | "baseline"). |
shared_scale | bool | False | Share one scale across all children. |
anchor | str | "edge" | Whether spacing is measured between facing edges (edge), or as a fixed pitch between the named anchor point on each child. |
reverse | bool | False | Reverse the children's order along dir. |
glue | bool | False | Stack semantics: children glued, sizes sum; spacing forced to 0. |
axes | Any | ||
x | int | float | str | Left edge of this operator's box, in the parent's space (pixels). Omitted, the parent places it. | |
y | int | float | str | Start edge on y (top where y reads top-down, bottom where it grows upward) of this operator's box, in the parent's space (pixels). Omitted, the parent places it. | |
w | int | float | str | Data-driven cross-axis extent (field/datum-sized children). | |
h | int | float | str | Data-driven cross-axis extent (field/datum-sized children). | |
size | int | float | str | Per-entry stack-axis extent (field/datum-sized children); a field(...).normalize() accessor makes it a space-filling spine. |
spread — Combinator form
Low-level combinator form of spread. Same fields as the operator form (OPERATORS.spread) plus key and the full box-dims group.
| Option | Type | Default | Description |
|---|---|---|---|
x | int | float | str | Left edge of this operator's box, in the parent's space (pixels). Omitted, the parent places it. | |
w | int | float | str | Data-driven cross-axis extent (field/datum-sized children). | |
y | int | float | str | Start edge on y (top where y reads top-down, bottom where it grows upward) of this operator's box, in the parent's space (pixels). Omitted, the parent places it. | |
h | int | float | str | Data-driven cross-axis extent (field/datum-sized children). | |
by | str | Any | Field to partition rows by; also accepts a field(...) accessor carrying domain ops (sort/reverse/bin). | |
dir | str | Axis to spread along: x, y, or an axis name the enclosing coordinate space declares (polar theta/r, geo lon/lat). | |
spacing | float | 8 | Gap between children, px. |
alignment | str | "baseline" | Cross-axis alignment ("start" | "middle" | "end" | "baseline"). |
shared_scale | bool | False | Share one scale across all children. |
anchor | str | "edge" | Whether spacing is measured between facing edges (edge), or as a fixed pitch between the named anchor point on each child. |
reverse | bool | False | Reverse the children's order along dir. |
glue | bool | False | Stack semantics: children glued, sizes sum; spacing forced to 0. |
axes | Any | ||
size | int | float | str | Per-entry stack-axis extent (field/datum-sized children); a field(...).normalize() accessor makes it a space-filling spine. | |
key | str | Internal per-node key override. |
Box dimensions
| Option | Type | Default | Description |
|---|---|---|---|
cx | int | float | str | Center x. | |
x2 | int | float | str | Right edge position. | |
em_x | bool | Embed x in the parent's x space. | |
cy | int | float | str | Center y. | |
y2 | int | float | str | Other y edge position. | |
em_y | bool | Embed y in the parent's y space. | |
dims | dict | Box dimensions by axis name: x/y, or a name the enclosing coordinate space declares (polar theta/r, geo lon/lat). Each value is a position (like x) or an interval {min, center, max, size, embedded}. |
Returns an Operator for use inside .flow().
Examples
# One bar per lake
chart(seafood).flow(spread(by="lake", dir="x")).mark(rect(h="count"))
# Wider gaps
chart(seafood).flow(spread(by="lake", dir="x", spacing=64)).mark(rect(h="count"))
# Nest spreads for grouped layouts
chart(seafood).flow(
spread(by="lake", dir="x"),
spread(by="species", dir="x", spacing=2),
).mark(rect(h="count", fill="species"))Naming the axis with dir
dir="x" and dir="y" mean the first and second axis in any coordinate space. dir also takes the names the enclosing coordinate space declares, so under polar you can write dir="theta" or dir="r":
chart(data, coord=polar()) \
.flow(spread(by="month", dir="theta", spacing=0)) \
.mark(rect(w=0.5, h="value"))dir="theta" lays the months out exactly as dir="x" does. A name the enclosing space does not declare raises an error that lists the names it does. stack takes the same dir.
Path-aware by
by accepts a field name, a dotted path string, or a field(...) accessor:
spread(by="species", dir="x") # field name
spread(by="origin.country", dir="x") # nested path
spread(by=field("species").sort(), dir="x") # field(...) accessorThe same bare field name works after a ref / select_all selection. The stream items are then refs, not raw records, but a ref is read through its .datum rows automatically, so you still write by="species". Do not add a datum. prefix: by="datum.species" looks for a field named datum inside each row, finds nothing, and puts every ref in one group.
How by resolves on a ref (homogeneity collapse)
A ref's .datum is the raw bag of rows that flowed into the node (a list; a fully-split leaf is a 1-row list). A by="field" path on a ref does not just read the field off the first row — it projects with homogeneity collapse:
fieldresolves to a scalar iff every row in the node's bag agrees on that field; otherwise it isNone— the "this field is multi-valued here, grouping by it is ill-posed" signal.
This is exactly SQL's ONLY_FULL_GROUP_BY / functional-dependency rule: you may only group by a column that is constant within each row-bag.
Example. After select_all("bars") where each ref is a lake aggregate of 5 species rows:
group(by="lake") # resolves — all 5 rows share one lake
group(by="species") # None — 5 distinct species; ill-posedTo group by a field that is multi-valued in the current bag, disaggregate first (split the bag so each child is homogeneous in that field). A fully-split cell (1 row) trivially collapses, so by="species" works once each node holds a single record.
So a ribbon chart reads:
chart(select_all("bars")) \
.flow(group(by="species")) \
.mark(ribbon(opacity=0.8))Field-expression pipeline
field(name) returns a chainable accessor — a builder where each method appends one op to an ordered pipeline. It works in two disjoint places:
- As
by(a domain slot):.sort(by=None, order=None),.sort(values)(an explicit order list),.reverse(),.bin(thresholds=None), and.drop_nulls()decide which groups exist and in what order. - As a mark or
sizechannel value (a value slot):.sum(),.mean(),.count(), and.distinct()fold a group's rows to one number, overriding the channel's own default aggregation (sum for size, mean for position).
Mixing the two — an aggregate op on by, or a domain op on a value channel — raises.
Sort a stack's groups by another field's total, instead of data order:
# Bars ordered ascending by their own total `value`
chart(data).flow(
spread(by=field("category").sort("value"), dir="x")
).mark(rect(h="value"))Omit the by argument to sort by the group key itself; pass order="desc" for descending.
Sort groups by an explicit order, for a domain-specific sequence no aggregate expresses (severity, calendar order, a fixed ranking):
# Weather categories in a fixed order, not alphabetical or by aggregate
chart(data).flow(
stack(by=field("weather").sort(["sun", "fog", "drizzle", "rain", "snow"]), dir="y")
)Groups whose key isn't in the list are appended after, in natural sort order.
Bin a numeric field into groups — a histogram, with no precomputed bins:
# One bar per ~10 auto-computed bins of `age`, height = row count per bin
chart(data).flow(
spread(by=field("age").bin(), dir="x")
).mark(rect(h=field("age").count()))Empty bins are dropped, like an ordinary group-by. Pass field("age").bin(5) (a count) or explicit thresholds (a list) to control the binning.
Drop rows with a missing/null grouping field, before grouping:
# Rows with a null/missing "genre" are excluded entirely, instead of
# forming their own "null" group
chart(data).flow(
treemap(by=field("genre").drop_nulls(), size="worldwideGross")
)Override a channel's default aggregation:
# The bar height is each species' MEAN weight, not the sum
chart(data).flow(
spread(by="species", dir="x")
).mark(rect(h=field("weight").mean())).count() and .distinct() report measure "count" (they're counts, not the source field's own units) unless you annotate the accessor explicitly: field("id", measure="my-measure").distinct().
Space-filling spines (mosaic / marimekko)
size=field(<name>).normalize() turns a stack's entries into a space-filling spine: each entry's size becomes its SHARE of the window (the operator's own split entries, v_e / Σv_e), so the axis reads as a local 0–100% conditional distribution. It replaces the removed normalize=True layout flag — .normalize() is a data transform on the size channel (a windowed share, computed once up front), not a layout mode, so the same field can drive both a cross-axis marginal size and the normalized fill with no preprocessing.
Nest two stacks on alternating axes for a mosaic: the outer sizes each column by its raw total (size="count" — the marginal), the inner sizes by each entry's share (the conditional):
(
chart(passengers, axes=True)
.flow(
# columns by class — width ∝ each class's count (marginal)
stack(by="pclass", dir="x", size="count"),
# survival share within each column (conditional), filling height
stack(by="survived", dir="y", size=field("count").normalize()),
)
.mark(rect(fill="survived", stroke="white", stroke_width=1))
.render(w=400, h=300)
).normalize() composes to any depth: a third alternating level gives a nested mosaic (class → sex → survived). Because each level's stacking axis is a local self-scaling region and the count is read raw everywhere, the marginal × conditional × … area factorization holds all the way down.
WARNING
Inner conditional axes are local scopes, so they don't yet render 0–1 tick labels — only the outermost marginal axis is labeled.
Notes
diris required —spread()raises aValueErrorwithout it.- Use
stackwhen you want groups touching edge-to-edge with no gap. - Data order determines group order; sort your data first if order matters.
