Skip to content

spread ​

Partitions the data and lays the groups out along an axis, with a gap between them. The workhorse operator for bar charts and small multiples.

python
from gofish import chart, spread, rect

chart(seafood).flow(spread(by="lake", dir="x")).mark(rect(h="count")).render(
    w=500, h=300, axes=True
)

Signature ​

python
spread(*, by=None, dir, **options) -> Operator

Parameters ​

spread — Operator form ​

Arrange children along dir with spacing, aligning them on the cross axis.

OptionTypeDefaultDescription
bystr | AnyField to partition rows by; also accepts a field(...) accessor carrying domain ops (sort/reverse/bin).
dirstrAxis to spread along: x, y, or an axis name the enclosing coordinate space declares (polar theta/r, geo lon/lat).
spacingfloat8Gap between children, px.
alignmentstr"baseline"Cross-axis alignment ("start" | "middle" | "end" | "baseline").
shared_scaleboolFalseShare one scale across all children.
anchorstr"edge"Whether spacing is measured between facing edges (edge), or as a fixed pitch between the named anchor point on each child.
reverseboolFalseReverse the children's order along dir.
glueboolFalseStack semantics: children glued, sizes sum; spacing forced to 0.
axesAny
xint | float | strLeft edge of this operator's box, in the parent's space (pixels). Omitted, the parent places it.
yint | float | strStart edge on y (top where y reads top-down, bottom where it grows upward) of this operator's box, in the parent's space (pixels). Omitted, the parent places it.
wint | float | strData-driven cross-axis extent (field/datum-sized children).
hint | float | strData-driven cross-axis extent (field/datum-sized children).
sizeint | float | strPer-entry stack-axis extent (field/datum-sized children); a field(...).normalize() accessor makes it a space-filling spine.

spread — Combinator form ​

Low-level combinator form of spread. Same fields as the operator form (OPERATORS.spread) plus key and the full box-dims group.

OptionTypeDefaultDescription
xint | float | strLeft edge of this operator's box, in the parent's space (pixels). Omitted, the parent places it.
wint | float | strData-driven cross-axis extent (field/datum-sized children).
yint | float | strStart edge on y (top where y reads top-down, bottom where it grows upward) of this operator's box, in the parent's space (pixels). Omitted, the parent places it.
hint | float | strData-driven cross-axis extent (field/datum-sized children).
bystr | AnyField to partition rows by; also accepts a field(...) accessor carrying domain ops (sort/reverse/bin).
dirstrAxis to spread along: x, y, or an axis name the enclosing coordinate space declares (polar theta/r, geo lon/lat).
spacingfloat8Gap between children, px.
alignmentstr"baseline"Cross-axis alignment ("start" | "middle" | "end" | "baseline").
shared_scaleboolFalseShare one scale across all children.
anchorstr"edge"Whether spacing is measured between facing edges (edge), or as a fixed pitch between the named anchor point on each child.
reverseboolFalseReverse the children's order along dir.
glueboolFalseStack semantics: children glued, sizes sum; spacing forced to 0.
axesAny
sizeint | float | strPer-entry stack-axis extent (field/datum-sized children); a field(...).normalize() accessor makes it a space-filling spine.
keystrInternal per-node key override.

Box dimensions ​

OptionTypeDefaultDescription
cxint | float | strCenter x.
x2int | float | strRight edge position.
em_xboolEmbed x in the parent's x space.
cyint | float | strCenter y.
y2int | float | strOther y edge position.
em_yboolEmbed y in the parent's y space.
dimsdictBox dimensions by axis name: x/y, or a name the enclosing coordinate space declares (polar theta/r, geo lon/lat). Each value is a position (like x) or an interval {min, center, max, size, embedded}.

Returns an Operator for use inside .flow().

Examples ​

python
# One bar per lake
chart(seafood).flow(spread(by="lake", dir="x")).mark(rect(h="count"))

# Wider gaps
chart(seafood).flow(spread(by="lake", dir="x", spacing=64)).mark(rect(h="count"))

# Nest spreads for grouped layouts
chart(seafood).flow(
    spread(by="lake", dir="x"),
    spread(by="species", dir="x", spacing=2),
).mark(rect(h="count", fill="species"))

Naming the axis with dir ​

dir="x" and dir="y" mean the first and second axis in any coordinate space. dir also takes the names the enclosing coordinate space declares, so under polar you can write dir="theta" or dir="r":

python
chart(data, coord=polar()) \
    .flow(spread(by="month", dir="theta", spacing=0)) \
    .mark(rect(w=0.5, h="value"))

dir="theta" lays the months out exactly as dir="x" does. A name the enclosing space does not declare raises an error that lists the names it does. stack takes the same dir.

Path-aware by ​

by accepts a field name, a dotted path string, or a field(...) accessor:

python
spread(by="species", dir="x")                # field name
spread(by="origin.country", dir="x")         # nested path
spread(by=field("species").sort(), dir="x")  # field(...) accessor

The same bare field name works after a ref / select_all selection. The stream items are then refs, not raw records, but a ref is read through its .datum rows automatically, so you still write by="species". Do not add a datum. prefix: by="datum.species" looks for a field named datum inside each row, finds nothing, and puts every ref in one group.

How by resolves on a ref (homogeneity collapse) ​

A ref's .datum is the raw bag of rows that flowed into the node (a list; a fully-split leaf is a 1-row list). A by="field" path on a ref does not just read the field off the first row — it projects with homogeneity collapse:

field resolves to a scalar iff every row in the node's bag agrees on that field; otherwise it is None — the "this field is multi-valued here, grouping by it is ill-posed" signal.

This is exactly SQL's ONLY_FULL_GROUP_BY / functional-dependency rule: you may only group by a column that is constant within each row-bag.

Example. After select_all("bars") where each ref is a lake aggregate of 5 species rows:

python
group(by="lake")     # resolves — all 5 rows share one lake
group(by="species")  # None — 5 distinct species; ill-posed

To group by a field that is multi-valued in the current bag, disaggregate first (split the bag so each child is homogeneous in that field). A fully-split cell (1 row) trivially collapses, so by="species" works once each node holds a single record.

So a ribbon chart reads:

python
chart(select_all("bars")) \
    .flow(group(by="species")) \
    .mark(ribbon(opacity=0.8))

Field-expression pipeline ​

field(name) returns a chainable accessor — a builder where each method appends one op to an ordered pipeline. It works in two disjoint places:

  • As by (a domain slot): .sort(by=None, order=None), .sort(values) (an explicit order list), .reverse(), .bin(thresholds=None), and .drop_nulls() decide which groups exist and in what order.
  • As a mark or size channel value (a value slot): .sum(), .mean(), .count(), and .distinct() fold a group's rows to one number, overriding the channel's own default aggregation (sum for size, mean for position).

Mixing the two — an aggregate op on by, or a domain op on a value channel — raises.

Sort a stack's groups by another field's total, instead of data order:

python
# Bars ordered ascending by their own total `value`
chart(data).flow(
    spread(by=field("category").sort("value"), dir="x")
).mark(rect(h="value"))

Omit the by argument to sort by the group key itself; pass order="desc" for descending.

Sort groups by an explicit order, for a domain-specific sequence no aggregate expresses (severity, calendar order, a fixed ranking):

python
# Weather categories in a fixed order, not alphabetical or by aggregate
chart(data).flow(
    stack(by=field("weather").sort(["sun", "fog", "drizzle", "rain", "snow"]), dir="y")
)

Groups whose key isn't in the list are appended after, in natural sort order.

Bin a numeric field into groups — a histogram, with no precomputed bins:

python
# One bar per ~10 auto-computed bins of `age`, height = row count per bin
chart(data).flow(
    spread(by=field("age").bin(), dir="x")
).mark(rect(h=field("age").count()))

Empty bins are dropped, like an ordinary group-by. Pass field("age").bin(5) (a count) or explicit thresholds (a list) to control the binning.

Drop rows with a missing/null grouping field, before grouping:

python
# Rows with a null/missing "genre" are excluded entirely, instead of
# forming their own "null" group
chart(data).flow(
    treemap(by=field("genre").drop_nulls(), size="worldwideGross")
)

Override a channel's default aggregation:

python
# The bar height is each species' MEAN weight, not the sum
chart(data).flow(
    spread(by="species", dir="x")
).mark(rect(h=field("weight").mean()))

.count() and .distinct() report measure "count" (they're counts, not the source field's own units) unless you annotate the accessor explicitly: field("id", measure="my-measure").distinct().

Space-filling spines (mosaic / marimekko) ​

size=field(<name>).normalize() turns a stack's entries into a space-filling spine: each entry's size becomes its SHARE of the window (the operator's own split entries, v_e / Σv_e), so the axis reads as a local 0–100% conditional distribution. It replaces the removed normalize=True layout flag — .normalize() is a data transform on the size channel (a windowed share, computed once up front), not a layout mode, so the same field can drive both a cross-axis marginal size and the normalized fill with no preprocessing.

Nest two stacks on alternating axes for a mosaic: the outer sizes each column by its raw total (size="count" — the marginal), the inner sizes by each entry's share (the conditional):

python
(
    chart(passengers, axes=True)
    .flow(
        # columns by class — width ∝ each class's count (marginal)
        stack(by="pclass", dir="x", size="count"),
        # survival share within each column (conditional), filling height
        stack(by="survived", dir="y", size=field("count").normalize()),
    )
    .mark(rect(fill="survived", stroke="white", stroke_width=1))
    .render(w=400, h=300)
)

.normalize() composes to any depth: a third alternating level gives a nested mosaic (class → sex → survived). Because each level's stacking axis is a local self-scaling region and the count is read raw everywhere, the marginal × conditional × … area factorization holds all the way down.

WARNING

Inner conditional axes are local scopes, so they don't yet render 0–1 tick labels — only the outermost marginal axis is labeled.

Notes ​

  • dir is required — spread() raises a ValueError without it.
  • Use stack when you want groups touching edge-to-edge with no gap.
  • Data order determines group order; sort your data first if order matters.