minipivot

The purpose of minipivot is to provide a zero-dependency reimplementation of two functions from the tidyr package: pivot_longer() and pivot_wider(). Together they reshape a data frame between “long” and “wide” layouts. The README provides scope notes; this page walks through both functions on a small toy data frame, including the column-selection styles .pivot_longer() supports and what happens when .pivot_wider() is asked to spread duplicate rows into one cell.

Functions

The minipivot.R script supplies two core functions:

  • .pivot_longer(.data, cols, names_to = "name", values_to = "value") is used to melt cols into two new columns holding the pivoted column names and their values.
  • .pivot_wider(.data, names_from, values_from, values_fill = NULL) is used to spread the unique values of names_from into new columns, filled from values_from.

It contains no internal helpers.

A minimal data set

fish <- data.frame(
  location = c("lake", "lake", "sea", "sea"),
  species  = c("trout", "bass", "trout", "bass"),
  count    = c(5, 3, 7, 2)
)
fish
  location species count
1     lake   trout     5
2     lake    bass     3
3      sea   trout     7
4      sea    bass     2

Wide to long

The .pivot_longer() function needs to know which columns to pivot. Here, count is the only column with values to melt, and location/species come along for the ride as id columns:

.pivot_longer(fish, count, names_to = "metric", values_to = "n")
  location species metric n
1     lake   trout  count 5
2     lake    bass  count 3
3      sea   trout  count 7
4      sea    bass  count 2

Long to wide

The .pivot_wider() function spreads species’s unique values (trout/bass) into new columns, filled from count. There’s no id_cols argument to set – every column other than names_from/values_from (location, here) is treated as an id column automatically:

wide <- .pivot_wider(fish, names_from = "species", values_from = "count")
wide
  location trout bass
1     lake     5    3
2      sea     7    2

Round-tripping

Pivoting the wide result back out reproduces the original long data (row order aside, since .pivot_wider() reorders rows by id column):

.pivot_longer(wide, c(trout, bass), names_to = "species", values_to = "count")
  location species count
1     lake   trout     5
2     lake    bass     3
3      sea   trout     7
4      sea    bass     2

Selecting cols a few different ways

The cols argument to .pivot_longer() accepts more than just a single bare name. Bare names combined with c(...), a plain character vector, and a negated -c(...) (“every column except these”) all agree, since trout/bass are the only two columns not already used as an id column above:

identical(
  .pivot_longer(wide, c(trout, bass)),
  .pivot_longer(wide, c("trout", "bass"))
)
[1] TRUE
identical(
  .pivot_longer(wide, c(trout, bass)),
  .pivot_longer(wide, -location)
)
[1] TRUE

There is no tidyselect-style starts_with()/everything() support – just bare names, c(...), character/numeric vectors, and unary - negation.

Missing combinations: values_fill

Dropping a row before widening leaves a combination unrepresented – sea/bass is missing below, so it comes back as NA by default:

.pivot_wider(fish[-4, ], names_from = "species", values_from = "count")
  location trout bass
1     lake     5    3
2      sea     7   NA

The values_fill argument supplies a value to use instead:

.pivot_wider(fish[-4, ], names_from = "species", values_from = "count", values_fill = 0)
  location trout bass
1     lake     5    3
2      sea     7    0

Duplicate combinations are an error, not a list-column

Real tidyr::pivot_wider() handles a duplicate names_from/id combination by warning and packing the extra values into a list-column. This mini has no list-column support (matching minitable’s scope), so the same situation is a hard error instead:

dup <- rbind(fish, data.frame(location = "lake", species = "trout", count = 99))
.pivot_wider(dup, names_from = "species", values_from = "count")
Error in `.pivot_wider()`:
! .pivot_wider(): `names_from`/id-column combination is not unique (duplicate rows would need to be aggregated); this mini does not implement a `values_fn`-style aggregation step.