miniverb

The purpose of miniverb is to provide a zero-dependency reimplementation of five one-table verbs from the dplyr package: filter(), select(), mutate(), arrange(), and summarise(). The README provides full scope notes; this page walks through each verb in turn on a small toy data frame, including a few things worth knowing about before relying on them: how missing values are handled, what happens when no conditions are supplied at all, and how the shared .by argument groups rows for filter(), mutate(), and summarise().

Functions

The miniverb.R script supplies:

  • .verb_filter(.data, ..., .by = NULL) keeps rows of .data where every condition in ... evaluates to TRUE. Conditions are evaluated in the scope of .data, so columns can be referred to by bare name.
  • .verb_select(.data, ...) keeps (and optionally renames) columns by bare name, integer position, or - exclusion.
  • .verb_mutate(.data, ..., .by = NULL) adds or overwrites columns via named name = expr arguments.
  • .verb_arrange(.data, ...) reorders rows by one or more columns.
  • .verb_desc(x) marks a column for descending order inside .verb_arrange().
  • .verb_summarise(.data, ..., .by = NULL) collapses to one row per group (or one row overall) via named name = expr arguments.

All five share a .by grouping convention (where applicable): a plain character vector of column names, not dplyr’s unquoted/tidyselect-lite syntax.

There are also two internal helpers:

  • .verb_dotdotdot() captures the unevaluated ... conditions passed to .verb_filter(), so each one can be evaluated against the data frame individually.
  • .verb_split_by() splits row indices into groups by .by’s column names, shared by .verb_filter(), .verb_mutate(), and .verb_summarise().

A minimal data set

fruits <- data.frame(
  name     = c("apple", "banana", "cherry", "date", "elderberry"),
  price    = c(1.20, 0.50, 3.00, NA, 4.50),
  in_stock = c(TRUE, TRUE, FALSE, TRUE, TRUE)
)
fruits
        name price in_stock
1      apple   1.2     TRUE
2     banana   0.5     TRUE
3     cherry   3.0    FALSE
4       date    NA     TRUE
5 elderberry   4.5     TRUE

filter()

A single condition

Columns are referred to by bare name, evaluated in the scope of the data frame – there’s no need to write fruits$in_stock or wrap the call in with(), as you would for minicase’s .case_when():

.verb_filter(fruits, in_stock)
        name price in_stock
1      apple   1.2     TRUE
2     banana   0.5     TRUE
4       date    NA     TRUE
5 elderberry   4.5     TRUE

Multiple conditions combine with AND

Passing more than one condition is equivalent to combining them with &, so a row is kept only if every condition holds. This is often more readable than a single expression joined with &&/&, especially as the number of conditions grows:

.verb_filter(fruits, in_stock, price < 4)
    name price in_stock
1  apple   1.2     TRUE
2 banana   0.5     TRUE

That’s equivalent to writing the conditions out explicitly and combining them yourself:

identical(
  .verb_filter(fruits, in_stock, price < 4),
  fruits[fruits$in_stock & fruits$price < 4 & !is.na(fruits$price), ]
)
[1] TRUE

NA conditions drop the row

The date row has a missing price, so any condition involving price is NA for that row rather than TRUE/FALSE. The .verb_filter() function treats NA the same way dplyr::filter() does: as “not proven to be kept”, so the row is dropped rather than kept by default or causing an error:

.verb_filter(fruits, price > 1)
        name price in_stock
1      apple   1.2     TRUE
3     cherry   3.0    FALSE
5 elderberry   4.5     TRUE

Notice that date disappears from the result entirely, rather than appearing with NA values or being kept because “we don’t know” – this is a common surprise for anyone used to base R’s [ subsetting, where an NA index instead produces a row of NAs:

fruits[fruits$price > 1, ]
         name price in_stock
1       apple   1.2     TRUE
3      cherry   3.0    FALSE
NA       <NA>    NA       NA
5  elderberry   4.5     TRUE

No conditions returns the data frame unchanged

Calling .verb_filter() with no conditions at all is legitimate and simply leaves the data frame unchanged – useful when conditions are built up programmatically and might end up empty:

identical(.verb_filter(fruits), fruits)
[1] TRUE

Grouped conditions with .by

The .by argument takes a character vector of column names and evaluates the conditions separately within each group – useful when a condition depends on a per-group quantity like a group mean:

.verb_filter(fruits, price > mean(price, na.rm = TRUE), .by = "in_stock")
        name price in_stock
5 elderberry   4.5     TRUE

select()

Keeping and renaming columns

.verb_select() keeps columns in the order they’re listed, and a new = old argument renames a column on the way through:

.verb_select(fruits, name, cost = price)
        name cost
1      apple  1.2
2     banana  0.5
3     cherry  3.0
4       date   NA
5 elderberry  4.5

Dropping columns instead

A - in front of a column name drops it rather than keeping it, which reads more naturally than listing every column to keep when only one or two need to go. Positive and negative selections can’t be mixed in the same call:

.verb_select(fruits, -in_stock)
        name price
1      apple   1.2
2     banana   0.5
3     cherry   3.0
4       date    NA
5 elderberry   4.5

mutate()

Adding a new column

.verb_mutate() takes named name = expr arguments and adds each one as a new column, evaluated in the scope of fruits:

.verb_mutate(fruits, discounted = round(price * 0.9, 2))
        name price in_stock discounted
1      apple   1.2     TRUE       1.08
2     banana   0.5     TRUE       0.45
3     cherry   3.0    FALSE       2.70
4       date    NA     TRUE         NA
5 elderberry   4.5     TRUE       4.05

Later columns can use earlier ones

Arguments are evaluated left to right, so an expression can refer to a column added by an earlier argument in the same call. Below, savings uses discounted even though it doesn’t exist in fruits itself:

.verb_mutate(
  fruits,
  discounted = round(price * 0.9, 2),
  savings = round(price - discounted, 2)
)
        name price in_stock discounted savings
1      apple   1.2     TRUE       1.08    0.12
2     banana   0.5     TRUE       0.45    0.05
3     cherry   3.0    FALSE       2.70    0.30
4       date    NA     TRUE         NA      NA
5 elderberry   4.5     TRUE       4.05    0.45

Grouped columns with .by

Passing .by evaluates each expression separately within each group, then puts the results back in the original row order – here, each fruit’s price is compared to the average price within its own in_stock group rather than the average across all five fruits:

.verb_mutate(
  fruits,
  price_vs_group = round(price - mean(price, na.rm = TRUE), 2),
  .by = "in_stock"
)
        name price in_stock price_vs_group
1      apple   1.2     TRUE          -0.87
2     banana   0.5     TRUE          -1.57
3     cherry   3.0    FALSE           0.00
4       date    NA     TRUE             NA
5 elderberry   4.5     TRUE           2.43

arrange()

Ascending by default, descending with .verb_desc()

.verb_arrange() sorts ascending by default. Wrapping a column in .verb_desc() sorts it descending instead, and NAs sort last either way, matching dplyr::arrange():

.verb_arrange(fruits, .verb_desc(price))
        name price in_stock
5 elderberry   4.5     TRUE
3     cherry   3.0    FALSE
1      apple   1.2     TRUE
2     banana   0.5     TRUE
4       date    NA     TRUE

Breaking ties with a second column

Extra arguments break ties left over from earlier ones, the same way order() does. Sorting by in_stock first and price second groups the out-of-stock cherry row on its own, then orders the remaining rows by descending price:

.verb_arrange(fruits, in_stock, .verb_desc(price))
        name price in_stock
3     cherry   3.0    FALSE
5 elderberry   4.5     TRUE
1      apple   1.2     TRUE
2     banana   0.5     TRUE
4       date    NA     TRUE

summarise()

One row overall

With no .by, .verb_summarise() collapses fruits down to a single row, computing each named expression once over every row:

.verb_summarise(fruits, average_price = round(mean(price, na.rm = TRUE), 2), n = length(name))
  average_price n
1           2.3 5

One row per group with .by

With .by, one summary row is produced per group instead, and the grouping columns appear first in the result:

.verb_summarise(
  fruits,
  average_price = round(mean(price, na.rm = TRUE), 2),
  n = length(name),
  .by = "in_stock"
)
  in_stock average_price n
1     TRUE          2.07 4
2    FALSE          3.00 1