minicase

The purpose of minicase is to provide a zero-dependency reimplementation of a single function from the dplyr package: case_when(). It performs a vectorised if/else across a sequence of condition ~ value formulas, where the first match wins. The README provides scope notes; this page walks through it on the same toy fruits data frame used in the miniverb vignette, building up to how missing values and default values interact.

Functions

The minicase.R script supplies one core function:

  • .case_when(...) is used to take a sequence of condition ~ value formulas and return a vector the same length as the (recycled) conditions/values, taking the value from the first formula whose condition is TRUE at each position, and NA wherever no condition matched.

It is supported by five internal helpers, all of which support .case_when()’s recycling and type-consistency checks – there’s no need to call any of them directly. The helpers are these:

  • .case_validate_length()
  • .case_replace_with()
  • .case_check_length()
  • .case_check_type()
  • .case_check_class()

A minimal data set

fruits <- data.frame(
  name     = c("apple", "banana", "cherry", "date", "elderberry"),
  price    = c(1.20, 0.50, 3.00, NA, 4.50),
  in_stock = c(TRUE, TRUE, FALSE, TRUE, TRUE)
)
fruits
        name price in_stock
1      apple   1.2     TRUE
2     banana   0.5     TRUE
3     cherry   3.0    FALSE
4       date    NA     TRUE
5 elderberry   4.5     TRUE

No data-masking

There is no data-masking in this mini: you must reference columns with $, or wrap in with(). Unlike dplyr::case_when(), conditions and values are evaluated in the caller’s environment, not the data frame’s – so bare column names don’t resolve on their own. The idiomatic way around that, mirrored here from base R, is to wrap the whole call in with(data, ...), which temporarily makes the data frame’s columns available as bare names for the duration of the call:

with(fruits, .case_when(
  price < 1 ~ "cheap",
  price < 3 ~ "mid",
  price >= 3 ~ "pricey"
))
[1] "mid"    "cheap"  "pricey" NA       "pricey"

Referring to columns via fruits$price directly, without with(), works exactly the same way – with() is a convenience, not a requirement:

identical(
  with(fruits, .case_when(price < 3 ~ "cheap", TRUE ~ "not cheap")),
  .case_when(fruits$price < 3 ~ "cheap", TRUE ~ "not cheap")
)
[1] TRUE

First match wins

Formulas are checked in order, and the first TRUE condition for each position determines its value, even if a later condition would also match. The cherry row (price 3.00) is a good illustration: it fails price < 3, but matches price >= 3, so it’s classified as "pricey" above. Conditions can combine several tests with &, which is useful for layering more specific rules ahead of more general ones:

with(fruits, .case_when(
  in_stock & price < 1 ~ "cheap and in stock",
  in_stock             ~ "in stock",
  TRUE                 ~ "out of stock or unknown price"
))
[1] "in stock"                      "cheap and in stock"           
[3] "out of stock or unknown price" "in stock"                     
[5] "in stock"                     

The banana row matches both the first and second formula here (it’s in stock and under $1), but because the first-match-wins rule checks formulas top to bottom, it’s classified by the more specific "cheap and in stock" rule rather than the more general "in stock" one. Ordering formulas from most to least specific is exactly how you get that behaviour deliberately, rather than by accident.

Unmatched positions are NA

The date row has a missing price, so every condition involving price evaluates to NA for that row – neither TRUE nor FALSE – and date itself gets NA in the result, since there’s no condition it definitively satisfies:

with(fruits, .case_when(
  price < 1 ~ "cheap",
  price < 3 ~ "mid",
  price >= 3 ~ "pricey"
))
[1] "mid"    "cheap"  "pricey" NA       "pricey"

Add a TRUE ~ ... formula at the end as a catch-all default, the same way you would in dplyr, and every remaining position – including ones where every prior condition was NA rather than FALSE – picks up that value instead:

with(fruits, .case_when(
  price < 1  ~ "cheap",
  price < 3  ~ "mid",
  price >= 3 ~ "pricey",
  TRUE       ~ "unknown"
))
[1] "mid"     "cheap"   "pricey"  "unknown" "pricey" 

Factors must share the same levels, not just the same class

Values across formulas are also checked for type/class consistency. For factors specifically, sharing class "factor" isn’t enough – the levels must match too, since assigning a factor label that isn’t among an existing factor’s levels would otherwise silently produce NA rather than raising a visible error:

f1 <- factor(c("a", "b"), levels = c("a", "b"))
f2 <- factor(c("x", "y"), levels = c("x", "y"))
.case_when(c(TRUE, FALSE) ~ f1, c(FALSE, TRUE) ~ f2)
Error in `.case_check_class()`:
! Factor levels must match across `.case_when()` values: expected levels `a`, `b`, got `x`, `y`.

Two factors with the same levels (in the same order) combine without any trouble:

f3 <- factor(c("a", "b"), levels = c("a", "b"))
f4 <- factor(c("b", "a"), levels = c("a", "b"))
.case_when(c(TRUE, FALSE) ~ f3, c(FALSE, TRUE) ~ f4)
[1] a a
Levels: a b

Scope: no .default/.ptype/.size

The mini does not implement any of these arguments. The trailing TRUE ~ "unknown" formula in the previous example is the idiomatic way to get a default value, and is in fact very often used when writing code with the real dplyr::case_when() function. There was no need for this mini to implement dplyr’s newer, separate .default argument as an alternative spelling of the same thing, and little to be gained by adding .ptype/.size to control the output type/length explicitly. If your use case needs those, a plain TRUE ~ value formula covers .default, and the output length/type simply follow from the conditions. If that still isn’t sufficient, then it’s likely you need the real version, not the mini.