minimap

The purpose of minimap is to provide zero-dependency mapping functions analogous to those supplied by the purrr package. Each function iterates over a collection of inputs, applying a user-supplied function to each of the inputs. The README provides full API reference and scope notes; this page walks through the functionality supplied, providing a short tutorial that starts from the basic case and building up to the two-input and side-effect variants.

Functions

The minimap.R script supplies the following functions:

  • .iter_map(.x, .f) is the backbone function in the script, built atop the base lapply() function. It applies the single-argument function .f to every element of .x, always returning a list, and preserving names if they are present.
  • .iter_map_dbl(.x, .f), .iter_map_lgl(.x, .f), and .iter_map_chr(.x, .f) are the typed variants, built atop the base vapply() function. They apply .f to every element of .x and return an atomic vector of the named type (double, logical, or character) instead of a list.
  • .iter_map2(.x, .y, .f) is the two-input variation, where the two inputs .x and .y are the same length, and .f is a function that takes two inputs. The output is a list that applies .f to each pair of input values.
  • .iter_imap(.x, .f) is the variation in which the user passes a named vector or list .x, and .f is a two-argument function that takes the value as its first argument and the name as the second argument, returning a list.
  • .iter_walk(.x, .f) behaves analogously to .iter_map() but is used only for the side effects of calling the function .f. It returns .x invisibly instead of the results.
  • .iter_iwalk(.x, .f) behaves analogously to .iter_imap() but is used only for the side effects of calling the function .f. It returns .x invisibly instead of the results.

There is one internal function:

  • .iter_assert(expr, message) is a small argument-checking helper built on base stop(), used to validate inputs like matching lengths or the presence of names.

Basic mapping

.iter_map() is the workhorse of the mini: apply a function to each element of a list or atomic vector, and get back a list of results the same length, with any names from the input preserved on the output. It’s a thin wrapper around lapply(), so anything you already know about lapply()’s behaviour carries over directly:

.iter_map(1:3, function(x) x + 1)
[[1]]
[1] 2

[[2]]
[1] 3

[[3]]
[1] 4

If the input is named, the names are preserved in the output:

.iter_map(c(a = 1, b = 2), function(x) x * 10)
$a
[1] 10

$b
[1] 20

Because the return type is always a list, .iter_map() is safe to use even when .f returns different types or lengths for different elements:

.iter_map(1:3, function(x) if (x == 2) c(x, x) else x)
[[1]]
[1] 1

[[2]]
[1] 2 2

[[3]]
[1] 3

Type-stable variants

Often you know in advance that every element of the result will be a single number, logical, or string – in which case getting a list back is inconvenient. .iter_map_dbl(), .iter_map_lgl(), and .iter_map_chr() cover that case: they behave like .iter_map(), but return an atomic vector of the named type via vapply(), which also means they validate that every call to .f really does produce a length-1 value of the expected type:

.iter_map_dbl(1:3, function(x) x + 1)
[1] 2 3 4
.iter_map_lgl(1:5, function(x) x %% 2 == 0)
[1] FALSE  TRUE FALSE  TRUE FALSE
.iter_map_chr(c(a = 1, b = 2), function(x) sprintf("value is %s", x))
           a            b 
"value is 1" "value is 2" 

If .f breaks that promise – by returning a vector of length 2 instead of 1, say – the call errors immediately rather than silently producing a malformed or recycled result:

.iter_map_dbl(1:3, function(x) c(x, x))
Error in `vapply()`:
! values must be length 1,
 but FUN(X[[1]]) result is length 2

Two-input variants

Sometimes one input isn’t enough. .iter_map2() walks two same-length vectors in parallel, calling .f with one element from each:

.iter_map2(1:3, 4:6, function(x, y) x + y)
[[1]]
[1] 5

[[2]]
[1] 7

[[3]]
[1] 9

Passing vectors of mismatched length is a deliberate error rather than silent recycling:

.iter_map2(1:3, 1:2, function(x, y) x + y)
Error:
! `.x` and `.y` must have the same length

.iter_imap() covers a related but different need: iterating over a single named vector together with its own names. Strictly speaking it is a two-input mapping function, but the inputs are the values of .x and the names of .x, and .f is called with the value first and the name second:

.iter_imap(c(a = 1, b = 2), function(val, name) paste0(name, "=", val))
$a
[1] "a=1"

$b
[1] "b=2"

Note that .x must actually be named for .iter_imap() to make sense, and that’s enforced. An error is produced if .x does not have names:

.iter_imap(1:3, function(val, name) paste0(name, "=", val))
Error:
! `.x` must be named

Iterating for the side-effects only

The .iter_map() and .iter_imap() functions exist for the situation where you want the results of applying a function to every input. Sometimes, however, you’re not actually interested in what .f returns, but are instead interested in what it does as a side effect (e.g., save a file). In this situation, you may prefer to use the .iter_walk() and .iter_iwalk() variations. These “walk” functions iterate over the inputs in the same way that the analogous “map” functions do, producing side-effects as the mapping unfolds, but they don’t capture the results. Instead, they always return the original .x input invisibly.

.iter_walk(1:3, function(x) cat("value:", x, "\n"))
value: 1 
value: 2 
value: 3 
.iter_iwalk(c(a = 1, b = 2), function(val, name) cat(name, "->", val, "\n"))
a -> 1 
b -> 2 

Because the input is returned invisibly, .iter_walk()/.iter_iwalk() can be dropped into the middle of a longer expression purely for their printing side effect, without changing what the expression evaluates to:

result <- .iter_walk(1:3, function(x) cat("processing", x, "\n"))
processing 1 
processing 2 
processing 3 
identical(result, 1:3)
[1] TRUE