minitable

The purpose of minitable is to provide a zero-dependency reimplementation of a few construction/coercion helpers from the tibble package, built entirely on data frames. The results are plain data frames, not actual tibbles. The README provides the full scope notes; this page walks through each function, including the cross-column-reference behaviour that’s the trickiest part of the mini to get right.

Functions

The minitable.R script supplies four core functions:

  • .table_tibble(...) is used to build a data frame column by column, sequentially, so later columns can refer to earlier ones by name.
  • .table_as_tibble(x, ...) is used to coerce x to a plain data frame via as.data.frame(), preserving literal (non-syntactic) column names.
  • .table_rownames_to_column(.data, var) is used to move .data’s row names into an explicit column named var; it leaves .data unchanged if the row names are just the default sequential ones.
  • .table_add_row(.data, ...) is used to append a single row to .data, with the same cross-column-reference support as .table_tibble().

There is also one internal helper:

  • .table_drop_dup_list() is used by both .table_tibble() and .table_add_row() to resolve what a column reference means when a name is reused more than once while constructing columns in sequence (see below).

Sequential, cross-referencing construction

Building a data frame with data.frame() directly requires every column to already exist as a complete vector before the call – you can’t refer to a column you’re building in the same call. Like tibble::tibble(), .table_tibble() lifts that restriction: later columns can refer to earlier ones by name, because each column expression is evaluated in order, in an environment built up from the columns already constructed:

.table_tibble(a = 1:3, b = a * 2, c = a + b)
  a b c
1 1 2 3
2 2 4 6
3 3 6 9

Here b refers to a, and c refers to both a and b – neither would resolve to anything if evaluated outside the call, since a, b, and c don’t otherwise exist in this session. Getting this right took real care in the implementation, but you can look at minitable.R directly if you’re interested in the mechanics.

Unnamed arguments are named after their deparsed expression

Just as in tibble::tibble(), an argument with no name = is named after its own source expression, converted to a string:

.table_tibble(1:3, sqrt(1:3))
  1:3 sqrt(1:3)
1   1  1.000000
2   2  1.414214
3   3  1.732051

This is mostly useful for quick, throwaway construction; naming columns explicitly (.table_tibble(x = 1:3, y = sqrt(1:3))) is more readable for anything that will be used again later.

Minimalist coercion

There’s no tibble class or special printing here – it’s just as.data.frame() with check.names = FALSE, so a literal, non-syntactic column name (like "1:3" from the previous section) survives unchanged instead of being sanitised by make.names():

.table_as_tibble(list(a = 1:2, b = c("x", "y")))
  a b
1 1 x
2 2 y

The difference from calling as.data.frame() directly shows up with a name that isn’t a syntactically valid R identifier, like the "1:3" name .table_tibble() generated automatically two sections ago:

names(.table_as_tibble(list(`1:3` = 1:3)))
[1] "1:3"
names(as.data.frame(list(`1:3` = 1:3)))
[1] "X1.3"

By default, check.names = TRUE in as.data.frame() rewrites "1:3" into the syntactically valid "X1.3"; .table_as_tibble() passes check.names = FALSE through instead, so the literal name survives.

Rowname to column conversion

A data frame with only the default sequential row names ("1", "2", …) is returned unchanged – there’s nothing meaningful to move into a column:

identical(
  .table_rownames_to_column(data.frame(x = 1:3), var = "id"),
  data.frame(x = 1:3)
)
[1] TRUE

Real (non-default) row names do get moved into a new column, placed first:

.table_rownames_to_column(head(mtcars, 3), var = "model")
          model  mpg cyl disp  hp drat    wt  qsec vs am gear carb
1     Mazda RX4 21.0   6  160 110 3.90 2.620 16.46  0  1    4    4
2 Mazda RX4 Wag 21.0   6  160 110 3.90 2.875 17.02  0  1    4    4
3    Datsun 710 22.8   4  108  93 3.85 2.320 18.61  1  1    4    1

Adding rows

Appending a row with .table_add_row() supports the same sequential evaluation as .table_tibble(), so a later value in the new row can refer to an earlier one: either a column already in .data, or one that is added in the same call.

df <- .table_tibble(a = 1:2, b = a * 10)
df
  a  b
1 1 10
2 2 20
.table_add_row(df, a = 3, b = a * 10)
  a  b
1 1 10
2 2 20
3 3 30

Here a * 10 in the new row refers to the a = 3 just supplied in that same call, not to the existing column a in df – the most recently constructed value for a given name always wins, which is what .table_drop_dup_list() is for internally.

Adding rows: column names and types must match

The new row’s columns must exactly match .data’s – no more, no fewer – and a mismatch is reported by name rather than surfacing rbind()’s own low-level error:

.table_add_row(df, a = 3)
Error in `.table_add_row()`:
! .table_add_row(): new row's columns must exactly match `.data`'s. Missing: b. 

Each value’s type must also match its column’s existing type (integer and double are treated as interchangeable, since mixing those two is unremarkable). Without this check, rbind() would silently coerce an entire existing column to a broader type just to accommodate one new, differently-typed value – e.g. one character value turning a whole numeric column into character – which is exactly the kind of silent surprise this mini avoids elsewhere too:

.table_add_row(df, a = "not a number", b = 999)
Error in `.table_add_row()`:
! .table_add_row(): column `a` is `integer` in `.data` but `character` in the new row; rbind() would silently coerce the whole column to accommodate it. Convert the new value to `integer` first if this is intentional.

Scope: data frame dropping still exists

Nothing in this mini changes the fact that the return objects are genuine data frames, which means that [ still behaves exactly the same as usual for a data frame, and subsetting is therefore subject to the data frame dropping rule. A future extension could in principle layer a very thin S3 class to address this, but that step is deliberately avoided here: the risk of introducing unexpected behaviour if a “quasi-tibble” classed object escapes into the wild seems too much to countenance in a mini.