set.seed(2984)
exposure <- c(rep(0, 12), round(rexp(48, rate = 1 / 20), 1))
head(exposure, 15) [1] 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 6.6 29.6 1.0
.cut_quantile()exclude.cut_evenly()startNAexclude and labeller work the same way for both functionsThe purpose of minicuts is to provide a zero-dependency reimplementation of cut_quantile()/cut_exposure_quantile() from erplots, generalised into .cut_quantile(), alongside a second function, .cut_evenly(), inspired by the chop_evenly()/chop_width() pair from santoku. Together they cover the two common ways of cutting a numeric vector into bins: fixed group size with data-dependent boundaries (.cut_quantile()), or fixed bin geometry with a data-dependent group size (.cut_evenly()). The README provides scope notes; this page walks through both functions on a small simulated exposure variable, including tie-breaking, custom labels, and the exclude argument the two functions share.
The minicuts.R script supplies two core functions:
.cut_quantile(x, n_bins = 4, ties = "upward", ...) is used to cut x into n_bins quantile bins, each holding (approximately) the same number of values. The remaining arguments control tie-break reproducibility (seed), the quantile algorithm (quantile_type), custom labels (labeller), and how excluded values are labelled (exclude/exclude_label)..cut_evenly(x, n_bins = NULL, width = NULL, ties = "upward", ...) is used to cut x into equal-width bins, either a fixed number of them spanning range(x) (n_bins) or a fixed width with however many bins the data needs (width). The remaining arguments control the bin anchor (start), custom labels (labeller), and how excluded values are labelled (exclude/exclude_label).They’re supported by five internal helpers, shared by both functions – there’s no need to call any of them directly. The helpers are these:
.cuts_resolve_exclude().cuts_bin_num().cuts_resolve_ties().cuts_with_seed().cuts_resolve_labels()A stand-in for a drug exposure variable from a clinical trial: some patients received a placebo (exposure exactly 0), the rest a range of doses.
.cut_quantile()By default, .cut_quantile() cuts x into four bins labelled "Q1"-"Q4", each holding a quarter of the values:
Because the placebo zeroes are exact ties clustered at the low end, they all land in Q1 here, and the lowest break point is pulled down to 0 as a result – something the exclude argument (below) exists to fix.
The labeller argument accepts either a plain character vector or a function of n_bins and the computed breaks:
Low Mid-low Mid-high High
15 15 15 15
Q1 (>0.0) Q2 (>1.2) Q3 (>5.9) Q4 (>18.6)
15 15 15 15
Repeated values sitting exactly on a quantile break point are common with skewed or rounded data. The ties argument controls which bin they land in:
[1] Q1 Q1 Q1 Q2 Q2 Q2 Q2 Q2 Q2 Q2
attr(,"breaks")
0% 50% 100%
1 4 7
attr(,"ties")
[1] upward
attr(,"quantile_type")
[1] 7
Levels: Q1 Q2
[1] Q1 Q1 Q1 Q1 Q1 Q1 Q1 Q2 Q2 Q2
attr(,"breaks")
0% 50% 100%
1 4 7
attr(,"ties")
[1] downward
attr(,"quantile_type")
[1] 7
Levels: Q1 Q2
The "split-even" option instead divides a tied group between its two candidate bins, aiming for equal bin sizes rather than sending every tied value the same way. It’s random, so it takes a seed to be reproducible:
excludePassing a predicate to exclude removes the matching values from the quantile calculation entirely, so they no longer distort the break points, but keeps them visible in the result under their own label rather than dropping them or coding them NA:
result
Placebo Q1 Q2 Q3 Q4
12 12 12 12 12
0% 25% 50% 75% 100%
0.300 4.175 9.850 23.550 53.200
The breaks are now computed purely from the 48 non-placebo doses, and the quantile bins split that group into four roughly equal quarters, with "Placebo" as a fifth, separate level.
.cut_evenly()Where .cut_quantile() fixes the group size and lets the bin edges follow from the data, .cut_evenly() does the opposite: it fixes the bin geometry and lets the group sizes follow from the data. Giving it n_bins spans range(x) with that many equal-width bins:
Q1 Q2 Q3 Q4
30 8 7 3
[1] 0.300 13.525 26.750 39.975 53.200
Giving it width instead fixes the bin width and works out how many bins are needed to cover the data, anchored at min(x) by default:
startAn explicit start moves that anchor. Here, bins are anchored at 0 rather than min(doses), so the bin edges land on round numbers:
A negative width builds bins downward from start (max(x) by default) instead of upward from it:
NAIf start is set explicitly and doesn’t reach one edge of range(x), some values fall outside every bin. .cut_evenly() codes those NA and warns, rather than silently widening the outermost bin to cover them:
exclude and labeller work the same way for both functionsThe placebo patients can be kept separate here too, and custom labels work identically to .cut_quantile():
There’s no ties = "split-even" option for .cut_evenly(), and consequently no seed argument either – there’s no “equal group size” goal to chase when the bins themselves are fixed by geometry rather than derived from data-dependent quantiles. Neither function supports santoku’s full generality: arbitrary breaks, non-numeric x such as dates, or weighted quantiles. If that generality is needed, reach for santoku itself.