Splitting
Splits-table generators — ports of Weka's CrossValidation /
LeaveOneOut / TestOnTrainingData / LearningCurve /
Holdout / HoldoutOrdered split producers.
Every generator returns a DataFrame with the Java ArffMapping schema:
type (TRAIN/TEST), rowid (original 0-based row), repeat,
fold, plus sample for learning-curve tasks only.
crossvalidation_splits(df, procedure, target=None, seed=1)
Generate cross-validation splits .
For each repeat:
- shuffle the dataset with the seed;
- stratify if the target is nominal;
- carve into
foldsslices and emit every instance asTRAINorTESTdepending on whether it belongs to fold f.
StratifiedKFold(shuffle=True) is the scikit-learn equivalent of Weka's
randomize + stratify + trainCV/testCV. We instantiate one
splitter per repeat with a distinct child seed so each repeat gets a fresh
permutation consecutively).
Source code in src/process_dataset/splitting.py
29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 | |
holdout_ordered_splits(df, procedure)
Generate holdout-ordered splits .
No shuffling — the first (100 - percentage)% of instances (in file order)
become TRAIN, the tail becomes TEST. Exactly one repeat and one fold.
Source code in src/process_dataset/splitting.py
210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 | |
holdout_splits(df, procedure, seed=1)
Generate holdout splits .
For each repeat, shuffle the data with the seed, then take the first
round(N * percentage / 100) instances as TEST, the rest as
TRAIN. sklearn.model_selection.ShuffleSplit does exactly this and
yields repeats independent shuffles in one call.
Source code in src/process_dataset/splitting.py
178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 | |
learning_curve_splits(df, procedure, target=None, seed=1)
Generate learning-curve splits.
Same CV skeleton as crossvalidation_splits, but inside each fold the training side is
subsampled at geometrically growing sizes 2 ** (6 + 0.5 * s) (capped at
the full train size), and each subsample is emitted under its own sample
index. The full test fold is repeated for every sample.
This is the only generator that emits a sample column.
Source code in src/process_dataset/splitting.py
124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 | |
leave_one_out_splits(df)
Generate leave-one-out splits .
For each fold f in 0..N-1, instance f is TEST and all others are
TRAIN. The output has N * N rows. sklearn.model_selection.LeaveOneOut
yields the train/test index pairs directly.
Source code in src/process_dataset/splitting.py
80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 | |
num_samples(train_size)
Number of subsamples for a training set of this size (incl. the full set).
Source code in src/process_dataset/splitting.py
116 117 118 119 120 121 | |
sample_size(number, train_size)
Return 2 ** (6 + 0.5 * number) capped at train_size.
Source code in src/process_dataset/splitting.py
111 112 113 | |
train_on_test_splits(df)
Generate test-on-training-data splits .
Every instance appears twice in fold 0 — once as TRAIN and once as
TEST. Trains and tests on the same data; used for in-sample evaluation.
Output size is 2 * N.
Source code in src/process_dataset/splitting.py
97 98 99 100 101 102 103 104 105 106 107 108 | |