Skip to content

Splitting

Splits-table generators — ports of Weka's CrossValidation / LeaveOneOut / TestOnTrainingData / LearningCurve / Holdout / HoldoutOrdered split producers.

Every generator returns a DataFrame with the Java ArffMapping schema: type (TRAIN/TEST), rowid (original 0-based row), repeat, fold, plus sample for learning-curve tasks only.

crossvalidation_splits(df, procedure, target=None, seed=1)

Generate cross-validation splits .

For each repeat:

  1. shuffle the dataset with the seed;
  2. stratify if the target is nominal;
  3. carve into folds slices and emit every instance as TRAIN or TEST depending on whether it belongs to fold f.

StratifiedKFold(shuffle=True) is the scikit-learn equivalent of Weka's randomize + stratify + trainCV/testCV. We instantiate one splitter per repeat with a distinct child seed so each repeat gets a fresh permutation consecutively).

Source code in src/process_dataset/splitting.py
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
def crossvalidation_splits(
    df: pd.DataFrame,
    procedure: EstimationProcedure,
    target: Optional[str] = None,
    seed: int = 1,
) -> pd.DataFrame:
    """Generate cross-validation splits .

    For each repeat:

    1. shuffle the dataset with the seed;
    2. stratify if the target is nominal;
    3. carve into ``folds`` slices and emit every instance as ``TRAIN`` or
       ``TEST`` depending on whether it belongs to fold *f*.

    ``StratifiedKFold(shuffle=True)`` is the scikit-learn equivalent of Weka's
    ``randomize`` + ``stratify`` + ``trainCV``/``testCV``. We instantiate one
    splitter per repeat with a distinct child seed so each repeat gets a fresh
    permutation consecutively).
    """
    n = len(df)
    stratify = _is_nominal(df, target)
    y = df[target].to_numpy() if stratify else None
    index = np.arange(n)

    rows: list[tuple] = []
    for repeat in range(procedure.repeats):
        if stratify:
            splitter = StratifiedKFold(
                n_splits=procedure.folds,
                shuffle=True,
                random_state=seed + repeat,
            )
            splits = splitter.split(index, y)
        else:
            splitter = KFold(
                n_splits=procedure.folds,
                shuffle=True,
                random_state=seed + repeat,
            )
            splits = splitter.split(index)

        for fold, (train_idx, test_idx) in enumerate(splits):
            for i in train_idx:
                rows.append(("TRAIN", int(i), repeat, fold))
            for i in test_idx:
                rows.append(("TEST", int(i), repeat, fold))

    return pd.DataFrame(rows, columns=["type", "rowid", "repeat", "fold"])

holdout_ordered_splits(df, procedure)

Generate holdout-ordered splits .

No shuffling — the first (100 - percentage)% of instances (in file order) become TRAIN, the tail becomes TEST. Exactly one repeat and one fold.

Source code in src/process_dataset/splitting.py
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
def holdout_ordered_splits(
    df: pd.DataFrame,
    procedure: EstimationProcedure,
) -> pd.DataFrame:
    """Generate holdout-ordered splits .

    No shuffling — the first ``(100 - percentage)%`` of instances (in file order)
    become ``TRAIN``, the tail becomes ``TEST``. Exactly one repeat and one fold.
    """
    n = len(df)
    test_size = n * procedure.percentage / 100.0
    threshold = n - test_size  # rows at index <= threshold are TRAIN

    rows = [("TRAIN" if i <= threshold else "TEST", i, 0, 0) for i in range(n)]
    return pd.DataFrame(rows, columns=["type", "rowid", "repeat", "fold"])

holdout_splits(df, procedure, seed=1)

Generate holdout splits .

For each repeat, shuffle the data with the seed, then take the first round(N * percentage / 100) instances as TEST, the rest as TRAIN. sklearn.model_selection.ShuffleSplit does exactly this and yields repeats independent shuffles in one call.

Source code in src/process_dataset/splitting.py
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
def holdout_splits(
    df: pd.DataFrame,
    procedure: EstimationProcedure,
    seed: int = 1,
) -> pd.DataFrame:
    """Generate holdout splits .

    For each repeat, shuffle the data with the seed, then take the first
    ``round(N * percentage / 100)`` instances as ``TEST``, the rest as
    ``TRAIN``. ``sklearn.model_selection.ShuffleSplit`` does exactly this and
    yields ``repeats`` independent shuffles in one call.
    """
    n = len(df)
    test_size = procedure.percentage / 100.0
    ss = ShuffleSplit(
        n_splits=procedure.repeats,
        test_size=test_size,
        random_state=seed,
    )

    rows: list[tuple] = []
    # rowid is the ORIGINAL row index, so we split an index array, not df rows.
    index = np.arange(n)
    for repeat, (train_idx, test_idx) in enumerate(ss.split(index)):
        for i in train_idx:
            rows.append(("TRAIN", int(i), repeat, 0))
        for i in test_idx:
            rows.append(("TEST", int(i), repeat, 0))

    return pd.DataFrame(rows, columns=["type", "rowid", "repeat", "fold"])

learning_curve_splits(df, procedure, target=None, seed=1)

Generate learning-curve splits.

Same CV skeleton as crossvalidation_splits, but inside each fold the training side is subsampled at geometrically growing sizes 2 ** (6 + 0.5 * s) (capped at the full train size), and each subsample is emitted under its own sample index. The full test fold is repeated for every sample.

This is the only generator that emits a sample column.

Source code in src/process_dataset/splitting.py
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
def learning_curve_splits(
    df: pd.DataFrame,
    procedure: EstimationProcedure,
    target: Optional[str] = None,
    seed: int = 1,
) -> pd.DataFrame:
    """Generate learning-curve splits.


     Same CV skeleton as `crossvalidation_splits`, but inside each fold the *training* side is
    subsampled at geometrically growing sizes ``2 ** (6 + 0.5 * s)`` (capped at
    the full train size), and each subsample is emitted under its own ``sample``
    index. The full test fold is repeated for every sample.

    This is the only generator that emits a ``sample`` column.
    """
    n = len(df)
    stratify = _is_nominal(df, target)
    y = df[target].to_numpy() if stratify else None
    index = np.arange(n)

    rows: list[tuple] = []
    for repeat in range(procedure.repeats):
        if stratify:
            splitter = StratifiedKFold(
                n_splits=procedure.folds,
                shuffle=True,
                random_state=seed + repeat,
            )
            splits = splitter.split(index, y)
        else:
            splitter = KFold(
                n_splits=procedure.folds,
                shuffle=True,
                random_state=seed + repeat,
            )
            splits = splitter.split(index)

        for fold, (train_idx, test_idx) in enumerate(splits):
            train_size = len(train_idx)
            for s in range(num_samples(train_size)):
                k = sample_size(s, train_size)
                # first k (already-shuffled) training rows at this sample size
                for i in train_idx[:k]:
                    rows.append(("TRAIN", int(i), repeat, fold, s))
                for i in test_idx:
                    rows.append(("TEST", int(i), repeat, fold, s))

    return pd.DataFrame(
        rows,
        columns=["type", "rowid", "repeat", "fold", "sample"],
    )

leave_one_out_splits(df)

Generate leave-one-out splits .

For each fold f in 0..N-1, instance f is TEST and all others are TRAIN. The output has N * N rows. sklearn.model_selection.LeaveOneOut yields the train/test index pairs directly.

Source code in src/process_dataset/splitting.py
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
def leave_one_out_splits(df: pd.DataFrame) -> pd.DataFrame:
    """Generate leave-one-out splits .

    For each fold *f* in ``0..N-1``, instance *f* is ``TEST`` and all others are
    ``TRAIN``. The output has ``N * N`` rows. ``sklearn.model_selection.LeaveOneOut``
    yields the train/test index pairs directly.
    """
    index = np.arange(len(df))
    rows: list[tuple] = []
    for fold, (train_idx, test_idx) in enumerate(LeaveOneOut().split(index)):
        for i in train_idx:
            rows.append(("TRAIN", int(i), 0, fold))
        for i in test_idx:
            rows.append(("TEST", int(i), 0, fold))
    return pd.DataFrame(rows, columns=["type", "rowid", "repeat", "fold"])

num_samples(train_size)

Number of subsamples for a training set of this size (incl. the full set).

Source code in src/process_dataset/splitting.py
116
117
118
119
120
121
def num_samples(train_size: int) -> int:
    """Number of subsamples for a training set of this size (incl. the full set)."""
    i = 0
    while sample_size(i, train_size) < train_size:
        i += 1
    return i + 1

sample_size(number, train_size)

Return 2 ** (6 + 0.5 * number) capped at train_size.

Source code in src/process_dataset/splitting.py
111
112
113
def sample_size(number: int, train_size: int) -> int:
    """Return ``2 ** (6 + 0.5 * number)`` capped at ``train_size``."""
    return int(min(train_size, round(2 ** (6 + number * 0.5))))

train_on_test_splits(df)

Generate test-on-training-data splits .

Every instance appears twice in fold 0 — once as TRAIN and once as TEST. Trains and tests on the same data; used for in-sample evaluation. Output size is 2 * N.

Source code in src/process_dataset/splitting.py
 97
 98
 99
100
101
102
103
104
105
106
107
108
def train_on_test_splits(df: pd.DataFrame) -> pd.DataFrame:
    """Generate test-on-training-data splits .

    Every instance appears twice in fold 0 — once as ``TRAIN`` and once as
    ``TEST``. Trains and tests on the same data; used for in-sample evaluation.
    Output size is ``2 * N``.
    """
    rows: list[tuple] = []
    for i in range(len(df)):
        rows.append(("TEST", i, 0, 0))
        rows.append(("TRAIN", i, 0, 0))
    return pd.DataFrame(rows, columns=["type", "rowid", "repeat", "fold"])