Module
OpenML EvaluationEngine fold generators.
Each generator produces a splits table: one row per
(instance, repeat, fold, [sample]) assignment, telling OpenML whether that
original instance belongs to TRAIN or TEST for that combination.
The splits table schema mirrors the Java ArffMapping:
column meaning
------ -------------------------------------------------------
type 'TRAIN' or 'TEST'
rowid original 0-based row index in the dataset
repeat repeat index (0-based)
fold fold index (0-based)
sample subsample index (learning-curve tasks only)
FOLD_GENERATION_SEED = 0
module-attribute
generate_folds(did, procedure, seed=1, base_url=DEFAULT_API_BASE, *, data_format='arff')
Download dataset did and compute its splits table for an explicit
procedure. Low-level entry point; the CLI uses
:func:generate_folds_for_task, which mirrors Java's task-driven
GenerateFolds.
Source code in src/process_dataset/module.py
106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 | |
generate_folds_for_task(task_id, *, base_url=DEFAULT_API_BASE, seed=FOLD_GENERATION_SEED, data_format='arff')
Port of GenerateFolds.java — task_id is a TASK id (matching
Java's -f generate_folds -id <task_id>).
Downloads the task's source dataset, reads the estimation-procedure type
and number_folds / number_repeats / percentage from the task,
and generates the splits with seed (default FOLD_GENERATION_SEED,
matching Java's Main.FOLD_GENERATION_SEED). Multitask tasks (which Java
serves from a pre-merged dataset) are not handled here.
Source code in src/process_dataset/module.py
158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 | |
load_dataset(did, base_url=DEFAULT_API_BASE, *, data_format='arff')
Download an OpenML dataset by id and return (DataFrame, target).
Downloads the dataset (ARFF or Parquet, per data_format) and parses it
via :class:~src.data_loader.DataLoader into a DataFrame, preserving file
order so that row indices are stable rowid s. Nominal columns become
pd.Categorical with their declared categories.
Source code in src/process_dataset/module.py
49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 | |