Skip to content

Module

OpenML EvaluationEngine fold generators.

Each generator produces a splits table: one row per (instance, repeat, fold, [sample]) assignment, telling OpenML whether that original instance belongs to TRAIN or TEST for that combination.

The splits table schema mirrors the Java ArffMapping:

column   meaning
------   -------------------------------------------------------
type     'TRAIN' or 'TEST'
rowid    original 0-based row index in the dataset
repeat   repeat index (0-based)
fold     fold index (0-based)
sample   subsample index (learning-curve tasks only)

FOLD_GENERATION_SEED = 0 module-attribute

generate_folds(did, procedure, seed=1, base_url=DEFAULT_API_BASE, *, data_format='arff')

Download dataset did and compute its splits table for an explicit procedure. Low-level entry point; the CLI uses :func:generate_folds_for_task, which mirrors Java's task-driven GenerateFolds.

Source code in src/process_dataset/module.py
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
def generate_folds(
    did: int,
    procedure: EstimationProcedure,
    seed: int = 1,
    base_url: str = DEFAULT_API_BASE,
    *,
    data_format: DataFormat = "arff",
) -> tuple[pd.DataFrame, pd.DataFrame, Optional[str]]:
    """Download dataset ``did`` and compute its splits table for an explicit
    ``procedure``. Low-level entry point; the CLI uses
    :func:`generate_folds_for_task`, which mirrors Java's task-driven
    ``GenerateFolds``.
    """
    df, target = load_dataset(did, base_url, data_format=data_format)
    return _splits_for_procedure(df, procedure, target=target, seed=seed), df, target

generate_folds_for_task(task_id, *, base_url=DEFAULT_API_BASE, seed=FOLD_GENERATION_SEED, data_format='arff')

Port of GenerateFolds.java — task_id is a TASK id (matching Java's -f generate_folds -id <task_id>).

Downloads the task's source dataset, reads the estimation-procedure type and number_folds / number_repeats / percentage from the task, and generates the splits with seed (default FOLD_GENERATION_SEED, matching Java's Main.FOLD_GENERATION_SEED). Multitask tasks (which Java serves from a pre-merged dataset) are not handled here.

Source code in src/process_dataset/module.py
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
def generate_folds_for_task(
    task_id: int,
    *,
    base_url: str = DEFAULT_API_BASE,
    seed: int = FOLD_GENERATION_SEED,
    data_format: DataFormat = "arff",
) -> tuple[pd.DataFrame, pd.DataFrame, Optional[str]]:
    """Port of ``GenerateFolds.java`` — ``task_id`` is a TASK id (matching
    Java's ``-f generate_folds -id <task_id>``).

    Downloads the task's source dataset, reads the estimation-procedure type
    and ``number_folds`` / ``number_repeats`` / ``percentage`` from the task,
    and generates the splits with ``seed`` (default ``FOLD_GENERATION_SEED``,
    matching Java's ``Main.FOLD_GENERATION_SEED``). Multitask tasks (which Java
    serves from a pre-merged dataset) are not handled here.
    """
    task_xml = get_task_xml(task_id, base_url)
    source_data = task_source_data(task_xml)
    did = int(source_data["oml:data_set_id"])
    target = source_data.get("oml:target_feature")

    ep = task_estimation_procedure(task_xml)
    if ep is None:
        raise ValueError(
            "Task has no estimation_procedure input; cannot generate folds."
        )
    procedure = _estimation_procedure_from_task(ep)

    df, _ = load_dataset(did, base_url, data_format=data_format)
    splits = _splits_for_procedure(df, procedure, target=target, seed=seed)
    return splits, df, target

load_dataset(did, base_url=DEFAULT_API_BASE, *, data_format='arff')

Download an OpenML dataset by id and return (DataFrame, target).

Downloads the dataset (ARFF or Parquet, per data_format) and parses it via :class:~src.data_loader.DataLoader into a DataFrame, preserving file order so that row indices are stable rowid s. Nominal columns become pd.Categorical with their declared categories.

Source code in src/process_dataset/module.py
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
def load_dataset(
    did: int,
    base_url: str = DEFAULT_API_BASE,
    *,
    data_format: DataFormat = "arff",
) -> tuple[pd.DataFrame, Optional[str]]:
    """Download an OpenML dataset by id and return ``(DataFrame, target)``.

    Downloads the dataset (ARFF or Parquet, per ``data_format``) and parses it
    via :class:`~src.data_loader.DataLoader` into a DataFrame, preserving file
    order so that row indices are stable ``rowid`` s. Nominal columns become
    ``pd.Categorical`` with their declared categories.
    """
    info: DatasetDownloadInfo = get_data_and_meta_information_from_did(
        did, dataset_type=data_format, base_url=base_url
    )

    attributes, rows = DataLoader(data_format).load(info)

    columns = [name for name, _ in attributes]
    df = pd.DataFrame(rows, columns=columns)

    for name, type_spec in attributes:
        if isinstance(type_spec, list):
            df[name] = pd.Categorical(df[name], categories=type_spec)

    return df, info.default_target_attribute