Data loader
Generic OpenML dataset loader.
Reads a downloaded dataset file (ARFF or Parquet) into a single normalized
(attributes, rows) structure that the feature / quality extractors and the
fold generators all consume. Picking the format happens once at the edge (CLI
flag, notebook cell, or explicit caller) — everything downstream is
format-agnostic.
The normalized form deliberately mirrors the liac-arff schema so the
existing extractors work unchanged:
attributes
list of (name, type_spec) pairs where type_spec is
* ``"NUMERIC"`` — numeric columns,
* ``[label, ...]`` (a list) — nominal columns (declared
categories),
* ``"STRING"`` / ``"DATE"`` — free-form / temporal columns.
rows
list of row lists; missing values are None (matching ARFF's ?).
DataLoader
Load a downloaded OpenML dataset into a normalized (attributes, rows) form.
Parameters
fmt:
"arff" (default) or "parquet". ARFF support is the incumbent
format; Parquet is being phased in as the dataset storage backend.
Examples
loader = DataLoader("parquet") attributes, rows = loader.load(dataset_info)
Source code in src/data_loader.py
39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 | |
fmt = fmt
instance-attribute
__init__(fmt='arff')
Source code in src/data_loader.py
54 55 56 57 58 59 60 61 | |
load(dataset)
Return (attributes, rows) for dataset.file_path.
attributes and rows follow the schema documented at the top of
this module, identical for both backends.
Source code in src/data_loader.py
67 68 69 70 71 72 73 74 75 76 77 78 | |