Skip to content

Data loader

Generic OpenML dataset loader.

Reads a downloaded dataset file (ARFF or Parquet) into a single normalized (attributes, rows) structure that the feature / quality extractors and the fold generators all consume. Picking the format happens once at the edge (CLI flag, notebook cell, or explicit caller) — everything downstream is format-agnostic.

The normalized form deliberately mirrors the liac-arff schema so the existing extractors work unchanged:

attributes list of (name, type_spec) pairs where type_spec is

  * ``"NUMERIC"``                       — numeric columns,
  * ``[label, ...]`` (a list)           — nominal columns (declared
                                          categories),
  * ``"STRING"`` / ``"DATE"``           — free-form / temporal columns.

rows list of row lists; missing values are None (matching ARFF's ?).

DataLoader

Load a downloaded OpenML dataset into a normalized (attributes, rows) form.

Parameters

fmt: "arff" (default) or "parquet". ARFF support is the incumbent format; Parquet is being phased in as the dataset storage backend.

Examples

loader = DataLoader("parquet") attributes, rows = loader.load(dataset_info)

Source code in src/data_loader.py
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
class DataLoader:
    """Load a downloaded OpenML dataset into a normalized ``(attributes, rows)`` form.

    Parameters
    ----------
    fmt:
        ``"arff"`` (default) or ``"parquet"``. ARFF support is the incumbent
        format; Parquet is being phased in as the dataset storage backend.

    Examples
    --------
    >>> loader = DataLoader("parquet")
    >>> attributes, rows = loader.load(dataset_info)
    """

    def __init__(self, fmt: DataFormat = "arff") -> None:
        fmt = fmt.lower()
        if fmt not in _SUPPORTED_FORMATS:
            raise ValueError(
                f"Unsupported data format: {fmt!r} "
                f"(expected one of {_SUPPORTED_FORMATS})."
            )
        self.fmt: DataFormat = fmt

    # ------------------------------------------------------------------
    # Public API
    # ------------------------------------------------------------------

    def load(
        self,
        dataset: DatasetDownloadInfo,
    ) -> tuple[list[tuple[str, object]], list[list]]:
        """Return ``(attributes, rows)`` for ``dataset.file_path``.

        ``attributes`` and ``rows`` follow the schema documented at the top of
        this module, identical for both backends.
        """
        if self.fmt == "arff":
            return self._load_arff(dataset.file_path)
        return self._load_parquet(dataset.file_path)

    # ------------------------------------------------------------------
    # Backends
    # ------------------------------------------------------------------

    @staticmethod
    def _load_arff(
        path: str,
    ) -> tuple[list[tuple[str, object]], list[list]]:
        with open(path, "r", encoding="utf-8", errors="replace") as f:
            payload = arff.load(f)
        return payload["attributes"], payload["data"]

    @staticmethod
    def _load_parquet(
        path: str,
    ) -> tuple[list[tuple[str, object]], list[list]]:
        df = pd.read_parquet(path)
        return _df_to_attributes_and_rows(df)

fmt = fmt instance-attribute

__init__(fmt='arff')

Source code in src/data_loader.py
54
55
56
57
58
59
60
61
def __init__(self, fmt: DataFormat = "arff") -> None:
    fmt = fmt.lower()
    if fmt not in _SUPPORTED_FORMATS:
        raise ValueError(
            f"Unsupported data format: {fmt!r} "
            f"(expected one of {_SUPPORTED_FORMATS})."
        )
    self.fmt: DataFormat = fmt

load(dataset)

Return (attributes, rows) for dataset.file_path.

attributes and rows follow the schema documented at the top of this module, identical for both backends.

Source code in src/data_loader.py
67
68
69
70
71
72
73
74
75
76
77
78
def load(
    self,
    dataset: DatasetDownloadInfo,
) -> tuple[list[tuple[str, object]], list[list]]:
    """Return ``(attributes, rows)`` for ``dataset.file_path``.

    ``attributes`` and ``rows`` follow the schema documented at the top of
    this module, identical for both backends.
    """
    if self.fmt == "arff":
        return self._load_arff(dataset.file_path)
    return self._load_parquet(dataset.file_path)