Skip to content

Loader

Dataset → DataFeature extraction (format-agnostic).

The actual file parsing lives in :class:src.data_loader.DataLoader; this module owns the per-column statistical filling (Feature objects) and is unaware of whether the source was ARFF or Parquet.

load_features(dataset, *, data_format='arff', did=None, evaluation_engine_id=None)

Extract a :class:DataFeature from a downloaded dataset.

Parameters

dataset: Already-downloaded dataset (file_path points at an .arff or .parquet file depending on how it was fetched). data_format: "arff" (default) or "parquet" — selects the :class:~src.data_loader.DataLoader backend.

Source code in src/features/loader.py
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
def load_features(
    dataset: DatasetDownloadInfo,
    *,
    data_format: DataFormat = "arff",
    did: int | None = None,
    evaluation_engine_id: int | None = None,
) -> DataFeature:
    """Extract a :class:`DataFeature` from a downloaded dataset.

    Parameters
    ----------
    dataset:
        Already-downloaded dataset (``file_path`` points at an ``.arff`` or
        ``.parquet`` file depending on how it was fetched).
    data_format:
        ``"arff"`` (default) or ``"parquet"`` — selects the
        :class:`~src.data_loader.DataLoader` backend.
    """
    try:
        attributes, rows = DataLoader(data_format).load(dataset)
    except Exception as exc:
        return DataFeature(
            did=did,
            evaluation_engine_id=evaluation_engine_id,
            error=str(exc),
        )

    target_names = normalize_target_names(dataset.default_target_attribute)

    if rows:
        columns = list(zip(*rows))
    else:
        columns = [tuple() for _ in attributes]

    features: list[Feature] = []

    for idx, ((name, type_spec), col) in enumerate(zip(attributes, columns)):
        type_name, type_range = _liac_type(type_spec)

        feat = Feature(
            index=idx,
            name=name,
            data_type=type_name,
            is_target=name in target_names,
        )

        if type_name == "numeric":
            _fill_numeric_feature(col, feat)

        elif type_name == "nominal":
            _fill_nominal_feature(col, feat, type_range)

        else:
            feat.number_of_values = len(col)

        features.append(feat)

    return DataFeature(
        did=did,
        evaluation_engine_id=evaluation_engine_id,
        features=features,
    )