Skip to content

ARFF

Serialize splits DataFrames (see src.process_dataset.splitting) to the OpenML splits ARFF format.

arff_head(text, n=15)

First n lines of an ARFF string — handy for quick inspection.

Source code in src/process_dataset/arff.py
48
49
50
def arff_head(text: str, n: int = 15) -> str:
    """First ``n`` lines of an ARFF string — handy for quick inspection."""
    return "\n".join(text.splitlines()[:n])

save_splits_arff(splits, path, relation='splits')

Write a splits DataFrame to path as OpenML-format ARFF.

Source code in src/process_dataset/arff.py
37
38
39
40
41
42
43
44
45
def save_splits_arff(
    splits: pd.DataFrame,
    path: str,
    relation: str = "splits",
) -> None:
    """Write a splits DataFrame to ``path`` as OpenML-format ARFF."""
    text = splits_to_arff(splits, relation=relation)
    with open(path, "w", encoding="utf-8") as f:
        f.write(text)

splits_to_arff(splits, relation='splits')

Serialize a splits DataFrame to an OpenML-format ARFF string.

Java's ArffMapping emits a 4- or 5-column ARFF (type, rowid, repeat, fold, optional sample). We reproduce that schema with liac-arff so the output is byte-compatible with what the OpenML server accepts. type is nominal {TRAIN, TEST}; every other column is NUMERIC. Columns are emitted in declaration order regardless of DataFrame column order.

Source code in src/process_dataset/arff.py
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
def splits_to_arff(splits: pd.DataFrame, relation: str = "splits") -> str:
    """Serialize a splits DataFrame to an OpenML-format ARFF string.

    Java's ``ArffMapping`` emits a 4- or 5-column ARFF (``type``, ``rowid``,
    ``repeat``, ``fold``, optional ``sample``). We reproduce that schema with
    ``liac-arff`` so the output is byte-compatible with what the OpenML server
    accepts. ``type`` is nominal ``{TRAIN, TEST}``; every other column is
    ``NUMERIC``. Columns are emitted in declaration order regardless of
    DataFrame column order.
    """
    attributes: list[tuple[str, object]] = [("type", ["TRAIN", "TEST"])]
    for col in ("rowid", "repeat", "fold", "sample"):
        if col in splits.columns:
            attributes.append((col, "NUMERIC"))

    ordered = splits[[name for name, _ in attributes]]
    data = ordered.astype(object).to_numpy().tolist()

    return arff.dumps(
        {
            "relation": relation,
            "attributes": attributes,
            "data": data,
        }
    )