Skip to content

Landmarkers

sklearn port of Weka's GenericLandmarker characterizers.

Each landmarker runs N-fold cross-validation of a classifier on the dataset and reports three meta-features — {name}AUC, {name}ErrRate, {name}Kappa — mirroring org.openml.webapplication.fantail.dc.landmarking.GenericLandmarker.

Faithfulness gaps (Weka classifiers without sklearn equivalents): * J48.* — Weka's C4.5 with confidence-based pruning (-C). sklearn's DecisionTreeClassifier is CART with no equivalent pruning flag, so all three J48 variants below use the same plain tree and emit identical values. The Java metric IDs are preserved so the server schema matches. * REPTree* / RandomTree* — approximated via DecisionTreeClassifier with the matching max_depth. Different split logic, so values diverge from Java. * CfsSubsetEval_* — SKIPPED. CFS (Correlation-based Feature Selection) is a Weka-specific subset evaluator with no sklearn equivalent; faking one would misrepresent the meta-feature. ALL_LANDMARKERS omits the three CFS entries that Java's CharacterizerFactory.all() includes.

The CV harness itself is faithful: Weka's Evaluation.crossValidateModel(cls, data, 2, new Random(1)) randomizes, stratifies (nominal target), and accumulates predictions across folds before computing weightedAreaUnderROC() / errorRate() / kappa(). Here, StratifiedKFold(2, shuffle=True, random_state=1) + pooled cross_val_predict plays the same role. Numeric targets return all-None (Java parity via UnassignedClassException).

ALL_LANDMARKERS = [('kNN1N', lambda: KNeighborsClassifier(n_neighbors=1)), ('NaiveBayes', lambda: GaussianNB()), ('DecisionStump', lambda: DecisionTreeClassifier(max_depth=1)), ('J48.001.', lambda: DecisionTreeClassifier()), ('J48.0001.', lambda: DecisionTreeClassifier()), ('J48.00001.', lambda: DecisionTreeClassifier()), ('REPTreeDepth1', lambda: DecisionTreeClassifier(max_depth=1)), ('REPTreeDepth2', lambda: DecisionTreeClassifier(max_depth=2)), ('REPTreeDepth3', lambda: DecisionTreeClassifier(max_depth=3)), ('RandomTreeDepth1', lambda: DecisionTreeClassifier(max_depth=1)), ('RandomTreeDepth2', lambda: DecisionTreeClassifier(max_depth=2)), ('RandomTreeDepth3', lambda: DecisionTreeClassifier(max_depth=3))] module-attribute

compute_all_landmarkers(X, y)

Run every landmarker in ALL_LANDMARKERS and merge results into one dict. X must be a clean numeric matrix; the caller (ExtractFeatures) is responsible for encoding via the same path pymfe uses.

Source code in src/qualities/landmarkers.py
141
142
143
144
145
146
147
148
149
150
151
def compute_all_landmarkers(
    X: np.ndarray,
    y: np.ndarray,
) -> dict[str, Optional[float]]:
    """Run every landmarker in ``ALL_LANDMARKERS`` and merge results into one
    dict. X must be a clean numeric matrix; the caller (``ExtractFeatures``)
    is responsible for encoding via the same path pymfe uses."""
    result: dict[str, Optional[float]] = {}
    for name, factory in ALL_LANDMARKERS:
        result.update(generic_landmarker(name, factory(), X, y))
    return result

expected_landmarker_ids()

All metric IDs the ALL set emits — port of CharacterizerFactory.getExpectedQualities(all(null)) (minus CFS).

Source code in src/qualities/landmarkers.py
135
136
137
138
def expected_landmarker_ids() -> list[str]:
    """All metric IDs the ALL set emits — port of
    ``CharacterizerFactory.getExpectedQualities(all(null))`` (minus CFS)."""
    return [f"{name}{suffix}" for name, _ in ALL_LANDMARKERS for suffix in _MEASURES]

generic_landmarker(name, estimator, X, y)

Compute one landmarker's three meta-features. Returns {name}AUC, {name}ErrRate, {name}Kappa — matching GenericLandmarker.getIDs.

Source code in src/qualities/landmarkers.py
100
101
102
103
104
105
106
107
108
109
110
111
112
113
def generic_landmarker(
    name: str,
    estimator,
    X: np.ndarray,
    y: np.ndarray,
) -> dict[str, Optional[float]]:
    """Compute one landmarker's three meta-features. Returns ``{name}AUC``,
    ``{name}ErrRate``, ``{name}Kappa`` — matching ``GenericLandmarker.getIDs``."""
    auc, err, kappa = _cross_validated_metrics(estimator, X, y)
    return {
        f"{name}AUC": auc,
        f"{name}ErrRate": err,
        f"{name}Kappa": kappa,
    }