Landmarkers
sklearn port of Weka's GenericLandmarker characterizers.
Each landmarker runs N-fold cross-validation of a classifier on the dataset
and reports three meta-features — {name}AUC, {name}ErrRate,
{name}Kappa — mirroring
org.openml.webapplication.fantail.dc.landmarking.GenericLandmarker.
Faithfulness gaps (Weka classifiers without sklearn equivalents):
* J48.* — Weka's C4.5 with confidence-based pruning (-C). sklearn's
DecisionTreeClassifier is CART with no equivalent pruning flag, so all
three J48 variants below use the same plain tree and emit identical
values. The Java metric IDs are preserved so the server schema matches.
* REPTree* / RandomTree* — approximated via DecisionTreeClassifier
with the matching max_depth. Different split logic, so values diverge
from Java.
* CfsSubsetEval_* — SKIPPED. CFS (Correlation-based Feature
Selection) is a Weka-specific subset evaluator with no sklearn equivalent;
faking one would misrepresent the meta-feature. ALL_LANDMARKERS omits
the three CFS entries that Java's CharacterizerFactory.all() includes.
The CV harness itself is faithful: Weka's
Evaluation.crossValidateModel(cls, data, 2, new Random(1)) randomizes,
stratifies (nominal target), and accumulates predictions across folds before
computing weightedAreaUnderROC() / errorRate() / kappa(). Here,
StratifiedKFold(2, shuffle=True, random_state=1) + pooled
cross_val_predict plays the same role. Numeric targets return all-None
(Java parity via UnassignedClassException).
ALL_LANDMARKERS = [('kNN1N', lambda: KNeighborsClassifier(n_neighbors=1)), ('NaiveBayes', lambda: GaussianNB()), ('DecisionStump', lambda: DecisionTreeClassifier(max_depth=1)), ('J48.001.', lambda: DecisionTreeClassifier()), ('J48.0001.', lambda: DecisionTreeClassifier()), ('J48.00001.', lambda: DecisionTreeClassifier()), ('REPTreeDepth1', lambda: DecisionTreeClassifier(max_depth=1)), ('REPTreeDepth2', lambda: DecisionTreeClassifier(max_depth=2)), ('REPTreeDepth3', lambda: DecisionTreeClassifier(max_depth=3)), ('RandomTreeDepth1', lambda: DecisionTreeClassifier(max_depth=1)), ('RandomTreeDepth2', lambda: DecisionTreeClassifier(max_depth=2)), ('RandomTreeDepth3', lambda: DecisionTreeClassifier(max_depth=3))]
module-attribute
compute_all_landmarkers(X, y)
Run every landmarker in ALL_LANDMARKERS and merge results into one
dict. X must be a clean numeric matrix; the caller (ExtractFeatures)
is responsible for encoding via the same path pymfe uses.
Source code in src/qualities/landmarkers.py
141 142 143 144 145 146 147 148 149 150 151 | |
expected_landmarker_ids()
All metric IDs the ALL set emits — port of
CharacterizerFactory.getExpectedQualities(all(null)) (minus CFS).
Source code in src/qualities/landmarkers.py
135 136 137 138 | |
generic_landmarker(name, estimator, X, y)
Compute one landmarker's three meta-features. Returns {name}AUC,
{name}ErrRate, {name}Kappa — matching GenericLandmarker.getIDs.
Source code in src/qualities/landmarkers.py
100 101 102 103 104 105 106 107 108 109 110 111 112 113 | |