JMI¶
skfeature.function.information_theoretical_based.jmi
Description¶
JMI (Joint Mutual Information) maximizes the joint information that each candidate feature and the already selected features carry about the class labels, balancing relevance and redundancy.
Note
This information-theoretic method requires discrete input features. Discretize continuous data first, for example with sklearn.preprocessing.KBinsDiscretizer.
Usage¶
import numpy as np
from sklearn.datasets import load_iris
from sklearn.feature_selection import SelectKBest
from sklearn.preprocessing import KBinsDiscretizer
from skfeature.function.information_theoretical_based import jmi
X, y = load_iris(return_X_y=True)
# information-theoretic scores require discrete features
X = KBinsDiscretizer(n_bins=5, encode="ordinal").fit_transform(X).astype(float)
# integrate with scikit-learn pipelines via SelectKBest
selector = SelectKBest(score_func=jmi.jmi, k=5)
X_selected = selector.fit_transform(X, y)
Parameters¶
mode:{{"rank", "index"}}, default"rank"—"rank"returns an array of feature indices ordered by importance and aligned withsklearn.feature_selection.SelectKBest;"index"returns the indices of the selected features with the most important one firstX:numpy array, shape(n_samples, n_features)— input data, must be discretey:numpy array, shape(n_samples,)— class labels**kwargs: additional parameters (seen_selected_featuresbelow)
Optional keyword arguments:
n_selected_features:int— number of features to select
Returns¶
score:numpy array, shape(n_features,)— ranking score of every feature, aligned withsklearn.feature_selection.SelectKBest
References¶
- Yang, Howard H. and Moody, John. "Data visualization and feature selection: New algorithms for nongaussian data." NIPS 1999.