Kaggle-Intermediate Machine Learning(2)
Pipelines
大家都对Linux之中的管道技术不陌生,机器学习之中也有这种类似的技术。
管道是使数据预处理和建模代码井井有条的一种简单方法。 具体来说,管道捆绑了预处理和建模步骤,因此您可以像使用单个步骤一样使用整个捆绑。
许多数据科学家无需管道就可以一起破解模型,但是管道有一些重要的好处。 这些包括:
1.更干净的代码:在预处理的每个步骤中考虑数据可能会变得混乱。 使用管道,您无需在每个步骤中手动跟踪培训和验证数据。
2.错误更少:错误地使用步骤或忘记预处理步骤的机会更少。
3.易于实现生产:很难将模型从原型过渡到可大规模部署的模型。 我们在这里不会涉及许多相关问题,但是管道可以提供帮助。
4.模型验证的更多选项:在下一个教程中,您将看到一个示例,其中涉及交叉验证。
实例
我们可以使用前面的一个例子,我们不会专注于数据加载步骤。 相反,您可以想象您正处在X_train,X_valid,y_train和y_valid中具有训练和验证数据的位置。我们使用下面的head()方法来查看训练数据。 请注意,数据包含分类数据和缺少值的列。 使用管道,可以轻松处理这两个问题!
X_train.head()
Output:
Type | Method | Regionname | Rooms | Distance | Postcode | Bedroom2 | Bathroom | Car | Landsize | BuildingArea | YearBuilt | Lattitude | Longtitude | Propertycount | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 12167 | u | S | Southern Metropolitan | 1 | 5.0 | 3182.0 | 1.0 | 1.0 | 1.0 | 0.0 | NaN | 1940.0 | -37.85984 | 144.9867 | 13240.0 |
| 6524 | h | SA | Western Metropolitan | 2 | 8.0 | 3016.0 | 2.0 | 2.0 | 1.0 | 193.0 | NaN | NaN | -37.85800 | 144.9005 | 6380.0 |
| 8413 | h | S | Western Metropolitan | 3 | 12.6 | 3020.0 | 3.0 | 1.0 | 1.0 | 555.0 | NaN | NaN | -37.79880 | 144.8220 | 3755.0 |
| 2919 | u | SP | Northern Metropolitan | 3 | 13.0 | 3046.0 | 3.0 | 1.0 | 1.0 | 265.0 | NaN | 1995.0 | -37.70830 | 144.9158 | 8870.0 |
| 6043 | h | S | Western Metropolitan | 3 | 13.3 | 3020.0 | 3.0 | 1.0 | 2.0 | 673.0 | 673.0 | 1970.0 | -37.76230 | 144.8272 | 4217.0 |
我们分三步构建完整的管道。
步骤1:定义预处理步骤
类似于管道如何将预处理步骤和建模步骤捆绑在一起,我们使用ColumnTransformer类将不同的预处理步骤捆绑在一起。 下面的代码:
在数字数据中估算缺失值,以及插补缺失值,并对分类数据应用一键编码。
from sklearn.compose import ColumnTransformer from sklearn.pipeline import Pipeline from sklearn.impute import SimpleImputer from sklearn.preprocessing import OneHotEncoder # Preprocessing for numerical data numerical_transformer = SimpleImputer(strategy='constant') # Preprocessing for categorical data categorical_transformer = Pipeline(steps=[ ('imputer', SimpleImputer(strategy='most_frequent')), ('onehot', OneHotEncoder(handle_unknown='ignore')) ]) # Bundle preprocessing for numerical and categorical data preprocessor = ColumnTransformer( transformers=[ ('num', numerical_transformer, numerical_cols), ('cat', categorical_transformer, categorical_cols) ])
步骤2:定义模型
接下来,我们使用熟悉的RandomForestRegressor类定义一个随机森林模型。
from sklearn.ensemble import RandomForestRegressor model = RandomForestRegressor(n_estimators=100, random_state=0)
步骤3:创建和评估管道
最后,我们使用Pipeline类定义捆绑了预处理和建模步骤的管道。 有一些重要的注意事项:
通过管道,我们可以对训练数据进行预处理,并在一行代码中拟合模型。 (相比之下,没有管道,我们必须在单独的步骤中进行插补,单次编码和模型训练。如果我们必须同时处理数字变量和分类变量,这将变得特别混乱!)
通过管道,我们将X_valid中未处理的特征提供给predict()命令,并且管道在生成预测之前自动对特征进行预处理。 (但是,在没有管道的情况下,我们必须记住在进行预测之前对验证数据进行预处理。)
from sklearn.metrics import mean_absolute_error # Bundle preprocessing and modeling code in a pipeline my_pipeline = Pipeline(steps=[('preprocessor', preprocessor), ('model', model) ]) # Preprocessing of training data, fit model my_pipeline.fit(X_train, y_train) # Preprocessing of validation data, get predictions preds = my_pipeline.predict(X_valid) # Evaluate the model score = mean_absolute_error(y_valid, preds) print('MAE:', score)
Output:
MAE: 160679.18917034855
结论
管道对于清理机器学习代码和避免错误很有价值,对于具有复杂数据预处理功能的工作流特别有用
真是很方便,不用考虑繁琐的细节,而让人们可以更多的关注思想。
实践
首先得到数据,并提取其标签
import pandas as pd from sklearn.model_selection import train_test_split # Read the data X_full = pd.read_csv('../input/train.csv', index_col='Id') X_test_full = pd.read_csv('../input/test.csv', index_col='Id') # Remove rows with missing target, separate target from predictors X_full.dropna(axis=0, subset=['SalePrice'], inplace=True) y = X_full.SalePrice X_full.drop(['SalePrice'], axis=1, inplace=True) # Break off validation set from training data X_train_full, X_valid_full, y_train, y_valid = train_test_split(X_full, y, train_size=0.8, test_size=0.2, random_state=0) # "Cardinality" means the number of unique values in a column # Select categorical columns with relatively low cardinality (convenient but arbitrary) categorical_cols = [cname for cname in X_train_full.columns if X_train_full[cname].nunique() < 10 and X_train_full[cname].dtype == "object"] # Select numerical columns numerical_cols = [cname for cname in X_train_full.columns if X_train_full[cname].dtype in ['int64', 'float64']] # Keep selected columns only my_cols = categorical_cols + numerical_cols X_train = X_train_full[my_cols].copy() X_valid = X_valid_full[my_cols].copy() X_test = X_test_full[my_cols].copy()
INPUT:
X_train.head()
OUTPUT:
MSZoning Street Alley LotShape LandContour Utilities LotConfig LandSlope Condition1 Condition2 ... GarageArea WoodDeckSF OpenPorchSF EnclosedPorch 3SsnPorch ScreenPorch PoolArea MiscVal MoSold YrSold
Id
619 RL Pave NaN Reg Lvl AllPub Inside Gtl Norm Norm ... 774 0 108 0 0 260 0 0 7 2007
871 RL Pave NaN Reg Lvl AllPub Inside Gtl PosN Norm ... 308 0 0 0 0 0 0 0 8 2009
93 RL Pave Grvl IR1 HLS AllPub Inside Gtl Norm Norm ... 432 0 0 44 0 0 0 0 8 2009
818 RL Pave NaN IR1 Lvl AllPub CulDSac Gtl Norm Norm ... 857 150 59 0 0 0 0 0 7 2008
303 RL Pave NaN IR1 Lvl AllPub Corner Gtl Norm Norm ... 843 468 81 0 0 0 0 0 1 2006
下一个代码单元将使用教程中的代码来预处理数据并训练模型。 无需更改即可运行此代码。
from sklearn.compose import ColumnTransformer from sklearn.pipeline import Pipeline from sklearn.impute import SimpleImputer from sklearn.preprocessing import OneHotEncoder from sklearn.ensemble import RandomForestRegressor from sklearn.metrics import mean_absolute_error # Preprocessing for numerical data numerical_transformer = SimpleImputer(strategy='constant') # Preprocessing for categorical data categorical_transformer = Pipeline(steps=[ ('imputer', SimpleImputer(strategy='most_frequent')), ('onehot', OneHotEncoder(handle_unknown='ignore')) ]) # Bundle preprocessing for numerical and categorical data preprocessor = ColumnTransformer( transformers=[ ('num', numerical_transformer, numerical_cols), ('cat', categorical_transformer, categorical_cols) ]) # Define model model = RandomForestRegressor(n_estimators=100, random_state=0) # Bundle preprocessing and modeling code in a pipeline clf = Pipeline(steps=[('preprocessor', preprocessor), ('model', model) ]) # Preprocessing of training data, fit model clf.fit(X_train, y_train) # Preprocessing of validation data, get predictions preds = clf.predict(X_valid) print('MAE:', mean_absolute_error(y_valid, preds))
Output
MAE: 17861.780102739725
至此数据的预处理和模型的建立已经完毕
该代码产生的平均绝对误差(MAE)值约为17862。 在下一步中,您将修改代码以做得更好。
Step 1: Improve the performance
Part A
现在轮到你了! 在下面的代码单元中,定义您自己的预处理步骤和随机森林模型。 填写以下变量的值:
numerical_transformercategorical_transformermodel
# Preprocessing for numerical data numerical_transformer = SimpleImputer(strategy='constant') # Preprocessing for categorical data categorical_transformer = Pipeline(steps=[ ('imputer', SimpleImputer(strategy='constant')), ('onehot', OneHotEncoder(handle_unknown='ignore')) ]) # Bundle preprocessing for numerical and categorical data preprocessor = ColumnTransformer( transformers=[ ('num', numerical_transformer, numerical_cols), ('cat', categorical_transformer, categorical_cols) ]) # Define model model = RandomForestRegressor(n_estimators=100, random_state=0) # Check your answer step_1.a.check()
Part B
运行下面的代码单元,无需更改。
要通过此步骤,您需要在A部分中定义一个管道,以实现比上述代码更低的MAE。 我们鼓励您花时间在这里并尝试许多不同的方法,以了解获得MAE的最低价格! (_如果您的代码未通过,请修改A部分中的预处理步骤和模型。_)
# Bundle preprocessing and modeling code in a pipeline my_pipeline = Pipeline(steps=[('preprocessor', preprocessor), ('model', model) ]) # Preprocessing of training data, fit model my_pipeline.fit(X_train, y_train) # Preprocessing of validation data, get predictions preds = my_pipeline.predict(X_valid) # Evaluate the model score = mean_absolute_error(y_valid, preds) print('MAE:', score) # Check your answer step_1.b.check()
MAE: 17621.3197260274
Step 2: Generate test predictions
现在,您将使用训练有素的模型来生成带有测试数据的预测。
# Preprocessing of test data, fit model preds_test = my_pipeline.predict(X_test) # Your code here # Check your answer step_2.check()
运行下一个代码单元,而不进行任何更改,将结果保存到CSV文件中,该文件可以直接提交给比赛。
# Save test predictions to file output = pd.DataFrame({'Id': X_test.index, 'SalePrice': preds_test}) output.to_csv('submission.csv', index=False)

浙公网安备 33010602011771号