Kaggle-Intermediate Machine Learning(2)

Pipelines

大家都对Linux之中的管道技术不陌生,机器学习之中也有这种类似的技术。

管道是使数据预处理和建模代码井井有条的一种简单方法。 具体来说,管道捆绑了预处理和建模步骤,因此您可以像使用单个步骤一样使用整个捆绑。
许多数据科学家无需管道就可以一起破解模型,但是管道有一些重要的好处。 这些包括:

1.更干净的代码:在预处理的每个步骤中考虑数据可能会变得混乱。 使用管道,您无需在每个步骤中手动跟踪培训和验证数据。
2.错误更少:错误地使用步骤或忘记预处理步骤的机会更少。
3.易于实现生产:很难将模型从原型过渡到可大规模部署的模型。 我们在这里不会涉及许多相关问题,但是管道可以提供帮助。
4.模型验证的更多选项:在下一个教程中,您将看到一个示例,其中涉及交叉验证。

 

实例

我们可以使用前面的一个例子,我们不会专注于数据加载步骤。 相反,您可以想象您正处在X_train,X_valid,y_train和y_valid中具有训练和验证数据的位置。我们使用下面的head()方法来查看训练数据。 请注意,数据包含分类数据和缺少值的列。 使用管道,可以轻松处理这两个问题!

X_train.head()

Output:


Type
MethodRegionnameRoomsDistancePostcodeBedroom2BathroomCarLandsizeBuildingAreaYearBuiltLattitudeLongtitudePropertycount
12167 u S Southern Metropolitan 1 5.0 3182.0 1.0 1.0 1.0 0.0 NaN 1940.0 -37.85984 144.9867 13240.0
6524 h SA Western Metropolitan 2 8.0 3016.0 2.0 2.0 1.0 193.0 NaN NaN -37.85800 144.9005 6380.0
8413 h S Western Metropolitan 3 12.6 3020.0 3.0 1.0 1.0 555.0 NaN NaN -37.79880 144.8220 3755.0
2919 u SP Northern Metropolitan 3 13.0 3046.0 3.0 1.0 1.0 265.0 NaN 1995.0 -37.70830 144.9158 8870.0
6043 h S Western Metropolitan 3 13.3 3020.0 3.0 1.0 2.0 673.0 673.0 1970.0 -37.76230 144.8272 4217.0

我们分三步构建完整的管道。

步骤1:定义预处理步骤

类似于管道如何将预处理步骤和建模步骤捆绑在一起,我们使用ColumnTransformer类将不同的预处理步骤捆绑在一起。 下面的代码:
在数字数据中估算缺失值,以及插补缺失值,并对分类数据应用一键编码。

from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import OneHotEncoder

# Preprocessing for numerical data
numerical_transformer = SimpleImputer(strategy='constant')

# Preprocessing for categorical data
categorical_transformer = Pipeline(steps=[
    ('imputer', SimpleImputer(strategy='most_frequent')),
    ('onehot', OneHotEncoder(handle_unknown='ignore'))
])

# Bundle preprocessing for numerical and categorical data
preprocessor = ColumnTransformer(
    transformers=[
        ('num', numerical_transformer, numerical_cols),
        ('cat', categorical_transformer, categorical_cols)
    ])

步骤2:定义模型

接下来,我们使用熟悉的RandomForestRegressor类定义一个随机森林模型。

from sklearn.ensemble import RandomForestRegressor

model = RandomForestRegressor(n_estimators=100, random_state=0)

步骤3:创建和评估管道

最后,我们使用Pipeline类定义捆绑了预处理和建模步骤的管道。 有一些重要的注意事项:

通过管道,我们可以对训练数据进行预处理,并在一行代码中拟合模型。 (相比之下,没有管道,我们必须在单独的步骤中进行插补,单次编码和模型训练。如果我们必须同时处理数字变量和分类变量,这将变得特别混乱!)
通过管道,我们将X_valid中未处理的特征提供给predict()命令,并且管道在生成预测之前自动对特征进行预处理。 (但是,在没有管道的情况下,我们必须记住在进行预测之前对验证数据进行预处理。)

from sklearn.metrics import mean_absolute_error

# Bundle preprocessing and modeling code in a pipeline
my_pipeline = Pipeline(steps=[('preprocessor', preprocessor),
                              ('model', model)
                             ])

# Preprocessing of training data, fit model 
my_pipeline.fit(X_train, y_train)

# Preprocessing of validation data, get predictions
preds = my_pipeline.predict(X_valid)

# Evaluate the model
score = mean_absolute_error(y_valid, preds)
print('MAE:', score)

Output:

MAE: 160679.18917034855

结论

管道对于清理机器学习代码和避免错误很有价值,对于具有复杂数据预处理功能的工作流特别有用

真是很方便,不用考虑繁琐的细节,而让人们可以更多的关注思想。

 

实践

首先得到数据,并提取其标签

import pandas as pd
from sklearn.model_selection import train_test_split

# Read the data
X_full = pd.read_csv('../input/train.csv', index_col='Id')
X_test_full = pd.read_csv('../input/test.csv', index_col='Id')

# Remove rows with missing target, separate target from predictors
X_full.dropna(axis=0, subset=['SalePrice'], inplace=True)
y = X_full.SalePrice
X_full.drop(['SalePrice'], axis=1, inplace=True)

# Break off validation set from training data
X_train_full, X_valid_full, y_train, y_valid = train_test_split(X_full, y, 
                                                                train_size=0.8, test_size=0.2,
                                                                random_state=0)

# "Cardinality" means the number of unique values in a column
# Select categorical columns with relatively low cardinality (convenient but arbitrary)
categorical_cols = [cname for cname in X_train_full.columns if
                    X_train_full[cname].nunique() < 10 and 
                    X_train_full[cname].dtype == "object"]

# Select numerical columns
numerical_cols = [cname for cname in X_train_full.columns if 
                X_train_full[cname].dtype in ['int64', 'float64']]

# Keep selected columns only
my_cols = categorical_cols + numerical_cols
X_train = X_train_full[my_cols].copy()
X_valid = X_valid_full[my_cols].copy()
X_test = X_test_full[my_cols].copy()

INPUT:

X_train.head()

OUTPUT:

    MSZoning    Street    Alley    LotShape    LandContour    Utilities    LotConfig    LandSlope    Condition1    Condition2    ...    GarageArea    WoodDeckSF    OpenPorchSF    EnclosedPorch    3SsnPorch    ScreenPorch    PoolArea    MiscVal    MoSold    YrSold
Id                                                                                    
619    RL    Pave    NaN    Reg    Lvl    AllPub    Inside    Gtl    Norm    Norm    ...    774    0    108    0    0    260    0    0    7    2007
871    RL    Pave    NaN    Reg    Lvl    AllPub    Inside    Gtl    PosN    Norm    ...    308    0    0    0    0    0    0    0    8    2009
93    RL    Pave    Grvl    IR1    HLS    AllPub    Inside    Gtl    Norm    Norm    ...    432    0    0    44    0    0    0    0    8    2009
818    RL    Pave    NaN    IR1    Lvl    AllPub    CulDSac    Gtl    Norm    Norm    ...    857    150    59    0    0    0    0    0    7    2008
303    RL    Pave    NaN    IR1    Lvl    AllPub    Corner    Gtl    Norm    Norm    ...    843    468    81    0    0    0    0    0    1    2006

下一个代码单元将使用教程中的代码来预处理数据并训练模型。 无需更改即可运行此代码。

from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import OneHotEncoder
from sklearn.ensemble import RandomForestRegressor
from sklearn.metrics import mean_absolute_error

# Preprocessing for numerical data
numerical_transformer = SimpleImputer(strategy='constant')

# Preprocessing for categorical data
categorical_transformer = Pipeline(steps=[
    ('imputer', SimpleImputer(strategy='most_frequent')),
    ('onehot', OneHotEncoder(handle_unknown='ignore'))
])

# Bundle preprocessing for numerical and categorical data
preprocessor = ColumnTransformer(
    transformers=[
        ('num', numerical_transformer, numerical_cols),
        ('cat', categorical_transformer, categorical_cols)
    ])

# Define model
model = RandomForestRegressor(n_estimators=100, random_state=0)

# Bundle preprocessing and modeling code in a pipeline
clf = Pipeline(steps=[('preprocessor', preprocessor),
                      ('model', model)
                     ])

# Preprocessing of training data, fit model 
clf.fit(X_train, y_train)

# Preprocessing of validation data, get predictions
preds = clf.predict(X_valid)

print('MAE:', mean_absolute_error(y_valid, preds))

Output

MAE: 17861.780102739725

至此数据的预处理和模型的建立已经完毕

该代码产生的平均绝对误差(MAE)值约为17862。 在下一步中,您将修改代码以做得更好。

Step 1: Improve the performance

Part A

现在轮到你了! 在下面的代码单元中,定义您自己的预处理步骤和随机森林模型。 填写以下变量的值:

 

  • numerical_transformer
  • categorical_transformer
  • model
# Preprocessing for numerical data
numerical_transformer = SimpleImputer(strategy='constant')

# Preprocessing for categorical data
categorical_transformer = Pipeline(steps=[
    ('imputer', SimpleImputer(strategy='constant')),
    ('onehot', OneHotEncoder(handle_unknown='ignore'))
])

# Bundle preprocessing for numerical and categorical data
preprocessor = ColumnTransformer(
    transformers=[
        ('num', numerical_transformer, numerical_cols),
        ('cat', categorical_transformer, categorical_cols)
    ])

# Define model
model = RandomForestRegressor(n_estimators=100, random_state=0)

# Check your answer
step_1.a.check()

Part B

运行下面的代码单元,无需更改。

要通过此步骤,您需要在A部分中定义一个管道,以实现比上述代码更低的MAE。 我们鼓励您花时间在这里并尝试许多不同的方法,以了解获得MAE的最低价格! (_如果您的代码未通过,请修改A部分中的预处理步骤和模型。_)

# Bundle preprocessing and modeling code in a pipeline
my_pipeline = Pipeline(steps=[('preprocessor', preprocessor),
                              ('model', model)
                             ])

# Preprocessing of training data, fit model 
my_pipeline.fit(X_train, y_train)

# Preprocessing of validation data, get predictions
preds = my_pipeline.predict(X_valid)

# Evaluate the model
score = mean_absolute_error(y_valid, preds)
print('MAE:', score)

# Check your answer
step_1.b.check()
MAE: 17621.3197260274

Step 2: Generate test predictions

现在,您将使用训练有素的模型来生成带有测试数据的预测。

# Preprocessing of test data, fit model
preds_test = my_pipeline.predict(X_test) # Your code here

# Check your answer
step_2.check()

运行下一个代码单元,而不进行任何更改,将结果保存到CSV文件中,该文件可以直接提交给比赛。

# Save test predictions to file
output = pd.DataFrame({'Id': X_test.index,
                       'SalePrice': preds_test})
output.to_csv('submission.csv', index=False)

 

 

posted @ 2020-08-29 10:43  caishunzhe  阅读(190)  评论(0)    收藏  举报