书籍数据科学技术与应用_回归分析
Sklearn模块
- 无监督:cluster(聚类)、decomposition(因子分解)、mixture(高斯混合模型)、neural_network(无监督的神经网络)、covariance(协方差估计)
- 有监督:tree(决策树)、svm(支持向量机)、neighbors(近邻算法)、linear_model(广义线性模型)、neural_network(神经网络)、kernel_ridge(岭回归)、naive_bayes(朴素贝叶斯)
- 数据转换:feature_extraction(特征提取)、feature_selection(特征选择)、preprocessing(预处理)
回归分析
- 线性回归
- 逻辑回归
- 多项式回归
线性回归
- 使用 Linear_Regression 类构建
- 模型初始化
- linreg = Linear_Regression()
- 模型学习
- linreg. fit (x, y)
- 模型预测
- new_y = linreg. predict (new_x),其中 new_x = [..., ..., ...]
- model_selection 提供数据集的切分方法
- metrics 类实现各类机器学习算法的性能评估 训练集划分
- x_train, x_test, y_train, y_test = model_selection. train_test_split (x, y, test_size, random_size)
- 模型预测
- y_train_pred = linreg_tr. predict (x_train)
- y_test_pred = linreg_tr . predict (x_test)
- 均方根误差
- y_train_error = metrics. mean_squared_error (y_train, y_train_pred)
- y_test_error = metrics. mean_squared_error (y_test, y_test_pred)
- 决定系数
- r_square = linreg_tr. score (x_test, y_test)
线性回归例子
由广告收益历史数据,建立广告投入和销量的关系模型,并由下个月的广告投入,预测销量
# 读取数据
import pandas as pd
df = pd. read_csv ('advertising. csv', index_col = 0)
df. head()

# 绘制相关性散点图
# 销量与电视广告投入
import matplotlib. pyplot as plt
df. plot (kind = 'scatter', y = 'Sales', x = 'TV')
plt. title ('Sales with Advertising on TV')
plt. ylabel ('Sales')
plt. xlabel ('TV')

# 销量与微博广告投入
import matplotlib. pyplot as plt
df. plot (kind = 'scatter', y = 'Sales', x = 'Weibo')
plt. title ('Sales with Advertising on Weibo')
plt. ylabel ('Sales')
plt. xlabel ('Weibo')

# 销量与微信广告投入
df. plot (kind = 'scatter', y = 'Sales', x = 'WeChat')
plt. title ('Sales with Advertising on WeChat')
plt. ylabel ('Sales')
plt. xlabel ('WeChat')

# 建立回归模型,注意 new_x 传入形式为二维表
from sklearn. linear_model import LinearRegression
linreg = LinearRegression()
y = df ['Sales']
x = df [['TV', 'Weibo', 'WeChat']]
linreg. fit (x,y)
new_x = pd. DataFrame ([[130.1,87.8,69.2]])
new_y = linreg. predict (new_x)
intercept = linreg. intercept_
coef = linreg. coef_
print ('预测值:', new_y)
print ('截距项:', intercept)
print ('回归系数:', coef)

# 训练集划分
# 参数 test_size 表示划分百分之...作为测试集,参数 random_state = 1 表示每次得到的样本划分相同,否则不同
from sklearn import model_selection
x_train, x_test, y_train, y_test = model_selection. train_test_split (x, y, test_size = 0.35, random_state = 1)
# 在训练集上建模学习
# 因为在训练集上得到的性能指标不能反映在未知数据上的真实预测性能,这里的性能评估是指在测试集上
from sklearn import linear_model
linreg_tr = linear_model. LinearRegression()
linreg_tr .fit (x_train, y_train)
print ('训练集截距为:', linreg_tr. intercept_)
print ('训练集系数为:', linreg_tr. coef_)
![]()
# 性能评估
from sklearn import metrics
y_train_pred = linreg_tr. predict (x_train)
train_error = metrics. mean_squared_error (y_train, y_train_pred)
print ('训练集均方根误差为:', train_error)
y_test_pred = linreg_tr. predict (x_test)
test_error = metrics. mean_squared_error (y_test, y_test_pred)
print ('测试集均方根误差为:', test_error)
r_squared = linreg_tr. score (x_test, y_test)
print ('测试集决定系数为:', r_squared)

-END

浙公网安备 33010602011771号