午夜刷屏:睡前屏幕时间与睡眠负债
Python 3
🌙 午夜刷屏
睡前屏幕时间如何悄悄偷走你的睡眠 —— 一场完整的端到端分析
本 notebook 调查了横跨不同职业、时型(chronotype)与应用使用习惯的 8,500 名个体, 旨在回答一个问题:在睡前刷屏的时代,究竟是什么在驱动睡眠负债? 我们将从原始数据 → 清洗 → 探索性分析 → 统计检验 → 轻量级预测模型 → 可落地的洞察,逐步展开。
- 环境配置 & 数据加载
- 数据质量检查
- 描述性概览
- 单变量分析
- 分类变量深入分析
- 双变量 & 相关性分析
- 睡眠负债分群
- 统计假设检验
- 一个简单的预测模型
- 关键洞察 & 建议
1. 环境配置与数据加载
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
from scipy import stats
sns.set_style("whitegrid")
plt.rcParams["figure.figsize"] = (9, 5)
plt.rcParams["axes.titlesize"] = 14
plt.rcParams["axes.titleweight"] = "bold"
PALETTE = ["#5C7AEA", "#F7B32B", "#F45B69", "#3AAFA9", "#8E7DBE"]
sns.set_palette(PALETTE)
df = pd.read_csv("/kaggle/input/datasets/samartalwar/sleep-debt-and-screen-time-late-night-phone-habits/bedtime_screentime_sleep_debt.csv")
print("Rows, Columns:", df.shape)
df.head()
Rows, Columns: (8500, 18)
user_id age gender occupation_type chronotype \
0 USR-00001 29 Female Healthcare / Shift Worker Intermediate
1 USR-00002 58 Female Student Intermediate
2 USR-00003 41 Female Healthcare / Shift Worker Night Owl
3 USR-00004 36 Non-Binary Remote Tech Intermediate
4 USR-00005 23 Female Healthcare / Shift Worker Intermediate
bedtime_phone_minutes primary_bedtime_app screen_brightness_pct \
0 179 Instagram / Reddit 54
1 163 YouTube 76
2 100 TikTok / Reels 59
3 34 YouTube 52
4 27 TikTok / Reels 82
blue_light_filter_active caffeine_post_5pm_mg physical_activity_min \
0 1 0 21
1 1 92 72
2 0 45 23
3 0 97 49
4 1 0 38
sleep_latency_min total_sleep_hours deep_sleep_pct rem_sleep_pct \
0 84.4 3.20 23.5 22.2
1 86.7 4.06 18.0 21.5
2 70.5 3.20 18.1 20.5
3 38.1 6.99 21.6 22.4
4 32.3 4.58 24.1 22.8
morning_alarm_snoozes next_day_fatigue_score sleep_debt_category
0 7 10.0 Severe Sleep Debt
1 7 10.0 Severe Sleep Debt
2 7 10.0 Severe Sleep Debt
3 1 1.7 Mild Deficit
4 4 5.3 Moderate Debt
df.info()
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 8500 entries, 0 to 8499
Data columns (total 18 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 user_id 8500 non-null object
1 age 8500 non-null int64
2 gender 8500 non-null object
3 occupation_type 8500 non-null object
4 chronotype 8500 non-null object
5 bedtime_phone_minutes 8500 non-null int64
6 primary_bedtime_app 8500 non-null object
7 screen_brightness_pct 8500 non-null int64
8 blue_light_filter_active 8500 non-null int64
9 caffeine_post_5pm_mg 8500 non-null int64
10 physical_activity_min 8500 non-null int64
11 sleep_latency_min 8500 non-null float64
12 total_sleep_hours 8500 non-null float64
13 deep_sleep_pct 8500 non-null float64
14 rem_sleep_pct 8500 non-null float64
15 morning_alarm_snoozes 8500 non-null int64
16 next_day_fatigue_score 8500 non-null float64
17 sleep_debt_category 8500 non-null object
dtypes: float64(5), int64(7), object(6)
memory usage: 1.2+ MB
2. 数据质量检查
在信任任何洞察之前,我们先确认数据是干净的:没有缺失值、没有重复用户,且各取值范围合理。
# Missing values
print("Missing values per column:")
print(df.isnull().sum().sum(), "total missing cells")
# Duplicate users
print("Duplicate user_id rows:", df["user_id"].duplicated().sum())
# Sanity check on ranges
df.describe().T[["min", "max", "mean"]]
Missing values per column:
0 total missing cells
Duplicate user_id rows: 0
min max mean
age 18.0 65.0 34.355647
bedtime_phone_minutes 1.0 180.0 59.250000
screen_brightness_pct 10.0 100.0 55.153882
blue_light_filter_active 0.0 1.0 0.467765
caffeine_post_5pm_mg 0.0 250.0 33.222471
physical_activity_min 0.0 112.0 35.701412
sleep_latency_min 6.0 123.3 40.671471
total_sleep_hours 3.2 9.8 6.266209
deep_sleep_pct 8.1 28.0 21.735741
rem_sleep_pct 9.6 27.0 19.105859
morning_alarm_snoozes 0.0 7.0 2.766000
next_day_fatigue_score 1.0 10.0 3.794671
3. 描述性概览
对我们将要探索的数值型习惯与结果变量做一个快速的统计快照。
numeric_cols = ["age", "bedtime_phone_minutes", "screen_brightness_pct",
"caffeine_post_5pm_mg", "physical_activity_min", "sleep_latency_min",
"total_sleep_hours", "deep_sleep_pct", "rem_sleep_pct",
"morning_alarm_snoozes", "next_day_fatigue_score"]
df[numeric_cols].describe().T.style.background_gradient(cmap="Blues", subset=["mean", "std"])
<pandas.io.formats.style.Styler at 0x7a091442e0f0>
4. 单变量分析
我们的关键变量各自表现如何?
fig, axes = plt.subplots(2, 2, figsize=(13, 9))
sns.histplot(df["total_sleep_hours"], bins=30, kde=True, ax=axes[0, 0], color=PALETTE[0])
axes[0, 0].set_title("Distribution of Total Sleep Hours")
sns.histplot(df["bedtime_phone_minutes"], bins=30, kde=True, ax=axes[0, 1], color=PALETTE[2])
axes[0, 1].set_title("Distribution of Bedtime Phone Minutes")
sns.histplot(df["next_day_fatigue_score"], bins=20, kde=True, ax=axes[1, 0], color=PALETTE[1])
axes[1, 0].set_title("Distribution of Next-Day Fatigue Score")
sns.histplot(df["age"], bins=25, kde=True, ax=axes[1, 1], color=PALETTE[4])
axes[1, 1].set_title("Distribution of Age")
plt.tight_layout()
plt.show()

观察: 相当大一部分参与者的总睡眠时长集中在推荐的 7–9 小时区间之下,而睡前手机使用时长则呈现明显的右偏 —— 大多数人的刷屏时间还算适度,但存在一条长尾,有一部分人睡前刷屏远超一小时。
5. 分类变量深入分析
这些参与者都是谁?他们睡前都在刷什么?
fig, axes = plt.subplots(2, 2, figsize=(14, 10))
df["occupation_type"].value_counts().plot(kind="barh", ax=axes[0, 0], color=PALETTE[0])
axes[0, 0].set_title("Participants by Occupation")
df["chronotype"].value_counts().plot(kind="bar", ax=axes[0, 1], color=PALETTE[1])
axes[0, 1].set_title("Participants by Chronotype")
axes[0, 1].tick_params(axis="x", rotation=0)
df["primary_bedtime_app"].value_counts().plot(kind="barh", ax=axes[1, 0], color=PALETTE[2])
axes[1, 0].set_title("Primary Bedtime App")
df["sleep_debt_category"].value_counts().plot(kind="bar", ax=axes[1, 1], color=PALETTE[3])
axes[1, 1].set_title("Sleep Debt Category")
axes[1, 1].tick_params(axis="x", rotation=20)
plt.tight_layout()
plt.show()

观察: TikTok/Reels 和 YouTube 是睡前首选应用,占据主导地位。朝九晚五的公司职员是最大的职业群体,且大多数参与者落在 “Moderate Debt”(中度负债)而非极端类别 —— 这是一个普遍存在的日常问题,而非罕见问题。
6. 双变量与相关性分析
睡前屏幕时间更长,真的意味着睡眠更少吗?
plt.figure(figsize=(8, 6))
sns.scatterplot(data=df, x="bedtime_phone_minutes", y="total_sleep_hours",
hue="sleep_debt_category", palette=PALETTE, alpha=0.6, s=25)
plt.title("Bedtime Phone Minutes vs Total Sleep Hours")
plt.xlabel("Minutes on Phone Before Bed")
plt.ylabel("Total Sleep Hours")
plt.legend(title="Sleep Debt Category", bbox_to_anchor=(1.02, 1), loc="upper left")
plt.tight_layout()
plt.show()
/tmp/ipykernel_16/2951321385.py:2: UserWarning: The palette list has more values (5) than needed (4), which may not be intended.
sns.scatterplot(data=df, x="bedtime_phone_minutes", y="total_sleep_hours",

corr = df[numeric_cols].corr()
plt.figure(figsize=(10, 8))
sns.heatmap(corr, annot=True, fmt=".2f", cmap="coolwarm", center=0, linewidths=0.5)
plt.title("Correlation Heatmap of Numeric Variables")
plt.tight_layout()
plt.show()

corr["total_sleep_hours"].sort_values().drop("total_sleep_hours")
next_day_fatigue_score -0.879461
morning_alarm_snoozes -0.831289
bedtime_phone_minutes -0.554407
sleep_latency_min -0.527086
caffeine_post_5pm_mg -0.038262
screen_brightness_pct -0.027445
physical_activity_min 0.004350
age 0.016108
rem_sleep_pct 0.033031
deep_sleep_pct 0.177925
Name: total_sleep_hours, dtype: float64
观察: 次日疲劳评分和早晨闹钟贪睡次数与总睡眠时长呈最强的负相关 —— 这是一个符合直觉的合理性校验,表明数据的表现符合预期。睡前手机使用分钟数和入睡潜伏期也与睡眠时长呈负相关,支持了睡前屏幕时间会推迟并缩短睡眠的观点。
蓝光过滤真的有用吗?
plt.figure(figsize=(7, 5))
sns.boxplot(data=df, x="blue_light_filter_active", y="total_sleep_hours", palette=PALETTE)
plt.xticks([0, 1], ["Filter Off", "Filter On"])
plt.title("Total Sleep Hours: Blue Light Filter On vs Off")
plt.xlabel("")
plt.tight_layout()
plt.show()
/tmp/ipykernel_16/3356736453.py:2: FutureWarning:
Passing `palette` without assigning `hue` is deprecated and will be removed in v0.14.0. Assign the `x` variable to `hue` and set `legend=False` for the same effect.
sns.boxplot(data=df, x="blue_light_filter_active", y="total_sleep_hours", palette=PALETTE)
/tmp/ipykernel_16/3356736453.py:2: UserWarning: The palette list has more values (5) than needed (2), which may not be intended.
sns.boxplot(data=df, x="blue_light_filter_active", y="total_sleep_hours", palette=PALETTE)

7. 睡眠负债分群
不同睡眠负债类别之间的习惯有何差异?
order = ["Optimal Recovery", "Mild Deficit", "Moderate Debt", "Severe Sleep Debt"]
fig, axes = plt.subplots(1, 2, figsize=(14, 5))
sns.boxplot(data=df, x="sleep_debt_category", y="bedtime_phone_minutes", order=order, ax=axes[0], palette=PALETTE)
axes[0].set_title("Bedtime Phone Minutes by Sleep Debt Category")
axes[0].tick_params(axis="x", rotation=20)
sns.boxplot(data=df, x="sleep_debt_category", y="next_day_fatigue_score", order=order, ax=axes[1], palette=PALETTE)
axes[1].set_title("Next-Day Fatigue by Sleep Debt Category")
axes[1].tick_params(axis="x", rotation=20)
plt.tight_layout()
plt.show()
/tmp/ipykernel_16/2227326578.py:5: FutureWarning:
Passing `palette` without assigning `hue` is deprecated and will be removed in v0.14.0. Assign the `x` variable to `hue` and set `legend=False` for the same effect.
sns.boxplot(data=df, x="sleep_debt_category", y="bedtime_phone_minutes", order=order, ax=axes[0], palette=PALETTE)
/tmp/ipykernel_16/2227326578.py:5: UserWarning: The palette list has more values (5) than needed (4), which may not be intended.
sns.boxplot(data=df, x="sleep_debt_category", y="bedtime_phone_minutes", order=order, ax=axes[0], palette=PALETTE)
/tmp/ipykernel_16/2227326578.py:9: FutureWarning:
Passing `palette` without assigning `hue` is deprecated and will be removed in v0.14.0. Assign the `x` variable to `hue` and set `legend=False` for the same effect.
sns.boxplot(data=df, x="sleep_debt_category", y="next_day_fatigue_score", order=order, ax=axes[1], palette=PALETTE)
/tmp/ipykernel_16/2227326578.py:9: UserWarning: The palette list has more values (5) than needed (4), which may not be intended.
sns.boxplot(data=df, x="sleep_debt_category", y="next_day_fatigue_score", order=order, ax=axes[1], palette=PALETTE)

pd.crosstab(df["chronotype"], df["sleep_debt_category"], normalize="index").loc[:, order].style.background_gradient(cmap="Reds", axis=1)
<pandas.io.formats.style.Styler at 0x7a08ed4429c0>
观察: 与 “Optimal Recovery”(最佳恢复)组相比,“Severe Sleep Debt”(重度睡眠负债)组的睡前手机使用分钟数和次日疲劳程度明显更高。与 “Morning Larks”(早起型)和 “Intermediates”(中间型)相比,“Night Owls”(夜猫子型)更多地偏向负债类别,这暗示时型与屏幕使用习惯之间存在交互作用。
8. 统计假设检验
超越可视化 —— 这些差异在统计上显著吗?
检验 1 —— 独立样本 t 检验: 蓝光过滤是否会显著改变总睡眠时长?
group_on = df[df["blue_light_filter_active"] == 1]["total_sleep_hours"]
group_off = df[df["blue_light_filter_active"] == 0]["total_sleep_hours"]
t_stat, p_value = stats.ttest_ind(group_on, group_off)
print(f"Mean sleep (filter ON): {group_on.mean():.2f} hrs")
print(f"Mean sleep (filter OFF): {group_off.mean():.2f} hrs")
print(f"t-statistic = {t_stat:.3f}, p-value = {p_value:.4f}")
Mean sleep (filter ON): 6.33 hrs
Mean sleep (filter OFF): 6.21 hrs
t-statistic = 3.555, p-value = 0.0004
检验 2 —— 单因素方差分析(One-way ANOVA): 总睡眠时长在不同职业类型之间是否存在显著差异?
groups = [df[df["occupation_type"] == occ]["total_sleep_hours"] for occ in df["occupation_type"].unique()]
f_stat, p_value_anova = stats.f_oneway(*groups)
print(f"F-statistic = {f_stat:.3f}, p-value = {p_value_anova:.4f}")
F-statistic = 235.035, p-value = 0.0000
检验 3 —— 卡方检验(Chi-square test): 时型与睡眠负债类别是否相互独立?
contingency = pd.crosstab(df["chronotype"], df["sleep_debt_category"])
chi2, p_value_chi2, dof, expected = stats.chi2_contingency(contingency)
print(f"Chi-square = {chi2:.3f}, p-value = {p_value_chi2:.4f}, degrees of freedom = {dof}")
Chi-square = 2865.856, p-value = 0.0000, degrees of freedom = 6
9. 一个简单的预测模型
我们能用一个简单、可解释的模型,根据一个人的习惯来预测其睡眠负债类别吗?
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import LabelEncoder
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score, classification_report
# Simple label encoding for categorical features
model_df = df.copy()
cat_features = ["gender", "occupation_type", "chronotype", "primary_bedtime_app"]
encoders = {}
for col in cat_features:
le = LabelEncoder()
model_df[col] = le.fit_transform(model_df[col])
encoders[col] = le
target_encoder = LabelEncoder()
model_df["sleep_debt_category"] = target_encoder.fit_transform(model_df["sleep_debt_category"])
feature_cols = cat_features + ["age", "bedtime_phone_minutes", "screen_brightness_pct",
"blue_light_filter_active", "caffeine_post_5pm_mg",
"physical_activity_min", "sleep_latency_min"]
X = model_df[feature_cols]
y = model_df["sleep_debt_category"]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
model = RandomForestClassifier(n_estimators=200, max_depth=8, random_state=42)
model.fit(X_train, y_train)
y_pred = model.predict(X_test)
print("Accuracy:", round(accuracy_score(y_test, y_pred), 3))
print()
print(classification_report(y_test, y_pred, target_names=target_encoder.classes_))
Accuracy: 0.722
precision recall f1-score support
Mild Deficit 0.51 0.52 0.52 374
Moderate Debt 0.79 0.88 0.83 903
Optimal Recovery 0.72 0.58 0.64 279
Severe Sleep Debt 0.84 0.55 0.66 144
accuracy 0.72 1700
macro avg 0.72 0.63 0.66 1700
weighted avg 0.72 0.72 0.72 1700
importances = pd.Series(model.feature_importances_, index=feature_cols).sort_values()
plt.figure(figsize=(8, 6))
importances.plot(kind="barh", color=PALETTE[0])
plt.title("Feature Importance — Predicting Sleep Debt Category")
plt.xlabel("Importance")
plt.tight_layout()
plt.show()

观察: 该模型旨在作为一个简单、可解释的基线(baseline),而非可直接投入生产环境的预测器。特征重要性图告诉我们,模型在推测某人的睡眠负债类别时最依赖哪些习惯 —— 这可以作为一个指引,提示干预措施可能在哪些方面最为有效。
🔑 核心洞察
- 睡前屏幕时间与睡眠减少相关 —— 睡前手机使用分钟数越多,总睡眠时长越短,睡眠负债类别越高。
- 疲劳和贪睡是最明显的睡眠负债信号 —— 在该数据集中,次日疲劳评分和早晨闹钟贪睡次数是与总睡眠时长相关性最强的因素。
- 时型(Chronotype)很重要 —— 与早起鸟型(Morning Larks)相比,夜猫子型(Night Owls)在较高睡眠负债类别中的占比明显偏高。
- 短视频应用主导睡前屏幕使用 —— TikTok/Reels 和 YouTube 是最常见的睡前应用,超过即时通讯或阅读类应用。
- 一个简单的 Random Forest 基线模型仅凭习惯数据就能捕捉到有意义的信号,无需任何深度睡眠实验室数据——请查看上方的特征重要性图表,了解哪些习惯最为关键。
- 收集纵向数据(同一用户跨越多个夜晚),从相关性迈向因果性。
- 测试一项针对性干预措施(例如睡前屏幕使用时间截止提醒),并重新测量睡眠负债类别。
- 使用更大规模的多因素 ANOVA 更深入地探索时型与应用类型之间的交互效应。
感谢阅读——如果这个 notebook 对你有用,欢迎点赞!🌙📊
浙公网安备 33010602011771号