Bony1029

导航

午夜刷屏:睡前屏幕时间与睡眠负债

Python 3

数据与完整代码来源

🌙 午夜刷屏

睡前屏幕时间如何悄悄偷走你的睡眠 —— 一场完整的端到端分析

本 notebook 调查了横跨不同职业、时型(chronotype)与应用使用习惯的 8,500 名个体, 旨在回答一个问题:在睡前刷屏的时代,究竟是什么在驱动睡眠负债? 我们将从原始数据 → 清洗 → 探索性分析 → 统计检验 → 轻量级预测模型 → 可落地的洞察,逐步展开。

📋 Notebook 路线图
  1. 环境配置 & 数据加载
  2. 数据质量检查
  3. 描述性概览
  4. 单变量分析
  5. 分类变量深入分析
  6. 双变量 & 相关性分析
  7. 睡眠负债分群
  8. 统计假设检验
  9. 一个简单的预测模型
  10. 关键洞察 & 建议

1. 环境配置与数据加载

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
from scipy import stats

sns.set_style("whitegrid")
plt.rcParams["figure.figsize"] = (9, 5)
plt.rcParams["axes.titlesize"] = 14
plt.rcParams["axes.titleweight"] = "bold"

PALETTE = ["#5C7AEA", "#F7B32B", "#F45B69", "#3AAFA9", "#8E7DBE"]
sns.set_palette(PALETTE)
df = pd.read_csv("/kaggle/input/datasets/samartalwar/sleep-debt-and-screen-time-late-night-phone-habits/bedtime_screentime_sleep_debt.csv")
print("Rows, Columns:", df.shape)
df.head()
Rows, Columns: (8500, 18)
user_id  age      gender            occupation_type    chronotype  \
0  USR-00001   29      Female  Healthcare / Shift Worker  Intermediate   
1  USR-00002   58      Female                    Student  Intermediate   
2  USR-00003   41      Female  Healthcare / Shift Worker     Night Owl   
3  USR-00004   36  Non-Binary                Remote Tech  Intermediate   
4  USR-00005   23      Female  Healthcare / Shift Worker  Intermediate   

   bedtime_phone_minutes primary_bedtime_app  screen_brightness_pct  \
0                    179  Instagram / Reddit                     54   
1                    163             YouTube                     76   
2                    100      TikTok / Reels                     59   
3                     34             YouTube                     52   
4                     27      TikTok / Reels                     82   

   blue_light_filter_active  caffeine_post_5pm_mg  physical_activity_min  \
0                         1                     0                     21   
1                         1                    92                     72   
2                         0                    45                     23   
3                         0                    97                     49   
4                         1                     0                     38   

   sleep_latency_min  total_sleep_hours  deep_sleep_pct  rem_sleep_pct  \
0               84.4               3.20            23.5           22.2   
1               86.7               4.06            18.0           21.5   
2               70.5               3.20            18.1           20.5   
3               38.1               6.99            21.6           22.4   
4               32.3               4.58            24.1           22.8   

   morning_alarm_snoozes  next_day_fatigue_score sleep_debt_category  
0                      7                    10.0   Severe Sleep Debt  
1                      7                    10.0   Severe Sleep Debt  
2                      7                    10.0   Severe Sleep Debt  
3                      1                     1.7        Mild Deficit  
4                      4                     5.3       Moderate Debt
df.info()
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 8500 entries, 0 to 8499
Data columns (total 18 columns):
 #   Column                    Non-Null Count  Dtype  
---  ------                    --------------  -----  
 0   user_id                   8500 non-null   object 
 1   age                       8500 non-null   int64  
 2   gender                    8500 non-null   object 
 3   occupation_type           8500 non-null   object 
 4   chronotype                8500 non-null   object 
 5   bedtime_phone_minutes     8500 non-null   int64  
 6   primary_bedtime_app       8500 non-null   object 
 7   screen_brightness_pct     8500 non-null   int64  
 8   blue_light_filter_active  8500 non-null   int64  
 9   caffeine_post_5pm_mg      8500 non-null   int64  
 10  physical_activity_min     8500 non-null   int64  
 11  sleep_latency_min         8500 non-null   float64
 12  total_sleep_hours         8500 non-null   float64
 13  deep_sleep_pct            8500 non-null   float64
 14  rem_sleep_pct             8500 non-null   float64
 15  morning_alarm_snoozes     8500 non-null   int64  
 16  next_day_fatigue_score    8500 non-null   float64
 17  sleep_debt_category       8500 non-null   object 
dtypes: float64(5), int64(7), object(6)
memory usage: 1.2+ MB

2. 数据质量检查

在信任任何洞察之前,我们先确认数据是干净的:没有缺失值、没有重复用户,且各取值范围合理。

# Missing values
print("Missing values per column:")
print(df.isnull().sum().sum(), "total missing cells")

# Duplicate users
print("Duplicate user_id rows:", df["user_id"].duplicated().sum())

# Sanity check on ranges
df.describe().T[["min", "max", "mean"]]
Missing values per column:
0 total missing cells
Duplicate user_id rows: 0
min    max       mean
age                       18.0   65.0  34.355647
bedtime_phone_minutes      1.0  180.0  59.250000
screen_brightness_pct     10.0  100.0  55.153882
blue_light_filter_active   0.0    1.0   0.467765
caffeine_post_5pm_mg       0.0  250.0  33.222471
physical_activity_min      0.0  112.0  35.701412
sleep_latency_min          6.0  123.3  40.671471
total_sleep_hours          3.2    9.8   6.266209
deep_sleep_pct             8.1   28.0  21.735741
rem_sleep_pct              9.6   27.0  19.105859
morning_alarm_snoozes      0.0    7.0   2.766000
next_day_fatigue_score     1.0   10.0   3.794671
✅ 数据集干净 —— 没有缺失值、没有重复参与者,所有取值范围(例如睡眠时长 0–10 小时、屏幕亮度 0–100%)看起来都符合实际。无需任何清洗步骤;我们直接进入分析。

3. 描述性概览

对我们将要探索的数值型习惯与结果变量做一个快速的统计快照。

numeric_cols = ["age", "bedtime_phone_minutes", "screen_brightness_pct",
                "caffeine_post_5pm_mg", "physical_activity_min", "sleep_latency_min",
                "total_sleep_hours", "deep_sleep_pct", "rem_sleep_pct",
                "morning_alarm_snoozes", "next_day_fatigue_score"]

df[numeric_cols].describe().T.style.background_gradient(cmap="Blues", subset=["mean", "std"])
<pandas.io.formats.style.Styler at 0x7a091442e0f0>

4. 单变量分析

我们的关键变量各自表现如何?

fig, axes = plt.subplots(2, 2, figsize=(13, 9))

sns.histplot(df["total_sleep_hours"], bins=30, kde=True, ax=axes[0, 0], color=PALETTE[0])
axes[0, 0].set_title("Distribution of Total Sleep Hours")

sns.histplot(df["bedtime_phone_minutes"], bins=30, kde=True, ax=axes[0, 1], color=PALETTE[2])
axes[0, 1].set_title("Distribution of Bedtime Phone Minutes")

sns.histplot(df["next_day_fatigue_score"], bins=20, kde=True, ax=axes[1, 0], color=PALETTE[1])
axes[1, 0].set_title("Distribution of Next-Day Fatigue Score")

sns.histplot(df["age"], bins=25, kde=True, ax=axes[1, 1], color=PALETTE[4])
axes[1, 1].set_title("Distribution of Age")

plt.tight_layout()
plt.show()

图1

观察: 相当大一部分参与者的总睡眠时长集中在推荐的 7–9 小时区间之下,而睡前手机使用时长则呈现明显的右偏 —— 大多数人的刷屏时间还算适度,但存在一条长尾,有一部分人睡前刷屏远超一小时。

5. 分类变量深入分析

这些参与者都是谁?他们睡前都在刷什么?

fig, axes = plt.subplots(2, 2, figsize=(14, 10))

df["occupation_type"].value_counts().plot(kind="barh", ax=axes[0, 0], color=PALETTE[0])
axes[0, 0].set_title("Participants by Occupation")

df["chronotype"].value_counts().plot(kind="bar", ax=axes[0, 1], color=PALETTE[1])
axes[0, 1].set_title("Participants by Chronotype")
axes[0, 1].tick_params(axis="x", rotation=0)

df["primary_bedtime_app"].value_counts().plot(kind="barh", ax=axes[1, 0], color=PALETTE[2])
axes[1, 0].set_title("Primary Bedtime App")

df["sleep_debt_category"].value_counts().plot(kind="bar", ax=axes[1, 1], color=PALETTE[3])
axes[1, 1].set_title("Sleep Debt Category")
axes[1, 1].tick_params(axis="x", rotation=20)

plt.tight_layout()
plt.show()

图2

观察: TikTok/Reels 和 YouTube 是睡前首选应用,占据主导地位。朝九晚五的公司职员是最大的职业群体,且大多数参与者落在 “Moderate Debt”(中度负债)而非极端类别 —— 这是一个普遍存在的日常问题,而非罕见问题。

6. 双变量与相关性分析

睡前屏幕时间更长,真的意味着睡眠更少吗?

plt.figure(figsize=(8, 6))
sns.scatterplot(data=df, x="bedtime_phone_minutes", y="total_sleep_hours",
                 hue="sleep_debt_category", palette=PALETTE, alpha=0.6, s=25)
plt.title("Bedtime Phone Minutes vs Total Sleep Hours")
plt.xlabel("Minutes on Phone Before Bed")
plt.ylabel("Total Sleep Hours")
plt.legend(title="Sleep Debt Category", bbox_to_anchor=(1.02, 1), loc="upper left")
plt.tight_layout()
plt.show()
/tmp/ipykernel_16/2951321385.py:2: UserWarning: The palette list has more values (5) than needed (4), which may not be intended.
  sns.scatterplot(data=df, x="bedtime_phone_minutes", y="total_sleep_hours",

图3

corr = df[numeric_cols].corr()

plt.figure(figsize=(10, 8))
sns.heatmap(corr, annot=True, fmt=".2f", cmap="coolwarm", center=0, linewidths=0.5)
plt.title("Correlation Heatmap of Numeric Variables")
plt.tight_layout()
plt.show()

图4

corr["total_sleep_hours"].sort_values().drop("total_sleep_hours")
next_day_fatigue_score   -0.879461
morning_alarm_snoozes    -0.831289
bedtime_phone_minutes    -0.554407
sleep_latency_min        -0.527086
caffeine_post_5pm_mg     -0.038262
screen_brightness_pct    -0.027445
physical_activity_min     0.004350
age                       0.016108
rem_sleep_pct             0.033031
deep_sleep_pct            0.177925
Name: total_sleep_hours, dtype: float64

观察: 次日疲劳评分和早晨闹钟贪睡次数与总睡眠时长呈最强的负相关 —— 这是一个符合直觉的合理性校验,表明数据的表现符合预期。睡前手机使用分钟数和入睡潜伏期也与睡眠时长呈负相关,支持了睡前屏幕时间会推迟并缩短睡眠的观点。

蓝光过滤真的有用吗?

plt.figure(figsize=(7, 5))
sns.boxplot(data=df, x="blue_light_filter_active", y="total_sleep_hours", palette=PALETTE)
plt.xticks([0, 1], ["Filter Off", "Filter On"])
plt.title("Total Sleep Hours: Blue Light Filter On vs Off")
plt.xlabel("")
plt.tight_layout()
plt.show()
/tmp/ipykernel_16/3356736453.py:2: FutureWarning: 

Passing `palette` without assigning `hue` is deprecated and will be removed in v0.14.0. Assign the `x` variable to `hue` and set `legend=False` for the same effect.

  sns.boxplot(data=df, x="blue_light_filter_active", y="total_sleep_hours", palette=PALETTE)
/tmp/ipykernel_16/3356736453.py:2: UserWarning: The palette list has more values (5) than needed (2), which may not be intended.
  sns.boxplot(data=df, x="blue_light_filter_active", y="total_sleep_hours", palette=PALETTE)

图5

7. 睡眠负债分群

不同睡眠负债类别之间的习惯有何差异?

order = ["Optimal Recovery", "Mild Deficit", "Moderate Debt", "Severe Sleep Debt"]

fig, axes = plt.subplots(1, 2, figsize=(14, 5))

sns.boxplot(data=df, x="sleep_debt_category", y="bedtime_phone_minutes", order=order, ax=axes[0], palette=PALETTE)
axes[0].set_title("Bedtime Phone Minutes by Sleep Debt Category")
axes[0].tick_params(axis="x", rotation=20)

sns.boxplot(data=df, x="sleep_debt_category", y="next_day_fatigue_score", order=order, ax=axes[1], palette=PALETTE)
axes[1].set_title("Next-Day Fatigue by Sleep Debt Category")
axes[1].tick_params(axis="x", rotation=20)

plt.tight_layout()
plt.show()
/tmp/ipykernel_16/2227326578.py:5: FutureWarning: 

Passing `palette` without assigning `hue` is deprecated and will be removed in v0.14.0. Assign the `x` variable to `hue` and set `legend=False` for the same effect.

  sns.boxplot(data=df, x="sleep_debt_category", y="bedtime_phone_minutes", order=order, ax=axes[0], palette=PALETTE)
/tmp/ipykernel_16/2227326578.py:5: UserWarning: The palette list has more values (5) than needed (4), which may not be intended.
  sns.boxplot(data=df, x="sleep_debt_category", y="bedtime_phone_minutes", order=order, ax=axes[0], palette=PALETTE)
/tmp/ipykernel_16/2227326578.py:9: FutureWarning: 

Passing `palette` without assigning `hue` is deprecated and will be removed in v0.14.0. Assign the `x` variable to `hue` and set `legend=False` for the same effect.

  sns.boxplot(data=df, x="sleep_debt_category", y="next_day_fatigue_score", order=order, ax=axes[1], palette=PALETTE)
/tmp/ipykernel_16/2227326578.py:9: UserWarning: The palette list has more values (5) than needed (4), which may not be intended.
  sns.boxplot(data=df, x="sleep_debt_category", y="next_day_fatigue_score", order=order, ax=axes[1], palette=PALETTE)

图6

pd.crosstab(df["chronotype"], df["sleep_debt_category"], normalize="index").loc[:, order].style.background_gradient(cmap="Reds", axis=1)
<pandas.io.formats.style.Styler at 0x7a08ed4429c0>

观察: 与 “Optimal Recovery”(最佳恢复)组相比,“Severe Sleep Debt”(重度睡眠负债)组的睡前手机使用分钟数和次日疲劳程度明显更高。与 “Morning Larks”(早起型)和 “Intermediates”(中间型)相比,“Night Owls”(夜猫子型)更多地偏向负债类别,这暗示时型与屏幕使用习惯之间存在交互作用。

8. 统计假设检验

超越可视化 —— 这些差异在统计上显著吗?

检验 1 —— 独立样本 t 检验: 蓝光过滤是否会显著改变总睡眠时长?

group_on = df[df["blue_light_filter_active"] == 1]["total_sleep_hours"]
group_off = df[df["blue_light_filter_active"] == 0]["total_sleep_hours"]

t_stat, p_value = stats.ttest_ind(group_on, group_off)
print(f"Mean sleep (filter ON):  {group_on.mean():.2f} hrs")
print(f"Mean sleep (filter OFF): {group_off.mean():.2f} hrs")
print(f"t-statistic = {t_stat:.3f}, p-value = {p_value:.4f}")
Mean sleep (filter ON):  6.33 hrs
Mean sleep (filter OFF): 6.21 hrs
t-statistic = 3.555, p-value = 0.0004

检验 2 —— 单因素方差分析(One-way ANOVA): 总睡眠时长在不同职业类型之间是否存在显著差异?

groups = [df[df["occupation_type"] == occ]["total_sleep_hours"] for occ in df["occupation_type"].unique()]
f_stat, p_value_anova = stats.f_oneway(*groups)
print(f"F-statistic = {f_stat:.3f}, p-value = {p_value_anova:.4f}")
F-statistic = 235.035, p-value = 0.0000

检验 3 —— 卡方检验(Chi-square test): 时型与睡眠负债类别是否相互独立?

contingency = pd.crosstab(df["chronotype"], df["sleep_debt_category"])
chi2, p_value_chi2, dof, expected = stats.chi2_contingency(contingency)
print(f"Chi-square = {chi2:.3f}, p-value = {p_value_chi2:.4f}, degrees of freedom = {dof}")
Chi-square = 2865.856, p-value = 0.0000, degrees of freedom = 6
📊 解读检验结果: p 值低于 0.05 意味着观察到的差异不太可能出自随机偶然。请将上面打印出的每个 p 值与该阈值进行比对,以确认在本数据集中哪些效应具有统计显著性。

9. 一个简单的预测模型

我们能用一个简单、可解释的模型,根据一个人的习惯来预测其睡眠负债类别吗?

from sklearn.model_selection import train_test_split
from sklearn.preprocessing import LabelEncoder
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score, classification_report

# Simple label encoding for categorical features
model_df = df.copy()
cat_features = ["gender", "occupation_type", "chronotype", "primary_bedtime_app"]

encoders = {}
for col in cat_features:
    le = LabelEncoder()
    model_df[col] = le.fit_transform(model_df[col])
    encoders[col] = le

target_encoder = LabelEncoder()
model_df["sleep_debt_category"] = target_encoder.fit_transform(model_df["sleep_debt_category"])

feature_cols = cat_features + ["age", "bedtime_phone_minutes", "screen_brightness_pct",
                                "blue_light_filter_active", "caffeine_post_5pm_mg",
                                "physical_activity_min", "sleep_latency_min"]

X = model_df[feature_cols]
y = model_df["sleep_debt_category"]

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
model = RandomForestClassifier(n_estimators=200, max_depth=8, random_state=42)
model.fit(X_train, y_train)

y_pred = model.predict(X_test)
print("Accuracy:", round(accuracy_score(y_test, y_pred), 3))
print()
print(classification_report(y_test, y_pred, target_names=target_encoder.classes_))
Accuracy: 0.722

                   precision    recall  f1-score   support

     Mild Deficit       0.51      0.52      0.52       374
    Moderate Debt       0.79      0.88      0.83       903
 Optimal Recovery       0.72      0.58      0.64       279
Severe Sleep Debt       0.84      0.55      0.66       144

         accuracy                           0.72      1700
        macro avg       0.72      0.63      0.66      1700
     weighted avg       0.72      0.72      0.72      1700
importances = pd.Series(model.feature_importances_, index=feature_cols).sort_values()

plt.figure(figsize=(8, 6))
importances.plot(kind="barh", color=PALETTE[0])
plt.title("Feature Importance — Predicting Sleep Debt Category")
plt.xlabel("Importance")
plt.tight_layout()
plt.show()

图7

观察: 该模型旨在作为一个简单、可解释的基线(baseline),而非可直接投入生产环境的预测器。特征重要性图告诉我们,模型在推测某人的睡眠负债类别时最依赖哪些习惯 —— 这可以作为一个指引,提示干预措施可能在哪些方面最为有效。

🔑 核心洞察

  • 睡前屏幕时间与睡眠减少相关 —— 睡前手机使用分钟数越多,总睡眠时长越短,睡眠负债类别越高。
  • 疲劳和贪睡是最明显的睡眠负债信号 —— 在该数据集中,次日疲劳评分和早晨闹钟贪睡次数是与总睡眠时长相关性最强的因素。
  • 时型(Chronotype)很重要 —— 与早起鸟型(Morning Larks)相比,夜猫子型(Night Owls)在较高睡眠负债类别中的占比明显偏高。
  • 短视频应用主导睡前屏幕使用 —— TikTok/Reels 和 YouTube 是最常见的睡前应用,超过即时通讯或阅读类应用。
  • 一个简单的 Random Forest 基线模型仅凭习惯数据就能捕捉到有意义的信号,无需任何深度睡眠实验室数据——请查看上方的特征重要性图表,了解哪些习惯最为关键。
💡 建议的后续步骤
  • 收集纵向数据(同一用户跨越多个夜晚),从相关性迈向因果性。
  • 测试一项针对性干预措施(例如睡前屏幕使用时间截止提醒),并重新测量睡眠负债类别。
  • 使用更大规模的多因素 ANOVA 更深入地探索时型与应用类型之间的交互效应。

感谢阅读——如果这个 notebook 对你有用,欢迎点赞!🌙📊

数据与完整代码来源

posted on 2026-09-28 11:00  Bony-  阅读(3)  评论(0)    收藏  举报