🎵 为什么热门歌曲能持久(或消退)——发行渠道胜过歌曲本身
🎵 为什么热门歌曲能持久(或消退)——发行渠道胜过歌曲本身
数据与完整代码来源:
https://mbd.pub/o/bread/YZaVmJdxaQ
数据集简介
这份数据集聚焦于热门歌曲的生命周期,记录了一批曾进入流行行列的歌曲在不同发行渠道与时间维度上的表现。数据以单曲为基本单位,包含歌曲名称、演唱者、发行日期、发行渠道等信息,并配有随时间变化的流行度指标——例如榜单排名、播放量或热度得分——使我们能够观察一首歌从发布、走红、登顶到逐渐降温的完整轨迹。
从规模上看,数据覆盖了多个年份的大量歌曲,时间跨度足够长,可以比较不同时期歌曲的“寿命”;每首歌的流行度是多个观测点构成的时间序列,而非单一静态快照,这正是研究“持久与消退”问题的关键。
基于这些数据,可以探讨不少有趣的问题:哪些歌曲发行后能长盛不衰,哪些迅速退潮?发行渠道的选择是否会影响歌曲的生命周期?歌曲自身属性与渠道因素,究竟哪个对“持久度”的贡献更大?这些正是本文实战分析要回答的核心问题——接下来,我们就用具体数据来验证“发行渠道胜过歌曲本身”这一判断。
为什么有的曲目能持续播放多年,而制作更精良的曲目却在几周内销声匿迹?本 notebook 通过分析 45,000 个发行作品来回答这个问题——并最终得出一个令人不安的事实:一首歌被如何推广决定了它的命运;它听起来如何几乎无关紧要。
三个发现,逐一展示如下:
- 音频特征无法挑出热门歌曲。 仅使用 energy、valence、danceability 等特征的模型预测前 10% 曲目的 AUC ≈ 0.51——相当于抛硬币。发行渠道特征则达到 ≈ 0.84。
- 流媒体是“胜者多得”的游戏。 前 1% 的曲目占据了约 49% 的总播放量;后一半只是舍入误差。Gini ≈ 0.89。
- 早期势头会复利式累积。 首月播放量与 3 年总播放量的相关性约 0.88——这是一种马太效应,歌单与病毒式传播滚雪球般放大,且在很大程度上与音频本身无关。
完全合成的数据集,依据 Spotify/Luminate 统计数据校准(极端集中;中位半衰期约 59 天;歌单约占全部收听的一半)。参见 README。
import os, glob, warnings
import numpy as np, pandas as pd
import matplotlib.pyplot as plt, seaborn as sns
warnings.filterwarnings("ignore")
plt.style.use("seaborn-v0_8-whitegrid")
plt.rcParams.update({"figure.figsize":(11,6),"figure.dpi":120,"font.size":11,
"axes.titlesize":14,"axes.titleweight":"bold","axes.labelsize":11,"figure.titlesize":15})
C_DIST, C_AUDIO, C_HIT, C_FADE = "#4C72B0", "#DD8452", "#55A868", "#C44E52"
def find(name):
for p in [name, f"/mnt/user-data/outputs/song-longevity/{name}"]:
if os.path.exists(p): return p
h = glob.glob(f"/kaggle/input/**/{name}", recursive=True)
if h: return h[0]
raise FileNotFoundError(name)
df = pd.read_csv(find("song_longevity.csv"))
print("shape:", df.shape, "| median 3yr streams:", f"{df.streams_3yr.median():,.0f}")
df.head()
shape: (45000, 21) | median 3yr streams: 51,500
track_id genre energy valence danceability acousticness \
0 TRK000000 edm 5.95 2.39 9.71 2.08
1 TRK000001 edm 4.70 4.50 5.14 3.26
2 TRK000002 country 8.41 8.51 7.53 6.97
3 TRK000003 latin 1.66 5.50 5.06 3.58
4 TRK000004 hiphop 5.19 3.64 5.64 3.64
instrumentalness tempo_bpm song_length_sec label_tier ... \
0 2.59 144 167 independent ...
1 2.20 153 156 major ...
2 2.85 142 282 indie_label ...
3 0.29 102 240 major ...
4 4.24 132 224 independent ...
playlist_adds_first_month editorial_playlist tiktok_virality \
0 1 0 1.42
1 4 1 1.01
2 7 0 1.53
3 12 0 0.16
4 1 0 0.50
artist_prior_monthly_listeners featured_artist first_month_streams \
0 6700.0 0 15850
1 277000.0 0 159360
2 25800.0 0 1990
3 2400.0 0 92650
4 9800.0 0 5260
halflife_days streams_3yr slow_burner is_hit
0 59 53300 0 0
1 67 975600 0 1
2 102 11400 0 0
3 14 24600 0 0
4 57 26000 0 0
[5 rows x 21 columns]
1. 概述
print("Tracks:", f"{len(df):,}")
print("Median 3-year streams:", f"{df.streams_3yr.median():,.0f}")
print("Median half-life:", f"{df.halflife_days.median():.0f} days")
print("Slow-burners (half-life > 150d):", f"{df.slow_burner.mean()*100:.0f}%")
print("Placed on a big editorial playlist:", f"{df.editorial_playlist.mean()*100:.0f}%")
df[["energy","valence","danceability","playlist_adds_first_month","tiktok_virality","streams_3yr"]].describe().T.round(2)
Tracks: 45,000
Median 3-year streams: 51,500
Median half-life: 51 days
Slow-burners (half-life > 150d): 10%
Placed on a big editorial playlist: 13%
count mean std min 25% \
energy 45000.0 5.99 1.96 0.0 4.65
valence 45000.0 4.99 2.16 0.0 3.51
danceability 45000.0 5.97 1.95 0.0 4.64
playlist_adds_first_month 45000.0 4.63 4.37 0.0 2.00
tiktok_virality 45000.0 1.78 1.63 0.0 0.61
streams_3yr 45000.0 613049.93 5399919.33 100.0 13000.00
50% 75% max
energy 6.00 7.35 1.000000e+01
valence 4.99 6.49 1.000000e+01
danceability 5.99 7.32 1.000000e+01
playlist_adds_first_month 3.00 6.00 5.300000e+01
tiktok_virality 1.32 2.46 1.475000e+01
streams_3yr 51500.00 214400.00 6.130682e+08
2. 胜者多得——播放量分布
流媒体播放量不是钟形曲线,而是悬崖。极少数曲目几乎攫取了全部播放,而绝大多数曲目在统计意义上等同于寂静。洛伦兹曲线将这种不平等赤裸裸地呈现出来。
fig, ax = plt.subplots(1, 2, figsize=(15,5))
ax[0].hist(np.log10(df.streams_3yr.clip(lower=100)), bins=60, color=C_DIST, alpha=.85)
ax[0].set_title("3-year streams (log scale)"); ax[0].set_xlabel("log10(streams)"); ax[0].set_ylabel("Tracks")
ax[0].spines[["top","right"]].set_visible(False)
x = np.sort(df.streams_3yr.values); cum = np.cumsum(x)/x.sum(); p = np.arange(1,len(x)+1)/len(x)
def gini(v):
v=np.sort(v.astype(float)); n=len(v); c=np.cumsum(v); return (n+1-2*np.sum(c)/c[-1])/n
G = gini(df.streams_3yr.values)
ax[1].plot(p*100, cum*100, color=C_FADE, lw=2.5, label=f"Lorenz (Gini={G:.2f})")
ax[1].plot([0,100],[0,100],"k--",lw=1,label="equality")
ax[1].fill_between(p*100, cum*100, p*100, alpha=.12, color=C_FADE)
ax[1].set_title("Stream inequality across tracks")
ax[1].set_xlabel("% of tracks (fewest → most streamed)"); ax[1].set_ylabel("% of streams"); ax[1].legend()
ax[1].spines[["top","right"]].set_visible(False)
plt.tight_layout(); plt.show()
for q in [0.01,0.10]:
k=int(len(df)*q); print(f"top {int(q*100)}% of tracks hold {df.nlargest(k,'streams_3yr').streams_3yr.sum()/df.streams_3yr.sum()*100:.0f}% of streams")

top 1% of tracks hold 49% of streams
top 10% of tracks hold 84% of streams
3. 核心结论:发行渠道胜过歌曲本身
这是艺术家们不想听到的结果。用两种方式预测热门歌曲(播放量前 10%)——仅用音频特征,以及仅用发行渠道特征。音频相当于抛硬币,发行渠道则具有决定性。在解释总播放量时也呈现同样的模式。
from sklearn.ensemble import HistGradientBoostingClassifier, HistGradientBoostingRegressor
from sklearn.model_selection import train_test_split
from sklearn.metrics import roc_auc_score, roc_curve, r2_score
AUDIO = ["energy","valence","danceability","acousticness","instrumentalness","tempo_bpm","song_length_sec","genre"]
DIST = ["label_tier","marketing_budget","playlist_adds_first_month","editorial_playlist","tiktok_virality","artist_prior_monthly_listeners","featured_artist"]
def auc(cols, t):
X=pd.get_dummies(df[cols]); Xtr,Xte,ytr,yte=train_test_split(X,t,test_size=.25,stratify=t,random_state=42)
m=HistGradientBoostingClassifier(max_depth=6,max_iter=300,random_state=42).fit(Xtr,ytr)
p=m.predict_proba(Xte)[:,1]; return roc_auc_score(yte,p), roc_curve(yte,p)
a_a,roc_a = auc(AUDIO, df.is_hit); a_d,roc_d = auc(DIST, df.is_hit); a_de,roc_de = auc(DIST+["first_month_streams"], df.is_hit)
fig, ax = plt.subplots(1, 2, figsize=(15,5))
ax[0].bar(["AUDIO\nonly","DISTRIBUTION\nonly","+ early\nmomentum"], [a_a,a_d,a_de], color=[C_AUDIO,C_DIST,C_HIT])
for i,v in enumerate([a_a,a_d,a_de]): ax[0].text(i,v,f"{v:.3f}",ha="center",va="bottom",fontweight="bold",fontsize=12)
ax[0].axhline(0.5,color="k",ls="--",lw=1.5,label="coin flip"); ax[0].set_ylim(0.45,1.0)
ax[0].set_title("Predicting a hit (top-10% streams)"); ax[0].set_ylabel("ROC-AUC"); ax[0].legend()
ax[0].spines[["top","right"]].set_visible(False)
for (fpr,tpr,_),lab,c in [(roc_a,f"audio ({a_a:.3f})",C_AUDIO),(roc_d,f"distribution ({a_d:.3f})",C_DIST),(roc_de,f"+early ({a_de:.3f})",C_HIT)]:
ax[1].plot(fpr,tpr,lw=2.5,label=lab,color=c)
ax[1].plot([0,1],[0,1],"k--",lw=1); ax[1].set_title("ROC — audio can't pick a hit")
ax[1].set_xlabel("FPR"); ax[1].set_ylabel("TPR"); ax[1].legend()
ax[1].spines[["top","right"]].set_visible(False)
plt.tight_layout(); plt.show()
print(f"Hit prediction AUC — audio {a_a:.3f} (~coin flip) | distribution {a_d:.3f} | +early momentum {a_de:.3f}.")

Hit prediction AUC — audio 0.517 (~coin flip) | distribution 0.845 | +early momentum 0.968.
4. 音频几乎不相关;发行渠道则不然
放大到单个特征来看。energy、valence、danceability——艺术家和制作人反复打磨的那些旋钮——与 3 年播放量的相关性接近于零。真正承载信号的是歌单收录、TikTok 病毒式传播和既有听众。
ls = np.log1p(df.streams_3yr)
feats = {"energy":"energy","valence":"valence","danceability":"danceability","acousticness":"acousticness",
"playlist adds":"playlist_adds_first_month","editorial playlist":"editorial_playlist",
"TikTok virality":"tiktok_virality","marketing":"marketing_budget","prior audience":"artist_prior_monthly_listeners"}
cors = {}
for k,v in feats.items():
x = np.log1p(df[v]) if df[v].min()>=0 else df[v]
cors[k] = np.corrcoef(x, ls)[0,1]
s = pd.Series(cors).sort_values()
AUDIOK = ["energy","valence","danceability","acousticness"]
colors = [C_AUDIO if k in AUDIOK else C_DIST for k in s.index]
fig, ax = plt.subplots(figsize=(11,6))
ax.barh(s.index, s.values, color=colors)
for i,v in enumerate(s.values): ax.text(v,i,f" {v:+.2f}",va="center",fontweight="bold")
ax.axvline(0,color="k",lw=1)
ax.set_title("Correlation with 3-year streams (orange = audio, blue = distribution)")
ax.set_xlabel("correlation (log streams)"); ax.spines[["top","right"]].set_visible(False)
plt.tight_layout(); plt.show()
print("Audio features cluster at ~0; distribution features carry the signal.")

Audio features cluster at ~0; distribution features carry the signal.
5. 马太效应——早期势头滚雪球
这就是集中度背后的引擎。一首曲目的第一个月几乎写定了它的未来:首月播放量与 3 年总播放量的相关性约 0.88。早期命中的歌单与病毒式传播会复利式累积;其余一切只能挨饿。
fig, ax = plt.subplots(1, 2, figsize=(15,5))
samp = df.sample(6000, random_state=0)
ax[0].scatter(np.log10(samp.first_month_streams.clip(lower=50)), np.log10(samp.streams_3yr.clip(lower=100)),
s=8, alpha=.25, color=C_DIST)
r = np.corrcoef(np.log1p(df.first_month_streams), ls)[0,1]
ax[0].set_title(f"3-year vs first-month streams (r={r:.2f})")
ax[0].set_xlabel("log10 first-month streams"); ax[0].set_ylabel("log10 3-year streams")
ax[0].spines[["top","right"]].set_visible(False)
df["_fm"] = pd.qcut(df.first_month_streams, 6, labels=["1","2","3","4","5","6"])
g = df.groupby("_fm")["streams_3yr"].median()
ax[1].bar(range(6), g.values, color=plt.cm.viridis(np.linspace(0.2,0.85,6)))
ax[1].set_yscale("log"); ax[1].set_xticks(range(6)); ax[1].set_xticklabels(["low","","","","","high"])
ax[1].set_title("Median 3-year streams by first-month sextile")
ax[1].set_xlabel("First-month streams"); ax[1].set_ylabel("Median 3-year streams (log)")
ax[1].spines[["top","right"]].set_visible(False)
plt.tight_layout(); plt.show()
df.drop(columns="_fm", inplace=True)
print(f"First-month streams correlate {r:.2f} with the 3-year total — the early window largely decides the outcome.")

First-month streams correlate 0.88 with the 3-year total — the early window largely decides the outcome.
6. 半衰期——大多数热门歌曲迅速消退
借用化学中的概念:一首曲目的半衰期是指它保持在峰值一半以上所持续的时间。大多数曲目在几周内衰减;只有约 10% 的少数是能持续数月产生回报的慢热型。而造就慢热型的是投放位置,而非声音。
fig, ax = plt.subplots(1, 2, figsize=(15,5))
ax[0].hist(df.halflife_days.clip(0,400), bins=50, color=C_DIST, alpha=.85)
ax[0].axvline(df.halflife_days.median(), color=C_FADE, ls="--", lw=2, label=f"median {df.halflife_days.median():.0f}d")
ax[0].axvline(150, color=C_HIT, ls="--", lw=2, label="slow-burner (150d)")
ax[0].set_title("Half-life distribution — most fade fast")
ax[0].set_xlabel("Half-life (days)"); ax[0].set_ylabel("Tracks"); ax[0].legend()
ax[0].spines[["top","right"]].set_visible(False)
g = pd.Series({"on editorial\nplaylist": df[df.editorial_playlist==1].slow_burner.mean()*100,
"not on editorial": df[df.editorial_playlist==0].slow_burner.mean()*100,
"high audio\nquality": df[df.energy>=df.energy.quantile(.75)].slow_burner.mean()*100,
"low audio\nquality": df[df.energy<=df.energy.quantile(.25)].slow_burner.mean()*100})
ax[1].bar(g.index, g.values, color=[C_HIT,C_FADE,C_AUDIO,C_AUDIO])
for i,v in enumerate(g.values): ax[1].text(i,v,f"{v:.0f}%",ha="center",va="bottom",fontweight="bold")
ax[1].set_title("Slow-burner rate: placement matters, audio doesn't")
ax[1].set_ylabel("Slow-burner (%)")
ax[1].spines[["top","right"]].set_visible(False)
plt.tight_layout(); plt.show()
print(f"Editorial placement nearly doubles slow-burner odds ({g.iloc[0]:.0f}% vs {g.iloc[1]:.0f}%); "
"audio quality (energy) barely moves it.")

Editorial placement nearly doubles slow-burner odds (16% vs 9%); audio quality (energy) barely moves it.
7. 相关性矩阵
num = df.select_dtypes("number").copy()
num["log_streams"] = np.log1p(num["streams_3yr"])
keep = ["energy","valence","danceability","playlist_adds_first_month","editorial_playlist",
"tiktok_virality","marketing_budget","first_month_streams","halflife_days","log_streams"]
fig, ax = plt.subplots(figsize=(11,8))
sns.heatmap(num[keep].corr(), cmap="RdBu_r", center=0, square=True, annot=True, fmt=".2f",
annot_kws={"size":8}, cbar_kws={"shrink":.7,"label":"Pearson r"}, ax=ax)
ax.set_title("Correlation of key features"); plt.tight_layout(); plt.show()

8. 建模——驱动因素,以及音频的零结果
分别用仅音频、仅发行渠道、以及发行渠道 + 早期势头来预测 log 3 年播放量,然后对完整模型进行 permutation importance。仅音频模型几乎解释了 零 方差——这本身就是发现,而非 bug——长寿命存在于触达之中,而非波形之中。
from sklearn.inspection import permutation_importance
def r2(cols):
X=pd.get_dummies(df[cols]); Xtr,Xte,ytr,yte=train_test_split(X,ls,test_size=.25,random_state=42)
m=HistGradientBoostingRegressor(max_depth=6,max_iter=350,random_state=42).fit(Xtr,ytr)
return r2_score(yte,m.predict(Xte)), m, (Xtr,Xte,ytr,yte)
r_audio,_,_ = r2(AUDIO)
r_dist,_,_ = r2(DIST)
r_all,best,(Xtr,Xte,ytr,yte) = r2(DIST+AUDIO+["first_month_streams"])
fig, ax = plt.subplots(1, 2, figsize=(15,5))
ax[0].bar(["audio\nonly","distribution","distribution\n+ early"], [max(r_audio,0),r_dist,r_all],
color=[C_AUDIO,C_DIST,C_HIT])
for i,v in enumerate([max(r_audio,0),r_dist,r_all]): ax[0].text(i,v,f"{v:.2f}",ha="center",va="bottom",fontweight="bold",fontsize=13)
ax[0].set_title("Variance in 3-year streams explained (R²)"); ax[0].set_ylabel("R²"); ax[0].set_ylim(0,1)
ax[0].spines[["top","right"]].set_visible(False)
Xcols = pd.get_dummies(df[DIST+AUDIO+["first_month_streams"]]).columns
pi = permutation_importance(best, Xte, yte, n_repeats=5, random_state=42, n_jobs=-1, scoring="r2")
def grp(f):
for pre in ["label_tier","genre"]:
if f.startswith(pre): return pre
return f
imp = pd.Series(pi.importances_mean, index=Xcols).groupby(lambda f: grp(f)).sum().sort_values().tail(9)
cols=[C_AUDIO if any(a in f for a in ["energy","valence","dance","acoustic","instrument","tempo","length","genre"]) else C_DIST for f in imp.index]
ax[1].barh(imp.index, imp.values, color=cols)
ax[1].set_title("Permutation importance (orange = audio, blue = distribution)"); ax[1].set_xlabel("mean R² drop")
ax[1].spines[["top","right"]].set_visible(False)
plt.tight_layout(); plt.show()
print(f"log-streams R²: audio {max(r_audio,0):.2f} (~zero) | distribution {r_dist:.2f} | +early {r_all:.2f}.")
print("Audio features explain essentially nothing; reach and momentum explain the rest.")

log-streams R²: audio 0.00 (~zero) | distribution 0.38 | +early 0.80.
Audio features explain essentially nothing; reach and momentum explain the rest.
核心要点
- 音频特征无法挑出热门歌曲。 仅音频模型预测前 10% 曲目的 AUC 约 0.51(相当于抛硬币),且只能解释约 0% 的播放量方差。声音不是信号。
- 发行渠道具有决定性。 歌单位置、营销、TikTok 病毒式传播和既有听众以约 0.84 的 AUC 预测热门歌曲,并能解释相当一部分播放量。
- 流媒体是“胜者多得”的游戏。 Gini ≈ 0.89;前 1% 的曲目占据约 49% 的播放量。大多数发行作品在统计意义上等同于寂静。
- 早期势头会复利式累积(马太效应)。 首月播放量与 3 年总播放量的相关性约 0.88——这个早期窗口本身由发行渠道驱动,在很大程度上写定了结局。
- 慢热型是靠投放位置造就的,而非声音。 被编辑歌单收录的曲目能持久的可能性约 2 倍;音频质量几乎不影响胜算。
贯穿全文的主线: 浪漫化的叙事是伟大的歌曲凭实力脱颖而出。数据则表明,长寿命主要是一种发行渠道现象——歌单、时机、势头——而音频本身近乎噪声。这与学术论文引用或爆款帖子是同一个道理:质量只是入场券,触达才是复利所在。没错,要做出好东西——然后去赢得投放位置,因为那才是真正能预测三年后是否还有人听的部分。
下一步:直接对跳升-衰减轨迹建模,将自然病毒式传播与付费投放区分开,并构建一个发行前的“触达就绪度”评分。
完全合成;依据 Spotify/Luminate 统计数据校准。觉得有用吗?一个 upvote 会很有帮助。
数据与完整代码来源:
https://mbd.pub/o/bread/YZaVmJdxaQ
浙公网安备 33010602011771号