IPL的19个赛季 对1243 场比赛的完整端到端分析
Python 3
INDIAN PREMIER LEAGUE | 2008 – 2026
IPL 的 19 个赛季
对 1,243 场比赛的完整端到端分析
数据清洗 • 特征工程 • 探索性分析 • 统计检验 • 洞察结论
本笔记本做了什么
IPL 是全球收视率最高的特许经营板球联赛。本笔记本取 从 2008 年揭幕战到 2026 年决赛的每一场 IPL 比赛 的原始比赛级记录,并像分析师在实际工作中那样,对其进行端到端的处理:
| 步骤 | 具体内容 |
|---|---|
| 1. 加载与检查 | 形状、数据类型、缺失值,以及对原始文件的初步查看 |
| 2. 清洗 | 修正赛季标签、合并更名后的球队、场地去重、修复城市信息 |
| 3. 特征工程 | 推导先攻方、追分目标、获胜分差以及赛事阶段 |
| 4. 探索 | 15+ 张关于时代、球队、掷币、场地、分差与球员的可视化图表 |
| 5. 检验 | 用卡方检验和 t 检验判断这些模式究竟是真实规律还是噪声 |
| 6. 结论 | 用平实的语言总结数据到底说了什么 |
我们要回答的问题:
- IPL 真的已经变成一个得分更高的联赛了吗,还是只是我们的错觉?
- 哪支球队最成功——按冠军数衡量,还是按胜率衡量?
- 赢得掷币真的能帮你赢下比赛吗?
- 追分真的比守分更容易吗?这一点随时间推移发生了变化吗?
- 哪些球场是击球天堂,哪些又是追分陷阱?
- 决定最多比赛胜负的球员都有谁?
1 · 环境设置
库、绘图主题以及贯穿全程使用的配色方案
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
from scipy import stats
import warnings
warnings.filterwarnings("ignore")
pd.set_option("display.max_columns", 50)
pd.set_option("display.width", 200)
配色方案只定义一次,然后在每张图表中复用。统一的配色是让笔记本看起来像一件完整的作品、而不是二十张互不相干的图表的最省力的办法。
NAVY = "#12263F"
BLUE = "#2E86AB"
ORANGE = "#FF6B35"
GOLD = "#F4A261"
GREEN = "#2A9D8F"
RED = "#E63946"
GREY = "#8D99AE"
SEQ = [NAVY, BLUE, ORANGE, GREEN, GOLD, RED, GREY]
plt.rcParams["figure.figsize"] = (11, 5.5)
plt.rcParams["figure.facecolor"] = "white"
plt.rcParams["axes.facecolor"] = "white"
plt.rcParams["axes.edgecolor"] = "#D9DEE5"
plt.rcParams["axes.titlesize"] = 14
plt.rcParams["axes.titleweight"] = "bold"
plt.rcParams["axes.titlecolor"] = NAVY
plt.rcParams["axes.labelcolor"] = NAVY
plt.rcParams["axes.grid"] = True
plt.rcParams["grid.color"] = "#EDF0F4"
plt.rcParams["xtick.color"] = "#5A6B7D"
plt.rcParams["ytick.color"] = "#5A6B7D"
plt.rcParams["font.size"] = 11
sns.set_palette(SEQ)
2 · 加载数据并初步查看
在看清自己手里到底是什么之前,永远不要动手清洗
df = pd.read_csv("/kaggle/input/datasets/arjunsinghgangwar/ipl-20082026-matches-dataset/IPL_Matches_Data_2008_2026.csv")
print("Rows :", df.shape[0])
print("Columns :", df.shape[1])
df.head(3)
Rows : 1243
Columns : 31
event_name season match_number date city venue team1 team2 toss_winner \
0 Indian Premier League 2007/08 1.0 18-04-2008 Bangalore M Chinnaswamy Stadium Royal Challengers Bangalore Kolkata Knight Riders Royal Challengers Bangalore
1 Indian Premier League 2007/08 2.0 19-04-2008 Chandigarh Punjab Cricket Association Stadium, Mohali Kings XI Punjab Chennai Super Kings Chennai Super Kings
2 Indian Premier League 2007/08 3.0 19-04-2008 Delhi Feroz Shah Kotla Delhi Daredevils Rajasthan Royals Rajasthan Royals
toss_decision team1_runs team1_wickets team2_runs team2_wickets winner result_type win_by_runs win_by_wickets player_of_match match_referee umpire1 umpire2 \
0 field 82 10 222 3 Kolkata Knight Riders complete 140 0 BB McCullum J Srinath Asad Rauf RE Koertzen
1 bat 207 4 240 5 Chennai Super Kings complete 33 0 MEK Hussey S Venkataraghavan MR Benson SL Shastri
2 bat 132 1 129 8 Delhi Daredevils complete 0 9 MF Maharoof GR Viswanath Aleem Dar GA Pratapkumar
tv_umpire reserve_umpire match_type overs_limit balls_per_over gender team_type team1_players team2_players
0 AM Saheba VN Kulkarni T20 20 6 male club R Dravid, W Jaffer, V Kohli, JH Kallis, CL Whi... SC Ganguly, BB McCullum, RT Ponting, DJ Hussey...
1 RB Tiffin MSS Ranawat T20 20 6 male club K Goel, JR Hopes, KC Sangakkara, Yuvraj Singh,... PA Patel, ML Hayden, MEK Hussey, MS Dhoni, SK ...
2 IL Howell Unknown T20 20 6 male club G Gambhir, V Sehwag, S Dhawan, MK Tiwary, KD K... T Kohli, YK Pathan, SR Watson, M Kaif, DS Lehm...
df.info()
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 1243 entries, 0 to 1242
Data columns (total 31 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 event_name 1243 non-null object
1 season 1243 non-null object
2 match_number 1169 non-null float64
3 date 1243 non-null object
4 city 1243 non-null object
5 venue 1243 non-null object
6 team1 1243 non-null object
7 team2 1243 non-null object
8 toss_winner 1243 non-null object
9 toss_decision 1243 non-null object
10 team1_runs 1243 non-null int64
11 team1_wickets 1243 non-null int64
12 team2_runs 1243 non-null int64
13 team2_wickets 1243 non-null int64
14 winner 1218 non-null object
15 result_type 1243 non-null object
16 win_by_runs 1243 non-null int64
17 win_by_wickets 1243 non-null int64
18 player_of_match 1234 non-null object
19 match_referee 1243 non-null object
20 umpire1 1243 non-null object
21 umpire2 1243 non-null object
22 tv_umpire 1243 non-null object
23 reserve_umpire 1243 non-null object
24 match_type 1243 non-null object
25 overs_limit 1243 non-null int64
26 balls_per_over 1243 non-null int64
27 gender 1243 non-null object
28 team_type 1243 non-null object
29 team1_players 1243 non-null object
30 team2_players 1243 non-null object
dtypes: float64(1), int64(8), object(22)
memory usage: 301.2+ KB
缺失值
只有四列存在缺失,而每一列的缺失最终都被证明是有意义的,而非损坏的数据。
missing = pd.DataFrame({
"missing": df.isna().sum(),
"percent": (df.isna().sum() / len(df) * 100).round(2)
})
missing = missing[missing["missing"] > 0].sort_values("missing", ascending=False)
missing
missing percent
match_number 74 5.95
winner 25 2.01
player_of_match 9 0.72
# Why are they missing? Let's look at the two important ones.
print("Result types where 'winner' is missing:")
print(df.loc[df["winner"].isna(), "result_type"].value_counts())
print()
print("Seasons where 'match_number' is missing (matches per season):")
print(df.loc[df["match_number"].isna(), "season"].value_counts().head())
Result types where 'winner' is missing:
result_type
tie 16
no result 9
Name: count, dtype: int64
Seasons where 'match_number' is missing (matches per season):
season
2009/10 4
2012 4
2011 4
2019 4
2020/21 4
Name: count, dtype: int64
• winner — 在 16 场 平局(由 super over 决出胜负,而本文件并未记录)和 9 场 无结果(因雨)的比赛中缺失。这些是真正未分出胜负的比赛,而不是错误。
• match_number — 每个赛季恰好有 3–4 场比赛缺失。那些正是 季后赛:资格赛、淘汰赛和决赛没有编号。这是一个免费得来的特征,而不是问题。
• player_of_match — 在 9 场被取消的比赛中缺失。这是正确的行为。
这里没有任何内容需要插补。我们要做的是把这些缺失转化为信息。
单一取值的列
在这个数据集中,有些列完全不携带任何信息——每一行的取值都相同。
constant_cols = [c for c in df.columns if df[c].nunique(dropna=False) == 1]
for c in constant_cols:
print(f"{c:<15} -> {df[c].unique()[0]}")
event_name -> Indian Premier League
match_type -> T20
overs_limit -> 20
balls_per_over -> 6
gender -> male
team_type -> club
3 · 数据清洗
这个文件中的五个真实问题,逐一修复
原始体育数据几乎从不具备直接用于分析的条件。这个文件存在五个问题,其中任何一个如果被跳过,都会悄悄地污染分析结果:
- 赛季标签不一致 —
2007/08、2009/10和2020/21与纯数字年份混排在一起。 - 球队曾更名 — Delhi Daredevils 改名为 Delhi Capitals;Kings XI Punjab 改名为 Punjab Kings;Royal Challengers Bangalore 改名为 Bengaluru。如果把它们当作不同的球队处理,就会把球队的历史拦腰截断。
- 场地名称重复 —
Wankhede Stadium和Wankhede Stadium, Mumbai其实是同一座球场。 - 51 场比赛的 city = "Unknown" — 但场地信息确切地告诉我们这些比赛是在哪里进行的。
- 日期是字符串,因此目前还无法进行任何基于时间的分析。
我们在副本上进行操作,这样原始的 df 保持原样、随时可比。
3.1 日期与干净的数值型赛季
data = df.copy()
# Dates -> real datetime
data["date"] = pd.to_datetime(data["date"], format="%d-%m-%Y")
# The season label is inconsistent ("2007/08"), but the first match of a season
# always falls in the correct calendar year. So take the year of the earliest match.
season_year = data.groupby("season")["date"].transform("min").dt.year
data["season_year"] = season_year
print(data.groupby("season")["season_year"].first().to_string())
season
2007/08 2008
2009 2009
2009/10 2010
2011 2011
2012 2012
2013 2013
2014 2014
2015 2015
2016 2016
2017 2017
2018 2018
2019 2019
2020/21 2020
2021 2021
2022 2022
2023 2023
2024 2024
2025 2025
2026 2026
3.2 合并更名后的球队
team_renames = {
"Delhi Daredevils": "Delhi Capitals",
"Kings XI Punjab": "Punjab Kings",
"Royal Challengers Bangalore": "Royal Challengers Bengaluru",
"Rising Pune Supergiants": "Rising Pune Supergiant",
}
for col in ["team1", "team2", "toss_winner", "winner"]:
data[col] = data[col].replace(team_renames)
print("Teams after merging renames:", data["team1"].nunique())
print()
print(pd.concat([data["team1"], data["team2"]]).value_counts().to_string())
Teams after merging renames: 15
Mumbai Indians 291
Royal Challengers Bengaluru 286
Delhi Capitals 281
Punjab Kings 278
Kolkata Knight Riders 278
Chennai Super Kings 266
Rajasthan Royals 251
Sunrisers Hyderabad 211
Gujarat Titans 77
Deccan Chargers 75
Lucknow Super Giants 72
Pune Warriors 46
Rising Pune Supergiant 30
Gujarat Lions 30
Kochi Tuskers Kerala 14
3.3 简短球队代码,让图表更易读
short = {
"Mumbai Indians": "MI", "Chennai Super Kings": "CSK", "Kolkata Knight Riders": "KKR",
"Royal Challengers Bengaluru": "RCB", "Rajasthan Royals": "RR", "Delhi Capitals": "DC",
"Sunrisers Hyderabad": "SRH", "Punjab Kings": "PBKS", "Gujarat Titans": "GT",
"Lucknow Super Giants": "LSG", "Deccan Chargers": "DCH", "Pune Warriors": "PWI",
"Gujarat Lions": "GL", "Rising Pune Supergiant": "RPS", "Kochi Tuskers Kerala": "KTK",
}
data["team1_short"] = data["team1"].map(short)
data["team2_short"] = data["team2"].map(short)
data["winner_short"] = data["winner"].map(short)
data["toss_winner_short"] = data["toss_winner"].map(short)
data[["team1", "team1_short", "team2", "team2_short"]].head(3)
team1 team1_short team2 team2_short
0 Royal Challengers Bengaluru RCB Kolkata Knight Riders KKR
1 Punjab Kings PBKS Chennai Super Kings CSK
2 Delhi Capitals DC Rajasthan Royals RR
3.4 场馆去重并修复城市
# Most duplicates are just a city appended after a comma -> keep the part before the first comma
data["venue"] = data["venue"].str.split(",").str[0].str.strip()
# The remaining few are spelling variants or genuine stadium renamings
venue_fixes = {
"M.Chinnaswamy Stadium": "M Chinnaswamy Stadium",
"Punjab Cricket Association Stadium": "Punjab Cricket Association IS Bindra Stadium",
"Zayed Cricket Stadium": "Sheikh Zayed Stadium",
"Sardar Patel Stadium": "Narendra Modi Stadium",
"Feroz Shah Kotla": "Arun Jaitley Stadium",
"Subrata Roy Sahara Stadium": "Maharashtra Cricket Association Stadium",
}
data["venue"] = data["venue"].replace(venue_fixes)
print("Unique venues:", df["venue"].nunique(), "->", data["venue"].nunique())
Unique venues: 60 -> 36
# Cities: merge the Bangalore/Bengaluru spelling and fill the 'Unknown' entries from the venue
data["city"] = data["city"].replace({"Bangalore": "Bengaluru", "Navi Mumbai": "Mumbai"})
city_from_venue = {
"Dubai International Cricket Stadium": "Dubai",
"Sharjah Cricket Stadium": "Sharjah",
}
unknown = data["city"] == "Unknown"
data.loc[unknown, "city"] = data.loc[unknown, "venue"].map(city_from_venue)
print("Cities still unknown:", data["city"].isna().sum())
print("Unique cities:", df["city"].nunique(), "->", data["city"].nunique())
Cities still unknown: 0
Unique cities: 38 -> 35
4 · 特征工程
这个数据集中最有价值的列,恰恰是那些不在文件里的列
原始文件记录了 team1_runs 和 team2_runs——但 team1 不一定是先攻的那支球队。不修正这一点,所有"第一局得分"分析都会沦为无稽之谈。
掷币结果告诉了我们击球顺序:
- 掷币获胜方选择 bat → 掷币获胜方先攻
- 掷币获胜方选择 field → 另一支球队先攻
凭借这一条规则,我们就能解锁各局得分、追分目标、胜负分差,以及比赛是靠守分获胜还是靠追分获胜。
# Who batted first?
toss_is_team1 = data["toss_winner"] == data["team1"]
chose_bat = data["toss_decision"] == "bat"
data["bat_first"] = np.where(chose_bat == toss_is_team1, data["team1"], data["team2"])
data["bat_second"] = np.where(data["bat_first"] == data["team1"], data["team2"], data["team1"])
# Quick sanity check against the very first match of IPL history
data.loc[[0], ["team1", "team2", "toss_winner", "toss_decision",
"team1_runs", "team2_runs", "bat_first", "winner"]]
team1 team2 toss_winner toss_decision team1_runs team2_runs bat_first winner
0 Royal Challengers Bengaluru Kolkata Knight Riders Royal Challengers Bengaluru field 82 222 Kolkata Knight Riders Kolkata Knight Riders
2008 年第 1 场比赛:RCB 赢得掷币并选择先守,于是 Kolkata 先攻并拿下 222 分。bat_first 正确地返回了 Kolkata Knight Riders。这条规则奏效。
first_is_team1 = data["bat_first"] == data["team1"]
data["first_innings_runs"] = np.where(first_is_team1, data["team1_runs"], data["team2_runs"])
data["first_innings_wkts"] = np.where(first_is_team1, data["team1_wickets"], data["team2_wickets"])
data["second_innings_runs"] = np.where(first_is_team1, data["team2_runs"], data["team1_runs"])
data["second_innings_wkts"] = np.where(first_is_team1, data["team2_wickets"], data["team1_wickets"])
data["target"] = data["first_innings_runs"] + 1
data.loc[:, ["bat_first", "first_innings_runs", "bat_second", "second_innings_runs", "target"]].head()
bat_first first_innings_runs bat_second second_innings_runs target
0 Kolkata Knight Riders 222 Royal Challengers Bengaluru 82 223
1 Chennai Super Kings 240 Punjab Kings 207 241
2 Rajasthan Royals 129 Delhi Capitals 132 130
3 Deccan Chargers 110 Kolkata Knight Riders 112 111
4 Mumbai Indians 165 Royal Challengers Bengaluru 166 166
# How was the match won - defending a total, or chasing one?
data["result_side"] = np.select(
[data["winner"] == data["bat_first"], data["winner"] == data["bat_second"]],
["Batting first", "Chasing"],
default="No result / Tie"
)
# Stage of the tournament: unnumbered matches are the playoffs
data["stage"] = np.where(data["match_number"].isna(), "Playoff", "League")
# A single margin column, plus a decided-matches flag we'll reuse constantly
data["margin"] = np.where(data["win_by_runs"] > 0,
data["win_by_runs"].astype(str) + " runs",
data["win_by_wickets"].astype(str) + " wickets")
data["decided"] = data["winner"].notna()
data["result_side"].value_counts()
result_side
Chasing 666
Batting first 552
No result / Tie 25
Name: count, dtype: int64
分析表
keep = ["season_year", "date", "city", "venue", "stage",
"team1", "team2", "team1_short", "team2_short",
"toss_winner", "toss_decision", "bat_first", "bat_second",
"first_innings_runs", "first_innings_wkts",
"second_innings_runs", "second_innings_wkts", "target",
"winner", "winner_short", "result_type", "result_side",
"win_by_runs", "win_by_wickets", "margin", "decided",
"player_of_match"]
ipl = data[keep].copy()
print("Clean analysis table:", ipl.shape)
print("Decided matches :", ipl["decided"].sum())
print("Seasons :", ipl["season_year"].min(), "-", ipl["season_year"].max())
ipl.head()
Clean analysis table: (1243, 27)
Decided matches : 1218
Seasons : 2008 - 2026
season_year date city venue stage team1 team2 team1_short team2_short \
0 2008 2008-04-18 Bengaluru M Chinnaswamy Stadium League Royal Challengers Bengaluru Kolkata Knight Riders RCB KKR
1 2008 2008-04-19 Chandigarh Punjab Cricket Association IS Bindra Stadium League Punjab Kings Chennai Super Kings PBKS CSK
2 2008 2008-04-19 Delhi Arun Jaitley Stadium League Delhi Capitals Rajasthan Royals DC RR
3 2008 2008-04-20 Kolkata Eden Gardens League Kolkata Knight Riders Deccan Chargers KKR DCH
4 2008 2008-04-20 Mumbai Wankhede Stadium League Mumbai Indians Royal Challengers Bengaluru MI RCB
toss_winner toss_decision bat_first bat_second first_innings_runs first_innings_wkts second_innings_runs second_innings_wkts target \
0 Royal Challengers Bengaluru field Kolkata Knight Riders Royal Challengers Bengaluru 222 3 82 10 223
1 Chennai Super Kings bat Chennai Super Kings Punjab Kings 240 5 207 4 241
2 Rajasthan Royals bat Rajasthan Royals Delhi Capitals 129 8 132 1 130
3 Deccan Chargers bat Deccan Chargers Kolkata Knight Riders 110 10 112 5 111
4 Mumbai Indians bat Mumbai Indians Royal Challengers Bengaluru 165 7 166 5 166
winner winner_short result_type result_side win_by_runs win_by_wickets margin decided player_of_match
0 Kolkata Knight Riders KKR complete Batting first 140 0 140 runs True BB McCullum
1 Chennai Super Kings CSK complete Batting first 33 0 33 runs True MEK Hussey
2 Delhi Capitals DC complete Chasing 0 9 9 wickets True MF Maharoof
3 Kolkata Knight Riders KKR complete Chasing 0 5 5 wickets True DJ Hussey
4 Royal Challengers Bengaluru RCB complete Chasing 0 5 5 wickets True MV Boucher
5 · 联盟的历年变迁
问题 1:IPL 真的变成了一个得分更高的赛事吗?
per_season = ipl.groupby("season_year").agg(
matches = ("date", "count"),
avg_first = ("first_innings_runs", "mean"),
avg_second = ("second_innings_runs", "mean"),
highest = ("first_innings_runs", "max"),
venues = ("venue", "nunique"),
).round(1)
per_season.head(19)
matches avg_first avg_second highest venues
season_year
2008 58 161.0 148.3 240 9
2009 57 150.6 136.3 211 8
2010 60 165.0 149.8 246 12
2011 73 152.4 137.4 232 12
2012 74 157.5 145.9 222 12
2013 76 156.2 141.2 263 12
2014 60 163.2 152.3 231 13
2015 59 166.4 144.7 235 13
2016 60 162.6 151.8 248 11
2017 59 165.9 152.5 230 10
2018 60 172.5 159.2 245 10
2019 60 167.0 156.9 232 9
2020 60 170.0 153.6 228 3
2021 60 159.4 151.2 235 7
2022 74 171.1 158.5 222 6
2023 74 182.7 164.4 257 12
2024 71 189.6 176.2 287 13
2025 74 189.0 169.5 286 13
2026 74 193.1 177.9 264 13
fig, ax = plt.subplots(figsize=(12, 5))
bars = ax.bar(per_season.index, per_season["matches"], color=BLUE, edgecolor="white")
# Highlight the two pandemic-affected seasons
for i, yr in enumerate(per_season.index):
if yr in (2020, 2021):
bars[i].set_color(ORANGE)
for x, y in zip(per_season.index, per_season["matches"]):
ax.text(x, y + 1, int(y), ha="center", fontsize=9, color=NAVY)
ax.set_title("Matches played each IPL season")
ax.set_xlabel("Season")
ax.set_ylabel("Matches")
ax.set_xticks(per_season.index)
ax.tick_params(axis="x", rotation=45)
ax.set_ylim(0, 90)
plt.tight_layout()
plt.show()

联盟的规模经历了三次清晰的跃升:头一个十年的大部分时间里是 58–60 场比赛的赛事;当参赛队伍扩军到 9–10 支时(2011–13)跃升至 70+ 场;而在 Gujarat Titans 和 Lucknow Super Giants 加入后,从 2022 年起永久固定为 74 场比赛。橙色柱形标出了在 UAE 举行的 COVID 赛季。
fig, ax = plt.subplots(figsize=(12, 5.5))
ax.plot(per_season.index, per_season["avg_first"], marker="o", lw=2.5,
color=NAVY, label="1st innings average")
ax.plot(per_season.index, per_season["avg_second"], marker="s", lw=2.5,
color=ORANGE, label="2nd innings average")
ax.fill_between(per_season.index, per_season["avg_second"], per_season["avg_first"],
color=BLUE, alpha=0.10)
ax.set_title("Average innings score by season")
ax.set_xlabel("Season")
ax.set_ylabel("Runs")
ax.set_xticks(per_season.index)
ax.tick_params(axis="x", rotation=45)
ax.legend(frameon=False)
plt.tight_layout()
plt.show()

# How often do teams post a really big total?
big = ipl[ipl["first_innings_runs"] >= 200].groupby("season_year").size()
big = big.reindex(per_season.index, fill_value=0)
rate = (big / per_season["matches"] * 100).round(1)
fig, ax = plt.subplots(figsize=(12, 5))
ax.bar(rate.index, rate.values, color=GOLD, edgecolor="white")
for x, y in zip(rate.index, rate.values):
ax.text(x, y + 0.4, f"{y:.0f}%", ha="center", fontsize=9, color=NAVY)
ax.set_title("Share of matches with a 200+ first innings total")
ax.set_xlabel("Season")
ax.set_ylabel("% of matches")
ax.set_xticks(rate.index)
ax.tick_params(axis="x", rotation=45)
plt.tight_layout()
plt.show()

十多年来,第一局平均得分一直徘徊在 160 附近的狭窄区间内。从 2023 年起,它们突破了这一区间。200+ 的比率是更敏锐的信号:过去打出 200 分是整个赛事才出现一次的事,如今已是常态。impact-player 规则、更平坦的球道,以及伴随 T20 长大的一代击球手,都在同一时间到来。
6 · 各特许球队
问题 2:IPL 历史上真正最成功的球队是谁?
played = pd.concat([ipl["team1"], ipl["team2"]]).value_counts()
won = ipl["winner"].value_counts()
teams = pd.DataFrame({"played": played, "won": won}).fillna(0)
teams["won"] = teams["won"].astype(int)
teams["lost"] = teams["played"] - teams["won"]
teams["win_pct"] = (teams["won"] / teams["played"] * 100).round(1)
teams["code"] = teams.index.map(short)
teams = teams.sort_values("win_pct", ascending=False)
teams[["code", "played", "won", "lost", "win_pct"]]
code played won lost win_pct
Gujarat Titans GT 77 47 30 61.0
Chennai Super Kings CSK 266 148 118 55.6
Mumbai Indians MI 291 155 136 53.3
Kolkata Knight Riders KKR 278 140 138 50.4
Rising Pune Supergiant RPS 30 15 15 50.0
Royal Challengers Bengaluru RCB 286 143 143 50.0
Rajasthan Royals RR 251 123 128 49.0
Sunrisers Hyderabad SRH 211 102 109 48.3
Lucknow Super Giants LSG 72 34 38 47.2
Punjab Kings PBKS 278 126 152 45.3
Delhi Capitals DC 281 125 156 44.5
Gujarat Lions GL 30 13 17 43.3
Kochi Tuskers Kerala KTK 14 6 8 42.9
Deccan Chargers DCH 75 29 46 38.7
Pune Warriors PWI 46 12 34 26.1
core = teams[teams["played"] >= 50].sort_values("win_pct")
fig, ax = plt.subplots(figsize=(11, 6))
colors = [ORANGE if v >= 50 else BLUE for v in core["win_pct"]]
ax.barh(core["code"], core["win_pct"], color=colors, edgecolor="white")
ax.axvline(50, color=RED, ls="--", lw=1.5)
ax.text(50.4, -0.4, "50%", color=RED, fontsize=10)
for code_, v, p in zip(core["code"], core["win_pct"], core["played"]):
ax.text(v + 0.6, code_, f"{v}% ({int(p)} matches)", va="center", fontsize=9, color=NAVY)
ax.set_title("Win percentage — teams with 50+ matches played")
ax.set_xlabel("Win %")
ax.set_xlim(0, 75)
plt.tight_layout()
plt.show()

冠军头衔:特许球队真正在乎的唯一数字
# The final is the last match played in each season
finals_idx = ipl.groupby("season_year")["date"].idxmax()
finals = ipl.loc[finals_idx, ["season_year", "bat_first", "bat_second",
"first_innings_runs", "second_innings_runs",
"winner", "margin", "venue", "player_of_match"]]
finals = finals.rename(columns={"winner": "champion"}).set_index("season_year")
finals[["champion", "margin", "venue", "player_of_match"]]
champion margin venue player_of_match
season_year
2008 Rajasthan Royals 3 wickets Dr DY Patil Sports Academy YK Pathan
2009 Deccan Chargers 6 runs New Wanderers Stadium A Kumble
2010 Chennai Super Kings 22 runs Dr DY Patil Sports Academy SK Raina
2011 Chennai Super Kings 58 runs MA Chidambaram Stadium M Vijay
2012 Kolkata Knight Riders 5 wickets MA Chidambaram Stadium MS Bisla
2013 Mumbai Indians 23 runs Eden Gardens KA Pollard
2014 Kolkata Knight Riders 3 wickets M Chinnaswamy Stadium MK Pandey
2015 Mumbai Indians 41 runs Eden Gardens RG Sharma
2016 Sunrisers Hyderabad 8 runs M Chinnaswamy Stadium BCJ Cutting
2017 Mumbai Indians 1 runs Rajiv Gandhi International Stadium KH Pandya
2018 Chennai Super Kings 8 wickets Wankhede Stadium SR Watson
2019 Mumbai Indians 1 runs Rajiv Gandhi International Stadium JJ Bumrah
2020 Mumbai Indians 5 wickets Dubai International Cricket Stadium TA Boult
2021 Chennai Super Kings 27 runs Dubai International Cricket Stadium F du Plessis
2022 Gujarat Titans 7 wickets Narendra Modi Stadium HH Pandya
2023 Chennai Super Kings 5 wickets Narendra Modi Stadium DP Conway
2024 Kolkata Knight Riders 8 wickets MA Chidambaram Stadium MA Starc
2025 Royal Challengers Bengaluru 6 runs Narendra Modi Stadium KH Pandya
2026 Royal Challengers Bengaluru 5 wickets Narendra Modi Stadium V Kohli
titles = finals["champion"].value_counts()
runners = finals.apply(
lambda r: r["bat_second"] if r["champion"] == r["bat_first"] else r["bat_first"], axis=1
).value_counts()
honours = pd.DataFrame({"titles": titles, "runner_up": runners}).fillna(0).astype(int)
honours["finals"] = honours["titles"] + honours["runner_up"]
honours = honours.sort_values(["titles", "finals"], ascending=False)
honours
titles runner_up finals
Chennai Super Kings 5 5 10
Mumbai Indians 5 1 6
Kolkata Knight Riders 3 1 4
Royal Challengers Bengaluru 2 3 5
Gujarat Titans 1 2 3
Sunrisers Hyderabad 1 2 3
Rajasthan Royals 1 1 2
Deccan Chargers 1 0 1
Punjab Kings 0 2 2
Delhi Capitals 0 1 1
Rising Pune Supergiant 0 1 1
h = honours.sort_values("titles")
idx = np.arange(len(h))
fig, ax = plt.subplots(figsize=(11, 6))
ax.barh(idx, h["titles"], color=ORANGE, edgecolor="white", label="Titles")
ax.barh(idx, h["runner_up"], left=h["titles"], color=GREY, edgecolor="white", label="Runner-up")
ax.set_yticks(idx)
ax.set_yticklabels([short.get(t, t) for t in h.index])
ax.set_title("Finals record by franchise")
ax.set_xlabel("Number of finals reached")
ax.legend(frameon=False, loc="lower right")
plt.tight_layout()
plt.show()

Chennai 和 Mumbai 在两份榜单上都位居榜首,这正是它们被视为联盟王朝球队的原因。但请注意那些胜率很高、奖杯柜却空空如也的球队——高常规赛胜率能把你送进季后赛,却赢不下那一场真正算数的比赛。
常规赛表现 vs 季后赛表现
playoffs = ipl[(ipl["stage"] == "Playoff") & ipl["decided"]]
po_played = pd.concat([playoffs["team1"], playoffs["team2"]]).value_counts()
po_won = playoffs["winner"].value_counts()
po = pd.DataFrame({"playoff_games": po_played, "playoff_wins": po_won}).fillna(0).astype(int)
po["playoff_win_pct"] = (po["playoff_wins"] / po["playoff_games"] * 100).round(1)
po = po.join(teams["win_pct"].rename("league_win_pct"))
po["clutch_gap"] = (po["playoff_win_pct"] - po["league_win_pct"]).round(1)
po.sort_values("playoff_games", ascending=False).head(10)
playoff_games playoff_wins playoff_win_pct league_win_pct clutch_gap
Chennai Super Kings 26 17 65.4 55.6 9.8
Mumbai Indians 22 14 63.6 53.3 10.3
Royal Challengers Bengaluru 20 10 50.0 50.0 0.0
Kolkata Knight Riders 15 10 66.7 50.4 16.3
Sunrisers Hyderabad 15 6 40.0 48.3 -8.3
Rajasthan Royals 13 6 46.2 49.0 -2.8
Delhi Capitals 11 2 18.2 44.5 -26.3
Gujarat Titans 9 4 44.4 61.0 -16.6
Punjab Kings 7 2 28.6 45.3 -16.7
Deccan Chargers 4 2 50.0 38.7 11.3
clutch_gap 是球队季后赛胜率与整体胜率之差。正数意味着球队在淘汰赛到来时会提升水准;负数则意味着它会掉链子。
7 · 掷币
问题 3:掷硬币真的能决定什么吗?
print(ipl["toss_decision"].value_counts())
print()
print((ipl["toss_decision"].value_counts(normalize=True) * 100).round(1))
toss_decision
field 825
bat 418
Name: count, dtype: int64
toss_decision
field 66.4
bat 33.6
Name: proportion, dtype: float64
toss_trend = pd.crosstab(ipl["season_year"], ipl["toss_decision"], normalize="index") * 100
fig, ax = plt.subplots(figsize=(12, 5.5))
ax.stackplot(toss_trend.index, toss_trend["field"], toss_trend["bat"],
colors=[BLUE, GOLD], labels=["Chose to field", "Chose to bat"], alpha=0.9)
ax.axhline(50, color="white", ls="--", lw=1.5)
ax.set_title("Toss decision by season (% of matches)")
ax.set_xlabel("Season")
ax.set_ylabel("% of tosses")
ax.set_xlim(toss_trend.index.min(), toss_trend.index.max())
ax.set_ylim(0, 100)
ax.set_xticks(toss_trend.index)
ax.tick_params(axis="x", rotation=45)
ax.legend(loc="lower left", frameon=False)
plt.tight_layout()
plt.show()

队长们的选择起初大致五五开,随后急剧倒向先守。从 2015 年起,每四场比赛中就有三场默认选择追分——晚间比赛中的露水,以及提前知道目标分数的那份踏实感,都把天平推向同一个方向。
decided = ipl[ipl["decided"]].copy()
decided["toss_won_match"] = decided["toss_winner"] == decided["winner"]
toss_advantage = decided["toss_won_match"].mean() * 100
print(f"Toss winner also won the match: {toss_advantage:.1f}% of {len(decided)} decided matches")
Toss winner also won the match: 51.6% of 1218 decided matches
by_season = decided.groupby("season_year")["toss_won_match"].mean() * 100
fig, ax = plt.subplots(figsize=(12, 5.5))
colors = [GREEN if v >= 50 else RED for v in by_season.values]
ax.bar(by_season.index, by_season.values, color=colors, edgecolor="white")
ax.axhline(50, color=NAVY, ls="--", lw=2)
ax.axhline(toss_advantage, color=ORANGE, lw=2,
label=f"All-time average {toss_advantage:.1f}%")
ax.set_title("How often did the toss winner win the match?")
ax.set_xlabel("Season")
ax.set_ylabel("% of matches")
ax.set_ylim(0, 80)
ax.set_xticks(by_season.index)
ax.tick_params(axis="x", rotation=45)
ax.legend(frameon=False)
plt.tight_layout()
plt.show()

在近 1,220 场分出胜负的比赛中,掷币获胜方的获胜次数只是略微过半。各个赛季在 50% 上下剧烈波动,而这正是随机噪声该有的样子。我们将在第 10 节用卡方检验正式证实这一点。
8 · 先攻 vs 追分
问题 4:追分真的更容易吗?
sides = decided["result_side"].value_counts()
print(sides)
print()
print((sides / sides.sum() * 100).round(1))
result_side
Chasing 666
Batting first 552
Name: count, dtype: int64
result_side
Chasing 54.7
Batting first 45.3
Name: count, dtype: float64
chase_trend = pd.crosstab(decided["season_year"], decided["result_side"], normalize="index") * 100
fig, ax = plt.subplots(figsize=(12, 5.5))
ax.plot(chase_trend.index, chase_trend["Chasing"], marker="o", lw=2.5,
color=ORANGE, label="Chasing team won")
ax.plot(chase_trend.index, chase_trend["Batting first"], marker="s", lw=2.5,
color=NAVY, label="Defending team won")
ax.axhline(50, color=GREY, ls="--", lw=1.5)
ax.set_title("Win rate: chasing vs defending, by season")
ax.set_xlabel("Season")
ax.set_ylabel("% of decided matches")
ax.set_xticks(chase_trend.index)
ax.tick_params(axis="x", rotation=45)
ax.legend(frameon=False)
plt.tight_layout()
plt.show()

目标分数的大小会改变答案吗?
bins = [0, 140, 160, 180, 200, 400]
labels = ["Under 140", "140-159", "160-179", "180-199", "200+"]
decided["target_band"] = pd.cut(decided["first_innings_runs"], bins=bins, labels=labels)
band = decided.groupby("target_band", observed=True).agg(
matches = ("winner", "count"),
chase_win_pct = ("result_side", lambda s: (s == "Chasing").mean() * 100)
).round(1)
band
matches chase_win_pct
target_band
Under 140 225 82.2
140-159 255 71.8
160-179 311 52.4
180-199 222 39.6
200+ 205 22.9
fig, ax = plt.subplots(figsize=(11, 5.5))
bars = ax.bar(band.index.astype(str), band["chase_win_pct"], color=BLUE, edgecolor="white")
ax.axhline(50, color=RED, ls="--", lw=1.5)
for b, v, n in zip(bars, band["chase_win_pct"], band["matches"]):
ax.text(b.get_x() + b.get_width()/2, v + 1.5, f"{v:.0f}%", ha="center", color=NAVY, fontweight="bold")
ax.text(b.get_x() + b.get_width()/2, 3, f"n={n}", ha="center", color="white", fontsize=9)
ax.set_title("Chasing success rate by size of the first innings total")
ax.set_xlabel("First innings score")
ax.set_ylabel("Chases won (%)")
ax.set_ylim(0, 85)
plt.tight_layout()
plt.show()

低于 160 分时,追分一方明显占优。这一优势在 160–179 区间急剧衰减,到 180 分时已荡然无存, 而一旦有球队拿到 200+ 的总分,天平便会果断倒向防守一方。那些不假思索便选择先投球的队长 在大多数比赛中做出了正确的决定——只是在那些球道平坦如公路的比赛中并非如此。
9 · 场地与条件
问题 5:哪些场地是击球手的乐园,哪些更有利于追分?
venue_stats = ipl.groupby("venue").agg(
matches = ("date", "count"),
avg_first = ("first_innings_runs", "mean"),
highest = ("first_innings_runs", "max"),
).round(1)
chase_pct = decided.groupby("venue")["result_side"].apply(lambda s: (s == "Chasing").mean() * 100)
venue_stats["chase_win_pct"] = chase_pct.round(1)
top_venues = venue_stats[venue_stats["matches"] >= 25].sort_values("matches", ascending=False)
top_venues
matches avg_first highest chase_win_pct
venue
Wankhede Stadium 132 173.2 243 55.0
Eden Gardens 107 168.6 261 57.1
M Chinnaswamy Stadium 104 174.2 287 56.0
Arun Jaitley Stadium 104 172.0 278 53.5
MA Chidambaram Stadium 98 165.7 246 46.9
Rajiv Gandhi International Stadium 90 168.5 286 54.5
Sawai Mansingh Stadium 68 169.1 229 64.7
Punjab Cricket Association IS Bindra Stadium 61 168.3 257 55.7
Narendra Modi Stadium 53 180.8 243 50.0
Maharashtra Cricket Association Stadium 51 162.4 211 45.1
Dubai International Cricket Stadium 46 164.4 219 51.2
Dr DY Patil Sports Academy 37 159.6 216 54.1
Sheikh Zayed Stadium 37 159.3 235 60.0
Bharat Ratna Shri Atal Bihari Vajpayee Ekana Cr... 29 174.9 235 59.3
Sharjah Cricket Stadium 28 159.0 228 64.3
Brabourne Stadium 27 178.5 217 48.1
v = top_venues.sort_values("matches")
labels = [name if len(name) <= 30 else name[:28] + "..." for name in v.index]
fig, ax = plt.subplots(figsize=(11, 6.5))
ax.barh(labels, v["matches"], color=NAVY, edgecolor="white")
for lab, n in zip(labels, v["matches"]):
ax.text(n + 1, lab, int(n), va="center", fontsize=9, color=NAVY)
ax.set_title("Busiest IPL grounds (25+ matches)")
ax.set_xlabel("Matches hosted")
plt.tight_layout()
plt.show()

v = top_venues.sort_values("avg_first")
fig, axes = plt.subplots(1, 2, figsize=(14, 6.5), sharey=True)
labels = [name if len(name) <= 28 else name[:26] + "..." for name in v.index]
axes[0].barh(labels, v["avg_first"], color=GOLD, edgecolor="white")
axes[0].set_title("Average first innings score")
axes[0].set_xlim(130, v["avg_first"].max() + 12)
for lab, x in zip(labels, v["avg_first"]):
axes[0].text(x + 1, lab, f"{x:.0f}", va="center", fontsize=9, color=NAVY)
colors = [GREEN if x >= 50 else RED for x in v["chase_win_pct"]]
axes[1].barh(labels, v["chase_win_pct"], color=colors, edgecolor="white")
axes[1].axvline(50, color=NAVY, ls="--", lw=1.5)
axes[1].set_title("Chasing team win rate (%)")
axes[1].set_xlim(0, 75)
for lab, x in zip(labels, v["chase_win_pct"]):
axes[1].text(x + 1, lab, f"{x:.0f}%", va="center", fontsize=9, color=NAVY)
plt.tight_layout()
plt.show()

右侧的绿色条形代表追分通常能成功的场地;红色条形代表防守总分更有效的场地。把两个面板结合起来看才是关键——平均得分高且条形为红色的场地,正是适合先击球并拿下高分的球场。
10 · 分差、大胜与险胜
IPL 比赛究竟有多胶着?
run_wins = decided[decided["win_by_runs"] > 0]
wicket_wins = decided[decided["win_by_wickets"] > 0]
fig, axes = plt.subplots(1, 2, figsize=(13.5, 5))
axes[0].hist(run_wins["win_by_runs"], bins=25, color=NAVY, edgecolor="white")
axes[0].axvline(run_wins["win_by_runs"].median(), color=ORANGE, lw=2.5,
label=f"Median {run_wins['win_by_runs'].median():.0f} runs")
axes[0].set_title(f"Wins while defending (n={len(run_wins)})")
axes[0].set_xlabel("Margin (runs)")
axes[0].set_ylabel("Matches")
axes[0].legend(frameon=False)
axes[1].hist(wicket_wins["win_by_wickets"], bins=np.arange(0.5, 11.5, 1),
color=BLUE, edgecolor="white")
axes[1].axvline(wicket_wins["win_by_wickets"].median(), color=ORANGE, lw=2.5,
label=f"Median {wicket_wins['win_by_wickets'].median():.0f} wickets")
axes[1].set_title(f"Wins while chasing (n={len(wicket_wins)})")
axes[1].set_xlabel("Margin (wickets)")
axes[1].set_xticks(range(1, 11))
axes[1].legend(frameon=False)
plt.tight_layout()
plt.show()

# The five most one-sided results in each direction
print("BIGGEST WINS BY RUNS")
cols = ["season_year", "bat_first", "first_innings_runs", "bat_second", "second_innings_runs", "margin"]
print(decided.nlargest(5, "win_by_runs")[cols].to_string(index=False))
print()
print("BIGGEST WINS BY WICKETS (fewest balls to spare not in this file, so ties broken by target)")
print(decided[decided["win_by_wickets"] == 10].nlargest(5, "target")[cols].to_string(index=False))
BIGGEST WINS BY RUNS
season_year bat_first first_innings_runs bat_second second_innings_runs margin
2017 Mumbai Indians 212 Delhi Capitals 66 146 runs
2016 Royal Challengers Bengaluru 248 Gujarat Lions 104 144 runs
2008 Kolkata Knight Riders 222 Royal Challengers Bengaluru 82 140 runs
2015 Royal Challengers Bengaluru 226 Punjab Kings 88 138 runs
2013 Royal Challengers Bengaluru 263 Pune Warriors 133 130 runs
BIGGEST WINS BY WICKETS (fewest balls to spare not in this file, so ties broken by target)
season_year bat_first first_innings_runs bat_second second_innings_runs margin
2025 Delhi Capitals 199 Gujarat Titans 205 10 wickets
2017 Gujarat Lions 183 Kolkata Knight Riders 184 10 wickets
2020 Punjab Kings 178 Chennai Super Kings 181 10 wickets
2021 Rajasthan Royals 177 Royal Challengers Bengaluru 181 10 wickets
2024 Lucknow Super Giants 165 Sunrisers Hyderabad 167 10 wickets
# The nail-biters
print(f"Matches won by 1-3 runs: {(decided['win_by_runs'].between(1,3)).sum()}")
print(f"Tied matches (super over): {(ipl['result_type'] == 'tie').sum()}")
print(f"Abandoned / no result: {(ipl['result_type'] == 'no result').sum()}")
print()
decided[decided["win_by_runs"] == 1][["season_year", "bat_first", "bat_second",
"first_innings_runs", "second_innings_runs", "winner"]]
Matches won by 1-3 runs: 37
Tied matches (super over): 16
Abandoned / no result: 9
season_year bat_first bat_second first_innings_runs second_innings_runs winner
44 2008 Punjab Kings Mumbai Indians 189 188 Punjab Kings
104 2009 Punjab Kings Deccan Chargers 134 133 Punjab Kings
284 2012 Delhi Capitals Rajasthan Royals 152 151 Delhi Capitals
290 2012 Mumbai Indians Pune Warriors 120 119 Mumbai Indians
459 2015 Chennai Super Kings Delhi Capitals 150 149 Chennai Super Kings
539 2016 Gujarat Lions Delhi Capitals 172 171 Gujarat Lions
555 2016 Royal Challengers Bengaluru Punjab Kings 175 174 Royal Challengers Bengaluru
635 2017 Mumbai Indians Rising Pune Supergiant 129 128 Mumbai Indians
734 2019 Royal Challengers Bengaluru Chennai Super Kings 161 160 Royal Challengers Bengaluru
755 2019 Mumbai Indians Chennai Super Kings 149 148 Mumbai Indians
837 2021 Royal Challengers Bengaluru Delhi Capitals 171 170 Royal Challengers Bengaluru
1017 2023 Lucknow Super Giants Kolkata Knight Riders 176 175 Lucknow Super Giants
1059 2024 Kolkata Knight Riders Royal Challengers Bengaluru 222 221 Kolkata Knight Riders
1073 2024 Sunrisers Hyderabad Rajasthan Royals 201 200 Sunrisers Hyderabad
1147 2025 Kolkata Knight Riders Rajasthan Royals 206 205 Kolkata Knight Riders
1182 2026 Gujarat Titans Delhi Capitals 210 209 Gujarat Titans
11 · 正面交锋
哪些宿敌对决实际上是一边倒的?
big_teams = teams[teams["played"] >= 150].index.tolist()
h2h = pd.DataFrame(0.0, index=big_teams, columns=big_teams)
for a in big_teams:
for b in big_teams:
if a == b:
h2h.loc[a, b] = np.nan
continue
games = decided[((decided["team1"] == a) & (decided["team2"] == b)) |
((decided["team1"] == b) & (decided["team2"] == a))]
if len(games) > 0:
h2h.loc[a, b] = round((games["winner"] == a).mean() * 100, 1)
h2h.index = [short[t] for t in h2h.index]
h2h.columns = [short[t] for t in h2h.columns]
h2h
CSK MI KKR RCB RR SRH PBKS DC
CSK NaN 48.8 65.6 60.0 50.0 62.5 50.0 63.6
MI 51.2 NaN 67.6 54.3 50.0 56.0 51.4 55.3
KKR 34.4 32.4 NaN 55.6 58.6 64.5 61.8 57.1
RCB 40.0 45.7 44.4 NaN 53.1 46.2 52.6 60.6
RR 50.0 50.0 41.4 46.9 NaN 41.7 60.0 48.4
SRH 37.5 44.0 35.5 53.8 58.3 NaN 69.2 56.0
PBKS 50.0 48.6 38.2 47.4 40.0 30.8 NaN 51.4
DC 36.4 44.7 42.9 39.4 51.6 44.0 48.6 NaN
fig, ax = plt.subplots(figsize=(9, 7))
sns.heatmap(h2h, annot=True, fmt=".0f", cmap="RdYlGn", center=50,
linewidths=1.5, linecolor="white", cbar_kws={"label": "Win % of row team"},
vmin=25, vmax=75, ax=ax)
ax.set_title("Head-to-head win % (row team vs column team)")
plt.tight_layout()
plt.show()

逐行阅读热力图:绿色单元格表示行所对应的球队压制该对手。大多数对抗都接近中性的 50 刻度,这悄然提醒着我们这个联赛竞争何等激烈——真正一边倒的对决只是例外,而非常态。
12 · 制胜功臣
问题 6:谁决定了最多的 IPL 比赛?
potm = ipl["player_of_match"].value_counts()
top_potm = potm.head(15).sort_values()
fig, ax = plt.subplots(figsize=(11, 6.5))
colors = [ORANGE if i >= len(top_potm) - 3 else BLUE for i in range(len(top_potm))]
ax.barh(top_potm.index, top_potm.values, color=colors, edgecolor="white")
for name, v in zip(top_potm.index, top_potm.values):
ax.text(v + 0.2, name, int(v), va="center", fontsize=9, color=NAVY)
ax.set_title("Most Player of the Match awards, 2008-2026")
ax.set_xlabel("Awards")
plt.tight_layout()
plt.show()

# Who won the most awards in each individual season?
season_potm = (ipl.dropna(subset=["player_of_match"])
.groupby(["season_year", "player_of_match"])
.size()
.reset_index(name="awards")
.sort_values(["season_year", "awards"], ascending=[True, False])
.groupby("season_year")
.head(1)
.set_index("season_year"))
season_potm
player_of_match awards
season_year
2008 SE Marsh 5
2009 YK Pathan 3
2010 SR Tendulkar 4
2011 CH Gayle 6
2012 CH Gayle 5
2013 MEK Hussey 5
2014 GJ Maxwell 4
2015 DA Warner 4
2016 V Kohli 5
2017 BA Stokes 3
2018 Rashid Khan 4
2019 AD Russell 4
2020 AB de Villiers 3
2021 RD Gaikwad 4
2022 Kuldeep Yadav 4
2023 Shubman Gill 4
2024 Abhishek Sharma 3
2025 KH Pandya 3
2026 Ishan Kishan 3
print(f"Different players have won at least one award : {len(potm)}")
print(f"Players with exactly one award : {(potm == 1).sum()}")
print(f"Awards taken by the top 15 players : {potm.head(15).sum()} "
f"({potm.head(15).sum() / potm.sum() * 100:.1f}% of all awards)")
Different players have won at least one award : 321
Players with exactly one award : 123
Awards taken by the top 15 players : 268 (21.7% of all awards)
13 · 统计检验
图表提出猜想,检验予以证实。
上述探索得出了四个论断。现在,每一个论断都将在标准的 α = 0.05 显著性水平下接受正式检验。
检验 1 — 赢得掷币会提高赢下比赛的概率吗?
# H0: the toss winner wins the match exactly 50% of the time
n = len(decided)
wins = (decided["toss_winner"] == decided["winner"]).sum()
# Binomial test against a fair 50/50
res = stats.binomtest(wins, n, p=0.5)
print(f"Toss winners who also won the match : {wins} / {n} ({wins/n*100:.2f}%)")
print(f"Binomial test p-value : {res.pvalue:.4f}")
print()
print("Reject H0" if res.pvalue < 0.05 else "Fail to reject H0 - the toss gives no real advantage")
Toss winners who also won the match : 628 / 1218 (51.56%)
Binomial test p-value : 0.2891
Fail to reject H0 - the toss gives no real advantage
检验 2 — 掷币时的选择与获胜是否相关?
# H0: the toss decision (bat / field) is independent of the match result for the toss winner
table = pd.crosstab(decided["toss_decision"],
decided["toss_winner"] == decided["winner"])
table.columns = ["Toss winner lost", "Toss winner won"]
print(table)
print()
chi2, p, dof, expected = stats.chi2_contingency(table)
print(f"Chi-square : {chi2:.3f}")
print(f"p-value : {p:.4f}")
print()
print("Reject H0 - the decision matters" if p < 0.05 else "Fail to reject H0 - no significant link")
Toss winner lost Toss winner won
toss_decision
bat 223 185
field 367 443
Chi-square : 9.123
p-value : 0.0025
Reject H0 - the decision matters
检验 3 — 近年来的第一局得分真的更高吗?
early = ipl[ipl["season_year"] <= 2015]["first_innings_runs"]
modern = ipl[ipl["season_year"] >= 2022]["first_innings_runs"]
print(f"2008-2015 : mean {early.mean():.1f} (n={len(early)})")
print(f"2022-2026 : mean {modern.mean():.1f} (n={len(modern)})")
print(f"Difference: {modern.mean() - early.mean():.1f} runs")
print()
t, p = stats.ttest_ind(modern, early, equal_var=False) # Welch's t-test
print(f"Welch t-statistic : {t:.3f}")
print(f"p-value : {p:.6f}")
pooled = np.sqrt((early.var() + modern.var()) / 2)
print(f"Cohen's d : {(modern.mean() - early.mean()) / pooled:.3f}")
print()
print("Reject H0 - scoring really has risen" if p < 0.05 else "Fail to reject H0")
2008-2015 : mean 158.8 (n=517)
2022-2026 : mean 185.1 (n=367)
Difference: 26.3 runs
Welch t-statistic : 11.329
p-value : 0.000000
Cohen's d : 0.786
Reject H0 - scoring really has risen
显著的 p 值只能告诉我们差异并非出于运气。Cohen's d 则告诉我们这个差异是否大到值得在意——大致而言,0.2 为小,0.5 为中,0.8 为大。
检验 4 — 场地会改变追分优势吗?
# H0: chase success is independent of the ground
busy = decided[decided["venue"].isin(top_venues.index)]
table = pd.crosstab(busy["venue"], busy["result_side"])
print(table)
print()
chi2, p, dof, expected = stats.chi2_contingency(table)
print(f"Chi-square : {chi2:.3f} dof: {dof}")
print(f"p-value : {p:.4f}")
print()
print("Reject H0 - venue matters" if p < 0.05 else "Fail to reject H0 - venues behave alike")
result_side Batting first Chasing
venue
Arun Jaitley Stadium 47 54
Bharat Ratna Shri Atal Bihari Vajpayee Ekana Cr... 11 16
Brabourne Stadium 14 13
Dr DY Patil Sports Academy 17 20
Dubai International Cricket Stadium 21 22
Eden Gardens 45 60
M Chinnaswamy Stadium 44 56
MA Chidambaram Stadium 51 45
Maharashtra Cricket Association Stadium 28 23
Narendra Modi Stadium 26 26
Punjab Cricket Association IS Bindra Stadium 27 34
Rajiv Gandhi International Stadium 40 48
Sawai Mansingh Stadium 24 44
Sharjah Cricket Stadium 10 18
Sheikh Zayed Stadium 14 21
Wankhede Stadium 59 72
Chi-square : 10.218 dof: 15
p-value : 0.8058
Fail to reject H0 - venues behave alike
第 9 节的条形图看起来很有说服力:一些场地似乎明显利于追分,另一些则利于防守。但在比赛场次最多的那些场地上,卡方检验的结果远高于 0.05,因此我们无法否定“所有场地表现相似”这一观点。每个场地仅有 30–130 场比赛,我们所能看到的差距完全处于纯靠运气就能产生的范围之内。
这是整个 notebook 中最有价值的一个结果,因为它与图表相矛盾。仅凭肉眼观察条形图就断言某个场地是“追分福地”,恰恰是这项检验要揪出的错误。
数值型比赛特征之间的相关性
num = ipl[["first_innings_runs", "first_innings_wkts",
"second_innings_runs", "second_innings_wkts",
"win_by_runs", "win_by_wickets", "season_year"]]
fig, ax = plt.subplots(figsize=(9, 7))
mask = np.triu(np.ones_like(num.corr(), dtype=bool))
sns.heatmap(num.corr(), mask=mask, annot=True, fmt=".2f", cmap="RdBu_r", center=0,
linewidths=1.5, linecolor="white", vmin=-1, vmax=1, ax=ax)
ax.set_title("Correlation matrix")
plt.tight_layout()
plt.show()

矩阵中最强的关系也是最显而易见的一个——第一局总分越大,成功防守该总分时的获胜分差就越大。season_year 与 first_innings_runs 之间温和的正相关,正是我们在检验 3 中测量到的得分膨胀,从另一个角度再次显现。
14 · 执行摘要
以上全部内容,一屏呈现
summary = pd.DataFrame({
"Metric": [
"Seasons analysed",
"Matches analysed",
"Matches with a decided result",
"Teams that have played",
"Grounds used",
"Cities hosted in",
"Average first innings score",
"Average first innings score (2008-2015)",
"Average first innings score (2022-2026)",
"Highest first innings total",
"Matches won chasing",
"Toss winner also won the match",
"Captains choosing to field",
"Tied matches (super over)",
"Abandoned matches",
],
"Value": [
ipl["season_year"].nunique(),
len(ipl),
int(ipl["decided"].sum()),
pd.concat([ipl["team1"], ipl["team2"]]).nunique(),
ipl["venue"].nunique(),
ipl["city"].nunique(),
f"{ipl['first_innings_runs'].mean():.1f}",
f"{early.mean():.1f}",
f"{modern.mean():.1f}",
f"{ipl['first_innings_runs'].max()}",
f"{(decided['result_side'] == 'Chasing').mean() * 100:.1f}%",
f"{(decided['toss_winner'] == decided['winner']).mean() * 100:.1f}%",
f"{(ipl['toss_decision'] == 'field').mean() * 100:.1f}%",
int((ipl["result_type"] == "tie").sum()),
int((ipl["result_type"] == "no result").sum()),
]
})
summary
Metric Value
0 Seasons analysed 19
1 Matches analysed 1243
2 Matches with a decided result 1218
3 Teams that have played 15
4 Grounds used 36
5 Cities hosted in 35
6 Average first innings score 168.7
7 Average first innings score (2008-2015) 158.8
8 Average first innings score (2022-2026) 185.1
9 Highest first innings total 287
10 Matches won chasing 54.7%
11 Toss winner also won the match 51.6%
12 Captains choosing to field 66.4%
13 Tied matches (super over) 16
14 Abandoned matches 9
六个答案
1. 得分真的上涨了吗?
是的,而且涨幅并不含蓄。第一局平均得分在十多年里一直稳定在 160 附近,随后从 2023 年起急剧攀升。这一上升在统计上显著,200+ 的出现频率也从罕见变成了常态。
2. 哪支球队最成功?
Chennai 和 Mumbai 是王朝级的球队——它们既霸榜冠军数,也占据胜率表的顶端。有几支球队交出了相当不错的联赛胜率,却从未将其转化为奖杯,这恰恰说明强势的联赛阶段几乎保证不了什么。
3. 掷币能决定比赛吗?
基本上不能。掷币获胜方的比赛胜率在统计上与抛硬币无异,而各赛季在 50% 上下摇摆的幅度,正是随机性所能产生的结果。
4. 追分更容易吗?
通常是的——但有条件。总分低于 160 时,追分一方明显占优;到 180 左右优势消失,超过 200 后则发生逆转。队长们压倒性地选择先投球,而对大多数总分而言,这正是正确的直觉。
5. 哪些球场偏爱谁?
没有图表显示的那么明显。各球场之间的平均得分确实存在真实差异,但表面上“利于追分”与“利于防守”的划分并未通过卡方检验——在当前样本量下,这些差距仍在随机波动所能产生的范围之内。目标分数的大小,远比球场的名字更能预测追分的成败。
6. 谁最能凭一己之力赢下比赛?
Player of the Match(全场最佳球员)榜单由一小群职业生涯漫长的击球手主导。数据集中绝大多数球员要么只获得过一次该奖项,要么从未获得——在这个层面上,决定比赛胜负的能力高度集中。
这份分析无法告诉你的事
对局限保持坦诚是分内之事:
- 没有逐球(ball-by-ball)数据。 打击率、投球经济率、Powerplay 与 death over 阶段的表现在这里全都不可见。若能拿到逐轮(over-by-over)数据,得分膨胀这一发现会有力得多。
- Super Over 的结果未被记录,因此 16 场战平的比赛没有胜者,在所有胜率计算中均被排除。
- 相关不等于因果。 更高的得分与 impact-player 规则在时间上重合,但这份文件无法将原因从更平坦的球道、更出色的击球手或更短的边界中剥离出来。
- 阵容列未被使用。
team1_players和team2_players保存着完整的十一人名单——一个自然的后续步骤是进行球员出场情况与阵容延续性分析。
感谢阅读。如果这个 notebook 对你有所帮助,一个 upvote 能让更多人看到它 — 我也随时欢迎在评论区讨论方法论。
端到端分析 · IPL 2008–2026
浙公网安备 33010602011771号