Bony1029

导航

IPL的19个赛季 对1243 场比赛的完整端到端分析

Python 3

数据与完整代码来源

INDIAN PREMIER LEAGUE  |  2008 – 2026

IPL 的 19 个赛季

对 1,243 场比赛的完整端到端分析

数据清洗  •  特征工程  •  探索性分析  •  统计检验  •  洞察结论


本笔记本做了什么

IPL 是全球收视率最高的特许经营板球联赛。本笔记本取 从 2008 年揭幕战到 2026 年决赛的每一场 IPL 比赛 的原始比赛级记录,并像分析师在实际工作中那样,对其进行端到端的处理:

步骤 具体内容
1. 加载与检查 形状、数据类型、缺失值,以及对原始文件的初步查看
2. 清洗 修正赛季标签、合并更名后的球队、场地去重、修复城市信息
3. 特征工程 推导先攻方、追分目标、获胜分差以及赛事阶段
4. 探索 15+ 张关于时代、球队、掷币、场地、分差与球员的可视化图表
5. 检验 用卡方检验和 t 检验判断这些模式究竟是真实规律还是噪声
6. 结论 用平实的语言总结数据到底说了什么

我们要回答的问题:

  1. IPL 真的已经变成一个得分更高的联赛了吗,还是只是我们的错觉?
  2. 哪支球队最成功——按冠军数衡量,还是按胜率衡量?
  3. 赢得掷币真的能帮你赢下比赛吗?
  4. 追分真的比守分更容易吗?这一点随时间推移发生了变化吗?
  5. 哪些球场是击球天堂,哪些又是追分陷阱?
  6. 决定最多比赛胜负的球员都有谁?

1  ·  环境设置

库、绘图主题以及贯穿全程使用的配色方案

import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
from scipy import stats

import warnings
warnings.filterwarnings("ignore")

pd.set_option("display.max_columns", 50)
pd.set_option("display.width", 200)

配色方案只定义一次,然后在每张图表中复用。统一的配色是让笔记本看起来像一件完整的作品、而不是二十张互不相干的图表的最省力的办法。

NAVY   = "#12263F"
BLUE   = "#2E86AB"
ORANGE = "#FF6B35"
GOLD   = "#F4A261"
GREEN  = "#2A9D8F"
RED    = "#E63946"
GREY   = "#8D99AE"

SEQ = [NAVY, BLUE, ORANGE, GREEN, GOLD, RED, GREY]

plt.rcParams["figure.figsize"]   = (11, 5.5)
plt.rcParams["figure.facecolor"] = "white"
plt.rcParams["axes.facecolor"]   = "white"
plt.rcParams["axes.edgecolor"]   = "#D9DEE5"
plt.rcParams["axes.titlesize"]   = 14
plt.rcParams["axes.titleweight"] = "bold"
plt.rcParams["axes.titlecolor"]  = NAVY
plt.rcParams["axes.labelcolor"]  = NAVY
plt.rcParams["axes.grid"]        = True
plt.rcParams["grid.color"]       = "#EDF0F4"
plt.rcParams["xtick.color"]      = "#5A6B7D"
plt.rcParams["ytick.color"]      = "#5A6B7D"
plt.rcParams["font.size"]        = 11

sns.set_palette(SEQ)

2  ·  加载数据并初步查看

在看清自己手里到底是什么之前,永远不要动手清洗

df = pd.read_csv("/kaggle/input/datasets/arjunsinghgangwar/ipl-20082026-matches-dataset/IPL_Matches_Data_2008_2026.csv")

print("Rows    :", df.shape[0])
print("Columns :", df.shape[1])
df.head(3)
Rows    : 1243
Columns : 31
event_name   season  match_number        date        city                                       venue                        team1                  team2                  toss_winner  \
0  Indian Premier League  2007/08           1.0  18-04-2008   Bangalore                       M Chinnaswamy Stadium  Royal Challengers Bangalore  Kolkata Knight Riders  Royal Challengers Bangalore   
1  Indian Premier League  2007/08           2.0  19-04-2008  Chandigarh  Punjab Cricket Association Stadium, Mohali              Kings XI Punjab    Chennai Super Kings          Chennai Super Kings   
2  Indian Premier League  2007/08           3.0  19-04-2008       Delhi                            Feroz Shah Kotla             Delhi Daredevils       Rajasthan Royals             Rajasthan Royals   

  toss_decision  team1_runs  team1_wickets  team2_runs  team2_wickets                 winner result_type  win_by_runs  win_by_wickets player_of_match      match_referee    umpire1         umpire2  \
0         field          82             10         222              3  Kolkata Knight Riders    complete          140               0     BB McCullum          J Srinath  Asad Rauf     RE Koertzen   
1           bat         207              4         240              5    Chennai Super Kings    complete           33               0      MEK Hussey  S Venkataraghavan  MR Benson      SL Shastri   
2           bat         132              1         129              8       Delhi Daredevils    complete            0               9     MF Maharoof       GR Viswanath  Aleem Dar  GA Pratapkumar   

   tv_umpire reserve_umpire match_type  overs_limit  balls_per_over gender team_type                                      team1_players                                      team2_players  
0  AM Saheba    VN Kulkarni        T20           20               6   male      club  R Dravid, W Jaffer, V Kohli, JH Kallis, CL Whi...  SC Ganguly, BB McCullum, RT Ponting, DJ Hussey...  
1  RB Tiffin    MSS Ranawat        T20           20               6   male      club  K Goel, JR Hopes, KC Sangakkara, Yuvraj Singh,...  PA Patel, ML Hayden, MEK Hussey, MS Dhoni, SK ...  
2  IL Howell        Unknown        T20           20               6   male      club  G Gambhir, V Sehwag, S Dhawan, MK Tiwary, KD K...  T Kohli, YK Pathan, SR Watson, M Kaif, DS Lehm...
df.info()
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 1243 entries, 0 to 1242
Data columns (total 31 columns):
 #   Column           Non-Null Count  Dtype  
---  ------           --------------  -----  
 0   event_name       1243 non-null   object 
 1   season           1243 non-null   object 
 2   match_number     1169 non-null   float64
 3   date             1243 non-null   object 
 4   city             1243 non-null   object 
 5   venue            1243 non-null   object 
 6   team1            1243 non-null   object 
 7   team2            1243 non-null   object 
 8   toss_winner      1243 non-null   object 
 9   toss_decision    1243 non-null   object 
 10  team1_runs       1243 non-null   int64  
 11  team1_wickets    1243 non-null   int64  
 12  team2_runs       1243 non-null   int64  
 13  team2_wickets    1243 non-null   int64  
 14  winner           1218 non-null   object 
 15  result_type      1243 non-null   object 
 16  win_by_runs      1243 non-null   int64  
 17  win_by_wickets   1243 non-null   int64  
 18  player_of_match  1234 non-null   object 
 19  match_referee    1243 non-null   object 
 20  umpire1          1243 non-null   object 
 21  umpire2          1243 non-null   object 
 22  tv_umpire        1243 non-null   object 
 23  reserve_umpire   1243 non-null   object 
 24  match_type       1243 non-null   object 
 25  overs_limit      1243 non-null   int64  
 26  balls_per_over   1243 non-null   int64  
 27  gender           1243 non-null   object 
 28  team_type        1243 non-null   object 
 29  team1_players    1243 non-null   object 
 30  team2_players    1243 non-null   object 
dtypes: float64(1), int64(8), object(22)
memory usage: 301.2+ KB

缺失值

只有四列存在缺失,而每一列的缺失最终都被证明是有意义的,而非损坏的数据。

missing = pd.DataFrame({
    "missing": df.isna().sum(),
    "percent": (df.isna().sum() / len(df) * 100).round(2)
})
missing = missing[missing["missing"] > 0].sort_values("missing", ascending=False)
missing
missing  percent
match_number          74     5.95
winner                25     2.01
player_of_match        9     0.72
# Why are they missing? Let's look at the two important ones.
print("Result types where 'winner' is missing:")
print(df.loc[df["winner"].isna(), "result_type"].value_counts())
print()
print("Seasons where 'match_number' is missing (matches per season):")
print(df.loc[df["match_number"].isna(), "season"].value_counts().head())
Result types where 'winner' is missing:
result_type
tie          16
no result     9
Name: count, dtype: int64

Seasons where 'match_number' is missing (matches per season):
season
2009/10    4
2012       4
2011       4
2019       4
2020/21    4
Name: count, dtype: int64
这些缺失究竟意味着什么

• winner — 在 16 场 平局(由 super over 决出胜负,而本文件并未记录)和 9 场 无结果(因雨)的比赛中缺失。这些是真正未分出胜负的比赛,而不是错误。
• match_number — 每个赛季恰好有 3–4 场比赛缺失。那些正是 季后赛:资格赛、淘汰赛和决赛没有编号。这是一个免费得来的特征,而不是问题。
• player_of_match — 在 9 场被取消的比赛中缺失。这是正确的行为。

这里没有任何内容需要插补。我们要做的是把这些缺失转化为信息。

单一取值的列

在这个数据集中,有些列完全不携带任何信息——每一行的取值都相同。

constant_cols = [c for c in df.columns if df[c].nunique(dropna=False) == 1]
for c in constant_cols:
    print(f"{c:<15} -> {df[c].unique()[0]}")
event_name      -> Indian Premier League
match_type      -> T20
overs_limit     -> 20
balls_per_over  -> 6
gender          -> male
team_type       -> club

3  ·  数据清洗

这个文件中的五个真实问题,逐一修复

原始体育数据几乎从不具备直接用于分析的条件。这个文件存在五个问题,其中任何一个如果被跳过,都会悄悄地污染分析结果:

  1. 赛季标签不一致 — 2007/08、2009/10 和 2020/21 与纯数字年份混排在一起。
  2. 球队曾更名 — Delhi Daredevils 改名为 Delhi Capitals;Kings XI Punjab 改名为 Punjab Kings;Royal Challengers Bangalore 改名为 Bengaluru。如果把它们当作不同的球队处理,就会把球队的历史拦腰截断。
  3. 场地名称重复 — Wankhede Stadium 和 Wankhede Stadium, Mumbai 其实是同一座球场。
  4. 51 场比赛的 city = "Unknown" — 但场地信息确切地告诉我们这些比赛是在哪里进行的。
  5. 日期是字符串,因此目前还无法进行任何基于时间的分析。

我们在副本上进行操作,这样原始的 df 保持原样、随时可比。

3.1  日期与干净的数值型赛季

data = df.copy()

# Dates -> real datetime
data["date"] = pd.to_datetime(data["date"], format="%d-%m-%Y")

# The season label is inconsistent ("2007/08"), but the first match of a season
# always falls in the correct calendar year. So take the year of the earliest match.
season_year = data.groupby("season")["date"].transform("min").dt.year
data["season_year"] = season_year

print(data.groupby("season")["season_year"].first().to_string())
season
2007/08    2008
2009       2009
2009/10    2010
2011       2011
2012       2012
2013       2013
2014       2014
2015       2015
2016       2016
2017       2017
2018       2018
2019       2019
2020/21    2020
2021       2021
2022       2022
2023       2023
2024       2024
2025       2025
2026       2026

3.2  合并更名后的球队

team_renames = {
    "Delhi Daredevils":            "Delhi Capitals",
    "Kings XI Punjab":             "Punjab Kings",
    "Royal Challengers Bangalore": "Royal Challengers Bengaluru",
    "Rising Pune Supergiants":     "Rising Pune Supergiant",
}

for col in ["team1", "team2", "toss_winner", "winner"]:
    data[col] = data[col].replace(team_renames)

print("Teams after merging renames:", data["team1"].nunique())
print()
print(pd.concat([data["team1"], data["team2"]]).value_counts().to_string())
Teams after merging renames: 15

Mumbai Indians                 291
Royal Challengers Bengaluru    286
Delhi Capitals                 281
Punjab Kings                   278
Kolkata Knight Riders          278
Chennai Super Kings            266
Rajasthan Royals               251
Sunrisers Hyderabad            211
Gujarat Titans                  77
Deccan Chargers                 75
Lucknow Super Giants            72
Pune Warriors                   46
Rising Pune Supergiant          30
Gujarat Lions                   30
Kochi Tuskers Kerala            14
一个判断性决定:Deccan Chargers 并未被合并进 Sunrisers Hyderabad。Hyderabad 特许球队于 2012 年被终止,2013 年售出了一支新的——不同的老板、不同的阵容、不同的球队。Gujarat Lions (2016–17) 和 Gujarat Titans (2022–) 同样彼此独立。更名会被合并;重新组建则不会。

3.3  简短球队代码,让图表更易读

short = {
    "Mumbai Indians": "MI", "Chennai Super Kings": "CSK", "Kolkata Knight Riders": "KKR",
    "Royal Challengers Bengaluru": "RCB", "Rajasthan Royals": "RR", "Delhi Capitals": "DC",
    "Sunrisers Hyderabad": "SRH", "Punjab Kings": "PBKS", "Gujarat Titans": "GT",
    "Lucknow Super Giants": "LSG", "Deccan Chargers": "DCH", "Pune Warriors": "PWI",
    "Gujarat Lions": "GL", "Rising Pune Supergiant": "RPS", "Kochi Tuskers Kerala": "KTK",
}

data["team1_short"]      = data["team1"].map(short)
data["team2_short"]      = data["team2"].map(short)
data["winner_short"]     = data["winner"].map(short)
data["toss_winner_short"] = data["toss_winner"].map(short)

data[["team1", "team1_short", "team2", "team2_short"]].head(3)
team1 team1_short                  team2 team2_short
0  Royal Challengers Bengaluru         RCB  Kolkata Knight Riders         KKR
1                 Punjab Kings        PBKS    Chennai Super Kings         CSK
2               Delhi Capitals          DC       Rajasthan Royals          RR

3.4  场馆去重并修复城市

# Most duplicates are just a city appended after a comma -> keep the part before the first comma
data["venue"] = data["venue"].str.split(",").str[0].str.strip()

# The remaining few are spelling variants or genuine stadium renamings
venue_fixes = {
    "M.Chinnaswamy Stadium":              "M Chinnaswamy Stadium",
    "Punjab Cricket Association Stadium":  "Punjab Cricket Association IS Bindra Stadium",
    "Zayed Cricket Stadium":               "Sheikh Zayed Stadium",
    "Sardar Patel Stadium":                "Narendra Modi Stadium",
    "Feroz Shah Kotla":                    "Arun Jaitley Stadium",
    "Subrata Roy Sahara Stadium":          "Maharashtra Cricket Association Stadium",
}
data["venue"] = data["venue"].replace(venue_fixes)

print("Unique venues:", df["venue"].nunique(), "->", data["venue"].nunique())
Unique venues: 60 -> 36
# Cities: merge the Bangalore/Bengaluru spelling and fill the 'Unknown' entries from the venue
data["city"] = data["city"].replace({"Bangalore": "Bengaluru", "Navi Mumbai": "Mumbai"})

city_from_venue = {
    "Dubai International Cricket Stadium": "Dubai",
    "Sharjah Cricket Stadium":             "Sharjah",
}
unknown = data["city"] == "Unknown"
data.loc[unknown, "city"] = data.loc[unknown, "venue"].map(city_from_venue)

print("Cities still unknown:", data["city"].isna().sum())
print("Unique cities:", df["city"].nunique(), "->", data["city"].nunique())
Cities still unknown: 0
Unique cities: 38 -> 35

4  ·  特征工程

这个数据集中最有价值的列,恰恰是那些不在文件里的列

原始文件记录了 team1_runs 和 team2_runs——但 team1 不一定是先攻的那支球队。不修正这一点,所有"第一局得分"分析都会沦为无稽之谈。

掷币结果告诉了我们击球顺序:

  • 掷币获胜方选择 bat → 掷币获胜方先攻
  • 掷币获胜方选择 field → 另一支球队先攻

凭借这一条规则,我们就能解锁各局得分、追分目标、胜负分差,以及比赛是靠守分获胜还是靠追分获胜。

# Who batted first?
toss_is_team1 = data["toss_winner"] == data["team1"]
chose_bat     = data["toss_decision"] == "bat"

data["bat_first"]  = np.where(chose_bat == toss_is_team1, data["team1"], data["team2"])
data["bat_second"] = np.where(data["bat_first"] == data["team1"], data["team2"], data["team1"])

# Quick sanity check against the very first match of IPL history
data.loc[[0], ["team1", "team2", "toss_winner", "toss_decision",
               "team1_runs", "team2_runs", "bat_first", "winner"]]
team1                  team2                  toss_winner toss_decision  team1_runs  team2_runs              bat_first                 winner
0  Royal Challengers Bengaluru  Kolkata Knight Riders  Royal Challengers Bengaluru         field          82         222  Kolkata Knight Riders  Kolkata Knight Riders

2008 年第 1 场比赛:RCB 赢得掷币并选择先守,于是 Kolkata 先攻并拿下 222 分。bat_first 正确地返回了 Kolkata Knight Riders。这条规则奏效。

first_is_team1 = data["bat_first"] == data["team1"]

data["first_innings_runs"]  = np.where(first_is_team1, data["team1_runs"],    data["team2_runs"])
data["first_innings_wkts"]  = np.where(first_is_team1, data["team1_wickets"], data["team2_wickets"])
data["second_innings_runs"] = np.where(first_is_team1, data["team2_runs"],    data["team1_runs"])
data["second_innings_wkts"] = np.where(first_is_team1, data["team2_wickets"], data["team1_wickets"])

data["target"] = data["first_innings_runs"] + 1
data.loc[:, ["bat_first", "first_innings_runs", "bat_second", "second_innings_runs", "target"]].head()
bat_first  first_innings_runs                   bat_second  second_innings_runs  target
0  Kolkata Knight Riders                 222  Royal Challengers Bengaluru                   82     223
1    Chennai Super Kings                 240                 Punjab Kings                  207     241
2       Rajasthan Royals                 129               Delhi Capitals                  132     130
3        Deccan Chargers                 110        Kolkata Knight Riders                  112     111
4         Mumbai Indians                 165  Royal Challengers Bengaluru                  166     166
# How was the match won - defending a total, or chasing one?
data["result_side"] = np.select(
    [data["winner"] == data["bat_first"], data["winner"] == data["bat_second"]],
    ["Batting first", "Chasing"],
    default="No result / Tie"
)

# Stage of the tournament: unnumbered matches are the playoffs
data["stage"] = np.where(data["match_number"].isna(), "Playoff", "League")

# A single margin column, plus a decided-matches flag we'll reuse constantly
data["margin"] = np.where(data["win_by_runs"] > 0,
                          data["win_by_runs"].astype(str) + " runs",
                          data["win_by_wickets"].astype(str) + " wickets")
data["decided"] = data["winner"].notna()

data["result_side"].value_counts()
result_side
Chasing            666
Batting first      552
No result / Tie     25
Name: count, dtype: int64

分析表

keep = ["season_year", "date", "city", "venue", "stage",
        "team1", "team2", "team1_short", "team2_short",
        "toss_winner", "toss_decision", "bat_first", "bat_second",
        "first_innings_runs", "first_innings_wkts",
        "second_innings_runs", "second_innings_wkts", "target",
        "winner", "winner_short", "result_type", "result_side",
        "win_by_runs", "win_by_wickets", "margin", "decided",
        "player_of_match"]

ipl = data[keep].copy()

print("Clean analysis table:", ipl.shape)
print("Decided matches     :", ipl["decided"].sum())
print("Seasons             :", ipl["season_year"].min(), "-", ipl["season_year"].max())
ipl.head()
Clean analysis table: (1243, 27)
Decided matches     : 1218
Seasons             : 2008 - 2026
season_year       date        city                                         venue   stage                        team1                        team2 team1_short team2_short  \
0         2008 2008-04-18   Bengaluru                         M Chinnaswamy Stadium  League  Royal Challengers Bengaluru        Kolkata Knight Riders         RCB         KKR   
1         2008 2008-04-19  Chandigarh  Punjab Cricket Association IS Bindra Stadium  League                 Punjab Kings          Chennai Super Kings        PBKS         CSK   
2         2008 2008-04-19       Delhi                          Arun Jaitley Stadium  League               Delhi Capitals             Rajasthan Royals          DC          RR   
3         2008 2008-04-20     Kolkata                                  Eden Gardens  League        Kolkata Knight Riders              Deccan Chargers         KKR         DCH   
4         2008 2008-04-20      Mumbai                              Wankhede Stadium  League               Mumbai Indians  Royal Challengers Bengaluru          MI         RCB   

                   toss_winner toss_decision              bat_first                   bat_second  first_innings_runs  first_innings_wkts  second_innings_runs  second_innings_wkts  target  \
0  Royal Challengers Bengaluru         field  Kolkata Knight Riders  Royal Challengers Bengaluru                 222                   3                   82                   10     223   
1          Chennai Super Kings           bat    Chennai Super Kings                 Punjab Kings                 240                   5                  207                    4     241   
2             Rajasthan Royals           bat       Rajasthan Royals               Delhi Capitals                 129                   8                  132                    1     130   
3              Deccan Chargers           bat        Deccan Chargers        Kolkata Knight Riders                 110                  10                  112                    5     111   
4               Mumbai Indians           bat         Mumbai Indians  Royal Challengers Bengaluru                 165                   7                  166                    5     166   

                        winner winner_short result_type    result_side  win_by_runs  win_by_wickets     margin  decided player_of_match  
0        Kolkata Knight Riders          KKR    complete  Batting first          140               0   140 runs     True     BB McCullum  
1          Chennai Super Kings          CSK    complete  Batting first           33               0    33 runs     True      MEK Hussey  
2               Delhi Capitals           DC    complete        Chasing            0               9  9 wickets     True     MF Maharoof  
3        Kolkata Knight Riders          KKR    complete        Chasing            0               5  5 wickets     True       DJ Hussey  
4  Royal Challengers Bengaluru          RCB    complete        Chasing            0               5  5 wickets     True      MV Boucher

5  ·  联盟的历年变迁

问题 1:IPL 真的变成了一个得分更高的赛事吗?

per_season = ipl.groupby("season_year").agg(
    matches      = ("date", "count"),
    avg_first    = ("first_innings_runs", "mean"),
    avg_second   = ("second_innings_runs", "mean"),
    highest      = ("first_innings_runs", "max"),
    venues       = ("venue", "nunique"),
).round(1)

per_season.head(19)
matches  avg_first  avg_second  highest  venues
season_year                                                 
2008              58      161.0       148.3      240       9
2009              57      150.6       136.3      211       8
2010              60      165.0       149.8      246      12
2011              73      152.4       137.4      232      12
2012              74      157.5       145.9      222      12
2013              76      156.2       141.2      263      12
2014              60      163.2       152.3      231      13
2015              59      166.4       144.7      235      13
2016              60      162.6       151.8      248      11
2017              59      165.9       152.5      230      10
2018              60      172.5       159.2      245      10
2019              60      167.0       156.9      232       9
2020              60      170.0       153.6      228       3
2021              60      159.4       151.2      235       7
2022              74      171.1       158.5      222       6
2023              74      182.7       164.4      257      12
2024              71      189.6       176.2      287      13
2025              74      189.0       169.5      286      13
2026              74      193.1       177.9      264      13
fig, ax = plt.subplots(figsize=(12, 5))
bars = ax.bar(per_season.index, per_season["matches"], color=BLUE, edgecolor="white")

# Highlight the two pandemic-affected seasons
for i, yr in enumerate(per_season.index):
    if yr in (2020, 2021):
        bars[i].set_color(ORANGE)

for x, y in zip(per_season.index, per_season["matches"]):
    ax.text(x, y + 1, int(y), ha="center", fontsize=9, color=NAVY)

ax.set_title("Matches played each IPL season")
ax.set_xlabel("Season")
ax.set_ylabel("Matches")
ax.set_xticks(per_season.index)
ax.tick_params(axis="x", rotation=45)
ax.set_ylim(0, 90)
plt.tight_layout()
plt.show()

图1

联盟的规模经历了三次清晰的跃升:头一个十年的大部分时间里是 58–60 场比赛的赛事;当参赛队伍扩军到 9–10 支时(2011–13)跃升至 70+ 场;而在 Gujarat Titans 和 Lucknow Super Giants 加入后,从 2022 年起永久固定为 74 场比赛。橙色柱形标出了在 UAE 举行的 COVID 赛季。

fig, ax = plt.subplots(figsize=(12, 5.5))

ax.plot(per_season.index, per_season["avg_first"], marker="o", lw=2.5,
        color=NAVY, label="1st innings average")
ax.plot(per_season.index, per_season["avg_second"], marker="s", lw=2.5,
        color=ORANGE, label="2nd innings average")
ax.fill_between(per_season.index, per_season["avg_second"], per_season["avg_first"],
                color=BLUE, alpha=0.10)

ax.set_title("Average innings score by season")
ax.set_xlabel("Season")
ax.set_ylabel("Runs")
ax.set_xticks(per_season.index)
ax.tick_params(axis="x", rotation=45)
ax.legend(frameon=False)
plt.tight_layout()
plt.show()

图2

# How often do teams post a really big total?
big = ipl[ipl["first_innings_runs"] >= 200].groupby("season_year").size()
big = big.reindex(per_season.index, fill_value=0)
rate = (big / per_season["matches"] * 100).round(1)

fig, ax = plt.subplots(figsize=(12, 5))
ax.bar(rate.index, rate.values, color=GOLD, edgecolor="white")
for x, y in zip(rate.index, rate.values):
    ax.text(x, y + 0.4, f"{y:.0f}%", ha="center", fontsize=9, color=NAVY)

ax.set_title("Share of matches with a 200+ first innings total")
ax.set_xlabel("Season")
ax.set_ylabel("% of matches")
ax.set_xticks(rate.index)
ax.tick_params(axis="x", rotation=45)
plt.tight_layout()
plt.show()

图3

发现 1 — 得分爆发是真实存在的,而且是最近才出现的。
十多年来,第一局平均得分一直徘徊在 160 附近的狭窄区间内。从 2023 年起,它们突破了这一区间。200+ 的比率是更敏锐的信号:过去打出 200 分是整个赛事才出现一次的事,如今已是常态。impact-player 规则、更平坦的球道,以及伴随 T20 长大的一代击球手,都在同一时间到来。

6  ·  各特许球队

问题 2:IPL 历史上真正最成功的球队是谁?

played = pd.concat([ipl["team1"], ipl["team2"]]).value_counts()
won    = ipl["winner"].value_counts()

teams = pd.DataFrame({"played": played, "won": won}).fillna(0)
teams["won"]     = teams["won"].astype(int)
teams["lost"]    = teams["played"] - teams["won"]
teams["win_pct"] = (teams["won"] / teams["played"] * 100).round(1)
teams["code"]    = teams.index.map(short)
teams = teams.sort_values("win_pct", ascending=False)

teams[["code", "played", "won", "lost", "win_pct"]]
code  played  won  lost  win_pct
Gujarat Titans                 GT      77   47    30     61.0
Chennai Super Kings           CSK     266  148   118     55.6
Mumbai Indians                 MI     291  155   136     53.3
Kolkata Knight Riders         KKR     278  140   138     50.4
Rising Pune Supergiant        RPS      30   15    15     50.0
Royal Challengers Bengaluru   RCB     286  143   143     50.0
Rajasthan Royals               RR     251  123   128     49.0
Sunrisers Hyderabad           SRH     211  102   109     48.3
Lucknow Super Giants          LSG      72   34    38     47.2
Punjab Kings                 PBKS     278  126   152     45.3
Delhi Capitals                 DC     281  125   156     44.5
Gujarat Lions                  GL      30   13    17     43.3
Kochi Tuskers Kerala          KTK      14    6     8     42.9
Deccan Chargers               DCH      75   29    46     38.7
Pune Warriors                 PWI      46   12    34     26.1
core = teams[teams["played"] >= 50].sort_values("win_pct")

fig, ax = plt.subplots(figsize=(11, 6))
colors = [ORANGE if v >= 50 else BLUE for v in core["win_pct"]]
ax.barh(core["code"], core["win_pct"], color=colors, edgecolor="white")

ax.axvline(50, color=RED, ls="--", lw=1.5)
ax.text(50.4, -0.4, "50%", color=RED, fontsize=10)

for code_, v, p in zip(core["code"], core["win_pct"], core["played"]):
    ax.text(v + 0.6, code_, f"{v}%  ({int(p)} matches)", va="center", fontsize=9, color=NAVY)

ax.set_title("Win percentage — teams with 50+ matches played")
ax.set_xlabel("Win %")
ax.set_xlim(0, 75)
plt.tight_layout()
plt.show()

图4

冠军头衔:特许球队真正在乎的唯一数字

# The final is the last match played in each season
finals_idx = ipl.groupby("season_year")["date"].idxmax()
finals = ipl.loc[finals_idx, ["season_year", "bat_first", "bat_second",
                              "first_innings_runs", "second_innings_runs",
                              "winner", "margin", "venue", "player_of_match"]]
finals = finals.rename(columns={"winner": "champion"}).set_index("season_year")
finals[["champion", "margin", "venue", "player_of_match"]]
champion     margin                                venue player_of_match
season_year                                                                                             
2008                    Rajasthan Royals  3 wickets           Dr DY Patil Sports Academy       YK Pathan
2009                     Deccan Chargers     6 runs                New Wanderers Stadium        A Kumble
2010                 Chennai Super Kings    22 runs           Dr DY Patil Sports Academy        SK Raina
2011                 Chennai Super Kings    58 runs               MA Chidambaram Stadium         M Vijay
2012               Kolkata Knight Riders  5 wickets               MA Chidambaram Stadium        MS Bisla
2013                      Mumbai Indians    23 runs                         Eden Gardens      KA Pollard
2014               Kolkata Knight Riders  3 wickets                M Chinnaswamy Stadium       MK Pandey
2015                      Mumbai Indians    41 runs                         Eden Gardens       RG Sharma
2016                 Sunrisers Hyderabad     8 runs                M Chinnaswamy Stadium     BCJ Cutting
2017                      Mumbai Indians     1 runs   Rajiv Gandhi International Stadium       KH Pandya
2018                 Chennai Super Kings  8 wickets                     Wankhede Stadium       SR Watson
2019                      Mumbai Indians     1 runs   Rajiv Gandhi International Stadium       JJ Bumrah
2020                      Mumbai Indians  5 wickets  Dubai International Cricket Stadium        TA Boult
2021                 Chennai Super Kings    27 runs  Dubai International Cricket Stadium    F du Plessis
2022                      Gujarat Titans  7 wickets                Narendra Modi Stadium       HH Pandya
2023                 Chennai Super Kings  5 wickets                Narendra Modi Stadium       DP Conway
2024               Kolkata Knight Riders  8 wickets               MA Chidambaram Stadium        MA Starc
2025         Royal Challengers Bengaluru     6 runs                Narendra Modi Stadium       KH Pandya
2026         Royal Challengers Bengaluru  5 wickets                Narendra Modi Stadium         V Kohli
titles = finals["champion"].value_counts()
runners = finals.apply(
    lambda r: r["bat_second"] if r["champion"] == r["bat_first"] else r["bat_first"], axis=1
).value_counts()

honours = pd.DataFrame({"titles": titles, "runner_up": runners}).fillna(0).astype(int)
honours["finals"] = honours["titles"] + honours["runner_up"]
honours = honours.sort_values(["titles", "finals"], ascending=False)
honours
titles  runner_up  finals
Chennai Super Kings               5          5      10
Mumbai Indians                    5          1       6
Kolkata Knight Riders             3          1       4
Royal Challengers Bengaluru       2          3       5
Gujarat Titans                    1          2       3
Sunrisers Hyderabad               1          2       3
Rajasthan Royals                  1          1       2
Deccan Chargers                   1          0       1
Punjab Kings                      0          2       2
Delhi Capitals                    0          1       1
Rising Pune Supergiant            0          1       1
h = honours.sort_values("titles")
idx = np.arange(len(h))

fig, ax = plt.subplots(figsize=(11, 6))
ax.barh(idx, h["titles"],    color=ORANGE, edgecolor="white", label="Titles")
ax.barh(idx, h["runner_up"], left=h["titles"], color=GREY, edgecolor="white", label="Runner-up")

ax.set_yticks(idx)
ax.set_yticklabels([short.get(t, t) for t in h.index])
ax.set_title("Finals record by franchise")
ax.set_xlabel("Number of finals reached")
ax.legend(frameon=False, loc="lower right")
plt.tight_layout()
plt.show()

图5

发现 2 — 稳定性与奖杯是两回事。
Chennai 和 Mumbai 在两份榜单上都位居榜首,这正是它们被视为联盟王朝球队的原因。但请注意那些胜率很高、奖杯柜却空空如也的球队——高常规赛胜率能把你送进季后赛,却赢不下那一场真正算数的比赛。

常规赛表现 vs 季后赛表现

playoffs = ipl[(ipl["stage"] == "Playoff") & ipl["decided"]]
po_played = pd.concat([playoffs["team1"], playoffs["team2"]]).value_counts()
po_won    = playoffs["winner"].value_counts()

po = pd.DataFrame({"playoff_games": po_played, "playoff_wins": po_won}).fillna(0).astype(int)
po["playoff_win_pct"] = (po["playoff_wins"] / po["playoff_games"] * 100).round(1)
po = po.join(teams["win_pct"].rename("league_win_pct"))
po["clutch_gap"] = (po["playoff_win_pct"] - po["league_win_pct"]).round(1)
po.sort_values("playoff_games", ascending=False).head(10)
playoff_games  playoff_wins  playoff_win_pct  league_win_pct  clutch_gap
Chennai Super Kings                     26            17             65.4            55.6         9.8
Mumbai Indians                          22            14             63.6            53.3        10.3
Royal Challengers Bengaluru             20            10             50.0            50.0         0.0
Kolkata Knight Riders                   15            10             66.7            50.4        16.3
Sunrisers Hyderabad                     15             6             40.0            48.3        -8.3
Rajasthan Royals                        13             6             46.2            49.0        -2.8
Delhi Capitals                          11             2             18.2            44.5       -26.3
Gujarat Titans                           9             4             44.4            61.0       -16.6
Punjab Kings                             7             2             28.6            45.3       -16.7
Deccan Chargers                          4             2             50.0            38.7        11.3

clutch_gap 是球队季后赛胜率与整体胜率之差。正数意味着球队在淘汰赛到来时会提升水准;负数则意味着它会掉链子。

7  ·  掷币

问题 3:掷硬币真的能决定什么吗?

print(ipl["toss_decision"].value_counts())
print()
print((ipl["toss_decision"].value_counts(normalize=True) * 100).round(1))
toss_decision
field    825
bat      418
Name: count, dtype: int64

toss_decision
field    66.4
bat      33.6
Name: proportion, dtype: float64
toss_trend = pd.crosstab(ipl["season_year"], ipl["toss_decision"], normalize="index") * 100

fig, ax = plt.subplots(figsize=(12, 5.5))
ax.stackplot(toss_trend.index, toss_trend["field"], toss_trend["bat"],
             colors=[BLUE, GOLD], labels=["Chose to field", "Chose to bat"], alpha=0.9)
ax.axhline(50, color="white", ls="--", lw=1.5)

ax.set_title("Toss decision by season (% of matches)")
ax.set_xlabel("Season")
ax.set_ylabel("% of tosses")
ax.set_xlim(toss_trend.index.min(), toss_trend.index.max())
ax.set_ylim(0, 100)
ax.set_xticks(toss_trend.index)
ax.tick_params(axis="x", rotation=45)
ax.legend(loc="lower left", frameon=False)
plt.tight_layout()
plt.show()

图6

队长们的选择起初大致五五开,随后急剧倒向先守。从 2015 年起,每四场比赛中就有三场默认选择追分——晚间比赛中的露水,以及提前知道目标分数的那份踏实感,都把天平推向同一个方向。

decided = ipl[ipl["decided"]].copy()
decided["toss_won_match"] = decided["toss_winner"] == decided["winner"]

toss_advantage = decided["toss_won_match"].mean() * 100
print(f"Toss winner also won the match: {toss_advantage:.1f}% of {len(decided)} decided matches")
Toss winner also won the match: 51.6% of 1218 decided matches
by_season = decided.groupby("season_year")["toss_won_match"].mean() * 100

fig, ax = plt.subplots(figsize=(12, 5.5))
colors = [GREEN if v >= 50 else RED for v in by_season.values]
ax.bar(by_season.index, by_season.values, color=colors, edgecolor="white")
ax.axhline(50, color=NAVY, ls="--", lw=2)
ax.axhline(toss_advantage, color=ORANGE, lw=2,
           label=f"All-time average {toss_advantage:.1f}%")

ax.set_title("How often did the toss winner win the match?")
ax.set_xlabel("Season")
ax.set_ylabel("% of matches")
ax.set_ylim(0, 80)
ax.set_xticks(by_season.index)
ax.tick_params(axis="x", rotation=45)
ax.legend(frameon=False)
plt.tight_layout()
plt.show()

图7

发现 3 — 掷币不仅在机制上像抛硬币,其实际效果也近乎如此。
在近 1,220 场分出胜负的比赛中,掷币获胜方的获胜次数只是略微过半。各个赛季在 50% 上下剧烈波动,而这正是随机噪声该有的样子。我们将在第 10 节用卡方检验正式证实这一点。

8  ·  先攻 vs 追分

问题 4:追分真的更容易吗?

sides = decided["result_side"].value_counts()
print(sides)
print()
print((sides / sides.sum() * 100).round(1))
result_side
Chasing          666
Batting first    552
Name: count, dtype: int64

result_side
Chasing          54.7
Batting first    45.3
Name: count, dtype: float64
chase_trend = pd.crosstab(decided["season_year"], decided["result_side"], normalize="index") * 100

fig, ax = plt.subplots(figsize=(12, 5.5))
ax.plot(chase_trend.index, chase_trend["Chasing"], marker="o", lw=2.5,
        color=ORANGE, label="Chasing team won")
ax.plot(chase_trend.index, chase_trend["Batting first"], marker="s", lw=2.5,
        color=NAVY, label="Defending team won")
ax.axhline(50, color=GREY, ls="--", lw=1.5)

ax.set_title("Win rate: chasing vs defending, by season")
ax.set_xlabel("Season")
ax.set_ylabel("% of decided matches")
ax.set_xticks(chase_trend.index)
ax.tick_params(axis="x", rotation=45)
ax.legend(frameon=False)
plt.tight_layout()
plt.show()

图8

目标分数的大小会改变答案吗?

bins   = [0, 140, 160, 180, 200, 400]
labels = ["Under 140", "140-159", "160-179", "180-199", "200+"]
decided["target_band"] = pd.cut(decided["first_innings_runs"], bins=bins, labels=labels)

band = decided.groupby("target_band", observed=True).agg(
    matches      = ("winner", "count"),
    chase_win_pct = ("result_side", lambda s: (s == "Chasing").mean() * 100)
).round(1)
band
matches  chase_win_pct
target_band                        
Under 140        225           82.2
140-159          255           71.8
160-179          311           52.4
180-199          222           39.6
200+             205           22.9
fig, ax = plt.subplots(figsize=(11, 5.5))
bars = ax.bar(band.index.astype(str), band["chase_win_pct"], color=BLUE, edgecolor="white")
ax.axhline(50, color=RED, ls="--", lw=1.5)

for b, v, n in zip(bars, band["chase_win_pct"], band["matches"]):
    ax.text(b.get_x() + b.get_width()/2, v + 1.5, f"{v:.0f}%", ha="center", color=NAVY, fontweight="bold")
    ax.text(b.get_x() + b.get_width()/2, 3, f"n={n}", ha="center", color="white", fontsize=9)

ax.set_title("Chasing success rate by size of the first innings total")
ax.set_xlabel("First innings score")
ax.set_ylabel("Chases won (%)")
ax.set_ylim(0, 85)
plt.tight_layout()
plt.show()

图9

发现 4 — 追分更容易,但仅限于一定程度。
低于 160 分时,追分一方明显占优。这一优势在 160–179 区间急剧衰减,到 180 分时已荡然无存, 而一旦有球队拿到 200+ 的总分,天平便会果断倒向防守一方。那些不假思索便选择先投球的队长 在大多数比赛中做出了正确的决定——只是在那些球道平坦如公路的比赛中并非如此。

9  ·  场地与条件

问题 5:哪些场地是击球手的乐园,哪些更有利于追分?

venue_stats = ipl.groupby("venue").agg(
    matches    = ("date", "count"),
    avg_first  = ("first_innings_runs", "mean"),
    highest    = ("first_innings_runs", "max"),
).round(1)

chase_pct = decided.groupby("venue")["result_side"].apply(lambda s: (s == "Chasing").mean() * 100)
venue_stats["chase_win_pct"] = chase_pct.round(1)

top_venues = venue_stats[venue_stats["matches"] >= 25].sort_values("matches", ascending=False)
top_venues
matches  avg_first  highest  chase_win_pct
venue                                                                                         
Wankhede Stadium                                        132      173.2      243           55.0
Eden Gardens                                            107      168.6      261           57.1
M Chinnaswamy Stadium                                   104      174.2      287           56.0
Arun Jaitley Stadium                                    104      172.0      278           53.5
MA Chidambaram Stadium                                   98      165.7      246           46.9
Rajiv Gandhi International Stadium                       90      168.5      286           54.5
Sawai Mansingh Stadium                                   68      169.1      229           64.7
Punjab Cricket Association IS Bindra Stadium             61      168.3      257           55.7
Narendra Modi Stadium                                    53      180.8      243           50.0
Maharashtra Cricket Association Stadium                  51      162.4      211           45.1
Dubai International Cricket Stadium                      46      164.4      219           51.2
Dr DY Patil Sports Academy                               37      159.6      216           54.1
Sheikh Zayed Stadium                                     37      159.3      235           60.0
Bharat Ratna Shri Atal Bihari Vajpayee Ekana Cr...       29      174.9      235           59.3
Sharjah Cricket Stadium                                  28      159.0      228           64.3
Brabourne Stadium                                        27      178.5      217           48.1
v = top_venues.sort_values("matches")
labels = [name if len(name) <= 30 else name[:28] + "..." for name in v.index]

fig, ax = plt.subplots(figsize=(11, 6.5))
ax.barh(labels, v["matches"], color=NAVY, edgecolor="white")
for lab, n in zip(labels, v["matches"]):
    ax.text(n + 1, lab, int(n), va="center", fontsize=9, color=NAVY)

ax.set_title("Busiest IPL grounds (25+ matches)")
ax.set_xlabel("Matches hosted")
plt.tight_layout()
plt.show()

图10

v = top_venues.sort_values("avg_first")

fig, axes = plt.subplots(1, 2, figsize=(14, 6.5), sharey=True)
labels = [name if len(name) <= 28 else name[:26] + "..." for name in v.index]

axes[0].barh(labels, v["avg_first"], color=GOLD, edgecolor="white")
axes[0].set_title("Average first innings score")
axes[0].set_xlim(130, v["avg_first"].max() + 12)
for lab, x in zip(labels, v["avg_first"]):
    axes[0].text(x + 1, lab, f"{x:.0f}", va="center", fontsize=9, color=NAVY)

colors = [GREEN if x >= 50 else RED for x in v["chase_win_pct"]]
axes[1].barh(labels, v["chase_win_pct"], color=colors, edgecolor="white")
axes[1].axvline(50, color=NAVY, ls="--", lw=1.5)
axes[1].set_title("Chasing team win rate (%)")
axes[1].set_xlim(0, 75)
for lab, x in zip(labels, v["chase_win_pct"]):
    axes[1].text(x + 1, lab, f"{x:.0f}%", va="center", fontsize=9, color=NAVY)

plt.tight_layout()
plt.show()

图11

右侧的绿色条形代表追分通常能成功的场地;红色条形代表防守总分更有效的场地。把两个面板结合起来看才是关键——平均得分高且条形为红色的场地,正是适合先击球并拿下高分的球场。

10  ·  分差、大胜与险胜

IPL 比赛究竟有多胶着?

run_wins    = decided[decided["win_by_runs"] > 0]
wicket_wins = decided[decided["win_by_wickets"] > 0]

fig, axes = plt.subplots(1, 2, figsize=(13.5, 5))

axes[0].hist(run_wins["win_by_runs"], bins=25, color=NAVY, edgecolor="white")
axes[0].axvline(run_wins["win_by_runs"].median(), color=ORANGE, lw=2.5,
                label=f"Median {run_wins['win_by_runs'].median():.0f} runs")
axes[0].set_title(f"Wins while defending  (n={len(run_wins)})")
axes[0].set_xlabel("Margin (runs)")
axes[0].set_ylabel("Matches")
axes[0].legend(frameon=False)

axes[1].hist(wicket_wins["win_by_wickets"], bins=np.arange(0.5, 11.5, 1),
             color=BLUE, edgecolor="white")
axes[1].axvline(wicket_wins["win_by_wickets"].median(), color=ORANGE, lw=2.5,
                label=f"Median {wicket_wins['win_by_wickets'].median():.0f} wickets")
axes[1].set_title(f"Wins while chasing  (n={len(wicket_wins)})")
axes[1].set_xlabel("Margin (wickets)")
axes[1].set_xticks(range(1, 11))
axes[1].legend(frameon=False)

plt.tight_layout()
plt.show()

图12

# The five most one-sided results in each direction
print("BIGGEST WINS BY RUNS")
cols = ["season_year", "bat_first", "first_innings_runs", "bat_second", "second_innings_runs", "margin"]
print(decided.nlargest(5, "win_by_runs")[cols].to_string(index=False))
print()
print("BIGGEST WINS BY WICKETS (fewest balls to spare not in this file, so ties broken by target)")
print(decided[decided["win_by_wickets"] == 10].nlargest(5, "target")[cols].to_string(index=False))
BIGGEST WINS BY RUNS
 season_year                   bat_first  first_innings_runs                  bat_second  second_innings_runs   margin
        2017              Mumbai Indians                 212              Delhi Capitals                   66 146 runs
        2016 Royal Challengers Bengaluru                 248               Gujarat Lions                  104 144 runs
        2008       Kolkata Knight Riders                 222 Royal Challengers Bengaluru                   82 140 runs
        2015 Royal Challengers Bengaluru                 226                Punjab Kings                   88 138 runs
        2013 Royal Challengers Bengaluru                 263               Pune Warriors                  133 130 runs

BIGGEST WINS BY WICKETS (fewest balls to spare not in this file, so ties broken by target)
 season_year            bat_first  first_innings_runs                  bat_second  second_innings_runs     margin
        2025       Delhi Capitals                 199              Gujarat Titans                  205 10 wickets
        2017        Gujarat Lions                 183       Kolkata Knight Riders                  184 10 wickets
        2020         Punjab Kings                 178         Chennai Super Kings                  181 10 wickets
        2021     Rajasthan Royals                 177 Royal Challengers Bengaluru                  181 10 wickets
        2024 Lucknow Super Giants                 165         Sunrisers Hyderabad                  167 10 wickets
# The nail-biters
print(f"Matches won by 1-3 runs: {(decided['win_by_runs'].between(1,3)).sum()}")
print(f"Tied matches (super over): {(ipl['result_type'] == 'tie').sum()}")
print(f"Abandoned / no result:     {(ipl['result_type'] == 'no result').sum()}")
print()
decided[decided["win_by_runs"] == 1][["season_year", "bat_first", "bat_second",
                                      "first_innings_runs", "second_innings_runs", "winner"]]
Matches won by 1-3 runs: 37
Tied matches (super over): 16
Abandoned / no result:     9
season_year                    bat_first                   bat_second  first_innings_runs  second_innings_runs                       winner
44           2008                 Punjab Kings               Mumbai Indians                 189                  188                 Punjab Kings
104          2009                 Punjab Kings              Deccan Chargers                 134                  133                 Punjab Kings
284          2012               Delhi Capitals             Rajasthan Royals                 152                  151               Delhi Capitals
290          2012               Mumbai Indians                Pune Warriors                 120                  119               Mumbai Indians
459          2015          Chennai Super Kings               Delhi Capitals                 150                  149          Chennai Super Kings
539          2016                Gujarat Lions               Delhi Capitals                 172                  171                Gujarat Lions
555          2016  Royal Challengers Bengaluru                 Punjab Kings                 175                  174  Royal Challengers Bengaluru
635          2017               Mumbai Indians       Rising Pune Supergiant                 129                  128               Mumbai Indians
734          2019  Royal Challengers Bengaluru          Chennai Super Kings                 161                  160  Royal Challengers Bengaluru
755          2019               Mumbai Indians          Chennai Super Kings                 149                  148               Mumbai Indians
837          2021  Royal Challengers Bengaluru               Delhi Capitals                 171                  170  Royal Challengers Bengaluru
1017         2023         Lucknow Super Giants        Kolkata Knight Riders                 176                  175         Lucknow Super Giants
1059         2024        Kolkata Knight Riders  Royal Challengers Bengaluru                 222                  221        Kolkata Knight Riders
1073         2024          Sunrisers Hyderabad             Rajasthan Royals                 201                  200          Sunrisers Hyderabad
1147         2025        Kolkata Knight Riders             Rajasthan Royals                 206                  205        Kolkata Knight Riders
1182         2026               Gujarat Titans               Delhi Capitals                 210                  209               Gujarat Titans

11  ·  正面交锋

哪些宿敌对决实际上是一边倒的?

big_teams = teams[teams["played"] >= 150].index.tolist()
h2h = pd.DataFrame(0.0, index=big_teams, columns=big_teams)

for a in big_teams:
    for b in big_teams:
        if a == b:
            h2h.loc[a, b] = np.nan
            continue
        games = decided[((decided["team1"] == a) & (decided["team2"] == b)) |
                        ((decided["team1"] == b) & (decided["team2"] == a))]
        if len(games) > 0:
            h2h.loc[a, b] = round((games["winner"] == a).mean() * 100, 1)

h2h.index   = [short[t] for t in h2h.index]
h2h.columns = [short[t] for t in h2h.columns]
h2h
CSK    MI   KKR   RCB    RR   SRH  PBKS    DC
CSK    NaN  48.8  65.6  60.0  50.0  62.5  50.0  63.6
MI    51.2   NaN  67.6  54.3  50.0  56.0  51.4  55.3
KKR   34.4  32.4   NaN  55.6  58.6  64.5  61.8  57.1
RCB   40.0  45.7  44.4   NaN  53.1  46.2  52.6  60.6
RR    50.0  50.0  41.4  46.9   NaN  41.7  60.0  48.4
SRH   37.5  44.0  35.5  53.8  58.3   NaN  69.2  56.0
PBKS  50.0  48.6  38.2  47.4  40.0  30.8   NaN  51.4
DC    36.4  44.7  42.9  39.4  51.6  44.0  48.6   NaN
fig, ax = plt.subplots(figsize=(9, 7))
sns.heatmap(h2h, annot=True, fmt=".0f", cmap="RdYlGn", center=50,
            linewidths=1.5, linecolor="white", cbar_kws={"label": "Win % of row team"},
            vmin=25, vmax=75, ax=ax)
ax.set_title("Head-to-head win % (row team vs column team)")
plt.tight_layout()
plt.show()

图13

逐行阅读热力图:绿色单元格表示行所对应的球队压制该对手。大多数对抗都接近中性的 50 刻度,这悄然提醒着我们这个联赛竞争何等激烈——真正一边倒的对决只是例外,而非常态。

12  ·  制胜功臣

问题 6:谁决定了最多的 IPL 比赛?

potm = ipl["player_of_match"].value_counts()
top_potm = potm.head(15).sort_values()

fig, ax = plt.subplots(figsize=(11, 6.5))
colors = [ORANGE if i >= len(top_potm) - 3 else BLUE for i in range(len(top_potm))]
ax.barh(top_potm.index, top_potm.values, color=colors, edgecolor="white")
for name, v in zip(top_potm.index, top_potm.values):
    ax.text(v + 0.2, name, int(v), va="center", fontsize=9, color=NAVY)

ax.set_title("Most Player of the Match awards, 2008-2026")
ax.set_xlabel("Awards")
plt.tight_layout()
plt.show()

图14

# Who won the most awards in each individual season?
season_potm = (ipl.dropna(subset=["player_of_match"])
                  .groupby(["season_year", "player_of_match"])
                  .size()
                  .reset_index(name="awards")
                  .sort_values(["season_year", "awards"], ascending=[True, False])
                  .groupby("season_year")
                  .head(1)
                  .set_index("season_year"))
season_potm
player_of_match  awards
season_year                         
2008                SE Marsh       5
2009               YK Pathan       3
2010            SR Tendulkar       4
2011                CH Gayle       6
2012                CH Gayle       5
2013              MEK Hussey       5
2014              GJ Maxwell       4
2015               DA Warner       4
2016                 V Kohli       5
2017               BA Stokes       3
2018             Rashid Khan       4
2019              AD Russell       4
2020          AB de Villiers       3
2021              RD Gaikwad       4
2022           Kuldeep Yadav       4
2023            Shubman Gill       4
2024         Abhishek Sharma       3
2025               KH Pandya       3
2026            Ishan Kishan       3
print(f"Different players have won at least one award : {len(potm)}")
print(f"Players with exactly one award                : {(potm == 1).sum()}")
print(f"Awards taken by the top 15 players            : {potm.head(15).sum()} "
      f"({potm.head(15).sum() / potm.sum() * 100:.1f}% of all awards)")
Different players have won at least one award : 321
Players with exactly one award                : 123
Awards taken by the top 15 players            : 268 (21.7% of all awards)

13  ·  统计检验

图表提出猜想,检验予以证实。

上述探索得出了四个论断。现在,每一个论断都将在标准的 α = 0.05 显著性水平下接受正式检验。

检验 1 — 赢得掷币会提高赢下比赛的概率吗?

# H0: the toss winner wins the match exactly 50% of the time
n     = len(decided)
wins  = (decided["toss_winner"] == decided["winner"]).sum()

# Binomial test against a fair 50/50
res = stats.binomtest(wins, n, p=0.5)

print(f"Toss winners who also won the match : {wins} / {n}  ({wins/n*100:.2f}%)")
print(f"Binomial test p-value               : {res.pvalue:.4f}")
print()
print("Reject H0" if res.pvalue < 0.05 else "Fail to reject H0 - the toss gives no real advantage")
Toss winners who also won the match : 628 / 1218  (51.56%)
Binomial test p-value               : 0.2891

Fail to reject H0 - the toss gives no real advantage

检验 2 — 掷币时的选择与获胜是否相关?

# H0: the toss decision (bat / field) is independent of the match result for the toss winner
table = pd.crosstab(decided["toss_decision"],
                    decided["toss_winner"] == decided["winner"])
table.columns = ["Toss winner lost", "Toss winner won"]
print(table)
print()

chi2, p, dof, expected = stats.chi2_contingency(table)
print(f"Chi-square : {chi2:.3f}")
print(f"p-value    : {p:.4f}")
print()
print("Reject H0 - the decision matters" if p < 0.05 else "Fail to reject H0 - no significant link")
Toss winner lost  Toss winner won
toss_decision                                   
bat                         223              185
field                       367              443

Chi-square : 9.123
p-value    : 0.0025

Reject H0 - the decision matters
一个微妙但重要的结果。掷币本身毫无价值(检验 1),但你用它做出的决定却有显著影响。赢得掷币后选择先投球的球队,获胜频率明显高于选择先击球的球队。硬币帮不了你 — 之后的决定才有用。

检验 3 — 近年来的第一局得分真的更高吗?

early  = ipl[ipl["season_year"] <= 2015]["first_innings_runs"]
modern = ipl[ipl["season_year"] >= 2022]["first_innings_runs"]

print(f"2008-2015 : mean {early.mean():.1f}  (n={len(early)})")
print(f"2022-2026 : mean {modern.mean():.1f}  (n={len(modern)})")
print(f"Difference: {modern.mean() - early.mean():.1f} runs")
print()

t, p = stats.ttest_ind(modern, early, equal_var=False)   # Welch's t-test
print(f"Welch t-statistic : {t:.3f}")
print(f"p-value           : {p:.6f}")

pooled = np.sqrt((early.var() + modern.var()) / 2)
print(f"Cohen's d         : {(modern.mean() - early.mean()) / pooled:.3f}")
print()
print("Reject H0 - scoring really has risen" if p < 0.05 else "Fail to reject H0")
2008-2015 : mean 158.8  (n=517)
2022-2026 : mean 185.1  (n=367)
Difference: 26.3 runs

Welch t-statistic : 11.329
p-value           : 0.000000
Cohen's d         : 0.786

Reject H0 - scoring really has risen

显著的 p 值只能告诉我们差异并非出于运气。Cohen's d 则告诉我们这个差异是否大到值得在意——大致而言,0.2 为小,0.5 为中,0.8 为大。

检验 4 — 场地会改变追分优势吗?

# H0: chase success is independent of the ground
busy = decided[decided["venue"].isin(top_venues.index)]
table = pd.crosstab(busy["venue"], busy["result_side"])
print(table)
print()

chi2, p, dof, expected = stats.chi2_contingency(table)
print(f"Chi-square : {chi2:.3f}   dof: {dof}")
print(f"p-value    : {p:.4f}")
print()
print("Reject H0 - venue matters" if p < 0.05 else "Fail to reject H0 - venues behave alike")
result_side                                         Batting first  Chasing
venue                                                                     
Arun Jaitley Stadium                                           47       54
Bharat Ratna Shri Atal Bihari Vajpayee Ekana Cr...             11       16
Brabourne Stadium                                              14       13
Dr DY Patil Sports Academy                                     17       20
Dubai International Cricket Stadium                            21       22
Eden Gardens                                                   45       60
M Chinnaswamy Stadium                                          44       56
MA Chidambaram Stadium                                         51       45
Maharashtra Cricket Association Stadium                        28       23
Narendra Modi Stadium                                          26       26
Punjab Cricket Association IS Bindra Stadium                   27       34
Rajiv Gandhi International Stadium                             40       48
Sawai Mansingh Stadium                                         24       44
Sharjah Cricket Stadium                                        10       18
Sheikh Zayed Stadium                                           14       21
Wankhede Stadium                                               59       72

Chi-square : 10.218   dof: 15
p-value    : 0.8058

Fail to reject H0 - venues behave alike
发现 5 — 场地效应未能经受住检验。
第 9 节的条形图看起来很有说服力:一些场地似乎明显利于追分,另一些则利于防守。但在比赛场次最多的那些场地上,卡方检验的结果远高于 0.05,因此我们无法否定“所有场地表现相似”这一观点。每个场地仅有 30–130 场比赛,我们所能看到的差距完全处于纯靠运气就能产生的范围之内。

这是整个 notebook 中最有价值的一个结果,因为它与图表相矛盾。仅凭肉眼观察条形图就断言某个场地是“追分福地”,恰恰是这项检验要揪出的错误。

数值型比赛特征之间的相关性

num = ipl[["first_innings_runs", "first_innings_wkts",
           "second_innings_runs", "second_innings_wkts",
           "win_by_runs", "win_by_wickets", "season_year"]]

fig, ax = plt.subplots(figsize=(9, 7))
mask = np.triu(np.ones_like(num.corr(), dtype=bool))
sns.heatmap(num.corr(), mask=mask, annot=True, fmt=".2f", cmap="RdBu_r", center=0,
            linewidths=1.5, linecolor="white", vmin=-1, vmax=1, ax=ax)
ax.set_title("Correlation matrix")
plt.tight_layout()
plt.show()

图15

矩阵中最强的关系也是最显而易见的一个——第一局总分越大,成功防守该总分时的获胜分差就越大。season_year 与 first_innings_runs 之间温和的正相关,正是我们在检验 3 中测量到的得分膨胀,从另一个角度再次显现。

14  ·  执行摘要

以上全部内容,一屏呈现

summary = pd.DataFrame({
    "Metric": [
        "Seasons analysed",
        "Matches analysed",
        "Matches with a decided result",
        "Teams that have played",
        "Grounds used",
        "Cities hosted in",
        "Average first innings score",
        "Average first innings score (2008-2015)",
        "Average first innings score (2022-2026)",
        "Highest first innings total",
        "Matches won chasing",
        "Toss winner also won the match",
        "Captains choosing to field",
        "Tied matches (super over)",
        "Abandoned matches",
    ],
    "Value": [
        ipl["season_year"].nunique(),
        len(ipl),
        int(ipl["decided"].sum()),
        pd.concat([ipl["team1"], ipl["team2"]]).nunique(),
        ipl["venue"].nunique(),
        ipl["city"].nunique(),
        f"{ipl['first_innings_runs'].mean():.1f}",
        f"{early.mean():.1f}",
        f"{modern.mean():.1f}",
        f"{ipl['first_innings_runs'].max()}",
        f"{(decided['result_side'] == 'Chasing').mean() * 100:.1f}%",
        f"{(decided['toss_winner'] == decided['winner']).mean() * 100:.1f}%",
        f"{(ipl['toss_decision'] == 'field').mean() * 100:.1f}%",
        int((ipl["result_type"] == "tie").sum()),
        int((ipl["result_type"] == "no result").sum()),
    ]
})
summary
Metric  Value
0                          Seasons analysed     19
1                          Matches analysed   1243
2             Matches with a decided result   1218
3                    Teams that have played     15
4                              Grounds used     36
5                          Cities hosted in     35
6               Average first innings score  168.7
7   Average first innings score (2008-2015)  158.8
8   Average first innings score (2022-2026)  185.1
9               Highest first innings total    287
10                      Matches won chasing  54.7%
11           Toss winner also won the match  51.6%
12               Captains choosing to field  66.4%
13                Tied matches (super over)     16
14                        Abandoned matches      9

六个答案

1. 得分真的上涨了吗?
是的,而且涨幅并不含蓄。第一局平均得分在十多年里一直稳定在 160 附近,随后从 2023 年起急剧攀升。这一上升在统计上显著,200+ 的出现频率也从罕见变成了常态。

2. 哪支球队最成功?
Chennai 和 Mumbai 是王朝级的球队——它们既霸榜冠军数,也占据胜率表的顶端。有几支球队交出了相当不错的联赛胜率,却从未将其转化为奖杯,这恰恰说明强势的联赛阶段几乎保证不了什么。

3. 掷币能决定比赛吗?
基本上不能。掷币获胜方的比赛胜率在统计上与抛硬币无异,而各赛季在 50% 上下摇摆的幅度,正是随机性所能产生的结果。

4. 追分更容易吗?
通常是的——但有条件。总分低于 160 时,追分一方明显占优;到 180 左右优势消失,超过 200 后则发生逆转。队长们压倒性地选择先投球,而对大多数总分而言,这正是正确的直觉。

5. 哪些球场偏爱谁?
没有图表显示的那么明显。各球场之间的平均得分确实存在真实差异,但表面上“利于追分”与“利于防守”的划分并未通过卡方检验——在当前样本量下,这些差距仍在随机波动所能产生的范围之内。目标分数的大小,远比球场的名字更能预测追分的成败。

6. 谁最能凭一己之力赢下比赛?
Player of the Match(全场最佳球员)榜单由一小群职业生涯漫长的击球手主导。数据集中绝大多数球员要么只获得过一次该奖项,要么从未获得——在这个层面上,决定比赛胜负的能力高度集中。


这份分析无法告诉你的事

对局限保持坦诚是分内之事:

  • 没有逐球(ball-by-ball)数据。 打击率、投球经济率、Powerplay 与 death over 阶段的表现在这里全都不可见。若能拿到逐轮(over-by-over)数据,得分膨胀这一发现会有力得多。
  • Super Over 的结果未被记录,因此 16 场战平的比赛没有胜者,在所有胜率计算中均被排除。
  • 相关不等于因果。 更高的得分与 impact-player 规则在时间上重合,但这份文件无法将原因从更平坦的球道、更出色的击球手或更短的边界中剥离出来。
  • 阵容列未被使用。 team1_players 和 team2_players 保存着完整的十一人名单——一个自然的后续步骤是进行球员出场情况与阵容延续性分析。

感谢阅读。如果这个 notebook 对你有所帮助,一个 upvote 能让更多人看到它 — 我也随时欢迎在评论区讨论方法论。

端到端分析  ·  IPL 2008–2026

数据与完整代码来源

posted on 2026-09-29 16:14  Bony-  阅读(5)  评论(0)    收藏  举报