Prometheus 入门

参考资料

为什么

  • prometheus和SDK中间有一条鸿沟
  • prometheus只是SDK后端的一种实现方案
  • 对prometheus的误解可能会造成SDK使用上的问题
  • 带着问题学习:
    • SDK里面记录打点一个请求成功数据时,是直接把这条数据推送到服务器里存储起来吗?
    • SDK里面记录一个方法的调用耗时,是把方法名称,方法耗时等信息推送到服务器里存储起来吗?

是什么

Prometheus是一个开源的监控系统和时间序列数据库(TSDB),主要用于云原生和微服务架构的应用监控。

  • 核心功能:
    • 采集指标:通过 HTTP 拉取应用、服务器、容器的指标数据。
    • 时间序列存储:以时间序列形式存储数据,便于分析历史趋势。
    • 查询语言(PromQL):可以对指标数据进行灵活查询。
    • 报警与可视化:可配合 Alertmanager 发送告警,也可用 Grafana 可视化。
  • 特点:
    • 高度可扩展,适合大规模系统。
    • 与 Kubernetes 等现代云平台紧密集成。
    • 社区活跃,生态丰富。
      Prometheus受启发于Google的Brogmon监控系统(相似的Kubernetes是从Google的Brog系统演变而来),从2012年开始由前Google工程师在Soundcloud以开源软件的形式进行研发,并且于2015年早期对外发布早期版本。2016年5月继Kubernetes之后成为第二个正式加入CNCF基金会的项目,同年6月正式发布1.0版本。2017年底发布了基于全新存储层的2.0版本,能更好地与容器平台、云平台配合。
      image
      Prometheus作为新一代的云原生监控系统,目前已经有超过650+位贡献者参与到Prometheus的研发工作上,并且超过120+项的第三方集成。
      架构
      image
  • go语言开发

安装

  1. 下载
    https://prometheus.io/download/
  2. 解压
tar xvfz prometheus-xx.tar.gz
cd prometheus-xx
  1. 运行
./prometheus --config.file=prometheus.yml

配置文件内容

global:
  scrape_interval: 15s 
  evaluation_interval: 15s 

alerting:
  alertmanagers:
    - static_configs:
        - targets:
          # - alertmanager:9093

rule_files:
  # - "first_rules.yml"
  # - "second_rules.yml"

scrape_configs:
  - job_name: "prometheus"
    static_configs:
      - targets: ["localhost:9090"]
  1. 访问 http://localhost:9090

理解prometheus

⚠️监控数据是prometheus每隔一段时间来pull,并不是每条数据都直接推送到prometheus服务器去
拉取时间配置

拉取时间配置

global:
  scrape_interval: 15s 
  evaluation_interval: 15s 

拉取配置

scrape_configs:
  - job_name: "prometheus"
    static_configs:
      - targets: ["localhost:9090"]

指标推送

部署pushgateway,推送到pushgateway,再由prometheus定时拉取

暴露指标

广义上讲所有可以向Prometheus提供监控样本数据的程序都可以被称为一个Exporter。而Exporter的一个实例称为target,如下所示,Prometheus通过轮询的方式定期从这些target中获取样本数据:
image

Exporter的来源

从Exporter的来源上来讲,主要分为两类:
image
image

Exporter规范

  • 所有的Exporter程序都需要按照Prometheus的规范,返回监控的样本数据。以Node Exporter为例,当访问/metrics地址时会返回以下内容:
# HELP node_cpu Seconds the cpus spent in each mode.
# TYPE node_cpu counter
node_cpu{cpu="cpu0",mode="idle"} 362812.7890625
# HELP node_load1 1m load average.
# TYPE node_load1 gauge
node_load1 3.0703125

Exporter返回的样本数据,主要由三个部分组成:

  • 样本的一般注释信息(HELP)
    • HELP <metrics_name> <doc_string>
  • 样本的类型注释信息(TYPE)
    • TYPE <metrics_name> <metrics_type>
  • 样本
<metric name>{<label name>=<label value>, ...}

Metrics类型

Counter

Counter类型的指标其工作方式和计数器一样,只增不减(除非系统发生重置)。常见的监控指标,如http_requests_total,node_cpu都是Counter类型的监控指标。 一般在定义Counter类型指标的名称时推荐使用_total作为后缀。

# HELP go_gc_cycles_automatic_gc_cycles_total Count of completed GC cycles generated by the Go runtime.
# TYPE go_gc_cycles_automatic_gc_cycles_total counter
go_gc_cycles_automatic_gc_cycles_total 14

Gauge

与Counter不同,Gauge类型的指标侧重于反应系统的当前状态。因此这类指标的样本数据可增可减。常见指标如:node_memory_MemFree(主机当前空闲的内容大小)、node_memory_MemAvailable(可用内存大小)都是Gauge类型的监控指标。

# HELP go_gc_scan_globals_bytes The total amount of global variable space that is scannable.
# TYPE go_gc_scan_globals_bytes gauge
go_gc_scan_globals_bytes 813280

Histogram

  • 用于统计和分析样本的分布情况
# HELP prometheus_tsdb_compaction_chunk_range Final time range of chunks on their first compaction
# TYPE prometheus_tsdb_compaction_chunk_range histogram
prometheus_tsdb_compaction_chunk_range_bucket{le="100"} 0
prometheus_tsdb_compaction_chunk_range_bucket{le="400"} 0
prometheus_tsdb_compaction_chunk_range_bucket{le="1600"} 0
prometheus_tsdb_compaction_chunk_range_bucket{le="6400"} 0
prometheus_tsdb_compaction_chunk_range_bucket{le="25600"} 0
prometheus_tsdb_compaction_chunk_range_bucket{le="102400"} 0
prometheus_tsdb_compaction_chunk_range_bucket{le="409600"} 0
prometheus_tsdb_compaction_chunk_range_bucket{le="1.6384e+06"} 260
prometheus_tsdb_compaction_chunk_range_bucket{le="6.5536e+06"} 780
prometheus_tsdb_compaction_chunk_range_bucket{le="2.62144e+07"} 780
prometheus_tsdb_compaction_chunk_range_bucket{le="+Inf"} 780
prometheus_tsdb_compaction_chunk_range_sum 1.1540798e+09
prometheus_tsdb_compaction_chunk_range_count 780
  • Histogram通过histogram_quantile函数是在服务器端计算的分位数

Summary

  • 用于统计和分析样本的分布情况
# HELP prometheus_tsdb_wal_fsync_duration_seconds Duration of WAL fsync.
# TYPE prometheus_tsdb_wal_fsync_duration_seconds summary
prometheus_tsdb_wal_fsync_duration_seconds{quantile="0.5"} 0.012352463
prometheus_tsdb_wal_fsync_duration_seconds{quantile="0.9"} 0.014458005
prometheus_tsdb_wal_fsync_duration_seconds{quantile="0.99"} 0.017316173
prometheus_tsdb_wal_fsync_duration_seconds_sum 2.888716127000002
prometheus_tsdb_wal_fsync_duration_seconds_count 216
  • Sumamry的分位数则是直接在客户端计算完成

问题

  • 接口qps怎么监控?
    • 每条指标保存的时候会同时保存当前的时间点,查询时使用增量查询
    • rate(request_call_total[1m])
  • 接口耗时怎么监控?SDK里直接提供了计时器功能怎么查询结果
    • 使用两个counter,一个记录总耗时,一个记录总请求数,总耗时/总请求数=平均耗时
  • 不是每个请求都应该记录下来吗?
    • 不记录单条打点数据,只记录固定时间内的变化
  • 通过SDK上传的指标查询不到?
    • 比如SDK里的计时器,是通过两个counter实现的,查询是需要自己计算
    • 有的SDK会自动给标签加后缀,比如代码写的是counter类型的quest_success,实际上报的是quest_success_total

标签

# HELP node_cpu Seconds the cpus spent in each mode.
# TYPE node_cpu counter
node_cpu{cpu="cpu0",mode="idle"} 362812.7890625
  • node_cpu是指标名称
  • cpu="cpu0"是标签
  • mode="idle"是标签
  • 标签用于指标筛选,指标+标签+时间+值 形成一条数据保存下来

PromQL简介

单条指标查询

image

范围指标查询

image

聚合查询

image

posted @ 2026-04-10 15:21  java拌饭  阅读(21)  评论(0)    收藏  举报