Prometheus 入门
参考资料
为什么
- prometheus和SDK中间有一条鸿沟
- prometheus只是SDK后端的一种实现方案
- 对prometheus的误解可能会造成SDK使用上的问题
- 带着问题学习:
- SDK里面记录打点一个请求成功数据时,是直接把这条数据推送到服务器里存储起来吗?
- SDK里面记录一个方法的调用耗时,是把方法名称,方法耗时等信息推送到服务器里存储起来吗?
是什么
Prometheus是一个开源的监控系统和时间序列数据库(TSDB),主要用于云原生和微服务架构的应用监控。
- 核心功能:
- 采集指标:通过 HTTP 拉取应用、服务器、容器的指标数据。
- 时间序列存储:以时间序列形式存储数据,便于分析历史趋势。
- 查询语言(PromQL):可以对指标数据进行灵活查询。
- 报警与可视化:可配合 Alertmanager 发送告警,也可用 Grafana 可视化。
- 特点:
- 高度可扩展,适合大规模系统。
- 与 Kubernetes 等现代云平台紧密集成。
- 社区活跃,生态丰富。
Prometheus受启发于Google的Brogmon监控系统(相似的Kubernetes是从Google的Brog系统演变而来),从2012年开始由前Google工程师在Soundcloud以开源软件的形式进行研发,并且于2015年早期对外发布早期版本。2016年5月继Kubernetes之后成为第二个正式加入CNCF基金会的项目,同年6月正式发布1.0版本。2017年底发布了基于全新存储层的2.0版本,能更好地与容器平台、云平台配合。
![image]()
Prometheus作为新一代的云原生监控系统,目前已经有超过650+位贡献者参与到Prometheus的研发工作上,并且超过120+项的第三方集成。
架构
![image]()
- go语言开发
安装
tar xvfz prometheus-xx.tar.gz
cd prometheus-xx
- 运行
./prometheus --config.file=prometheus.yml
配置文件内容
global:
scrape_interval: 15s
evaluation_interval: 15s
alerting:
alertmanagers:
- static_configs:
- targets:
# - alertmanager:9093
rule_files:
# - "first_rules.yml"
# - "second_rules.yml"
scrape_configs:
- job_name: "prometheus"
static_configs:
- targets: ["localhost:9090"]
理解prometheus
⚠️监控数据是prometheus每隔一段时间来pull,并不是每条数据都直接推送到prometheus服务器去
拉取时间配置
拉取时间配置
global:
scrape_interval: 15s
evaluation_interval: 15s
拉取配置
scrape_configs:
- job_name: "prometheus"
static_configs:
- targets: ["localhost:9090"]
指标推送
部署pushgateway,推送到pushgateway,再由prometheus定时拉取
暴露指标
广义上讲所有可以向Prometheus提供监控样本数据的程序都可以被称为一个Exporter。而Exporter的一个实例称为target,如下所示,Prometheus通过轮询的方式定期从这些target中获取样本数据:

Exporter的来源
从Exporter的来源上来讲,主要分为两类:


Exporter规范
- 所有的Exporter程序都需要按照Prometheus的规范,返回监控的样本数据。以Node Exporter为例,当访问/metrics地址时会返回以下内容:
# HELP node_cpu Seconds the cpus spent in each mode.
# TYPE node_cpu counter
node_cpu{cpu="cpu0",mode="idle"} 362812.7890625
# HELP node_load1 1m load average.
# TYPE node_load1 gauge
node_load1 3.0703125
Exporter返回的样本数据,主要由三个部分组成:
- 样本的一般注释信息(HELP)
- HELP <metrics_name> <doc_string>
- 样本的类型注释信息(TYPE)
- TYPE <metrics_name> <metrics_type>
- 样本
<metric name>{<label name>=<label value>, ...}
Metrics类型
Counter
Counter类型的指标其工作方式和计数器一样,只增不减(除非系统发生重置)。常见的监控指标,如http_requests_total,node_cpu都是Counter类型的监控指标。 一般在定义Counter类型指标的名称时推荐使用_total作为后缀。
# HELP go_gc_cycles_automatic_gc_cycles_total Count of completed GC cycles generated by the Go runtime.
# TYPE go_gc_cycles_automatic_gc_cycles_total counter
go_gc_cycles_automatic_gc_cycles_total 14
Gauge
与Counter不同,Gauge类型的指标侧重于反应系统的当前状态。因此这类指标的样本数据可增可减。常见指标如:node_memory_MemFree(主机当前空闲的内容大小)、node_memory_MemAvailable(可用内存大小)都是Gauge类型的监控指标。
# HELP go_gc_scan_globals_bytes The total amount of global variable space that is scannable.
# TYPE go_gc_scan_globals_bytes gauge
go_gc_scan_globals_bytes 813280
Histogram
- 用于统计和分析样本的分布情况
# HELP prometheus_tsdb_compaction_chunk_range Final time range of chunks on their first compaction
# TYPE prometheus_tsdb_compaction_chunk_range histogram
prometheus_tsdb_compaction_chunk_range_bucket{le="100"} 0
prometheus_tsdb_compaction_chunk_range_bucket{le="400"} 0
prometheus_tsdb_compaction_chunk_range_bucket{le="1600"} 0
prometheus_tsdb_compaction_chunk_range_bucket{le="6400"} 0
prometheus_tsdb_compaction_chunk_range_bucket{le="25600"} 0
prometheus_tsdb_compaction_chunk_range_bucket{le="102400"} 0
prometheus_tsdb_compaction_chunk_range_bucket{le="409600"} 0
prometheus_tsdb_compaction_chunk_range_bucket{le="1.6384e+06"} 260
prometheus_tsdb_compaction_chunk_range_bucket{le="6.5536e+06"} 780
prometheus_tsdb_compaction_chunk_range_bucket{le="2.62144e+07"} 780
prometheus_tsdb_compaction_chunk_range_bucket{le="+Inf"} 780
prometheus_tsdb_compaction_chunk_range_sum 1.1540798e+09
prometheus_tsdb_compaction_chunk_range_count 780
- Histogram通过histogram_quantile函数是在服务器端计算的分位数
Summary
- 用于统计和分析样本的分布情况
# HELP prometheus_tsdb_wal_fsync_duration_seconds Duration of WAL fsync.
# TYPE prometheus_tsdb_wal_fsync_duration_seconds summary
prometheus_tsdb_wal_fsync_duration_seconds{quantile="0.5"} 0.012352463
prometheus_tsdb_wal_fsync_duration_seconds{quantile="0.9"} 0.014458005
prometheus_tsdb_wal_fsync_duration_seconds{quantile="0.99"} 0.017316173
prometheus_tsdb_wal_fsync_duration_seconds_sum 2.888716127000002
prometheus_tsdb_wal_fsync_duration_seconds_count 216
- Sumamry的分位数则是直接在客户端计算完成
问题
- 接口qps怎么监控?
- 每条指标保存的时候会同时保存当前的时间点,查询时使用增量查询
- rate(request_call_total[1m])
- 接口耗时怎么监控?SDK里直接提供了计时器功能怎么查询结果
- 使用两个counter,一个记录总耗时,一个记录总请求数,总耗时/总请求数=平均耗时
- 不是每个请求都应该记录下来吗?
- 不记录单条打点数据,只记录固定时间内的变化
- 通过SDK上传的指标查询不到?
- 比如SDK里的计时器,是通过两个counter实现的,查询是需要自己计算
- 有的SDK会自动给标签加后缀,比如代码写的是counter类型的quest_success,实际上报的是quest_success_total
标签
# HELP node_cpu Seconds the cpus spent in each mode.
# TYPE node_cpu counter
node_cpu{cpu="cpu0",mode="idle"} 362812.7890625
- node_cpu是指标名称
- cpu="cpu0"是标签
- mode="idle"是标签
- 标签用于指标筛选,指标+标签+时间+值 形成一条数据保存下来
PromQL简介
单条指标查询

范围指标查询

聚合查询




浙公网安备 33010602011771号