Go Prometheus Grafana 监控告警完整监控系统搭建没有监控的系统是黑盒。本文从指标采集、可视化、告警三个维度讲透 Go 端监控整套流程。一、Prometheus 三件套Prometheus服务器拉模式采集 metricsGrafana可视化Alertmanager告警路由Go 端使用client_golangimportgithub.com/prometheus/client_golang/prometheus/promhttphttp.Handle(/metrics,promhttp.Handler())二、自定义指标var(activeUsersprometheus.NewGauge(prometheus.GaugeOpts{Name:active_users,Help:current active users,})requestCountprometheus.NewCounterVec(prometheus.CounterOpts{Name:http_request_total},[]string{method,endpoint,status},))prometheus.MustRegister(activeUsers,requestCount)三、HTTP 中间件采集funcpromMiddleware()gin.HandlerFunc{returnfunc(c*gin.Context){c.Next()requestCount.WithLabelValues(c.Request.Method,c.FullPath(),fmt.Sprint(c.Writer.Status())).Inc()}}四、Go runtime 自动指标promhttp自动暴露go_goroutinesgo_gc_duration_secondsprocess_resident_memory_bytesgo_memstats_*五、Grafana 看板模版 11752 是 Go 应用的入口模版包含goroutine 趋势GC pauseheap profileHTTP p50/p95/p99六、PromQL# 5xx 错误率 sum(rate(http_request_total{status~5..}[5m])) by (service) / sum(rate(http_request_total[5m])) by (service) # p99 latency histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le))七、告警规则groups:-name:examplerules:-alert:HighErrorRateexpr:|sum(rate(http_request_total{status~5..}[5m])) / sum(rate(http_request_total[5m])) 0.05for:1mlabels:severity:criticalPrometheus 触发后Alertmanager 通过钉钉 / Webhook / 邮件告警。八、推送 vs 拉取Prometheus 默认拉模式。短期任务用 pushimportgithub.com/prometheus/client_golang/prometheus/pushpush.Add(https://pushgateway:9091,job1,map[string]string{},dto.Metric{...})九、自研 metrics 对接typeCounterstruct{val atomic.Int64 onfunc()float64}func(c*Counter)Describe(chchan-*prometheus.Desc){ch-prometheus.NewDesc(c.name,,nil,nil)}func(c*Counter)Collect(chchan-prometheus.Metric){ch-prometheus.MustNewConstMetric(c.desc,prometheus.GaugeValue,float64(c.val.Load()))}十、踩坑清单cardinality 爆炸label 过多 → 不同 label 巨多 → 内存炸长 label 非常消耗内存default registry 重复注册用独立 registry拉不动业务端 prometheus 是单点的5 min scrape十一、生产实战完整监控架构app → promhttp prometheus → scrape intervals → TSDB ↓ Grafana → 看板 ↓ Alertmanager → 钉钉 / 企微 / 邮件12-factor 标准。十二、eBPF 集成Prometheus 还可收集 BPF trace metricrequest latency at kernel level网络 stack 队列监控十三、总结与展望Prometheus Grafana 是事实标准。结合 OTel、Loki 等可观测数据可构建完整 observability 体系。未来VictoriaMetrics、Mimir 等云原生时序数据库针对大规模云高效处理。十四、参考文献prometheus client_golang 文档grafana 模板库prometheus book