27-追踪栈
链路追踪(Traces)记录一次请求穿过整个分布式系统的完整路径。单体应用看日志够用,一旦一个请求穿 5、6 个微服务,没有链路追踪排障就是盲人摸象。本章覆盖 OpenTelemetry 接入、Jaeger/Tempo 存储选型、TraceID 串联三支柱、采样策略。
核心数据模型#
- Trace:一次完整请求,有全局唯一
trace_id。 - Span:一个操作单元(一次 RPC、一次 DB 查询),有 start/end、父子关系、attributes。
- Trace Context:A 调 B 时把
trace_id + span_id透传过去(HTTP header)。任何一跳没透传,链路就断——落地链路追踪最常见的坑。
Trace ID: 4bf92f3577b34da6a3ce929d0e0e4736
[frontend] POST /checkout 0ms → 3102ms
[checkout-api] handleCheckout 2ms → 3100ms
[checkout-api] validateOrder (DB query) 3ms → 12ms
[checkout-api] reserveInventory (gRPC) 15ms → 45ms
[checkout-api] processPayment (HTTP) 48ms → 3098ms ← 慢!
[payment-svc] callBankAPI 50ms → 3096ms ← 超时看到这个 Flame Graph,3 秒延迟立刻定位到 payment-svc 调用银行 API 这一跳。
OpenTelemetry 接入#
OpenTelemetry(OTel)是 CNCF 的开源项目,目标是成为 Metrics/Logs/Traces 的统一采集标准。一次用 OTel SDK 埋点,三种信号统一产出,到处可导出——选型第一原则:埋点层绑定标准 OTel,不绑死任何后端厂商。
应用侧 SDK#
func initTracer(ctx context.Context) func() {
exp, _ := otlptracegrpc.New(ctx,
otlptracegrpc.WithEndpoint("otel-collector.obs.svc:4317"),
otlptracegrpc.WithInsecure())
tp := trace.NewTracerProvider(
trace.WithBatcher(exp),
trace.WithSampler(trace.ParentBased(trace.TraceIDRatioBased(0.05))),
trace.WithResource(resource.NewWithAttributes(
semconv.SchemaURL,
semconv.ServiceName("order-service"),
semconv.ServiceVersion(os.Getenv("APP_VERSION")),
)),
)
otel.SetTracerProvider(tp)
return func() { _ = tp.Shutdown(ctx) }
}OTel Collector#
OTel Collector 是三段式 pipeline 的独立进程:
Receivers 接收 → Processors 处理 → Exporters 导出
(收三种信号) (batch/采样/脱敏/ (发给 Prometheus/
加label/过滤) Loki/Tempo/任意后端)价值:埋点用 OTel SDK + 采集用 OTel Collector → 完全厂商中立,换后端只改 exporter 一行配置。
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
batch:
timeout: 1s
send_batch_size: 8192
memory_limiter:
check_interval: 1s
limit_percentage: 80
tail_sampling:
decision_wait: 10s
policies:
- name: error-traces
type: status_code
status_code: { status_codes: [ERROR] }
- name: slow-traces
type: latency
latency: { threshold_ms: 1000 }
- name: probabilistic
type: probabilistic
probabilistic: { sampling_percentage: 5 }
exporters:
otlp/tempo:
endpoint: tempo-distributor.tempo.svc:4317
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, tail_sampling, batch]
exporters: [otlp/tempo]部署模式#
- Agent 模式:每节点 DaemonSet,应用通过
host.ip:4317上报。网络跳转少。缺点:tail sampling 只能看到一个节点的 trace。 - Gateway 模式:集中部署几个 Collector,tail sampling 可做全局决策,但需要有状态(同一 trace 要打到同一 gateway)。
- Agent + Gateway 双层:Agent 负责 batch、enrichment,Gateway 负责 tail sampling、路由。大规模推荐。
存储选型:Jaeger vs Tempo#
| 维度 | Jaeger | Tempo |
|---|---|---|
| 存储 | Elasticsearch(重) | 对象存储(轻) |
| 索引 | span tag 倒排索引 | 只索引 trace_id + 少量元数据 |
| 查询 | tag 全文过滤 | trace_id 精确 + TraceQL |
| 运维成本 | 高(ES JVM) | 低(对象存储 + 无 JVM) |
| 成本 | 高 | 低(约 ES 的 1/10) |
| 与 Grafana/Loki 联动 | 一般 | 原生集成 |
Tempo 的设计哲学:"少索引省钱"——和 Loki 一脉相承。2.x 引入 Parquet 列式存储 + TraceQL,过滤能力补上来了。
TraceQL 查询#
# 按 service 过滤
{ resource.service.name = "order-service" }
# 按 status 过滤
{ .status = error }
# 按 duration
{ duration > 500ms }
# 组合
{ resource.service.name = "payment-api" && duration > 1s }
# 父子关系
{ span.http.target = "/checkout" } >> { .status = error }TraceID 串联三支柱#
三支柱单独看都是片面的,真正救命的是它们能联动——靠的是共享关联键:
Metrics(发现问题)→ Logs(理解细节)→ Traces(定位根因)
↓ ↓ ↓
exemplar trace_id trace_id
(trace_id) 字段 本身Exemplar:Metric → Trace 直接跳转#
Prometheus 的 Exemplar 功能在指标数据点上附加 trace_id。Grafana 看到延迟尖刺,点一下跳到那个慢请求的完整链路:
spanCtx := trace.SpanFromContext(r.Context()).SpanContext()
if spanCtx.IsValid() {
requestDuration.With(labels).(prometheus.ExemplarObserver).ObserveWithExemplar(
duration,
prometheus.Labels{"traceID": spanCtx.TraceID().String()},
)
}Loki → Tempo(日志查 trace)#
Loki datasource 配 derivedFields,从日志提取 trace_id,点击跳转 Tempo。
Tempo → Loki(trace 查日志)#
Tempo datasource 配 tracesToLogs,点击 span 自动跳转 Loki 查同时间段日志。
采样策略#
全量采集 trace 会有性能开销,但采样率设太低(如 1%)低流量时根本看不到 trace。
Head Sampling(头采样)#
应用上来就决定采不采,问题:错误 trace 可能因采样率低而错过。
Tail Sampling(尾采样)#
先全量收集,等 trace 完整后再决定要不要采。OTel Collector 的 tail_sampling processor:
tail_sampling:
decision_wait: 15s # 等 15s 让 trace 完整
num_traces: 200000 # 内存中同时保留的 trace 数
policies:
# 错误 trace 全保留
- name: keep-errors
type: status_code
status_code: { status_codes: [ERROR] }
# 慢请求全保留
- name: keep-slow
type: latency
latency: { threshold_ms: 800 }
# 其他按 5% 采样
- name: sample-rest
type: probabilistic
probabilistic: { sampling_percentage: 5 }核心优势:错误和慢请求不丢,其他省存储。这是 tail sampling 相比 head sampling 的本质区别——head sampling 在请求开始就决定,错误 trace 可能被错过;tail sampling 看完整条 trace 再决定。
坑:同一 trace 必须打到同一 Collector(tail_sampling 是单机状态),Gateway 前面要用一致性哈希 LB 按 trace_id hash。
小结#
链路追踪是微服务排障的命根子。OpenTelemetry 是厂商中立的埋点+采集标准——选型第一原则绑定 OTel 不绑后端。存储选型上 Tempo(对象存储 + 少索引)比 Jaeger(ES 倒排)便宜一个数量级,配合 TraceQL 过滤能力够用。TraceID 是三支柱联动的关联键——日志必带 trace_id、指标用 exemplar 挂 trace_id、Grafana 配 data source linking 实现一键互跳。采样用 tail sampling——错误和慢请求全采,其他按比例采,省钱又不丢关键链路。