路线图

27-追踪栈

星辉 2026-07-02 阅读 3 min 541 字 路线图
27-追踪栈 封面

链路追踪(Traces)记录一次请求穿过整个分布式系统的完整路径。单体应用看日志够用,一旦一个请求穿 5、6 个微服务,没有链路追踪排障就是盲人摸象。本章覆盖 OpenTelemetry 接入、Jaeger/Tempo 存储选型、TraceID 串联三支柱、采样策略。

核心数据模型#

  • Trace:一次完整请求,有全局唯一 trace_id。
  • Span:一个操作单元(一次 RPC、一次 DB 查询),有 start/end、父子关系、attributes。
  • Trace Context:A 调 B 时把 trace_id + span_id 透传过去(HTTP header)。任何一跳没透传,链路就断——落地链路追踪最常见的坑。
text
Trace ID: 4bf92f3577b34da6a3ce929d0e0e4736

[frontend]         POST /checkout                          0ms → 3102ms
  [checkout-api]   handleCheckout                          2ms → 3100ms
    [checkout-api] validateOrder (DB query)                3ms → 12ms
    [checkout-api] reserveInventory (gRPC)                15ms → 45ms
    [checkout-api] processPayment (HTTP)                   48ms → 3098ms  ← 慢!
      [payment-svc] callBankAPI                            50ms → 3096ms  ← 超时

看到这个 Flame Graph,3 秒延迟立刻定位到 payment-svc 调用银行 API 这一跳。

OpenTelemetry 接入#

OpenTelemetry(OTel)是 CNCF 的开源项目,目标是成为 Metrics/Logs/Traces 的统一采集标准。一次用 OTel SDK 埋点,三种信号统一产出,到处可导出——选型第一原则:埋点层绑定标准 OTel,不绑死任何后端厂商。

应用侧 SDK#

go
func initTracer(ctx context.Context) func() {
    exp, _ := otlptracegrpc.New(ctx,
        otlptracegrpc.WithEndpoint("otel-collector.obs.svc:4317"),
        otlptracegrpc.WithInsecure())

    tp := trace.NewTracerProvider(
        trace.WithBatcher(exp),
        trace.WithSampler(trace.ParentBased(trace.TraceIDRatioBased(0.05))),
        trace.WithResource(resource.NewWithAttributes(
            semconv.SchemaURL,
            semconv.ServiceName("order-service"),
            semconv.ServiceVersion(os.Getenv("APP_VERSION")),
        )),
    )
    otel.SetTracerProvider(tp)
    return func() { _ = tp.Shutdown(ctx) }
}

OTel Collector#

OTel Collector 是三段式 pipeline 的独立进程:

text
Receivers 接收  →  Processors 处理  →  Exporters 导出
(收三种信号)      (batch/采样/脱敏/   (发给 Prometheus/
                  加label/过滤)       Loki/Tempo/任意后端)

价值:埋点用 OTel SDK + 采集用 OTel Collector → 完全厂商中立,换后端只改 exporter 一行配置。

yaml
receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
      http:
        endpoint: 0.0.0.0:4318

processors:
  batch:
    timeout: 1s
    send_batch_size: 8192
  memory_limiter:
    check_interval: 1s
    limit_percentage: 80
  tail_sampling:
    decision_wait: 10s
    policies:
      - name: error-traces
        type: status_code
        status_code: { status_codes: [ERROR] }
      - name: slow-traces
        type: latency
        latency: { threshold_ms: 1000 }
      - name: probabilistic
        type: probabilistic
        probabilistic: { sampling_percentage: 5 }

exporters:
  otlp/tempo:
    endpoint: tempo-distributor.tempo.svc:4317

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, tail_sampling, batch]
      exporters: [otlp/tempo]

部署模式#

  • Agent 模式:每节点 DaemonSet,应用通过 host.ip:4317 上报。网络跳转少。缺点:tail sampling 只能看到一个节点的 trace。
  • Gateway 模式:集中部署几个 Collector,tail sampling 可做全局决策,但需要有状态(同一 trace 要打到同一 gateway)。
  • Agent + Gateway 双层:Agent 负责 batch、enrichment,Gateway 负责 tail sampling、路由。大规模推荐。

存储选型:Jaeger vs Tempo#

维度JaegerTempo
存储Elasticsearch(重)对象存储(轻)
索引span tag 倒排索引只索引 trace_id + 少量元数据
查询tag 全文过滤trace_id 精确 + TraceQL
运维成本高(ES JVM)低(对象存储 + 无 JVM)
成本高低(约 ES 的 1/10)
与 Grafana/Loki 联动一般原生集成

Tempo 的设计哲学:"少索引省钱"——和 Loki 一脉相承。2.x 引入 Parquet 列式存储 + TraceQL,过滤能力补上来了。

TraceQL 查询#

traceql
# 按 service 过滤
{ resource.service.name = "order-service" }

# 按 status 过滤
{ .status = error }

# 按 duration
{ duration > 500ms }

# 组合
{ resource.service.name = "payment-api" && duration > 1s }

# 父子关系
{ span.http.target = "/checkout" } >> { .status = error }

TraceID 串联三支柱#

三支柱单独看都是片面的,真正救命的是它们能联动——靠的是共享关联键:

text
Metrics(发现问题)→ Logs(理解细节)→ Traces(定位根因)
     ↓                     ↓                  ↓
  exemplar             trace_id           trace_id
  (trace_id)           字段               本身

Exemplar:Metric → Trace 直接跳转#

Prometheus 的 Exemplar 功能在指标数据点上附加 trace_id。Grafana 看到延迟尖刺,点一下跳到那个慢请求的完整链路:

go
spanCtx := trace.SpanFromContext(r.Context()).SpanContext()
if spanCtx.IsValid() {
    requestDuration.With(labels).(prometheus.ExemplarObserver).ObserveWithExemplar(
        duration,
        prometheus.Labels{"traceID": spanCtx.TraceID().String()},
    )
}

Loki → Tempo(日志查 trace)#

Loki datasource 配 derivedFields,从日志提取 trace_id,点击跳转 Tempo。

Tempo → Loki(trace 查日志)#

Tempo datasource 配 tracesToLogs,点击 span 自动跳转 Loki 查同时间段日志。

采样策略#

全量采集 trace 会有性能开销,但采样率设太低(如 1%)低流量时根本看不到 trace。

Head Sampling(头采样)#

应用上来就决定采不采,问题:错误 trace 可能因采样率低而错过。

Tail Sampling(尾采样)#

先全量收集,等 trace 完整后再决定要不要采。OTel Collector 的 tail_sampling processor:

yaml
tail_sampling:
  decision_wait: 15s       # 等 15s 让 trace 完整
  num_traces: 200000       # 内存中同时保留的 trace 数
  policies:
    # 错误 trace 全保留
    - name: keep-errors
      type: status_code
      status_code: { status_codes: [ERROR] }
    # 慢请求全保留
    - name: keep-slow
      type: latency
      latency: { threshold_ms: 800 }
    # 其他按 5% 采样
    - name: sample-rest
      type: probabilistic
      probabilistic: { sampling_percentage: 5 }

核心优势:错误和慢请求不丢,其他省存储。这是 tail sampling 相比 head sampling 的本质区别——head sampling 在请求开始就决定,错误 trace 可能被错过;tail sampling 看完整条 trace 再决定。

坑:同一 trace 必须打到同一 Collector(tail_sampling 是单机状态),Gateway 前面要用一致性哈希 LB 按 trace_id hash。

小结#

链路追踪是微服务排障的命根子。OpenTelemetry 是厂商中立的埋点+采集标准——选型第一原则绑定 OTel 不绑后端。存储选型上 Tempo(对象存储 + 少索引)比 Jaeger(ES 倒排)便宜一个数量级,配合 TraceQL 过滤能力够用。TraceID 是三支柱联动的关联键——日志必带 trace_id、指标用 exemplar 挂 trace_id、Grafana 配 data source linking 实现一键互跳。采样用 tail sampling——错误和慢请求全采,其他按比例采,省钱又不丢关键链路。