Lumberyardo11y

Documentation

Point an agent at us, learn six query operators, set up burn-rate alerts. That is most of it.

Endpoints

SignalPathProtocols
Logs/api/v1/logsOTLP, JSON, Elastic bulk
Metrics/api/v1/writePrometheus remote-write, OTLP
Traces/api/v1/tracesOTLP gRPC and HTTP

Drop rules

Applied at ingest, before metering. This is the cheapest possible place to remove noise.

json
{
  "rules": [
    {"match": "service=~'.*' and level='debug'",
     "action": "drop", "unless": "trace_id != ''"},
    {"match": "path='/healthz'", "action": "drop"},
    {"match": "service='batch'", "action": "sample",
     "rate": 0.05}
  ]
}

Language

Pipelines, left to right. Six operators cover most work: filter, parse, rate, group by, quantile and join.

bash
logs {service="api"} 
  | json 
  | status >= 500 
  | rate(1m) by (route) 
  | topk(10)

Examples

  • metrics {name="http_request_duration_seconds"} | quantile(0.99) by (route)
  • traces {duration > 5s} | group by (db.system) | count()
  • logs {} | pattern | topk(20) by (pattern_id) — find the twenty log shapes producing your volume

SLOs

Define the objective, not the alert. We generate multi-window multi-burn-rate alerts from it, which is the difference between being paged when it matters and being paged at 03:00 for a blip.

json
{
  "name": "checkout availability",
  "objective": 99.9,
  "window": "30d",
  "good": "logs {service='checkout'} | status < 500",
  "total": "logs {service='checkout'}"
}

Routing

Alerts route to the usual destinations, with per-team schedules and an explicit 'this is a warning, do not page' severity that is actually respected.