Cases

OPS-1836 · Incident Ops

payments-api latency above rubric

SEV-3 · Approved

  1. 1

    Signal

  2. 2

    Cause

  3. 3

    Plan

  4. 4

    Draft

  5. 5

    Person gate

  6. 6

    Record

Prometheus

payments-api p99 above 800ms for 20 minutes. Error rate is flat.

Cause

One replica is hot. The others are idle. This is capacity, not a bad deploy.

Plan

Add one replica. Leave the database alone.

Draft

Scale payments-api by one

Allow-list · scale-deployment · Jira OPS-1836

Footprint

CustomersChannelsAPIsServicesAsyncDataHostsRetail customersCustomer · 3Partner channelCustomer · 1storefrontChannel · 2checkoutChannel · 3mobile-appChannel · 4orders-apiAPI · 9payments-apiAPI · 5auth-apiAPI · 1pricing-apiAPI · 3catalog-apiAPI · 2inventory-apiAPI · 3orders.fulfillmentKafka · 4payments.eventsKafka · 2fulfill-workerWorker · 3orders-dbDatabase · 2payments-dbDatabase · 1catalog-dbDatabase · 2card-networkExternal · 1app-node-aHost · 2app-node-cHost · 1

Person gate

Approved

The allow-listed action stays on the case. Dry-run: nothing was sent.

29 Sep, 09:14

  • JiraDry-run
  • PrometheusDry-run
  • KubernetesDry-run