Auto-Scaling Strategy

Overview

This design defines how Link dynamically scales application processing and Kubernetes infrastructure in response to workload demand.

Capacity targets and burst requirements are defined by Performance Model. This design describes how the platform supplies the application, Kafka, and Kubernetes capacity necessary to meet those requirements.

This design assumes the Workload Mode capability defined by Service Workload Mode has been implemented so REST/API processing and asynchronous worker processing can be deployed and scaled independently.

Terminology

KEDA (Kubernetes Event-Driven Autoscaling) is a Kubernetes autoscaling component that can scale workloads based on external workload signals such as Kafka consumer lag. In this design, KEDA is the preferred mechanism for translating Kafka backlog into desired Worker replica counts, while Kubernetes HPA performs the underlying pod scaling behavior.

High-Level Guidance

  • Kafka topic partition counts should be provisioned according to the burst-processing capacity required by Performance Model, with appropriate headroom for expected growth.
  • Kafka partitions are planned capacity. Partition counts should not be automatically increased or decreased as part of pod autoscaling.
  • Services that combine REST APIs and asynchronous processing should be deployed as separate API and Worker Kubernetes Deployments using the same application image and different Workload Modes.
  • API and Worker Deployments should scale independently.
  • Kafka Worker deployments should scale primarily according to Kafka consumer lag/backlog.
  • REST/API deployments should scale according to CPU, memory, and where practical request-oriented utilization/latency metrics.
  • REST/API deployments must not scale to zero.
  • Worker deployments may scale to zero only when service-specific startup, latency, scheduling, and availability characteristics make doing so safe.
  • Kafka Worker maximum replicas must not exceed useful consumer concurrency provided by the relevant consumer-group/topic partition topology.
  • Kubernetes resource requests must accurately represent workload consumption because pod scheduling and node-pool autoscaling depend on those requests.
  • AKS node pools should automatically add capacity when pods cannot be scheduled because sufficient requested CPU/memory capacity is unavailable, and remove capacity when nodes become safely underutilized.
  • Scale-up should be comparatively aggressive when backlog develops. Scale-down should be conservative and require sustained low demand.
  • Autoscaling limits must account for downstream constraints and not merely available Kubernetes compute capacity.
  • Kubernetes should monitor Kafka consumer lag, including via the Kafka REST API consumer-group lag summary endpoint where available, to determine when to increase replicas for Kafka-backed services. See Confluent Documentation for more information.

Visual Model

Three-Layer Capacity Model

Partitions define available parallelism. Replicas consume that parallelism. Nodes provide the compute needed to run those replicas.

flowchart TB
    subgraph Kafka["Layer 1 - Kafka Capacity"]
        Partitions["Kafka Partitions<br>Available Consumer Parallelism"]
    end

    subgraph App["Layer 2 - Application Capacity"]
        Workers["Worker Replicas<br>Active Processing Capacity"]
    end

    subgraph Compute["Layer 3 - Infrastructure Capacity"]
        Nodes["AKS Nodes<br>CPU / Memory Capacity"]
    end

    Performance["Achievable Throughput"]

    Partitions -->|"Limits useful<br>consumer concurrency"| Workers
    Nodes -->|"Provides compute<br>for replicas"| Workers
    Workers --> Performance

    Partitions -. "Planned capacity" .-> Performance
    Nodes -. "Elastic infrastructure" .-> Performance

End-to-End Scaling Chain

Workload demand first changes pod demand. AKS node scaling occurs only when Kubernetes cannot schedule the desired pods with the currently available requested CPU and memory capacity.

flowchart LR
    Demand["Workload Demand<br>Kafka Lag / API Traffic"]

    subgraph WorkloadScaling["Workload Scaling"]
        Scaler["KEDA / HPA"]
        Pods["Desired Pod Replicas"]
    end

    subgraph Kubernetes["Kubernetes Scheduling"]
        Scheduler["Kubernetes Scheduler"]
        Pending["Pending / Unschedulable Pods"]
    end

    subgraph Infrastructure["AKS Infrastructure"]
        ClusterAutoscaler["AKS Cluster Autoscaler"]
        Nodes["Additional Node Capacity"]
    end

    Throughput["Higher Processing Throughput"]

    Demand --> Scaler
    Scaler --> Pods
    Pods --> Scheduler

    Scheduler -->|"Capacity available"| Throughput
    Scheduler -->|"Insufficient requested<br>CPU / memory"| Pending

    Pending --> ClusterAutoscaler
    ClusterAutoscaler --> Nodes
    Nodes --> Scheduler

    Throughput -. "Reduces backlog / demand" .-> Demand

Burst Scaling Timeline

Reporting workloads are intentionally handled asymmetrically: scale up quickly when backlog develops, drain the burst, then scale down conservatively after demand remains low.

flowchart LR
    Normal["1. Normal Load<br>Low Kafka Lag<br>Minimum Workers"]

    Burst["2. Reporting Burst<br>Messages Arrive Rapidly"]

    Lag["3. Kafka Lag Increases"]

    PodScale["4. Worker Autoscaler<br>Adds Replicas"]

    NodeScale["5. AKS Adds Nodes<br>If Pods Cannot Schedule"]

    Peak["6. Processing Capacity<br>Reaches Burst Demand"]

    Drain["7. Backlog Drains"]

    Stabilize["8. Sustained Low Demand<br>Scale-Down Window"]

    Contract["9. Worker Replicas Reduce<br>Unused Nodes Removed"]

    Normal --> Burst
    Burst --> Lag
    Lag --> PodScale
    PodScale --> NodeScale
    NodeScale --> Peak
    Peak --> Drain
    Drain --> Stabilize
    Stabilize --> Contract
    Contract --> Normal

Scaling Layers

Link autoscaling operates at three related layers.

1. Kafka Capacity

Kafka partitions define how much consumer parallelism is available.

Increasing the number of worker pods does not guarantee increased throughput if Kafka cannot assign useful partition work to those replicas.

Partition counts should therefore be derived from anticipated burst throughput in Performance Model and revised through explicit capacity-planning activities as production load grows.

Partition count is intentionally not a reactive autoscaling mechanism because:

  • Kafka partition increases are operationally significant.
  • Partition counts generally cannot be reduced in place.
  • Changing partition counts may change key-to-partition mapping for future messages.
  • Excessive partition counts introduce broker and operational overhead.
  • Capacity should be available before a reporting burst rather than created in reaction to one.

Partition plans should include reasonable growth headroom so Worker autoscaling does not immediately reach a Kafka concurrency ceiling during a burst.

Useful Consumer Concurrency

The maximum useful replica count must be determined from the actual consumer topology rather than using a single universal formula.

For the common case where one consumer group consumes one topic:

maximum useful worker replicas = topic partition count

For a consumer group subscribed to multiple topics, useful concurrency can reflect the partitions available across the subscribed topics.

If one service process hosts multiple independent consumer groups, the useful replica ceiling may differ again because each group receives its own partition assignments.

Therefore each independently scaled Kafka workload must document:

  • Consumer group or groups
  • Subscribed topic or topics
  • Partition count per topic
  • Message key / partitioning strategy
  • Maximum useful pod concurrency

Partition/key distribution must also be monitored. A topic with many partitions may still provide poor parallelism if workload is concentrated on a small number of keys.

Application Deployment Model

Services that expose application REST endpoints and perform Kafka/background processing should use separate Kubernetes Deployments.

Example:

  • report-api

    • Workload Mode: Api
    • Handles synchronous UI/API traffic
    • Kafka consumers disabled
    • Maintains a non-zero minimum replica count
  • report-worker

    • Workload Mode: Worker
    • Handles Kafka/background processing
    • Application REST controllers disabled
    • Scales independently according to asynchronous workload

Both Deployments use the same application image.

Services that have only one responsibility do not need to be artificially split.

Kafka Worker Autoscaling

Kafka consumer workloads should use Kafka backlog as their primary scaling signal.

KEDA or an equivalent Kubernetes external-metric mechanism should expose Kafka consumer lag to the Kubernetes autoscaling system.

CPU and memory remain important resource and safety signals but should not be the sole basis for scaling event-processing workers. A service may have substantial queued work while waiting on database, storage, network, garbage collection, or other dependencies without maintaining consistently high CPU utilization.

The preferred autoscaling objective is not simply to maintain zero lag. Short-lived backlog during reporting bursts is expected.

Scaling should instead be tuned to maintain an acceptable backlog age or drain time for each pipeline stage.

For example:

required replicas ≈ queued work / (measured work per replica per minute × acceptable drain window)

Production measurements should be used to refine the actual scaler configuration.

Worker Minimum Replicas

Worker minimum replica count is service-specific.

Scale-to-zero may be appropriate when:

  • No work is expected outside discrete bursts.
  • Startup is fast.
  • No required scheduler or coordination responsibility exists.
  • The first queued event may tolerate pod startup delay.

A non-zero minimum should be retained when:

  • The service has expensive warm-up or in-memory compilation/cache behavior.
  • Near-immediate event processing is required.
  • Startup creates significant load on databases or dependent systems.
  • The process owns background responsibilities that must remain available.

Measure Evaluation is a notable candidate for retaining warm capacity because measure evaluator compilation and caching can make cold starts more expensive.

REST/API Autoscaling

API deployments shall never scale to zero.

Interactive and operational APIs should retain an availability-appropriate minimum replica count independent of pipeline workload.

REST scaling should use Kubernetes HPA or equivalent metrics based on:

  • CPU utilization
  • Memory utilization where useful
  • Request rate/concurrency when available
  • Request latency when available and operationally reliable

API scaling limits should be independent from Kafka partition counts because API replicas do not consume partition assignments.

Separating API and Worker deployments prevents reporting bursts from controlling the amount of synchronous API capacity available to users.

Scale-Up and Scale-Down Behavior

Reporting workloads are burst-oriented. Scaling behavior should therefore be asymmetric.

Scale Up

Worker capacity should increase quickly when sustained backlog develops.

Scale-up stabilization should be minimal unless testing demonstrates that rapid increases overload a downstream dependency.

Multiple replicas may be added during a scaling interval when required to meet backlog-drain objectives.

Scale Down

Scale-down should be conservative.

Workers should only be removed after low demand has persisted for an appropriate stabilization period. Initial configurations should favor multi-minute stabilization rather than immediately removing replicas whenever lag falls.

Scale-down must account for work already being processed.

Services must support graceful termination so a pod can:

  1. Stop accepting new asynchronous work.
  2. Complete or safely relinquish in-flight work.
  3. Commit Kafka offsets only according to the service's successful-processing semantics.
  4. Cleanly leave the consumer group.
  5. Terminate within the Kubernetes termination grace period.

Consumers should tolerate at-least-once delivery and duplicate processing where appropriate.

Kubernetes Node Pool Autoscaling

Pod autoscaling and node autoscaling serve different purposes.

The expected scaling chain is:

  1. Kafka lag or API workload causes the workload autoscaler to request additional pods.
  2. Kubernetes schedules those pods onto existing node capacity where possible.
  3. If pods cannot be scheduled because adequate requested CPU/memory is unavailable, the AKS Cluster Autoscaler adds node capacity.
  4. Pending pods are scheduled onto the new nodes.
  5. When workloads later contract, underutilized nodes may be removed when their workloads can be safely rescheduled elsewhere.

Node scale-up should therefore be driven by Kubernetes scheduling demand rather than a direct rule such as "node CPU exceeds X percent."

Accurate pod resource requests are critical. If a workload substantially under-requests CPU or memory, Kubernetes may continue scheduling pods onto saturated nodes instead of creating the unschedulable-pod pressure needed to add capacity.

Resource Requests and Limits

Each scalable workload should define resource requests and limits based on measured production or performance-test behavior.

API and Worker deployments from the same service may use different resource profiles.

For example, a Measure Evaluation Worker performing CQL evaluation may require substantially more CPU than the API portion of the same application.

Resource settings should be periodically compared with actual telemetry and adjusted as the performance model evolves.

Node Pool Isolation

The platform should evaluate separating general/API workloads from compute-intensive reporting workers.

A possible model is:

  • General/system node pool

    • User-facing APIs
    • Admin BFF
    • tenant/configuration services
    • platform/system workloads
  • Processing node pool

    • Normalization workers
    • Measure Evaluation workers
    • Validation workers
    • other burst-processing workloads

This prevents a reporting burst from consuming all compute capacity required by interactive and operational services.

Dedicated workload pools should be adopted only where measurements demonstrate sufficient operational or cost benefit.

Scaling Ceilings and Downstream Constraints

Maximum replica count must not be based solely on available Kubernetes capacity.

For each workload:

maximum replicas = minimum of useful parallelism and all downstream safety limits

Potential constraints include:

  • Kafka partition/consumer concurrency
  • EHR query throttling
  • SQL/database capacity
  • Redis/cache capacity
  • Blob/storage throughput
  • terminology-service capacity
  • submission endpoint limitations
  • internal per-pod concurrency
  • cost guardrails

The Data Acquisition workload is especially sensitive to external EHR constraints. Increasing pod count must never bypass tenant/facility-specific request throttling.

Internal Concurrency

Replica count is only one dimension of concurrency.

Several Link services also have application-level concurrency controls, such as:

  • bounded internal work queues
  • worker-thread counts
  • maximum degree of parallelism
  • per-facility concurrency
  • SFTP parallel-processing limits

Autoscaling configuration and internal concurrency configuration must be evaluated together.

Doubling both pod count and per-pod parallelism may increase downstream traffic by substantially more than intended.

Observability

The Scale Dashboard described by Performance Model should correlate autoscaling behavior with actual throughput.

At minimum, monitor:

Kafka

  • Consumer lag by group/topic/partition
  • Oldest-message/backlog age where available
  • Produce rate
  • Consume rate
  • Partition/key distribution

Application

  • Desired/current/ready replicas
  • Processing throughput
  • Processing duration
  • Error rate
  • Retry rate
  • Dead-letter activity
  • Startup/warm-up duration

Kubernetes

  • CPU and memory requests
  • Actual CPU and memory utilization
  • Pending/unschedulable pods
  • Node count
  • Allocatable versus requested resources
  • Autoscaler decisions

Pipeline

  • Patients processed per minute
  • End-to-end patient processing duration
  • Time spent in each pipeline stage

A particularly useful operational correlation is:

Kafka lag → Worker replicas → Node count → Patients/minute

If replicas increase without a corresponding increase in throughput, another system dependency has become the limiting factor.

Capacity Planning

Autoscaling configuration should be derived from the performance model through the following relationship:

Required throughput → measured per-replica throughput → required consumer concurrency → required Kafka partitions → pod resources → required node capacity

These values should be revised from production measurements rather than assuming indefinitely linear scaling.

Open Questions

  1. Which services should initially be split into separate API and Worker Deployments?
  2. What is the complete service → consumer group → topic → partition topology for the current pipeline?
  3. How much Kafka partition headroom should be maintained above currently modeled burst requirements?
  4. Do the current message keys distribute burst workloads evenly enough to use the configured partition capacity?
  5. What backlog age or drain-time objective should apply to each major pipeline stage?
  6. Which worker services may safely scale to zero?
  7. What are the measured cold-start and warm-up costs for Measure Evaluation, Validation, Normalization, and other major workers?
  8. Can all Kafka consumers currently terminate safely during in-flight processing and consumer-group rebalancing?
  9. Which services have internal concurrency settings that must be coordinated with replica autoscaling?
  10. Which databases, caches, external endpoints, or downstream services establish maximum safe replica counts?
  11. Should API/general workloads and burst-processing workers use separate AKS node pools?
  12. How much pre-existing node capacity should be retained to reduce latency while AKS provisions new nodes?
  13. Should KEDA be the platform standard for Kafka scaling, or should Kafka metrics be routed through the existing Prometheus/HPA stack?
  14. What cost guardrails should prevent runaway scaling caused by malformed workloads, poison messages, or producer defects?

Relationships

flowchart LR
nauto_scaling_strategy_594E871C["Design: Auto-Scaling Strategy"]