Skip to main content
A scale target connects a capacity pool to infrastructure that can change the supply behind its destinations. Each target identifies one scalable resource, the range Arklow may operate within, and the connection used to observe and change it.

Scale targets and autoscaling

Many platforms already include an autoscaler. Arklow can work with that controller or request a replica count directly, depending on the target type. Admission control remains active while new capacity provisions. Work can wait during that interval and resume as the destination becomes ready.

Metric binding

A pool holds one scale target per resource that serves its work. The binding tells each target which measurements are about its resource: a query from one of your metric sources, and the tag value that identifies the resource in what the query returns.
Let’s follow one action from arrival to the target that serves it:
  1. An action tagged {customer_id=C1, model_id=M1} routes to a destination in this pool.
  2. Its metric policy matches those tags to series tags by name, selecting request_latency{customer_id=C1, model_id=M1}.
  3. The target matching model_id=M1 watches that series.
  4. When the work pushes that series into pressure, reflecting the actual part of that infrastructure, model_id=M1 grows.
The model_id on the action reached the target through the series: the policy carried the value across on the shared name, and the target matched it. The customer_id traveled the same way but does a different job; it says whose work the series describes. One resource serves many customers. Two actions, {customer_id=C1, model_id=M2} and {customer_id=C2, model_id=M2}, produce two series: request_latency{customer_id=C1, model_id=M2} and request_latency{customer_id=C2, model_id=M2}. Both carry model_id=M2, so both belong to target model-b, and pressure on either one grows it. A customer with a private resource has a different model_id on their series, and with it their own target.

Pool ownership

Every scale target belongs to one capacity pool. The target should change capacity that serves the destinations in that pool. For example, a pool containing destinations backed by one model deployment can own the target for that deployment. Destinations backed by unrelated deployments belong in separate pools with separate targets. One pool can have several scale targets when separate components contribute capacity to the same work. A metric binding can identify the component represented by each target.

Write control

While the scale-targets member of Write Control is off, adjustments are recorded and not written. While it is on, Arklow writes replica counts within the target’s configured range.

Providers

Baseten

Adjust the replica range for a model deployment.

Together AI

Adjust the replica range for a dedicated endpoint.

Google Cloud

Control a managed instance group through its autoscaler or target size.

HTTP scale API

Connect another platform through a small HTTPS contract.

Attach a provider target

1

Open the credential

Create or open the credential that can read the provider account.
2

Enable the provider

Under Scale target providers, enable Baseten, Together AI, or Google Cloud managed instance groups.
3

Sync the inventory

Click Sync now, then open Scale Targets.
4

Choose the resource

Select the deployment, endpoint, or managed instance group under Provider inventory and click Attach.
5

Select capacity

Choose the destinations this target serves, or an existing pool. Choosing destinations creates a new pool for the target, with those destinations as its members.
6

Set the replica range

Enter the minimum and maximum replica counts.
7

Bind a metric if needed

Select a metric query. Add a tag projection when the query covers several scalable resources.
8

Attach

Attach the target, then review its observations and adjustment history.
The two HTTP scale target kinds cannot enumerate customer-implemented endpoints. Create those targets with New manual target.