Scale targets and autoscaling
Many platforms already include an autoscaler. Arklow can work with that controller or request a replica count directly, depending on the target type.
Admission control remains active while new capacity provisions. Work can wait during that interval and resume as the destination becomes ready.
Metric binding
A pool holds one scale target per resource that serves its work. The binding tells each target which measurements are about its resource: a query from one of your metric sources, and the tag value that identifies the resource in what the query returns.- An action tagged
{customer_id=C1, model_id=M1}routes to a destination in this pool. - Its metric policy matches those tags to series tags by name, selecting
request_latency{customer_id=C1, model_id=M1}. - The target matching
model_id=M1watches that series. - When the work pushes that series into pressure, reflecting the actual part of that infrastructure,
model_id=M1grows.
model_id on the action reached the target through the series: the policy carried the value across on the shared name, and the target matched it. The customer_id traveled the same way but does a different job; it says whose work the series describes.
One resource serves many customers. Two actions, {customer_id=C1, model_id=M2} and {customer_id=C2, model_id=M2}, produce two series: request_latency{customer_id=C1, model_id=M2} and request_latency{customer_id=C2, model_id=M2}. Both carry model_id=M2, so both belong to target model-b, and pressure on either one grows it. A customer with a private resource has a different model_id on their series, and with it their own target.
Pool ownership
Every scale target belongs to one capacity pool. The target should change capacity that serves the destinations in that pool. For example, a pool containing destinations backed by one model deployment can own the target for that deployment. Destinations backed by unrelated deployments belong in separate pools with separate targets. One pool can have several scale targets when separate components contribute capacity to the same work. A metric binding can identify the component represented by each target.Write control
While the scale-targets member of Write Control is off, adjustments are recorded and not written. While it is on, Arklow writes replica counts within the target’s configured range.Providers
Baseten
Adjust the replica range for a model deployment.
Together AI
Adjust the replica range for a dedicated endpoint.
Google Cloud
Control a managed instance group through its autoscaler or target size.
HTTP scale API
Connect another platform through a small HTTPS contract.
Attach a provider target
1
Open the credential
Create or open the credential that can read the provider account.
2
Enable the provider
Under Scale target providers, enable Baseten, Together AI, or Google Cloud managed instance groups.
3
Sync the inventory
Click Sync now, then open Scale Targets.
4
Choose the resource
Select the deployment, endpoint, or managed instance group under Provider inventory and click Attach.
5
Select capacity
Choose the destinations this target serves, or an existing pool. Choosing destinations creates a new pool for the target, with those destinations as its members.
6
Set the replica range
Enter the minimum and maximum replica counts.
7
Bind a metric if needed
Select a metric query. Add a tag projection when the query covers several scalable resources.
8
Attach
Attach the target, then review its observations and adjustment history.