Causal evidence
X→V→Y,pe∼Beta(αe,βe) Decisions X affect runtime mediators V, which affect serving outcomes Y. Each graph edge carries a belief updated from supporting and contradictory deployment evidence. Its confidence and normalized uncertainty are:
ce=αe+βeαe,Ue=(αe+βe)2(αe+βe+1)12αeβe Scoped mechanisms group these relationships into testable hypotheses. Confidence guides future proposals; uncertainty helps identify informative deployments to explore.
Job outcomes and priorities
Let Li be job i’s deployment ladder: the set of configurations, environments, and replica groups assigned to it. Its predicted outcome vector and objective weights are:
yi(Li)=(cost/token,p99 TTFT,p99 TPOT,throughput,SLO margin) wij≥0,j∑wij=1 Cost, TTFT, and TPOT are minimized; throughput and SLO margin are maximized. The planning algorithm normalizes each outcome against a job-specific ideal reference zij⋆ using scaleσij:
gij(Li)={(yij(Li)−zij⋆)/σij,(zij⋆−yij(Li))/σij,metric minimized,metric maximized. The augmented Tchebycheff job score is:
Ji(Li)=−[jmaxwijgij(Li)+ρj∑wijgij(Li)] The largest weighted gap drives the score. The sum term, weighted byρ, distinguishes candidates with the same maximum gap. Higher scores indicate closer alignment with the job’s weighted objectives.
Learning value, serving risk, and switching cost
Exploration rewards testing uncertain edges and mechanisms. Activation indicators count whether any deployment in the plan tests a relationship:
Expl(P)=e∈E∑Ae(P)Ue+ωmm∈M∑Am(P)Um Ae(P) andAm(P) activate each edge or mechanism once per plan. ωm weights mechanism uncertainty separately from edge uncertainty.
Breach estimates come from prediction-error bands calibrated against observed residuals. Per-objective estimates are combined using the implementation’s independence approximation:
pibrk(Li)=1−j∈Si∏(1−pjbrk) Si contains the job’s constrained serving objectives. Changing a running deployment also incurs:
Switchi=Ccold+Cparallel+Ckill+Crisk These terms account for loading and warm-up, concurrent execution, draining and termination, and transient underperformance. They are zero when the deployment is retained.
Cluster objective and selection
A plan P={Li} assigns a ladder to every active and pending job. Retaining the previous ladder keeps an active deployment; an empty ladder defers a pending job. The cluster score is:
Φ(P)=(i,Li)∈P∑[Ji(Li)−γpibrk(Li)−λtSwitchi(Li(t−1),Li)]+βtExpl(P) γ weights serving risk,λt weights reconfiguration cost, and βt weights learning value. Exploration and switching weights adapt across planning cycles.
The optimization goal is:
Pt∈argP∈FtmaxΦ(P) Ft contains plans satisfying resource capacity, physical deployment, serving-risk, active-swap-budget, quota, and priority constraints. Search is bounded, so the planning algorithm commits the highest-scoring feasible plan it finds rather than an exhaustive global optimum. New telemetry updates the evidence for the next cycle.
Read the research behind the formulation