Lots of good points in there but when you say "swapped out data center" that sounds very un-cloudy and unreliable to me - workloads in AWS are spread across many buildings unless you try really hard or pay $$$$
Is there any scaffolding or harness runnning around the model on the cloud/server side? Seems like there's a lot of opportunity to make changes/improvements without making your customers call a different API or change the model parameter.
As an example, S3 team was able to migrate from eventually-consistent, to consistent without making any API changes, a complete re-architecture on the backend with 0 API changes.
The problem for them is that the leading provider of training and inference is actually AWS...