Model Switching in Production

Changing providers should not rewrite the agent. This note asks which seams still break. It does not publish a score.

What fails when a production system changes models?

Teams usually discover the break in production: a tool call the new model formats differently, a retrieval snippet the old model tolerated, an instruction the new model ignores. Those failures get blamed on “the model” because the agent, the corpus, and the tools were never separable enough to test one variable.

What the experiment will hold constant

A fair switch keeps the task, the knowledge scope, and the tool grants fixed, and changes the endpoint. The record is two executions of the same input. Quality, latency, cost, and groundedness are read from those traces. If the agent definition had to change for the new provider to run, the architecture failed before the model did.

Method, and why there is no score yet

This page is a research scaffold. Future work will publish the task set, the rubric, and the limitations in the same place as any number. Until then, do not treat product UI samples as this experiment. Obliq can already route and replay; that capability is not a published win rate.

What you can do before the paper exists

Run the comparison yourself. Freeze the agent, point it at the candidate endpoint, replay known inputs, and read the traces. That is the operating practice this research is trying to make measurable.

Questions

What is model switching in production?

It is changing the model endpoint behind a live workflow—provider, size, or a private deployment—without treating the rest of the system as disposable.

What should not break when the model changes?

Agent instructions, knowledge scope, and tool policy should stay. If those are written into a vendor SDK, the switch is a rewrite. The research question is which other seams still fail even when those are stable.

Does this page report a measured win rate?

No. It is a research scaffold. Future experiments will publish method and limitations. No win rate is claimed here.