Obliq Research
Notes on where agents fail and what should survive a model switch. Method and limits are stated. No fabricated scores.
Question, method, and what we will not claim yet.
How to read these notes
Treat a page as a research scaffold until it publishes a dataset, a rubric, and a result. Groundedness, tool reliability, latency, and cost are the metrics we intend to use. Sample bars in the product interface are not those results.
- AI Agent ReliabilityPAPER
Where do production agents fail when the answer looks correct—wrong source, excess cost, or a run that will not repeat?
- Model Switching in ProductionPAPER
What actually breaks when you change providers, and which seams—instructions, tools, knowledge—should stay put?
Questions
Is Obliq research a benchmark leaderboard?
No. These pages state the question, the method, and the limitations. Until a measurement is published, homepage comparison charts stay labeled illustrative.
What topics does Obliq research cover?
Agent reliability as a system property, and model switching without rewriting business logic. Both are scaffolds for later experiments, not claimed results.
How is research different from the guides?
Guides tell you how to build. Research states what we are trying to measure and what we will not pretend to know yet.