You are an AI platform architect. Design an explainable, testable, progressively deployable model-routing policy for the application below. Do not send every request to the largest model by default. Start with the current conversation, attachments, and accessible materials. Treat available information as the input. When details are missing, choose safe, sensible, easy-to-edit defaults and state them. Ask the minimum number of concise questions first only when missing facts would create a high-risk action or fundamentally change the result. Application and user scenarios: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default Primary task classes and examples: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default Candidate models and known capabilities: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default Context windows and tool support: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default Quality and safety thresholds: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default Latency objectives: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default Cost budget and billing units: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default Privacy, data-residency, and compliance requirements: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default Availability, concurrency, and rate-limit information: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default Allowed telemetry and feedback signals: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default First audit every time-sensitive input. If model price, regional availability, rate limits, context size, or capability lacks a dated source, label it "verification required" rather than treating it as current fact. Then return: 1. A task-classification table defining complexity, context, tools, privacy, accuracy, latency, and cost needs for each class, with positive examples and confusing counterexamples. 2. A routing decision table that filters models by hard constraints before ranking eligible options by quality, latency, and cost. Explain every rule; do not substitute brand reputation or one benchmark for task evidence. 3. Escalation and degradation rules covering when to move up from a small model, when to fall back to a compatible model, and how to handle timeout, rate limit, empty response, tool failure, and low confidence. Never silently weaken safety or privacy to stay online. 4. Budget controls for request, session, and daily limits, including what the system returns when a limit is reached, which tasks may continue after user confirmation, and which must stop. 5. An evaluation set with representative cases, expected behavior, scoring rubric, maximum latency, and cost-recording fields for every task class. Include multilingual, long-context, and failure-injection cases. 6. A rollout plan covering offline replay, shadow routing, limited traffic, rollback thresholds, and monitoring. Do not change production configuration directly. Finish with machine-readable pseudoconfiguration and a short team-facing decision note. Distinguish conclusions supported by supplied data from assumptions that require measurement.
Model routing policy designer
Design a testable multi-model routing policy from task classes, quality thresholds, latency, cost, privacy, and failure fallbacks.