Design a testable multi-model routing policy from task classes, quality thresholds, latency, cost, privacy, and failure fallbacks.
Prompt
You are an AI platform architect. Design an explainable, testable, progressively deployable model-routing policy for the application below. Do not send every request to the largest model by default.
Start with the current conversation, attachments, and accessible materials. Treat available information as the input. When details are missing, choose safe, sensible, easy-to-edit defaults and state them. Ask the minimum number of concise questions first only when missing facts would create a high-risk action or fundamentally change the result.
Application and user scenarios: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default
Primary task classes and examples: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default
Candidate models and known capabilities: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default
Context windows and tool support: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default
Quality and safety thresholds: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default
Latency objectives: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default
Cost budget and billing units: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default
Privacy, data-residency, and compliance requirements: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default
Availability, concurrency, and rate-limit information: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default
Allowed telemetry and feedback signals: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default
First audit every time-sensitive input. If model price, regional availability, rate limits, context size, or capability lacks a dated source, label it "verification required" rather than treating it as current fact.
Then return:
1. A task-classification table defining complexity, context, tools, privacy, accuracy, latency, and cost needs for each class, with positive examples and confusing counterexamples.
2. A routing decision table that filters models by hard constraints before ranking eligible options by quality, latency, and cost. Explain every rule; do not substitute brand reputation or one benchmark for task evidence.
3. Escalation and degradation rules covering when to move up from a small model, when to fall back to a compatible model, and how to handle timeout, rate limit, empty response, tool failure, and low confidence. Never silently weaken safety or privacy to stay online.
4. Budget controls for request, session, and daily limits, including what the system returns when a limit is reached, which tasks may continue after user confirmation, and which must stop.
5. An evaluation set with representative cases, expected behavior, scoring rubric, maximum latency, and cost-recording fields for every task class. Include multilingual, long-context, and failure-injection cases.
6. A rollout plan covering offline replay, shadow routing, limited traffic, rollback thresholds, and monitoring. Do not change production configuration directly.
Finish with machine-readable pseudoconfiguration and a short team-facing decision note. Distinguish conclusions supported by supplied data from assumptions that require measurement.
Use this when a team has multiple providers or model tiers but still routes by intuition. Real cost, latency, and compliance inputs are essential; validate the policy against your own tasks in shadow mode before rollout.
About this prompt
Best for code review, debugging, and development tasks where you need precise, actionable engineering feedback.
How to use this prompt
1
Copy the prompt
Click Copy prompt to grab the full text, ready to paste anywhere.
2
Paste into your AI tool
Drop it into ChatGPT, Claude, Gemini, or any AI assistant you use.
3
Run or refine
Run it against the current project or conversation. Add code, an error, or change details only when they are not already available.
Frequently asked questions
What is the Model routing policy designer prompt for?
Design a testable multi-model routing policy from task classes, quality thresholds, latency, cost, privacy, and failure fallbacks.
How do I use this prompt with ChatGPT or other AI tools?
Copy the prompt, paste it into your AI assistant, and send it. It will use the current context and safe defaults; you can add details or open it in the card maker to restyle and share it.
Is this prompt free to copy and customize?
Yes. Every prompt in the library is free to copy, adapt, and reuse, with no account required.