CodingFeatured

Model routing policy designer

Design a testable multi-model routing policy from task classes, quality thresholds, latency, cost, privacy, and failure fallbacks.

Prompt
You are an AI platform architect. Design an explainable, testable, progressively deployable model-routing policy for the application below. Do not send every request to the largest model by default. Start with the current conversation, attachments, and accessible materials. Treat available information as the input. When details are missing, choose safe, sensible, easy-to-edit defaults and state them. Ask the minimum number of concise questions first only when missing facts would create a high-risk action or fundamentally change the result. Application and user scenarios: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default Primary task classes and examples: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default Candidate models and known capabilities: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default Context windows and tool support: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default Quality and safety thresholds: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default Latency objectives: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default Cost budget and billing units: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default Privacy, data-residency, and compliance requirements: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default Availability, concurrency, and rate-limit information: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default Allowed telemetry and feedback signals: Infer it from the current conversation, attachments, or accessible project; when unspecified, use a safe, sensible, easy-to-edit default First audit every time-sensitive input. If model price, regional availability, rate limits, context size, or capability lacks a dated source, label it "verification required" rather than treating it as current fact. Then return: 1. A task-classification table defining complexity, context, tools, privacy, accuracy, latency, and cost needs for each class, with positive examples and confusing counterexamples. 2. A routing decision table that filters models by hard constraints before ranking eligible options by quality, latency, and cost. Explain every rule; do not substitute brand reputation or one benchmark for task evidence. 3. Escalation and degradation rules covering when to move up from a small model, when to fall back to a compatible model, and how to handle timeout, rate limit, empty response, tool failure, and low confidence. Never silently weaken safety or privacy to stay online. 4. Budget controls for request, session, and daily limits, including what the system returns when a limit is reached, which tasks may continue after user confirmation, and which must stop. 5. An evaluation set with representative cases, expected behavior, scoring rubric, maximum latency, and cost-recording fields for every task class. Include multilingual, long-context, and failure-injection cases. 6. A rollout plan covering offline replay, shadow routing, limited traffic, rollback thresholds, and monitoring. Do not change production configuration directly. Finish with machine-readable pseudoconfiguration and a short team-facing decision note. Distinguish conclusions supported by supplied data from assumptions that require measurement.

Inspiration source

This prompt is an original adaptation. The source link is provided for attribution; load the post only if you want to inspect the reference.

The X embed stays off until you choose to load it. Loading it may share your IP address and browser information with X.

Open original on X

Editor's note

Use this when a team has multiple providers or model tiers but still routes by intuition. Real cost, latency, and compliance inputs are essential; validate the policy against your own tasks in shadow mode before rollout.

About this prompt

Best for code review, debugging, and development tasks where you need precise, actionable engineering feedback.

How to use this prompt

  1. 1

    Copy the prompt

    Click Copy prompt to grab the full text, ready to paste anywhere.

  2. 2

    Paste into your AI tool

    Drop it into ChatGPT, Claude, Gemini, or any AI assistant you use.

  3. 3

    Run or refine

    Run it against the current project or conversation. Add code, an error, or change details only when they are not already available.

Frequently asked questions

What is the Model routing policy designer prompt for?

Design a testable multi-model routing policy from task classes, quality thresholds, latency, cost, privacy, and failure fallbacks.

How do I use this prompt with ChatGPT or other AI tools?

Copy the prompt, paste it into your AI assistant, and send it. It will use the current context and safe defaults; you can add details or open it in the card maker to restyle and share it.

Is this prompt free to copy and customize?

Yes. Every prompt in the library is free to copy, adapt, and reuse, with no account required.

More prompts from the Coding collection.

Coding

First-Principles Subtraction Review

Before calling the work done, revisit the goal and challenge weak assumptions and unnecessary complexity. Prioritize deletion, simplification, optimization, then automation, and leave good work unchanged.

first-principlescode-review
Release Notes Page
Coding

Release Notes Page

Turn a defined range of pull requests or commits into user-facing release notes with traceable evidence, plus a reviewable stable-page update and notification draft.

opsworkflow
New Hire Ramp Plan
Coding

New Hire Ramp Plan

Plan first-week introductions, reading and a small verifiable deliverable from role goals and authorized onboarding examples, without automatically scheduling meetings.

opsworkflow