الذكاء الاصطناعي وتعلّم الآلة
GPT-6 Model Routing: Optimize Cost per Successful Task
A practical GPT-6 routing policy for choosing Luna, Sol or Astra by task risk, success rate, latency and total workflow cost.
هذه المقالة متاحة حاليًا باللغة الإنجليزية.
A low-cost model call can make the workflow expensive when it triggers retries, manual review and delayed decisions. Sending every request to the strongest model creates a different cost and latency problem.
OpenAI's October 2, 2026 guide to the GPT-6 family recommends matching the model, reasoning effort and speed to the workload, then measuring task success, latency and cost per successful task. Enterprise teams need to turn that guidance into a routing policy they can test and govern.

Start with the accepted outcome
Token price is one input to the decision. The unit that matters to an operating team is the cost of an accepted result:
cost per accepted task = (model usage + tools + runtime + review + retries) / accepted tasks
This is an operating metric, not an OpenAI pricing formula. It forces the team to count the work around the model call. A cheap extraction that needs frequent correction may cost more than a higher-quality first pass. A complex analysis that uses a stronger model may still be economical if it reduces failed runs and senior review time.
Define “accepted” for each task family before comparing models. An invoice extraction may require every mandatory field, a valid total and no unsupported value. A code change may need passing tests and a reviewable diff. A research brief may need source coverage, correct claims and a clear record of uncertainty. One generic quality score will hide the failures that matter.
Route task families, not individual prompts
OpenAI currently describes GPT-6 Luna as suited to focused tasks at scale, GPT-6.1 Sol as a fit for complex coding, research and computer use, and GPT-6 Astra as the option for the hardest reasoning work. Its model-selection documentation treats these as starting points and recommends comparing the same inputs across settings.
Convert the workload into a small number of stable task families. Each family gets a default route, a measurable pass condition and an escalation rule.
| Task family | Sensible starting route | Evidence to collect | Escalate when | |---|---|---|---| | Structured extraction and triage | Luna with low or medium effort | Field accuracy, invalid output rate, review time | Required fields conflict or confidence checks fail | | Multi-source synthesis and technical delivery | GPT-6.1 Sol with medium effort | Factual accuracy, completion rate, tool failures, latency | Evidence conflicts or the task crosses a risk boundary | | Ambiguous, high-complexity analysis | Astra at the lowest effort that passes | Expert acceptance, unsupported claims, decision usefulness | A person must decide scope, policy or irreversible action |
These routes are evaluation hypotheses. They are not guarantees about a particular organisation's data, prompts or tools. Keep the lightest route that consistently meets the acceptance bar.
Separate four routing decisions
A production router should make four explicit choices instead of hiding them behind a single “use the best model” flag.
Model
Select the default model for the task family. The initial choice should reflect ambiguity, tool use and consequence, not the job title of the requester. A board-level document can contain simple extraction steps; an ordinary support ticket can contain a difficult policy exception.
Reasoning effort
OpenAI's guide maps low effort to routine work, medium to work requiring judgment, and high to difficult debugging or deeper analysis. Extra-high and maximum settings should stay only where tests show enough improvement to justify the additional time and cost. The API can change reasoning effort during a conversation without breaking the prompt cache, according to the guide.
Speed
Fast and Ultrafast modes change response speed and price independently of reasoning effort. Use a faster tier when waiting has measurable business cost, such as an interactive support flow. A nightly analysis normally needs a throughput and completion target rather than premium token speed.
Authority
Model capability does not decide whether the workflow may send, approve, delete or commit. Keep action authority in a separate policy. A route may escalate from Luna to Sol for better analysis while the final action still requires the same reviewer.
Use failure signals to escalate
Routing by prompt length or user seniority is easy to implement and weak in practice. Escalation should come from observable signals tied to the acceptance criteria.
Useful signals include:
- a schema validator rejects required fields;
- two authoritative sources disagree;
- a tool returns a partial or unknown outcome;
- the task needs data outside the approved boundary;
- the model reports an assumption that changes the recommendation;
- a deterministic calculation does not match the generated result;
- the requested action exceeds the route's authority.
Some failures need a stronger model. Others need a person, a corrected source or a retry of the same tool. Treating every failure as a model upgrade can raise cost without fixing the cause.
A worked routing example
Consider a hypothetical supplier-invoice workflow. This example illustrates the policy; it is not a Bayseian customer result.
Luna extracts the supplier, purchase order, dates, currency and totals into a schema. Deterministic checks compare the subtotal, tax and total and verify that the purchase order exists. A clean record moves to the review queue.
If the invoice and purchase order disagree, Sol receives the conflicting fields, source snippets and policy needed to prepare an explanation. It does not receive unrelated documents. If the issue involves a new bank account, a policy exception or an amount above the approval threshold, the workflow hands the case to a person. Astra might support a difficult analysis, but it does not inherit authority to approve the payment.
Record the cost and latency of every attempt, including the first extraction, any escalation and the review time. The result can then be compared with a single-model baseline.
Build the evaluation before the router
Create a representative set for each task family. Include ordinary work, edge cases and failures the team has already seen. Keep the expected outcome and review rubric separate from the prompt used by the model.
For every route, record:
- accepted-task rate;
- cost per accepted task;
- median and tail latency;
- retry and escalation rate;
- human review time;
- policy or permission violations.
Run the same cases across model and reasoning combinations. OpenAI's model-selection page recommends this direct comparison and advises keeping the lightest setting that meets the quality bar. Re-test when prompts, tools, reference data or model versions change.
The router also needs a fallback for outages and rate limits. Decide whether the workflow can wait, use an approved alternative or stop. Silent downgrade is risky when the backup route has not passed the same acceptance tests.
Control context and recurring cost
The October 2 guide also recommends prompt caching for recurring work and compaction for long conversations. OpenAI says cached input can cost less than uncached input, depending on the model. Put stable instructions, tool definitions and reference material before changing task details so the cache can be reused. Monitor where reuse breaks instead of assuming a high cache-hit rate.
Context reduction needs its own quality check. Remove material the task does not need while keeping the evidence required for the decision. For long-running work, verify that compaction preserves outstanding actions, approvals and source references.
Use the API deployment checklist alongside task-specific tests for monitoring, limits and production readiness. Availability, tools and reasoning settings can differ by product and model version, so confirm the current model catalogue and regional requirements before rollout.
A practical rollout sequence
Start with one task family and one default route. Run it in shadow mode beside the existing process. Review failures by cause, then add only the escalation paths that address those causes.
Publish the routing table as configuration with an owner and review date. Log the selected model, reasoning effort, speed tier, trigger and final outcome for each run. Set budgets on accepted-task cost and tail latency, not only on tokens. A route change should go through the same evaluation as a prompt or tool change.
Treat model routing as an operational policy with measurable routes, a named owner and action authority managed separately. Teams deciding which workflow to evaluate first can use our agent workflow selection guide. Bayseian's AI systems practice can help turn that candidate into a tested routing and control design.
مقالات ذات صلة
Google Cloud API Gateway MCP: What to Check Before Exposing REST APIs
Google Cloud can now expose REST operations as MCP tools through API Gateway. Before connecting an agent, review tool discovery, operation scope, authentication and the preview limits.
الذكاء الاصطناعي وتعلّم الآلةSecuring AI Agents With Runtime Boundaries: What NVIDIA OpenShell Adds
NVIDIA's new agent safety reference design puts policy enforcement outside the agent. Here is how enterprise teams can evaluate the boundary, its limits and the review steps that still matter.
هل تعمل على مشروع مماثل؟
لا عروض ترويجية، بل حوار عملي مع الفريق الذي يبني هذه الأنظمة ويشغّلها في بيئة الإنتاج.
ابدأ حوارًا معنا