AI Guide
OpenAI Decisions API: Support Triage for Indonesian Teams
Updated 2026-10-07 · by Fahmi Fahreza
Start a Decisions API support pilot with internal queue suggestions. Define clear categories, prepare labeled Indonesian-language examples, and measure routing errors and review workload before enabling automatic changes.
What is OpenAI Decisions API?
OpenAI Decisions API helps applications classify and assess text or images. For Indonesian support teams, a useful starting point is suggesting which queue should handle a ticket. Refunds, account recovery, and customer replies still need their own operating rules.
OpenAI announced the beta on October 6, 2026, in its API changelog. When checked on October 7, 2026, the Decisions guide described a public beta using POST /v1/decisions, with gpt-6-luna as the only available model.
This article proposes a pilot design and synthetic examples. The examples have not been run against the API, are not client implementation results, and do not demonstrate measured accuracy or savings.
Choose the answer your workflow needs
Decisions offers three question types: predicate for a condition's probability, choice for selecting an option, and score for rubric-based assessment. Scores are probability-weighted averages of zero-based level indices and can be fractional. See the Decisions question types.
For this pilot, use one choice question named support_queue. Define four queues:
- pembayaran (billing): invoices, payment confirmation, or refund requests.
- teknis (technical): problems with a feature or using the product.
- pengiriman (shipping): tracking, shipment status, or undelivered parcels.
- review: out-of-scope issues, missing context, or multiple issues without a clear priority.
Write down the boundaries. “I cannot pay because the checkout page is blank” goes to technical in this design; “payment succeeded but I was charged again” goes to billing. If team members disagree, improve the definitions before changing the prompt. The model cannot settle a queue policy the team has not agreed on.
Need an order number, summary, and explanation in a JSON object? Consider Structured Outputs with a supported schema. Need the model to request a tool invocation? Use function calling; your application still executes the tool's code and checks authorization.
Build an Indonesian-language evaluation set
Begin with synthetic tickets to agree on the rules. Then use de-identified samples approved by the data owner. Include abbreviations, typos, mixed Indonesian and English, and follow-up messages that need earlier context.
These are expected labels under the pilot rules above, not API predictions. Indonesian examples are retained so the test reflects the intended audience:
| Sample message | Expected queue | What it tests |
|---|---|---|
| “Kak, status paket belum berubah sejak Senin.” | pengiriman | Recognizing shipment status without an explicit tracking-number keyword. |
| “Udah bayar, kok invoice muncul lagi?” | pembayaran | Distinguishing a billing complaint from checkout failure. |
| “Pas klik bayar layarnya putih, belum kepotong.” | teknis | Preventing the word “pay” from outweighing context. |
| “Pesanan sudah datang, tapi akun saya gabisa login.” | teknis | Avoiding a shipping label merely because an order is mentioned. |
| “Yang kemarin itu gimana ya?” | review | Refusing to guess missing context. |
| “Paket belum datang dan ada tagihan ganda.” | review | Recognizing two issues without forcing a single route. |
Ask two team members to label part of the sample independently. Resolve disagreements and record the human rationale. Keep tuning data separate from the final test set so evaluation measures more than the examples already used to improve instructions.
Decide when the system should defer
A choice answer includes probabilities and confidence; do not treat them as interchangeable. An answer can also be a refusal. Images require inline base64 data URLs, rather than HTTP(S) URLs or file_id. See the Decisions API reference and image-input constraints.
Keep the first version text-only. The following is a suggested application policy:
- Reject empty input and include only context needed for triage.
- Read the answer matching the question name, then validate its type and allowed option.
- Send
review, refusals, incomplete answers, timeouts, and API errors to a human queue. Do not convert failures into a default category that looks successful. - Hold results that do not meet the team's evaluation threshold. Calibrate thresholds on labeled data; a number such as 0.9 does not automatically mean 90% of tickets are routed correctly.
- Record the suggested route, instruction version, reviewer decision, and correction reason without copying unnecessary customer data.
Distinguish error costs. Sending a tracking question to technical support wastes time; missing a possible account takeover could be more serious. Define a separate escalation path for sensitive cases regardless of the general classifier's confidence.
Connect the workflow gradually
The conceptual flow is: receive a ticket, minimize its data, request classification, validate the response, apply the review policy, and write the suggested route to an internal note.
For n8n, a pattern to test is HTTP Request for the API request, If for the review condition, then Switch to map accepted labels. Set Switch's fallback to Extra Output and connect it to a human queue; its default behavior ignores unmatched items. Request errors still need a separate error path.
This is a proposed integration based on the nodes' documented functions, not a test against the API, your version, or your credentials. This article does not supply a built-in Decisions node or an import-ready workflow.
Start in observation mode: people continue routing tickets while the system only records suggestions. Compare the two without changing customer queues. Once quality is demonstrated, enable internal changes for agreed low-risk categories. Keep a simple route back to manual triage.
Use the ticket identity to prevent duplicate notes during retries. Limit retries, make failures visible to operators, and assign an owner to the review queue. A technically safe queue can still fail operationally if nobody reads it.
Measure before expanding automation
Compare the pilot with the existing manual process or keyword rules. Track:
- Routing errors by category, including the most costly types of mistake.
- Tickets held for review and staff capacity to handle them.
- Time until the right team receives a ticket, including queueing and corrections.
- Ambiguous cases that need more context.
- Actual API charges, maintenance effort, and review time.
Set acceptance criteria before seeing results. Do not remove the review queue just to raise the automation percentage. Repeat evaluation when customer language, products, or categories change. Model-call speed alone does not explain ticket resolution time.
Practice support triage with Fahmi
Contact Fahmi Fahreza about AI automation training with one support process you want to work through. Bring queue definitions, safe examples, and success criteria so the discussion starts with real work. Practical topics can be agreed around your team's needs and tool access.
For workflow foundations, read AI agents with n8n for business. For a team session, explore corporate AI training.
Frequently asked questions
How should a support team first try Decisions API?
Choose one channel and have the system suggest an internal queue. Compare its suggestions with support-team labels before allowing automatic ticket changes.
What threshold is safe for Indonesian-language tickets?
There is no universal number. Choose thresholds using labeled data representing your team's language, categories, and error costs, then test on data not used for tuning.
Can a triage result automatically approve a refund?
Not in this design. A billing label only identifies the reviewing team. Refunds, account recovery, and promises to customers need separate checks and approval.
What should a team prepare for a workshop with Fahmi?
Bring synthetic or de-identified tickets, queue definitions, escalation rules, and success criteria. Discuss training needs and tool access before the practical session.
Want a session like this for your team or event?