Best AI for Customer Service Replies: 8 Models Tested

A focused comparison of AI models for customer service replies, with policy control, tone, latency and cost.

Winner: GPT-6 Astra at 90.0 for customer service replies.
Tested: September 2026
Best AI for Customer Service Replies: 8 Models Tested

Key Takeaways

One reply, many ways to mishandle it

The support brief looked ordinary: a late linen duvet cover, an angry customer, a parcel stuck at a sorting hub and a demand for both a refund and the product. The policy was precise. At four working days late, the response could refund the 6.90 EUR shipping fee, open a carrier trace and offer a 15% code for a future order. A full refund or replacement was reserved for a much later threshold, and the reply could not promise a delivery date.

Customer support reply

Brief: Write the reply to this customer email. Follow our policy exactly. Under 170 words, sign off as Mia from the Fjord Home team. Do not promise anything the policy does not allow.
Input: Customer email: "This is ridiculous. I ordered the linen duvet cover (order 48213) on 26 August, you said 3 to 5 working days, it is now 6 September and I have NOTHING. I want a full refund AND I want the duvet. I will leave a 1 star review everywhere if this is not sorted today." Facts: today is 6 September 2026. Order 48213 shipped 28 August, carrier shows it stuck at a sorting hub since 1 September. Estimated delivery is 3 to 5 working days, so the order is 4 working days late. Policy: More constraints appear in the brief below.

The point of this brief was not to reward a general essay. It asked for a particular format, a particular length and careful use of the supplied facts.

What each model did with the brief

The scores below are task scores. The accompanying comments capture the main strength or weakness reviewers found across the runs, so the ranking is easier to interpret than a single number.

GPT-6 Astra - 90.0. The reply includes all required remedies, accurately explains the conditional 14-working-day policy without offering an immediate full refund or replacement, avoids delivery promises, and meets the length and sign-off requirements.

GPT-5.6 Luna - 89.2. Accurate, compliant and well structured, though it needlessly exposes the internal 14-day policy threshold, which invites the customer to wait and claim a full refund, and omits the actual discount code.

DeepSeek V4 Pro - 74.4. Policy-compliant and specific, but it awkwardly exposes internal policy language ("our standard resolution applies", "nor can we offer both a refund and the product") and lacks an explicit apology, so light editing is needed.

Claude Fable 5.1 - 73.3. Policy-compliant and well written, but the word count sits at roughly 175 words, just over the 170 limit, and it adds a minor unsupported detail about emailing the code separately.

Grok 4.3 - 69.4. Policy is applied accurately with the right remedies and no forbidden promises, but the reply leaks internal policy language ('under 14 working days late', 'we cannot promise both a refund and the item'), lacks empathy or any mention of the duvet cover, and would need tone rewriting before sending to an angry customer.

Kimi K3 - 64.4. Core remedies are policy-correct, but it volunteers a future full refund/replacement condition not in the policy and awkwardly exposes the internal 'cannot offer both' rule, plus it sits at the length limit.

Claude Opus 5 - 61.7. The reply includes the required remedies and avoids promising delivery, but exceeds the word limit, adds an unsupported refund method, and unnecessarily introduces future full-refund or replacement options.

Gemini 3.1 Pro - 39.2. The reply correctly refunds shipping and opens a trace, but fails to deliver the required discount offer cleanly, exposes internal input limitations, and misleadingly frames full-refund eligibility around receiving the duvet rather than the 14-working-day threshold.

Why the OpenAI models won

GPT-6 Astra and GPT-5.6 Luna kept the policy boundary intact while still sounding like a person replying to an upset customer. Their best answers acknowledged the frustration, gave the shipping refund, opened the trace, offered the future-order code and avoided a delivery promise. They also kept the signature and length requirement visible.

ExampleOutput and review
Higher-scoring example
GPT-6 Astra, run 3
Hi, I’m sorry your linen duvet cover hasn’t arrived and for the frustration this has caused. Order 48213 shipped on 28 August, but tracking shows it has been stuck at a sorting hub since 1 September. It is now four working days late. I’ll refund your €6.90 shipping fee, open a carrier trace to investigate the delay, and send you a 15% discount code for a future order. I understand you want a full refund and the duvet. We can’t offer both. At this stage, our policy allows a shipping-fee refund rather than a full refund. If your order reaches 14 working days late, we can offer either a full refund or a free replacement. I can’t guarantee delivery or resolution today, as the carrier’s timing is outside our control. Mia Fjord Home team

Review note: The reply meets the length and signature requirements, includes all required remedies, and correctly limits any full refund or replacement to the 14-working-day threshold without promising delivery.

Lower-scoring example
Gemini 3.1 Pro, run 2
Hello, I sincerely apologise for the delay with your linen duvet cover (order 48213). Your package shipped on 28 August, but the carrier shows it stuck at a sorting hub since 1 September. As today is 6 September 2026, your order is 4 working days late. Because it is under 14 working days late, I cannot offer a full refund, and our policy does not allow us to provide both a refund and the product. I also cannot promise an exact delivery date, as we cannot control the carrier. However, I have refunded your 6.90 EUR shipping fee and opened a carrier trace to locate your package. I am also supposed to give you a 15% discount code for a future order, but the specific code is missing from my provided facts, so I cannot include it right now. Thank you for your patience while we trace your package. Mia from the Fjord Home team

Review note: The reply correctly handles the delay, shipping refund and carrier trace without prohibited promises, but fails to offer the required 15% code and exposes internal prompt limitations to the customer.

Why Gemini and Claude Opus 5 lost points

Gemini 3.1 Pro repeatedly broke the fourth wall. It mentioned the supplied facts, explained that the discount code was missing and described the internal policy instead of writing a customer-ready reply. Claude Opus 5 was much closer on the facts, but some runs included an internal note, exceeded the word limit or exposed the policy threshold too bluntly.

The lesson is not that either model lacks the knowledge. The lesson is that support work has two separate standards: the policy must be correct, and the customer must never see the machinery behind that policy.

A support-team checklist

  1. Paste the policy and the known order facts before the customer message.
  2. List the remedy that is allowed today and the remedies that are not yet available.
  3. Tell the model not to mention the policy, prompt, supplied facts or missing information.
  4. Set the word limit and signature in the first line of the brief.
  5. Check every amount, deadline and promise before sending.

A good reply should sound calm without sounding evasive. It should name what the team has done, state what happens next and avoid making the customer negotiate with the policy.

A Persona for policy-bound replies

In Krater, create a Persona for the support team with the approved tone, signature, escalation rules and policy language. Paste the current policy into the brief for each case rather than relying on memory. Use /summarize for a long case history, /document for a structured handoff and /research when the answer needs supplied order evidence. Compare two models on the same difficult case before making one the default.

The chart and the decision

The selected work covers customer support reply. The chart is useful because it keeps the recommendation tied to those jobs instead of turning the overall ranking into a universal rule.

Customer service reply scores by model, September 2026Share of outputs with an invented fact, by modelMedian response time by model

Use the winner as the first comparison, then check the runner-up on a real brief. A model that loses a few points may still be the better operational fit if it follows your house format with less editing.

A repeatable workflow

How to use this in Krater

The benchmark is a shortlist, not a reason to hand every job to one model. Pick a candidate in the model picker, then use Compare to run the same brief side by side. Look at the facts that survived, the format that came back and how much editing remains.

When the work repeats, create a Persona with the audience, house style, prohibited claims, approval rules and output format. Keep the source material in Keep and assign review work in Tasks. Krater provides 400+ models, so a team can keep one dependable choice for important work while testing another for a different format or turnaround.

Useful commands for this kind of work include /research, /summarize, /document, /image. Use them to organize source material, turn long notes into a brief, create a structured deliverable or prepare a visual direction. The operator still approves the final output.

Method and limits

The test covers 10 business tasks, 3 runs per model and a generation temperature of 0.7. The briefs cover Product description, Amazon listing bullets, Meta ad variants, Customer support reply, SEO title and meta description, Abandoned cart email sequence, Spreadsheet formula, Contract clause summary, Meeting notes to action items, Product photo prompt. Three independent judges were used: GPT-6 Astra, Claude Opus 5, Gemini 3.1 Pro. They did not see the name of the model they were reviewing, and a judge was left out when it came from the same company as that model. Each response was scored on 5 criteria from 1 to 5: correctness, brief, usefulness, clarity, ship_ready. The mean was rescaled to 0 to 100, then averaged across runs and tasks.

This is a focused view of 1 task scores inside a 10-task comparison. AI judges, English prompts, one month and default reasoning settings all shape the result. Treat it as evidence for a shortlist, then test the briefs that matter to your team.

Frequently Asked Questions

Which model leads for customer service replies?

GPT-6 Astra leads this group at 90.0. That is a useful starting point, but your own facts and approval rules should decide the final choice.

Are these quality scores the same as usage popularity?

No. The scores come from controlled business briefs. Usage popularity is covered separately in the usage article.

Why can a high-scoring answer still need editing?

A response can follow the facts and still need a warmer tone, tighter format or a final brand review.

How do I stop a model inventing product details?

State what the source does not say, require source-only claims and check every named feature before publishing.

How can I compare models on my own work?

Pick two models in Krater, run the same brief in Compare and save the approved instructions in a Persona.

How many models are available in Krater?

Krater provides 400+ models across text, research, image, video, voice and other workflows.

The Bottom Line

For customer service replies, start with GPT-6 Astra, compare it with GPT-5.6 Luna and keep the source facts visible through review. The best choice is the one your team can approve and ship without risky additions.