19EU-region deployment rows
2short and long context suites
60reviewed scenarios per candidate
1,140model-scenario runs compared

Executive result

The public ranking prioritizes intent accuracy, then p95 end-to-end response time. In the short-context suite, Bedrock NVIDIA Nemotron Super 3 120B priority was the fastest of nine deployments with 30/30 accuracy at 338 ms p95. In long context, Regolo GPT-OSS 20B was the fastest of seven perfect-accuracy deployments at 401 ms p95.

The result is an important reminder: the right comparison begins with task success, then asks which equally accurate deployment responds most reliably at the slower end of real API traffic.

Neur-X production choice

Neur-X chose Infercom GPT-OSS 120B

Infercom GPT-OSS 120B delivered the strongest production balance within the consistently perfect-accuracy tier. It classified all 60 reviewed scenarios correctly—30/30 in both short and long conversational contexts—while keeping p95 end-to-end response time to 508 ms and 596 ms respectively at the 1.00x cost baseline.

Some deployments were faster in an individual suite, but either carried a higher relative cost or did not preserve perfect accuracy across both context lengths. For Neur-X, Infercom provided the most convincing combination of accuracy consistency, responsive tail latency and cost efficiency for a latency-sensitive production workload.

Why conversational intent detection is harder than it looks

A user rarely states an intent in a clean, isolated command. They correct themselves, refer to something said several turns earlier, mix a request with background information, or change direction midway through a sentence. In a voice interaction, the classifier also sits on a latency-critical path: even a correct answer can feel wrong if it arrives after the natural turn-taking window.

That creates three simultaneous requirements. The model must infer the right intent from conversational evidence, return output that obeys a strict machine-readable contract, and do both quickly enough for the surrounding agent to respond without an awkward pause. Cost matters too because this decision may happen on every turn, across every active conversation.

How Neur-X ran the benchmark

We benchmark models on production-shaped conversations rather than toy prompts or generic leaderboard questions. Every candidate in this snapshot was exercised through the deployed request path with the same reviewed scenarios, the real intent-detection prompt contract, model-specific settings and structured-output rules.

Controlled comparison

  • Two suites: 30 short-context and 30 long-context evaluations expose how added conversation history changes accuracy and latency.
  • Equivalent task: candidates receive the same scenario set and must satisfy the same output contract.
  • Real path: latency is measured across the deployed request route, not from an isolated model-only call.
  • Deployment-aware: standard and priority-routed variants are measured as distinct production paths.
  • Failure-aware: temporary API or provider failures are tracked separately from a model's reasoning capability.

The published tables report p50 latency for the typical response, p95 latency for the slower tail, intent accuracy out of 30 and a relative price multiplier. Rows are sorted first by intent success rate, then by the fastest p95 end-to-end API response time. Relative price breaks any remaining tie. Price multipliers are benchmark-relative, not public list-price quotations.

What the results tell us

1. Accuracy-first ordering changes the headline

Regolo GPT-OSS 20B recorded a very fast 333 ms p95 and the lowest relative price in short context, but its 28/30 intent result places it below every perfect-accuracy deployment. With longer history it improved to 30/30 and became the fastest perfect-accuracy result at 401 ms p95. More context did not slow its typical response in this test.

2. Perfect accuracy was common; production-grade balance was not

Nine deployments reached 30/30 in the short suite, and seven did so in the long suite. Their speed and cost profiles varied sharply. Infercom GPT-OSS 120B achieved 30/30 in both suites with p95 below 600 ms. Azure GPT-5 Mini also achieved 30/30 in both, but its p95 measured 2,483 ms short and 2,637 ms long. Accuracy alone would hide a difference users can feel.

3. Tail latency separates superficially similar candidates

Bedrock GPT-OSS 120B and Infercom GPT-OSS 120B both achieved perfect short-context accuracy. Their p50 figures were relatively close, but the Bedrock route recorded a 1,754 ms p95 versus 508 ms for Infercom. Measuring only an average or median would miss that reliability gap.

4. Longer context changed models differently

Several candidates improved with more history, including Regolo GPT-OSS 20B and Azure GPT-5.4 Mini. Others lost intent accuracy: the standard Bedrock Nemotron Super 3 result moved from 30/30 to 27/30, and its priority route moved from 30/30 to 25/30. Context-window capacity is not the same thing as task stability under realistic history.

5. Priority routing does not guarantee a better outcome

Priority variants were not consistently faster or more accurate than standard routes in this snapshot. Nemotron priority produced the fastest perfect-accuracy short result, yet its long-context accuracy was 25/30 versus 27/30 for the standard route. A deployment path should therefore be benchmarked as its own system, rather than treated as an interchangeable wrapper around a model name.

A practical shortlist by operating priority

  • Fastest perfect short-context result: Bedrock NVIDIA Nemotron Super 3 120B priority, at 338 ms p95 and 30/30 accuracy.
  • Fastest perfect long-context result: Regolo GPT-OSS 20B, at 401 ms p95, 30/30 accuracy and 0.50x relative price.
  • Neur-X production choice: Infercom GPT-OSS 120B, with 30/30 accuracy in both suites, 508/596 ms p95 and a 1.00x relative price.
  • Lowest measured latency: Scaleway Qwen3 Coder 30B A3B, at 173/169 ms p50 and 275/227 ms p95, with 29/30 short and 28/30 long accuracy.
  • Accuracy-first, slower option: Azure GPT-5.6 Luna, with 30/30 twice but p50 above one second in both suites.

Complete benchmark results

Rows are ordered by intent success rate from highest to lowest, then by p95 end-to-end API response time from fastest to slowest. Relative price is the final tie-breaker. The trophy identifies the deployment chosen by Neur-X for production. All latency values are milliseconds.

Intent Short — 30 reviewed scenarios
Model deploymentIntentp50p95Price
Bedrock NVIDIA Nemotron Super 3 120B (EU, priority)30/302523382.24x
Bedrock NVIDIA Nemotron Super 3 120B (EU)30/302583541.28x
Infercom GPT-OSS 120B (EU) Neur-X choice30/303505081.00x
Gemini 3.5 Flash Lite (EU, priority)30/305066491.98x
Bedrock GPT-OSS 120B (EU, priority)30/303537731.33x
Infercom Gemma 4 31B (EU)30/308421,7100.81x
Bedrock GPT-OSS 120B (EU)30/304001,7540.78x
Azure GPT-5.6 Luna (EU)30/301,2441,7970.82x
Azure GPT-5 Mini (EU)30/301,3942,4832.43x
Scaleway Qwen3 Coder 30B A3B (EU, FR)29/301732751.51x
Gemini 3.1 Flash Lite (EU, priority, legacy)29/308691,5423.82x
Regolo GPT-OSS 20B (EU, IT)28/301933330.50x
Regolo Qwen 3.8 27B (EU, IT, no reasoning)28/303874773.71x
Azure GPT-5.4 Mini (EU)28/306541,9903.35x
Regolo Mistral Small 4 119B (EU, IT)28/309372,3972.76x
Melious GLM 5.2 (EU router)28/301,3083,0276.72x
Regolo Gemma 4 31B (EU, IT)28/302,1646,4071.84x
Melious DeepSeek V4 Flash 0731 (EU router)27/301,0112,3250.89x
Bedrock Nova 2 Lite (EU)27/301,6232,9615.60x
Intent Long — 30 reviewed scenarios with extended history
Model deploymentIntentp50p95Price
Regolo GPT-OSS 20B (EU, IT)30/301764010.50x
Infercom GPT-OSS 120B (EU) Neur-X choice30/303815961.00x
Gemini 3.5 Flash Lite (EU, priority)30/305069341.98x
Infercom Gemma 4 31B (EU)30/308269990.81x
Azure GPT-5.4 Mini (EU)30/307391,2913.35x
Azure GPT-5.6 Luna (EU)30/301,1351,5370.82x
Azure GPT-5 Mini (EU)30/301,5162,6372.43x
Regolo Qwen 3.8 27B (EU, IT, no reasoning)29/304145773.71x
Bedrock GPT-OSS 120B (EU, priority)29/304078091.33x
Bedrock GPT-OSS 120B (EU)29/304081,3250.78x
Gemini 3.1 Flash Lite (EU, priority, legacy)29/301,0211,4073.82x
Scaleway Qwen3 Coder 30B A3B (EU, FR)28/301692271.51x
Melious GLM 5.2 (EU router)28/301,8644,1986.72x
Bedrock NVIDIA Nemotron Super 3 120B (EU)27/302433481.28x
Regolo Gemma 4 31B (EU, IT)27/302,3216,5011.84x
Melious DeepSeek V4 Flash 0731 (EU router)26/309631,5730.89x
Regolo Mistral Small 4 119B (EU, IT)26/308841,8872.76x
Bedrock Nova 2 Lite (EU)26/302,0163,3555.60x
Bedrock NVIDIA Nemotron Super 3 120B (EU, priority)25/302434192.24x

Model and deployment names follow the benchmark export, with redundant route variants consolidated for the public comparison. Settings were low reasoning unless the label states otherwise; Nemotron used the benchmark's Shape v2 contract.

How to read—and not overread—this benchmark

This is a task-specific snapshot, not a universal ranking of model intelligence. It answers a narrower and more useful question: which measured deployment offered the best operating balance for Neur-X's intent contract, scenarios and request path at the time of testing?

  • Thirty scenarios per suite can expose meaningful regressions, but cannot represent every language, domain or conversational edge case.
  • Latency varies with provider load, region, routing, network conditions and service updates. Repeated refreshes matter more than one permanent number.
  • Relative cost reflects this benchmark's workload and normalization. It should not be read as a general price comparison.
  • A model update, provider change or prompt-contract revision can materially change the ranking.
  • Temporary API failures are operationally important, but they should not be mislabeled as a model reasoning error.

What this changes for production AI teams

First, evaluate the complete deployment, not only the model family. Second, keep a long-context suite: a model that is reliable on an isolated utterance can behave differently once real history is attached. Third, optimize for tail latency as well as median latency. Finally, maintain a repeatable test contract so model switching becomes an evidence-based engineering decision rather than a brand preference.

At Neur-X, this benchmark is part of a continuous evaluation loop. Candidates can be retested through the same path as providers, model versions, prompts and traffic conditions evolve. That is how we pursue fast customer interactions without treating accuracy, reliability or cost as an afterthought.

Build and test reliable AI interactions

Neur-X helps teams design voice and support agents, test production-shaped conversations and improve the complete customer interaction loop.

Discuss your use case