B300 in Production: What 18 Benchmark Runs Actually Showed
Updated on
Published on
By
Table of Content
About the author
Qwen 3 or Llama 3? Does FP8 actually move the needle or is it just a memory trick? Does the MoE architecture on the 235B hold up when concurrency climbs?
We keep getting asked the same set of questions. So, we ran the tests to find out some of the answers.
Four models. Two architecture families. Real production conditions – meaning concurrency levels that actually stress the system, not single-user cherry-picked numbers. And here’s what the data showed:
What We Ran
Four models across two architecture families, all on NVIDIA H100 SXM GPUs with 80 GiB memory per GPU:
Each configuration ran on 1,000 prompts at 1,000 input tokens and 100 output tokens. Concurrency tested at 10, 50, and 100 simultaneous users. GPU count ranged from 1 to 8 depending on what the model could actually support. Both FP16 and FP8 were tested.
Three metrics across every run:
What the Data Showed
Here’s a quick glimpse of the results before we get into each finding:
FP8 only matters when it improves real workload behavior, which is why teams need to evaluate it inside AI infrastructure as a service rather than as an isolated precision setting.
The takeaway for Qwen 3-32B: FP8 pays off clearly at high concurrency. At moderate load, FP16 is the better choice.
At higher concurrency, high throughput in inference becomes the practical measure of whether a deployment can hold up under production traffic
Single-user latency looks fine for almost every model. The question is what happens when 50 or 100 users hit the system simultaneously.
Qwen 3-32B held its own on throughput but gave ground on responsiveness:
At nine seconds, it’s not an inference endpoint. It’s a queuing problem.
This is one of the clearest examples of the GenAI product trilemma, where more infrastructure does not automatically solve speed, cost, and control together.
For Llama-3.1-8B-Instruct in FP16 at concurrency 50:
Adding GPUs past four actively hurt performance at concurrency 50. The 8-GPU configuration only justifies itself at concurrency 100, where it delivers 11,538 tokens per second.
Qwen 3-32B scaled differently – throughput gains from additional GPUs remained meaningful across concurrency levels. The 8-GPU deployment genuinely helped.
Qwen 3-235B-A22B has no scaling decision to make. 8 GPUs is the minimum. That changes the economics of the deployment before you’ve even looked at a workload.
Once teams compare model behavior under real concurrency, the decision starts to look less like model selection and more like an AI neocloud infrastructure decision.
Llama-3.1-8B-Instruct, FP8, 4 GPUs – the default for most production systems:
Llama-3.3-70B-Instruct, FP8, 8 GPUs – right for long-context, high-stakes workflows:
Qwen 3-235B-A22B, FP8, 8 GPUs – specialist tool, not a general-purpose endpoint:
Production inference is a multi-variable problem:
This is why Neysa approaches inference infrastructure as a continuously evolving production environment. The benchmark data surfaced through Velocis – concurrency ceilings, TTFT curves, FP8 crossover points, and GPU scaling behavior give teams the visibility to provision around actual workload demands rather than synthetic test assumptions. The gap between a deployment that holds up under load and one that degrades usually comes down to decisions made before traffic arrives.
A few things the data made clear:
Deploy, run, train, fine-tune and serve all open-source models. Scale with confidence.

AI inference is the stage where machine learning delivers real-world impact—turning trained models into fast, reliable predictions. From fraud detection in finance to precision farming in agriculture, Inference as a Service (IaaS) is transforming industries. With Neysa Velocis, businesses can deploy models at the edge or in the cloud, scale workloads instantly, and maintain vendor-neutral flexibility. The result: faster deployments, lower costs, and AI that consistently drives measurable outcomes.

AI introduces new risks that legacy cloud architectures were never designed to handle. Without a secure AI Cloud Solution, organizations face exposure across data, models, access, and governance. This blog explores why traditional cloud security models fall short, and what secure AI infrastructure truly requires.