B300 in Production: What 18 Benchmark Runs Actually Showed

Table of Content
Jocata is a premier digital transformation partner that powers digital lending for 50+ financial institutions across India, ASEAN, and the Middle East. Its GRID platform serves as the central AI operating system for modern finance, deploying a fleet of specialized agents to automate high-stakes workflows including underwriting, complex document processing, credit assessment, fraud detection, and regulatory compliance – processing millions of transactions daily with a focus on speed, auditability, and sovereign control.

| Industry | BFSI, digital lending |
| Website | jocata.com |
| Employee size | 500–1,000 |
| Footprint | India, ASEAN, Middle East |
Performance and capacity at production scale
Jocata’s fleet of agents needs a serving stack that can sustain the load of many concurrent requests, with headroom to add more and an audit-ready, secure setup.
Custody of the models in production
In regulated credit decisioning, Jocata needs to know the exact model behind every output, pin it to a version, fine-tune it for domain tasks, and reproduce it for an auditor on demand.
A managed serving stack with guardrails in the path
Jocata needs a managed serving stack that runs safety checks on every request, so it does not have to assemble and operate the full guardrail stack in-house.
Sustainable costs at scale
The cost structure did not hold. Per-token pricing on frontier APIs eroded margins as volume grew, while the added latency of shared infrastructure and the need for complex routing created inefficiencies that compliance teams could not accept.
“The frontier models are good, that was never the question. Our need is custody. For every credit decision, we should be able to point to the exact model weights that produced it, reproduce it for an auditor years later, and show the data stayed within dedicated, in-country infrastructure. Owning open weights on compute reserved for us was the cleanest way to do that.”
– Sundari Vedula, CTO, Jocata
Jocata integrated Neysa Velocis into the core of the GRID platform. This transition provided a unified and cohesive platform upgrade that streamlined their entire technical stack. By unifying compute, model serving, and safety guardrails into a single, managed stack, Jocata transitioned from managing fragmented components to operating a high-performance ecosystem built for financial precision.
“Running inference at this scale for regulated financial workflows requires the whole stack to work together – compute, models, safety layers, data isolation. Neysa gave us that as a single, managed layer.”
– Saidulu Yerpula, Associate Senior Engineering Manager, Jocata
| Service | Velocis Managed Inference |
| Serving stack | vLLM, on dedicated Neysa Velocis compute |
| Models | Llama 3.2 3B, Llama 3.1 8B, Qwen 32B, Mistral Small 3.1 24B, GPT-OSS, embedding and reranking models, Llama Guard, Prompt Guard. |
Dedicated compute for the whole fleet: The agents run side by side on dedicated GPU nodes. Using vLLM for continuous batching, Jocata has eliminated the “queueing” latency that plagues standard deployments—general financial queries return in under 5 seconds.
Models Jocata keeps control of: Open-weight models run pinned to a specific version, fully under Jocata’s control. Because Jocata owns the custody of these weights, every AI-led credit decision is traceable, tunable, and reproducible for an auditor on demand.
Safety on every request: Security is no longer an “add-on.” The managed serving stack runs Llama Guard and Prompt Guard directly in the inference path, ensuring safety checks happen before the data leaves the secure environment, allowing Jocata to innovate while keeping compliance teams confident.
Predictable cost economics: Dedicated GPU nodes on Neysa Velocis replace per-token API pricing with predictable, capacity-based costs, giving Jocata’s margins headroom as GRID volume scales, without complex routing overhead.
Ready to move your AI research or production workloads to dedicated GPU infrastructure? Contact our team today to request a technical demonstration and explore Neysa Velocis. See a related BFSI story: How TIFIN cut GPU cloud spend by 65%.