Blog · Benchmark
Same model. Same GPU. 32× faster.
AI teams spend a lot of time choosing the right model, GPU, and serving stack. Once the system is in production, another question matters just as much: are you running that model in the best way for your workload?
LOFT long-context benchmark
Qwen 3.8 27B · vLLM · 1× A100-80GB · same model, same hardware
Time to first token
32× faster
Est. cost per 1,000 requests
97.5% lower
Output quality
12 / 12
outputs byte-identical under deterministic settings
We tested this using LOFT, Google DeepMind’s long-context evaluation framework. On the same 27B model (Qwen 3.8, served with vLLM) and A100-80GB GPU, WisePlan reduced time to first token from 85.8 seconds to 2.7 seconds. Estimated cost fell from $68.93 to $1.70 per 1,000 requests. Under deterministic settings, 12 out of 12 outputs were byte-identical.
The model didn’t change. The hardware didn’t change.
The execution plan did.
Why long context?
Long context is becoming part of everyday AI: documents, codebases, enterprise knowledge, agent sessions. It is also expensive. As context grows, the model has more input to process before it can respond, increasing compute, latency, and cost.
That’s why we chose LOFT. We didn’t want to design a benchmark around our technology. We wanted a public framework built to stress long-context workloads.
What we found points to a broader problem: running a good model and running it efficiently are not the same thing.
Platform optimization is not workload optimization
Inference infrastructure has improved enormously. Modern platforms offer faster runtimes, better scheduling, higher GPU utilization, and sophisticated serving systems.
But these improvements largely optimize the platform across many applications. They don’t answer a different question: what is the right execution plan for your workload, given its traffic patterns, latency requirements, cost constraints, and behavior?
Platform optimization
Makes the serving stack faster for everyone, across many applications.
Workload optimization
Finds the right execution plan for your traffic, latency targets and budget.
Two companies can use the same model and hardware and still need very different execution plans.
That’s why we built WisePlan, TokenWise’s proprietary optimization engine. WisePlan optimizes how your workload runs for the outcomes that matter to you.
The underlying optimization is deliberately abstracted away. You shouldn’t need a team of inference experts continuously retuning infrastructure as your workload changes.
The benchmark that matters is yours
LOFT gave us an independent way to test WisePlan. But ultimately, we care about a different benchmark: your production workload.
If you’re self-deploying models and running your own inference infrastructure, we’d like to benchmark it. What application you’re building doesn’t matter. What matters is whether your existing inference can run better.
- Keep your model.
- Keep your application.
- Keep your quality bar.
WisePlan optimizes how your workload runs for the outcomes that matter to you.
If there isn’t meaningful room to improve, that’s useful to know. If there is, we’ll show you the difference in latency and dollars.
Results reflect one controlled configuration using the LOFT evaluation framework and are not representative of expected performance for every workload. Performance varies by model, hardware, traffic patterns, context, and application requirements.