← Back
Contact
Request a diagnostic
Tell us what you are running and what it is doing wrong. We will tell you whether we think there is anything to find — including when the answer is no.
What helps us answer quickly
- Hardware — GPU model and count, memory, interconnect
- Model — exact model and precision, and how it is quantized
- Serving stack — vLLM / TensorRT-LLM / other, and the version
- Load — context length, output length, concurrency
- What you observe — throughput, time to first token, where it hurts
None of that is mandatory. It just shortens the first conversation.
Before you write
We are not the right people if you already run a dedicated performance team, or if the decision is going to be made purely on the lowest quoted number. We are a fit if you own the hardware, need it to actually perform, and want numbers you can check.