Inferex: A benchmark-driven LLM inference engineering platform
Jun 2026
Inference Engineering
LLMOps
An LLM inference platform built layer by layer, from a FastAPI gateway and Quantisation benchmarks to vLLM vs SGLang engine comparison on an RTX 5090 GPU, Observability with Grafana and Prometheus, Docker and Kubernetes for Deployment.