FastAPI vs Triton Inference Server for healthcare inference on Kubernetes¶
Scope: the serving-stack choice between a Python web service (FastAPI) and NVIDIA Triton Inference Server for a small NLP model behind a regulated-data boundary, as reported in arXiv 2602.00053 (Ali, 2026), and the hybrid pattern it recommends (FastAPI gateway and de-identification in front, Triton for inference). Covers what dynamic batching changes, the reported latency and throughput table, an arithmetic audit of that table, a runnable closed-loop batching model, and the state of the reference architecture it builds on. It is a single-author, single-configuration study; general serving internals are in inference serving and continuous batching internals.
All benchmark figures below come from one Zenodo-sourced "Benchmarking Note" that the paper cites for its only results table. They were not re-measured here. The Python block checks the paper's arithmetic and models batching with assumed constants; it does not reproduce the paper's hardware results.
What it is¶
The paper compares two deployment shapes for a DistilBERT sentiment model on an AWS g4dn.xlarge (one NVIDIA T4, 4 vCPUs, 16 GB RAM), both containerised in one Kubernetes cluster:
- FastAPI baseline. PyTorch model on CPU behind a Python REST service.
- Triton. The model exported to ONNX, served on the T4 with dynamic batching, tested at batch size 1 (no batching) and batch size 16 (dynamic).
The recommended architecture splits roles: a FastAPI gateway handles OAuth2/JWT authentication, routing, and retries with bounded timeouts; independent preprocessing pods normalise input and strip or obfuscate protected health information (PHI); Triton pods execute inference and scale between 2 and 10 replicas with a Horizontal Pod Autoscaler at 60% CPU utilisation. The paper takes this design from a public reference architecture by D. Gopalan (Hugging Face Space digopala/ai-inference-architecture-healthcare, and a Zenodo record the paper cites as 10.5281/zenodo.16946461, which returned HTTP 403 to an automated fetch and was not checked).
Why use it¶
- Batching is the lever. A GPU serving one request at a time is mostly idle. Triton's dynamic batcher groups concurrent requests into one execution, which is where the reported 780 requests per second comes from.
- Separation of duties. Keeping authentication, schema validation, and PHI removal in a cheap CPU tier means the inference tier never sees raw identifiers and its logs stay cleaner; the paper states this as the reason to keep FastAPI in front.
- Model versioning and probes. Triton's model repository layout (
/models/<name>/<version>) supports hot reload and rollback, and readiness and liveness probes gate traffic during updates.
When to use it (and when not)¶
| Situation | Reasonable choice | Basis |
|---|---|---|
| Prototype, low concurrency, small CPU-runnable model | FastAPI alone | Lowest single-request latency in the paper's table (22 ms p50). |
| Sustained concurrent load on a GPU-worthy model | Triton with dynamic batching | Highest throughput row (780 req/s). |
| Regulated data in requests | Gateway and de-identification tier in front of the inference tier | Architectural argument in the paper; no measured hybrid numbers. |
| LLM serving | Neither as written | The paper is about a 66M-parameter encoder; LLM engines are in inference serving. |
The paper's conclusion that Triton is "mandatory" for production-scale throughput is stronger than a three-row comparison of CPU FastAPI against GPU Triton can support. The FastAPI row runs on CPU and the Triton rows on GPU, so the comparison mixes serving framework with hardware. No row serves the same model on the same device through both frameworks, and no row exercises the hybrid path the paper endorses.
Architecture¶
flowchart LR
C["Clinical client"] --> GW["FastAPI gateway<br/>OAuth2/JWT, routing, retries"]
GW --> PP["Preprocessor pods<br/>PHI de-identification, tensor normalisation"]
PP --> TR["Triton pods (2 to 10, HPA on CPU 60%)<br/>dynamic batching, ONNX on T4"]
TR --> PR["Prometheus metrics :8002"]
REG["Model repository<br/>/models/name/version"] --> TR
Triton listens on 8000 (HTTP), 8001 (gRPC), and 8002 (metrics).
How to use it¶
Reported results (paper Table I, peak-load test; 10, 50, and 100 concurrent users were tested but the table does not say which level it reports):
| Framework | Hardware | Batch mode | Batch size | p50 (ms) | p95 (ms) | Throughput (req/s) |
|---|---|---|---|---|---|---|
| FastAPI | CPU | none | 1 | 22 | 45 | 450 |
| Triton | GPU | none | 1 | 28 | 52 | 420 |
| Triton | GPU | dynamic | 16 | 34 | 60 | 780 |
The paper attributes Triton's single-request disadvantage to scheduling and host-to-device transfer overhead, and the wider p95 to requests waiting for a batch to fill or a window to expire. Batching in Triton is configured per model with dynamic_batching { preferred_batch_size, max_queue_delay_microseconds }; the queue delay is the direct trade between throughput and tail latency (see Triton's model configuration guide in the references). The paper does not state the preferred batch sizes or queue delay it used.
How to develop with it¶
The block below does two things. First it checks the quoted table: the ratios in the text, and Little's law (requests in flight equal throughput times latency) to see what concurrency each row implies. Second it runs a closed-loop model of one GPU with a batching window, with service time a + c * b milliseconds for batch size b. The constants a = 10, c = 1, and a 5 ms window are assumptions chosen for illustration, not fitted to the paper.
"""Table I arithmetic checks plus a closed-loop dynamic-batching model.
The model's service-time constants are illustrative assumptions, not measurements."""
from __future__ import annotations
import heapq
import numpy as np
# Table I of arXiv 2602.00053: (p50 ms, p95 ms, req/s)
TABLE = {"fastapi_cpu": (22, 45, 450), "triton_b1": (28, 52, 420), "triton_b16": (34, 60, 780)}
# 1. Ratios quoted in the text.
r = {k: v[2] for k, v in TABLE.items()}
print("triton_b16 / triton_b1 = %.3f" % (r["triton_b16"] / r["triton_b1"]))
print("triton_b16 / fastapi = %.3f" % (r["triton_b16"] / r["fastapi_cpu"]))
assert round(100 * (r["triton_b16"] / r["triton_b1"] - 1)) == 86 # text says 85%
assert round(100 * (r["triton_b16"] / r["fastapi_cpu"] - 1)) == 73 # text says 73%
assert r["triton_b16"] / r["fastapi_cpu"] < 2 # abstract: "nearly double"
# 2. Little's law: in-flight requests L = throughput * latency (p50 used as a proxy for mean).
for k, (p50, p95, tput) in TABLE.items():
print("%-12s in flight ~ %5.1f (implied mean latency at 50 users %.0f ms, 100 users %.0f ms; p95 %d ms)"
% (k, tput * p50 / 1000, 50000 / tput, 100000 / tput, p95))
# A 50- or 100-user cell would need mean latency above the reported p95, so the row is not that cell.
assert 50000 / r["triton_b16"] > TABLE["triton_b16"][1] and 100000 / r["triton_b16"] > TABLE["triton_b16"][1]
def simulate(users: int, max_batch: int, max_delay: float, a: float, c: float, horizon: float = 60.0):
"""Closed loop, zero think time. One GPU; service time a + c*b ms for batch b."""
ready = [(0.0, u) for u in range(users)]
heapq.heapify(ready)
t, done, lat = 0.0, 0, []
while t < horizon * 1000:
t = max(t, ready[0][0]) # wait for the first request
deadline = t + max_delay
while len(ready) > 0 and sum(1 for x in ready if x[0] <= t) < max_batch and t < deadline:
nxt = min(x[0] for x in ready if x[0] > t) if any(x[0] > t for x in ready) else deadline
t = min(nxt, deadline)
batch = [heapq.heappop(ready) for _ in range(min(max_batch, sum(1 for x in ready if x[0] <= t)))]
t += a + c * len(batch)
for arr, u in batch:
lat.append(t - arr)
heapq.heappush(ready, (t, u))
done += len(batch)
lat = np.array(lat)
return done / (t / 1000), np.percentile(lat, 50), np.percentile(lat, 95)
a, c = 10.0, 1.0
for users in (1, 10, 50):
for b, d in ((1, 0.0), (16, 5.0)):
tput, p50, p95 = simulate(users, b, d, a, c)
print("users=%2d max_batch=%2d tput=%6.1f req/s p50=%6.1f ms p95=%6.1f ms" % (users, b, tput, p50, p95))
# 3. Equivalences and boundaries.
t1, l50, _ = simulate(1, 1, 0.0, a, c)
assert abs(t1 - 1000 / (a + c)) < 1.0 and abs(l50 - (a + c)) < 1e-6 # one user, no batching: 1/s
u1 = simulate(1, 16, 5.0, a, c)
assert u1[1] >= a + c + 5.0 - 1e-6 # lone request pays the full batch window
hi_b, hi_nb = simulate(50, 16, 5.0, a, c), simulate(50, 1, 0.0, a, c)
assert hi_b[0] > 3 * hi_nb[0] # batching wins at saturation
assert simulate(50, 16, 5.0, a, c)[0] <= 1000 * 16 / (a + c * 16) + 1e-6 # never exceeds the full-batch ceiling
print("checks passed")
Executed output (Python 3, numpy):
triton_b16 / triton_b1 = 1.857
triton_b16 / fastapi = 1.733
fastapi_cpu in flight ~ 9.9 (implied mean latency at 50 users 111 ms, 100 users 222 ms; p95 45 ms)
triton_b1 in flight ~ 11.8 (implied mean latency at 50 users 119 ms, 100 users 238 ms; p95 52 ms)
triton_b16 in flight ~ 26.5 (implied mean latency at 50 users 64 ms, 100 users 128 ms; p95 60 ms)
users= 1 max_batch= 1 tput= 90.9 req/s p50= 11.0 ms p95= 11.0 ms
users= 1 max_batch=16 tput= 62.5 req/s p50= 16.0 ms p95= 16.0 ms
users=10 max_batch= 1 tput= 90.9 req/s p50= 110.0 ms p95= 110.0 ms
users=10 max_batch=16 tput= 400.0 req/s p50= 25.0 ms p95= 25.0 ms
users=50 max_batch= 1 tput= 90.9 req/s p50= 550.0 ms p95= 550.0 ms
users=50 max_batch=16 tput= 615.4 req/s p50= 78.0 ms p95= 104.0 ms
checks passed
What the output supports:
- The table's own ratios are 1.857 (batched over unbatched Triton) and 1.733 (batched Triton over FastAPI). The text writes "85%" for the first, which is 85.7% truncated rather than rounded; the abstract's "nearly double that of the baseline" describes 1.73x.
- Using p50 as a stand-in for mean latency (an approximation), Little's law puts the FastAPI row near 9.9 requests in flight, consistent with the 10-user level, and the unbatched Triton row near 11.8. The batched Triton row implies 26.5 in flight, which matches none of the stated 10, 50, or 100 user levels. At 50 or 100 users that row's mean latency would have to be 64 ms or 128 ms, above the reported p95 of 60 ms. The three rows therefore were probably not measured at one common concurrency, or the latencies are not what the labels say. The table does not let a reader decide which.
- In the model, a lone request under a 5 ms batching window pays the full window (16 ms against 11 ms), and at 50 users batching lifts throughput 6.8x with lower tail latency than the unbatched queue. That last point differs from the paper's observation that batching raises p50 from 28 to 34 ms; it follows from the model's single-instance closed loop, where the unbatched server queues, and is not evidence about the paper's setup.
How to maintain it¶
- Scale on the right signal. The paper's own limitations section says CPU utilisation lags for GPU work and proposes scaling on DCGM GPU metrics; the exporter side is covered in DCGM exporter manifest.
- Add drift monitoring. The paper states the pipeline has none.
- Re-benchmark after any model, runtime, or Triton version change; the reference manifest pins
nvcr.io/nvidia/tritonserver:23.01-py3, an old release. - Keep de-identification tests in CI: a preprocessor regression is a compliance regression, and the paper reports no de-identification accuracy.
How to run it in production¶
- Give Triton pods a GPU request. The reference
k8s.yaml(Hugging Face Space, read 2026-09-30) requests only CPU and memory and sets nonvidia.com/gpuresource, so as written no GPU would be allocated to the pods. See Kubernetes GPU scheduling. - Replace the example gateway. The reference
app.main.example.pyaccepts any non-empty bearer token inverify_tokenand forwards to a literal/v2/models/<model>/inferplaceholder; the paper's "OAuth2 and JWT" enforcement is described inSECURITY.mdbut is not implemented in that file. The preprocessor manifest uses the placeholder imageyourrepo/preprocessor:latest. - The reference
hpa.yamltargets CPU 60% and memory 70%; the paper mentions only CPU. - 2 to 10 Triton replicas on a single-T4
g4dn.xlargedoes not describe a testbed the paper actually used; how the HPA range relates to the one-GPU benchmark node is not stated. - Compliance claims (HIPAA alignment) belong to the deployment, not the benchmark. The paper reports no audit, no penetration test, and no PHI-removal measurement.
Failure modes¶
- Cross-hardware comparison read as a framework comparison. CPU FastAPI against GPU Triton says little about FastAPI versus Triton on equal hardware.
- Batch window on sparse traffic. With few concurrent requests, each one waits for the queue delay before execution (the model shows 16 ms versus 11 ms for a single user).
- Lagging autoscaler. CPU-based HPA misses GPU saturation, the paper's first listed limitation.
- Placeholder manifests shipped as production. An unvalidated bearer check and an unset preprocessor image pass a dry-run and fail review.
- PHI in logs. De-identification must run before any component that logs request bodies, including the gateway itself.
Open questions and validation¶
- No code or scripts for the benchmark (Locust plus a custom client, ONNX export, Triton model config) are released; the referenced Hugging Face Space contains manifests and a README but no benchmark harness or model files. The numbers are not reproducible from public artifacts.
- The Zenodo "Benchmarking Note" was not reachable for this page (HTTP 403), so whether Table I is reproduced from it or measured anew is unconfirmed.
- Missing: repeat counts, confidence intervals, tested-user-level per row, Triton instance count, preferred batch sizes, queue delay, the hybrid path's latency cost, and GPU utilisation.
- Differences between the abstract and the body: the abstract says FastAPI has "lower overhead for single-request workloads with a p50 latency of 22 ms" (true only against Triton with GPU transfer overhead, per the body); the body then reports FastAPI with a lower p95 too (45 versus 60 ms), which the abstract omits.
References¶
- Ali, "Scalable and Secure AI Inference in Healthcare: A Comparative Benchmarking of FastAPI and Triton Inference Server on Kubernetes", arXiv 2602.00053: https://arxiv.org/abs/2602.00053
- Gopalan, reference architecture (Hugging Face Space): https://huggingface.co/spaces/digopala/ai-inference-architecture-healthcare
- NVIDIA Triton Inference Server model configuration (dynamic batcher): https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/user_guide/model_configuration.html
- Triton Inference Server repository: https://github.com/triton-inference-server/server
- FastAPI repository: https://github.com/fastapi/fastapi
- Kubernetes Horizontal Pod Autoscaling: https://kubernetes.io/docs/tasks/run-application/horizontal-pod-autoscale/
- Sanh et al., "DistilBERT", arXiv 1910.01108: https://arxiv.org/abs/1910.01108
- Locust load testing: https://docs.locust.io
Related: Inference serving · Continuous batching internals · Kubernetes GPU · Inference QoS and admission control · Observability and monitoring · DCGM exporter manifest · Split inference privacy over WAN