EDGE INFERENCE METRICS
LLM Quantization & FPS Benchmark
Toggle model quantization bits (FP16, INT8, INT4) and batch sizes to benchmark token throughput speed, VRAM usage, and latency.
SELECT QUANTIZATION BITS:
THROUGHPUT SPEED 24.5 tokens/sec
VRAM ALLOCATION 14.8 GB
INFERENCE LATENCY 40.8 ms