Benchmarking Ollama and vLLM for Concurrent LLM Serving: A Multi-Scenario Evaluation of Performance and Scalability

dc.contributor.authorAy, Betul
dc.contributor.authorDemirdag, Yunus Emre
dc.date.accessioned2026-09-08T07:11:55Z
dc.date.issued2026
dc.departmentFırat Üniveristesi
dc.description.abstractServing LLMs to many concurrent users gives rise to significant challenges for latency, throughput, and scalability. This paper presents a systematic and reproducible benchmark of two widely used large language model (LLM) serving frameworks, Ollama and vLLM, under concurrent workloads. We conducted a controlled comparison by running both frameworks on a single NVIDIA H100 80 GB GPU with the same model weights (Qwen3-4B) and inference configurations. We then evaluated them using five open benchmark datasets across four scenarios consisting of baseline question answering, complex reasoning, streaming interaction, and stress testing. Each scenario was executed under increasing concurrency levels, ranging from light loads to high-concurrency stress, to measure end-to-end latency, throughput, time-to-first-token (TTFT), success rate, and resource usage. Our experiments show that vLLM clearly outperforms Ollama across all four scenarios, achieving a 100% request success rate. It delivers 20-29 times higher throughput and 8-19 times lower P95 latency, and it completes every request successfully. Additionally, vLLM produced the first token within 0.5-3.5 s and remained stable at up to 100 concurrent users. Ollama, by contrast, required 54-122 s for the first token, hit a concurrency bottleneck near 10 users, and exhibited a timeout-based error rate of 13-30.06% under heavier loads. Notably, both frameworks demonstrated significant memory growth during extended endurance tests, necessitating careful monitoring in long-running deployments. Overall, for a single model on a single GPU, vLLM is highly suitable for high-concurrency serving, while Ollama remains a practical choice for lightweight, local, or developmental workflows.
dc.identifier.doi10.3390/app16115435
dc.identifier.issn2076-3417
dc.identifier.issue11
dc.identifier.scopus2-s2.0-105041481094
dc.identifier.scopusqualityQ1
dc.identifier.urihttps://doi.org/10.3390/app16115435
dc.identifier.urihttps://hdl.handle.net/11508/65214
dc.identifier.volume16
dc.identifier.wosWOS:001789774300001
dc.identifier.wosqualityQ2
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherMdpi
dc.relation.ispartofApplied Sciences-Basel
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WOS_20250903
dc.subjectLarge Language Model Serving
dc.subjectOllama
dc.subjectVllm
dc.subjectLatency Optimization
dc.subjectConcurrent Request Handling
dc.subjectContinuous Batching
dc.subjectFp8 Quantization
dc.subjectPagedattention
dc.subjectStress Testing
dc.subjectPerformance Evaluation
dc.subjectQwen3
dc.subjectNvidia H100
dc.titleBenchmarking Ollama and vLLM for Concurrent LLM Serving: A Multi-Scenario Evaluation of Performance and Scalability
dc.typeArticle

Dosyalar