<rss xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title>Llama.cpp - Tag - Jorgen Bergstrom</title><link>https://bergstrom.org/tags/llama.cpp/</link><description>Llama.cpp - Tag - Jorgen Bergstrom</description><generator>Hugo -- gohugo.io</generator><language>en-us</language><copyright>Jorgen Bergstrom</copyright><lastBuildDate>Wed, 07 Oct 2026 00:00:00 -0400</lastBuildDate><atom:link href="https://bergstrom.org/tags/llama.cpp/" rel="self" type="application/rss+xml"/><item><title>Benchmarking Local LLMs on MMLU: What 14,042 Questions Taught Me</title><link>https://bergstrom.org/posts/mmlu_local_llms/</link><pubDate>Wed, 07 Oct 2026 00:00:00 -0400</pubDate><author>Jorgen Bergstrom</author><guid>https://bergstrom.org/posts/mmlu_local_llms/</guid><description>Introduction I have a collected a handful of GGUF models sitting in ~/.cache/llama.cpp. I wanted a straight answer to the simple question: how good are these local models at recalling facts and doing simple reasoning?
So I built a small harness that runs my local models against the full MMLU test set (14,042 questions across 57 subjects) using llama.cpp, one model at a time, and records every single prediction.
This post covers what MMLU is, how to get the questions, how the benchmark works, the example questions, and the results.</description></item></channel></rss>