Performance
This page has two sets of benchmarks. Both are measured on Windows x64, and both are committed with the numbers they produced.
Every row times the benchmarked code only, without interpreter startup and without compilation. Each workload repeats the load inside one process and keeps its best time, because the best time has the least noise from other work on the machine.
Each cell in Benchmarks below is measured three times. The harness starts three processes one after another, all pinned to the same logical CPU, and the page shows the fastest of the three.
The page shows the fastest run because noise only makes a run slower. A process can share its core with other work, land on an efficiency core, or get an unlucky code alignment. None of these makes it faster. The core alone changes the result by tens of percent on this hardware.
The harness also checks how far the three runs are apart. It reports a spread of more than five percent when the runs also differ by more than five milliseconds, because on a row that takes tens of milliseconds the clock alone accounts for a few percent. A reported row is measured again on an idle machine.
LuaJIT is measured with the JIT off. Quirrel, Lua and QuickJS are built for platforms where generating code at runtime is not allowed, such as consoles and phones, so a comparison against a JIT would not be useful.
If you need the fastest interpreter, or an AOT language, use Daslang, which has more benchmarks of this kind. Quirrel, Lua and JS are highly dynamic and much quicker to learn, so this page compares those.
Benchmarks
Shorter bars are better.
n-bodies
| LuaJIT2.1-joff | |
|---|---|
| Luau-0.735 | |
| Lua-5.5.1 | |
| Quirrel-4.38.0 | |
| Squirrel-3.2 | |
| QuickJS-ng-0.16.2 |
particles-kinematics
| LuaJIT2.1-joff | |
|---|---|
| Quirrel-4.38.0 | |
| Luau-0.735 | |
| Squirrel-3.2 | |
| Lua-5.5.1 | |
| QuickJS-ng-0.16.2 |
exp-loop
| LuaJIT2.1-joff | |
|---|---|
| Luau-0.735 | |
| Lua-5.5.1 | |
| Quirrel-4.38.0 | |
| Squirrel-3.2 | |
| QuickJS-ng-0.16.2 |
dictionary
| LuaJIT2.1-joff | |
|---|---|
| Luau-0.735 | |
| Lua-5.5.1 | |
| Quirrel-4.38.0 | |
| Squirrel-3.2 | |
| QuickJS-ng-0.16.2 |
darg-ui-benchmark
| Luau-0.735 | |
|---|---|
| LuaJIT2.1-joff | |
| Lua-5.5.1 | |
| QuickJS-ng-0.16.2 | |
| Quirrel-4.38.0 | |
| Squirrel-3.2 |
fibonacci-recursive
| LuaJIT2.1-joff | |
|---|---|
| Luau-0.735 | |
| Lua-5.5.1 | |
| QuickJS-ng-0.16.2 | |
| Quirrel-4.38.0 | |
| Squirrel-3.2 |
fibonacci-loop
| LuaJIT2.1-joff | |
|---|---|
| Luau-0.735 | |
| Quirrel-4.38.0 | |
| Lua-5.5.1 | |
| Squirrel-3.2 | |
| QuickJS-ng-0.16.2 |
primes-loop
| LuaJIT2.1-joff | |
|---|---|
| Lua-5.5.1 | |
| Luau-0.735 | |
| QuickJS-ng-0.16.2 | |
| Quirrel-4.38.0 | |
| Squirrel-3.2 |
float2string
| Squirrel-3.2 | |
|---|---|
| Luau-0.735 | |
| Quirrel-4.38.0 | |
| LuaJIT2.1-joff | |
| QuickJS-ng-0.16.2 | |
| Lua-5.5.1 |
queen
| LuaJIT2.1-joff | |
|---|---|
| Lua-5.5.1 | |
| Luau-0.735 | |
| Quirrel-4.38.0 | |
| Squirrel-3.2 |
sort
| Luau-0.735 | |
|---|---|
| Lua-5.5.1 | |
| LuaJIT2.1-joff | |
| Quirrel-4.38.0 | |
| Squirrel-3.2 |
spectral-norm
| LuaJIT2.1-joff | |
|---|---|
| Lua-5.5.1 | |
| Luau-0.735 | |
| Quirrel-4.38.0 | |
| Squirrel-3.2 | |
| QuickJS-ng-0.16.2 |
string2float
| Luau-0.735 | |
|---|---|
| LuaJIT2.1-joff | |
| QuickJS-ng-0.16.2 | |
| Lua-5.5.1 | |
| Quirrel-4.38.0 | |
| Squirrel-3.2 |
Measured on Intel64 Family 6 Model 183 Stepping 1, GenuineIntel, Windows-11-10.0.26200-SP0 (11), 2026-09-03. Best iteration of 3 process runs.
VM acceptance
This is the head-to-head harness used for interpreter work. It runs Quirrel against Lua and Luau on paired workloads with identical algorithms: general interpreter loads (fib, binarytrees, life, mandel, strings) and daRg-UI-shaped loads (desc_churn, probe_storm, nullable_probe, closure_storm, method_calls). Run it before and after any change to the VM.
Times are milliseconds. The ratio is Quirrel over the other interpreter, so a value above 1.0 means Quirrel is slower.
| workload | quirrel | lua55 | luau |
|---|---|---|---|
| fib | 61 | 30 (2.03x) | 27 (2.25x) |
| binarytrees | 20 | 25 (0.80x) | 8 (2.48x) |
| life | 76 | 40 (1.90x) | 36 (2.11x) |
| mandel | 39 | 19 (2.05x) | 16 (2.50x) |
| strings | 48 | 47 (1.02x) | 45 (1.06x) |
| desc_churn | 68 | 95 (0.72x) | 27 (2.49x) |
| probe_storm | 61 | 43 (1.42x) | 40 (1.54x) |
| nullable_probe | 65 | 20 (3.25x) | 15 (4.30x) |
| closure_storm | 64 | 102 (0.63x) | 21 (3.12x) |
| method_calls | 259 | 207 (1.25x) | 163 (1.59x) |
Measured on Intel64 Family 6 Model 183 Stepping 1, GenuineIntel, Windows-10-10.0.26200-SP0 (10), 2026-08-31. Median of the best rep over 3 process runs.
What was built, and how
Everything is built for Windows 64-bit with clang-cl where the source allows it. The Quirrel rows measure the release build of the interpreter from the Dagor engine tree, which is the shipped runtime configuration: clang and mimalloc. A build on the CRT heap, such as a plain cmake sq.exe, is up to twice as slow on table-sweep workloads, so it would not measure shipped performance.
The third-party interpreters are committed next to their workloads. Each is a release console exe that imports KERNEL32 only. Two of them need a change to get their jump-table interpreter loop, because both gate it on __GNUC__, which clang-cl does not define. Lua takes -DLUA_USE_JUMPTABLE=1. In QuickJS the line #if defined(EMSCRIPTEN) || defined(_MSC_VER) becomes #if defined(EMSCRIPTEN) || (defined(_MSC_VER) && !defined(__clang__)). Without the QuickJS change the primes workload takes twice as long, which measures the MSVC command line and not QuickJS.
LuaJIT has one source change, so that a short repetition can be timed: os.clock() reads the performance counter, because the CRT clock() ticks once per millisecond, which is about the time one rep takes.
Sources
Every workload, harness and prebuilt interpreter is in bench/ next to doc/ in the Dagor engine tree, under prog/1stPartyLibs/quirrel/quirrel, one folder per language: bench/quirrel, bench/lua, bench/luau, bench/js.
bench/benchmarks.py- the cross-language suite above. It writesdoc/content/_bench.json, which this page renders.bench/run_vm_bench.py- the VM acceptance harness. It writesdoc/content/_vm_bench.json.
Both need the Windows interpreters, so the numbers are refreshed by hand and committed. The site itself builds anywhere. bench/README.md says how to run them and how each committed binary was built.