Quirrel4.41.0

Performance

This page has two sets of benchmarks. Both are measured on Windows x64, and both are committed with the numbers they produced.

Every row times the benchmarked code only, without interpreter startup and without compilation. Each workload repeats the load inside one process and keeps its best time, because the best time has the least noise from other work on the machine.

Each cell in Benchmarks below is measured three times. The harness starts three processes one after another, all pinned to the same logical CPU, and the page shows the fastest of the three.

The page shows the fastest run because noise only makes a run slower. A process can share its core with other work, land on an efficiency core, or get an unlucky code alignment. None of these makes it faster. The core alone changes the result by tens of percent on this hardware.

The harness also checks how far the three runs are apart. It reports a spread of more than five percent when the runs also differ by more than five milliseconds, because on a row that takes tens of milliseconds the clock alone accounts for a few percent. A reported row is measured again on an idle machine.

LuaJIT is measured with the JIT off. Quirrel, Lua and QuickJS are built for platforms where generating code at runtime is not allowed, such as consoles and phones, so a comparison against a JIT would not be useful.

If you need the fastest interpreter, or an AOT language, use Daslang, which has more benchmarks of this kind. Quirrel, Lua and JS are highly dynamic and much quicker to learn, so this page compares those.

Benchmarks

Shorter bars are better.

n-bodies

LuaJIT2.1-joff
0.318s
Luau-0.735
0.401s
Lua-5.5.1
0.505s
Quirrel-4.38.0
Squirrel-3.2
1.075s
QuickJS-ng-0.16.2
1.251s

particles-kinematics

LuaJIT2.1-joff
0.209s
Quirrel-4.38.0
Luau-0.735
0.230s
Squirrel-3.2
0.360s
Lua-5.5.1
0.381s
QuickJS-ng-0.16.2
0.546s

exp-loop

LuaJIT2.1-joff
0.105s
Luau-0.735
0.109s
Lua-5.5.1
0.202s
Quirrel-4.38.0
Squirrel-3.2
0.380s
QuickJS-ng-0.16.2
0.441s

dictionary

LuaJIT2.1-joff
0.112s
Luau-0.735
0.191s
Lua-5.5.1
0.352s
Quirrel-4.38.0
Squirrel-3.2
0.459s
QuickJS-ng-0.16.2
0.953s

darg-ui-benchmark

Luau-0.735
0.099s
LuaJIT2.1-joff
0.103s
Lua-5.5.1
0.244s
QuickJS-ng-0.16.2
0.308s
Quirrel-4.38.0
Squirrel-3.2
0.468s

fibonacci-recursive

LuaJIT2.1-joff
0.038s
Luau-0.735
0.044s
Lua-5.5.1
0.048s
QuickJS-ng-0.16.2
0.078s
Quirrel-4.38.0
Squirrel-3.2
0.133s

fibonacci-loop

LuaJIT2.1-joff
0.027s
Luau-0.735
0.027s
Quirrel-4.38.0
Lua-5.5.1
0.037s
Squirrel-3.2
0.048s
QuickJS-ng-0.16.2
0.142s

primes-loop

LuaJIT2.1-joff
0.042s
Lua-5.5.1
0.058s
Luau-0.735
0.062s
QuickJS-ng-0.16.2
0.091s
Quirrel-4.38.0
Squirrel-3.2
0.173s

float2string

Squirrel-3.2
0.041s
Luau-0.735
0.042s
Quirrel-4.38.0
LuaJIT2.1-joff
0.104s
QuickJS-ng-0.16.2
0.200s
Lua-5.5.1
0.458s

queen

LuaJIT2.1-joff
0.001s
Lua-5.5.1
0.001s
Luau-0.735
0.002s
Quirrel-4.38.0
Squirrel-3.2
0.002s

sort

Luau-0.735
0.030s
Lua-5.5.1
0.044s
LuaJIT2.1-joff
0.048s
Quirrel-4.38.0
Squirrel-3.2
0.110s

spectral-norm

LuaJIT2.1-joff
0.176s
Lua-5.5.1
0.244s
Luau-0.735
0.245s
Quirrel-4.38.0
Squirrel-3.2
0.657s
QuickJS-ng-0.16.2
0.715s

string2float

Luau-0.735
0.066s
LuaJIT2.1-joff
0.084s
QuickJS-ng-0.16.2
0.094s
Lua-5.5.1
0.105s
Quirrel-4.38.0
Squirrel-3.2
0.142s

Measured on Intel64 Family 6 Model 183 Stepping 1, GenuineIntel, Windows-11-10.0.26200-SP0 (11), 2026-09-03. Best iteration of 3 process runs.

VM acceptance

This is the head-to-head harness used for interpreter work. It runs Quirrel against Lua and Luau on paired workloads with identical algorithms: general interpreter loads (fib, binarytrees, life, mandel, strings) and daRg-UI-shaped loads (desc_churn, probe_storm, nullable_probe, closure_storm, method_calls). Run it before and after any change to the VM.

Times are milliseconds. The ratio is Quirrel over the other interpreter, so a value above 1.0 means Quirrel is slower.

workloadquirrellua55luau
fib6130 (2.03x)27 (2.25x)
binarytrees2025 (0.80x)8 (2.48x)
life7640 (1.90x)36 (2.11x)
mandel3919 (2.05x)16 (2.50x)
strings4847 (1.02x)45 (1.06x)
desc_churn6895 (0.72x)27 (2.49x)
probe_storm6143 (1.42x)40 (1.54x)
nullable_probe6520 (3.25x)15 (4.30x)
closure_storm64102 (0.63x)21 (3.12x)
method_calls259207 (1.25x)163 (1.59x)

Measured on Intel64 Family 6 Model 183 Stepping 1, GenuineIntel, Windows-10-10.0.26200-SP0 (10), 2026-08-31. Median of the best rep over 3 process runs.

What was built, and how

Everything is built for Windows 64-bit with clang-cl where the source allows it. The Quirrel rows measure the release build of the interpreter from the Dagor engine tree, which is the shipped runtime configuration: clang and mimalloc. A build on the CRT heap, such as a plain cmake sq.exe, is up to twice as slow on table-sweep workloads, so it would not measure shipped performance.

The third-party interpreters are committed next to their workloads. Each is a release console exe that imports KERNEL32 only. Two of them need a change to get their jump-table interpreter loop, because both gate it on __GNUC__, which clang-cl does not define. Lua takes -DLUA_USE_JUMPTABLE=1. In QuickJS the line #if defined(EMSCRIPTEN) || defined(_MSC_VER) becomes #if defined(EMSCRIPTEN) || (defined(_MSC_VER) && !defined(__clang__)). Without the QuickJS change the primes workload takes twice as long, which measures the MSVC command line and not QuickJS.

LuaJIT has one source change, so that a short repetition can be timed: os.clock() reads the performance counter, because the CRT clock() ticks once per millisecond, which is about the time one rep takes.

Sources

Every workload, harness and prebuilt interpreter is in bench/ next to doc/ in the Dagor engine tree, under prog/1stPartyLibs/quirrel/quirrel, one folder per language: bench/quirrel, bench/lua, bench/luau, bench/js.

Both need the Windows interpreters, so the numbers are refreshed by hand and committed. The site itself builds anywhere. bench/README.md says how to run them and how each committed binary was built.