Performance
The compiler is measured on twenty programs, each written four times: in the
xc language, in Objective-C with ARC, in C++ and in Swift, doing the same work
with the same algorithm and the same data. All are built with optimisation
(-O3 for xc, Objective-C and C++, -O for Swift), every version prints a
checksum, and a run only counts if all four checksums agree. Every program, in
all four languages, is on the benchmark sources
page; each benchmark’s name below links to its own. On x86-64 the C++ source
is built twice, with clang and with GCC (g++ -O3), so the Linux figures carry
a column for each compiler.
Each figure is measured with the released 0.74 xcc on two machines:
- arm64: an Apple MacBook Pro with an M4 Max (12 performance and 4 efficiency cores), macOS;
- x86-64: an AMD Ryzen 9 9955HX (16 cores, 32 threads), Linux.
Times are seconds for the timed region, the best of five runs, each run waiting until the machine is otherwise idle.
Summary
Section titled “Summary”How much faster xc’s code runs than each language’s, as the geometric mean over nineteen of the twenty benchmarks:
| xc compared with | arm64 (Apple M4 Max) | x86-64 (AMD Ryzen 9 9955HX) |
|---|---|---|
| Objective-C | 1.41× faster | 1.64× faster |
| C++ (clang) | 1.09× faster | 1.31× faster |
| C++ (GCC) | – | 1.35× faster |
| Swift | 1.48× faster | 1.60× faster |
On both targets xc is ahead of all of them.
The twentieth benchmark, matrix_mul_f32, is kept out of those means. xcc
recognises its loop nest and replaces it with a matrix kernel of its own, which
makes xc tens to hundreds of times faster at that one operation; counted in, it would
dominate means that stand for ordinary code, and most programs do not multiply
matrices. Its figures are under Matrix multiplies.
Per benchmark
Section titled “Per benchmark”Each row is a benchmark, each dot one language: how much faster or slower xc
is than that language. Dots left of the centre line are benchmarks xc wins. Choose a
release above a chart to see how it stood then; the rows stay in 0.74’s
order. Releases before 0.65 were measured against Objective-C alone, apart
from matrix_mul_f32, which was added later and measured against all three.
| benchmark | xc | Objective-C | C++ (clang) | C++ (GCC) | Swift | xc vs the fastest of the others |
|---|---|---|---|---|---|---|
int_muldiv | 1.14 | 0.94 | 0.94 | – | 1.23 | 1.22× slower |
float_math | 0.97 | 0.82 | 0.81 | – | 0.83 | 1.19× slower |
sort_small | 1.01 | 0.94 | 0.98 | – | 1.29 | 1.08× slower |
mem_copy | 0.94 | 1.53 | 0.89 | – | 0.99 | 1.05× slower |
arc_alloc | 1.01 | 1.43 | 0.96 | – | 2.21 | 1.05× slower |
arc_array | 0.54 | 5.21 | 0.52 | – | 0.88 | 1.04× slower |
call_depth | 0.93 | 0.99 | 0.99 | – | 0.91 | 1.02× slower |
matrix_mul | 0.22 | 0.22 | 0.22 | – | 1.12 | 1.01× slower |
hash_mix | 1.16 | 1.15 | 1.15 | – | 1.16 | 1.01× slower |
branch_mix | 0.99 | 0.99 | 0.99 | – | 1.01 | the same |
int_accum | 1.11 | 1.11 | 1.11 | – | 1.12 | the same |
array_map | 1.05 | 1.42 | 1.06 | – | 1.06 | the same |
bit_ops | 1.11 | 1.15 | 1.15 | – | 1.17 | 1.03× faster |
poly_dispatch | 0.52 | 1.12 | 0.55 | – | 2.03 | 1.05× faster |
array_sum | 0.81 | 0.89 | 0.89 | – | 1.67 | 1.10× faster |
method_call | 0.84 | 2.29 | 1.10 | – | 1.08 | 1.28× faster |
sieve | 1.07 | 1.44 | 1.46 | – | 1.91 | 1.35× faster |
struct_copy | 0.69 | 0.99 | 0.99 | – | 1.00 | 1.44× faster |
string_scan | 0.91 | 2.31 | 2.35 | – | 2.37 | 2.5× faster |
matrix_mul_f32 | 0.03 | 9.80 | 9.81 | – | 9.71 | 322× faster |
x86-64
Section titled “x86-64”| benchmark | xc | Objective-C | C++ (clang) | C++ (GCC) | Swift | xc vs the fastest of the others |
|---|---|---|---|---|---|---|
sort_small | 1.92 | 1.51 | 1.50 | 0.82 | 1.02 | 2.3× slower |
arc_array | 1.18 | 3.73 | 0.64 | 0.63 | 0.61 | 1.92× slower |
sieve | 1.13 | 0.85 | 0.83 | 1.81 | 1.22 | 1.36× slower |
struct_copy | 0.87 | 0.79 | 0.79 | 0.65 | 0.79 | 1.34× slower |
float_math | 0.76 | 0.74 | 0.74 | 0.69 | 0.74 | 1.09× slower |
poly_dispatch | 0.97 | 1.24 | 0.93 | 0.93 | 6.69 | 1.04× slower |
method_call | 1.31 | 1.90 | 1.59 | 1.58 | 1.27 | 1.03× slower |
matrix_mul | 0.64 | 0.62 | 0.62 | 0.74 | 2.06 | 1.02× slower |
branch_mix | 0.82 | 0.81 | 0.81 | 0.81 | 0.81 | 1.01× slower |
int_accum | 0.98 | 0.97 | 0.97 | 0.97 | 0.97 | the same |
bit_ops | 0.98 | 0.98 | 0.98 | 0.98 | 0.98 | the same |
hash_mix | 1.01 | 1.08 | 1.08 | 1.01 | 1.08 | the same |
arc_alloc | 0.38 | 2.18 | 0.44 | 0.45 | 0.76 | 1.16× faster |
int_muldiv | 0.51 | 1.70 | 1.70 | 0.69 | 1.70 | 1.35× faster |
string_scan | 1.62 | 3.15 | 3.17 | 2.71 | 3.16 | 1.67× faster |
mem_copy | 0.53 | 1.05 | 1.02 | 1.42 | 0.97 | 1.82× faster |
array_sum | 0.37 | 1.44 | 1.09 | 2.84 | 1.09 | 3.0× faster |
array_map | 0.42 | 1.57 | 1.47 | 2.03 | 2.06 | 3.5× faster |
call_depth | 0.26 | 0.95 | 0.95 | 0.96 | 0.97 | 3.6× faster |
matrix_mul_f32 | 0.22 | 6.28 | 6.26 | 1.95 | 8.76 | 8.7× faster |
The fastest results are where the runtime does the work: string_scan,
method_call and struct_copy are byte scanning, dynamic dispatch and
aggregate copies. The slowest show where xcc’s code generation has most to
gain:
- Vectorisation. On arm64,
int_muldivandfloat_mathare vectorised by both compilers, and clang’s loops are tighter. (matrix_mulreached clang’s speed in 0.73.) - Reference counting on x86-64.
arc_array, where C++ reads elements without retaining them. - x86-64 loops.
sieve,sort_smallandstruct_copy, where the other compilers’ loops are faster.
Matrix multiplies
Section titled “Matrix multiplies”A dense matrix multiply written as three plain loops, C[i][j] the sum over
k of A[i][k] · B[k][j] in float or double, runs as a kernel the
compiler writes itself (from 0.71). The results are the loops’, to the last
bit, and -fno-matmul turns it off. The kernel depends on the machine:
- Apple M4 and later (SME): the multiply runs on the matrix unit. From
0.72 it works on a 2×2 block of the unit’s tiles at once, and under
:goal(speed), whichmatrix_mul_f32uses, it leaves out a NaN check that only affects the bits of a NaN (the inputs have none). - x86-64: the widest vector unit the processor has (SSE2, AVX2 or AVX-512), chosen when the program starts; the Ryzen 9 9955HX picks AVX-512.
matrix_mul_f32 measures it: a 128×128 float multiply, 10,000 times.
(matrix_mul multiplies u32 values, which the matrix unit does not take, so
it stays a vectorised loop.)
matrix_mul_f32 | xc | Objective-C | C++ (clang) | C++ (GCC) | Swift |
|---|---|---|---|---|---|
| arm64, Apple M4 Max | 30 ms | 9.8 s | 9.8 s | – | 9.7 s |
| xc is | 325× faster | 325× faster | – | 322× faster | |
| x86-64, AMD Ryzen 9 9955HX | 225 ms | 6.3 s | 6.3 s | 2.0 s | 8.8 s |
| xc is | 28× faster | 28× faster | 8.7× faster | 39× faster |
Parallel blocks and the GPU
Section titled “Parallel blocks and the GPU”A par block runs a loop’s iterations at once,
across every CPU thread or on the GPU (from 0.7). The GPU is reached through:
- Metal on Apple silicon;
- CUDA on an NVIDIA GPU under Windows;
- Vulkan on Linux, Windows and Android (from 0.72);
- WebGPU in a browser (from 0.72).
Four programs in benchmark/par measure it:
mandelbrot: a 2048×2048 escape-time image, at most 256 iterations a pixel;perlin: 2048×2048 improved noise, four octaves;nbody: the force on each of 8192 bodies from all the others;saxpy:y = a·x + yover 16 million integers, with a sum.
Each runs its block eight times and reports the best run, in four modes: one
thread (XC_PAR=cpu XC_PAR_THREADS=1), all threads (XC_PAR=cpu), GPU
(XC_PAR=gpu) and auto, the default, which times the CPU and the GPU and
keeps the faster (from 0.73 it remembers the choice, per program and per
machine, so later runs do not measure again). Every mode’s checksum must agree. On Windows GPU is CUDA
and Vulkan the same card through Vulkan (XC_PAR_GPU=vulkan).
Apple MacBook Pro, M4 Max (12 performance and 4 efficiency cores), its GPU through Metal (ms; best of eight runs)
| benchmark | one thread | all threads | GPU | auto | GPU vs all threads | GPU’s first run |
|---|---|---|---|---|---|---|
mandelbrot | 481 | 100 | 1.6 | 2.1 | 65× faster | 22.5 |
nbody | 67.7 | 6.9 | 0.7 | 0.9 | 9.9× faster | 22.6 |
perlin | 327 | 31.3 | 1.5 | 0.9 | 20× faster | 23.4 |
saxpy | 7.5 | 1.0 | 12.5 | 1.0 | 13.2× slower | 33.4 |
AMD Ryzen 9 9955HX (16 cores, 32 threads), Linux, its integrated Radeon 610M (2 compute units, 128 shader lanes) through Vulkan (ms; best of eight runs)
| benchmark | one thread | all threads | GPU | auto | GPU vs all threads | GPU’s first run |
|---|---|---|---|---|---|---|
mandelbrot | 398 | 40.3 | 20.0 | 19.2 | 2.0× faster | 40.1 |
nbody | 95.1 | 7.1 | 3.6 | 3.6 | 1.97× faster | 23.3 |
perlin | 374 | 25.8 | 7.1 | 7.0 | 3.6× faster | 30.8 |
saxpy | 12.8 | 2.5 | 218 | 2.5 | 86× slower | 253 |
AMD Ryzen 9 5950X (16 cores, 32 threads), Windows, an NVIDIA RTX 3090 through CUDA and Vulkan (ms; best of eight runs)
| benchmark | one thread | all threads | GPU | Vulkan | auto | GPU vs all threads | GPU’s first run | Vulkan’s first run |
|---|---|---|---|---|---|---|---|---|
mandelbrot | 455 | 46.7 | 1.6 | 1.1 | 1.6 | 30× faster | 202 | 122 |
nbody | 120 | 8.8 | 1.9 | 0.9 | 1.8 | 4.7× faster | 207 | 124 |
perlin | 498 | 29.3 | 2.0 | 2.0 | 2.0 | 14.6× faster | 220 | 128 |
saxpy | 13.5 | 5.6 | 39.9 | 44.8 | 5.4 | 7.2× slower | 258 | 178 |
Chrome on the Apple M4 Max: wasm32 (one thread) and WebGPU (from 0.73) (ms; best of eight runs)
| benchmark | one thread | all threads | GPU | auto | GPU vs one thread | GPU’s first run |
|---|---|---|---|---|---|---|
mandelbrot | 364 | – | 2.6 | 1.9 | 140× faster | 12.1 |
nbody | 53.9 | – | 3.2 | 3.2 | 16.8× faster | 7.3 |
perlin | 149 | – | 2.6 | 2.5 | 57× faster | 8.6 |
saxpy | 15.7 | – | 22.0 | 16.3 | 1.40× slower | 56.1 |
What the tables show:
- A discrete GPU wins by a wide margin on the three programs that do real
work per element, and
autofinds that and uses it. saxpystays on the CPU. It does one multiply-add for every eight bytes it moves, so copying to and from the GPU outweighs the arithmetic, andautokeeps it on the CPU everywhere.- The Ryzen’s integrated Radeon has 128 shader lanes against the CPU’s 32
threads. It just beats them on
mandelbrotandnbodyand loses on the other two; its copies ofsaxpy’s arrays are slower still, because the runtime does not yet use the CPU’s cache for them there. - Vulkan against CUDA: on the RTX 3090, Vulkan is faster on
mandelbrot(1.1 ms against 1.6) andnbody(0.9 ms against 1.9), level onperlinand slower onsaxpy(44.8 ms against 39.9). Up to 0.74autoused CUDA on an NVIDIA GPU; from 0.75 it measures both and keeps the faster for each block. - The first run carries one-off costs (building the kernel, and on NVIDIA
creating the driver context and compiling the PTX or SPIR-V), so it is shown
apart: a few tens of milliseconds on Metal and the integrated GPU, and on the
3090 202 to 258 ms through CUDA and 122 to 178 ms through Vulkan.
autopays it once, while it measures.
Release to release
Section titled “Release to release”How fast xc’s code is in each release, relative to 0.62: above 1× is faster.
The thin lines are single benchmarks (labelled where they moved by more than
twelve percent), the thick line the geometric mean (without matrix_mul_f32). Releases shown: 0.62, 0.63, 0.64, 0.65, 0.66, 0.7, 0.71, 0.72, 0.73, 0.74. arc_alloc, method_call changed in
0.64’s benchmark set and are left out of this history.
What is being compared, and what is not
Section titled “What is being compared, and what is not”The compiler that ships. The xc numbers come from the xcc in the download.
The same work in every language. Every version keeps its data where the xc
version keeps it (local arrays, not static ones, which clang optimises
differently) and leaves nothing a compiler can remove: an object that could be
put on the stack outlives its iteration, a call whose target could be resolved
at compile time takes its class at run time, and results are folded in so no
loop has a closed form. Where the original’s point is reference-counted
objects, the C++ version uses std::shared_ptr and the Swift version a class,
so they pay for reference counting too.
Different runtimes on the two targets. The Objective-C column is Apple’s
Foundation on arm64 and GNUstep with libobjc2 on x86-64, and Swift is 6.2 on the
Mac and 6.1 on Linux. These are different implementations, so a language’s
times compare within a target and not across one. Swift on Linux is markedly
slower on poly_dispatch and array_map than on the Mac; the runs were
repeated on an idle machine and reproduce.
Timed regions of about one second. Each benchmark times its own inner loop rather than the process, so start-up and data set-up are excluded.
Alignment noise on x86-64. xcc aligns every loop head on x86-64 to a 32-byte
boundary and the start of .text to 64 bytes, so an unrelated change elsewhere
cannot move a loop across a fetch boundary; see
Optimisation.
Reproducing
Section titled “Reproducing”The sources are in benchmark/src, one .xc, .m, .cpp and .swift per
program, and the runner builds and times them all:
python3 benchmark/run.py --opt O3 --repeats 5python3 benchmark/page.pyThe x86-64 legs cross-build xc here and build the other languages on the
configured Linux host; without one the runner measures this machine only.
--langs measures some of the languages and adds them to the results already
there. Results land in benchmark/<version>/results.json, and page.py turns
them into this page.