Skip to content

Performance

The compiler is measured on twenty programs, each written four times: in the xc language, in Objective-C with ARC, in C++ and in Swift, doing the same work with the same algorithm and the same data. All are built with optimisation (-O3 for xc, Objective-C and C++, -O for Swift), every version prints a checksum, and a run only counts if all four checksums agree. Every program, in all four languages, is on the benchmark sources page; each benchmark’s name below links to its own. On x86-64 the C++ source is built twice, with clang and with GCC (g++ -O3), so the Linux figures carry a column for each compiler.

Each figure is measured with the released 0.74 xcc on two machines:

  • arm64: an Apple MacBook Pro with an M4 Max (12 performance and 4 efficiency cores), macOS;
  • x86-64: an AMD Ryzen 9 9955HX (16 cores, 32 threads), Linux.

Times are seconds for the timed region, the best of five runs, each run waiting until the machine is otherwise idle.

How much faster xc’s code runs than each language’s, as the geometric mean over nineteen of the twenty benchmarks:

xc compared witharm64 (Apple M4 Max)x86-64 (AMD Ryzen 9 9955HX)
Objective-C1.41× faster1.64× faster
C++ (clang)1.09× faster1.31× faster
C++ (GCC)–1.35× faster
Swift1.48× faster1.60× faster

On both targets xc is ahead of all of them.

The twentieth benchmark, matrix_mul_f32, is kept out of those means. xcc recognises its loop nest and replaces it with a matrix kernel of its own, which makes xc tens to hundreds of times faster at that one operation; counted in, it would dominate means that stand for ordinary code, and most programs do not multiply matrices. Its figures are under Matrix multiplies.

Each row is a benchmark, each dot one language: how much faster or slower xc is than that language. Dots left of the centre line are benchmarks xc wins. Choose a release above a chart to see how it stood then; the rows stay in 0.74’s order. Releases before 0.65 were measured against Objective-C alone, apart from matrix_mul_f32, which was added later and measured against all three.

Release:
arm64: xc against each language, per benchmark (0.62)xc vs Objective-Cxc vs C++ (clang)xc vs C++ (GCC)xc vs Swift8× faster4× faster2× fastersame2× slower4× slower8× slower← xc fasterxc slower →int_muldivint_muldiv: the four programsint_muldiv: xc is 1.21× slower than Objective-Cfloat_mathfloat_math: the four programsfloat_math: xc is 1.48× slower than Objective-Csort_smallsort_small: the four programssort_small: xc is 1.29× slower than Objective-Cmem_copymem_copy: the four programsmem_copy: xc is 1.80× slower than Objective-Carc_arrayarc_array: the four programsarc_array: xc is 5.6× faster than Objective-Ccall_depthcall_depth: the four programscall_depth: xc is 1.06× faster than Objective-Cmatrix_mulmatrix_mul: the four programsmatrix_mul: xc is 1.83× slower than Objective-Chash_mixhash_mix: the four programshash_mix: xc is 1.03× slower than Objective-Cbranch_mixbranch_mix: the four programsbranch_mix: xc is 1.11× slower than Objective-Cint_accumint_accum: the four programsint_accum: xc is 1.12× slower than Objective-Carray_maparray_map: the four programsarray_map: xc is 1.43× slower than Objective-Cbit_opsbit_ops: the four programsbit_ops: xc is 1.13× slower than Objective-Cpoly_dispatchpoly_dispatch: the four programspoly_dispatch: xc is 1.14× slower than Objective-Carray_sumarray_sum: the four programsarray_sum: xc is the same than Objective-Csievesieve: the four programssieve: xc is 1.30× faster than Objective-Cstruct_copystruct_copy: the four programsstruct_copy: xc is 1.21× slower than Objective-Cstring_scanstring_scan: the four programsstring_scan: xc is 2.6× faster than Objective-Cmatrix_mul_f32matrix_mul_f32: the four programsmatrix_mul_f32: xc is 1.03× faster than Objective-Cmatrix_mul_f32: xc is 1.03× faster than C++ (clang)matrix_mul_f32: xc is 1.03× slower than Swift
arm64: xc against each language, per benchmark (0.62)
arm64: xc against each language, per benchmark (0.63)xc vs Objective-Cxc vs C++ (clang)xc vs C++ (GCC)xc vs Swift8× faster4× faster2× fastersame2× slower4× slower8× slower← xc fasterxc slower →int_muldivint_muldiv: the four programsint_muldiv: xc is 1.22× slower than Objective-Cfloat_mathfloat_math: the four programsfloat_math: xc is 1.45× slower than Objective-Csort_smallsort_small: the four programssort_small: xc is 1.28× slower than Objective-Cmem_copymem_copy: the four programsmem_copy: xc is 1.78× slower than Objective-Carc_arrayarc_array: the four programsarc_array: xc is 5.6× faster than Objective-Ccall_depthcall_depth: the four programscall_depth: xc is 1.07× faster than Objective-Cmatrix_mulmatrix_mul: the four programsmatrix_mul: xc is 1.83× slower than Objective-Chash_mixhash_mix: the four programshash_mix: xc is 1.04× slower than Objective-Cbranch_mixbranch_mix: the four programsbranch_mix: xc is 1.11× slower than Objective-Cint_accumint_accum: the four programsint_accum: xc is 1.12× slower than Objective-Carray_maparray_map: the four programsarray_map: xc is 1.38× slower than Objective-Cbit_opsbit_ops: the four programsbit_ops: xc is 1.12× slower than Objective-Cpoly_dispatchpoly_dispatch: the four programspoly_dispatch: xc is 1.14× slower than Objective-Carray_sumarray_sum: the four programsarray_sum: xc is the same than Objective-Csievesieve: the four programssieve: xc is 1.30× faster than Objective-Cstruct_copystruct_copy: the four programsstruct_copy: xc is 1.17× slower than Objective-Cstring_scanstring_scan: the four programsstring_scan: xc is 2.6× faster than Objective-Cmatrix_mul_f32matrix_mul_f32: the four programsmatrix_mul_f32: xc is 1.05× faster than Objective-Cmatrix_mul_f32: xc is 1.05× faster than C++ (clang)matrix_mul_f32: xc is 1.02× slower than Swift
arm64: xc against each language, per benchmark (0.63)
arm64: xc against each language, per benchmark (0.64)xc vs Objective-Cxc vs C++ (clang)xc vs C++ (GCC)xc vs Swift8× faster4× faster2× fastersame2× slower4× slower8× slower← xc fasterxc slower →int_muldivint_muldiv: the four programsint_muldiv: xc is 1.22× slower than Objective-Cfloat_mathfloat_math: the four programsfloat_math: xc is 1.48× slower than Objective-Csort_smallsort_small: the four programssort_small: xc is 1.34× slower than Objective-Cmem_copymem_copy: the four programsmem_copy: xc is 1.78× slower than Objective-Carc_arrayarc_array: the four programsarc_array: xc is 5.6× faster than Objective-Ccall_depthcall_depth: the four programscall_depth: xc is 1.07× faster than Objective-Cmatrix_mulmatrix_mul: the four programsmatrix_mul: xc is 1.82× slower than Objective-Chash_mixhash_mix: the four programshash_mix: xc is 1.06× slower than Objective-Cbranch_mixbranch_mix: the four programsbranch_mix: xc is 1.11× slower than Objective-Cint_accumint_accum: the four programsint_accum: xc is 1.12× slower than Objective-Carray_maparray_map: the four programsarray_map: xc is 1.37× slower than Objective-Cbit_opsbit_ops: the four programsbit_ops: xc is 1.12× slower than Objective-Cpoly_dispatchpoly_dispatch: the four programspoly_dispatch: xc is 1.14× slower than Objective-Carray_sumarray_sum: the four programsarray_sum: xc is the same than Objective-Csievesieve: the four programssieve: xc is 1.30× faster than Objective-Cstruct_copystruct_copy: the four programsstruct_copy: xc is 1.25× slower than Objective-Cstring_scanstring_scan: the four programsstring_scan: xc is 2.6× faster than Objective-Cmatrix_mul_f32matrix_mul_f32: the four programsmatrix_mul_f32: xc is 1.04× faster than Objective-Cmatrix_mul_f32: xc is 1.05× faster than C++ (clang)matrix_mul_f32: xc is 1.02× slower than Swift
arm64: xc against each language, per benchmark (0.64)
arm64: xc against each language, per benchmark (0.65)xc vs Objective-Cxc vs C++ (clang)xc vs C++ (GCC)xc vs Swift8× faster4× faster2× fastersame2× slower4× slower8× slower← xc fasterxc slower →int_muldivint_muldiv: the four programsint_muldiv: xc is 1.22× slower than Objective-Cint_muldiv: xc is 1.22× slower than C++ (clang)int_muldiv: xc is 1.07× faster than Swiftfloat_mathfloat_math: the four programsfloat_math: xc is 1.19× slower than Objective-Cfloat_math: xc is 1.19× slower than C++ (clang)float_math: xc is 1.16× slower than Swiftsort_smallsort_small: the four programssort_small: xc is 1.18× slower than Objective-Csort_small: xc is 1.14× slower than C++ (clang)sort_small: xc is 1.19× faster than Swiftmem_copymem_copy: the four programsmem_copy: xc is 1.63× faster than Objective-Cmem_copy: xc is 1.05× slower than C++ (clang)mem_copy: xc is 1.06× faster than Swiftarc_arrayarc_array: the four programsarc_array: xc is 9.6× faster than Objective-Carc_array: xc is 1.04× slower than C++ (clang)arc_array: xc is 1.64× faster than Swiftcall_depthcall_depth: the four programscall_depth: xc is 1.08× faster than Objective-Ccall_depth: xc is 1.08× faster than C++ (clang)call_depth: xc is 1.01× slower than Swiftmatrix_mulmatrix_mul: the four programsmatrix_mul: xc is 1.75× slower than Objective-Cmatrix_mul: xc is 1.73× slower than C++ (clang)matrix_mul: xc is 2.9× faster than Swifthash_mixhash_mix: the four programshash_mix: xc is the same than Objective-Chash_mix: xc is the same than C++ (clang)hash_mix: xc is the same than Swiftbranch_mixbranch_mix: the four programsbranch_mix: xc is the same than Objective-Cbranch_mix: xc is the same than C++ (clang)branch_mix: xc is 1.02× faster than Swiftint_accumint_accum: the four programsint_accum: xc is the same than Objective-Cint_accum: xc is the same than C++ (clang)int_accum: xc is 1.01× faster than Swiftarray_maparray_map: the four programsarray_map: xc is 1.36× faster than Objective-Carray_map: xc is 1.02× slower than C++ (clang)array_map: xc is 1.01× faster than Swiftbit_opsbit_ops: the four programsbit_ops: xc is 1.02× faster than Objective-Cbit_ops: xc is 1.03× faster than C++ (clang)bit_ops: xc is 1.03× faster than Swiftpoly_dispatchpoly_dispatch: the four programspoly_dispatch: xc is 1.72× faster than Objective-Cpoly_dispatch: xc is 1.18× slower than C++ (clang)poly_dispatch: xc is 3.2× faster than Swiftarray_sumarray_sum: the four programsarray_sum: xc is 1.05× slower than Objective-Carray_sum: xc is 1.05× slower than C++ (clang)array_sum: xc is 1.80× faster than Swiftsievesieve: the four programssieve: xc is 1.34× faster than Objective-Csieve: xc is 1.35× faster than C++ (clang)sieve: xc is 1.77× faster than Swiftstruct_copystruct_copy: the four programsstruct_copy: xc is 1.46× faster than Objective-Cstruct_copy: xc is 1.45× faster than C++ (clang)struct_copy: xc is 1.46× faster than Swiftstring_scanstring_scan: the four programsstring_scan: xc is 2.5× faster than Objective-Cstring_scan: xc is 2.6× faster than C++ (clang)string_scan: xc is 2.6× faster than Swiftmatrix_mul_f32matrix_mul_f32: the four programsmatrix_mul_f32: xc is 1.77× faster than Objective-Cmatrix_mul_f32: xc is 1.77× faster than C++ (clang)matrix_mul_f32: xc is 1.66× faster than Swift
arm64: xc against each language, per benchmark (0.65)
arm64: xc against each language, per benchmark (0.66)xc vs Objective-Cxc vs C++ (clang)xc vs C++ (GCC)xc vs Swift8× faster4× faster2× fastersame2× slower4× slower8× slower← xc fasterxc slower →int_muldivint_muldiv: the four programsint_muldiv: xc is 1.22× slower than Objective-Cint_muldiv: xc is 1.21× slower than C++ (clang)int_muldiv: xc is 1.07× faster than Swiftfloat_mathfloat_math: the four programsfloat_math: xc is 1.18× slower than Objective-Cfloat_math: xc is 1.19× slower than C++ (clang)float_math: xc is 1.16× slower than Swiftsort_smallsort_small: the four programssort_small: xc is 1.18× slower than Objective-Csort_small: xc is 1.13× slower than C++ (clang)sort_small: xc is 1.19× faster than Swiftmem_copymem_copy: the four programsmem_copy: xc is 1.62× faster than Objective-Cmem_copy: xc is 1.05× slower than C++ (clang)mem_copy: xc is 1.06× faster than Swiftarc_arrayarc_array: the four programsarc_array: xc is 9.7× faster than Objective-Carc_array: xc is the same than C++ (clang)arc_array: xc is 1.71× faster than Swiftcall_depthcall_depth: the four programscall_depth: xc is 1.08× faster than Objective-Ccall_depth: xc is 1.08× faster than C++ (clang)call_depth: xc is the same than Swiftmatrix_mulmatrix_mul: the four programsmatrix_mul: xc is 1.77× slower than Objective-Cmatrix_mul: xc is 1.74× slower than C++ (clang)matrix_mul: xc is 2.9× faster than Swifthash_mixhash_mix: the four programshash_mix: xc is the same than Objective-Chash_mix: xc is the same than C++ (clang)hash_mix: xc is the same than Swiftbranch_mixbranch_mix: the four programsbranch_mix: xc is 1.02× slower than Objective-Cbranch_mix: xc is 1.02× slower than C++ (clang)branch_mix: xc is the same than Swiftint_accumint_accum: the four programsint_accum: xc is the same than Objective-Cint_accum: xc is the same than C++ (clang)int_accum: xc is 1.01× faster than Swiftarray_maparray_map: the four programsarray_map: xc is 1.39× faster than Objective-Carray_map: xc is 1.03× slower than C++ (clang)array_map: xc is the same than Swiftbit_opsbit_ops: the four programsbit_ops: xc is 1.03× faster than Objective-Cbit_ops: xc is 1.03× faster than C++ (clang)bit_ops: xc is 1.05× faster than Swiftpoly_dispatchpoly_dispatch: the four programspoly_dispatch: xc is 1.72× faster than Objective-Cpoly_dispatch: xc is 1.18× slower than C++ (clang)poly_dispatch: xc is 3.2× faster than Swiftarray_sumarray_sum: the four programsarray_sum: xc is 1.04× slower than Objective-Carray_sum: xc is 1.04× slower than C++ (clang)array_sum: xc is 1.79× faster than Swiftsievesieve: the four programssieve: xc is 1.34× faster than Objective-Csieve: xc is 1.36× faster than C++ (clang)sieve: xc is 1.78× faster than Swiftstruct_copystruct_copy: the four programsstruct_copy: xc is 1.45× faster than Objective-Cstruct_copy: xc is 1.45× faster than C++ (clang)struct_copy: xc is 1.46× faster than Swiftstring_scanstring_scan: the four programsstring_scan: xc is 2.5× faster than Objective-Cstring_scan: xc is 2.6× faster than C++ (clang)string_scan: xc is 2.6× faster than Swiftmatrix_mul_f32matrix_mul_f32: the four programsmatrix_mul_f32: xc is 1.76× faster than Objective-Cmatrix_mul_f32: xc is 1.76× faster than C++ (clang)matrix_mul_f32: xc is 1.64× faster than Swift
arm64: xc against each language, per benchmark (0.66)
arm64: xc against each language, per benchmark (0.7)xc vs Objective-Cxc vs C++ (clang)xc vs C++ (GCC)xc vs Swift8× faster4× faster2× fastersame2× slower4× slower8× slower← xc fasterxc slower →int_muldivint_muldiv: the four programsint_muldiv: xc is 1.22× slower than Objective-Cint_muldiv: xc is 1.22× slower than C++ (clang)int_muldiv: xc is 1.07× faster than Swiftfloat_mathfloat_math: the four programsfloat_math: xc is 1.21× slower than Objective-Cfloat_math: xc is 1.22× slower than C++ (clang)float_math: xc is 1.19× slower than Swiftsort_smallsort_small: the four programssort_small: xc is 1.22× slower than Objective-Csort_small: xc is 1.17× slower than C++ (clang)sort_small: xc is 1.22× faster than Swiftmem_copymem_copy: the four programsmem_copy: xc is 1.65× faster than Objective-Cmem_copy: xc is 1.05× slower than C++ (clang)mem_copy: xc is 1.06× faster than Swiftarc_arrayarc_array: the four programsarc_array: xc is 9.7× faster than Objective-Carc_array: xc is 1.07× slower than C++ (clang)arc_array: xc is 1.59× faster than Swiftcall_depthcall_depth: the four programscall_depth: xc is 1.06× faster than Objective-Ccall_depth: xc is 1.06× faster than C++ (clang)call_depth: xc is 1.02× slower than Swiftmatrix_mulmatrix_mul: the four programsmatrix_mul: xc is 1.75× slower than Objective-Cmatrix_mul: xc is 1.72× slower than C++ (clang)matrix_mul: xc is 2.9× faster than Swifthash_mixhash_mix: the four programshash_mix: xc is 1.12× faster than Objective-Chash_mix: xc is 1.04× faster than C++ (clang)hash_mix: xc is the same than Swiftbranch_mixbranch_mix: the four programsbranch_mix: xc is 1.09× faster than Objective-Cbranch_mix: xc is 1.09× faster than C++ (clang)branch_mix: xc is 1.11× faster than Swiftint_accumint_accum: the four programsint_accum: xc is 1.02× faster than Objective-Cint_accum: xc is 1.02× faster than C++ (clang)int_accum: xc is 1.03× faster than Swiftarray_maparray_map: the four programsarray_map: xc is 1.42× faster than Objective-Carray_map: xc is the same than C++ (clang)array_map: xc is the same than Swiftbit_opsbit_ops: the four programsbit_ops: xc is 1.01× faster than Objective-Cbit_ops: xc is 1.01× faster than C++ (clang)bit_ops: xc is 1.01× faster than Swiftpoly_dispatchpoly_dispatch: the four programspoly_dispatch: xc is 1.73× faster than Objective-Cpoly_dispatch: xc is 1.17× slower than C++ (clang)poly_dispatch: xc is 3.2× faster than Swiftarray_sumarray_sum: the four programsarray_sum: xc is 1.01× slower than Objective-Carray_sum: xc is 1.01× slower than C++ (clang)array_sum: xc is 1.93× faster than Swiftsievesieve: the four programssieve: xc is 1.36× faster than Objective-Csieve: xc is 1.37× faster than C++ (clang)sieve: xc is 1.79× faster than Swiftstruct_copystruct_copy: the four programsstruct_copy: xc is 1.44× faster than Objective-Cstruct_copy: xc is 1.43× faster than C++ (clang)struct_copy: xc is 1.45× faster than Swiftstring_scanstring_scan: the four programsstring_scan: xc is 2.5× faster than Objective-Cstring_scan: xc is 2.6× faster than C++ (clang)string_scan: xc is 2.6× faster than Swiftmatrix_mul_f32matrix_mul_f32: the four programsmatrix_mul_f32: xc is 1.72× faster than Objective-Cmatrix_mul_f32: xc is 1.72× faster than C++ (clang)matrix_mul_f32: xc is 1.61× faster than Swift
arm64: xc against each language, per benchmark (0.7)
arm64: xc against each language, per benchmark (0.71)xc vs Objective-Cxc vs C++ (clang)xc vs C++ (GCC)xc vs Swift8× faster4× faster2× fastersame2× slower4× slower8× slower← xc fasterxc slower →int_muldivint_muldiv: the four programsint_muldiv: xc is 1.22× slower than Objective-Cint_muldiv: xc is 1.22× slower than C++ (clang)int_muldiv: xc is 1.07× faster than Swiftfloat_mathfloat_math: the four programsfloat_math: xc is 1.17× slower than Objective-Cfloat_math: xc is 1.18× slower than C++ (clang)float_math: xc is 1.16× slower than Swiftsort_smallsort_small: the four programssort_small: xc is 1.21× slower than Objective-Csort_small: xc is 1.16× slower than C++ (clang)sort_small: xc is 1.18× faster than Swiftmem_copymem_copy: the four programsmem_copy: xc is 1.63× faster than Objective-Cmem_copy: xc is 1.05× slower than C++ (clang)mem_copy: xc is 1.05× faster than Swiftarc_arrayarc_array: the four programsarc_array: xc is 9.7× faster than Objective-Carc_array: xc is 1.06× slower than C++ (clang)arc_array: xc is 1.61× faster than Swiftcall_depthcall_depth: the four programscall_depth: xc is 1.08× faster than Objective-Ccall_depth: xc is 1.08× faster than C++ (clang)call_depth: xc is 1.01× slower than Swiftmatrix_mulmatrix_mul: the four programsmatrix_mul: xc is 1.75× slower than Objective-Cmatrix_mul: xc is 1.72× slower than C++ (clang)matrix_mul: xc is 2.9× faster than Swifthash_mixhash_mix: the four programshash_mix: xc is the same than Objective-Chash_mix: xc is 1.01× slower than C++ (clang)hash_mix: xc is the same than Swiftbranch_mixbranch_mix: the four programsbranch_mix: xc is the same than Objective-Cbranch_mix: xc is 1.01× faster than C++ (clang)branch_mix: xc is 1.02× faster than Swiftint_accumint_accum: the four programsint_accum: xc is the same than Objective-Cint_accum: xc is 1.01× faster than C++ (clang)int_accum: xc is 1.02× faster than Swiftarray_maparray_map: the four programsarray_map: xc is 1.37× faster than Objective-Carray_map: xc is 1.04× slower than C++ (clang)array_map: xc is the same than Swiftbit_opsbit_ops: the four programsbit_ops: xc is 1.03× faster than Objective-Cbit_ops: xc is 1.03× faster than C++ (clang)bit_ops: xc is 1.04× faster than Swiftpoly_dispatchpoly_dispatch: the four programspoly_dispatch: xc is 1.73× faster than Objective-Cpoly_dispatch: xc is 1.18× slower than C++ (clang)poly_dispatch: xc is 3.1× faster than Swiftarray_sumarray_sum: the four programsarray_sum: xc is 1.02× slower than Objective-Carray_sum: xc is 1.03× slower than C++ (clang)array_sum: xc is 1.85× faster than Swiftsievesieve: the four programssieve: xc is 1.36× faster than Objective-Csieve: xc is 1.37× faster than C++ (clang)sieve: xc is 1.80× faster than Swiftstruct_copystruct_copy: the four programsstruct_copy: xc is 1.43× faster than Objective-Cstruct_copy: xc is 1.44× faster than C++ (clang)struct_copy: xc is 1.44× faster than Swiftstring_scanstring_scan: the four programsstring_scan: xc is 2.5× faster than Objective-Cstring_scan: xc is 2.6× faster than C++ (clang)string_scan: xc is 2.6× faster than Swiftmatrix_mul_f32matrix_mul_f32: the four programsmatrix_mul_f32: xc is 153× faster than Objective-Cmatrix_mul_f32: xc is 153× faster than C++ (clang)matrix_mul_f32: xc is 143× faster than Swift
arm64: xc against each language, per benchmark (0.71)
arm64: xc against each language, per benchmark (0.72)xc vs Objective-Cxc vs C++ (clang)xc vs C++ (GCC)xc vs Swift8× faster4× faster2× fastersame2× slower4× slower8× slower← xc fasterxc slower →int_muldivint_muldiv: the four programsint_muldiv: xc is 1.22× slower than Objective-Cint_muldiv: xc is 1.21× slower than C++ (clang)int_muldiv: xc is 1.08× faster than Swiftfloat_mathfloat_math: the four programsfloat_math: xc is 1.19× slower than Objective-Cfloat_math: xc is 1.19× slower than C++ (clang)float_math: xc is 1.17× slower than Swiftsort_smallsort_small: the four programssort_small: xc is 1.23× slower than Objective-Csort_small: xc is 1.18× slower than C++ (clang)sort_small: xc is 1.15× faster than Swiftmem_copymem_copy: the four programsmem_copy: xc is 1.63× faster than Objective-Cmem_copy: xc is 1.05× slower than C++ (clang)mem_copy: xc is 1.06× faster than Swiftarc_arrayarc_array: the four programsarc_array: xc is 9.7× faster than Objective-Carc_array: xc is 1.05× slower than C++ (clang)arc_array: xc is 1.63× faster than Swiftcall_depthcall_depth: the four programscall_depth: xc is 1.08× faster than Objective-Ccall_depth: xc is 1.08× faster than C++ (clang)call_depth: xc is 1.01× slower than Swiftmatrix_mulmatrix_mul: the four programsmatrix_mul: xc is 1.76× slower than Objective-Cmatrix_mul: xc is 1.73× slower than C++ (clang)matrix_mul: xc is 2.9× faster than Swifthash_mixhash_mix: the four programshash_mix: xc is the same than Objective-Chash_mix: xc is the same than C++ (clang)hash_mix: xc is the same than Swiftbranch_mixbranch_mix: the four programsbranch_mix: xc is the same than Objective-Cbranch_mix: xc is the same than C++ (clang)branch_mix: xc is 1.02× faster than Swiftint_accumint_accum: the four programsint_accum: xc is 1.01× faster than Objective-Cint_accum: xc is 1.01× faster than C++ (clang)int_accum: xc is 1.02× faster than Swiftarray_maparray_map: the four programsarray_map: xc is 1.38× faster than Objective-Carray_map: xc is 1.01× faster than C++ (clang)array_map: xc is 1.01× faster than Swiftbit_opsbit_ops: the four programsbit_ops: xc is 1.03× faster than Objective-Cbit_ops: xc is 1.03× faster than C++ (clang)bit_ops: xc is 1.05× faster than Swiftpoly_dispatchpoly_dispatch: the four programspoly_dispatch: xc is 1.73× faster than Objective-Cpoly_dispatch: xc is 1.17× slower than C++ (clang)poly_dispatch: xc is 3.2× faster than Swiftarray_sumarray_sum: the four programsarray_sum: xc is 1.04× slower than Objective-Carray_sum: xc is 1.04× slower than C++ (clang)array_sum: xc is 1.81× faster than Swiftsievesieve: the four programssieve: xc is 1.26× faster than Objective-Csieve: xc is 1.27× faster than C++ (clang)sieve: xc is 1.67× faster than Swiftstruct_copystruct_copy: the four programsstruct_copy: xc is 1.42× faster than Objective-Cstruct_copy: xc is 1.43× faster than C++ (clang)struct_copy: xc is 1.45× faster than Swiftstring_scanstring_scan: the four programsstring_scan: xc is 2.5× faster than Objective-Cstring_scan: xc is 2.6× faster than C++ (clang)string_scan: xc is 2.6× faster than Swiftmatrix_mul_f32matrix_mul_f32: the four programsmatrix_mul_f32: xc is 325× faster than Objective-Cmatrix_mul_f32: xc is 327× faster than C++ (clang)matrix_mul_f32: xc is 313× faster than Swift
arm64: xc against each language, per benchmark (0.72)
arm64: xc against each language, per benchmark (0.73)xc vs Objective-Cxc vs C++ (clang)xc vs C++ (GCC)xc vs Swift8× faster4× faster2× fastersame2× slower4× slower8× slower← xc fasterxc slower →int_muldivint_muldiv: the four programsint_muldiv: xc is 1.22× slower than Objective-Cint_muldiv: xc is 1.22× slower than C++ (clang)int_muldiv: xc is 1.08× faster than Swiftfloat_mathfloat_math: the four programsfloat_math: xc is 1.19× slower than Objective-Cfloat_math: xc is 1.19× slower than C++ (clang)float_math: xc is 1.17× slower than Swiftsort_smallsort_small: the four programssort_small: xc is 1.08× slower than Objective-Csort_small: xc is 1.04× slower than C++ (clang)sort_small: xc is 1.28× faster than Swiftmem_copymem_copy: the four programsmem_copy: xc is 1.64× faster than Objective-Cmem_copy: xc is 1.05× slower than C++ (clang)mem_copy: xc is 1.06× faster than Swiftarc_arrayarc_array: the four programsarc_array: xc is 9.7× faster than Objective-Carc_array: xc is 1.05× slower than C++ (clang)arc_array: xc is 1.63× faster than Swiftcall_depthcall_depth: the four programscall_depth: xc is 1.06× faster than Objective-Ccall_depth: xc is 1.07× faster than C++ (clang)call_depth: xc is 1.02× slower than Swiftmatrix_mulmatrix_mul: the four programsmatrix_mul: xc is 1.02× slower than Objective-Cmatrix_mul: xc is the same than C++ (clang)matrix_mul: xc is 5.0× faster than Swifthash_mixhash_mix: the four programshash_mix: xc is the same than Objective-Chash_mix: xc is the same than C++ (clang)hash_mix: xc is the same than Swiftbranch_mixbranch_mix: the four programsbranch_mix: xc is the same than Objective-Cbranch_mix: xc is the same than C++ (clang)branch_mix: xc is 1.02× faster than Swiftint_accumint_accum: the four programsint_accum: xc is the same than Objective-Cint_accum: xc is 1.01× faster than C++ (clang)int_accum: xc is 1.01× faster than Swiftarray_maparray_map: the four programsarray_map: xc is 1.37× faster than Objective-Carray_map: xc is 1.01× slower than C++ (clang)array_map: xc is 1.01× faster than Swiftbit_opsbit_ops: the four programsbit_ops: xc is 1.03× faster than Objective-Cbit_ops: xc is 1.04× faster than C++ (clang)bit_ops: xc is 1.04× faster than Swiftpoly_dispatchpoly_dispatch: the four programspoly_dispatch: xc is 2.1× faster than Objective-Cpoly_dispatch: xc is 1.05× faster than C++ (clang)poly_dispatch: xc is 3.9× faster than Swiftarray_sumarray_sum: the four programsarray_sum: xc is 1.10× faster than Objective-Carray_sum: xc is 1.09× faster than C++ (clang)array_sum: xc is 2.1× faster than Swiftsievesieve: the four programssieve: xc is 1.39× faster than Objective-Csieve: xc is 1.40× faster than C++ (clang)sieve: xc is 1.83× faster than Swiftstruct_copystruct_copy: the four programsstruct_copy: xc is 1.43× faster than Objective-Cstruct_copy: xc is 1.44× faster than C++ (clang)struct_copy: xc is 1.44× faster than Swiftstring_scanstring_scan: the four programsstring_scan: xc is 2.5× faster than Objective-Cstring_scan: xc is 2.6× faster than C++ (clang)string_scan: xc is 2.6× faster than Swiftmatrix_mul_f32matrix_mul_f32: the four programsmatrix_mul_f32: xc is 326× faster than Objective-Cmatrix_mul_f32: xc is 327× faster than C++ (clang)matrix_mul_f32: xc is 323× faster than Swift
arm64: xc against each language, per benchmark (0.73)
arm64: xc against each language, per benchmark (0.74)xc vs Objective-Cxc vs C++ (clang)xc vs C++ (GCC)xc vs Swift8× faster4× faster2× fastersame2× slower4× slower8× slower← xc fasterxc slower →int_muldivint_muldiv: the four programsint_muldiv: xc is 1.22× slower than Objective-Cint_muldiv: xc is 1.21× slower than C++ (clang)int_muldiv: xc is 1.08× faster than Swiftfloat_mathfloat_math: the four programsfloat_math: xc is 1.19× slower than Objective-Cfloat_math: xc is 1.19× slower than C++ (clang)float_math: xc is 1.17× slower than Swiftsort_smallsort_small: the four programssort_small: xc is 1.08× slower than Objective-Csort_small: xc is 1.03× slower than C++ (clang)sort_small: xc is 1.27× faster than Swiftmem_copymem_copy: the four programsmem_copy: xc is 1.63× faster than Objective-Cmem_copy: xc is 1.05× slower than C++ (clang)mem_copy: xc is 1.06× faster than Swiftarc_arrayarc_array: the four programsarc_array: xc is 9.7× faster than Objective-Carc_array: xc is 1.04× slower than C++ (clang)arc_array: xc is 1.64× faster than Swiftcall_depthcall_depth: the four programscall_depth: xc is 1.06× faster than Objective-Ccall_depth: xc is 1.06× faster than C++ (clang)call_depth: xc is 1.02× slower than Swiftmatrix_mulmatrix_mul: the four programsmatrix_mul: xc is 1.01× slower than Objective-Cmatrix_mul: xc is the same than C++ (clang)matrix_mul: xc is 5.1× faster than Swifthash_mixhash_mix: the four programshash_mix: xc is 1.01× slower than Objective-Chash_mix: xc is the same than C++ (clang)hash_mix: xc is the same than Swiftbranch_mixbranch_mix: the four programsbranch_mix: xc is the same than Objective-Cbranch_mix: xc is the same than C++ (clang)branch_mix: xc is 1.02× faster than Swiftint_accumint_accum: the four programsint_accum: xc is the same than Objective-Cint_accum: xc is the same than C++ (clang)int_accum: xc is 1.01× faster than Swiftarray_maparray_map: the four programsarray_map: xc is 1.35× faster than Objective-Carray_map: xc is the same than C++ (clang)array_map: xc is the same than Swiftbit_opsbit_ops: the four programsbit_ops: xc is 1.03× faster than Objective-Cbit_ops: xc is 1.04× faster than C++ (clang)bit_ops: xc is 1.05× faster than Swiftpoly_dispatchpoly_dispatch: the four programspoly_dispatch: xc is 2.1× faster than Objective-Cpoly_dispatch: xc is 1.05× faster than C++ (clang)poly_dispatch: xc is 3.9× faster than Swiftarray_sumarray_sum: the four programsarray_sum: xc is 1.10× faster than Objective-Carray_sum: xc is 1.10× faster than C++ (clang)array_sum: xc is 2.1× faster than Swiftsievesieve: the four programssieve: xc is 1.35× faster than Objective-Csieve: xc is 1.36× faster than C++ (clang)sieve: xc is 1.79× faster than Swiftstruct_copystruct_copy: the four programsstruct_copy: xc is 1.44× faster than Objective-Cstruct_copy: xc is 1.44× faster than C++ (clang)struct_copy: xc is 1.45× faster than Swiftstring_scanstring_scan: the four programsstring_scan: xc is 2.5× faster than Objective-Cstring_scan: xc is 2.6× faster than C++ (clang)string_scan: xc is 2.6× faster than Swiftmatrix_mul_f32matrix_mul_f32: the four programsmatrix_mul_f32: xc is 325× faster than Objective-Cmatrix_mul_f32: xc is 325× faster than C++ (clang)matrix_mul_f32: xc is 322× faster than Swift
arm64: xc against each language, per benchmark (0.74)
Release:
x86-64: xc against each language, per benchmark (0.62)xc vs Objective-Cxc vs C++ (clang)xc vs C++ (GCC)xc vs Swift8× faster4× faster2× fastersame2× slower4× slower8× slower← xc fasterxc slower →sort_smallsort_small: the four programssort_small: xc is 1.50× slower than Objective-Carc_arrayarc_array: the four programsarc_array: xc is 3.2× faster than Objective-Csievesieve: the four programssieve: xc is 2.3× slower than Objective-Cstruct_copystruct_copy: the four programsstruct_copy: xc is 1.44× slower than Objective-Cfloat_mathfloat_math: the four programsfloat_math: xc is 2.0× slower than Objective-Cpoly_dispatchpoly_dispatch: the four programspoly_dispatch: xc is 1.27× faster than Objective-Cmatrix_mulmatrix_mul: the four programsmatrix_mul: xc is 2.0× slower than Objective-Cbranch_mixbranch_mix: the four programsbranch_mix: xc is 1.01× slower than Objective-Cint_accumint_accum: the four programsint_accum: xc is the same than Objective-Cbit_opsbit_ops: the four programsbit_ops: xc is the same than Objective-Chash_mixhash_mix: the four programshash_mix: xc is 1.10× faster than Objective-Cint_muldivint_muldiv: the four programsint_muldiv: xc is 1.19× slower than Objective-Cstring_scanstring_scan: the four programsstring_scan: xc is 1.94× faster than Objective-Cmem_copymem_copy: the four programsmem_copy: xc is 1.46× slower than Objective-Carray_sumarray_sum: the four programsarray_sum: xc is 1.20× faster than Objective-Carray_maparray_map: the four programsarray_map: xc is 1.35× slower than Objective-Ccall_depthcall_depth: the four programscall_depth: xc is 1.11× slower than Objective-Cmatrix_mul_f32matrix_mul_f32: the four programsmatrix_mul_f32: xc is 2.5× slower than Objective-Cmatrix_mul_f32: xc is 2.6× slower than C++ (clang)matrix_mul_f32: xc is 1.79× slower than Swift
x86-64: xc against each language, per benchmark (0.62)
x86-64: xc against each language, per benchmark (0.63)xc vs Objective-Cxc vs C++ (clang)xc vs C++ (GCC)xc vs Swift8× faster4× faster2× fastersame2× slower4× slower8× slower← xc fasterxc slower →sort_smallsort_small: the four programssort_small: xc is 1.25× slower than Objective-Carc_arrayarc_array: the four programsarc_array: xc is 3.2× faster than Objective-Csievesieve: the four programssieve: xc is 2.3× slower than Objective-Cstruct_copystruct_copy: the four programsstruct_copy: xc is 1.44× slower than Objective-Cfloat_mathfloat_math: the four programsfloat_math: xc is 2.0× slower than Objective-Cpoly_dispatchpoly_dispatch: the four programspoly_dispatch: xc is 1.27× faster than Objective-Cmatrix_mulmatrix_mul: the four programsmatrix_mul: xc is 2.0× slower than Objective-Cbranch_mixbranch_mix: the four programsbranch_mix: xc is 1.13× slower than Objective-Cint_accumint_accum: the four programsint_accum: xc is the same than Objective-Cbit_opsbit_ops: the four programsbit_ops: xc is the same than Objective-Chash_mixhash_mix: the four programshash_mix: xc is 1.07× faster than Objective-Cint_muldivint_muldiv: the four programsint_muldiv: xc is 1.19× slower than Objective-Cstring_scanstring_scan: the four programsstring_scan: xc is 1.94× faster than Objective-Cmem_copymem_copy: the four programsmem_copy: xc is 1.46× slower than Objective-Carray_sumarray_sum: the four programsarray_sum: xc is 1.19× faster than Objective-Carray_maparray_map: the four programsarray_map: xc is 1.37× slower than Objective-Ccall_depthcall_depth: the four programscall_depth: xc is 1.11× slower than Objective-Cmatrix_mul_f32matrix_mul_f32: the four programsmatrix_mul_f32: xc is 2.4× slower than Objective-Cmatrix_mul_f32: xc is 2.4× slower than C++ (clang)matrix_mul_f32: xc is 1.70× slower than Swift
x86-64: xc against each language, per benchmark (0.63)
x86-64: xc against each language, per benchmark (0.64)xc vs Objective-Cxc vs C++ (clang)xc vs C++ (GCC)xc vs Swift8× faster4× faster2× fastersame2× slower4× slower8× slower← xc fasterxc slower →sort_smallsort_small: the four programssort_small: xc is 1.21× slower than Objective-Carc_arrayarc_array: the four programsarc_array: xc is 3.2× faster than Objective-Csievesieve: the four programssieve: xc is 1.82× slower than Objective-Cstruct_copystruct_copy: the four programsstruct_copy: xc is 1.09× slower than Objective-Cfloat_mathfloat_math: the four programsfloat_math: xc is 1.01× faster than Objective-Cpoly_dispatchpoly_dispatch: the four programspoly_dispatch: xc is 1.27× faster than Objective-Cmatrix_mulmatrix_mul: the four programsmatrix_mul: xc is 1.08× slower than Objective-Cbranch_mixbranch_mix: the four programsbranch_mix: xc is the same than Objective-Cint_accumint_accum: the four programsint_accum: xc is the same than Objective-Cbit_opsbit_ops: the four programsbit_ops: xc is the same than Objective-Chash_mixhash_mix: the four programshash_mix: xc is 1.07× faster than Objective-Cint_muldivint_muldiv: the four programsint_muldiv: xc is 1.19× slower than Objective-Cstring_scanstring_scan: the four programsstring_scan: xc is 1.95× faster than Objective-Cmem_copymem_copy: the four programsmem_copy: xc is 1.02× slower than Objective-Carray_sumarray_sum: the four programsarray_sum: xc is 1.55× faster than Objective-Carray_maparray_map: the four programsarray_map: xc is 1.01× faster than Objective-Ccall_depthcall_depth: the four programscall_depth: xc is 1.11× slower than Objective-Cmatrix_mul_f32matrix_mul_f32: the four programsmatrix_mul_f32: xc is 1.98× slower than Objective-Cmatrix_mul_f32: xc is 2.00× slower than C++ (clang)matrix_mul_f32: xc is 1.40× slower than Swift
x86-64: xc against each language, per benchmark (0.64)
x86-64: xc against each language, per benchmark (0.65)xc vs Objective-Cxc vs C++ (clang)xc vs C++ (GCC)xc vs Swift8× faster4× faster2× fastersame2× slower4× slower8× slower← xc fasterxc slower →sort_smallsort_small: the four programssort_small: xc is 1.19× slower than Objective-Csort_small: xc is 1.22× slower than C++ (clang)sort_small: xc is 1.75× slower than Swiftarc_arrayarc_array: the four programsarc_array: xc is 3.2× faster than Objective-Carc_array: xc is 1.84× slower than C++ (clang)arc_array: xc is 1.92× slower than Swiftsievesieve: the four programssieve: xc is 1.49× slower than Objective-Csieve: xc is 1.48× slower than C++ (clang)sieve: xc is 1.01× slower than Swiftstruct_copystruct_copy: the four programsstruct_copy: xc is 1.10× slower than Objective-Cstruct_copy: xc is 1.11× slower than C++ (clang)struct_copy: xc is 1.11× slower than Swiftfloat_mathfloat_math: the four programsfloat_math: xc is 1.05× slower than Objective-Cfloat_math: xc is 1.05× slower than C++ (clang)float_math: xc is 1.05× slower than Swiftpoly_dispatchpoly_dispatch: the four programspoly_dispatch: xc is 1.28× faster than Objective-Cpoly_dispatch: xc is 1.04× slower than C++ (clang)poly_dispatch: xc is 6.9× faster than Swiftmatrix_mulmatrix_mul: the four programsmatrix_mul: xc is 1.10× faster than Objective-Cmatrix_mul: xc is 1.10× faster than C++ (clang)matrix_mul: xc is 3.6× faster than Swiftbranch_mixbranch_mix: the four programsbranch_mix: xc is the same than Objective-Cbranch_mix: xc is the same than C++ (clang)branch_mix: xc is the same than Swiftint_accumint_accum: the four programsint_accum: xc is the same than Objective-Cint_accum: xc is the same than C++ (clang)int_accum: xc is the same than Swiftbit_opsbit_ops: the four programsbit_ops: xc is the same than Objective-Cbit_ops: xc is the same than C++ (clang)bit_ops: xc is the same than Swifthash_mixhash_mix: the four programshash_mix: xc is 1.08× faster than Objective-Chash_mix: xc is 1.07× faster than C++ (clang)hash_mix: xc is 1.06× faster than Swiftint_muldivint_muldiv: the four programsint_muldiv: xc is 1.19× slower than Objective-Cint_muldiv: xc is 1.19× slower than C++ (clang)int_muldiv: xc is 1.19× slower than Swiftstring_scanstring_scan: the four programsstring_scan: xc is 1.94× faster than Objective-Cstring_scan: xc is 1.95× faster than C++ (clang)string_scan: xc is 1.95× faster than Swiftmem_copymem_copy: the four programsmem_copy: xc is 1.02× slower than Objective-Cmem_copy: xc is 1.06× slower than C++ (clang)mem_copy: xc is 1.12× slower than Swiftarray_sumarray_sum: the four programsarray_sum: xc is 1.58× faster than Objective-Carray_sum: xc is 1.19× faster than C++ (clang)array_sum: xc is 1.19× faster than Swiftarray_maparray_map: the four programsarray_map: xc is 1.01× faster than Objective-Carray_map: xc is 1.06× slower than C++ (clang)array_map: xc is 1.32× faster than Swiftcall_depthcall_depth: the four programscall_depth: xc is 1.11× slower than Objective-Ccall_depth: xc is 1.11× slower than C++ (clang)call_depth: xc is 1.09× slower than Swiftmatrix_mul_f32matrix_mul_f32: the four programsmatrix_mul_f32: xc is 1.76× slower than Objective-Cmatrix_mul_f32: xc is 1.78× slower than C++ (clang)matrix_mul_f32: xc is 1.24× slower than Swift
x86-64: xc against each language, per benchmark (0.65)
x86-64: xc against each language, per benchmark (0.66)xc vs Objective-Cxc vs C++ (clang)xc vs C++ (GCC)xc vs Swift8× faster4× faster2× fastersame2× slower4× slower8× slower← xc fasterxc slower →sort_smallsort_small: the four programssort_small: xc is 1.14× slower than Objective-Csort_small: xc is 1.14× slower than C++ (clang)sort_small: xc is 1.67× slower than Swiftarc_arrayarc_array: the four programsarc_array: xc is 3.0× faster than Objective-Carc_array: xc is 1.94× slower than C++ (clang)arc_array: xc is 2.0× slower than Swiftsievesieve: the four programssieve: xc is 1.67× slower than Objective-Csieve: xc is 1.69× slower than C++ (clang)sieve: xc is 1.13× slower than Swiftstruct_copystruct_copy: the four programsstruct_copy: xc is 1.09× slower than Objective-Cstruct_copy: xc is 1.09× slower than C++ (clang)struct_copy: xc is 1.09× slower than Swiftfloat_mathfloat_math: the four programsfloat_math: xc is 1.02× slower than Objective-Cfloat_math: xc is 1.02× slower than C++ (clang)float_math: xc is 1.02× slower than Swiftpoly_dispatchpoly_dispatch: the four programspoly_dispatch: xc is 1.27× faster than Objective-Cpoly_dispatch: xc is 1.04× slower than C++ (clang)poly_dispatch: xc is 6.9× faster than Swiftmatrix_mulmatrix_mul: the four programsmatrix_mul: xc is 1.10× faster than Objective-Cmatrix_mul: xc is 1.10× faster than C++ (clang)matrix_mul: xc is 3.6× faster than Swiftbranch_mixbranch_mix: the four programsbranch_mix: xc is the same than Objective-Cbranch_mix: xc is the same than C++ (clang)branch_mix: xc is the same than Swiftint_accumint_accum: the four programsint_accum: xc is the same than Objective-Cint_accum: xc is the same than C++ (clang)int_accum: xc is the same than Swiftbit_opsbit_ops: the four programsbit_ops: xc is the same than Objective-Cbit_ops: xc is the same than C++ (clang)bit_ops: xc is the same than Swifthash_mixhash_mix: the four programshash_mix: xc is 1.06× faster than Objective-Chash_mix: xc is 1.06× faster than C++ (clang)hash_mix: xc is 1.06× faster than Swiftint_muldivint_muldiv: the four programsint_muldiv: xc is 1.67× faster than Objective-Cint_muldiv: xc is 1.66× faster than C++ (clang)int_muldiv: xc is 1.67× faster than Swiftstring_scanstring_scan: the four programsstring_scan: xc is 1.94× faster than Objective-Cstring_scan: xc is 1.95× faster than C++ (clang)string_scan: xc is 1.95× faster than Swiftmem_copymem_copy: the four programsmem_copy: xc is 1.52× faster than Objective-Cmem_copy: xc is 1.47× faster than C++ (clang)mem_copy: xc is 1.40× faster than Swiftarray_sumarray_sum: the four programsarray_sum: xc is 2.9× faster than Objective-Carray_sum: xc is 2.2× faster than C++ (clang)array_sum: xc is 2.2× faster than Swiftarray_maparray_map: the four programsarray_map: xc is 2.0× faster than Objective-Carray_map: xc is 1.91× faster than C++ (clang)array_map: xc is 2.7× faster than Swiftcall_depthcall_depth: the four programscall_depth: xc is 1.57× slower than Objective-Ccall_depth: xc is 1.57× slower than C++ (clang)call_depth: xc is 1.54× slower than Swiftmatrix_mul_f32matrix_mul_f32: the four programsmatrix_mul_f32: xc is 1.53× slower than Objective-Cmatrix_mul_f32: xc is 1.54× slower than C++ (clang)matrix_mul_f32: xc is 1.08× slower than Swift
x86-64: xc against each language, per benchmark (0.66)
x86-64: xc against each language, per benchmark (0.7)xc vs Objective-Cxc vs C++ (clang)xc vs C++ (GCC)xc vs Swift8× faster4× faster2× fastersame2× slower4× slower8× slower← xc fasterxc slower →sort_smallsort_small: the four programssort_small: xc is 1.19× slower than Objective-Csort_small: xc is 1.19× slower than C++ (clang)sort_small: xc is 1.72× slower than Swiftarc_arrayarc_array: the four programsarc_array: xc is 3.0× faster than Objective-Carc_array: xc is 1.89× slower than C++ (clang)arc_array: xc is 2.0× slower than Swiftsievesieve: the four programssieve: xc is 1.63× slower than Objective-Csieve: xc is 1.65× slower than C++ (clang)sieve: xc is 1.12× slower than Swiftstruct_copystruct_copy: the four programsstruct_copy: xc is 1.09× slower than Objective-Cstruct_copy: xc is 1.09× slower than C++ (clang)struct_copy: xc is 1.09× slower than Swiftfloat_mathfloat_math: the four programsfloat_math: xc is 1.02× slower than Objective-Cfloat_math: xc is 1.01× slower than C++ (clang)float_math: xc is 1.02× slower than Swiftpoly_dispatchpoly_dispatch: the four programspoly_dispatch: xc is 1.27× faster than Objective-Cpoly_dispatch: xc is 1.04× slower than C++ (clang)poly_dispatch: xc is 6.9× faster than Swiftmatrix_mulmatrix_mul: the four programsmatrix_mul: xc is 1.10× faster than Objective-Cmatrix_mul: xc is 1.10× faster than C++ (clang)matrix_mul: xc is 3.6× faster than Swiftbranch_mixbranch_mix: the four programsbranch_mix: xc is the same than Objective-Cbranch_mix: xc is the same than C++ (clang)branch_mix: xc is the same than Swiftint_accumint_accum: the four programsint_accum: xc is the same than Objective-Cint_accum: xc is the same than C++ (clang)int_accum: xc is the same than Swiftbit_opsbit_ops: the four programsbit_ops: xc is the same than Objective-Cbit_ops: xc is the same than C++ (clang)bit_ops: xc is the same than Swifthash_mixhash_mix: the four programshash_mix: xc is 1.06× faster than Objective-Chash_mix: xc is 1.06× faster than C++ (clang)hash_mix: xc is 1.06× faster than Swiftint_muldivint_muldiv: the four programsint_muldiv: xc is 1.69× faster than Objective-Cint_muldiv: xc is 1.69× faster than C++ (clang)int_muldiv: xc is 1.69× faster than Swiftstring_scanstring_scan: the four programsstring_scan: xc is 1.95× faster than Objective-Cstring_scan: xc is 1.96× faster than C++ (clang)string_scan: xc is 1.95× faster than Swiftmem_copymem_copy: the four programsmem_copy: xc is 1.87× faster than Objective-Cmem_copy: xc is 1.81× faster than C++ (clang)mem_copy: xc is 1.73× faster than Swiftarray_sumarray_sum: the four programsarray_sum: xc is 3.1× faster than Objective-Carray_sum: xc is 2.3× faster than C++ (clang)array_sum: xc is 2.3× faster than Swiftarray_maparray_map: the four programsarray_map: xc is 2.0× faster than Objective-Carray_map: xc is 1.91× faster than C++ (clang)array_map: xc is 2.7× faster than Swiftcall_depthcall_depth: the four programscall_depth: xc is 1.83× faster than Objective-Ccall_depth: xc is 1.83× faster than C++ (clang)call_depth: xc is 1.87× faster than Swiftmatrix_mul_f32matrix_mul_f32: the four programsmatrix_mul_f32: xc is 1.56× slower than Objective-Cmatrix_mul_f32: xc is 1.57× slower than C++ (clang)matrix_mul_f32: xc is 1.10× slower than Swift
x86-64: xc against each language, per benchmark (0.7)
x86-64: xc against each language, per benchmark (0.71)xc vs Objective-Cxc vs C++ (clang)xc vs C++ (GCC)xc vs Swift8× faster4× faster2× fastersame2× slower4× slower8× slower← xc fasterxc slower →sort_smallsort_small: the four programssort_small: xc is 1.24× slower than Objective-Csort_small: xc is 1.24× slower than C++ (clang)sort_small: xc is 1.81× slower than Swiftarc_arrayarc_array: the four programsarc_array: xc is 3.0× faster than Objective-Carc_array: xc is 1.84× slower than C++ (clang)arc_array: xc is 2.0× slower than Swiftsievesieve: the four programssieve: xc is 1.45× slower than Objective-Csieve: xc is 1.51× slower than C++ (clang)sieve: xc is 1.02× slower than Swiftstruct_copystruct_copy: the four programsstruct_copy: xc is 1.11× slower than Objective-Cstruct_copy: xc is 1.11× slower than C++ (clang)struct_copy: xc is 1.11× slower than Swiftfloat_mathfloat_math: the four programsfloat_math: xc is 1.05× slower than Objective-Cfloat_math: xc is 1.05× slower than C++ (clang)float_math: xc is 1.06× slower than Swiftpoly_dispatchpoly_dispatch: the four programspoly_dispatch: xc is 1.28× faster than Objective-Cpoly_dispatch: xc is 1.04× slower than C++ (clang)poly_dispatch: xc is 6.9× faster than Swiftmatrix_mulmatrix_mul: the four programsmatrix_mul: xc is 1.10× faster than Objective-Cmatrix_mul: xc is 1.10× faster than C++ (clang)matrix_mul: xc is 3.6× faster than Swiftbranch_mixbranch_mix: the four programsbranch_mix: xc is the same than Objective-Cbranch_mix: xc is the same than C++ (clang)branch_mix: xc is the same than Swiftint_accumint_accum: the four programsint_accum: xc is the same than Objective-Cint_accum: xc is the same than C++ (clang)int_accum: xc is the same than Swiftbit_opsbit_ops: the four programsbit_ops: xc is the same than Objective-Cbit_ops: xc is the same than C++ (clang)bit_ops: xc is the same than Swifthash_mixhash_mix: the four programshash_mix: xc is 1.06× faster than Objective-Chash_mix: xc is 1.06× faster than C++ (clang)hash_mix: xc is 1.06× faster than Swiftint_muldivint_muldiv: the four programsint_muldiv: xc is 3.3× faster than Objective-Cint_muldiv: xc is 3.3× faster than C++ (clang)int_muldiv: xc is 3.3× faster than Swiftstring_scanstring_scan: the four programsstring_scan: xc is 1.94× faster than Objective-Cstring_scan: xc is 1.95× faster than C++ (clang)string_scan: xc is 1.95× faster than Swiftmem_copymem_copy: the four programsmem_copy: xc is 2.8× faster than Objective-Cmem_copy: xc is 2.7× faster than C++ (clang)mem_copy: xc is 2.6× faster than Swiftarray_sumarray_sum: the four programsarray_sum: xc is 4.3× faster than Objective-Carray_sum: xc is 3.3× faster than C++ (clang)array_sum: xc is 3.3× faster than Swiftarray_maparray_map: the four programsarray_map: xc is 3.7× faster than Objective-Carray_map: xc is 3.4× faster than C++ (clang)array_map: xc is 4.8× faster than Swiftcall_depthcall_depth: the four programscall_depth: xc is 3.6× faster than Objective-Ccall_depth: xc is 3.6× faster than C++ (clang)call_depth: xc is 3.7× faster than Swiftmatrix_mul_f32matrix_mul_f32: the four programsmatrix_mul_f32: xc is 22× faster than Objective-Cmatrix_mul_f32: xc is 22× faster than C++ (clang)matrix_mul_f32: xc is 31× faster than Swift
x86-64: xc against each language, per benchmark (0.71)
x86-64: xc against each language, per benchmark (0.72)xc vs Objective-Cxc vs C++ (clang)xc vs C++ (GCC)xc vs Swift8× faster4× faster2× fastersame2× slower4× slower8× slower← xc fasterxc slower →sort_smallsort_small: the four programssort_small: xc is 1.23× slower than Objective-Csort_small: xc is 1.23× slower than C++ (clang)sort_small: xc is 1.79× slower than Swiftarc_arrayarc_array: the four programsarc_array: xc is 3.2× faster than Objective-Carc_array: xc is 1.78× slower than C++ (clang)arc_array: xc is 1.90× slower than Swiftsievesieve: the four programssieve: xc is 1.71× slower than Objective-Csieve: xc is 1.71× slower than C++ (clang)sieve: xc is 1.14× slower than Swiftstruct_copystruct_copy: the four programsstruct_copy: xc is 1.11× slower than Objective-Cstruct_copy: xc is 1.11× slower than C++ (clang)struct_copy: xc is 1.11× slower than Swiftfloat_mathfloat_math: the four programsfloat_math: xc is 1.05× slower than Objective-Cfloat_math: xc is 1.05× slower than C++ (clang)float_math: xc is 1.06× slower than Swiftpoly_dispatchpoly_dispatch: the four programspoly_dispatch: xc is 1.27× faster than Objective-Cpoly_dispatch: xc is 1.04× slower than C++ (clang)poly_dispatch: xc is 6.9× faster than Swiftmatrix_mulmatrix_mul: the four programsmatrix_mul: xc is 1.10× faster than Objective-Cmatrix_mul: xc is 1.10× faster than C++ (clang)matrix_mul: xc is 3.6× faster than Swiftbranch_mixbranch_mix: the four programsbranch_mix: xc is the same than Objective-Cbranch_mix: xc is the same than C++ (clang)branch_mix: xc is the same than Swiftint_accumint_accum: the four programsint_accum: xc is the same than Objective-Cint_accum: xc is the same than C++ (clang)int_accum: xc is the same than Swiftbit_opsbit_ops: the four programsbit_ops: xc is the same than Objective-Cbit_ops: xc is the same than C++ (clang)bit_ops: xc is the same than Swifthash_mixhash_mix: the four programshash_mix: xc is 1.06× faster than Objective-Chash_mix: xc is 1.06× faster than C++ (clang)hash_mix: xc is 1.06× faster than Swiftint_muldivint_muldiv: the four programsint_muldiv: xc is 3.3× faster than Objective-Cint_muldiv: xc is 3.3× faster than C++ (clang)int_muldiv: xc is 3.3× faster than Swiftstring_scanstring_scan: the four programsstring_scan: xc is 1.94× faster than Objective-Cstring_scan: xc is 1.95× faster than C++ (clang)string_scan: xc is 1.95× faster than Swiftmem_copymem_copy: the four programsmem_copy: xc is 3.0× faster than Objective-Cmem_copy: xc is 2.9× faster than C++ (clang)mem_copy: xc is 2.7× faster than Swiftarray_sumarray_sum: the four programsarray_sum: xc is 3.8× faster than Objective-Carray_sum: xc is 2.9× faster than C++ (clang)array_sum: xc is 2.9× faster than Swiftarray_maparray_map: the four programsarray_map: xc is 3.6× faster than Objective-Carray_map: xc is 3.4× faster than C++ (clang)array_map: xc is 4.8× faster than Swiftcall_depthcall_depth: the four programscall_depth: xc is 3.6× faster than Objective-Ccall_depth: xc is 3.6× faster than C++ (clang)call_depth: xc is 3.7× faster than Swiftmatrix_mul_f32matrix_mul_f32: the four programsmatrix_mul_f32: xc is 28× faster than Objective-Cmatrix_mul_f32: xc is 28× faster than C++ (clang)matrix_mul_f32: xc is 40× faster than Swift
x86-64: xc against each language, per benchmark (0.72)
x86-64: xc against each language, per benchmark (0.73)xc vs Objective-Cxc vs C++ (clang)xc vs C++ (GCC)xc vs Swift8× faster4× faster2× fastersame2× slower4× slower8× slower← xc fasterxc slower →sort_smallsort_small: the four programssort_small: xc is 1.27× slower than Objective-Csort_small: xc is 1.28× slower than C++ (clang)sort_small: xc is 1.86× slower than Swiftarc_arrayarc_array: the four programsarc_array: xc is 3.2× faster than Objective-Carc_array: xc is 1.79× slower than C++ (clang)arc_array: xc is 1.88× slower than Swiftsievesieve: the four programssieve: xc is 1.35× slower than Objective-Csieve: xc is 1.32× slower than C++ (clang)sieve: xc is 1.09× faster than Swiftstruct_copystruct_copy: the four programsstruct_copy: xc is 1.11× slower than Objective-Cstruct_copy: xc is 1.11× slower than C++ (clang)struct_copy: xc is 1.11× slower than Swiftfloat_mathfloat_math: the four programsfloat_math: xc is 1.02× slower than Objective-Cfloat_math: xc is 1.01× slower than C++ (clang)float_math: xc is 1.02× slower than Swiftpoly_dispatchpoly_dispatch: the four programspoly_dispatch: xc is 1.27× faster than Objective-Cpoly_dispatch: xc is 1.04× slower than C++ (clang)poly_dispatch: xc is 6.9× faster than Swiftmatrix_mulmatrix_mul: the four programsmatrix_mul: xc is 1.02× slower than Objective-Cmatrix_mul: xc is 1.02× slower than C++ (clang)matrix_mul: xc is 3.2× faster than Swiftbranch_mixbranch_mix: the four programsbranch_mix: xc is 1.01× slower than Objective-Cbranch_mix: xc is 1.01× slower than C++ (clang)branch_mix: xc is the same than Swiftint_accumint_accum: the four programsint_accum: xc is the same than Objective-Cint_accum: xc is the same than C++ (clang)int_accum: xc is the same than Swiftbit_opsbit_ops: the four programsbit_ops: xc is the same than Objective-Cbit_ops: xc is the same than C++ (clang)bit_ops: xc is the same than Swifthash_mixhash_mix: the four programshash_mix: xc is 1.06× faster than Objective-Chash_mix: xc is 1.06× faster than C++ (clang)hash_mix: xc is 1.06× faster than Swiftint_muldivint_muldiv: the four programsint_muldiv: xc is 3.3× faster than Objective-Cint_muldiv: xc is 3.3× faster than C++ (clang)int_muldiv: xc is 3.3× faster than Swiftstring_scanstring_scan: the four programsstring_scan: xc is 1.94× faster than Objective-Cstring_scan: xc is 1.95× faster than C++ (clang)string_scan: xc is 1.95× faster than Swiftmem_copymem_copy: the four programsmem_copy: xc is 3.0× faster than Objective-Cmem_copy: xc is 2.9× faster than C++ (clang)mem_copy: xc is 2.7× faster than Swiftarray_sumarray_sum: the four programsarray_sum: xc is 6.0× faster than Objective-Carray_sum: xc is 4.5× faster than C++ (clang)array_sum: xc is 4.5× faster than Swiftarray_maparray_map: the four programsarray_map: xc is 3.7× faster than Objective-Carray_map: xc is 3.4× faster than C++ (clang)array_map: xc is 4.9× faster than Swiftcall_depthcall_depth: the four programscall_depth: xc is 3.6× faster than Objective-Ccall_depth: xc is 3.6× faster than C++ (clang)call_depth: xc is 3.7× faster than Swiftmatrix_mul_f32matrix_mul_f32: the four programsmatrix_mul_f32: xc is 27× faster than Objective-Cmatrix_mul_f32: xc is 27× faster than C++ (clang)matrix_mul_f32: xc is 40× faster than Swift
x86-64: xc against each language, per benchmark (0.73)
x86-64: xc against each language, per benchmark (0.74)xc vs Objective-Cxc vs C++ (clang)xc vs C++ (GCC)xc vs Swift8× faster4× faster2× fastersame2× slower4× slower8× slower← xc fasterxc slower →sort_smallsort_small: the four programssort_small: xc is 1.27× slower than Objective-Csort_small: xc is 1.28× slower than C++ (clang)sort_small: xc is 2.3× slower than C++ (GCC)sort_small: xc is 1.87× slower than Swiftarc_arrayarc_array: the four programsarc_array: xc is 3.2× faster than Objective-Carc_array: xc is 1.83× slower than C++ (clang)arc_array: xc is 1.86× slower than C++ (GCC)arc_array: xc is 1.92× slower than Swiftsievesieve: the four programssieve: xc is 1.32× slower than Objective-Csieve: xc is 1.36× slower than C++ (clang)sieve: xc is 1.61× faster than C++ (GCC)sieve: xc is 1.08× faster than Swiftstruct_copystruct_copy: the four programsstruct_copy: xc is 1.11× slower than Objective-Cstruct_copy: xc is 1.11× slower than C++ (clang)struct_copy: xc is 1.34× slower than C++ (GCC)struct_copy: xc is 1.11× slower than Swiftfloat_mathfloat_math: the four programsfloat_math: xc is 1.02× slower than Objective-Cfloat_math: xc is 1.01× slower than C++ (clang)float_math: xc is 1.09× slower than C++ (GCC)float_math: xc is 1.02× slower than Swiftpoly_dispatchpoly_dispatch: the four programspoly_dispatch: xc is 1.27× faster than Objective-Cpoly_dispatch: xc is 1.04× slower than C++ (clang)poly_dispatch: xc is 1.04× slower than C++ (GCC)poly_dispatch: xc is 6.9× faster than Swiftmatrix_mulmatrix_mul: the four programsmatrix_mul: xc is 1.02× slower than Objective-Cmatrix_mul: xc is 1.02× slower than C++ (clang)matrix_mul: xc is 1.16× faster than C++ (GCC)matrix_mul: xc is 3.2× faster than Swiftbranch_mixbranch_mix: the four programsbranch_mix: xc is 1.01× slower than Objective-Cbranch_mix: xc is the same than C++ (clang)branch_mix: xc is the same than C++ (GCC)branch_mix: xc is 1.01× slower than Swiftint_accumint_accum: the four programsint_accum: xc is the same than Objective-Cint_accum: xc is the same than C++ (clang)int_accum: xc is the same than C++ (GCC)int_accum: xc is the same than Swiftbit_opsbit_ops: the four programsbit_ops: xc is the same than Objective-Cbit_ops: xc is the same than C++ (clang)bit_ops: xc is the same than C++ (GCC)bit_ops: xc is the same than Swifthash_mixhash_mix: the four programshash_mix: xc is 1.06× faster than Objective-Chash_mix: xc is 1.06× faster than C++ (clang)hash_mix: xc is the same than C++ (GCC)hash_mix: xc is 1.06× faster than Swiftint_muldivint_muldiv: the four programsint_muldiv: xc is 3.3× faster than Objective-Cint_muldiv: xc is 3.3× faster than C++ (clang)int_muldiv: xc is 1.35× faster than C++ (GCC)int_muldiv: xc is 3.3× faster than Swiftstring_scanstring_scan: the four programsstring_scan: xc is 1.94× faster than Objective-Cstring_scan: xc is 1.95× faster than C++ (clang)string_scan: xc is 1.67× faster than C++ (GCC)string_scan: xc is 1.95× faster than Swiftmem_copymem_copy: the four programsmem_copy: xc is 1.97× faster than Objective-Cmem_copy: xc is 1.91× faster than C++ (clang)mem_copy: xc is 2.7× faster than C++ (GCC)mem_copy: xc is 1.82× faster than Swiftarray_sumarray_sum: the four programsarray_sum: xc is 3.9× faster than Objective-Carray_sum: xc is 3.0× faster than C++ (clang)array_sum: xc is 7.7× faster than C++ (GCC)array_sum: xc is 3.0× faster than Swiftarray_maparray_map: the four programsarray_map: xc is 3.7× faster than Objective-Carray_map: xc is 3.5× faster than C++ (clang)array_map: xc is 4.8× faster than C++ (GCC)array_map: xc is 4.9× faster than Swiftcall_depthcall_depth: the four programscall_depth: xc is 3.6× faster than Objective-Ccall_depth: xc is 3.6× faster than C++ (clang)call_depth: xc is 3.7× faster than C++ (GCC)call_depth: xc is 3.7× faster than Swiftmatrix_mul_f32matrix_mul_f32: the four programsmatrix_mul_f32: xc is 28× faster than Objective-Cmatrix_mul_f32: xc is 28× faster than C++ (clang)matrix_mul_f32: xc is 8.7× faster than C++ (GCC)matrix_mul_f32: xc is 39× faster than Swift
x86-64: xc against each language, per benchmark (0.74)
benchmarkxcObjective-CC++ (clang)C++ (GCC)Swiftxc vs the fastest of the others
int_muldiv1.140.940.94–1.231.22× slower
float_math0.970.820.81–0.831.19× slower
sort_small1.010.940.98–1.291.08× slower
mem_copy0.941.530.89–0.991.05× slower
arc_alloc1.011.430.96–2.211.05× slower
arc_array0.545.210.52–0.881.04× slower
call_depth0.930.990.99–0.911.02× slower
matrix_mul0.220.220.22–1.121.01× slower
hash_mix1.161.151.15–1.161.01× slower
branch_mix0.990.990.99–1.01the same
int_accum1.111.111.11–1.12the same
array_map1.051.421.06–1.06the same
bit_ops1.111.151.15–1.171.03× faster
poly_dispatch0.521.120.55–2.031.05× faster
array_sum0.810.890.89–1.671.10× faster
method_call0.842.291.10–1.081.28× faster
sieve1.071.441.46–1.911.35× faster
struct_copy0.690.990.99–1.001.44× faster
string_scan0.912.312.35–2.372.5× faster
matrix_mul_f320.039.809.81–9.71322× faster
benchmarkxcObjective-CC++ (clang)C++ (GCC)Swiftxc vs the fastest of the others
sort_small1.921.511.500.821.022.3× slower
arc_array1.183.730.640.630.611.92× slower
sieve1.130.850.831.811.221.36× slower
struct_copy0.870.790.790.650.791.34× slower
float_math0.760.740.740.690.741.09× slower
poly_dispatch0.971.240.930.936.691.04× slower
method_call1.311.901.591.581.271.03× slower
matrix_mul0.640.620.620.742.061.02× slower
branch_mix0.820.810.810.810.811.01× slower
int_accum0.980.970.970.970.97the same
bit_ops0.980.980.980.980.98the same
hash_mix1.011.081.081.011.08the same
arc_alloc0.382.180.440.450.761.16× faster
int_muldiv0.511.701.700.691.701.35× faster
string_scan1.623.153.172.713.161.67× faster
mem_copy0.531.051.021.420.971.82× faster
array_sum0.371.441.092.841.093.0× faster
array_map0.421.571.472.032.063.5× faster
call_depth0.260.950.950.960.973.6× faster
matrix_mul_f320.226.286.261.958.768.7× faster

The fastest results are where the runtime does the work: string_scan, method_call and struct_copy are byte scanning, dynamic dispatch and aggregate copies. The slowest show where xcc’s code generation has most to gain:

  • Vectorisation. On arm64, int_muldiv and float_math are vectorised by both compilers, and clang’s loops are tighter. (matrix_mul reached clang’s speed in 0.73.)
  • Reference counting on x86-64. arc_array, where C++ reads elements without retaining them.
  • x86-64 loops. sieve, sort_small and struct_copy, where the other compilers’ loops are faster.

A dense matrix multiply written as three plain loops, C[i][j] the sum over k of A[i][k] · B[k][j] in float or double, runs as a kernel the compiler writes itself (from 0.71). The results are the loops’, to the last bit, and -fno-matmul turns it off. The kernel depends on the machine:

  • Apple M4 and later (SME): the multiply runs on the matrix unit. From 0.72 it works on a 2×2 block of the unit’s tiles at once, and under :goal(speed), which matrix_mul_f32 uses, it leaves out a NaN check that only affects the bits of a NaN (the inputs have none).
  • x86-64: the widest vector unit the processor has (SSE2, AVX2 or AVX-512), chosen when the program starts; the Ryzen 9 9955HX picks AVX-512.

matrix_mul_f32 measures it: a 128×128 float multiply, 10,000 times. (matrix_mul multiplies u32 values, which the matrix unit does not take, so it stays a vectorised loop.)

matrix_mul_f32xcObjective-CC++ (clang)C++ (GCC)Swift
arm64, Apple M4 Max30 ms9.8 s9.8 s–9.7 s
xc is325× faster325× faster–322× faster
x86-64, AMD Ryzen 9 9955HX225 ms6.3 s6.3 s2.0 s8.8 s
xc is28× faster28× faster8.7× faster39× faster

A par block runs a loop’s iterations at once, across every CPU thread or on the GPU (from 0.7). The GPU is reached through:

  • Metal on Apple silicon;
  • CUDA on an NVIDIA GPU under Windows;
  • Vulkan on Linux, Windows and Android (from 0.72);
  • WebGPU in a browser (from 0.72).

Four programs in benchmark/par measure it:

  • mandelbrot: a 2048×2048 escape-time image, at most 256 iterations a pixel;
  • perlin: 2048×2048 improved noise, four octaves;
  • nbody: the force on each of 8192 bodies from all the others;
  • saxpy: y = a·x + y over 16 million integers, with a sum.

Each runs its block eight times and reports the best run, in four modes: one thread (XC_PAR=cpu XC_PAR_THREADS=1), all threads (XC_PAR=cpu), GPU (XC_PAR=gpu) and auto, the default, which times the CPU and the GPU and keeps the faster (from 0.73 it remembers the choice, per program and per machine, so later runs do not measure again). Every mode’s checksum must agree. On Windows GPU is CUDA and Vulkan the same card through Vulkan (XC_PAR_GPU=vulkan).

Apple MacBook Pro, M4 Max (12 performance and 4 efficiency cores), its GPU through Metal (ms; best of eight runs)

benchmarkone threadall threadsGPUautoGPU vs all threadsGPU’s first run
mandelbrot4811001.62.165× faster22.5
nbody67.76.90.70.99.9× faster22.6
perlin32731.31.50.920× faster23.4
saxpy7.51.012.51.013.2× slower33.4

AMD Ryzen 9 9955HX (16 cores, 32 threads), Linux, its integrated Radeon 610M (2 compute units, 128 shader lanes) through Vulkan (ms; best of eight runs)

benchmarkone threadall threadsGPUautoGPU vs all threadsGPU’s first run
mandelbrot39840.320.019.22.0× faster40.1
nbody95.17.13.63.61.97× faster23.3
perlin37425.87.17.03.6× faster30.8
saxpy12.82.52182.586× slower253

AMD Ryzen 9 5950X (16 cores, 32 threads), Windows, an NVIDIA RTX 3090 through CUDA and Vulkan (ms; best of eight runs)

benchmarkone threadall threadsGPUVulkanautoGPU vs all threadsGPU’s first runVulkan’s first run
mandelbrot45546.71.61.11.630× faster202122
nbody1208.81.90.91.84.7× faster207124
perlin49829.32.02.02.014.6× faster220128
saxpy13.55.639.944.85.47.2× slower258178

Chrome on the Apple M4 Max: wasm32 (one thread) and WebGPU (from 0.73) (ms; best of eight runs)

benchmarkone threadall threadsGPUautoGPU vs one threadGPU’s first run
mandelbrot364–2.61.9140× faster12.1
nbody53.9–3.23.216.8× faster7.3
perlin149–2.62.557× faster8.6
saxpy15.7–22.016.31.40× slower56.1

What the tables show:

  • A discrete GPU wins by a wide margin on the three programs that do real work per element, and auto finds that and uses it.
  • saxpy stays on the CPU. It does one multiply-add for every eight bytes it moves, so copying to and from the GPU outweighs the arithmetic, and auto keeps it on the CPU everywhere.
  • The Ryzen’s integrated Radeon has 128 shader lanes against the CPU’s 32 threads. It just beats them on mandelbrot and nbody and loses on the other two; its copies of saxpy’s arrays are slower still, because the runtime does not yet use the CPU’s cache for them there.
  • Vulkan against CUDA: on the RTX 3090, Vulkan is faster on mandelbrot (1.1 ms against 1.6) and nbody (0.9 ms against 1.9), level on perlin and slower on saxpy (44.8 ms against 39.9). Up to 0.74 auto used CUDA on an NVIDIA GPU; from 0.75 it measures both and keeps the faster for each block.
  • The first run carries one-off costs (building the kernel, and on NVIDIA creating the driver context and compiling the PTX or SPIR-V), so it is shown apart: a few tens of milliseconds on Metal and the integrated GPU, and on the 3090 202 to 258 ms through CUDA and 122 to 178 ms through Vulkan. auto pays it once, while it measures.

How fast xc’s code is in each release, relative to 0.62: above 1× is faster. The thin lines are single benchmarks (labelled where they moved by more than twelve percent), the thick line the geometric mean (without matrix_mul_f32). Releases shown: 0.62, 0.63, 0.64, 0.65, 0.66, 0.7, 0.71, 0.72, 0.73, 0.74. arc_alloc, method_call changed in 0.64’s benchmark set and are left out of this history.

arm64: xc's speed relative to 0.620.80×0.90×1.00×1.10×1.25×1.50×2.00×3.00×0.620.630.640.650.660.70.710.720.730.74arc_array: 1.00, 1.00, 1.03, 1.86, 1.85, 1.77, 1.79, 1.83, 1.83, 1.85array_map: 1.00, 1.02, 1.15, 1.59, 1.62, 1.61, 1.52, 1.63, 1.50, 1.61array_sum: 1.00, 1.01, 1.02, 1.04, 1.03, 1.01, 1.02, 1.03, 1.17, 1.18bit_ops: 1.00, 1.01, 1.03, 1.27, 1.26, 1.17, 1.24, 1.27, 1.26, 1.28branch_mix: 1.00, 1.01, 1.03, 1.25, 1.22, 1.24, 1.21, 1.23, 1.23, 1.23call_depth: 1.00, 1.01, 1.03, 1.11, 1.10, 1.02, 1.08, 1.09, 1.08, 1.08float_math: 1.00, 0.95, 0.96, 1.30, 1.32, 1.28, 1.31, 1.30, 1.30, 1.30hash_mix: 1.00, 0.93, 0.96, 1.09, 1.10, 1.10, 1.09, 1.10, 1.10, 1.09int_accum: 1.00, 0.99, 1.03, 1.26, 1.28, 1.27, 1.27, 1.28, 1.27, 1.28int_muldiv: 1.00, 1.00, 1.02, 1.08, 1.07, 1.05, 1.05, 1.07, 1.07, 1.08matrix_mul: 1.00, 1.00, 1.03, 4.76, 4.75, 4.73, 4.70, 4.69, 8.12, 8.29matrix_mul_f32: 1.00, 1.01, 1.01, 1.72, 1.70, 1.66, 147.84, 351.04, 349.66, 350.32mem_copy: 1.00, 1.01, 1.05, 1.79, 1.79, 1.79, 1.77, 1.80, 1.78, 1.80poly_dispatch: 1.00, 0.99, 1.02, 2.13, 2.15, 2.12, 2.11, 2.14, 2.56, 2.61sieve: 1.00, 1.00, 1.03, 1.12, 1.14, 1.13, 1.13, 1.05, 1.13, 1.14sort_small: 1.00, 1.00, 1.01, 1.20, 1.21, 1.16, 1.14, 1.15, 1.26, 1.31string_scan: 1.00, 0.99, 1.05, 1.08, 1.09, 1.08, 1.06, 1.07, 1.06, 1.09struct_copy: 1.00, 1.00, 1.07, 1.93, 1.94, 1.91, 1.88, 1.91, 1.90, 1.931.000.991.031.451.451.421.421.431.511.53matrix_mul_f32 350.32×matrix_mul 8.29×poly_dispatch 2.61×struct_copy 1.93×arc_array 1.85×mem_copy 1.80×array_map 1.61×sort_small 1.31×float_math 1.30×bit_ops 1.28×int_accum 1.28×branch_mix 1.23×array_sum 1.18×sieve 1.14×geometric mean
arm64: xc's speed relative to 0.62
x86-64: xc's speed relative to 0.620.80×0.90×1.00×1.10×1.25×1.50×2.00×3.00×0.620.630.640.650.660.70.710.720.730.74arc_array: 1.00, 0.98, 1.01, 0.99, 0.93, 0.94, 0.94, 1.00, 1.01, 0.99array_map: 1.00, 0.99, 1.36, 1.37, 2.76, 2.76, 4.98, 4.93, 5.01, 5.01array_sum: 1.00, 1.00, 1.29, 1.32, 2.41, 2.56, 3.60, 3.21, 5.01, 3.27bit_ops: 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00branch_mix: 1.00, 0.89, 1.01, 1.01, 1.01, 1.01, 1.00, 1.00, 1.00, 1.00call_depth: 1.00, 1.00, 1.00, 1.00, 0.71, 2.03, 4.05, 4.03, 4.03, 4.04float_math: 1.00, 1.00, 2.03, 1.94, 2.00, 2.00, 1.93, 1.93, 2.00, 2.00hash_mix: 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00int_accum: 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00int_muldiv: 1.00, 1.00, 1.01, 1.00, 1.98, 2.02, 3.98, 3.98, 3.96, 3.95matrix_mul: 1.00, 1.00, 1.87, 4.31, 4.30, 4.29, 4.28, 4.28, 3.83, 3.81matrix_mul_f32: 1.00, 1.05, 1.28, 1.44, 1.66, 1.63, 55.40, 70.09, 69.85, 70.84mem_copy: 1.00, 1.00, 1.43, 1.43, 2.21, 2.73, 4.12, 4.33, 4.32, 2.88poly_dispatch: 1.00, 1.00, 1.00, 1.01, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00sieve: 1.00, 1.00, 1.24, 1.50, 1.34, 1.36, 1.50, 1.33, 1.66, 1.64sort_small: 1.00, 1.21, 1.25, 1.25, 1.31, 1.27, 1.21, 1.22, 1.18, 1.17string_scan: 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00struct_copy: 1.00, 1.00, 1.32, 1.30, 1.32, 1.31, 1.29, 1.29, 1.29, 1.291.001.001.191.261.421.531.801.791.851.75matrix_mul_f32 70.84×array_map 5.01×call_depth 4.04×int_muldiv 3.95×matrix_mul 3.81×array_sum 3.27×mem_copy 2.88×float_math 2.00×sieve 1.64×struct_copy 1.29×sort_small 1.17×geometric mean
x86-64: xc's speed relative to 0.62

The compiler that ships. The xc numbers come from the xcc in the download.

The same work in every language. Every version keeps its data where the xc version keeps it (local arrays, not static ones, which clang optimises differently) and leaves nothing a compiler can remove: an object that could be put on the stack outlives its iteration, a call whose target could be resolved at compile time takes its class at run time, and results are folded in so no loop has a closed form. Where the original’s point is reference-counted objects, the C++ version uses std::shared_ptr and the Swift version a class, so they pay for reference counting too.

Different runtimes on the two targets. The Objective-C column is Apple’s Foundation on arm64 and GNUstep with libobjc2 on x86-64, and Swift is 6.2 on the Mac and 6.1 on Linux. These are different implementations, so a language’s times compare within a target and not across one. Swift on Linux is markedly slower on poly_dispatch and array_map than on the Mac; the runs were repeated on an idle machine and reproduce.

Timed regions of about one second. Each benchmark times its own inner loop rather than the process, so start-up and data set-up are excluded.

Alignment noise on x86-64. xcc aligns every loop head on x86-64 to a 32-byte boundary and the start of .text to 64 bytes, so an unrelated change elsewhere cannot move a loop across a fetch boundary; see Optimisation.

The sources are in benchmark/src, one .xc, .m, .cpp and .swift per program, and the runner builds and times them all:

python3 benchmark/run.py --opt O3 --repeats 5
python3 benchmark/page.py

The x86-64 legs cross-build xc here and build the other languages on the configured Linux host; without one the runner measures this machine only. --langs measures some of the languages and adds them to the results already there. Results land in benchmark/<version>/results.json, and page.py turns them into this page.