ChangeLog
Version 0.74 — debugging
Section titled “Version 0.74 — debugging”xcc -g writes DWARF debug information into the executable, so lldb and gdb
can stop on a source line, step, show a backtrace with each frame’s file and
line and arguments, and print a function’s variables.
-gonarm64(macOS),x86_64(Linux, dynamic and-static) andwin64: a line table for the program and the library files it imports, a function entry for every function and method (a breakpoint on a name stops at its first statement), call frames (a backtrace is right in every frame, leaf functions included), and the integer, floating-point,booland pointer variables and parameters of each function, soframe variable,print xandbtshowingadd(a=2, b=3)work.-gkeeps those variables in memory for the whole function, as C compilers do at-O0; without-gthe output is unchanged. See Debugging.- Windows executables have a symbol table, so a debugger or a crash report
names their functions with or without
-g.
Version 0.73 — faster arm64, GPU blocks that remember, and libraries on Linux
Section titled “Version 0.73 — faster arm64, GPU blocks that remember, and libraries on Linux”arm64 code is faster again, reversing the drift since 0.65. A par block keeps
what auto learns about the CPU and the GPU, so a machine measures once. On
Vulkan and WebGPU, blocks may divide 64-bit integers, reduce 8- and 16-bit
values, and divide and take square roots exactly. A library built for Linux
links against the system’s C library and names the libraries it needs, and on
Windows #import reads a MinGW-built DLL’s own debug information.
parremembers: whatautomeasures is kept in the program’s own settings store, as a size threshold for each block: a run outside the learned bounds is decided without measuring. The keys include a hash of the GPU and CPU, so a new GPU is measured afresh, and of the block’s GPU version. The user’s own settings (par.<block> = cpu,gpu,autoor a number of items from which the GPU is used) come before them. See Parallel blocks.:goal(accuracy)on Vulkan and WebGPU: a block whose goal is accuracy divides and takes square roots offloats in integer arithmetic, correctly rounded, so its results are the CPU’s; up to 0.72 it ran on the CPU.- More on WebGPU: 64-bit integer division and remainder, and conversions
between 64-bit integers and
float. On Vulkan and WebGPU, reductions of 8- and 16-bit values andbools. WebGPU keeps a block’s buffers between runs. - Linux libraries:
--emit-libon-A x86_64builds a library linked against glibc, as executables are: each-llibrary becomes one of its dependencies, and it exports only its own API.-statickeeps the musl library. #import <clib>on Windows: a DLL built by a MinGW toolchain with-gis read for its functions, types and enum constants, as on Linux and macOS.- iOS and Android libraries: the third-party tree is searched in the
target’s own directory (
ios,ios-sim,android) before the architecture’s;#import <UIKit>finds the iOS SDK withoutSDKROOT;--with-libmay be given more than once.
Faster
Section titled “Faster”- arm64:
matrix_multakes 225 ms instead of 389 (1.73× faster; clang’s C++ takes 224),poly_dispatchis 1.20× faster,array_sum1.13× andsort_small1.10×. A dense integer matrix multiply now loads four elements of a row at once and multiplies by lanes, and a 32-bit copy feeding a call or a loop is a 64-bit one, which costs nothing on Apple’s cores. - The benchmark suite runs 1.04× faster than 0.72 on arm64 and 1.08× faster on
x86-64, where
arc_allocis 2.6× faster,array_sum1.6× andsieve1.25× (matrix_mulis 11% slower). xc’s code is 1.45× faster than C++ on arm64 and 1.60× faster on x86-64 (geometric mean); withoutmatrix_mul_f32, 1.09× and 1.38× faster. See Performance.
- On x86-64 and Windows, converting a
floattou64in an unrolled loop could overwrite another value the loop was carrying, giving a wrong result at-O2. - A WebGPU block with more than eight arrays returned zeros; it now runs where the device allows, and on the CPU where it does not.
- On Vulkan, an 8- or 16-bit or
boolfield of a block written by its body overwrote its neighbours’ bytes.
Version 0.72 — the GPU everywhere, and Foundation grows
Section titled “Version 0.72 — the GPU everywhere, and Foundation grows”par blocks now reach the GPU on every desktop and mobile target and in the
browser: Vulkan on Linux, Windows and Android, and WebGPU on the web, beside
0.7’s Metal and CUDA. A block can cover a 2-D or 3-D grid. A for loop says
whether it wants speed or exactness. Linux programs link dynamically against
the system’s C library by default, as C programs do, and #import of a C
library finds its debug information in a separate debug file. Foundation gains
twenty-one classes, many of them moved from UXKit.
paron Vulkan and WebGPU: onx86_64Linux,win64andandroideach block also gets a Vulkan version, and onwasm32a WebGPU one, run in the browser while the program waits (JSPI). On Windows, CUDA is tried first on an NVIDIA GPU and Vulkan on any other;XC_PAR_GPU=vulkan|cudachooses. 8- and 16-bit values andbools are exact on every Vulkan and WebGPU device. See Parallel blocks.par :grid(w, h[, d]): a block over the points of a grid, reading its point aspar.x,par.yandpar.z, on the CPU and every GPU.:goalon aforloop::goal(speed)lets the compiler drop a loop’s exactness checks (a matrix multiply then skips its NaN check); plain loops keep:goal(accuracy), while aparblock’s loops follow its own goal.- Linux links dynamically:
-A x86_64executables link against glibc by default and can load any system library;-staticgives the self-contained musl executable earlier releases made. #import <clib>reads a stripped library’s types and functions from its separate debug file (by build ID or.gnu_debuglink, under/usr/lib/debugor$XCC_DEBUG_DIR), compressed DWARF included, and says so when a library has no debug information at all.- Foundation:
Bag,Range,BinaryHeap,Cache,Null,JSON,Expression,NumberFormatter,NotificationCenter,UndoManager,Progress,StateMachine,SearchIndex,IndexSet,AttributedString,Regex,Predicate,Validator,SortDescriptor,SocketandCSV; log levels, subsystems and monitors inLog; file URLs inUrl; path components inString; moreCharacterSets;Number.withBool. See Foundation. - UXKit:
UXTextView, an editable rich-text view: the platform’s own editor on macOS, iOS, Android, Windows, GTK and the web, and drawn and edited by the view elsewhere. UXKit’s copies of the classes Foundation now has are gone.
Faster
Section titled “Faster”- The SME matrix kernel works on a 2×2 block of tiles, and under
:goal(speed)skips its NaN check:matrix_mul_f32takes 30 ms on an Apple M4 Max, 2.4× faster than in 0.71 and 327× faster than clang’s C++. - The benchmark suite runs 1.05× faster than 0.71 on arm64 and 1.01× faster on
x86-64, where it is now measured as shipped, linked dynamically. xc’s code
is 1.37× faster than C++ on arm64 and 1.47× faster on x86-64 (geometric
mean); without
matrix_mul_f32, 1.03× and 1.27× faster. - GPU kernels for Vulkan are structured loops and ifs, so a SIMD group’s
threads reconverge, and on a discrete GPU they work in its own memory: on an
RTX 3090, Vulkan comes within 1.2–1.7× of CUDA (
mandelbrot2.2 ms against 1.9). See Performance.
- A variable assigned in a
tryblock kept its old value after acatchthat returned. #importof a C library: 8-byte integers in its DWARF were read as 32-bit, and a program using one in another directory could not find it at run time.- A 64-bit shift by a 32-bit count printed a CUDA kernel NVIDIA’s driver refused, so the block ran on the CPU.
Version 0.71 — AVX-512, and matrix multiplies on the matrix unit
Section titled “Version 0.71 — AVX-512, and matrix multiplies on the matrix unit”A dense matrix multiply written as three loops now runs as a matrix kernel the compiler generates itself: on Apple silicon with SME (M4 and later) on the matrix unit, and on x86-64 with SSE2, AVX2 or AVX-512 — with the results of the loops, bit for bit. x86-64 and Windows gain AVX-512, as a flag and as a third level the program picks at load.
- AVX-512:
-mavx512(or-msimd=avx512) vectorises x86-64 and Windows code with 512-bit registers (F, DQ, BW and VL). Under-msimd=auto, the default, each vectorised function is also built for AVX-512 and the program picks it at load on a machine that has it;XC_SIMD=avx512forces it and-mnativeselects it where the build machine supports it. - Matrix kernels: at
-O2and above,C[i][j] = Σ A[i][k] · B[k][j]overfloatordouble, written as three loops, runs on the SME matrix unit on Apple silicon macOS and as a vector kernel on x86-64 Linux. The results are the loops’ to the last bit, NaNs and signed zeros included; on a Mac without SME, or when the matrices overlap or hold a NaN, the loops run as written.-fno-matmulturns it off. - A class that lists a protocol and leaves out one of its required methods is now refused, naming the class, the protocol and the method (it built, and crashed calling the missing method).
Faster
Section titled “Faster”matrix_mul_f32(a 128×128floatmultiply, 400 times): 3.5 ms on an Apple M4 Max against 391 ms for clang’s C++, and 12.5 ms on an AMD Zen 5 processor with AVX-512 against 252 ms.- The benchmark suite, now twenty programs with
matrix_mul_f32(measured back to 0.62 for the history chart), runs 1.23× faster than 0.7 on arm64 and 1.37× faster on x86-64, where AVX-512 also makesarray_map,call_depthandint_muldivtwice as fast. xc’s code is 1.30× faster than C++ on arm64 and 1.47× faster on x86-64 (geometric mean); withoutmatrix_mul_f32, 1.03× and 1.28× faster. - Compiling long x86-64 functions:
vectorize_wide_tailsbuilds in 20 s instead of 45 (copy propagation and loop-invariant motion no longer scale with the square of a function’s length).
Version 0.7 — parallel blocks, on every CPU thread and on the GPU
Section titled “Version 0.7 — parallel blocks, on every CPU thread and on the GPU”A par block runs a loop’s iterations at once: over every CPU thread, or on
the GPU — Metal on Apple silicon, and NVIDIA GPUs under Windows through the
driver alone, with no CUDA toolkit. The compiler holds every block to what a GPU
can run, on every target, so the same source runs anywhere; by default each
block measures the CPU and the GPU and keeps the faster. Foundation gains
Operation and OperationQueue. See Parallel blocks
and the new GPU section of Performance.
parblocks:par name { for (T i in a..b) { … } }runs independent iterations across every CPU thread;:reduce(op var)folds results, and an integer reduction matches a serial run exactly.- On the GPU: Metal on Apple silicon (macOS), and NVIDIA on Windows, where the program hands the driver PTX and needs no toolkit. Global and captured arrays become GPU buffers; helper functions of plain values run there too.
- The device is chosen per block (
auto, the default): the CPU is timed first, a block under 1 ms stays there, otherwise the GPU is timed too and the faster kept, each after a warm-up run.XC_PAR=cpu|gpu|auto,Par.device(name, choice)andXC_PAR_REPORT=1override and report it. :goal(speed)(the default) lets the GPU use its fast maths;:goal(accuracy)keeps precisesin,expand friends.- The
par-gpuwarning says why a block cannot run on the target’s GPU, and what to change; a block whose items read each other’s results is an error. OperationandOperationQueue: units of work with dependencies across queues, priorities, cancellation and completion blocks, run on worker threads, serially, or on the main run loop; on targets without threads a queue runs its operations when the program waits for them.--manifest-attr name=valuesets an attribute on an APK’s<application>.sizeofis the target’s size type:u64on 64-bit targets,u32on 32-bit ones,u16on the 6502, sosizeofof a 1 MB buffer is 1048576.- A loop variable can be captured by a block or a
parblock inside the loop.
Faster
Section titled “Faster”- A
parblock on the GPU:mandelbrotover 2048×2048 pixels takes 2.0 ms on Apple silicon’s GPU against 99 ms on every CPU thread and 482 ms on one, and 1.9 ms on an RTX 3090. A memory-bound block such assaxpystays on the CPU, whereautofinds it is faster. - x86-64 and Windows: an unrolled AVX2 loop no longer chains its iterations
through the integer splats, so
call_depthruns 2.9× faster than in 0.66 and the x86-64 benchmark suite 1.08× faster. xc’s code is now 1.11× faster than C++ on x86-64 (geometric mean), and 1.04× faster on arm64.
Wrong code fixed
Section titled “Wrong code fixed”- A static method called on a class that also has instances no longer runs
the class’s
initon its static storage (it crashed Windows programs and brokeOperationQueueon the 6502). - x86-64 and Windows, at
-O3: a checked downcast past an object’s own class could return null after loop rotation. - AVX2: an unrolled loop’s integer splats no longer chain one iteration into the next.
- wasm32: an indirect call that returns a struct declares its return slot.
- arm9: a DWARF import takes the declaration that has the parameters.
Errors that used to be silent
Section titled “Errors that used to be silent”- A protocol cast the 6502 or m68k cannot check at run time is an error unless the class declares the protocol.
sizeofa value too large for the target’s size type is an error.
Linking
Section titled “Linking”- A program links with
-cobjects passed through-Wl,orXTC_LDFLAGS. - Windows: a program calling POSIX names such as
closeorchdirloads, through the C runtime’s own exports.
Version 0.66 — AVX2 wherever it runs, and Linux GUI programs
Section titled “Version 0.66 — AVX2 wherever it runs, and Linux GUI programs”x86-64 and Windows programs now use AVX2 on machines that have it and SSE2 on
machines that don’t, from one binary, picked when the program starts: on a
Zen 5 machine an elementwise map and a sum take about half the time they did.
-A x86_64 -dynamic links against glibc, so a Linux program can load GTK 4,
libGL and other shared system libraries, still with xcc’s own linker.
- Runtime SIMD dispatch on
-A x86_64and-A win64, the default (-msimd=auto): each function the vectoriser widens is built for SSE2 and for AVX2, and the program picks one per function at startup from what the CPU and the operating system support.XC_SIMD=baseorXC_SIMD=avx2forces a level for one run; forcing one the machine lacks falls back, with a note on stderr.-cobjects and libraries keep one version. -mavx2(or-msimd=avx2) builds AVX2 code only,-msimd=baseSSE2 only, and-mnativethe level of the machine runningxcc.-dynamicon-A x86_64: a dynamically linked glibc executable that can load shared system libraries.-l<name>takeslib<name>.sofrom-Land the system library directories; a Mac can link for Linux with-Lto copies of the libraries. A symbol that neither glibc nor an-llibrary defines is a link error.libUXGtk.so, UXKit’s GTK 4 back end as a shared library, is a separate download for Linux GUI programs linked with-dynamic.- A library’s structs, enums and typedefs import into its clients.
defertakes a block.Bundle.main()finds the program’s executable throughPATHwhen it was started by name.
Faster
Section titled “Faster”- AVX2 code (dispatched, or with
-mavx2): 256-bit vector loops, withvzeroupperaround calls. On a Zen 5 machinearray_mapruns 2.0× andarray_sum1.9× faster than with SSE2. - The compiler holds less memory while it builds a large program.
Wrong code fixed
Section titled “Wrong code fixed”- arm64, in a program that starts a thread: a
deallocthat called a method onselfcould run again from inside itself until the stack overflowed. - An array ivar of class pointers, and an object local whose address is taken, now own what they hold.
- x86-64 and Windows:
expandpoware accurate to the last bit. - A single-precision denormal constant keeps its value.
printfof a typed collection’s element prints its value.- Windows: a frame larger than a page is probed one page at a time; a
floatpassed through...reaches the callee; a weak reference is unlinked only through a valid back-pointer. - wasm32: objects use the 40-byte header and 32-bit reference count the other targets use.
Filesreads/procand/sysfiles, which report no size.
Errors that used to be silent
Section titled “Errors that used to be silent”- macOS: calling a C function that no linked library exports is a link error.
- A static
-A x86_64link refuses an-lit cannot use (only a shared library, or nothing) instead of dropping it. - A typed collection of primitives without
Numberis an error, and so is a checked cast from a raw pointer. - A construct the lowering cannot handle is named, not numbered.
Install
Section titled “Install”make install PREFIX=<dir>keeps the third-party tree beside the prefix instead of writing/opt/xcc/3p.
Version 0.65 — faster loops, native settings and system frameworks
Section titled “Version 0.65 — faster loops, native settings and system frameworks”The optimiser and the arm64 back end close the distance to clang: across the
benchmark suite arm64 code is now 1.03× faster than clang’s C++ (geometric
mean), from 1.24× slower, and 1.38× faster than 0.64; matrix multiply is 4.7×
faster. x86-64 code is 1.08× slower than C++, from 1.13× slower. See
Performance.
Settings.standard() keeps its values in the platform’s own settings store, and
#import <Framework> links a macOS or iOS system framework.
Settings.standard(name)uses the platform’s own store: the user’s preferences on macOS and iOS, the registry on Windows,localStoragein a browser, and a text file under$XDG_CONFIG_HOME(or~/.config) elsewhere. The API is unchanged. A 0.64 settings file is read the first time and the nextsavewrites the new location.#import <Framework>links a macOS or iOS system framework, as-frameworkdoes.- The wasm32 loader passes the mouse wheel and the right button to the program.
Faster
Section titled “Faster”- A loop around a reduction loop is vectorised across the outer loop: four neighbouring outer iterations are four lanes, which is the shape of a matrix multiply.
- Pointers step through strided and offset array walks (
a[i*n + k]) instead of recomputing the index. - An integer reduction in an unrolled loop keeps one operation per iteration on the carried value; the rest run in parallel.
- More loops are rotated so the test sits at the bottom, and a value computed in a loop’s condition is not computed again in its body.
- arm64: a leaf function that never touches its frame has none; parameters and
call results stay in the registers they arrive in;
madd,msubandmlareplace multiply-then-add; local arrays and functions start on 16-byte boundaries. - x86-64: local arrays start on 16-byte boundaries.
Wrong code fixed
Section titled “Wrong code fixed”- arm64: a value live across a call could share an argument register with a loop-carried value at -O2 and above.
Version 0.64 — C’s printf, an HTTP client and a run loop
Section titled “Version 0.64 — C’s printf, an HTTP client and a run loop”printf and the other format functions now follow C, keeping %@ for objects.
The library gains settings files, bundles, an HTTP client, file operations on a
background thread and a run loop to deliver their results. The x86-64 back end
is faster, and a number of wrong-code bugs are fixed.
Changed
Section titled “Changed”Stdio.printf,String.withFormat,String.appendFormatandLog.error,Log.warningandLog.infotake C’s conversions, flags, field widths, precisions and*, with C’s output:%xhas no leading zeros,%fand%lfboth print six places,%eand%gare C’s, and the floating conversions round exactly as C does.%@still prints an object’sdescription(), and now an enum’s name too.- The arguments follow C: an integer narrower than
intis passed as anintand afloatas adouble.intis 32 bits, 16 on xt6502;longis 64 bits on the 64-bit targets. - With a literal format the compiler sizes each integer conversion to the
argument actually passed, so
%dprints ani64whole; the check now reports only an argument of the wrong kind or a wrong count. A variadic function that passes its format and...on to one of these is treated the same way at its own call sites.
Settings(a key/value file) andBundle(a program’s resources).Http: an HTTP/1.1 client, with TLS where the platform provides it, and the native transport forurl.fetch.AsyncFiles: theFilesoperations on a background thread, andRunLoop, which the completions can be delivered to.- Third-party libraries resolve and load with the compiler.
- A float literal may have an exponent without a point (
1e30), and thef/Fsuffix marks a single-precision literal. PLATFORM_androidis defined for-A android.- New warnings: a non-void function that can reach its closing brace, and a local that hides a field.
Faster
Section titled “Faster”- x86-64: values are loaded straight into their home registers, registers are
ranked by loop-weighted use and shared along live segments, a float phi copy
is one
movaps, and loops rotate so the test sits at the bottom. Across the benchmark suite x86-64 code is now 1.14× faster than clang’s (geometric mean), from 1.03× slower; see Performance.
Wrong code fixed
Section titled “Wrong code fixed”- A
%@argument toString.withFormatorappendFormatwas released once too often, so a local object could be freed while still in use. - A method called through the vtable with variadic arguments got its tail on arm64.
- A scalar argument converts to a
floatparameter; a local that shadows a field writes the local;sizeof(x)on a variable measures the variable; a string literal in a global aggregate initialiser is emitted; an implicit-self call boxes its arguments; an array bound folds in a struct field and a prefix type. - A
doubleliteral below the smallest normal value is kept rather than flushed to zero.
Version 0.63 — archiving, class names and Windows DLLs
Section titled “Version 0.63 — archiving, class names and Windows DLLs”This release adds keyed archiving to the language, gives every object a runtime class name, and lets win64 build a DLL in-house. It also fixes wrong code across the back ends, and arm9 gains load-time constructors and can build a program from objects.
CodableandCoder: keyed archiving. A class that adoptsCodablegainsencodeWith(coder)andinitWith(coder); aCoderwrites a keyed archive as JSON, optionally gzipped.ObjectadoptsCodable.Object.className()returns an object’s runtime class name, andObject.newInstanceOfClass()makes another instance of it. Class and block names are carried into the image.- Windows:
xcc -A win64 --emit-libbuilds a DLL in-house, with an entry point, imports and an export table. A DLL’s interface is read back, and a module dispatches protocols through the itable under the Microsoft ABI. - arm9: a load-time constructor in an object runs, and both compilers’
-cobjects and IR are the same, so a program can be built from objects and linked in-house. xcc-sim-6502andxcc-sim-68kreport their own name and the version they were built from.
Wrong code fixed
Section titled “Wrong code fixed”- Floating point: a comparison follows IEEE where either operand is a NaN; a negated float flips its sign bit instead of subtracting from zero; a negated integer literal wider than its type is built at the wider type; and an integer initialiser for a float global is folded into the image rather than built at run time.
- Conversions between floating point and unsigned or 64-bit integers are
complete on x86-64, win64, m68k and arm9. On xt6502 a 64-bit integer converts
straight to
floatand rounds once. - x86-64 and win64 no longer rename a function that is named after a register, which broke a call to one.
- A 64-bit select copies both words on m68k and arm9, and a 64-bit phi stays out of a register home on those targets.
- wasm32: a local the optimiser pinned and then removed is given no slot.
- x86-64: a weak undefined symbol in a dynamic link keeps its value 0.
- arm32: both assemblers write one literal-pool word per distinct expression, as
asdoes, so the linked images agree. - xt6502:
:mainand:bankedplace free functions and methods; every new block is sized to its contents with its header in a trailer; and a variadic reentrance error no longer names the 6502 packing-buffer address. - arm9 and android place the extra arguments of a variadic call where the callee
reads them. On android a library can also be
#imported. -fltokeeps an imported vtable and each object’s string literals.- The shipped compiler refuses a class whose parent is not a class, an
unreachable function end is
Unreachable, an open slice start is a zero of the count’s width, and preprocessor diagnostics carry their location. - The reference compiler catches up on several points of its own: a
floatresult returned through an indirect call comes froms0/d0; an autoboxed argument is not a raw pointer for a class parameter; an array ivar of an inline class-array element is at its address; an enum ivar counts as an integer ivar; a library interface does not re-export a protocol it imported; an x86-64 dynamic link exports what the shipped linker exports; and-Son arm9 writes the assembly a link would assemble.
Version 0.62 — wrong code fixed, shared libraries that work together
Section titled “Version 0.62 — wrong code fixed, shared libraries that work together”Most of this release fixes wrong code and makes xc libraries work with each other: a library can now import another library and be used from a program on arm64, x86_64 and wasm32. Several mistakes that used to compile are now errors. Libraries built by an earlier release must be rebuilt, and on x86_64 programs that use them must be rebuilt too.
New errors
Section titled “New errors”- Every call checks its argument count: static, instance, inherited, category, protocol and free-function calls, library imports, and calls through function pointers, callbacks and blocks. A variadic call must pass at least the fixed arguments.
- A call argument whose class is unrelated to the parameter’s (neither a
subclass nor an ancestor) is refused, for methods and functions alike.
Passing an ancestor, such as the
Object*a collection returns, where a subclass is declared is still accepted. A class reference still does not convert to a rawpointerimplicitly; cast it. deleteon a class instance is refused. ARC owns the object; freeing it by hand freed it twice.deleteon a struct or primitive array is unchanged.- On xt6502, a call to a function that is declared but never defined is an error naming the function and the call’s position. It used to become a jump to address 0.
- m68k refuses inline assembly it cannot compile. It used to drop the block.
- Errors in a call now point at the name being called, not at the closing
).
Wrong code fixed
Section titled “Wrong code fixed”- A virtual call to an overloaded method could run another overload’s body. Each overload now dispatches through its own slot.
- An overload that differs only by return type, passed directly as an
argument (
Math.ln(Math.E())), takes the parameter’s type. - A call through a callback or function pointer now converts each argument to the parameter’s type, as a direct call does. On xt6502 the callee read garbage; on arm64 a negative narrow integer passed to a 64-bit parameter arrived as a large positive number.
- A condition is tested on its whole value on every target. arm64, x86_64 and
win64 tested only the low 32 bits of a 64-bit value or pointer, arm9 and m68k
one half of a 64-bit value, and xt6502 the low byte of a pointer. The right
side of
&&and||was reduced to its low byte. A floating-point condition was tested by its bits, so-0.0counted as true. !pon a pointer tests the whole pointer. On arm64 a pointer on a 64 KB boundary read as null.- An integer converted to a pointer keeps the full pointer width. arm64 kept only 16 bits, so an address round-tripped through an integer crashed.
- On arm64 and x86_64 at
-O2and above, the first call to a static method of a class from a library could crash. - xt6502 at
-O3could lose a struct parameter across a call when only the addresses of its fields were still in use. - On arm9, m68k, wasm32 and xt6502 the 16-bit retain count now stops at
$FFFFin both directions. It used to wrap, freeing an object that was still in use. An object that reaches the limit is never freed. - Freeing
new C[0]ran one element’sdealloc, and an empty array created at run time reported a.lengthof 1. - xt6502 converts between floating point and 64-bit integers correctly.
- m68k converts a floating-point value to a narrow integer correctly: out of
range gives 0, and values from 2^31 to 2^32 reach
u32. A pointer converted to a 64-bit integer fills both halves. - On arm64, a call to a cloaked or banked variadic function put its extra arguments in the wrong place.
- A struct passed to a variadic function had all its fields written to its first byte.
- A static method could fill a protocol’s method slot, so a call through the protocol passed the object as an extra first argument.
- On arm9, inside a class or block body, a bare
printfcalled the C library instead of theStdio.printfthatuse Stdio;brings in. - A
weakfield of protocol type was not cleared when its object was freed. - A bound-method field was aligned to 8 bytes on 32-bit targets.
- win64 code is optimised as fully as x86_64 code.
Library and runtime
Section titled “Library and runtime”- xt6502:
Math.ln,Math.expandMath.powgive correct results fordoubleandfloat.Math.rand()returns a value in [0.5, 1.0) instead of 0.Time.secondsSinceno longer always returns 0, andTime.delaySecondsreads its argument correctly. - Soft-float m68k has
Math.sqrt, the trigonometric functions,ln,expandpow. They used to fail at assembly.
Libraries that import libraries
Section titled “Libraries that import libraries”- A library can subclass a class from another library. Its new methods took slots the parent library already used.
- Calls through
Hashable,Comparable,Object*andString*across a library boundary use each class’s protocol table, which every module numbers the same way. They used each module’s own slot numbers and gave wrong answers or crashed. This makes protocol calls in a multi-module program a short table walk, and a program that uses an xc library now carries every built-in method. - arm64: a library records the libraries it imports and binds its imports to them, so a program that names only the dependent library runs.
- x86_64: programs that use xc libraries link, including a client class that subclasses a library class, and a library that takes the address of another library’s function. A library records the libraries it imports, and its load-time constructors run, in dependency order, before the program’s own.
- wasm32: a library’s vtable entries that name another library’s methods are
filled in. The loader loads every library a program needs, including those
imported only by other libraries, in dependency order, and reports a missing
library or a cycle. A library’s
.jsonsidecar lists the libraries it imports. - Library builds are byte-identical between runs and machines, with the interface written in one canonical form.
xt6502 options
Section titled “xt6502 options”-Q rts|loopchooses what happens whenmainreturns: return to DOS withmain’s value (the default) or spin. Programs used to stop at aBRK.xcc-sim-6502still exits withmain’s value.--xtc-stackmoves return addresses and saved registers onto the software stack, and the:xtcStackand:hwStackannotations choose per function. The software stack is smaller than the hardware stack, so recursion runs out sooner, and there is no overflow check.-Fmb <n>keeps functions shorter thanninstructions in main RAM.-dpprints each function’s placement and-dueach region’s and bank’s usage.
Version 0.61 — optimiser work, measured, and wrong code fixed
Section titled “Version 0.61 — optimiser work, measured, and wrong code fixed”Most of this release is optimiser and back-end work, measured against clang on a new benchmark suite. It also fixes wrong-code bugs found by building real programs, and refuses several mistakes that used to compile without a word.
Licence
Section titled “Licence”- The compiler and its tools are GPLv3. The archives carry the text as
COPYING. - The standard library and runtime (
lib/xc/in an install) are GPLv3 with the GCC Runtime Library Exception (lib/xc/COPYING.RUNTIME). They are compiled into every program xcc builds, and the exception means those programs carry no obligation. Closed and commercial programs are fine.
New errors
Section titled “New errors”- A pointer to a scalar of a different width is refused as an argument.
Passing
&narrow(ani32) wherei64*is declared used to write eight bytes into four; the reverse left the high half unset. Differences of sign only,void*, function pointers, and struct and class pointers are not affected. A cast still overrides the check. - Two file-scope declarations of one name at different types, such as
u32 gX;in one file andu32 gX[64];in another, are an error naming both types. They used to share one object. Identical redeclarations still merge. new C(args)checks its arguments against the class’sinitmethods, including inherited ones. A subclass with noinitof its own used to run no initialiser, so every field read back zero. Arguments that match noinitare an error.new C()with no arguments is still the allocate-and-zero form.xccrefuses more than one source file. It used to compile only the last one and write a binary with nomain, which failed at load time withSymbol not found: _main. Objects and archives can still be listed beside the source file.
Wrong code fixed
Section titled “Wrong code fixed”.lengthon an array of more than 65535 elements returned the count modulo 65536 (120000 read as 54464).for (v in arr)used the same value, so long arrays were iterated short. The count is 32 bits on arm64, x86-64, win64, arm9 and wasm32.- A bodyless
externglobal, such as a framework constant or a global defined in a library, read a zeroed copy of its own instead of the real object, so a framework constant came back null. It is now a reference to the definition. new i64[N]andnew u64[N]called the allocator with a missing argument and could abort with a nonsense size.- arm64: a function containing a floating-point conditional, such as
if (c) x = -x;, could corrupt adoubleits caller held in a register. - x86-64 Linux: storing a callback into a slot holding stale data, such as a reused union member, could crash. arm64 already had this fix.
Pool.forRangeWithThreadsskipped any chunk whose thread failed to start and returned as if it had run. Those chunks now run on the calling thread.printffield widths now apply to%ld,%luand%con x86-64, win64, arm9, Atari ST and wasm32, as they already did on arm64.
Checked builds
Section titled “Checked builds”-fbounds-checkchecks fixed-size arrays (locals, globals, and arrays sized by their initialiser) against their declared length. Before, only heap allocations were checked and other subscripts passed unchecked.- A failed check reports the source position. Most sites used to print
?:0:0. -fbounds-checkon a target other than arm64 is an error.xccused to accept it and then fail at link on x86-64 and win64, or build a wasm32 module that checked nothing.
Performance
Section titled “Performance”The repository has a benchmark suite in benchmark/: nineteen programs, each
written in xc and in Objective-C with ARC, both built at -O3, with a
checksum that must agree. The figures come from the compiler that ships.
Against clang the geometric mean is 0.92x on arm64 and 1.04x on x86-64, where
lower is faster. The Performance page has the
per-program table and the caveats.
Optimiser:
- More loops vectorise: counting matches over bytes, loops that carry two accumulators, division by a constant, and reductions over the loop counter with no array.
- Small structs passed or copied by value are split into fields and kept in registers.
- Functions whose locals have their address taken can be inlined. On arm64, x86-64 and win64 so can functions taking a struct by value.
- A value stored to a field and read back in the same block is reused.
- An
if/elsechoosing between two values becomes a branch-free select, and a short-circuit&&no longer builds a boolean in memory. - Full unrolling is capped by the size of the result, so large unrolled loops no longer spill.
- On arm64, loop blocks are laid out so the hot path falls through.
arm64:
- Functions with large stack arrays get full register allocation and single-instruction frame access. Past 16 KB of frame, each access cost three instructions, and past 32 KB nothing was kept in a register.
- Functions with several loops reuse registers again. A live-range error made every value overlap every other.
- Floating-point code uses d16-d31 in functions that do not vectorise, and values between calls use x0-x7.
- More constants are encoded in the instruction: shifted 12-bit immediates
such as
#4096, logical immediates, shift counts and shifted-register operands.
x86-64:
- Floats are kept in registers instead of stack slots.
- The register allocator gains six caller-saved registers on Linux and four on win64.
- A block that ends by jumping to the next block falls through.
- Loop heads are aligned to 32 bytes, and ELF
.textis aligned to 64 bytes so that alignment holds. - A conditional select reuses the flags from its compare, and vector operations no longer copy a source register that dies at the instruction.
- Division by a constant vectorises, and constant array indices fold into the address.
Runtime:
- Small objects are cheaper to allocate. The macOS and x86-64 Linux runtimes no longer round every allocation up to 256 bytes, and x86-64 Linux keeps up to 64 freed blocks per 16-byte size class, up to 1 KB, for reuse. The allocation benchmark is now faster than clang on both hosts. win64 and arm9 are unchanged.
wasm32:
- Three optimisations are on: unrolling loops with a run-time trip count, inlining functions that take a struct by value, and hoisting global addresses. Over eight benchmarks under Node, code is 31.5% faster on the geometric mean for modules 1.9% larger.
- The loader provides
clock_gettime, so programs that time themselves run.
Assemblers
Section titled “Assemblers”- x86-64: SSE shifts by an immediate (
psrlw,psrld,psrlq,psllq) were encoded as the register form and produced invalid code.psrad,pslld,psrlqandpsllqby immediate, andpcmpeqb,pcmpeqw,pcmpgtbandpcmpgtw, were missing. - arm64:
umull2andushrare supported. - arm9:
//comments are accepted as well as@, and neither is treated as a comment inside a quoted string.
Install contents
Section titled “Install contents”- The install and every archive hold
xcc,xcc-sign,xcc-asand the two simulators,xcc-sim-6502andxcc-sim-68k, plus the support tree inlib/xc.xccruns every stage itself, from parsing to linking, so there are no separate stage programs inbin/.
Options
Section titled “Options”xcc in 0.6 answered “unrecognised option” to many options this site
documented. They all work now, and xcc -h lists every option in sections.
- Output and inspection:
--output,-a/--assemble-only,-E/--preprocessed <path>,--emit-ir,--emit-ir-opt,-fdce-trace. - Paths and definitions: joined
-DNAME[=VALUE],--include,--library-path,--xcc-home, and theXCC_HOME,XTC_HOMEandXTC_LDFLAGSenvironment variables.~/xccand~/xtcare searched for the support tree. - Targets:
--arch, the spellingsx86-64,amd64,windows,wasm,armv7andcortex-a9, and-A 68000/-A 68030with-mhard-float(68881) and-fpic(the GOT/a5model) on m68k. - 6502 layouts:
-mtakes the built-in layouts the back end supports and a.lnkfile by path, and names what is missing in a layout it cannot build (xt-heapandxt-test-falloverneed split banking, which the back end does not have).-ll/--list-layoutsand-dl/--dump-layoutprint layouts. - Optimisation: bare
-O,-Fli/--fn-leaf-inline. - Code generation:
-fthread-safe-arc,-fno-thread-safe-arc, and-fmalloc=mimallocon x86_64, which the in-house link now honours. - Linking:
-Wl,/-Xlinkertake linker flags as well as files (-rpathbecomes a run-path entry; other flags are reported and skipped),--no-self-hostlinks arm64, android and arm9 executables with the platform toolchain, and--emit-libbuilds android and iOS libraries. - Android packaging:
--needed,--with-lib,--lib-name,--with-dex. - Accepted with a warning that they have no effect:
-Q,--xtc-stack,-Fmb,-dp,-du,-falloc=bump,-gand the retired-farc.
xcc-sign exports a signing identity from the macOS keychain
(--export-identity) and creates one through App Store Connect
(--fetch-identity, with --list-certs and --revoke-cert), with no other
tool. xcc-as takes the full assembler command line: banked and split-bank
output, PRG, listings, -D, -I and multiple inputs.
Tools and documentation
Section titled “Tools and documentation”- An
xccrun from outside an install looked for/opt/xcc/0.6as its fallback library root, so a newer compiler could build against 0.6’s libraries. The fallback now follows the compiler’s own version. XTIR_OPT_STOP_AFTER=<pass>works with an installedxcc. It used to be ignored. An unknown pass name is refused with the list of valid names, and an empty value counts as unset.- The language reference has a Grammar page, and
docs/xtc.bnfholds the same grammar.
Version 0.6 — the compiler is written in xc
Section titled “Version 0.6 — the compiler is written in xc”The xcc in this release is the compiler written in xc, compiled by itself. 0.5 was
the internal line that led to this release and was never published, so its changes are
all listed here.
xcc is written in xc
Section titled “xcc is written in xc”xccis the compiler written in xc, compiled by itself. The whole toolchain rebuilds itself to a fixed point.- It ships for all three hosts: macOS on Apple silicon, Linux x86-64 and Windows x64.
Each is a single self-contained binary. With no
-A, it builds for the host it runs on. xcc-signis written in xc too.xccrejects an option it does not implement with an error rather than ignoring it.- The install goes to
/opt/xcc/0.6and leaves an installed 0.4 alone.xcc -vreportsxcc 0.6 (xc, self-hosted).
callback
Section titled “callback”A bound method now has a named type, spelled like block, with the signature inline:
callback onChange void(i32 value) = &controller.valueChanged;if (onChange) { onChange(3); }onChange = (callback void(i32))0;It works for locals, fields, parameters, globals and return types, and the standard
library uses it. A stored callback always auto-zeroes when its receiver dies, so writing
weak: on one is an error.
- A callback can be called from any expression: an array element, a struct field, another object’s field, or the result of a call.
- Arrays of callbacks, and global arrays of pointers, take initialisers such as
{ &dbl, &sq }. (pointer)cbgives the function’s code address, and(callback i32(i32))pmakes a callable callback from a C function pointer.- A C function can no longer declare a
callbackparameter. C expects a one-word function pointer, and the two-word callback shifted every argument after it. Declare the parameter aspointerand pass(pointer)&fn. - Assigning
&obj.methodto ablockis refused. Before, it compiled and the call did nothing.
Checked builds and analysis
Section titled “Checked builds and analysis”-fbounds-checkchecks every subscript against the array’s declared length or the allocation’s own count. A failing check prints the index, the real bound and a symbolised stack, then aborts. It is implemented for arm64.-Wanalyzeturns on static checks for unreachable code, dead stores, conditions that are always true or false, and unused locals. It is off by default.- The “
newin a loop will leak” warning is removed. Under ARC none of the cases it reported leaked.
Language
Section titled “Language”q - pon two pointers gives the distance in elements, as in C, as a signed integer of the target’s pointer width. It used to be rejected.- Dereferencing a value that is not a pointer is an error. Before, it compiled and read
from whatever address the value held. To write to an absolute address, cast first:
*(main:u8*)addr = v. - The
: unrollloop annotation now makes the optimiser unroll that loop beyond its usual limits. It used to be accepted and ignored. - A type that is never defined but used only through a pointer (
Handle*as a field, parameter, return type or cast) is an opaque handle, as in C. struct Foo;declares a struct that is defined later.
Library and runtime
Section titled “Library and runtime”Data.withCapacity(n)reserves space, asArray,SetandMapdo. It used to returnnzero bytes already counted as content.Data.withLength(n)is the sized buffer.cString()on an emptyStringreturns an empty string instead of null.printfhonours flags and field widths (%5d,%-10s). An unrecognised specification used to shift every argument after it. The 6502 copies accept widths but do not pad.- On 64-bit hosts the reference count is 32 bits, so an object can be retained more than 65,535 times without being freed while still in use.
- File and process functions (
Files,Process) work on x86-64 Linux, Windows, wasm32 under Node, and m68k. - Load-time constructors run on Android, x86-64 Linux and Windows.
- xt6502:
Math.TWO_PI()forfloatreturns 2π instead of 2/π, andMathseeds its random generator instead of writing to address$0000.
Targets
Section titled “Targets”- arm64: variadic functions use the native AAPCS convention, so a variadic xc function can be called through a prototype from another unit or from C. Structs and callbacks passed by value that do not fit in registers go on the stack as AAPCS requires.
- Separate compilation: uninitialised file-scope globals and
externglobals are common symbols on arm64 and x86-64, so every unit shares one copy and the largest definition wins. Initialised globals are exported on x86-64. - x86-64 linking: the static link drops unreachable functions and duplicate data from separately compiled objects, and uninitialised globals take no space in the file. A program built from separate objects is now about the size of the same program built as one unit.
- iOS:
xcc --sign <identity.pem>(with--sign-entitlements <plist>) signs the output as part of the build.xcc-sign --seal-resourceswrites an app bundle’s resource seal, and--info-plist,--code-resourcesand entitlements are bound into the signature, so a bundle signed without Apple’scodesigninstalls on a device.Url.fetchandLogwork on iOS.
Bug fixes
Section titled “Bug fixes”- A chain of constants such as
192 * 128 * 16folds at full precision. It wrapped at 16 bits, silently giving 0. - Integer literal division such as
840 / 56gives 15. The dividend was truncated to 8 bits. while (n-- > 0)terminates.- A
switchcase reached by fall-through sees the previous case’s updates to locals, and a local updated inside aswitchinside a loop keeps its value across iterations. - Swapping two class-pointer locals inside a loop is no longer lost at loop exit (arm64, arm9).
- Taking the address of a parameter no longer corrupts the caller’s arguments when the function is inlined.
- Values in very long functions are no longer corrupted across calls at
-O2and above (arm64). - A sum of two
doubleproducts (a*a + b*b) is computed correctly on arm64. doubletoi64/u64conversion keeps values above 2³² on arm64.- A function-local
staticarray keeps its contents between calls. !on afloatcompiles.s.n++,p->n++,s.n += kand++on an instance variable compile and update the field.- A struct’s array field decays to a pointer;
&localinside a ternary and&*pcompile. - A name declared as an array in one block and a scalar in another no longer shares one slot and crashes.
- A function returned as a callback no longer loses half its address and crashes.
p = c ? new P() : new P()no longer leaks the object.- Returning a
weak:field retains it. The caller used to release an object it did not own. - Storing a new object into a
weak:local, field, global or array element, and aweak:return type, no longer leak. - A function with a prototype in a shared header and a definition in one unit links from every unit.
- Virtual and protocol calls work in x86-64 programs built from separately compiled objects. They could crash at startup.
- String literals in two arm64 objects compiled with
-cno longer collide at link. - A static x86-64 program that uses OpenSSL (through libpq, for example) no longer crashes at exit.
mainreceivesargcandargvon Windows. On xt6502 they are zero rather than undefined.- wasm32:
main(argc, argv), float-heavy code, locals of one name in sibling blocks, 64-bit pointer offsets andnew T[n]with a 64-bit count all produce valid modules. A comparison inside a ternary is typedbool, and inline assembly is a hard error. A library can call a virtual method on an object the application created. - The m68k assembler no longer mis-resolves labels longer than 79 characters.
- The arm64 assemblers accept
fmsubandfnmadd. - In-house links that include Objective-C objects keep the data after them aligned.
- A code signature is the last thing in the file, as device install and
codesignrequire.
Version 0.4 — blocks, UTF-8 strings, the ambient platform, iOS and Android
Section titled “Version 0.4 — blocks, UTF-8 strings, the ambient platform, iOS and Android”The first release published as xcc archives.
Host builds for Linux and Windows
Section titled “Host builds for Linux and Windows”Every host build (macOS, Linux as static musl binaries that run on any x86_64 distribution, and Windows) carries all the code generators, so the host decides only where the compiler runs, never what it can produce. On a Windows host the default target is win64.
Blocks
Section titled “Blocks”Closures as first-class values, declared like variables
(block b u32(u16 x, u16 y) = { … };), with by-value snapshot captures. A block can
be passed as a parameter, returned, stored in an ivar, written inline as a method
argument, and given a bare { … } body that takes the declared signature. A named
literal can call itself. Locals declared block: are captured copy-in/write-back;
returning such a block or storing it through a member or subscript is a compile
error. Capturing self or an ivar is an error in this release: copy it into a local
first. A block passed where a bound method is expected is refused. Blocks are
lowered onto classes at parse time, so they work on every backend including the 6502.
See Blocks.
Strings are UTF-8, end to end
Section titled “Strings are UTF-8, end to end”String is UTF-8-native. Methods that work in bytes carry Byte in their name, and
methods that work in characters carry Char. This is a breaking change: length is
now byteLength, substring is substringBytes, indexOf is byteIndexOf, and
the other byte-position methods follow the same pattern. Two names keep their
spelling with a new meaning: charAt(n) returns the n-th code point, and
appendChar appends a code point. New members include charCount,
substringChars, isValidUtf8 and sanitizedUtf8 (invalid sequences become
U+FFFD). String.withEncodedBytes and Data.withStringEncoded transcode UTF-8,
ASCII, Latin-1 and UTF-16LE/BE at the edges.
String and char literals gain \uNNNN and \UNNNNNNNN escapes with fixed digit
counts, unlike C’s greedy \x. The code point is stored as UTF-8. \xNN is limited
to ASCII (\x7F and below): use \u for a character, or appendByte for a raw
byte.
Moving code to 0.4
Section titled “Moving code to 0.4”Renamed methods fail to compile, so the compiler finds those call sites for you.
charAt and appendChar still compile with their new meaning, so library members
added in 0.4 carry since("0.4"), and xcc --migrate=0.3:0.4 compiles as if the
library were still 0.3. Newer members drop out of lookup, and every call that relied
on the old meaning fails with a position. Fix those, then build without the flag.
The ambient platform surface
Section titled “The ambient platform surface”Url (with a fetch whose completion is a block), Log with the Logger protocol
behind it, and the Platform facade with its PlatformDelegate are available with
zero imports on every target. Each target’s prelude wires its own transport and
logger (browser fetch and console on wasm32, a tty-coloured console on hosted
targets), and application source never names a platform. On wasm32 the generated
loader carries default browser implementations, and a page can replace any of them
through globalThis.xccImports.browser.
iOS, Android and code signing
Section titled “iOS, Android and code signing”-A iosand-A ios-simbuild arm64 Mach-O for iOS devices and the simulator.xcc --sign <identity.pem>(or the standalonexcc-sign) replaces the ad-hoc signature with a developer signature;--sign-entitlementsembeds an entitlements plist. The signer has no Apple dependency and runs on Linux and Windows hosts.-A androidbuilds aarch64 ELF for Android, and--emit-apkpackages an installable APK. Both link in-house, with no Android SDK, NDK or JDK.
The toolchain-free cross matrix
Section titled “The toolchain-free cross matrix”make install copies the musl and mingw link pools into the install, so a machine
with only xcc on it produces static Linux ELFs and Windows PEs. An x86-64 link that
finds no musl pool fails with an error naming where it looked, never a silent
fallback. A win64 link without the mingw pool uses the freestanding runtime, which
carries its own allocator and printf. The in-house Mach-O path is now the default
on Linux and Windows hosts too; linking -l against macOS system libraries still
needs an Apple SDK’s .tbd stubs. x86-64 shared libraries link and run in-house.
Sharper edges made safe
Section titled “Sharper edges made safe”- Raw and class pointers no longer convert silently in either direction; a sema error names the fix. An explicit cast still works.
- On wasm32, an
externdefinition exports its spelled name even when overloads mangle the symbol internally, and twoexterndefinitions of one name are an error. - A C-variadic import on wasm32 is a compile error instead of an invalid module. Use
Stdio.printfor a fixed-arity import. - A Linux binary’s
mainreturn flushes stdio throughexit(3), so piped output is no longer truncated at the buffer. - A 64-bit multiply by a constant wider than 32 bits keeps its top bits on x86_64 and win64.
- A global whose initialiser cannot be folded, such as a string literal, held zero.
It is now initialised before the first statement of
main. - A cyclic
#importcould silently corrupt field offsets at-O2and above. - A declaration that shadowed an outer name rebound it for the rest of the function.
The outer binding is now restored at scope exit, and a C-style
forvariable is scoped to the loop. - An enum constant now matches an enum-typed parameter in an overloaded call.
- Assigning a strong local to a class-pointer parameter freed the object while the parameter still used it.
Version 0.3 — wasm32, separate compilation, the in-house toolchain
Section titled “Version 0.3 — wasm32, separate compilation, the in-house toolchain”Renamed to xcc, and installable
Section titled “Renamed to xcc, and installable”The binaries are renamed: the driver is xcc, the assembler xcc-as, and the
simulators xcc-sim-6502 / xcc-sim-68k. The language is still called xc.
make install puts the toolchain in /opt/xcc/<version>, and the compiler finds its
libraries relative to its own binary. With no -A or -m, xcc builds for the host,
as cc does; the 6502 is -A 6502.
Source files use the .xc extension, and the pointer sigil is * (u8* p, *p).
The old .xt extension and @ sigil are still accepted.
wasm32
Section titled “wasm32”-A wasm32 produces a .wasm file and a loader that runs under Node or in a browser.
The WAT assembler and binary writer are in-house. Classes, ARC, protocols, weak
references, i64 and floats all work. It also has structured control flow at -O1
and above, v128 SIMD, and multi-module --emit-lib. #package and extern declare
wasm imports and exports.
The in-house toolchain
Section titled “The in-house toolchain”xcc has its own assembler, object writer and linker for every target: Mach-O (with an
ad-hoc code signer), ELF, PE/COFF and wasm. It is the default everywhere. A Mac builds
Linux and Windows executables with no other toolchain installed, and the compiler
builds and runs on Linux and Windows. The in-house linkers read static archives and
resolve -l themselves. A failed in-house link fails the build. --no-self-host
selects the external toolchain, and any build it finishes carries a warning.
Separate compilation
Section titled “Separate compilation”-cwrites a relocatable object with its interface (.xtc.iface) beside it. A client compiles against the interface, not the source, and virtual dispatch across objects uses the defining module’s slot numbering.-fltorecompiles the IR each object carries as one module, so inlining and dead-code removal work across objects again.- Both work on arm64, x86_64, win64 and arm9.
--emit-libcan build a library that wraps an external C library. Third-party libraries install under/opt/xcc/3p.
Categories and extensions
Section titled “Categories and extensions”class Shape (Drawing) { … } adds methods to any class in scope, including one inside
a prebuilt shared library. class Shape () { … } may also add fields, but only where
the class itself is compiled. Category methods that a subclass overrides dispatch
correctly across library boundaries, and several libraries may extend one class.
Threading
Section titled “Threading”Thread.spawn(&obj.method), Mutex, Cond, Sem, Atomic, ThreadLocal and
Pool.forRange are available on arm64, x86_64, win64 and arm9 (XTOS). ARC refcounts
become atomic automatically in modules that use threads; -fthread-safe-arc and
-fno-thread-safe-arc override the choice.
Language
Section titled “Language”i64/u64on every target, including xt6502, m68k and arm9.Numberholds 64-bit values andprintfprints them.defer { … }runs when the enclosing scope exits by any path, before that scope’s ARC releases.- Checked errors:
throws,throw,tryandcatch, with typedcatch (T e)arms. Calling athrowsfunction outside atryis a compile error. - Typed collections:
Array<String>*,Set<T>andMap<K, V>check what goes in and return the element type without a cast.for (i32 v in coll)unboxes. - A
staticfield has one copy per class. - Structs lay out with the target’s C alignment.
struct Name :packed { … }opts out. - A bodyless function declared with
...uses the C variadic ABI, andf(fmt, ...)forwards a variadic’s arguments. - Adjacent string literals concatenate.
return;in a function that declares a return value is an error.-farc=offis removed.
Foundation
Section titled “Foundation”A Copying protocol. String.appendFormat and String.withFormat, in-place string
editing, path helpers and CharacterSet. Array insert, remove and replace. Host file
I/O and argv. Map and Set iterate in insertion order.
Self-hosting
Section titled “Self-hosting”The front end, optimiser, every back end, the assemblers and the linkers are ported to xc. The port produces byte-identical output and builds itself to a fixed point.
Options
Section titled “Options”-fmalloc=mimalloc (x86_64), -Wunguarded-action (a callback called without being
tested), -x-<arch>,<option> for target-specific options, and xcc -v reports the
build identity.
- ARC:
breakandcontinuereleased nothing, the right arm of&&/||leaked a temporary, and a store through a pointer did not retain. - Returning a strong local through an upcast freed it.
- arm64 passes call arguments past the eighth on the stack, and its frame limit rises from 16 KB to 4 MB.
- Inline assembly was silently dropped on x86_64 and arm9.
- x86_64 returns a struct larger than 16 bytes through memory, as System V requires.
- A Mach-O dylib is no longer limited to 255 exported symbols.
- A failed checked downcast aborts on every target.
- A reduction loop that did not start at zero produced wrong results.
Version 0.2 — the IR compiler, shared libraries, bound methods
Section titled “Version 0.2 — the IR compiler, shared libraries, bound methods”A new version line. 0.12 was the last release of the AST code generator; 0.2 is the first of the IR compiler, which replaces it. The old code generator is removed.
One IR, six targets
Section titled “One IR, six targets”The compiler lowers to a single architecture-neutral IR, and each backend passes the full fixture corpus:
-A | Target | Output |
|---|---|---|
6502 (default) | banked xt6502: 4 KB hidden hardware stack, SP-relative addressing | banked 6502 executable (.xex), run under xts (now xcc-sim-6502) |
arm64 | native macOS / Linux host | Mach-O / ELF executable |
arm9 | AArch32 / XTOS | ELF executable, or a .so |
m68k | Motorola 680x0, -m atarist | GEMDOS .tos, run under xst (now xcc-sim-68k) |
x86_64 | Linux (musl) | ELF executable |
win64 | Windows x64 | PE executable or DLL |
win64 joined later in the line. It runs under Wine and has full C interop, including
struct arguments, callbacks from C, and #import <user32> for the Windows API.
Standard-library classes resolve by architecture × platform, so one source serves all of them. The compiler imports a per-platform prelude before every file, so application source does not name its platform.
The xl / xe flat and PORTB memory models, and the Commodore c64 target, are
retired.
On the m68k, floats are software by default; -mhard-float uses a 68881/68882. On
the xt6502, float and double are IEEE, computed by the MECH math coprocessor, which
also handles 32-bit multiply and divide. The 5-byte software float is removed.
Optimiser
Section titled “Optimiser”-O3 is the default. Every backend has a register allocator. The IR optimiser adds
inlining, loop-invariant code motion, loop unrolling, recursion-to-loop, if-conversion
and strength reduction. Loops auto-vectorise to NEON on arm64 and arm9 and to SSE on
x86_64.
Shared libraries: --emit-lib and #import <Lib>
Section titled “Shared libraries: --emit-lib and #import <Lib>”A program can be split into a library and its clients:
xtc -A arm9 --emit-lib -o libShapes.so shapes.xtxtc -A arm9 -L . -o app.so app.xtThis works on arm9 (.so), arm64 (.dylib), x86_64 (.so) and win64 (DLL). The
library carries its own interface inside the binary, so #import <Shapes>
type-checks the client against the real library, with no header to fall out of sync.
Classes (with inheritance, virtual dispatch back into a client subclass, and
downcasts), protocols, structs by value, enums (constants and type names), free
functions, typedefs, weak: fields, bound methods, and C types re-exported from
other libraries all cross the boundary.
#import <Foo> also reads a plain C library’s DWARF for its functions, types and
enum constants. Build the C library with -fno-eliminate-unused-debug-types, or gcc
drops the enum constants. See Modules.
Protocols across a shared library
Section titled “Protocols across a shared library”A protocol method is identified by its index within its own declaration, and the protocol by a hash of its name. Every module derives both identically with no coordination, so two independently built libraries compose, and a class conforming to a protocol from each dispatches correctly through both. An object can be downcast to a protocol at runtime.
Bound methods and optional protocol methods
Section titled “Bound methods and optional protocol methods”&obj.method yields a storable, callable {receiver, code} value. A plain function or
a static method widens into the same type, so one action field accepts any of
them. A stored callback never owns its receiver and auto-zeroes when the receiver dies.
The type was spelled with ^ in this release; it is written callback today, as below.
An optional protocol method may be left unimplemented, which leaves a null slot,
so testing a callback is equivalent to respondsTo:
callback resized void(i32 w, i32 h) = &delegate.didResize;if (resized) { resized(w, h); }Together these support the delegate and target/action patterns.
extern globals
Section titled “extern globals”Globals are scoped to the module they are compiled in. extern u16 gCounter; refers to
one defined elsewhere without reserving storage for a second copy, as an imported
library’s globals require.
weak: without a table
Section titled “weak: without a table”Weak slots are linked onto an intrusive list whose head lives in the referent’s own
heap header. There is no capacity limit (the bounded side table and its
[weak] entries setting are removed), stores are O(1), and destroying an object with
no weak references costs one null test instead of a full table scan. A stored
callback gets the same auto-zeroing with nothing to declare.
Removes a method from the vtable under --emit-lib, where whole-program
devirtualisation is unsound because the program is not whole.
Smaller changes
Section titled “Smaller changes”mainreturns the process exit code. Avoid mainreturns 0, and the 6502 simulator reports the code.- The preprocessor supports
#and##, and macro arguments substitute whole tokens only. printf%d,%uand%xformat an argument at its own width.- Foundation gains ordering and sorting, a full
String, set algebra, functionalArraymethods andData.hexString. The containers no longer leak.
Diagnostics
Section titled “Diagnostics”Cases that previously degraded silently are now errors: a store to a non-existent struct field, an unknown type name, an unresolvable imported type, and a construct the lowering cannot express. Before, these produced notes and the build succeeded with the code missing.
Version 0.12
Section titled “Version 0.12”New features
Section titled “New features”The main change is 3-byte heap pointers on banked-heap layouts. A heap pointer carries its bank byte alongside lo/hi, so a class instance, struct, or array allocated in any heap bank can be passed, returned, stored as an ivar, or kept in a collection without losing track of its bank. Every codegen path that moves a heap pointer was updated: ARC retains/releases, member access, ivar stores, multi-return tuples, downcasts, weak slots, stack-array zero-init / scope-exit walkers, subscript stores (const- and dyn-indexed), chained writes (o.mid.leaf = …), and Foundation Array / Map / Set storage. Programs on xt, rambo*, compy*, and xe-heap can spread their object graph across the full heap without trampolining through main RAM.
The bank-switch bracket optimiser covers more multi-byte field-access patterns:
- width=2 path-A bracket gate
- multi-byte heap-pointer field reads
- width=4 global-base banked field reads
- ARC field stores + struct copies
- multi-byte banked-store clusters (ExprAssign, ExprMembers)
- width=2 / width=4 dyn-banked-array reads
- xe-family bracket coverage
Each removes a save/restore around bank-select registers when the cluster shares a bank. On real programs this means fewer cycles per banked field access.
Bank-register addresses are layout-configurable. Layouts may place the bank-select hardware registers (previously hardcoded at $82/$83/$84/$85) at any address, for cartridge-mapped designs that expose the bank latches outside zero page. The compiler, the xcc-as preload-stub generator, and the xcc-sim-6502 simulator all use the layout’s addresses.
Graphics:
Gfx7: GR.7 (160×96 4-colour) with bulk-byte hline / vline fast pathsGfx15: GR.15 (160×192 4-colour) with the same bulk-byte pathgfxCreate(mode, textRows)factory inGfxFactory.xc, withGFX_<w>_<h>_<b>aliases (GFX_320_192_1, etc.). It picks the right subclass and returns aGfx@for polymorphic use. Call it asinline:gfxCreate(MODE, ROWS)when the mode is a compile-time constant: asm-level branch elimination then drops the unused subclass arms (~5 KB saved on a typical factory call)Gfx.clear()moved to the base class so it dispatches throughGfx@
Other:
inline:method()on banked-heap (xe) PORTB-brackets the inlined body- Vtable reachability uses the call-site × instantiation cross product, so dead vtable slots are zeroed instead of dangling
- Dead ARC retval stash/restore pairs are elided
xcc-aswarns on indirect-indexed addressing through a non-ZP operandxcc-asenforces split-bank size limits inwriteBankedXEX
Bug fixes
Section titled “Bug fixes”- codegen:
_virtual_dispatchtail switched fromJMP (__vt_call_vec)to self-modifyingJMP $0000(the indirect form hit the 6502JMP ($XXFF)page-crossing bug at -O3 on xl-shadow / xe-nobank) - codegen: pin vtable targets to
:main, because virtual dispatch is not bank-aware - codegen: pre-allocate ZP for inline-asm
(name),Yoperands - codegen:
_method_call_tramproutes region-C receivers via$84/$85 - codegen:
emitMethodDispatchreceiver bank source for heap-w3 - codegen:
_xcall_*_resumepreserves Y across the trampoline - codegen: bank packer estimator counts long-branch rewrites
- codegen: heap-w3 for-in stores result + bank source for spilled receiver
- codegen: heap-w3 ZP-resident struct field loads slot+2 bank
- codegen: heap-w3 pointer null-check tests lo+hi (was lo only)
- codegen: heap-w3 borrowed-init retain on 3-byte strong class pointer
- codegen: widen narrow call return when target type is wider
- codegen: gate
_cast_op_bankemit on heap-w3 cast site - foundation:
Map.containsdelegates toget;Set.containsuses if/else (avoids&&short-circuit bool-return path) - foundation:
Gfx7.vlinepen=0 erase + colour overwrite
Version 0.11
Section titled “Version 0.11”New features
Section titled “New features”The main addition is a Foundation-style class library:
- an
Objectroot class - primitive wrappers (
Number/String/Data) - heterogeneous collections:
Array, hash-basedMapandSet - the supporting
Comparable/Hashable/Enumerableprotocols
Autoboxing promotes primitives at Object@ call sites, with matching unboxing into primitive destinations. The language also gained:
- range-based
for-in(for (T i in start..end), with step and descending forms) - array slicing (
arr[m..n],arr[..n],arr[m..]) - range expressions as fixed-array initialisers
To obtain pointers to banks used as data, bank(BANK_TYPE, idx) is a builtin, and the raw:T@ pointer flavour is added.
In codegen, cloaked code regions extend across the full set of bank windows that a target’s memory-map layout defines. Calls across regions are transparent, an auto-overflow demote ladder handles full regions, and same-region bracket elision means a call from a bank to a function in the same bank pays no banked calling-convention penalty.
A new xt-shadow-heap-regC layout adds shadow main + region-C heap fallover, and the xt layouts are restructured to use banking by default.
The toolchain has a -v/--version flag, which helps diagnose why an include file is not found.
Bug fixes
Section titled “Bug fixes”- codegen: retbuf-aliasing and banked frame-save symbol leak
- codegen: per-region cloak tracker + xe-heap bank-0 cloak placement
- codegen: zero out vtable slots whose implementation was dropped by reachability
- codegen: preserve Z = retval-lo across banked-call trampolines
- codegen: float→int cast staging bugs
- codegen: drop stackRangeSet gate on auto-cloak; fix xe-heap dispatch
- driver: -H path sanitisation, search-path diagnostics, ASCII output mode
- driver: sanitise XTC_HOME env var on Windows (strip quotes, normalise backslashes)
- driver: use strtoull in parseLongLongAddr for GNUstep portability
- sema: preserve resolved return type on implicit-self bare calls
- arc: set Y to heap_bank_first before stashing _arc_retval_bank
- banked: nested method-call trampoline + Number cross-kind equals
- xl-shadow: reserve screen RAM at $8000-$9FFF; ship Array.dealloc
- xcc-sim-6502: keep SAVMSC at $8000 for explicit banked targets
- xcc-as: keep longbr trio together when previous line has its
; longbrcomment - xcc-as: bank-page overflow handling
- stdio: use BOTSCR (1-based row count), not BOTSCR-1
- stdio: port scroll() into cloaked Stdio variant
- optimiser: incorrect CMP #$00 elision in for-in range loops
- foundation: Number lazy cross-kind cache + float-cast ivar store fix