Skip to content

ChangeLog

xcc -g writes DWARF debug information into the executable, so lldb and gdb can stop on a source line, step, show a backtrace with each frame’s file and line and arguments, and print a function’s variables.

  • -g on arm64 (macOS), x86_64 (Linux, dynamic and -static) and win64: a line table for the program and the library files it imports, a function entry for every function and method (a breakpoint on a name stops at its first statement), call frames (a backtrace is right in every frame, leaf functions included), and the integer, floating-point, bool and pointer variables and parameters of each function, so frame variable, print x and bt showing add(a=2, b=3) work. -g keeps those variables in memory for the whole function, as C compilers do at -O0; without -g the output is unchanged. See Debugging.
  • Windows executables have a symbol table, so a debugger or a crash report names their functions with or without -g.

Version 0.73 — faster arm64, GPU blocks that remember, and libraries on Linux

Section titled “Version 0.73 — faster arm64, GPU blocks that remember, and libraries on Linux”

arm64 code is faster again, reversing the drift since 0.65. A par block keeps what auto learns about the CPU and the GPU, so a machine measures once. On Vulkan and WebGPU, blocks may divide 64-bit integers, reduce 8- and 16-bit values, and divide and take square roots exactly. A library built for Linux links against the system’s C library and names the libraries it needs, and on Windows #import reads a MinGW-built DLL’s own debug information.

  • par remembers: what auto measures is kept in the program’s own settings store, as a size threshold for each block: a run outside the learned bounds is decided without measuring. The keys include a hash of the GPU and CPU, so a new GPU is measured afresh, and of the block’s GPU version. The user’s own settings (par.<block> = cpu, gpu, auto or a number of items from which the GPU is used) come before them. See Parallel blocks.
  • :goal(accuracy) on Vulkan and WebGPU: a block whose goal is accuracy divides and takes square roots of floats in integer arithmetic, correctly rounded, so its results are the CPU’s; up to 0.72 it ran on the CPU.
  • More on WebGPU: 64-bit integer division and remainder, and conversions between 64-bit integers and float. On Vulkan and WebGPU, reductions of 8- and 16-bit values and bools. WebGPU keeps a block’s buffers between runs.
  • Linux libraries: --emit-lib on -A x86_64 builds a library linked against glibc, as executables are: each -l library becomes one of its dependencies, and it exports only its own API. -static keeps the musl library.
  • #import <clib> on Windows: a DLL built by a MinGW toolchain with -g is read for its functions, types and enum constants, as on Linux and macOS.
  • iOS and Android libraries: the third-party tree is searched in the target’s own directory (ios, ios-sim, android) before the architecture’s; #import <UIKit> finds the iOS SDK without SDKROOT; --with-lib may be given more than once.
  • arm64: matrix_mul takes 225 ms instead of 389 (1.73× faster; clang’s C++ takes 224), poly_dispatch is 1.20× faster, array_sum 1.13× and sort_small 1.10×. A dense integer matrix multiply now loads four elements of a row at once and multiplies by lanes, and a 32-bit copy feeding a call or a loop is a 64-bit one, which costs nothing on Apple’s cores.
  • The benchmark suite runs 1.04× faster than 0.72 on arm64 and 1.08× faster on x86-64, where arc_alloc is 2.6× faster, array_sum 1.6× and sieve 1.25× (matrix_mul is 11% slower). xc’s code is 1.45× faster than C++ on arm64 and 1.60× faster on x86-64 (geometric mean); without matrix_mul_f32, 1.09× and 1.38× faster. See Performance.
  • On x86-64 and Windows, converting a float to u64 in an unrolled loop could overwrite another value the loop was carrying, giving a wrong result at -O2.
  • A WebGPU block with more than eight arrays returned zeros; it now runs where the device allows, and on the CPU where it does not.
  • On Vulkan, an 8- or 16-bit or bool field of a block written by its body overwrote its neighbours’ bytes.

Version 0.72 — the GPU everywhere, and Foundation grows

Section titled “Version 0.72 — the GPU everywhere, and Foundation grows”

par blocks now reach the GPU on every desktop and mobile target and in the browser: Vulkan on Linux, Windows and Android, and WebGPU on the web, beside 0.7’s Metal and CUDA. A block can cover a 2-D or 3-D grid. A for loop says whether it wants speed or exactness. Linux programs link dynamically against the system’s C library by default, as C programs do, and #import of a C library finds its debug information in a separate debug file. Foundation gains twenty-one classes, many of them moved from UXKit.

  • par on Vulkan and WebGPU: on x86_64 Linux, win64 and android each block also gets a Vulkan version, and on wasm32 a WebGPU one, run in the browser while the program waits (JSPI). On Windows, CUDA is tried first on an NVIDIA GPU and Vulkan on any other; XC_PAR_GPU=vulkan|cuda chooses. 8- and 16-bit values and bools are exact on every Vulkan and WebGPU device. See Parallel blocks.
  • par :grid(w, h[, d]): a block over the points of a grid, reading its point as par.x, par.y and par.z, on the CPU and every GPU.
  • :goal on a for loop: :goal(speed) lets the compiler drop a loop’s exactness checks (a matrix multiply then skips its NaN check); plain loops keep :goal(accuracy), while a par block’s loops follow its own goal.
  • Linux links dynamically: -A x86_64 executables link against glibc by default and can load any system library; -static gives the self-contained musl executable earlier releases made.
  • #import <clib> reads a stripped library’s types and functions from its separate debug file (by build ID or .gnu_debuglink, under /usr/lib/debug or $XCC_DEBUG_DIR), compressed DWARF included, and says so when a library has no debug information at all.
  • Foundation: Bag, Range, BinaryHeap, Cache, Null, JSON, Expression, NumberFormatter, NotificationCenter, UndoManager, Progress, StateMachine, SearchIndex, IndexSet, AttributedString, Regex, Predicate, Validator, SortDescriptor, Socket and CSV; log levels, subsystems and monitors in Log; file URLs in Url; path components in String; more CharacterSets; Number.withBool. See Foundation.
  • UXKit: UXTextView, an editable rich-text view: the platform’s own editor on macOS, iOS, Android, Windows, GTK and the web, and drawn and edited by the view elsewhere. UXKit’s copies of the classes Foundation now has are gone.
  • The SME matrix kernel works on a 2×2 block of tiles, and under :goal(speed) skips its NaN check: matrix_mul_f32 takes 30 ms on an Apple M4 Max, 2.4× faster than in 0.71 and 327× faster than clang’s C++.
  • The benchmark suite runs 1.05× faster than 0.71 on arm64 and 1.01× faster on x86-64, where it is now measured as shipped, linked dynamically. xc’s code is 1.37× faster than C++ on arm64 and 1.47× faster on x86-64 (geometric mean); without matrix_mul_f32, 1.03× and 1.27× faster.
  • GPU kernels for Vulkan are structured loops and ifs, so a SIMD group’s threads reconverge, and on a discrete GPU they work in its own memory: on an RTX 3090, Vulkan comes within 1.2–1.7× of CUDA (mandelbrot 2.2 ms against 1.9). See Performance.
  • A variable assigned in a try block kept its old value after a catch that returned.
  • #import of a C library: 8-byte integers in its DWARF were read as 32-bit, and a program using one in another directory could not find it at run time.
  • A 64-bit shift by a 32-bit count printed a CUDA kernel NVIDIA’s driver refused, so the block ran on the CPU.

Version 0.71 — AVX-512, and matrix multiplies on the matrix unit

Section titled “Version 0.71 — AVX-512, and matrix multiplies on the matrix unit”

A dense matrix multiply written as three loops now runs as a matrix kernel the compiler generates itself: on Apple silicon with SME (M4 and later) on the matrix unit, and on x86-64 with SSE2, AVX2 or AVX-512 — with the results of the loops, bit for bit. x86-64 and Windows gain AVX-512, as a flag and as a third level the program picks at load.

  • AVX-512: -mavx512 (or -msimd=avx512) vectorises x86-64 and Windows code with 512-bit registers (F, DQ, BW and VL). Under -msimd=auto, the default, each vectorised function is also built for AVX-512 and the program picks it at load on a machine that has it; XC_SIMD=avx512 forces it and -mnative selects it where the build machine supports it.
  • Matrix kernels: at -O2 and above, C[i][j] = Σ A[i][k] · B[k][j] over float or double, written as three loops, runs on the SME matrix unit on Apple silicon macOS and as a vector kernel on x86-64 Linux. The results are the loops’ to the last bit, NaNs and signed zeros included; on a Mac without SME, or when the matrices overlap or hold a NaN, the loops run as written. -fno-matmul turns it off.
  • A class that lists a protocol and leaves out one of its required methods is now refused, naming the class, the protocol and the method (it built, and crashed calling the missing method).
  • matrix_mul_f32 (a 128×128 float multiply, 400 times): 3.5 ms on an Apple M4 Max against 391 ms for clang’s C++, and 12.5 ms on an AMD Zen 5 processor with AVX-512 against 252 ms.
  • The benchmark suite, now twenty programs with matrix_mul_f32 (measured back to 0.62 for the history chart), runs 1.23× faster than 0.7 on arm64 and 1.37× faster on x86-64, where AVX-512 also makes array_map, call_depth and int_muldiv twice as fast. xc’s code is 1.30× faster than C++ on arm64 and 1.47× faster on x86-64 (geometric mean); without matrix_mul_f32, 1.03× and 1.28× faster.
  • Compiling long x86-64 functions: vectorize_wide_tails builds in 20 s instead of 45 (copy propagation and loop-invariant motion no longer scale with the square of a function’s length).

Version 0.7 — parallel blocks, on every CPU thread and on the GPU

Section titled “Version 0.7 — parallel blocks, on every CPU thread and on the GPU”

A par block runs a loop’s iterations at once: over every CPU thread, or on the GPU — Metal on Apple silicon, and NVIDIA GPUs under Windows through the driver alone, with no CUDA toolkit. The compiler holds every block to what a GPU can run, on every target, so the same source runs anywhere; by default each block measures the CPU and the GPU and keeps the faster. Foundation gains Operation and OperationQueue. See Parallel blocks and the new GPU section of Performance.

  • par blocks: par name { for (T i in a..b) { … } } runs independent iterations across every CPU thread; :reduce(op var) folds results, and an integer reduction matches a serial run exactly.
  • On the GPU: Metal on Apple silicon (macOS), and NVIDIA on Windows, where the program hands the driver PTX and needs no toolkit. Global and captured arrays become GPU buffers; helper functions of plain values run there too.
  • The device is chosen per block (auto, the default): the CPU is timed first, a block under 1 ms stays there, otherwise the GPU is timed too and the faster kept, each after a warm-up run. XC_PAR=cpu|gpu|auto, Par.device(name, choice) and XC_PAR_REPORT=1 override and report it.
  • :goal(speed) (the default) lets the GPU use its fast maths; :goal(accuracy) keeps precise sin, exp and friends.
  • The par-gpu warning says why a block cannot run on the target’s GPU, and what to change; a block whose items read each other’s results is an error.
  • Operation and OperationQueue: units of work with dependencies across queues, priorities, cancellation and completion blocks, run on worker threads, serially, or on the main run loop; on targets without threads a queue runs its operations when the program waits for them.
  • --manifest-attr name=value sets an attribute on an APK’s <application>.
  • sizeof is the target’s size type: u64 on 64-bit targets, u32 on 32-bit ones, u16 on the 6502, so sizeof of a 1 MB buffer is 1048576.
  • A loop variable can be captured by a block or a par block inside the loop.
  • A par block on the GPU: mandelbrot over 2048×2048 pixels takes 2.0 ms on Apple silicon’s GPU against 99 ms on every CPU thread and 482 ms on one, and 1.9 ms on an RTX 3090. A memory-bound block such as saxpy stays on the CPU, where auto finds it is faster.
  • x86-64 and Windows: an unrolled AVX2 loop no longer chains its iterations through the integer splats, so call_depth runs 2.9× faster than in 0.66 and the x86-64 benchmark suite 1.08× faster. xc’s code is now 1.11× faster than C++ on x86-64 (geometric mean), and 1.04× faster on arm64.
  • A static method called on a class that also has instances no longer runs the class’s init on its static storage (it crashed Windows programs and broke OperationQueue on the 6502).
  • x86-64 and Windows, at -O3: a checked downcast past an object’s own class could return null after loop rotation.
  • AVX2: an unrolled loop’s integer splats no longer chain one iteration into the next.
  • wasm32: an indirect call that returns a struct declares its return slot.
  • arm9: a DWARF import takes the declaration that has the parameters.
  • A protocol cast the 6502 or m68k cannot check at run time is an error unless the class declares the protocol.
  • sizeof a value too large for the target’s size type is an error.
  • A program links with -c objects passed through -Wl, or XTC_LDFLAGS.
  • Windows: a program calling POSIX names such as close or chdir loads, through the C runtime’s own exports.

Version 0.66 — AVX2 wherever it runs, and Linux GUI programs

Section titled “Version 0.66 — AVX2 wherever it runs, and Linux GUI programs”

x86-64 and Windows programs now use AVX2 on machines that have it and SSE2 on machines that don’t, from one binary, picked when the program starts: on a Zen 5 machine an elementwise map and a sum take about half the time they did. -A x86_64 -dynamic links against glibc, so a Linux program can load GTK 4, libGL and other shared system libraries, still with xcc’s own linker.

  • Runtime SIMD dispatch on -A x86_64 and -A win64, the default (-msimd=auto): each function the vectoriser widens is built for SSE2 and for AVX2, and the program picks one per function at startup from what the CPU and the operating system support. XC_SIMD=base or XC_SIMD=avx2 forces a level for one run; forcing one the machine lacks falls back, with a note on stderr. -c objects and libraries keep one version.
  • -mavx2 (or -msimd=avx2) builds AVX2 code only, -msimd=base SSE2 only, and -mnative the level of the machine running xcc.
  • -dynamic on -A x86_64: a dynamically linked glibc executable that can load shared system libraries. -l<name> takes lib<name>.so from -L and the system library directories; a Mac can link for Linux with -L to copies of the libraries. A symbol that neither glibc nor an -l library defines is a link error.
  • libUXGtk.so, UXKit’s GTK 4 back end as a shared library, is a separate download for Linux GUI programs linked with -dynamic.
  • A library’s structs, enums and typedefs import into its clients.
  • defer takes a block.
  • Bundle.main() finds the program’s executable through PATH when it was started by name.
  • AVX2 code (dispatched, or with -mavx2): 256-bit vector loops, with vzeroupper around calls. On a Zen 5 machine array_map runs 2.0× and array_sum 1.9× faster than with SSE2.
  • The compiler holds less memory while it builds a large program.
  • arm64, in a program that starts a thread: a dealloc that called a method on self could run again from inside itself until the stack overflowed.
  • An array ivar of class pointers, and an object local whose address is taken, now own what they hold.
  • x86-64 and Windows: exp and pow are accurate to the last bit.
  • A single-precision denormal constant keeps its value.
  • printf of a typed collection’s element prints its value.
  • Windows: a frame larger than a page is probed one page at a time; a float passed through ... reaches the callee; a weak reference is unlinked only through a valid back-pointer.
  • wasm32: objects use the 40-byte header and 32-bit reference count the other targets use.
  • Files reads /proc and /sys files, which report no size.
  • macOS: calling a C function that no linked library exports is a link error.
  • A static -A x86_64 link refuses an -l it cannot use (only a shared library, or nothing) instead of dropping it.
  • A typed collection of primitives without Number is an error, and so is a checked cast from a raw pointer.
  • A construct the lowering cannot handle is named, not numbered.
  • make install PREFIX=<dir> keeps the third-party tree beside the prefix instead of writing /opt/xcc/3p.

Version 0.65 — faster loops, native settings and system frameworks

Section titled “Version 0.65 — faster loops, native settings and system frameworks”

The optimiser and the arm64 back end close the distance to clang: across the benchmark suite arm64 code is now 1.03× faster than clang’s C++ (geometric mean), from 1.24× slower, and 1.38× faster than 0.64; matrix multiply is 4.7× faster. x86-64 code is 1.08× slower than C++, from 1.13× slower. See Performance. Settings.standard() keeps its values in the platform’s own settings store, and #import <Framework> links a macOS or iOS system framework.

  • Settings.standard(name) uses the platform’s own store: the user’s preferences on macOS and iOS, the registry on Windows, localStorage in a browser, and a text file under $XDG_CONFIG_HOME (or ~/.config) elsewhere. The API is unchanged. A 0.64 settings file is read the first time and the next save writes the new location.
  • #import <Framework> links a macOS or iOS system framework, as -framework does.
  • The wasm32 loader passes the mouse wheel and the right button to the program.
  • A loop around a reduction loop is vectorised across the outer loop: four neighbouring outer iterations are four lanes, which is the shape of a matrix multiply.
  • Pointers step through strided and offset array walks (a[i*n + k]) instead of recomputing the index.
  • An integer reduction in an unrolled loop keeps one operation per iteration on the carried value; the rest run in parallel.
  • More loops are rotated so the test sits at the bottom, and a value computed in a loop’s condition is not computed again in its body.
  • arm64: a leaf function that never touches its frame has none; parameters and call results stay in the registers they arrive in; madd, msub and mla replace multiply-then-add; local arrays and functions start on 16-byte boundaries.
  • x86-64: local arrays start on 16-byte boundaries.
  • arm64: a value live across a call could share an argument register with a loop-carried value at -O2 and above.

Version 0.64 — C’s printf, an HTTP client and a run loop

Section titled “Version 0.64 — C’s printf, an HTTP client and a run loop”

printf and the other format functions now follow C, keeping %@ for objects. The library gains settings files, bundles, an HTTP client, file operations on a background thread and a run loop to deliver their results. The x86-64 back end is faster, and a number of wrong-code bugs are fixed.

  • Stdio.printf, String.withFormat, String.appendFormat and Log.error, Log.warning and Log.info take C’s conversions, flags, field widths, precisions and *, with C’s output: %x has no leading zeros, %f and %lf both print six places, %e and %g are C’s, and the floating conversions round exactly as C does. %@ still prints an object’s description(), and now an enum’s name too.
  • The arguments follow C: an integer narrower than int is passed as an int and a float as a double. int is 32 bits, 16 on xt6502; long is 64 bits on the 64-bit targets.
  • With a literal format the compiler sizes each integer conversion to the argument actually passed, so %d prints an i64 whole; the check now reports only an argument of the wrong kind or a wrong count. A variadic function that passes its format and ... on to one of these is treated the same way at its own call sites.
  • Settings (a key/value file) and Bundle (a program’s resources).
  • Http: an HTTP/1.1 client, with TLS where the platform provides it, and the native transport for url.fetch.
  • AsyncFiles: the Files operations on a background thread, and RunLoop, which the completions can be delivered to.
  • Third-party libraries resolve and load with the compiler.
  • A float literal may have an exponent without a point (1e30), and the f/F suffix marks a single-precision literal.
  • PLATFORM_android is defined for -A android.
  • New warnings: a non-void function that can reach its closing brace, and a local that hides a field.
  • x86-64: values are loaded straight into their home registers, registers are ranked by loop-weighted use and shared along live segments, a float phi copy is one movaps, and loops rotate so the test sits at the bottom. Across the benchmark suite x86-64 code is now 1.14× faster than clang’s (geometric mean), from 1.03× slower; see Performance.
  • A %@ argument to String.withFormat or appendFormat was released once too often, so a local object could be freed while still in use.
  • A method called through the vtable with variadic arguments got its tail on arm64.
  • A scalar argument converts to a float parameter; a local that shadows a field writes the local; sizeof(x) on a variable measures the variable; a string literal in a global aggregate initialiser is emitted; an implicit-self call boxes its arguments; an array bound folds in a struct field and a prefix type.
  • A double literal below the smallest normal value is kept rather than flushed to zero.

Version 0.63 — archiving, class names and Windows DLLs

Section titled “Version 0.63 — archiving, class names and Windows DLLs”

This release adds keyed archiving to the language, gives every object a runtime class name, and lets win64 build a DLL in-house. It also fixes wrong code across the back ends, and arm9 gains load-time constructors and can build a program from objects.

  • Codable and Coder: keyed archiving. A class that adopts Codable gains encodeWith(coder) and initWith(coder); a Coder writes a keyed archive as JSON, optionally gzipped. Object adopts Codable.
  • Object.className() returns an object’s runtime class name, and Object.newInstanceOfClass() makes another instance of it. Class and block names are carried into the image.
  • Windows: xcc -A win64 --emit-lib builds a DLL in-house, with an entry point, imports and an export table. A DLL’s interface is read back, and a module dispatches protocols through the itable under the Microsoft ABI.
  • arm9: a load-time constructor in an object runs, and both compilers’ -c objects and IR are the same, so a program can be built from objects and linked in-house.
  • xcc-sim-6502 and xcc-sim-68k report their own name and the version they were built from.
  • Floating point: a comparison follows IEEE where either operand is a NaN; a negated float flips its sign bit instead of subtracting from zero; a negated integer literal wider than its type is built at the wider type; and an integer initialiser for a float global is folded into the image rather than built at run time.
  • Conversions between floating point and unsigned or 64-bit integers are complete on x86-64, win64, m68k and arm9. On xt6502 a 64-bit integer converts straight to float and rounds once.
  • x86-64 and win64 no longer rename a function that is named after a register, which broke a call to one.
  • A 64-bit select copies both words on m68k and arm9, and a 64-bit phi stays out of a register home on those targets.
  • wasm32: a local the optimiser pinned and then removed is given no slot.
  • x86-64: a weak undefined symbol in a dynamic link keeps its value 0.
  • arm32: both assemblers write one literal-pool word per distinct expression, as as does, so the linked images agree.
  • xt6502: :main and :banked place free functions and methods; every new block is sized to its contents with its header in a trailer; and a variadic reentrance error no longer names the 6502 packing-buffer address.
  • arm9 and android place the extra arguments of a variadic call where the callee reads them. On android a library can also be #imported.
  • -flto keeps an imported vtable and each object’s string literals.
  • The shipped compiler refuses a class whose parent is not a class, an unreachable function end is Unreachable, an open slice start is a zero of the count’s width, and preprocessor diagnostics carry their location.
  • The reference compiler catches up on several points of its own: a float result returned through an indirect call comes from s0/d0; an autoboxed argument is not a raw pointer for a class parameter; an array ivar of an inline class-array element is at its address; an enum ivar counts as an integer ivar; a library interface does not re-export a protocol it imported; an x86-64 dynamic link exports what the shipped linker exports; and -S on arm9 writes the assembly a link would assemble.

Version 0.62 — wrong code fixed, shared libraries that work together

Section titled “Version 0.62 — wrong code fixed, shared libraries that work together”

Most of this release fixes wrong code and makes xc libraries work with each other: a library can now import another library and be used from a program on arm64, x86_64 and wasm32. Several mistakes that used to compile are now errors. Libraries built by an earlier release must be rebuilt, and on x86_64 programs that use them must be rebuilt too.

  • Every call checks its argument count: static, instance, inherited, category, protocol and free-function calls, library imports, and calls through function pointers, callbacks and blocks. A variadic call must pass at least the fixed arguments.
  • A call argument whose class is unrelated to the parameter’s (neither a subclass nor an ancestor) is refused, for methods and functions alike. Passing an ancestor, such as the Object* a collection returns, where a subclass is declared is still accepted. A class reference still does not convert to a raw pointer implicitly; cast it.
  • delete on a class instance is refused. ARC owns the object; freeing it by hand freed it twice. delete on a struct or primitive array is unchanged.
  • On xt6502, a call to a function that is declared but never defined is an error naming the function and the call’s position. It used to become a jump to address 0.
  • m68k refuses inline assembly it cannot compile. It used to drop the block.
  • Errors in a call now point at the name being called, not at the closing ).
  • A virtual call to an overloaded method could run another overload’s body. Each overload now dispatches through its own slot.
  • An overload that differs only by return type, passed directly as an argument (Math.ln(Math.E())), takes the parameter’s type.
  • A call through a callback or function pointer now converts each argument to the parameter’s type, as a direct call does. On xt6502 the callee read garbage; on arm64 a negative narrow integer passed to a 64-bit parameter arrived as a large positive number.
  • A condition is tested on its whole value on every target. arm64, x86_64 and win64 tested only the low 32 bits of a 64-bit value or pointer, arm9 and m68k one half of a 64-bit value, and xt6502 the low byte of a pointer. The right side of && and || was reduced to its low byte. A floating-point condition was tested by its bits, so -0.0 counted as true.
  • !p on a pointer tests the whole pointer. On arm64 a pointer on a 64 KB boundary read as null.
  • An integer converted to a pointer keeps the full pointer width. arm64 kept only 16 bits, so an address round-tripped through an integer crashed.
  • On arm64 and x86_64 at -O2 and above, the first call to a static method of a class from a library could crash.
  • xt6502 at -O3 could lose a struct parameter across a call when only the addresses of its fields were still in use.
  • On arm9, m68k, wasm32 and xt6502 the 16-bit retain count now stops at $FFFF in both directions. It used to wrap, freeing an object that was still in use. An object that reaches the limit is never freed.
  • Freeing new C[0] ran one element’s dealloc, and an empty array created at run time reported a .length of 1.
  • xt6502 converts between floating point and 64-bit integers correctly.
  • m68k converts a floating-point value to a narrow integer correctly: out of range gives 0, and values from 2^31 to 2^32 reach u32. A pointer converted to a 64-bit integer fills both halves.
  • On arm64, a call to a cloaked or banked variadic function put its extra arguments in the wrong place.
  • A struct passed to a variadic function had all its fields written to its first byte.
  • A static method could fill a protocol’s method slot, so a call through the protocol passed the object as an extra first argument.
  • On arm9, inside a class or block body, a bare printf called the C library instead of the Stdio.printf that use Stdio; brings in.
  • A weak field of protocol type was not cleared when its object was freed.
  • A bound-method field was aligned to 8 bytes on 32-bit targets.
  • win64 code is optimised as fully as x86_64 code.
  • xt6502: Math.ln, Math.exp and Math.pow give correct results for double and float. Math.rand() returns a value in [0.5, 1.0) instead of 0. Time.secondsSince no longer always returns 0, and Time.delaySeconds reads its argument correctly.
  • Soft-float m68k has Math.sqrt, the trigonometric functions, ln, exp and pow. They used to fail at assembly.
  • A library can subclass a class from another library. Its new methods took slots the parent library already used.
  • Calls through Hashable, Comparable, Object* and String* across a library boundary use each class’s protocol table, which every module numbers the same way. They used each module’s own slot numbers and gave wrong answers or crashed. This makes protocol calls in a multi-module program a short table walk, and a program that uses an xc library now carries every built-in method.
  • arm64: a library records the libraries it imports and binds its imports to them, so a program that names only the dependent library runs.
  • x86_64: programs that use xc libraries link, including a client class that subclasses a library class, and a library that takes the address of another library’s function. A library records the libraries it imports, and its load-time constructors run, in dependency order, before the program’s own.
  • wasm32: a library’s vtable entries that name another library’s methods are filled in. The loader loads every library a program needs, including those imported only by other libraries, in dependency order, and reports a missing library or a cycle. A library’s .json sidecar lists the libraries it imports.
  • Library builds are byte-identical between runs and machines, with the interface written in one canonical form.
  • -Q rts|loop chooses what happens when main returns: return to DOS with main’s value (the default) or spin. Programs used to stop at a BRK. xcc-sim-6502 still exits with main’s value.
  • --xtc-stack moves return addresses and saved registers onto the software stack, and the :xtcStack and :hwStack annotations choose per function. The software stack is smaller than the hardware stack, so recursion runs out sooner, and there is no overflow check.
  • -Fmb <n> keeps functions shorter than n instructions in main RAM.
  • -dp prints each function’s placement and -du each region’s and bank’s usage.

Version 0.61 — optimiser work, measured, and wrong code fixed

Section titled “Version 0.61 — optimiser work, measured, and wrong code fixed”

Most of this release is optimiser and back-end work, measured against clang on a new benchmark suite. It also fixes wrong-code bugs found by building real programs, and refuses several mistakes that used to compile without a word.

  • The compiler and its tools are GPLv3. The archives carry the text as COPYING.
  • The standard library and runtime (lib/xc/ in an install) are GPLv3 with the GCC Runtime Library Exception (lib/xc/COPYING.RUNTIME). They are compiled into every program xcc builds, and the exception means those programs carry no obligation. Closed and commercial programs are fine.
  • A pointer to a scalar of a different width is refused as an argument. Passing &narrow (an i32) where i64* is declared used to write eight bytes into four; the reverse left the high half unset. Differences of sign only, void*, function pointers, and struct and class pointers are not affected. A cast still overrides the check.
  • Two file-scope declarations of one name at different types, such as u32 gX; in one file and u32 gX[64]; in another, are an error naming both types. They used to share one object. Identical redeclarations still merge.
  • new C(args) checks its arguments against the class’s init methods, including inherited ones. A subclass with no init of its own used to run no initialiser, so every field read back zero. Arguments that match no init are an error. new C() with no arguments is still the allocate-and-zero form.
  • xcc refuses more than one source file. It used to compile only the last one and write a binary with no main, which failed at load time with Symbol not found: _main. Objects and archives can still be listed beside the source file.
  • .length on an array of more than 65535 elements returned the count modulo 65536 (120000 read as 54464). for (v in arr) used the same value, so long arrays were iterated short. The count is 32 bits on arm64, x86-64, win64, arm9 and wasm32.
  • A bodyless extern global, such as a framework constant or a global defined in a library, read a zeroed copy of its own instead of the real object, so a framework constant came back null. It is now a reference to the definition.
  • new i64[N] and new u64[N] called the allocator with a missing argument and could abort with a nonsense size.
  • arm64: a function containing a floating-point conditional, such as if (c) x = -x;, could corrupt a double its caller held in a register.
  • x86-64 Linux: storing a callback into a slot holding stale data, such as a reused union member, could crash. arm64 already had this fix.
  • Pool.forRangeWithThreads skipped any chunk whose thread failed to start and returned as if it had run. Those chunks now run on the calling thread.
  • printf field widths now apply to %ld, %lu and %c on x86-64, win64, arm9, Atari ST and wasm32, as they already did on arm64.
  • -fbounds-check checks fixed-size arrays (locals, globals, and arrays sized by their initialiser) against their declared length. Before, only heap allocations were checked and other subscripts passed unchecked.
  • A failed check reports the source position. Most sites used to print ?:0:0.
  • -fbounds-check on a target other than arm64 is an error. xcc used to accept it and then fail at link on x86-64 and win64, or build a wasm32 module that checked nothing.

The repository has a benchmark suite in benchmark/: nineteen programs, each written in xc and in Objective-C with ARC, both built at -O3, with a checksum that must agree. The figures come from the compiler that ships. Against clang the geometric mean is 0.92x on arm64 and 1.04x on x86-64, where lower is faster. The Performance page has the per-program table and the caveats.

Optimiser:

  • More loops vectorise: counting matches over bytes, loops that carry two accumulators, division by a constant, and reductions over the loop counter with no array.
  • Small structs passed or copied by value are split into fields and kept in registers.
  • Functions whose locals have their address taken can be inlined. On arm64, x86-64 and win64 so can functions taking a struct by value.
  • A value stored to a field and read back in the same block is reused.
  • An if/else choosing between two values becomes a branch-free select, and a short-circuit && no longer builds a boolean in memory.
  • Full unrolling is capped by the size of the result, so large unrolled loops no longer spill.
  • On arm64, loop blocks are laid out so the hot path falls through.

arm64:

  • Functions with large stack arrays get full register allocation and single-instruction frame access. Past 16 KB of frame, each access cost three instructions, and past 32 KB nothing was kept in a register.
  • Functions with several loops reuse registers again. A live-range error made every value overlap every other.
  • Floating-point code uses d16-d31 in functions that do not vectorise, and values between calls use x0-x7.
  • More constants are encoded in the instruction: shifted 12-bit immediates such as #4096, logical immediates, shift counts and shifted-register operands.

x86-64:

  • Floats are kept in registers instead of stack slots.
  • The register allocator gains six caller-saved registers on Linux and four on win64.
  • A block that ends by jumping to the next block falls through.
  • Loop heads are aligned to 32 bytes, and ELF .text is aligned to 64 bytes so that alignment holds.
  • A conditional select reuses the flags from its compare, and vector operations no longer copy a source register that dies at the instruction.
  • Division by a constant vectorises, and constant array indices fold into the address.

Runtime:

  • Small objects are cheaper to allocate. The macOS and x86-64 Linux runtimes no longer round every allocation up to 256 bytes, and x86-64 Linux keeps up to 64 freed blocks per 16-byte size class, up to 1 KB, for reuse. The allocation benchmark is now faster than clang on both hosts. win64 and arm9 are unchanged.

wasm32:

  • Three optimisations are on: unrolling loops with a run-time trip count, inlining functions that take a struct by value, and hoisting global addresses. Over eight benchmarks under Node, code is 31.5% faster on the geometric mean for modules 1.9% larger.
  • The loader provides clock_gettime, so programs that time themselves run.
  • x86-64: SSE shifts by an immediate (psrlw, psrld, psrlq, psllq) were encoded as the register form and produced invalid code. psrad, pslld, psrlq and psllq by immediate, and pcmpeqb, pcmpeqw, pcmpgtb and pcmpgtw, were missing.
  • arm64: umull2 and ushr are supported.
  • arm9: // comments are accepted as well as @, and neither is treated as a comment inside a quoted string.
  • The install and every archive hold xcc, xcc-sign, xcc-as and the two simulators, xcc-sim-6502 and xcc-sim-68k, plus the support tree in lib/xc. xcc runs every stage itself, from parsing to linking, so there are no separate stage programs in bin/.

xcc in 0.6 answered “unrecognised option” to many options this site documented. They all work now, and xcc -h lists every option in sections.

  • Output and inspection: --output, -a/--assemble-only, -E/--preprocessed <path>, --emit-ir, --emit-ir-opt, -fdce-trace.
  • Paths and definitions: joined -DNAME[=VALUE], --include, --library-path, --xcc-home, and the XCC_HOME, XTC_HOME and XTC_LDFLAGS environment variables. ~/xcc and ~/xtc are searched for the support tree.
  • Targets: --arch, the spellings x86-64, amd64, windows, wasm, armv7 and cortex-a9, and -A 68000 / -A 68030 with -mhard-float (68881) and -fpic (the GOT/a5 model) on m68k.
  • 6502 layouts: -m takes the built-in layouts the back end supports and a .lnk file by path, and names what is missing in a layout it cannot build (xt-heap and xt-test-fallover need split banking, which the back end does not have). -ll/--list-layouts and -dl/--dump-layout print layouts.
  • Optimisation: bare -O, -Fli/--fn-leaf-inline.
  • Code generation: -fthread-safe-arc, -fno-thread-safe-arc, and -fmalloc=mimalloc on x86_64, which the in-house link now honours.
  • Linking: -Wl,/-Xlinker take linker flags as well as files (-rpath becomes a run-path entry; other flags are reported and skipped), --no-self-host links arm64, android and arm9 executables with the platform toolchain, and --emit-lib builds android and iOS libraries.
  • Android packaging: --needed, --with-lib, --lib-name, --with-dex.
  • Accepted with a warning that they have no effect: -Q, --xtc-stack, -Fmb, -dp, -du, -falloc=bump, -g and the retired -farc.

xcc-sign exports a signing identity from the macOS keychain (--export-identity) and creates one through App Store Connect (--fetch-identity, with --list-certs and --revoke-cert), with no other tool. xcc-as takes the full assembler command line: banked and split-bank output, PRG, listings, -D, -I and multiple inputs.

  • An xcc run from outside an install looked for /opt/xcc/0.6 as its fallback library root, so a newer compiler could build against 0.6’s libraries. The fallback now follows the compiler’s own version.
  • XTIR_OPT_STOP_AFTER=<pass> works with an installed xcc. It used to be ignored. An unknown pass name is refused with the list of valid names, and an empty value counts as unset.
  • The language reference has a Grammar page, and docs/xtc.bnf holds the same grammar.

Version 0.6 — the compiler is written in xc

Section titled “Version 0.6 — the compiler is written in xc”

The xcc in this release is the compiler written in xc, compiled by itself. 0.5 was the internal line that led to this release and was never published, so its changes are all listed here.

  • xcc is the compiler written in xc, compiled by itself. The whole toolchain rebuilds itself to a fixed point.
  • It ships for all three hosts: macOS on Apple silicon, Linux x86-64 and Windows x64. Each is a single self-contained binary. With no -A, it builds for the host it runs on.
  • xcc-sign is written in xc too.
  • xcc rejects an option it does not implement with an error rather than ignoring it.
  • The install goes to /opt/xcc/0.6 and leaves an installed 0.4 alone. xcc -v reports xcc 0.6 (xc, self-hosted).

A bound method now has a named type, spelled like block, with the signature inline:

callback onChange void(i32 value) = &controller.valueChanged;
if (onChange) { onChange(3); }
onChange = (callback void(i32))0;

It works for locals, fields, parameters, globals and return types, and the standard library uses it. A stored callback always auto-zeroes when its receiver dies, so writing weak: on one is an error.

  • A callback can be called from any expression: an array element, a struct field, another object’s field, or the result of a call.
  • Arrays of callbacks, and global arrays of pointers, take initialisers such as { &dbl, &sq }.
  • (pointer)cb gives the function’s code address, and (callback i32(i32))p makes a callable callback from a C function pointer.
  • A C function can no longer declare a callback parameter. C expects a one-word function pointer, and the two-word callback shifted every argument after it. Declare the parameter as pointer and pass (pointer)&fn.
  • Assigning &obj.method to a block is refused. Before, it compiled and the call did nothing.
  • -fbounds-check checks every subscript against the array’s declared length or the allocation’s own count. A failing check prints the index, the real bound and a symbolised stack, then aborts. It is implemented for arm64.
  • -Wanalyze turns on static checks for unreachable code, dead stores, conditions that are always true or false, and unused locals. It is off by default.
  • The “new in a loop will leak” warning is removed. Under ARC none of the cases it reported leaked.
  • q - p on two pointers gives the distance in elements, as in C, as a signed integer of the target’s pointer width. It used to be rejected.
  • Dereferencing a value that is not a pointer is an error. Before, it compiled and read from whatever address the value held. To write to an absolute address, cast first: *(main:u8*)addr = v.
  • The : unroll loop annotation now makes the optimiser unroll that loop beyond its usual limits. It used to be accepted and ignored.
  • A type that is never defined but used only through a pointer (Handle* as a field, parameter, return type or cast) is an opaque handle, as in C.
  • struct Foo; declares a struct that is defined later.
  • Data.withCapacity(n) reserves space, as Array, Set and Map do. It used to return n zero bytes already counted as content. Data.withLength(n) is the sized buffer.
  • cString() on an empty String returns an empty string instead of null.
  • printf honours flags and field widths (%5d, %-10s). An unrecognised specification used to shift every argument after it. The 6502 copies accept widths but do not pad.
  • On 64-bit hosts the reference count is 32 bits, so an object can be retained more than 65,535 times without being freed while still in use.
  • File and process functions (Files, Process) work on x86-64 Linux, Windows, wasm32 under Node, and m68k.
  • Load-time constructors run on Android, x86-64 Linux and Windows.
  • xt6502: Math.TWO_PI() for float returns 2π instead of 2/π, and Math seeds its random generator instead of writing to address $0000.
  • arm64: variadic functions use the native AAPCS convention, so a variadic xc function can be called through a prototype from another unit or from C. Structs and callbacks passed by value that do not fit in registers go on the stack as AAPCS requires.
  • Separate compilation: uninitialised file-scope globals and extern globals are common symbols on arm64 and x86-64, so every unit shares one copy and the largest definition wins. Initialised globals are exported on x86-64.
  • x86-64 linking: the static link drops unreachable functions and duplicate data from separately compiled objects, and uninitialised globals take no space in the file. A program built from separate objects is now about the size of the same program built as one unit.
  • iOS: xcc --sign <identity.pem> (with --sign-entitlements <plist>) signs the output as part of the build. xcc-sign --seal-resources writes an app bundle’s resource seal, and --info-plist, --code-resources and entitlements are bound into the signature, so a bundle signed without Apple’s codesign installs on a device. Url.fetch and Log work on iOS.
  • A chain of constants such as 192 * 128 * 16 folds at full precision. It wrapped at 16 bits, silently giving 0.
  • Integer literal division such as 840 / 56 gives 15. The dividend was truncated to 8 bits.
  • while (n-- > 0) terminates.
  • A switch case reached by fall-through sees the previous case’s updates to locals, and a local updated inside a switch inside a loop keeps its value across iterations.
  • Swapping two class-pointer locals inside a loop is no longer lost at loop exit (arm64, arm9).
  • Taking the address of a parameter no longer corrupts the caller’s arguments when the function is inlined.
  • Values in very long functions are no longer corrupted across calls at -O2 and above (arm64).
  • A sum of two double products (a*a + b*b) is computed correctly on arm64.
  • double to i64/u64 conversion keeps values above 2³² on arm64.
  • A function-local static array keeps its contents between calls.
  • ! on a float compiles.
  • s.n++, p->n++, s.n += k and ++ on an instance variable compile and update the field.
  • A struct’s array field decays to a pointer; &local inside a ternary and &*p compile.
  • A name declared as an array in one block and a scalar in another no longer shares one slot and crashes.
  • A function returned as a callback no longer loses half its address and crashes.
  • p = c ? new P() : new P() no longer leaks the object.
  • Returning a weak: field retains it. The caller used to release an object it did not own.
  • Storing a new object into a weak: local, field, global or array element, and a weak: return type, no longer leak.
  • A function with a prototype in a shared header and a definition in one unit links from every unit.
  • Virtual and protocol calls work in x86-64 programs built from separately compiled objects. They could crash at startup.
  • String literals in two arm64 objects compiled with -c no longer collide at link.
  • A static x86-64 program that uses OpenSSL (through libpq, for example) no longer crashes at exit.
  • main receives argc and argv on Windows. On xt6502 they are zero rather than undefined.
  • wasm32: main(argc, argv), float-heavy code, locals of one name in sibling blocks, 64-bit pointer offsets and new T[n] with a 64-bit count all produce valid modules. A comparison inside a ternary is typed bool, and inline assembly is a hard error. A library can call a virtual method on an object the application created.
  • The m68k assembler no longer mis-resolves labels longer than 79 characters.
  • The arm64 assemblers accept fmsub and fnmadd.
  • In-house links that include Objective-C objects keep the data after them aligned.
  • A code signature is the last thing in the file, as device install and codesign require.

Version 0.4 — blocks, UTF-8 strings, the ambient platform, iOS and Android

Section titled “Version 0.4 — blocks, UTF-8 strings, the ambient platform, iOS and Android”

The first release published as xcc archives.

Every host build (macOS, Linux as static musl binaries that run on any x86_64 distribution, and Windows) carries all the code generators, so the host decides only where the compiler runs, never what it can produce. On a Windows host the default target is win64.

Closures as first-class values, declared like variables (block b u32(u16 x, u16 y) = { … };), with by-value snapshot captures. A block can be passed as a parameter, returned, stored in an ivar, written inline as a method argument, and given a bare { … } body that takes the declared signature. A named literal can call itself. Locals declared block: are captured copy-in/write-back; returning such a block or storing it through a member or subscript is a compile error. Capturing self or an ivar is an error in this release: copy it into a local first. A block passed where a bound method is expected is refused. Blocks are lowered onto classes at parse time, so they work on every backend including the 6502. See Blocks.

String is UTF-8-native. Methods that work in bytes carry Byte in their name, and methods that work in characters carry Char. This is a breaking change: length is now byteLength, substring is substringBytes, indexOf is byteIndexOf, and the other byte-position methods follow the same pattern. Two names keep their spelling with a new meaning: charAt(n) returns the n-th code point, and appendChar appends a code point. New members include charCount, substringChars, isValidUtf8 and sanitizedUtf8 (invalid sequences become U+FFFD). String.withEncodedBytes and Data.withStringEncoded transcode UTF-8, ASCII, Latin-1 and UTF-16LE/BE at the edges.

String and char literals gain \uNNNN and \UNNNNNNNN escapes with fixed digit counts, unlike C’s greedy \x. The code point is stored as UTF-8. \xNN is limited to ASCII (\x7F and below): use \u for a character, or appendByte for a raw byte.

Renamed methods fail to compile, so the compiler finds those call sites for you. charAt and appendChar still compile with their new meaning, so library members added in 0.4 carry since("0.4"), and xcc --migrate=0.3:0.4 compiles as if the library were still 0.3. Newer members drop out of lookup, and every call that relied on the old meaning fails with a position. Fix those, then build without the flag.

Url (with a fetch whose completion is a block), Log with the Logger protocol behind it, and the Platform facade with its PlatformDelegate are available with zero imports on every target. Each target’s prelude wires its own transport and logger (browser fetch and console on wasm32, a tty-coloured console on hosted targets), and application source never names a platform. On wasm32 the generated loader carries default browser implementations, and a page can replace any of them through globalThis.xccImports.browser.

  • -A ios and -A ios-sim build arm64 Mach-O for iOS devices and the simulator.
  • xcc --sign <identity.pem> (or the standalone xcc-sign) replaces the ad-hoc signature with a developer signature; --sign-entitlements embeds an entitlements plist. The signer has no Apple dependency and runs on Linux and Windows hosts.
  • -A android builds aarch64 ELF for Android, and --emit-apk packages an installable APK. Both link in-house, with no Android SDK, NDK or JDK.

make install copies the musl and mingw link pools into the install, so a machine with only xcc on it produces static Linux ELFs and Windows PEs. An x86-64 link that finds no musl pool fails with an error naming where it looked, never a silent fallback. A win64 link without the mingw pool uses the freestanding runtime, which carries its own allocator and printf. The in-house Mach-O path is now the default on Linux and Windows hosts too; linking -l against macOS system libraries still needs an Apple SDK’s .tbd stubs. x86-64 shared libraries link and run in-house.

  • Raw and class pointers no longer convert silently in either direction; a sema error names the fix. An explicit cast still works.
  • On wasm32, an extern definition exports its spelled name even when overloads mangle the symbol internally, and two extern definitions of one name are an error.
  • A C-variadic import on wasm32 is a compile error instead of an invalid module. Use Stdio.printf or a fixed-arity import.
  • A Linux binary’s main return flushes stdio through exit(3), so piped output is no longer truncated at the buffer.
  • A 64-bit multiply by a constant wider than 32 bits keeps its top bits on x86_64 and win64.
  • A global whose initialiser cannot be folded, such as a string literal, held zero. It is now initialised before the first statement of main.
  • A cyclic #import could silently corrupt field offsets at -O2 and above.
  • A declaration that shadowed an outer name rebound it for the rest of the function. The outer binding is now restored at scope exit, and a C-style for variable is scoped to the loop.
  • An enum constant now matches an enum-typed parameter in an overloaded call.
  • Assigning a strong local to a class-pointer parameter freed the object while the parameter still used it.

Version 0.3 — wasm32, separate compilation, the in-house toolchain

Section titled “Version 0.3 — wasm32, separate compilation, the in-house toolchain”

The binaries are renamed: the driver is xcc, the assembler xcc-as, and the simulators xcc-sim-6502 / xcc-sim-68k. The language is still called xc. make install puts the toolchain in /opt/xcc/<version>, and the compiler finds its libraries relative to its own binary. With no -A or -m, xcc builds for the host, as cc does; the 6502 is -A 6502.

Source files use the .xc extension, and the pointer sigil is * (u8* p, *p). The old .xt extension and @ sigil are still accepted.

-A wasm32 produces a .wasm file and a loader that runs under Node or in a browser. The WAT assembler and binary writer are in-house. Classes, ARC, protocols, weak references, i64 and floats all work. It also has structured control flow at -O1 and above, v128 SIMD, and multi-module --emit-lib. #package and extern declare wasm imports and exports.

xcc has its own assembler, object writer and linker for every target: Mach-O (with an ad-hoc code signer), ELF, PE/COFF and wasm. It is the default everywhere. A Mac builds Linux and Windows executables with no other toolchain installed, and the compiler builds and runs on Linux and Windows. The in-house linkers read static archives and resolve -l themselves. A failed in-house link fails the build. --no-self-host selects the external toolchain, and any build it finishes carries a warning.

  • -c writes a relocatable object with its interface (.xtc.iface) beside it. A client compiles against the interface, not the source, and virtual dispatch across objects uses the defining module’s slot numbering.
  • -flto recompiles the IR each object carries as one module, so inlining and dead-code removal work across objects again.
  • Both work on arm64, x86_64, win64 and arm9.
  • --emit-lib can build a library that wraps an external C library. Third-party libraries install under /opt/xcc/3p.

class Shape (Drawing) { … } adds methods to any class in scope, including one inside a prebuilt shared library. class Shape () { … } may also add fields, but only where the class itself is compiled. Category methods that a subclass overrides dispatch correctly across library boundaries, and several libraries may extend one class.

Thread.spawn(&obj.method), Mutex, Cond, Sem, Atomic, ThreadLocal and Pool.forRange are available on arm64, x86_64, win64 and arm9 (XTOS). ARC refcounts become atomic automatically in modules that use threads; -fthread-safe-arc and -fno-thread-safe-arc override the choice.

  • i64 / u64 on every target, including xt6502, m68k and arm9. Number holds 64-bit values and printf prints them.
  • defer { … } runs when the enclosing scope exits by any path, before that scope’s ARC releases.
  • Checked errors: throws, throw, try and catch, with typed catch (T e) arms. Calling a throws function outside a try is a compile error.
  • Typed collections: Array<String>*, Set<T> and Map<K, V> check what goes in and return the element type without a cast. for (i32 v in coll) unboxes.
  • A static field has one copy per class.
  • Structs lay out with the target’s C alignment. struct Name :packed { … } opts out.
  • A bodyless function declared with ... uses the C variadic ABI, and f(fmt, ...) forwards a variadic’s arguments.
  • Adjacent string literals concatenate.
  • return; in a function that declares a return value is an error. -farc=off is removed.

A Copying protocol. String.appendFormat and String.withFormat, in-place string editing, path helpers and CharacterSet. Array insert, remove and replace. Host file I/O and argv. Map and Set iterate in insertion order.

The front end, optimiser, every back end, the assemblers and the linkers are ported to xc. The port produces byte-identical output and builds itself to a fixed point.

-fmalloc=mimalloc (x86_64), -Wunguarded-action (a callback called without being tested), -x-<arch>,<option> for target-specific options, and xcc -v reports the build identity.

  • ARC: break and continue released nothing, the right arm of &&/|| leaked a temporary, and a store through a pointer did not retain.
  • Returning a strong local through an upcast freed it.
  • arm64 passes call arguments past the eighth on the stack, and its frame limit rises from 16 KB to 4 MB.
  • Inline assembly was silently dropped on x86_64 and arm9.
  • x86_64 returns a struct larger than 16 bytes through memory, as System V requires.
  • A Mach-O dylib is no longer limited to 255 exported symbols.
  • A failed checked downcast aborts on every target.
  • A reduction loop that did not start at zero produced wrong results.

Version 0.2 — the IR compiler, shared libraries, bound methods

Section titled “Version 0.2 — the IR compiler, shared libraries, bound methods”

A new version line. 0.12 was the last release of the AST code generator; 0.2 is the first of the IR compiler, which replaces it. The old code generator is removed.

The compiler lowers to a single architecture-neutral IR, and each backend passes the full fixture corpus:

-ATargetOutput
6502 (default)banked xt6502: 4 KB hidden hardware stack, SP-relative addressingbanked 6502 executable (.xex), run under xts (now xcc-sim-6502)
arm64native macOS / Linux hostMach-O / ELF executable
arm9AArch32 / XTOSELF executable, or a .so
m68kMotorola 680x0, -m ataristGEMDOS .tos, run under xst (now xcc-sim-68k)
x86_64Linux (musl)ELF executable
win64Windows x64PE executable or DLL

win64 joined later in the line. It runs under Wine and has full C interop, including struct arguments, callbacks from C, and #import <user32> for the Windows API.

Standard-library classes resolve by architecture × platform, so one source serves all of them. The compiler imports a per-platform prelude before every file, so application source does not name its platform.

The xl / xe flat and PORTB memory models, and the Commodore c64 target, are retired.

On the m68k, floats are software by default; -mhard-float uses a 68881/68882. On the xt6502, float and double are IEEE, computed by the MECH math coprocessor, which also handles 32-bit multiply and divide. The 5-byte software float is removed.

-O3 is the default. Every backend has a register allocator. The IR optimiser adds inlining, loop-invariant code motion, loop unrolling, recursion-to-loop, if-conversion and strength reduction. Loops auto-vectorise to NEON on arm64 and arm9 and to SSE on x86_64.

Shared libraries: --emit-lib and #import <Lib>

Section titled “Shared libraries: --emit-lib and #import <Lib>”

A program can be split into a library and its clients:

Terminal window
xtc -A arm9 --emit-lib -o libShapes.so shapes.xt
xtc -A arm9 -L . -o app.so app.xt

This works on arm9 (.so), arm64 (.dylib), x86_64 (.so) and win64 (DLL). The library carries its own interface inside the binary, so #import <Shapes> type-checks the client against the real library, with no header to fall out of sync. Classes (with inheritance, virtual dispatch back into a client subclass, and downcasts), protocols, structs by value, enums (constants and type names), free functions, typedefs, weak: fields, bound methods, and C types re-exported from other libraries all cross the boundary.

#import <Foo> also reads a plain C library’s DWARF for its functions, types and enum constants. Build the C library with -fno-eliminate-unused-debug-types, or gcc drops the enum constants. See Modules.

A protocol method is identified by its index within its own declaration, and the protocol by a hash of its name. Every module derives both identically with no coordination, so two independently built libraries compose, and a class conforming to a protocol from each dispatches correctly through both. An object can be downcast to a protocol at runtime.

Bound methods and optional protocol methods

Section titled “Bound methods and optional protocol methods”

&obj.method yields a storable, callable {receiver, code} value. A plain function or a static method widens into the same type, so one action field accepts any of them. A stored callback never owns its receiver and auto-zeroes when the receiver dies. The type was spelled with ^ in this release; it is written callback today, as below.

An optional protocol method may be left unimplemented, which leaves a null slot, so testing a callback is equivalent to respondsTo:

callback resized void(i32 w, i32 h) = &delegate.didResize;
if (resized) { resized(w, h); }

Together these support the delegate and target/action patterns.

Globals are scoped to the module they are compiled in. extern u16 gCounter; refers to one defined elsewhere without reserving storage for a second copy, as an imported library’s globals require.

Weak slots are linked onto an intrusive list whose head lives in the referent’s own heap header. There is no capacity limit (the bounded side table and its [weak] entries setting are removed), stores are O(1), and destroying an object with no weak references costs one null test instead of a full table scan. A stored callback gets the same auto-zeroing with nothing to declare.

Removes a method from the vtable under --emit-lib, where whole-program devirtualisation is unsound because the program is not whole.

  • main returns the process exit code. A void main returns 0, and the 6502 simulator reports the code.
  • The preprocessor supports # and ##, and macro arguments substitute whole tokens only.
  • printf %d, %u and %x format an argument at its own width.
  • Foundation gains ordering and sorting, a full String, set algebra, functional Array methods and Data.hexString. The containers no longer leak.

Cases that previously degraded silently are now errors: a store to a non-existent struct field, an unknown type name, an unresolvable imported type, and a construct the lowering cannot express. Before, these produced notes and the build succeeded with the code missing.

The main change is 3-byte heap pointers on banked-heap layouts. A heap pointer carries its bank byte alongside lo/hi, so a class instance, struct, or array allocated in any heap bank can be passed, returned, stored as an ivar, or kept in a collection without losing track of its bank. Every codegen path that moves a heap pointer was updated: ARC retains/releases, member access, ivar stores, multi-return tuples, downcasts, weak slots, stack-array zero-init / scope-exit walkers, subscript stores (const- and dyn-indexed), chained writes (o.mid.leaf = …), and Foundation Array / Map / Set storage. Programs on xt, rambo*, compy*, and xe-heap can spread their object graph across the full heap without trampolining through main RAM.

The bank-switch bracket optimiser covers more multi-byte field-access patterns:

  • width=2 path-A bracket gate
  • multi-byte heap-pointer field reads
  • width=4 global-base banked field reads
  • ARC field stores + struct copies
  • multi-byte banked-store clusters (ExprAssign, ExprMembers)
  • width=2 / width=4 dyn-banked-array reads
  • xe-family bracket coverage

Each removes a save/restore around bank-select registers when the cluster shares a bank. On real programs this means fewer cycles per banked field access.

Bank-register addresses are layout-configurable. Layouts may place the bank-select hardware registers (previously hardcoded at $82/$83/$84/$85) at any address, for cartridge-mapped designs that expose the bank latches outside zero page. The compiler, the xcc-as preload-stub generator, and the xcc-sim-6502 simulator all use the layout’s addresses.

Graphics:

  • Gfx7: GR.7 (160×96 4-colour) with bulk-byte hline / vline fast paths
  • Gfx15: GR.15 (160×192 4-colour) with the same bulk-byte path
  • gfxCreate(mode, textRows) factory in GfxFactory.xc, with GFX_<w>_<h>_<b> aliases (GFX_320_192_1, etc.). It picks the right subclass and returns a Gfx@ for polymorphic use. Call it as inline:gfxCreate(MODE, ROWS) when the mode is a compile-time constant: asm-level branch elimination then drops the unused subclass arms (~5 KB saved on a typical factory call)
  • Gfx.clear() moved to the base class so it dispatches through Gfx@

Other:

  • inline:method() on banked-heap (xe) PORTB-brackets the inlined body
  • Vtable reachability uses the call-site × instantiation cross product, so dead vtable slots are zeroed instead of dangling
  • Dead ARC retval stash/restore pairs are elided
  • xcc-as warns on indirect-indexed addressing through a non-ZP operand
  • xcc-as enforces split-bank size limits in writeBankedXEX
  • codegen: _virtual_dispatch tail switched from JMP (__vt_call_vec) to self-modifying JMP $0000 (the indirect form hit the 6502 JMP ($XXFF) page-crossing bug at -O3 on xl-shadow / xe-nobank)
  • codegen: pin vtable targets to :main, because virtual dispatch is not bank-aware
  • codegen: pre-allocate ZP for inline-asm (name),Y operands
  • codegen: _method_call_tramp routes region-C receivers via $84/$85
  • codegen: emitMethodDispatch receiver bank source for heap-w3
  • codegen: _xcall_*_resume preserves Y across the trampoline
  • codegen: bank packer estimator counts long-branch rewrites
  • codegen: heap-w3 for-in stores result + bank source for spilled receiver
  • codegen: heap-w3 ZP-resident struct field loads slot+2 bank
  • codegen: heap-w3 pointer null-check tests lo+hi (was lo only)
  • codegen: heap-w3 borrowed-init retain on 3-byte strong class pointer
  • codegen: widen narrow call return when target type is wider
  • codegen: gate _cast_op_bank emit on heap-w3 cast site
  • foundation: Map.contains delegates to get; Set.contains uses if/else (avoids && short-circuit bool-return path)
  • foundation: Gfx7.vline pen=0 erase + colour overwrite

The main addition is a Foundation-style class library:

  • an Object root class
  • primitive wrappers (Number / String / Data)
  • heterogeneous collections: Array, hash-based Map and Set
  • the supporting Comparable / Hashable / Enumerable protocols

Autoboxing promotes primitives at Object@ call sites, with matching unboxing into primitive destinations. The language also gained:

  • range-based for-in (for (T i in start..end), with step and descending forms)
  • array slicing (arr[m..n], arr[..n], arr[m..])
  • range expressions as fixed-array initialisers

To obtain pointers to banks used as data, bank(BANK_TYPE, idx) is a builtin, and the raw:T@ pointer flavour is added.

In codegen, cloaked code regions extend across the full set of bank windows that a target’s memory-map layout defines. Calls across regions are transparent, an auto-overflow demote ladder handles full regions, and same-region bracket elision means a call from a bank to a function in the same bank pays no banked calling-convention penalty.

A new xt-shadow-heap-regC layout adds shadow main + region-C heap fallover, and the xt layouts are restructured to use banking by default.

The toolchain has a -v/--version flag, which helps diagnose why an include file is not found.

  • codegen: retbuf-aliasing and banked frame-save symbol leak
  • codegen: per-region cloak tracker + xe-heap bank-0 cloak placement
  • codegen: zero out vtable slots whose implementation was dropped by reachability
  • codegen: preserve Z = retval-lo across banked-call trampolines
  • codegen: float→int cast staging bugs
  • codegen: drop stackRangeSet gate on auto-cloak; fix xe-heap dispatch
  • driver: -H path sanitisation, search-path diagnostics, ASCII output mode
  • driver: sanitise XTC_HOME env var on Windows (strip quotes, normalise backslashes)
  • driver: use strtoull in parseLongLongAddr for GNUstep portability
  • sema: preserve resolved return type on implicit-self bare calls
  • arc: set Y to heap_bank_first before stashing _arc_retval_bank
  • banked: nested method-call trampoline + Number cross-kind equals
  • xl-shadow: reserve screen RAM at $8000-$9FFF; ship Array.dealloc
  • xcc-sim-6502: keep SAVMSC at $8000 for explicit banked targets
  • xcc-as: keep longbr trio together when previous line has its ; longbr comment
  • xcc-as: bank-page overflow handling
  • stdio: use BOTSCR (1-based row count), not BOTSCR-1
  • stdio: port scroll() into cloaked Stdio variant
  • optimiser: incorrect CMP #$00 elision in for-in range loops
  • foundation: Number lazy cross-kind cache + float-cast ivar store fix